We call this Machine Native Intelligence: AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost.
Building prod, not God
TypeSafe is not trying to build a model that does everything. It is designed for production systems where code needs a narrow decision it can inspect and act on. Our expectation is that large-scale AI automation will be closer to 99% machine-to-machine interactions and 1% human interaction. That shifts the design target from responses that feel good to read toward outputs that behave predictably inside software. Read the TypeSafe manifesto.Three post-training approaches
Pretrained language models have been adapted in two major ways. TypeSafe adds a third. RLHF and RLVR are shown here for context; TypeSafe’s training path is RLCD.
RLHF was used to train InstructGPT and ChatGPT and was co-invented by Diogo Almeida, cofounder of TypeSafe.


RLCD and calibrated decisions
RLCD optimizes for a different output contract:- The model does not generate text.
- It returns decisions and probabilities.
- Higher probability should correspond to a greater chance that the answer is correct.
- Outcomes assigned a probability of
0.2should occur about 20% of the time. - Outcomes assigned a probability of
0.8should occur about 80% of the time. - Outcomes assigned a probability of
1.0should occur 100% of the time.
The problems with RLHF
RLHF teaches a model to say things that people prefer. That objective works well for chatbots, but it can also reward sycophancy and confident-sounding hallucinations. Preference optimization also causes mode dropping: the model learns to favor a particular style, such as instruction following, while reducing the probability of other possible outputs.

Mode collapse analogy
Mode collapse analogy



