[Source: TypeSafe AI launch article by Diogo Almeida, 2026-09-14; documentation and evaluation site reviewed 2026-09-15.]
Jev is TypeSafe AI's first System One model, launched in early access on September 14, 2026. It accepts text or structured application state and answers predefined questions with typed values and probability distributions. TypeSafe describes a new architecture, parallel sampling and Reinforcement Learning for Calibrated Decisions (RLCD). The reviewed sources describe the approach but do not provide enough architecture or training detail to reproduce it. Announcement
The interface has three primitives: Choice selects from supplied options, Score returns a probability-weighted rubric score, and Noul returns a yes/no probability. Independent questions can share one request to POST /v1/systemone, using jev-latest. Python and JavaScript SDKs are documented. Jev does not generate prose, code or reasoning explanations. Extraction therefore requires a constrained representation or candidate selection rather than arbitrary text generation. The launch gives a maximum choice cardinality of 255, with larger sets handled in stages. API System One Announcement
TypeSafe quotes $0.042 per million input tokens, free output tokens and 70-500 ms end-to-end latency. Its headline 193.6x speed and 444.6x cost improvements come from its own workflow evaluation, and the company explicitly calls these the higher end of expected real-world gains. Measurements generally originate on the US West Coast, near the service. These are vendor measurements, not independently reproduced results or latency guarantees. Announcement
The evaluation covers four equally weighted workflows: security incidents, agent trace observability, invoice processing and customer service. It assumes the workflow code is correct and uses averaged GPT-6 Astra and Claude Fable 5.1 answers at high reasoning as reference labels. Other models run at provider defaults. This measures agreement with model consensus, not independently established business correctness. Workflows were authored by TypeSafe's capabilities team; the LLM adapter requests probabilities as well as decisions, increasing comparison cost and latency. A fair reproduction should also compare small models with constrained outputs, matching decision quality and probability requirements. Evaluation methodology Adapter repository
The claim that Jev cannot hallucinate needs a narrow reading. TypeSafe guarantees schema matching, not semantic correctness: an allowed option can still be the wrong one. Its documentation explicitly says calibration across groups of predictions does not guarantee an individual answer. Choice and Score confidence is a statistic derived from the output distribution, not an independent verification signal; Noul does not carry a separate confidence field. Calibration on a new domain, especially under distribution shift, remains something to measure. Confidence System One
A concrete agent example is skill suggestion using the Hermes catalog. Two calls first rank 182 skills and then inspect the top three, with the option to recommend none. TypeSafe reports 488 requests with Claude Haiku 4.5: wrong skill loads fall from 16.8% to 7.3% on 315 covered requests; unnecessary loads fall from 9.8% to 4.0% on 173 uncovered requests. This is a vendor experiment on one harness and model, not evidence that the gains generalize to every agent. Other documented patterns include RAG passage classification, citation checking and entity alignment. Skill suggestion cookbook Documentation index
The Doom demo runs about 10 queries per second for approximately $7 per hour, according to Almeida. It consumes structured game state represented as data and text, not screenshots or video. TypeSafe says a conventional Doom bot could play better; the demonstration illustrates reactive decisions and instruction following rather than superior game performance or visual control. The hourly figure is specific to this workload, not a fixed API rate. Demo post Demo details