SnackOnAI Engineering | Senior AI Systems Researcher | Technical Deep Dive | September 30, 2026
The Promise
Most production LLM calls ask for a paragraph when the code only needs a decision. You pay generation prices, wait on sequential decoding, then parse, validate, and retry your way back to a value you could have typed yourself.
Jev deletes that round trip. State in, typed decisions out, every question answered in the same forward pass, output tokens free because there is no output to generate.
The architecture is sound and the price is real. The marketing around hallucination and confidence is where this issue gets interesting, because both claims are narrower than they read, and both are checkable from TypeSafe's own published numbers.
What this covers: the System One request and response contract, the three primitives, how confidence is actually computed, the token economics of batched questions, and the documented failure modes. What this excludes: the RLCD training method beyond what the lab has published, and the Doom and Wikiracing demos except where they reveal cost structure.
This follows our run on structure in AI systems. MetaGPT made agents pass typed documents. Spec Kit made humans write typed specs. Jev pushes the same idea into the model itself: constrain the output space at the weights, not with a parser.
What It Actually Does
You send a state (text or JSON) and a map of questions. Each question is one of three primitives, and all of them are evaluated in parallel, in isolation, against the same state.
Primitive | Asks | Returns |
|---|---|---|
Choice | Which of these options |
|
Score | Where on this rubric |
|
Noul | Is this statement true |
|
"Noul" is short for Bernoulli. That is the whole API surface.
The company came out of stealth on September 15, 2026 with $40M led by DCVC, founded by Diogo Almeida, a co-author of the InstructGPT paper behind ChatGPT. The published numbers for jev-1.13:
Property | Value |
|---|---|
Price | $0.042 per million input tokens, output free |
Latency | 70ms to 500ms end to end |
Context | 64k tokens per request, 32k for state plus the longest question |
Rate limits | 100k tokens/sec, 40 requests/sec, adjusting dynamically |
Cardinality | 255 options per Choice |
Input | Text or JSON only, no image, audio, or video |
Training | Reinforcement Learning for Calibrated Decisions, no per-customer fine tuning |
Caption: Note what is missing. No temperature, no max tokens, no system prompt, no streaming. The request has three fields because a decision has no style parameters.
Jev reached Vercel AI Gateway within 36 hours of launch as typesafe-ai/jev, callable through AI SDK 7's evaluate, and OpenRouter lists it too. Vercel reported it hit more than twice as many paid teams in 24 hours as any previous model launch on the gateway.
The Architecture, Unpacked

Caption: Focus on the single downward arrow in the middle. The state crosses the wire once and the question count rides nearly free, which inverts the LLM habit of one call per judgment.
Three decisions carry this design, ranked:
One, the output space is the schema, enforced at sampling. A frontier model with JSON mode generates tokens that happen to parse. Jev has no token stream to constrain. The set of representable answers is the set of options you supplied, which is why schema violations are zero by construction rather than by validation.
Two, the state is encoded once and every question reads it in parallel. This is the economic core. Cost scales with the state, not the question count, and the docs state plainly that adding questions barely changes response time. It also kills a familiar failure: each question is evaluated in isolation, so question twelve cannot be contaminated by question three.
Three, the policy stays in your code. Jev returns a distribution, never an action. The thresholds, the weights across sub-judgments, and the fallback path are ordinary if statements you own, version, and test. That is a real architectural difference from an agent that decides and acts in one opaque step.
The Code, Annotated
One call, five questions, two of them speculative
# From TypeSafe's Choice docs. A support ticket touching three teams.
from typesafe_sdk import Choice, TypeSafeClient
TRIAGE_QUESTIONS = {
"department": Choice(
instructions="Which team should handle this?",
criteria={
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems",
},
),
# ← THIS is the trick: return_reason only matters if department == returns,
# and shipping_issue only if department == shipping. Both are asked anyway.
# A speculative question costs tens of tokens; a second round trip costs
# the whole state again plus another network hop.
"return_reason": Choice(
instructions="If the customer wants to return something, why?",
criteria={"wrong_size": "The item doesn't fit", "wrong_item": "A different product was delivered",
"damaged": "The item arrived broken or faulty",
"changed_mind": "The item is fine, the customer no longer wants it",
"other": "A return reason that fits none of the above"},
),
"requested_resolution": Choice(
instructions="What does the customer want to happen?",
criteria={"exchange": "Swap the item for a different one", "refund": "Money back",
"replacement": "The same item sent again", "information": "Just an answer, no action needed"},
),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None}, # names are self-evident
),
}
response = client.system_one(state=ticket, questions=TRIAGE_QUESTIONS)
answers = response.answers
if answers["department"].confidence < 0.3:
send_to_manual_triage(ticket) # the model said "I'm not sure", so don't act
elif answers["department"].choice == "returns":
assign(ticket, team="returns", issue=answers["return_reason"].choice)
# A second team with a real share of the probability gets a copy.
# ← THIS is the part JSON mode cannot give you: the runner-up is actionable.
for team, probability in answers["department"].probabilities.items():
if team != answers["department"].choice and probability > 0.25:
notify(ticket, team=team)
Caption: The loop at the bottom is the payoff for typed probabilities. With a text model you get the winning label and nothing about the 0.35 that went to billing.
The confidence score, reproduced in one line
TypeSafe's confidence docs show a demo widget that approximates confidence for three options as (3 × largest probability − 1) / 2. Generalize it and check it against every published response in the docs.
def confidence(probabilities):
"""TypeSafe's confidence, generalized from the formula in their own docs."""
n = len(probabilities)
return (n * max(probabilities) - 1) / (n - 1) # ← THIS is the trick: only the peak matters
CASES = [ # (probabilities, confidence as published by TypeSafe)
([0.85, 0.15, 0.00], 0.78), # quickstart, department
([0.61, 0.35, 0.04], 0.42), # choice docs, department
([0.74, 0.26, 0.0, 0.0, 0.0], 0.67), # choice docs, shipping_issue (5 options)
([0.40, 0.34, 0.24, 0.02], 0.20), # choice docs, requested_resolution (4 options)
([0.84, 0.16, 0.00], 0.76), # choice docs, tone
([1.00, 0.00], 1.00), # structured criteria example (2 options)
([0.00, 1.00, 0.00], 1.00), # quickstart, frustration (a Score, same rule)
]
for probabilities, published in CASES:
assert abs(confidence(probabilities) - published) < 0.005
# all seven reproduce: 0.775, 0.415, 0.675, 0.200, 0.760, 1.000, 1.000
# And the consequence, which the docs do not spell out:
confidence([0.4, 0.4, 0.2]) # 0.1 two live candidates
confidence([0.4, 0.3, 0.3]) # 0.1 one leader, two stragglers, same number
Caption: Seven for seven against TypeSafe's published responses, Scores included. Confidence is a rescaling of the top probability, so two distributions with different shapes and different entropies collapse to the same number.
Before and after, the integration you delete
# BEFORE: frontier LLM with JSON mode, the pattern this replaces
for attempt in range(3): # retry loop exists because parsing can fail
raw = llm.chat(model="...", response_format={"type": "json_object"},
messages=[{"role": "system", "content": SCHEMA_PROMPT},
{"role": "user", "content": ticket}])
try:
data = json.loads(raw.choices[0].message.content)
department = Department(data["department"]) # may raise: valid JSON, invalid enum
break
except (json.JSONDecodeError, ValueError, KeyError):
continue # pay again, wait again
else:
department = Department.UNKNOWN # and you still need a fallback
confidence = None # you can prompt for one, it won't be calibrated
# AFTER: the same decision, with the failure modes above made unrepresentable
answer = client.system_one(state=ticket, questions={"department": DEPARTMENT}).answers["department"]
department, confidence = answer.choice, answer.confidence
Caption: The deleted code is the point. No parse step, no enum validation, no retry budget, no fallback label. What survives is the question of whether the chosen option is right, which no schema can answer.
It In Action
Input: a deliberately messy ticket from TypeSafe's docs. Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this? Five Choice questions in one request.
Step one, the call. One POST. State plus all five questions comes to 589 input tokens.
Step two, the answers.
department returns 0.61 (billing 0.35, shipping 0.04) confidence 0.42
return_reason wrong_size 1.00 confidence 1.00
shipping_issue delayed 0.74 (other 0.26) confidence 0.67
requested_resolution refund 0.40 (replacement 0.34, exchange 0.24) confidence 0.20
tone frustrated 0.84 (angry 0.16) confidence 0.76
usage: input_tokens 589, output_tokens 212 (billed at zero)
Caption: The 0.20 on requested_resolution is the useful answer here. The ticket genuinely does not say whether the customer wants money or a swap, and the distribution says so instead of guessing.
Step three, what the code does. Routes to returns, copies billing because 0.35 clears a 0.25 threshold, ignores shipping_issue entirely, and asks the customer what they want because resolution confidence is under 0.5.
Step four, the bill. 589 input tokens at $0.042 per million is $0.0000247. The 212 output tokens are free.
Jev $0.0000247 (589 in, 212 out free)
same call at $0.20 in / $1.00 out $0.000330 13x
same call at $1.25 in / $10.00 out $0.002856 115x
same call at $3.00 in / $15.00 out $0.004947 200x
Caption: The multiplier is a function of which model you are replacing. Against a cheap small model the gap is 13x, not the 444x on TypeSafe's homepage. Both numbers are honest about different comparisons.
Step five, at volume. TypeSafe's parallel questions cookbook runs 13 questions over the 53,777-character GDPR Wikipedia article. Batched: one call, $0.000497, 0.27 seconds. Unbatched: 13 calls, $0.006090, 2.71 seconds. Same answers, with run to run standard deviation of exactly 0.0 on eleven of thirteen questions under both strategies.
Back out the tokens and the structure shows: $0.000497 at list price is about 11,800 input tokens, of which the article is roughly 11,150. The thirteen questions cost about 680 tokens combined. The document cost sixteen times more than every question asked of it.
Why This Design Works, And What It Trades Away
It works because it removes the two expensive parts of structured extraction at once: sequential decoding and output validation. Output tokens are free because there is no decode loop to pay for, and the retry loop disappears because the invalid states are unrepresentable rather than caught.
It works a second time because probabilities are a better interface than labels. The runner up at 0.35, the 0.20 confidence that means "the ticket is ambiguous", the Noul you can threshold at 0.7 instead of 0.5 for a destructive action: all of that is ordinary code acting on numbers.
What it trades away, from the lab's own jaggedness page:
No generation, and no arithmetic. It does not count reliably, does not compare dates as ordered values, and underperforms on numeric representations such as hex colors versus English color names. The fix is always the same: extract with Jev, compute in code.
Literal reading. It answers the question you wrote, not the one you meant. The docs offer the best heuristic I have read on this: when you look at a wrong answer and find yourself explaining what you really meant, that explanation is the missing half of the instruction.
Context rot is still real. Accuracy falls as the state grows with irrelevant detail, which means retrieval and filtering remain your problem. The 32k state ceiling enforces it anyway.
State is not treated as hostile. The docs say directly that injected instructions in the state can move the answer. A model that cannot emit a tool call is a smaller blast radius than one that can, but screening pages with Jev is a filter, not a security boundary.
English first, text only. CJK and other languages are handled but not equally well. Images and audio must be transcribed before they become state.
Technical Moats
The interface is not a moat. Within a day of launch, openjev reproduced the shape by reading option logits straight from Qwen3.5-4B. On one RTX 3090, 21 questions took 1.02 seconds as direct logits against 5.33 seconds as a generated JSON array. The 5x is the cost of generation, available to anyone with logprob access today.
The calibration is. The same project scored 0.845 modal agreement against Jev's published 0.883 on 102 aligned cases. Close, not equal, and the gap is exactly the part that took two years of training. Almeida's own position is that the bottleneck is calibration training data, not architecture.
Distribution is already compounding. Vercel AI Gateway, OpenRouter, Cloudflare Workers AI, a Claude Code plugin, an agent skill, and community wrappers for DSPy, Ruby, Elixir, and MCP inside the first week. The cheapest way to make a new primitive stick is to put it behind the SDKs engineers already call.
Insights
Insight One: "cannot hallucinate" is a claim about your type system, not about the world.
TypeSafe's own chart footnote is admirably direct: their 0% is not empirical, it is a mathematical property of schema matching. Jev cannot return an option you did not define. It can absolutely return the wrong one you did define, and that failure is harder to detect than a malformed JSON blob, because invalid output throws and a wrong enum just routes a ticket to the wrong team forever.
Two things in the docs sharpen this further. First, the state is not treated as adversarial, so a passage that argues for its own classification can move the answer. Second, and more interesting for anyone building thresholds, the model does not guarantee structural invariants. The docs publish this example: "Is the customer asking for a refund?" asked as a Noul returns 0.22 on a ticket, while the same question as a two-option Choice puts 0.01 on yes with confidence 0.97. And the same question with its negation, as two Nouls, sums to 1.19.
That is not a bug, because each question is evaluated in isolation and a Choice is relative while a Noul is absolute. But it means calibration and coherence are different properties, and only one of them is being trained. A threshold tuned on one phrasing does not transfer to another phrasing of the same judgment. Treat every question as its own classifier with its own operating point, because that is what it is.
Insight Two: confidence is not a second signal, it is a rescaling of the top probability.
The code section above reproduces TypeSafe's published confidence on all seven examples in the docs with (n · max(p) − 1) / (n − 1). Choices and Scores both. That tells you three things.
It is free to compute, so nothing is lost by the API returning it. It is not an independent measure of certainty, so "high confidence means higher accuracy" is a claim about the probabilities, and the confidence field adds no information to the distribution you already received. And it throws away shape: [0.4, 0.4, 0.2] and [0.4, 0.3, 0.3] both return 0.1, though the first is a genuine two-way tie and the second is a leader against a diffuse field. Entropy separates them. Confidence does not.
The practical consequence: when your routing decision depends on whether there is a credible second option, gate on the margin between the top two probabilities, not on confidence. The docs invite exactly this, noting you are never locked into their definition. And note that Noul, the primitive most people will reach for first, returns no confidence at all, so a Noul-based pipeline has no built-in gate. You are thresholding the raw probability, which is the right thing to do anyway.
Takeaway
In TypeSafe's own GDPR cookbook, the document costs about 11,150 input tokens and all thirteen questions asked of it cost about 680. The question is nearly free; the context is the entire bill.
That inverts the habit every LLM pipeline has trained into us. With a generative model you minimize calls by asking one careful question, because each extra question grows the output you pay a premium for and risks polluting the context. Here the state crosses the wire once and the marginal question costs about fifty tokens, evaluated in isolation.
So the discipline flips. Ask every question your code might conceivably branch on, including the speculative ones you will throw away, and spend your engineering effort on shrinking and filtering the state. The parallel questions cookbook measured 12.2x cheaper and 10.0x faster for the same answers, purely from not re-sending the document. That is a refactor, not a model upgrade.
TL;DR For Engineers
Jev takes a state plus typed questions and returns Choice, Score, and Noul answers with probabilities in one parallel pass. $0.042 per million input tokens, output free, 70ms to 500ms.
Schema violations are zero by construction, which is a type guarantee, not a correctness guarantee. A wrong valid enum is the real failure mode and it does not throw.
confidence = (n · max(p) − 1) / (n − 1), verified against all seven published examples. It ignores distribution shape, and Noul answers do not carry one. Gate on the top-two margin instead.Cost scales with state, not questions: 13 questions over a 54k-character article cost 12.2x less batched, with identical answers.
Keep arithmetic, counting, date comparison, and text generation out of it. The docs say so, and the jaggedness page is the most useful page they publish.
Explain It Like I'm New
Software is full of small judgment calls that code cannot quite make. Is this support ticket urgent? Does this page actually answer the question someone asked? Is this command about to delete something important? A programmer can write rules for the easy cases, but the hard ones need something closer to reading comprehension.
Chatbot-style AI can do that reading, but it answers the way a person would, in sentences. Software cannot use a sentence. So teams wrap the model in extra code that asks it to reply in a fixed format, then checks the reply, then asks again when the format comes back wrong. It works, and it is slow and expensive, because the model is writing an essay when all you wanted was a verdict.
Jev skips the writing. You tell it the exact set of answers it is allowed to give, like a multiple-choice test with no essay section, and it hands back one of those answers plus a number saying how sure it is. Because it is not composing sentences, it is roughly a hundred times faster and cheaper on this kind of work.
The catch is the same as with a multiple-choice test. The answer is always one of your options, which does not mean it is the right option. The useful part is that it tells you when it is unsure, so your code can send the hard cases to a person.
See It In Action
Side-by-side demo and workflow evals, TypeSafe (launch post, evals site). Watch all probabilities appear at once against a token-by-token LLM, then dig into the four published workflows with full queries and disagreements.
Doom and Wikiracing, TypeSafe (launch post). The Doom bot runs about ten queries a second at roughly $7 per hour, which backs out to about 4,600 input tokens of game state per call. That is the clearest demonstration of what sub-second, sub-cent decisions enable.
The playground, TypeSafe (console). Paste any text as state, add a Noul, and watch probabilities change as you reword the instruction. Fifteen minutes here teaches more about literal reading than any doc page.
Jev judged everything I have written, Every (write-up). The best independent test so far: 777 judgments in under 0.7 seconds for about a quarter of a cent, and a head to head where Jev caught six of seven planted defects against Fable 5.1's seven, at a median 0.35 seconds per passage versus 8.83.
pi-warden, DevMortimer (repo). A coding-agent guardrail that judges every tool call with four typed questions in about 250ms. Read it for how the thresholds are chosen, which is the part nobody writes down.
Community Conversation
Hacker News launch thread (discussion). The top comment went straight at the hallucination claim, and the consensus landed where it should: a schema-valid answer can still be a wrong answer. Commenter ramon156's apples to oranges objection on 70ms versus 329s is the fairest critique of the speed framing.
Diogo Almeida, TypeSafe founder, posting as CompleteSkeptic (reply). On using Jev for code via an AST, he says the hard part is state engineering, getting the right dependencies into context, which is the same bottleneck the jaggedness page describes.
openjev and r/LocalLLaMA (repo). Reproduced the interface on Qwen3.5-4B logits in a day and published the gap: 0.845 modal agreement versus Jev's 0.883. The strongest evidence that the moat is calibration rather than the API shape.
The Register (article). Makes the point that hallucination-free is not a fair comparison when the output is not natural language, and that structured responses can still be incorrect.
Vercel AI SDK team (changelog). Shipped
evaluatewith Jev behind it in 36 hours, and the SDK docs carry a warning worth repeating: swap an OpenAI or Anthropic model in as the evaluator and the probabilities are prompted estimates with no calibration guarantee, so your thresholds do not transfer.Firecrawl's teardown, Hiba Fathima (post). Six use cases that hold up, all with the same shape: Jev sits next to an LLM as a filter or a judge, never as a replacement.
Constrain The Output Space, Not The Prose
The durable idea here is older than Jev and sharper in it than anywhere else: if your software is going to act on a model's answer, bound the answer space before you ask. Everything good about this model follows from that one constraint, including the free output tokens, the parallel sampling, and the absent parse step.
What does not follow is correctness. The schema guarantee removes a whole class of integration bugs and leaves the hard one untouched, and the confidence field, being a rescaled peak probability, will not catch it for you. Use the distribution, keep the thresholds in code, test every question as its own classifier, and treat the state as the thing you engineer.
Ask more questions. Send less context. That is the whole operating manual.
References
Training language models to follow instructions with human feedback, Ouyang et al., NeurIPS 2022. The InstructGPT paper TypeSafe's founder co-authored, and the RLHF lineage RLCD defines itself against.
On Calibration of Modern Neural Networks, Guo et al., ICML 2017. The paper that made calibration a first-class metric, and the standard for what "calibrated" should mean.
Language Models (Mostly) Know What They Know, Kadavath et al., 2022. Evidence that the uncertainty signal exists inside models, which is the premise of training a model to emit it directly.
Just Ask for Calibration, Tian et al., EMNLP 2023. Measures how badly prompted confidence scores are calibrated, the baseline Jev's confidence has to beat.
Efficient Guided Generation for Large Language Models, Willard and Louf, 2023. The Outlines approach to constrained decoding, the generative alternative to giving up strings.
Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models, Tam et al., EMNLP 2024. Finds that format constraints can cost reasoning quality in generative models, which is the tradeoff a System One model sidesteps by not reasoning in tokens at all.
GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Zaratiana et al., NAACL 2024. The encoder-side lineage that r/LocalLLaMA pointed at, worth reading before concluding this category is new.
Jev is a non-generative model that takes a state plus typed questions and returns choices, scores, and probabilities in one parallel pass, at $0.042 per million input tokens with free output and sub-second latency. The schema guarantee is real and removes the parse-validate-retry loop, but it says nothing about whether the chosen option is correct, and the confidence score turns out to be a closed-form rescaling of the top probability rather than an independent signal. Ask many questions per call, send less state, and keep your thresholds in code.
Keep Going
The habit from this issue: before your next structured-output call, count how many of the tokens you are paying for are the document and how many are the question. If the document dominates, you are paying to re-send it every time you think of something else to ask.
SnackOnAI runs this teardown weekly on the systems engineers actually deploy: new model classes, agent frameworks, serving stacks, and the launch claims whose footnotes disagree with their headlines. No announcements, no press release summaries. Subscribe at snackonai.com and join 10,000+ engineers reading it.
Forward this to whoever on your team is still running a JSON-mode retry loop to get a yes or no.
Sponsored Ad If you enjoy practical AI insights, check out SnackOnAI and support the newsletter by subscribing, sharing, and exploring our sponsored ad, it helps us keep building and delivering value 🚀
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.
With Attio, you’ll get:
Leads automatically prioritised and routed to the right rep
Expansion and risk signals caught the moment they land
Follow-ups written in your voice, already there when you arrive
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?


