Jev in the Break Queue: Your Break Queue Doesn't Need a Model That Talks
TypeSafe's Jev returns a typed decision with a probability instead of prose, in milliseconds and for a fraction of a cent. Applied to NAV reconciliation break triage, with a live simulator showing where to set the confidence gate and why calibration decides whether that gate means anything.
Let me set the scene. It's 6:40 a.m. and the overnight reconciliation has dropped twelve thousand breaks into your queue. Cash, positions, trades, income. Every one needs the same first answer before anyone can fix it: what kind of break is this, and who owns it?
You've probably already tried putting an LLM on that question. It works. It also writes you a paragraph you never read, takes seconds you don't have, and bills you for output tokens on a decision that was always going to be one of six labels.
On 15 September 2026, TypeSafe AI shipped a model built for exactly that shape of problem. It's called Jev, and it can't write a sentence.
A decision model, not a smaller chatbot
Thesis: most AI calls inside back-office software are judgments, not conversations.
TypeSafe calls Jev the first System One model, borrowing Kahneman's fast, intuitive System 1 against slow, deliberate System 2. You send it a state (text or JSON your software already has) and a set of typed questions whose possible answers you define up front. It returns one of your answers, with a probability attached. No prose. No explanation. No malformed JSON, because there's no JSON being written.
Three question types make up the whole surface, and one request can mix all of them over the same state:
no separate confidence: the probability is the certainty
Each question sees the same state but not the other questions, and TypeSafe says they're answered in one parallel pass rather than token by token. That's where the speed comes from. TypeSafe quotes 70 to 500 milliseconds end to end against seconds for frontier models, and charges $0.042 per million input tokens with output free. Version 1.13 takes text and JSON only, up to roughly 32,000 tokens of state plus question.
The failure mode changes. An LLM can invent a seventh break category. Jev can't. It can only pick the wrong one of your six, confidently.
That's the part the launch posts underplay. "Can't hallucinate" covers the shape of the answer, not whether it's right. TypeSafe's own FAQ concedes this, and the first ten days of outside testing were blunt about it.
NAV break triage, one call per break
Thesis: split the break into what code computes and what needs judgment, and ask Jev only the second half.
Here's a cash break coming off a custodian feed. Notice what's not in the questions. Materiality is a threshold comparison, and TypeSafe's own limitations page for 1.13 says plainly that Jev isn't a calculator and reads dates as text. So code computes basis points against NAV and days outstanding. Jev gets the judgment calls.
# POST https://api.typesafe.ai/v1/systemone (shape abbreviated; field names per TypeSafe docs) { "model": "jev-1.13.0", # pin the version, never jev-latest, for MRM "state": { "fund": "GLB-EQ-INC-07", "break_type": "cash", "ccy": "EUR", "book_narrative": "DIV ACCR NESN 3.00 CHF", "custodian_narrative": "CASH DIV NESTLE NET WHT 35%", "materiality_bps": 0.6, # computed in code "days_open": 2 # computed in code }, "questions": { "root_cause": { "type": "choice", "options": ["opt_1: timing", "opt_2: price", "opt_3: corporate action / tax", "opt_4: fx", "opt_5: trade capture", "opt_6: other"] }, "same_economic_event": { "type": "noul", "statement": "Both narratives describe the same cash event" }, "urgency": { "type": "score", "levels": ["0 resolves itself", "1 fix today", "2 fix before NAV strike", "3 escalate now"] } } }
The answer comes back as root_cause = opt_3 with a probability per option, a confidence, a P(yes) for the same-event check, and an urgency level. Your harness turns those into a queue assignment. Jev never touches the ledger.
Name options neutrally. A preprint by Sun, Xu, Shi and Yang swapped which loaded names (yes/no) went with which meaning and watched Jev's AUC fall from 0.81 to 0.58, while invalid answers stayed at zero. Neutral labels showed little effect. In a break queue, that means opt_3 beats "probably a WHT issue".
Where do you set the gate?
Thesis: the threshold is a risk decision, and calibration decides whether it means anything.
Confidence-gated routing is the pattern everyone serious converges on: let Jev act when it's sure, send the middle to an LLM, send what the LLM can't settle to a person. Move the gate and watch three things trade against each other. Misroutes. Spend. Analyst hours.
Then move the calibration gap. That's how far Jev's stated confidence sits above its real hit rate on your data. An independent out-of-distribution test on invented support tickets measured an average gap around 0.107, roughly three times its gap on public benchmarks. Your break narratives look a lot more like invented tickets than like public benchmarks.
Play with it for a minute and one thing jumps out. With raw confidence and a 0.10 gap, the gate that looks safe at 0.85 is quietly sending over a thousand breaks a day to the wrong queue, nearly three times what Jev's own numbers imply. Flip to recalibrated and the gate's promise and its result line up again, and now you can raise τ knowing what you're buying. The model didn't change. Your labelled data did.
The threshold isn't a tuning knob. It's a control, and SR 11-7 will ask you to evidence it.
What in fund services is Jev-shaped
Thesis: if the answer is a label you already route on, it's a candidate. If the answer is a number, a date or a sentence, it isn't.
| Decision | Fit | Why |
|---|---|---|
| Recon break root cause and owner | Choice | Fixed label set, huge volume, confidence-routable |
| Corporate action event type from a custodian notice | Choice | Classification over free text; code still extracts ratios and dates |
| Book vs custodian narrative describe the same event | Noul | Pairwise match judgment, the same shape Pricogni ran 9,081 times for $0.32 |
| Price challenge urgency before NAV strike | Score | Ordered rubric; inspect mass on the top level, not just the mean |
| Agent guardrail: may this tool call post to the GL? | Noul + human | Good gate, but confidence alone never authorizes money movement |
| Investor onboarding risk rating | Human first | Decision about people; a bare probability with no reason is hard to defend |
| Accruals, NAV arithmetic, settlement date logic | No | Not a calculator; reads dates as text |
| Pull the dividend rate out of a PDF notice | No | Can't extract values. Code finds candidates, Jev can pick among them |
| Explain the break to the client or auditor | No | Needs words. That's the LLM's job in the escalation lane |
What your model risk team will ask
Thesis: the speed and price claims have held up; the accuracy and vendor story still need evidence from you.
The largest independent test so far, a not-yet-reviewed preprint by Ibrahim and Zaki, put Jev against 19 LLMs on 15 labelling tasks. Jev trailed the best LLM for each task on 14 of them, by a median of 11.6 macro-F1 points, at a median 44 times lower cost. The interesting result was the hybrid: send only Jev's low-confidence items to an LLM and you matched or beat the LLM alone, at a quarter to half its cost. That's the architecture in the diagram above.
Before Jev touches a live queue
Pin the version. jev-1.13.0, not the moving alias. A silent model update is a model change.
Build the labelled set first. A few hundred past breaks with known root causes. That set recalibrates the gate, benchmarks Jev against a cheap LLM and a two-line rule, and survives a vendor switch.
Test the dumbest baseline. One open benchmark found a two-line text-matching rule within a few points of both Jev and Haiku on phishing. Half of your break narratives may be regex-shaped.
Ask the vendor questions. It's proprietary with no weights or self-hosting, processes data in the US only for now, and rate limits can change during early access. For client data under a servicing agreement, that's the first conversation, not the last.
Treat the state as untrusted. TypeSafe's own limitations page says text written to steer the model can move the answer. A custodian narrative is input from outside your perimeter.
The model is cheap. The ruler isn't.
Jev is the first time a judgment costs a fraction of a cent and comes back before the analyst has finished reading the break ID. That changes the design question. You stop asking whether you can afford to put a model on the queue and start asking how many narrow questions to ask per break.
But the thing that makes it safe to deploy isn't Jev. It's the labelled set, the calibration map and the gate you can defend to model risk. That's the same argument I've been making about FundOps-Bench: the verifier is the product. Rivals rebuilt Jev's interface on open models within a week. Nobody can rebuild your five hundred labelled breaks.
So here's my question for you. If you put a gate on your break queue tomorrow, where would you set τ, and could you show anyone why?
- TypeSafe: Introducing System One models and Jev
- TypeSafe docs: Primitives and Jev 1.13 jaggedness
- Pooya Golchian: What Is Jev, and Does It Replace LLMs?
- Ibrahim and Zaki: Jev vs 19 LLMs on labelling tasks (preprint)
- Sun, Xu, Shi and Yang: label-name sensitivity (preprint)
- scienthoon: Jev out-of-distribution calibration
- jev-phishing-bench
- Pricogni: the thirty-cent judge
- Confidence-aware fallback with Choice, Score and Noul