FinTech

Jev in the Break Queue: Your Break Queue Doesn't Need a Model That Talks

TypeSafe's Jev returns a typed decision with a probability instead of prose, in milliseconds and for a fraction of a cent. Applied to NAV reconciliation break triage, with a live simulator showing where to set the confidence gate and why calibration decides whether that gate means anything.

October 1, 2026·11 min read
AIFund AccountingReconciliationJevModel RiskCalibration
LLM writes a verdict → Jev returns a typed decision
seconds per break → milliseconds per break
analyst triages everything → analyst sees the unsure middle

Let me set the scene. It's 6:40 a.m. and the overnight reconciliation has dropped twelve thousand breaks into your queue. Cash, positions, trades, income. Every one needs the same first answer before anyone can fix it: what kind of break is this, and who owns it?

You've probably already tried putting an LLM on that question. It works. It also writes you a paragraph you never read, takes seconds you don't have, and bills you for output tokens on a decision that was always going to be one of six labels.

On 15 September 2026, TypeSafe AI shipped a model built for exactly that shape of problem. It's called Jev, and it can't write a sentence.

01The model

A decision model, not a smaller chatbot

Thesis: most AI calls inside back-office software are judgments, not conversations.

TypeSafe calls Jev the first System One model, borrowing Kahneman's fast, intuitive System 1 against slow, deliberate System 2. You send it a state (text or JSON your software already has) and a set of typed questions whose possible answers you define up front. It returns one of your answers, with a probability attached. No prose. No explanation. No malformed JSON, because there's no JSON being written.

Three question types make up the whole surface, and one request can mix all of them over the same state:

Noul
Does this condition hold?
returns P(yes), 0 to 1
no separate confidence: the probability is the certainty
Choice
Which one of these options?
returns the pick, a probability per option (up to 255) and a confidence
Score
Where on this ordered rubric?
returns a level, per-level probabilities, legend and a confidence

Each question sees the same state but not the other questions, and TypeSafe says they're answered in one parallel pass rather than token by token. That's where the speed comes from. TypeSafe quotes 70 to 500 milliseconds end to end against seconds for frontier models, and charges $0.042 per million input tokens with output free. Version 1.13 takes text and JSON only, up to roughly 32,000 tokens of state plus question.

The failure mode changes. An LLM can invent a seventh break category. Jev can't. It can only pick the wrong one of your six, confidently.

That's the part the launch posts underplay. "Can't hallucinate" covers the shape of the answer, not whether it's right. TypeSafe's own FAQ concedes this, and the first ten days of outside testing were blunt about it.

02The use case

NAV break triage, one call per break

Thesis: split the break into what code computes and what needs judgment, and ask Jev only the second half.

Here's a cash break coming off a custodian feed. Notice what's not in the questions. Materiality is a threshold comparison, and TypeSafe's own limitations page for 1.13 says plainly that Jev isn't a calculator and reads dates as text. So code computes basis points against NAV and days outstanding. Jev gets the judgment calls.

# POST https://api.typesafe.ai/v1/systemone   (shape abbreviated; field names per TypeSafe docs)
{
  "model": "jev-1.13.0",            # pin the version, never jev-latest, for MRM
  "state": {
    "fund": "GLB-EQ-INC-07", "break_type": "cash", "ccy": "EUR",
    "book_narrative": "DIV ACCR NESN 3.00 CHF",
    "custodian_narrative": "CASH DIV NESTLE NET WHT 35%",
    "materiality_bps": 0.6,                # computed in code
    "days_open": 2                         # computed in code
  },
  "questions": {
    "root_cause": { "type": "choice", "options":
      ["opt_1: timing", "opt_2: price", "opt_3: corporate action / tax",
       "opt_4: fx", "opt_5: trade capture", "opt_6: other"] },
    "same_economic_event": { "type": "noul",
      "statement": "Both narratives describe the same cash event" },
    "urgency": { "type": "score", "levels":
      ["0 resolves itself", "1 fix today", "2 fix before NAV strike", "3 escalate now"] }
  }
}

The answer comes back as root_cause = opt_3 with a probability per option, a confidence, a P(yes) for the same-event check, and an urgency level. Your harness turns those into a queue assignment. Jev never touches the ledger.

Applied

Name options neutrally. A preprint by Sun, Xu, Shi and Yang swapped which loaded names (yes/no) went with which meaning and watched Jev's AUC fall from 0.81 to 0.58, while invalid answers stayed at zero. Neutral labels showed little effect. In a break queue, that means opt_3 beats "probably a WHT issue".

03Live simulation

Where do you set the gate?

Thesis: the threshold is a risk decision, and calibration decides whether it means anything.

Confidence-gated routing is the pattern everyone serious converges on: let Jev act when it's sure, send the middle to an LLM, send what the LLM can't settle to a person. Move the gate and watch three things trade against each other. Misroutes. Spend. Analyst hours.

Then move the calibration gap. That's how far Jev's stated confidence sits above its real hit rate on your data. An independent out-of-distribution test on invented support tickets measured an average gap around 0.107, roughly three times its gap on public benchmarks. Your break narratives look a lot more like invented tickets than like public benchmarks.

Break triage router · illustrative model
60%
Routed by Jev alone
7,166 of 12,000 breaks
1,145
Misroutes per day
Jev would claim 429
$7.17
Model spend per day
LLM on every break: $17
145 h
Analyst triage hours
today: 1,200 h
Jev confidence across the queue (teal clears the gate)
τ 0.850.30.50.70.91.0n
Stated confidence vs real accuracy
0.50.70.90.50.70.9statedactual
Jev decidesescalated to LLMreaches an analyst
Illustrative assumptions, not measured bank or TypeSafe figures: 800 input tokens per break. Jev at $0.042/M input, output free. Escalation LLM at $1/M input and $5/M output with 120 output tokens. The LLM settles 70% of what it receives. Analyst triage takes 6 minutes per break. Confidence distribution is synthetic (bimodal: clean breaks and messy narratives). "Recalibrated" fits a monotone map on 500 labelled past breaks, which removes most of the gap and leaves a residual of 0.01.

Play with it for a minute and one thing jumps out. With raw confidence and a 0.10 gap, the gate that looks safe at 0.85 is quietly sending over a thousand breaks a day to the wrong queue, nearly three times what Jev's own numbers imply. Flip to recalibrated and the gate's promise and its result line up again, and now you can raise τ knowing what you're buying. The model didn't change. Your labelled data did.

The threshold isn't a tuning knob. It's a control, and SR 11-7 will ask you to evidence it.
04Fit

What in fund services is Jev-shaped

Thesis: if the answer is a label you already route on, it's a candidate. If the answer is a number, a date or a sentence, it isn't.

DecisionFitWhy
Recon break root cause and ownerChoiceFixed label set, huge volume, confidence-routable
Corporate action event type from a custodian noticeChoiceClassification over free text; code still extracts ratios and dates
Book vs custodian narrative describe the same eventNoulPairwise match judgment, the same shape Pricogni ran 9,081 times for $0.32
Price challenge urgency before NAV strikeScoreOrdered rubric; inspect mass on the top level, not just the mean
Agent guardrail: may this tool call post to the GL?Noul + humanGood gate, but confidence alone never authorizes money movement
Investor onboarding risk ratingHuman firstDecision about people; a bare probability with no reason is hard to defend
Accruals, NAV arithmetic, settlement date logicNoNot a calculator; reads dates as text
Pull the dividend rate out of a PDF noticeNoCan't extract values. Code finds candidates, Jev can pick among them
Explain the break to the client or auditorNoNeeds words. That's the LLM's job in the escalation lane
05Risk

What your model risk team will ask

Thesis: the speed and price claims have held up; the accuracy and vendor story still need evidence from you.

The largest independent test so far, a not-yet-reviewed preprint by Ibrahim and Zaki, put Jev against 19 LLMs on 15 labelling tasks. Jev trailed the best LLM for each task on 14 of them, by a median of 11.6 macro-F1 points, at a median 44 times lower cost. The interesting result was the hybrid: send only Jev's low-confidence items to an LLM and you matched or beat the LLM alone, at a quarter to half its cost. That's the architecture in the diagram above.

Before Jev touches a live queue

Pin the version. jev-1.13.0, not the moving alias. A silent model update is a model change.

Build the labelled set first. A few hundred past breaks with known root causes. That set recalibrates the gate, benchmarks Jev against a cheap LLM and a two-line rule, and survives a vendor switch.

Test the dumbest baseline. One open benchmark found a two-line text-matching rule within a few points of both Jev and Haiku on phishing. Half of your break narratives may be regex-shaped.

Ask the vendor questions. It's proprietary with no weights or self-hosting, processes data in the US only for now, and rate limits can change during early access. For client data under a servicing agreement, that's the first conversation, not the last.

Treat the state as untrusted. TypeSafe's own limitations page says text written to steer the model can move the answer. A custodian narrative is input from outside your perimeter.

06Where I land

The model is cheap. The ruler isn't.

Jev is the first time a judgment costs a fraction of a cent and comes back before the analyst has finished reading the break ID. That changes the design question. You stop asking whether you can afford to put a model on the queue and start asking how many narrow questions to ask per break.

But the thing that makes it safe to deploy isn't Jev. It's the labelled set, the calibration map and the gate you can defend to model risk. That's the same argument I've been making about FundOps-Bench: the verifier is the product. Rivals rebuilt Jev's interface on open models within a week. Nobody can rebuild your five hundred labelled breaks.

So here's my question for you. If you put a gate on your break queue tomorrow, where would you set τ, and could you show anyone why?

Sources
  1. TypeSafe: Introducing System One models and Jev
  2. TypeSafe docs: Primitives and Jev 1.13 jaggedness
  3. Pooya Golchian: What Is Jev, and Does It Replace LLMs?
  4. Ibrahim and Zaki: Jev vs 19 LLMs on labelling tasks (preprint)
  5. Sun, Xu, Shi and Yang: label-name sensitivity (preprint)
  6. scienthoon: Jev out-of-distribution calibration
  7. jev-phishing-bench
  8. Pricogni: the thirty-cent judge
  9. Confidence-aware fallback with Choice, Score and Noul
Views are my own. Break data in the simulation is synthetic.

Share this article