Inside Laya: The Architecture of a Decision Model, Rebuilt From Its Code
Decision models answer in one forward pass, with a probability for every option you hand them, and never write a word. Nobody has published a paper on how one is built. Laya's code is open, so here is the whole machine drawn out, layer by layer, in the style of the transformer diagram you already know.
01 · The machine in one picture
A decision model, start to finish
Put your question and its answer options into the input row, each option behind its own [MASK] marker. A bidirectional encoder reads the entire row, so information flows between the options and the case being judged. The model then pulls one vector out at each marker, turns each into a single score, and runs a softmax across only those scores. The highest probability is the answer. One trip through the network, nothing generated.
Step through it below. Each stage maps to a section of this article.
All stages shown. Press "Step through a forward pass" to walk the row from the bottom up.
02 · The case files
Jev is closed, Laya is open
This teardown reads Laya's source line by line and redraws it. Two models define the category right now, and they sit at opposite ends of the disclosure spectrum.
Jev published claims only
TypeSafe AI's "System One" model, launched September 2026. Closed weights, hosted API at POST /v1/systemone.
- Launch post describes a new architecture trained with RLCD
- No paper, no diagram, no code
- Third-party p50 latency 236-276 ms
Laya read from code
Convai Innovations. Apache 2.0, weights on Hugging Face, full source on GitHub.
- Claims to be RLCD-trained, with its own implementation
- Cites two earlier papers, neither of which draws this model
- Serves a Jev-compatible endpoint shape
Throughout the article, claims are tagged by where they come from: read from code for what the source or the model card states, inference for my own reading, published claims only for what the vendor said. Laya moves fast. This reading is pinned to v0.3.5; the README is at 0.3.21 as of this writing. The architecture has not changed between them; the runtime has.
03 · The known suspect
How an LLM answers the same question
Take one routine call: is this answer good enough to pass, or does it need rework? An LLM is a decoder. Each token can attend only to the tokens before it. It reads the prompt once (prefill), then writes one token, appends it, and runs again (decode). At every step the final layer scores the whole vocabulary, commonly 100,000 entries or more, and picks one.
To say "pass", it must write "pass", and your program must parse it. Structured outputs constrain the format, but generation is still one token per step. Ask for a confidence number and it writes that too. Xiong et al. (2023) found verbalized confidence tends to run high.
Keep one trick in mind. You can read the LLM's own probability for the token "pass" versus "rework" instead of asking it to write a number. That is the closest an LLM gets to what a decision model does natively, and the difference between the two is the whole story of this article.
04 · Step 1: the input row
Options are written into the input, not baked into the output layer
read from code Laya flattens a request into one long row of tokens. A start token, the question type and your instructions, a separator. Then the options, each preceded by its own [MASK] token. Another separator, then the state (the thing being judged), then a closing separator.
Three question types share this layout:
| Type | What it does | Output |
|---|---|---|
choice | Pick one of k labels | label, distribution, confidence |
score | Place the state on an ordinal scale (0, 1, 2 ...) | expected level, distribution |
noul | Yes or no, scored as two slots: false, true | P(true) |
The budget is a design constraint, not a detail
The question and all options share one head budget. When options overflow it, every option is trimmed to the same length and the trimmed words never reach the model. Past about 20 options with short descriptions, labels start to look alike to the encoder. This is why Laya scores 0.425 on Banking77's 77 intents while Jev, which accepts up to 255 options, posts 0.870. The repo's remedies are a wider head_max_len, or an embedding shortlist (predict_shortlist) before the forward pass.
inference For enterprise schemas, this argues for hierarchical questions: route coarse first (which desk), then fine (which break type), rather than one 60-way choice.
05 · Step 2: an encoder that reads both ways
Why the option markers can see the state that comes after them
This is the first structural break from an LLM. The encoder is bidirectional, so information flows forward and backward along the row. The [MASK] in front of "pass" can absorb the state that sits after it. In a causal decoder, a token placed before the state could never see it.
read from code The backbone is ModernBERT-large, the open encoder from Answer.AI and LightOn (Warner et al., 2024): 28 layers, each token represented as a 1,024-number vector. ModernBERT alternates full-row global attention with shorter local windows across its layers. The global layers are what let information cross the whole row.
Laya does not freeze it. All ~395M encoder parameters are trained together with the head on top. The multilingual sibling swaps in mmBERT-base (22 layers, 256k vocabulary, 322M total) to cover 100+ languages, and a router picks between checkpoints by detecting script and language before the forward pass.
06 · Step 3: where the answer comes from
Gather, score, softmax over your options only
read from code After the encoder, Laya adds a learned question-type embedding to every token vector, so the head knows whether it is handling a choice, a score or a yes/no. Then come two transformer layers trained from scratch. Together with the scorer, these form the decision head and bring the model to about 421M parameters.
The head then goes to each [MASK] position and gathers the vector there: one for "pass", one for "rework". Because the encoder reads both ways, each vector already reflects the state plus its own option. A small two-layer network turns each vector into one scalar. A softmax over that short list turns scores into probabilities that sum to one.
The answer space is whatever you typed into the row. Tomorrow you want different options, you write different options. Nothing is retrained, because no output layer is tied to a fixed label set.
One pass, many questions
Ask five questions and Laya builds five rows and batches them in the same forward pass. Each row carries its own copy of the state, so each question gets its own reading. The README lists 39.5 ms for one question on a T4 with the English checkpoint, and 7.2 ms per question batched on the multilingual one.
What "confidence" means here
It cannot answer outside your options, because it has no way to write anything else. It can still point at the wrong option, and those are different failure modes. For choice and score, the printed confidence is 1 minus the normalized entropy of the distribution. It measures how concentrated the probabilities are, not the probability that the chosen answer is right. The newer runtime adds answer_confidence, the max probability, which is the number to threshold on. Try both below.
Interactive: temperature, probability and the two confidence numbers
Scores are fixed example values for a 4-way routing question. Temperature divides the scores before the softmax. Drag it toward 0.1 and a small gap in scores becomes near-certainty, which is exactly why v0.3.5 began clamping stored temperatures.
07 · Training: RLCD end to end
Rewarding honest probabilities instead of preferred text
A typical chat LLM learns next-token prediction, is tuned on examples, then often goes through RLHF, which rewards answers people prefer. TypeSafe's argument is that preferred text is the right target for a chat product and the wrong one for software that has to make a decision. That argument is where the name comes from: Reinforcement Learning for Calibrated Decisions.
The reward: a strictly proper scoring rule
Picture a forecaster who honestly believes there is a 70% chance of rain tomorrow. Under a proper scoring rule, saying 90% loses points on average, and so does saying 50%. The only way to maximize expected score is to report the number you actually believe. read from code Laya's main rule is the log score, the log of the probability given to what actually happened. It adds a spherical score, and for ordinal score questions subtracts a ranked probability penalty.
read from code The loop: the model scores the options, makes four copies of those scores and adds a little random noise to each. Each noisy copy becomes a slightly different distribution, and each is rewarded against the target. Every copy is compared with the average of the four, and the model is pushed toward the copies that beat the average. GRPO, the RL method common in LLM post-training, uses the same group-mean idea.
The notebook also has a second term: an ordinary supervised loss that pulls probabilities straight toward the target. inference Because the log score rewards the same thing cross-entropy does, the RL part and the supervised part pull in the same direction. In practice this looks less like pure RL and more like supervised calibration training with an exploration term layered on.
08 · Calibration, and the limits
Training on honest probabilities does not, by itself, give honest probabilities
The model card says so plainly. Expected calibration error (ECE) measures the gap between how confident a model is and how often it is right. The shipped checkpoints are over-confident. Fitting one temperature per (question type, option count) on held-out data moved the English checkpoint's mean ECE from 0.466 to 0.081.
The config file shows why that step needs watching. It stored a fitted temperature near 0.1 for choices with 11 or more options, which turns small score gaps into near-certainty (try it in Figure 5). Since v0.3.5 the library clamps temperatures to the range 0.5 to 5.0 at load and warns that those answers are uncalibrated.
| typed-decisions benchmark (2,000 decisions) | accuracy |
|---|---|
| random guess | 0.318 |
| Laya base English checkpoint, zero-shot | 0.362 |
| always pick the per-question majority class | 0.461 |
| Jev 1.13.0 (published) | 0.727 |
| teacher self-agreement ceiling | 0.735 |
| laya-typed-decisions, fine-tuned on this benchmark's train split | 0.766 |
The base checkpoint sits below the majority-class baseline. All of the capability on this benchmark comes from fine-tuning. So the honest framing is that Laya is a fast base you specialise on your own data, not a zero-shot decision engine. The repo ships a notebook that runs the full loop (dataset, RLCD training, temperature fitting, evaluation) on Kaggle's free 2x T4 GPUs in roughly 4 to 5 hours.
Failure modes to test before production
| Failure | Symptom | Mitigation |
|---|---|---|
| High cardinality | Accuracy falls past ~20 options as labels are trimmed | Shortlist by embedding, or split coarse/fine |
Label leakage in noul | English checkpoint follows the false/true labels, not the state | Override labels (A/B), or use a 2-way choice |
| Position bias | Identical options get different logits by slot | Average over rotated option_order in one pass |
| Negation | "Do not cancel" routed to cancel_account | Hold-out tests on negated phrasings |
| Wrong language, high confidence | Khmer on English checkpoint: 0.000 accuracy at 0.952 raw confidence | Route before the forward pass |
09 · Why decision models exist
The gap between a random forest and an LLM
Software is full of small decisions over unstructured input: which team gets this ticket, is this answer good enough to publish, is this break a timing difference or a real one. Each is really a fuzzy if statement. Laya's design sits on a lineage that has been converging on one idea: put the labels in the input instead of fixing them in the output layer.
| Approach | Reads raw text | Labels set at call time | Passes per call | Native probability |
|---|---|---|---|---|
| Random forest / GBDT | No, hand-built features | No | 1 | Yes |
| Fine-tuned BERT classifier (2018) | Yes | No, tied to output layer | 1 | Yes |
| NLI zero-shot (Yin, Hay, Roth 2019) | Yes | Yes | 1 per label | Per label |
| GLiNER (2023), labels in encoder input | Yes | Yes | 1 | Per span |
| Laya / decision models | Yes | Yes | 1 for all questions | Softmax over options |
| LLM | Yes | Yes | 1 + n decode steps | Written, or token logprob |
The gap is in the middle. You want to write the question and options at call time, get the answer in one pass, and get a probability you can calibrate and then threshold. That threshold is what makes a cascade work.
Where this lands in fund operations
inference Asset servicing runs on exactly this kind of fuzzy if. A few places the shape fits, each with a short, stable option set:
| Decision | Type | Options |
|---|---|---|
| NAV reconciliation break: root cause | choice | pricing, FX, corporate action, timing, booking error, other |
| Trade exception: route to desk | choice | custody, TA, fund accounting, middle office |
| Client query urgency | score | routine, same day, blocks NAV strike |
| Does this email instruct a cash movement? | noul | P(true), gated hard to human review |
None of these benefit from a model that can write prose. All of them benefit from a calibrated number, a 40 ms latency budget inside a batch window, and weights that never leave the bank's network. The work is in the fine-tuning set and the temperature fit, not in the model.
Next time a decision model hands you "pass" with 0.94 next to it, you now have a picture, drawn from the code, of where that number came from. Check it against your own data before your code trusts it.
10 · Sources
References
- Laya source, github.com/NandhaKishorM/laya (v0.3.5 read for this article; README at 0.3.21).
- Laya model card, huggingface.co/convaiinnovations/laya, Convai Innovations.
- TypeSafe AI, Introducing System One Models and Jev (15 Sep 2026).
- Warner et al., ModernBERT (2024).
- Devlin et al., BERT (2018).
- Yin, Hay and Roth, Benchmarking Zero-shot Text Classification (2019).
- Zaratiana et al., GLiNER (2023).
- Xiong et al., Can LLMs Express Their Uncertainty? (2023).
Benchmark and latency figures are as published in the Laya README and model card; Jev figures there are third-party, not measured by the Laya author. Values labelled "example" in the figures are illustrations.