Development

Inside Laya: The Architecture of a Decision Model, Rebuilt From Its Code

Decision models answer in one forward pass, with a probability for every option you hand them, and never write a word. Nobody has published a paper on how one is built. Laya's code is open, so here is the whole machine drawn out, layer by layer, in the style of the transformer diagram you already know.

October 2, 2026·18 min read
AIDecision ModelsTransformersModernBERTCalibrationLayaJev

01 · The machine in one picture

A decision model, start to finish

Put your question and its answer options into the input row, each option behind its own [MASK] marker. A bidirectional encoder reads the entire row, so information flows between the options and the case being judged. The model then pulls one vector out at each marker, turns each into a single score, and runs a softmax across only those scores. The highest probability is the answer. One trip through the network, nothing generated.

Step through it below. Each stage maps to a section of this article.

Laya architecture: input row, ModernBERT-large encoder repeated 28 times, question-type embedding, two decision-head transformer layers, gather at mask positions, option scorer, softmax over options, with a dimmed act/escalate head. Typed answer choice: pass p = [0.88, 0.12] confidence = 1 - H(p)/log k (example values) Softmax over k options / temperature T(type, k) Option scorer (2-layer MLP) s_pass = 2.1 s_rework = 0.1 gather h at [MASK] positions act / escalate head no usable signal yet 2x new, from scratch Feed-forward + residual, norm Self-attention (full row) + residual Decision head transformer head params ~26M | model total ~421M + Question-type embedding choice | score | noul 28x fully fine-tuned GeGLU feed-forward + residual Bidirectional attention + RoPE global layers + local-window layers ModernBERT-large encoder d_model = 1024 | ~395M params every token sees every token, both directions Token embeddings (1024-d) INPUT ROW (max 512 tokens: head budget 192, rest for the state) [CLS] type + instructions [SEP] [MASK] pass [MASK] rework [SEP] state ... [SEP] options come BEFORE the state; the encoder reads both ways, so order does not hide the state from the markers

All stages shown. Press "Step through a forward pass" to walk the row from the bottom up.

Figure 1. Laya (English checkpoint), drawn from laya/common.py at v0.3.5. Scores and probabilities in the figure are example values, not measurements. The dashed act/escalate head exists in code but its own README says it carries no usable signal, so it is drawn dimmed.
inputpretrained encoderdecision head (new)outputoption marker

02 · The case files

Jev is closed, Laya is open

This teardown reads Laya's source line by line and redraws it. Two models define the category right now, and they sit at opposite ends of the disclosure spectrum.

Jev published claims only

TypeSafe AI's "System One" model, launched September 2026. Closed weights, hosted API at POST /v1/systemone.

  • Launch post describes a new architecture trained with RLCD
  • No paper, no diagram, no code
  • Third-party p50 latency 236-276 ms

Laya read from code

Convai Innovations. Apache 2.0, weights on Hugging Face, full source on GitHub.

  • Claims to be RLCD-trained, with its own implementation
  • Cites two earlier papers, neither of which draws this model
  • Serves a Jev-compatible endpoint shape

Throughout the article, claims are tagged by where they come from: read from code for what the source or the model card states, inference for my own reading, published claims only for what the vendor said. Laya moves fast. This reading is pinned to v0.3.5; the README is at 0.3.21 as of this writing. The architecture has not changed between them; the runtime has.

03 · The known suspect

How an LLM answers the same question

Take one routine call: is this answer good enough to pass, or does it need rework? An LLM is a decoder. Each token can attend only to the tokens before it. It reads the prompt once (prefill), then writes one token, appends it, and runs again (decode). At every step the final layer scores the whole vocabulary, commonly 100,000 entries or more, and picks one.

To say "pass", it must write "pass", and your program must parse it. Structured outputs constrain the format, but generation is still one token per step. Ask for a confidence number and it writes that too. Xiong et al. (2023) found verbalized confidence tends to run high.

Keep one trick in mind. You can read the LLM's own probability for the token "pass" versus "rework" instead of asking it to write a number. That is the closest an LLM gets to what a decision model does natively, and the difference between the two is the whole story of this article.

LLM (decoder, causal) prompt: question + answer prefill decode 1 decode 2 decode n text "pass" -> parse each step: softmax over ~100k vocabulary tokens cost and latency grow with tokens read AND written Laya (encoder, bidirectional) row: q [M]pass [M]rework state one forward pass (encoder + head) p(pass), p(rework) softmax over exactly the k options you supplied ~40 ms on a T4 for one question; zero output tokens; cannot answer outside your option set
Figure 2. Same question, two machines. The LLM's probability for "pass" is a slice of a vocabulary-wide softmax from a model trained to predict the next token. Laya's is a softmax over your options only, from a head trained to score options.

04 · Step 1: the input row

Options are written into the input, not baked into the output layer

read from code Laya flattens a request into one long row of tokens. A start token, the question type and your instructions, a separator. Then the options, each preceded by its own [MASK] token. Another separator, then the state (the thing being judged), then a closing separator.

[CLS]choiceIs this answer good enough to ship?[SEP][MASK]pass: meets the bar[MASK]rework: needs changes[SEP]The NAV variance is explained by a stale FX rate on the EUR share class...[SEP]
head_max_len 192: question + options
remainder of 512: the state
Figure 3. One row per question. Budget proportions are for the English checkpoint; the multilingual and typed-decisions checkpoints use 1,024 tokens with a 256-token head.

Three question types share this layout:

TypeWhat it doesOutput
choicePick one of k labelslabel, distribution, confidence
scorePlace the state on an ordinal scale (0, 1, 2 ...)expected level, distribution
noulYes or no, scored as two slots: false, trueP(true)

The budget is a design constraint, not a detail

The question and all options share one head budget. When options overflow it, every option is trimmed to the same length and the trimmed words never reach the model. Past about 20 options with short descriptions, labels start to look alike to the encoder. This is why Laya scores 0.425 on Banking77's 77 intents while Jev, which accepts up to 255 options, posts 0.870. The repo's remedies are a wider head_max_len, or an embedding shortlist (predict_shortlist) before the forward pass.

inference For enterprise schemas, this argues for hierarchical questions: route coarse first (which desk), then fine (which break type), rather than one 60-way choice.

05 · Step 2: an encoder that reads both ways

Why the option markers can see the state that comes after them

This is the first structural break from an LLM. The encoder is bidirectional, so information flows forward and backward along the row. The [MASK] in front of "pass" can absorb the state that sits after it. In a causal decoder, a token placed before the state could never see it.

Causal (LLM decoder)CCQQMMooMMooSSM rows: state column emptyBidirectional (Laya encoder)CCQQMMooMMooSSM rows: state column filled (orange)C = CLS Q = question M = option marker o = option text S = state
Figure 4. Who can attend to whom, for a seven-token row. Rows are the token doing the reading. Under the causal mask the option markers (M) are blind to the state (S). Under the bidirectional mask they read all of it before any score is computed.

read from code The backbone is ModernBERT-large, the open encoder from Answer.AI and LightOn (Warner et al., 2024): 28 layers, each token represented as a 1,024-number vector. ModernBERT alternates full-row global attention with shorter local windows across its layers. The global layers are what let information cross the whole row.

Laya does not freeze it. All ~395M encoder parameters are trained together with the head on top. The multilingual sibling swaps in mmBERT-base (22 layers, 256k vocabulary, 322M total) to cover 100+ languages, and a router picks between checkpoints by detecting script and language before the forward pass.

07 · Training: RLCD end to end

Rewarding honest probabilities instead of preferred text

A typical chat LLM learns next-token prediction, is tuned on examples, then often goes through RLHF, which rewards answers people prefer. TypeSafe's argument is that preferred text is the right target for a chat product and the wrong one for software that has to make a decision. That argument is where the name comes from: Reinforcement Learning for Calibrated Decisions.

The reward: a strictly proper scoring rule

Picture a forecaster who honestly believes there is a 70% chance of rain tomorrow. Under a proper scoring rule, saying 90% loses points on average, and so does saying 50%. The only way to maximize expected score is to report the number you actually believe. read from code Laya's main rule is the log score, the log of the probability given to what actually happened. It adds a spherical score, and for ordinal score questions subtracts a ranked probability penalty.

Policy logits z for k options z + noise 1 z + noise 2 z + noise 3 z + noise 4 zero-mean Gaussian on logits Proper scoring rule R = log score + spherical score - RPS (ordinal) vs target: label or teacher distribution R1 R2 R3 R4 Advantage A_i = R_i - mean(R) group baseline (GRPO) Loss REINFORCE + supervised term update both encoder and head; push toward copies that beat the group mean
Figure 6. The RLCD loop as implemented in the repo's fine-tuning notebook. In the typed-decisions run the target is a teacher model's probability distribution. Multi-turn conversations use TD(lambda = 1.0) over prefix slices, per the model card.

read from code The loop: the model scores the options, makes four copies of those scores and adds a little random noise to each. Each noisy copy becomes a slightly different distribution, and each is rewarded against the target. Every copy is compared with the average of the four, and the model is pushed toward the copies that beat the average. GRPO, the RL method common in LLM post-training, uses the same group-mean idea.

The notebook also has a second term: an ordinary supervised loss that pulls probabilities straight toward the target. inference Because the log score rewards the same thing cross-entropy does, the RL part and the supervised part pull in the same direction. In practice this looks less like pure RL and more like supervised calibration training with an exploration term layered on.

08 · Calibration, and the limits

Training on honest probabilities does not, by itself, give honest probabilities

The model card says so plainly. Expected calibration error (ECE) measures the gap between how confident a model is and how often it is right. The shipped checkpoints are over-confident. Fitting one temperature per (question type, option count) on held-out data moved the English checkpoint's mean ECE from 0.466 to 0.081.

The config file shows why that step needs watching. It stored a fitted temperature near 0.1 for choices with 11 or more options, which turns small score gaps into near-certainty (try it in Figure 5). Since v0.3.5 the library clamps temperatures to the range 0.5 to 5.0 at load and warns that those answers are uncalibrated.

typed-decisions benchmark (2,000 decisions)accuracy
random guess0.318
Laya base English checkpoint, zero-shot0.362
always pick the per-question majority class0.461
Jev 1.13.0 (published)0.727
teacher self-agreement ceiling0.735
laya-typed-decisions, fine-tuned on this benchmark's train split0.766

The base checkpoint sits below the majority-class baseline. All of the capability on this benchmark comes from fine-tuning. So the honest framing is that Laya is a fast base you specialise on your own data, not a zero-shot decision engine. The repo ships a notebook that runs the full loop (dataset, RLCD training, temperature fitting, evaluation) on Kaggle's free 2x T4 GPUs in roughly 4 to 5 hours.

Failure modes to test before production

FailureSymptomMitigation
High cardinalityAccuracy falls past ~20 options as labels are trimmedShortlist by embedding, or split coarse/fine
Label leakage in noulEnglish checkpoint follows the false/true labels, not the stateOverride labels (A/B), or use a 2-way choice
Position biasIdentical options get different logits by slotAverage over rotated option_order in one pass
Negation"Do not cancel" routed to cancel_accountHold-out tests on negated phrasings
Wrong language, high confidenceKhmer on English checkpoint: 0.000 accuracy at 0.952 raw confidenceRoute before the forward pass

09 · Why decision models exist

The gap between a random forest and an LLM

Software is full of small decisions over unstructured input: which team gets this ticket, is this answer good enough to publish, is this break a timing difference or a real one. Each is really a fuzzy if statement. Laya's design sits on a lineage that has been converging on one idea: put the labels in the input instead of fixing them in the output layer.

ApproachReads raw textLabels set at call timePasses per callNative probability
Random forest / GBDTNo, hand-built featuresNo1Yes
Fine-tuned BERT classifier (2018)YesNo, tied to output layer1Yes
NLI zero-shot (Yin, Hay, Roth 2019)YesYes1 per labelPer label
GLiNER (2023), labels in encoder inputYesYes1Per span
Laya / decision modelsYesYes1 for all questionsSoftmax over options
LLMYesYes1 + n decode stepsWritten, or token logprob

The gap is in the middle. You want to write the question and options at call time, get the answer in one pass, and get a probability you can calibrate and then threshold. That threshold is what makes a cascade work.

Request ticket, break, doc Decision model ~40 ms, self-hosted calibrated on your data max p >= tau ? yes Act automatically sample and audit what you act on no Escalate: LLM or human pay LLM cost only on hard cases
Figure 7. The cascade. tau is a policy you fit on held-out data at the option counts and dtype you serve, not a property of the model.

Where this lands in fund operations

inference Asset servicing runs on exactly this kind of fuzzy if. A few places the shape fits, each with a short, stable option set:

DecisionTypeOptions
NAV reconciliation break: root causechoicepricing, FX, corporate action, timing, booking error, other
Trade exception: route to deskchoicecustody, TA, fund accounting, middle office
Client query urgencyscoreroutine, same day, blocks NAV strike
Does this email instruct a cash movement?noulP(true), gated hard to human review

None of these benefit from a model that can write prose. All of them benefit from a calibrated number, a 40 ms latency budget inside a batch window, and weights that never leave the bank's network. The work is in the fine-tuning set and the temperature fit, not in the model.

Next time a decision model hands you "pass" with 0.94 next to it, you now have a picture, drawn from the code, of where that number came from. Check it against your own data before your code trusts it.

10 · Sources

References

  1. Laya source, github.com/NandhaKishorM/laya (v0.3.5 read for this article; README at 0.3.21).
  2. Laya model card, huggingface.co/convaiinnovations/laya, Convai Innovations.
  3. TypeSafe AI, Introducing System One Models and Jev (15 Sep 2026).
  4. Warner et al., ModernBERT (2024).
  5. Devlin et al., BERT (2018).
  6. Yin, Hay and Roth, Benchmarking Zero-shot Text Classification (2019).
  7. Zaratiana et al., GLiNER (2023).
  8. Xiong et al., Can LLMs Express Their Uncertainty? (2023).

Benchmark and latency figures are as published in the Laya README and model card; Jev figures there are third-party, not measured by the Laya author. Values labelled "example" in the figures are illustrations.

Independent technical writing. Views are my own.

Share this article