Adam Innes · Blog

Laya: A Free, Local Jev Alternative, and How to Connect It to Claude

· 7 min · ai, llm, ai agents, python, open source

The thing I said I would check about Jev was whether the confidence numbers hold up on data that is not the vendor’s. That measurement is still the one that matters, and it got more interesting now that there is an open-weights model making the same bet. Laya, from Convai Innovations, is Apache 2.0, runs on a laptop, costs nothing per call, and answers the same three question types. It also speaks Jev’s wire protocol, which means the integration code from my earlier post runs against it with a changed base URL and nothing else.

That is a genuinely useful thing to exist. It is not, as far as I can tell from the published numbers, a drop-in replacement for Jev’s judgment. Those are two separate claims and the gap between them is the whole story.

What it actually is

Laya is not a small language model with a JSON mode. It is an encoder with a decision head, non-autoregressive, producing a probability distribution over your options in a single forward pass. There is no sampling loop, which is why the latency is what it is.

Three checkpoints ship in one Hugging Face repo, and only the one you ask for is downloaded. The English checkpoint is ModernBERT-large at 421M parameters with a 512-token context. The multilingual one is mmBERT-base at 322M, 1024 tokens nominally and up to 8,192, covering 100+ languages. A third, laya-typed-decisions, is ModernBERT-large again at 1024 tokens, tuned for the typed-question workload specifically. A Router in front of them detects script and Latin stopword distribution to pick a checkpoint, which the docs measure at 0.09 ms of overhead for ordinary English and under 2 percent of total latency in general.

The primitives are the same three Jev defines. choice picks a key from a dictionary you supply. score places the input on an ordinal scale. noul returns P(true) for a statement. The Python is about as plain as it gets:

from laya import Router

router = Router(preload=True)

result = router.predict(
    {"body": "Billed twice this month. Cancelling if this isn't fixed today."},
    {
        "queue": {
            "type": "choice",
            "instructions": "Which engineering queue owns this ticket?",
            "criteria": {
                "billing": "refunds, SLA credits, invoice disputes",
                "infrastructure": "outages, network, database failures",
                "security": "breaches, vulnerability reports",
            },
        },
        "churn_risk": {
            "type": "noul",
            "instructions": "The customer threatens to cancel",
        },
    },
)

Convai measures 32.8 ms for a single question on a T4, and 72.3 ms for ten questions batched, which works out to roughly 7.2 ms each. Their comparison puts Jev at 236 to 276 ms over the network for the same work, so the speed claim is around six to seven times on a single question and far more when you batch. Treat that like any vendor benchmark, including TypeSafe’s: the direction is obviously right, because a 421M-parameter local forward pass beats a hosted call across the internet, and the exact multiple depends on where your code is sitting.

The fanning-out property survives, which is the architectural point. Ten questions against one state cost roughly double one question, not ten times, so the habit of asking speculative questions and letting your code decide afterwards what mattered still works.

Connecting it to Claude

There are two things people mean by this and they are worth keeping apart, same as with Jev.

The first is giving Claude Code a decision tool it can call during a session. Laya ships an MCP server for exactly that:

pip install "laya[mcp]"
claude mcp add laya laya-mcp-server --env LAYA_DEVICE=cpu

It talks stdio, so Claude Code owns the process. The tools it exposes are laya_predict, laya_predict_batch, laya_route, laya_route_batch, laya_shortlist, laya_preset, laya_decide and laya_status. LAYA_DEVICE picks cpu or cuda, LAYA_MODELS controls which checkpoints preload, and LAYA_PRELOAD=0 keeps startup fast if you would rather pay the load cost on first call.

Where that earns its place is bulk classification during a session. Asking Claude to label two thousand log lines or triage a backlog of tickets means two thousand decisions at frontier prices and frontier latency; handing that to a local tool it calls in a loop is a different shape of work entirely. If you would rather run one model server for several sessions, laya-serve plus LAYA_BASE_URL on the MCP process points them all at it, though laya_shortlist and the prediction hooks do not work in remote mode.

The second meaning is Laya inside the thing you are building, in front of Claude rather than behind it. This is where the Jev compatibility pays off:

pip install "laya[serve]"
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve

That exposes POST /v1/systemone on the Jev wire protocol. A client written against typesafe-sdk can point its base URL at it. Responses are schema-compatible with extra fields added, and LAYA_JEV_STRICT=1 projects them back down to the strict contract if the extras break your parser. There is a batch endpoint taking up to 64 states against one question set, a 50,000-character state limit, 100 options per choice, and LAYA_API_KEY for bearer auth if it is reachable from anywhere you do not control. Notably there is no OpenAI-compatible endpoint, which is the correct decision — pretending this is a chat completion would misrepresent what it does.

So the guardrail pattern — screen every message into and out of a Claude app for injection attempts, route it, check whether a retrieved passage is relevant, decide whether to escalate — becomes something you run on your own hardware with no per-call cost and no data leaving your network. For anyone who could not send user content to a third party at all, that is the difference between having this architecture and not having it.

The part that is not free

Here is what the published numbers say, and credit to Convai for publishing them, because they are not flattering.

Zero-shot on their own typed-decisions benchmark, the base English checkpoint scores 0.362 and the multilingual one 0.342, against a majority-class baseline of 0.461. Both are worse than always guessing the most common answer. The headline 0.766 comes from laya-typed-decisions after fine-tuning on that benchmark’s training split. Narrower, well-shaped tasks do fine out of the box — 0.953 on AG News with four labels, 0.993 reported on email spam — but 0.425 on Banking77’s 77 labels, 0.600 on six-way emotion, and 0.530 accuracy with a 0.400 macro-F1 on held-out toxicity. The documented advice is to keep choice under about 20 options, and the Banking77 number is what happens when you ignore it.

The calibration situation matters more, because calibration is the entire reason to prefer a System One model over an if-statement. Both checkpoints ship over-confident. Expected calibration error is 0.466 for the English checkpoint as downloaded, dropping to 0.081 once you fit a single temperature per question type on held-out data; multilingual goes 0.314 to 0.106. The docs call refitting temperature per question type and option count the single highest-value fix available, and on those numbers they are right — the probabilities are close to decorative until you do it.

There are sharper edges worth knowing before you build on them. Option order changes choice answers, with an option_order parameter to average the bias out. The English checkpoint stays confident on text it handles badly outside English. Accuracy on the multilingual encoder gets unreliable past about 4,000 tokens even though it accepts 8,192. And the API has two confidence fields that mean different things — answer_confidence is probability mass on the reported answer, while confidence is normalized entropy for choice and score but max(p_yes, p_no) for noul — so they must never share a threshold. I would read that twice before writing the gating logic.

The 512-token English context is the constraint that will bite first in practice. Jev budgets 64k across state and questions with 32k for state plus the longest question. Laya’s English checkpoint gives you 512 tokens total, 192 of which is the head budget for your instructions and criteria. A support email fits. A support thread does not. There is a predict_long() that scans overlapping windows and aggregates, and the typed-decisions and multilingual checkpoints give you 1024, but filtering the state down before you send it stops being advice and becomes a requirement.

What I would do with it

Use it for the labeling and eval work first, where being wrong is cheap and you learn what you need to know anyway. Connect the MCP server, point Claude at a pile of your own data, and get a few thousand decisions out of it next to labels you already trust. That exercise produces the two things you need: whether the base checkpoint is good enough on your particular task, and the held-out set you fit temperature on. Both are prerequisites for the architecture, and you were supposed to run the same measurement before trusting Jev.

If the answer comes back good, the Jev-compatible endpoint means you can run the guardrail and routing layer locally for nothing, which is a real position to be in. If it comes back bad, you have a fine-tuning run in front of you — a Kaggle notebook on 2x T4s in about four hours, by their reckoning, plus a calibration fit — and that is where “free” turns into an afternoon and some judgment about whether your task deserves it.

What I keep coming back to is that the two models are not really competing on the same axis. Jev sells calibrated judgment you rent by the call. Laya gives you the architecture, the wire protocol and the speed, and leaves the calibration as homework. Whether that is a bargain depends entirely on whether you were ever going to do the homework, and the uncomfortable thing is that if you were not, you should not be gating automatic decisions on confidence numbers from either one.

← all posts