Laya
Open-source System 1 decision engine - typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request. A self-hosted, Jev-compatible alternative.
What is this?
Laya is a multilingual, non-autoregressive System 1 decision engine. It answers typed questions —
choice,score, and yes/no (noul) — over any text in a single forward pass, in 100+ languages, at ~33 ms per question on a T4. A router picks the right checkpoint per request.
Open-source (Apache-2.0), self-hostable, and wire-compatible with TypeSafe's hosted Jev API — making it a drop-in open alternative to Jev for structured decisions over text.
What problem does it solve?
LLMs are slow, expensive, and unstructured for decisions that need a typed answer:
LLM: Text → prompt → generate tokens → parse → hope it's valid
(200ms+, costs tokens, can hallucinate, needs parsing)
Laya: Text → one forward pass → typed answer
(33ms, $0 self-hosted, calibrated probabilities, nothing to parse)
Jev (TypeSafe) does this well but is proprietary and metered at $0.042/1M tokens. Laya is the open-source alternative: same wire protocol, self-hosted, and measured 6-7x faster.
How does it work?
Your app
│ state (text, email, ticket, JSON)
│ questions: {choice, score, noul}
▼
┌─────────────────────────────┐
│ Router │
│ detects script & language │
└──────────┬──────────────────┘
│
┌─────┼─────────────┐
▼ ▼ ▼
laya laya- laya-typed-
(English) multilingual decisions
└─────┬─────────────┘
▼
single forward pass
(ModernBERT encoder)
▼
typed answers + calibrated
answer_confidence
- You pass a
state(any text) and a dict of typed questions. - The Router detects the script/language and routes to the right checkpoint.
- One forward pass over a ModernBERT encoder produces all answers at once — no text generation, nothing to parse, nothing to hallucinate.
- Every answer carries
answer_confidence— the calibrated probability of the reported answer — so you can gate decisions on one number.
Repository Structure
NandhaKishorM/laya/
├── laya/ # Python package
│ ├── integrations/ # LangChain, LlamaIndex, CrewAI
│ └── mcp/ # MCP server
├── laya-ts/ # TypeScript / Node / browser version
├── research/ # Training scripts, RLCD recipes
├── benchmarks/ # Benchmark suites + results
├── docker/ # Container images
├── nix/ # Nix packaging
├── scripts/ # CLI + utilities
├── tests/ # Test suite
└── notebooks/ # Examples
Key Features
- Typed decisions —
choice(pick from options),score(rate on a scale),noul(yes/no probability) — all in one call - 100+ languages — router auto-selects English or multilingual checkpoint by script
- Single forward pass — 33 ms/question on a T4, 7.2 ms/question batched; 6-7x faster than Jev
- Calibrated — trained with RL against strictly proper scoring rules (RLCD);
answer_confidenceis a real probability - Jev-compatible —
laya.serveexposesPOST /v1/systemonewith a schema-identical payload; repoint your Jev client's baseUrl - Self-hosted — $0 per token vs Jev's $0.042/1M; your data never leaves your machine
- Fine-tunable — specialize the base checkpoint on your own labeled decisions (beats Jev by 3.9 points when fine-tuned)
- Integrations — MCP server, LangChain, LlamaIndex, CrewAI, ONNX Runtime, CLI, HTTP server, TypeScript SDK
What You'll Learn
- How non-autoregressive "System 1" decision models work (vs autoregressive LLMs)
- How to make structured, calibrated decisions over text without generation or parsing
- How to route multilingual inputs to the right checkpoint
- How to self-host a Jev-compatible decision API
- How to fine-tune a decision engine on your own labeled data
- How to gate agent actions on calibrated confidence
See it in action
Install and run the quickstart:
pip install laya
from laya import Router
router = Router()
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["not urgent", "soon", "blocking"]},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"]) # billing
print(result["answers"]["churn_risk"]["noul"]) # probability the answer is yes
print(result["routing"]["model"]) # english
Try the HF Space demo, the Colab notebook, or the docs.
Benchmarks vs Jev 1.13.0 (third-party published): Laya wins accuracy on typed-decisions (0.766 vs 0.727), AG News (0.950 vs 0.910), DAIR Emotion (0.595 vs 0.480), calibration ECE (0.081 vs 0.246), and p50 latency (32.8 ms vs 236-276 ms). Jev leads on >20-option label spaces and soft distribution matching.
Want to learn more about this project?
Have questions or want to understand this technology more deeply? Send me a message.
Related Projects
Strix
Open-source AI penetration testing tool - autonomous AI hackers that run your code dynamically, find vulnerabilities, and validate them through actual proofs-of-concept.
Cua
Give AI agents computers they can use — open-source desktop automation drivers, isolated cloud desktops, local macOS VMs, and benchmarks for computer-use agents.
OpenWorker
Open-source AI coworker that lives on your desktop and delivers finished work — security reviews, documents, Slack replies — with your own model and 25+ integrations.
Discussion
Loading comments...