Laya 0.3.20 API
ActiveLaya 0.3.20 is a multilingual System 1 decision engine for typed classification, scoring and routing over text or JSON in one fast pass.
Laya 0.3.20 API Background
Overview
Laya 0.3.20 is a multilingual, non-autoregressive System 1 decision engine from Convai Innovations designed for typed decisions rather than text generation. The Laya 0.3.20 API evaluates choice, score, and noul questions over text or JSON-like state in a single forward pass, enabling fast classification, triage, moderation, routing, and schema-driven decisions. It supports 100+ languages through a router that selects among English, multilingual, and typed-decisions checkpoints, with measured single-question latency around 33 ms on a T4 GPU and batched throughput that improves significantly as question count grows.
Development History
Laya evolved as a self-hostable decision model family focused on structured outputs, calibrated probabilities, and multilingual routing. Version 0.3.20 represents a mature release with three production checkpoints, Router-based automatic checkpoint selection, batch inference, long-document handling up to 8,192 tokens on the multilingual encoder, and integrations for server, MCP, LangChain, LlamaIndex, and CrewAI. The research context shows the platform expanding from zero-shot typed decisions into fine-tuned domain specialization, where the typed-decisions checkpoint reached 0.766 accuracy on a 2,000-decision benchmark, materially improving over the base checkpoints.
Key Innovations
- Non-autoregressive typed decision inference that answers multiple structured questions in one forward pass without text generation or parsing
- Router-based checkpoint selection that detects script and language before inference, sending requests to English or multilingual models for better accuracy and reliability
- Training and confidence design centered on reinforcement learning against strictly proper scoring rules, enabling probability outputs that are more useful for confidence gating and abstention policies
Laya 0.3.20 API Technical Specifications
Architecture
Laya 0.3.20 uses encoder-based architectures optimized for structured decision tasks. The English and typed-decisions checkpoints are based on ModernBERT-large with 421M parameters, while the multilingual checkpoint uses mmBERT-base with 322M parameters and supports 100+ languages. The Laya 0.3.20 API exposes these through a Router that automatically chooses the appropriate checkpoint per request. It supports single-call prediction, batch inference, schema-driven decisions, and long-document scanning, with multilingual context up to 1,024 tokens by default and up to 8,192 tokens when configured for long inputs.
Parameters
The model family includes three checkpoints: laya and laya-typed-decisions at 421 million parameters each, and laya-multilingual at 322 million parameters. Context windows differ by checkpoint: the English model uses 512 tokens, the typed-decisions checkpoint uses 1,024 tokens, and the multilingual encoder supports 1,024 tokens by default with an extended path up to 8,192 tokens. In the Laya 0.3.20 API, per-request token budgets can be adjusted using max_len and head_max_len to balance document length against option capacity for high-cardinality decision tasks.
Capabilities
- Typed decisions across choice, ordinal score, and noul probability outputs with calibrated confidence metadata
- Automatic multilingual routing across 100+ languages, including non-Latin scripts, with single forward-pass inference
- Batch prediction, schema-driven structured outputs, long-document scanning, and self-hosted HTTP or MCP deployment through the Laya 0.3.20 API
- Prediction hooks for redaction, caching, logging, routing overrides, and confidence-based escalation in production workflows
Limitations
- Zero-shot performance is uneven for specialized typed-decision workflows; the strongest results come from fine-tuning on domain-specific data rather than relying on the base checkpoints alone
- High-cardinality choice tasks can degrade sharply because option texts share a fixed token budget, making 50+ label problems difficult without increasing head_max_len or using shortlist strategies
- The multilingual checkpoint is weaker on English than the English checkpoint, while the English checkpoint can fail badly outside English, so routing is essential rather than optional
- Some documented edge cases remain, including score bias in multilingual scoring, noul label sensitivity, and weak handling of certain negation-heavy choice formulations
Laya 0.3.20 API Performance
Strengths
- Very fast structured inference, with measured latency around 32.8 ms for one multilingual question and 72.3 ms for 10 batched questions on a T4 GPU
- Strong multilingual routing gains over using a single checkpoint, with the routed setup matching the best model on English and non-English benchmark slices
- Fine-tuned typed-decisions performance is highly competitive, reaching 0.766 accuracy on a 2,000-decision benchmark and outperforming the reported 0.727 result for Jev on that benchmark
- Long-document support is practical for operational use, with the multilingual model reading up to 8,192 tokens and answering 16 to 18 of 20 requests correctly up to about 4,000 preceding tokens in reported tests
Real-world Effectiveness
In practice, the Laya 0.3.20 API is most effective for high-volume, well-scoped business decisions such as support triage, moderation, guardrails, intent routing, and schema extraction where latency and deterministic structure matter more than free-form generation. Reported results show meaningful throughput gains from batching, strong routing behavior across 51 languages, and substantial quality improvements after domain fine-tuning. It is especially useful where teams need typed outputs with confidence scores, but it should be configured carefully for large option sets, calibrated on local data, and validated for score or noul edge cases before fully automated deployment.
Laya 0.3.20 API When to Use
Scenarios
- You have a high-volume customer support or operations workflow where every request must be classified into departments, urgency levels, refund status, or churn risk. The Laya 0.3.20 API fits because it answers multiple typed questions in one forward pass, supports multilingual input, and returns confidence-rich structured outputs without generation overhead. This gives operations teams faster routing, more predictable automation, and lower integration complexity for ticketing, inbox triage, and service desk workflows.
- You have a multilingual product receiving user messages, complaints, or moderation events across many scripts and languages. The Laya 0.3.20 API is ideal because its router detects whether the English or multilingual checkpoint should handle each request before inference, avoiding the severe confidence failures that can occur when English-only models process non-English text. This improves consistency across global traffic and makes automated triage, safety screening, and intent detection more reliable in mixed-language production systems.
- You have a business process that needs structured decisions rather than generated prose, such as mapping emails or JSON documents into schema fields, evaluating policy flags, or deciding whether a workflow requires human review. The Laya 0.3.20 API is well suited because it supports schema-driven decisions, batch inference, and deployment as a self-hosted service. Teams can turn unstructured inputs into typed outputs quickly, reduce parsing errors, and build deterministic downstream automation around stable response formats.
Best Practices
- Use Router-based inference by default, calibrate confidence on held-out data, and apply abstention or escalation thresholds rather than trusting raw probabilities out of the box
- For large label spaces or long documents, tune max_len and head_max_len carefully, consider shortlist strategies for 20+ options, and validate quality on your own workload before production rollout
- Fine-tune on domain-specific decisions whenever accuracy matters, because the biggest gains reported for the Laya 0.3.20 API come from specialization rather than zero-shot use