Jein
Sign in

Models and language routing

Jein runs two pinned checkpoints. /v1/models lists the public aliases; the model field in a System One response names the checkpoint that actually ran.

Request model Selection Response model
jev-latest, jev-preview, jev-1.13.0, laya-latest Detect language in the state when AUTO_ROUTE_MULTILINGUAL=true (the default) laya-0.3.20 for English; laya-multilingual-0.3.20 for detected non-English
laya-0.3.20 English checkpoint, including non-English states laya-0.3.20
laya-multilingual, laya-multilingual-0.3.20 Multilingual checkpoint, including English states laya-multilingual-0.3.20

The x-jein-routed-language response header is an ISO 639-1 code such as en, de, or fr; und means the text was too short or ambiguous to identify. Routing reads string values of state, not JSON field names or the question instructions. It uses a pinned, in-process, CPU-only classifier and makes no network request. Ambiguous text stays on English; choose laya-multilingual explicitly if you already know the state is non-English. Set AUTO_ROUTE_MULTILINGUAL=false to keep aliases on English.

Language and quality

The multilingual model card reports evaluations on about 50 languages; its “100+ languages” statement is unverified. The local quick check covers German, French and Spanish only: 11/12 correct versus 9/12 for the English checkpoint. On English, the multilingual checkpoint scored 8/12 versus 10/12. The larger, mostly English local corpus scored 59/78 (76%) on accepted short-state answers, below its provisional 80% gate. These samples are directional, and multilingual confidence is uncalibrated. Review decision thresholds on your own traffic before changing them.

Per-model limits

Limit English Multilingual
Laya sequence window per question 512 tokens 2048 tokens
Instruction and option head budget 192 tokens 256 tokens
Typical state room 437–478 tokens 1975–2014 tokens
Default processed-token cap per request 3,500 6,000
Questions per request 32 32

The worker rejects input that would be truncated. It bills the same Jev-style input-token count for either model; more questions increase processed compute tokens. Pricing is currently the same for both models. Both checkpoints have a 9-second API deadline, with 500 ms reserved for response and metering. The multilingual cap and 1.1 ms/token admission seed come from the shared-host CPU quick check; they need a fresh latency check on the deployment CPU.