Skip to contentZils
Back to model research

Completed experiment · 7 October 2026 UTC

Mate-in-one chess study.

A fresh chess adapter improved Zils’ ability to select an immediate checkmate. On the same 512 held-out positions, it outperformed both shared Zils and TypeSafe Jev.

Checkmate accuracySame 512 test positions · Higher is better
Chess-trained Zils
50.78%
260 / 512 correct · Fresh adapter · 2,048 training positions
TypeSafe Jev 1.13.0
43.55%
223 / 512 correct · Regular Jev · same test inputs
Shared Zils
42.97%
220 / 512 correct · No chess-specific adapter

All legal checking moves were offered. Every move that immediately checkmates received credit.

Learn from different games.

The adapter learned from public Lichess puzzles. All three models saw the same board, instructions, candidate moves, and order at test time.

Training positions
2,048
Calibration positions · reserved, unused
512
Held-out test positions
512

One position per source game. No repeated source games, exact positions, or color-swapped vertical mirrors across the full sample. Fifteen test positions have multiple winning moves, all accepted.

An improvement with measured uncertainty.

Compared with shared Zils

+7.81 percentage points

Paired 95% interval: +3.32 to +12.11 points.

Compared with TypeSafe Jev

+7.23 percentage points

Paired 95% interval: +2.15 to +12.30 points.

Both intervals stay above zero. We resampled the 512 test games 2,000 times, pairing each model’s result on the same position. These are individual intervals without adjustment for multiple comparisons. Only one adapter seed and one fixed recipe were tested.

A narrow task with an exact answer.

The task is to choose an immediate checkmate from every legal checking move. Test positions had 2–10 choices; the protocol allows 2–16. This measures restricted mate-in-one selection. Full-game strength, unrestricted move generation, and multi-move planning were not evaluated.

Uniform random choice · expected
33.93%
Simple capture heuristic
36.52%
Deterministic chess-rules oracle
100.00%

Chess software already solves every position in this task. The experiment asks whether specialization improves Zils on a new domain. No model was deployed or promoted.

All scores, training method, and reliability
All 512 test positions per model. Any immediate checkmate receives credit. Lower negative log likelihood (NLL) is better.
ModelCorrectAccuracyProbability on mating movesNLLConfident mistakes
Chess-trained Zils260 / 51250.78%43.25%1.0592 / 18 confident choices
TypeSafe Jev 1.13.0223 / 51243.55%38.70%1.2943 / 11 confident choices
Shared Zils220 / 51242.97%36.08%1.1640 / 0 confident choices

Inputs and answers. The opponent’s setup move was applied first. Models received the resulting board, side to move, and candidate moves with piece/from/to descriptions. Solutions, checkmate notation, ratings, and puzzle or game identifiers stayed out of model requests. A rules engine independently verified every mating alternative and executed all 1,536 selected moves for the final audit.

Data selection. The first 200,000 records of the pinned Lichess archive supplied the pool; this is not a uniform sample of the full database. Positions with fewer than two or more than sixteen checking moves, or where every candidate mated, were excluded. Source-game separation and exact/mirror deduplication still allow related tactical patterns and possible overlap with public pretraining data.

Model foundation. Shared Zils uses the published JevK5 4B weights; chess-trained Zils adds the adapter trained for this study.

Training. One fresh rank-16 adapter, one epoch, seed 557, and 512 optimizer steps. The final saved checkpoint was used without test-directed tuning. Training and saving took 22.69 minutes on an RTX 4090, excluding loading and input preparation. Peak training tensor allocation was 8.52 GiB. Both smoke and full adapters passed independent fresh-process reload checks. Native temperatures were retained; calibration data was unused.

API reliability. Jev returned valid answers on the first attempt for 506/512 positions. Six rejected responses were retained and recovered within the three-attempt limit: five probability-sum failures and one inconsistent choice. The reported scores use validated responses after retries. All three model runs completed; no failed or missing positions were dropped.

Additional measures. First-listed move accuracy was 34.77%. A confident choice assigns at least 90% probability to the selected move. The table’s mating probability is the total probability assigned to all correct moves. NLL is the negative log of that total, with a floor for zero probability.