Completed experiment · 7 October 2026 UTC
Mate-in-one chess study.
A fresh chess adapter improved Zils’ ability to select an immediate checkmate. On the same 512 held-out positions, it outperformed both shared Zils and TypeSafe Jev.
- Chess-trained Zils
- 50.78%
- 260 / 512 correct · Fresh adapter · 2,048 training positions
- TypeSafe Jev 1.13.0
- 43.55%
- 223 / 512 correct · Regular Jev · same test inputs
- Shared Zils
- 42.97%
- 220 / 512 correct · No chess-specific adapter
All legal checking moves were offered. Every move that immediately checkmates received credit.
Learn from different games.
The adapter learned from public Lichess puzzles. All three models saw the same board, instructions, candidate moves, and order at test time.
- Training positions
- 2,048
- Calibration positions · reserved, unused
- 512
- Held-out test positions
- 512
One position per source game. No repeated source games, exact positions, or color-swapped vertical mirrors across the full sample. Fifteen test positions have multiple winning moves, all accepted.
An improvement with measured uncertainty.
Compared with shared Zils
+7.81 percentage points
Paired 95% interval: +3.32 to +12.11 points.
Compared with TypeSafe Jev
+7.23 percentage points
Paired 95% interval: +2.15 to +12.30 points.
Both intervals stay above zero. We resampled the 512 test games 2,000 times, pairing each model’s result on the same position. These are individual intervals without adjustment for multiple comparisons. Only one adapter seed and one fixed recipe were tested.
A narrow task with an exact answer.
The task is to choose an immediate checkmate from every legal checking move. Test positions had 2–10 choices; the protocol allows 2–16. This measures restricted mate-in-one selection. Full-game strength, unrestricted move generation, and multi-move planning were not evaluated.
- Uniform random choice · expected
- 33.93%
- Simple capture heuristic
- 36.52%
- Deterministic chess-rules oracle
- 100.00%
Chess software already solves every position in this task. The experiment asks whether specialization improves Zils on a new domain. No model was deployed or promoted.
All scores, training method, and reliability
| Model | Correct | Accuracy | Probability on mating moves | NLL | Confident mistakes |
|---|---|---|---|---|---|
| Chess-trained Zils | 260 / 512 | 50.78% | 43.25% | 1.059 | 2 / 18 confident choices |
| TypeSafe Jev 1.13.0 | 223 / 512 | 43.55% | 38.70% | 1.294 | 3 / 11 confident choices |
| Shared Zils | 220 / 512 | 42.97% | 36.08% | 1.164 | 0 / 0 confident choices |
Inputs and answers. The opponent’s setup move was applied first. Models received the resulting board, side to move, and candidate moves with piece/from/to descriptions. Solutions, checkmate notation, ratings, and puzzle or game identifiers stayed out of model requests. A rules engine independently verified every mating alternative and executed all 1,536 selected moves for the final audit.
Data selection. The first 200,000 records of the pinned Lichess archive supplied the pool; this is not a uniform sample of the full database. Positions with fewer than two or more than sixteen checking moves, or where every candidate mated, were excluded. Source-game separation and exact/mirror deduplication still allow related tactical patterns and possible overlap with public pretraining data.
Model foundation. Shared Zils uses the published JevK5 4B weights; chess-trained Zils adds the adapter trained for this study.
Training. One fresh rank-16 adapter, one epoch, seed 557, and 512 optimizer steps. The final saved checkpoint was used without test-directed tuning. Training and saving took 22.69 minutes on an RTX 4090, excluding loading and input preparation. Peak training tensor allocation was 8.52 GiB. Both smoke and full adapters passed independent fresh-process reload checks. Native temperatures were retained; calibration data was unused.
API reliability. Jev returned valid answers on the first attempt for 506/512 positions. Six rejected responses were retained and recovered within the three-attempt limit: five probability-sum failures and one inconsistent choice. The reported scores use validated responses after retries. All three model runs completed; no failed or missing positions were dropped.
Additional measures. First-listed move accuracy was 34.77%. A confident choice assigns at least 90% probability to the selected move. The table’s mating probability is the total probability assigned to all correct moves. NLL is the negative log of that total, with a floor for zero probability.