RECORDED
▴JEVBENCH PUBLIC · ZILS 147/231 · KEV 147/231 · TIED▴CALIBRATION ECE · ZILS 0.114 VS 0.044 · REGRESSED▴P50 LATENCY · 42 MS · LOCALHOST▴TESTNET 579 · WEIGHTS VERIFIED · BLOCK 8081345▴MINER UID 1 · 183/224 HELD-OUT · U16 65535▴MINER UID 2 · 180/224 HELD-OUT · U16 64598▴MINER UID 3 · 180/224 HELD-OUT · U16 62125▴BASE · QWEN3.5-0.8B-BASE▴CUSTOMER DELIVERY · PLANNED

Specialized decision models · Early access

Your data.
Your decision model.

Train, evaluate, and deploy models for your business’s decisions.

Via the Zils Discord community. Customer delivery is planned.

ZILS · DECISION INSTRUMENTILLUSTRATIVE

CHOOSE THE RIGHT MODEL

context › Extract the total and due date from this invoice as JSON.

probabilities out · no text generated→ small model

CONTEXT IN · PROBABILITIES OUT · EXAMPLES ARE ILLUSTRATIVE, NOT MODEL OUTPUT

A DECISION PRIMITIVE FOR YOUR AI STACK

Small decisions.
A big part of your AI bill.

Model selection. Tool choice. Escalation. Zils is building a trainable decision primitive for these repeated judgments: context in, probabilities out. The goal is to replace full LLM calls where a specialized model can meet your quality bar.

Potential savings depend on call volume, serving costs, and quality. Customer cost savings have not yet been measured.

Training & evaluation implementedCustomer delivery plannedExplore the research ↗

Inside the AI stack

Put a decision model
where the calls add up.

Train on authorized labeled examples from your workflows. Evaluate against your current stack, including quality, latency, and total cost.

01↗

Choose the right model

Evaluate whether a request needs a larger model or can stay on a smaller one. Reserve expensive calls for the work that needs them.

Request → model
02⌘

Select the next tool

Turn context into a choice among the tools your agent can use, without generating a full text response for every selection.

Context → tool
03✳

Know when to escalate

Use labeled outcomes to evaluate whether an agent should continue, retry, or ask for review before the next step.

Agent state → next action

Proposed customer workflow

Good examples in.
Evidence behind every decision.

Customer-specific jobs and delivery are planned. Today’s research system already exercises training, checkpoint verification, and evaluation.

  1. 01

    Provide labeled examples

    One recurring decision, its possible outcomes, and data you’re authorized to train on, with a separate set held out.

  2. 02

    Train candidates

    Adapt a starting model to the task. Recipes and candidate versions stay traceable.

  3. 03

    Evaluate against a baseline

    Accuracy, probability quality, consequential mistakes, latency, and total cost vs. your current approach.

  4. 04

    Deploy a qualifying model

    Acceptance criteria are agreed before training. Only a candidate that meets them moves on.

IT ALREADY RAN ONCE

Train. Evaluate. Verify. On the record.

This terminal replays the recorded Bittensor testnet round, line for line, from its public JSON. Every number on it is in the file.

testnet-round-001.json ↗
zils · recorded round replay

Three miner processes and one validator were operated by one operator on one host. Fresh checkpoints; reused synthetic development benchmark, not an independent generalization test.

Planned deliverables

More than a model.
A basis for confidence.

Know what was trained, how it was tested, and what it would take to put it to work.

01 / PLANNED

A versioned model

A selected checkpoint with its identity, training configuration, and dependencies recorded.

02 / PLANNED

Evaluation evidence

Baseline comparison, dataset and rubric versions, measured tradeoffs, known limitations.

03 / PLANNED

A deployment plan

Self-hosting or a managed endpoint, assessed against your app, hardware, and data requirements.

Data, considered.

Current research worker bundles include copies of training data; confidential distributed training is not supported. Data handling is part of scoping. No Zils model release is publicly available yet.

JEVBENCH · PUBLIC ITEMS

147/231

zils = published kev · +6 / −6 items

Open about the evidence

Accuracy tied. Confidence quality regressed.

We publish the result as it came out, including where it got worse. This does not establish a general improvement.

Your next model starts
with a better question.

Where do repeated decisions add up in your AI stack? Tell us about the calls, your labeled examples, and the quality bar a smaller model would need to meet.

Discuss your use case

Opens Discord. Start with a task description; don’t share private data.