Completed experiment · · NVIDIA B300
Practical models for miners.
How much model does a support decision need? We tested task-specific adapters across 2B, 4B and 9B models. H2O 4B offered lower latency; JevK5 2B offered lower memory use. Neither established an accuracy advantage over our strongest JevK5 4B adapter.
The starting point: tuning an open model can make it competitive with official TypeSafe Jev. Our earlier direct support comparison measured 79.2% for JevK5 4B with an adapter versus 70.6% for hosted Jev 1.13.0 on 500 conversations. This follow-up explores the models and hardware costs behind that approach; it does not rerun the hosted Jev comparison.
B300 was temporary research hardware. These results guide miner experiments; they do not change the shared model or establish production performance. JevK5 here means locally run open weights, separate from the hosted TypeSafe Jev comparison.
H2O vs JevK5 4B adapter
2.99×
Lower median latency: 59.7 ms vs 178.4 ms, on B300.
JevK5 2B vs 4B adapter
54% less
Peak CUDA allocation: 3.76 GiB vs 8.24 GiB. This is not total card capacity.
Selected 9B vs 4B adapter
No clear gain
80.8% vs 82.4% accuracy on the original test cohort. The paired interval includes zero.
Miner-sized follow-ups
Speed, memory, and accuracy.
Each cohort contains 500 ABCD conversations, one decision per conversation, with all 30 actions and full input context. Both cohorts were reused for the H2O and 2B follow-ups: earlier results were already known. Their checkpoints were selected using development data only.
| Model / training | Test accuracy | Historical accuracy | Test Brier ↓ | Test median | CUDA peak |
|---|---|---|---|---|---|
| H2O 4B + adapter4,096 examples · epoch 1 | 83.0%415 / 500 | 83.6%418 / 500 | 0.2883 | 59.7 ms | 8.82 GiB |
| JevK5 4B + adapter4,096 examples · epoch 2 | 82.4%412 / 500 | 84.8%424 / 500 | 0.3115 | 178.4 ms | 8.24 GiB |
| JevK5 2B + adapter4,096 examples · epoch 2 | 80.8%404 / 500 | 84.4%422 / 500 | 0.3370 | 135.9 ms | 3.76 GiB |
| Earlier H2O 4B adapterReused 1,024-example adapter | 79.8%399 / 500 | 80.0%400 / 500 | 0.3501 | 59.4 ms | 8.82 GiB |
| H2O 4B stockNo ABCD adapter | 54.0%270 / 500 | 56.2%281 / 500 | 0.6574 | 43.8 ms | 8.77 GiB |
| JevK5 4B stockNo ABCD adapter | 52.6%263 / 500 | 57.2%286 / 500 | 0.6545 | 129.7 ms | 8.19 GiB |
| JevK5 2B stockNo ABCD adapter | 22.4%112 / 500 | 24.4%122 / 500 | 0.8696 | 99.6 ms | 3.71 GiB |
Brier measures probability error; lower is better. CUDA peaks come from isolated final inference processes. H2O resets its peak after loading and warmup; Jev includes loading allocations. These are different measurement scopes and exclude some driver/process overhead.
H2O: a speed candidate
Against the selected 4B adapter, H2O gained 0.6 percentage points on the test cohort (paired 95% interval −1.2 to +2.4), and lost 1.2 on the historical cohort (−3.6 to +1.4). The speed result is useful; the accuracy differences remain inconclusive.
2B: a memory candidate
The 2B adapter trailed 4B by 1.6 points on the test cohort (−4.4 to +1.2), and 0.4 on the historical cohort (−3.2 to +2.2). Lower allocation makes it worth testing on smaller GPUs. This does not prove equivalent accuracy or that a 4 GB card is sufficient.
Original 4B / 9B matrix
More parameters did not settle it.
The 500-case test cohort was fresh in identifiable local records when this original matrix ran. It became retrospective when we chose the later H2O and 2B trials. The historical cohort was already reused. The selected 9B and 4B adapters were chosen on development accuracy, with Brier as a tie-breaker.
| Model / training | Test accuracy | Historical accuracy |
|---|---|---|
| JevK5 4B stock | 52.6% | 57.2% |
| 4B full weights · 1,024 × 1 epoch | 65.8% | 65.4% |
| 4B adapter · 1,024 × 1 epoch | 75.0% | 79.6% |
| 4B adapter · 4,096 × 2 epochs | 82.4% | 84.8% |
| JevK5 9B stock | 55.2% | 57.8% |
| 9B adapter · 1,024 × 1 epoch | 76.4% | 80.2% |
| 9B adapter · 4,096 × 1 epoch | 80.8% | 83.6% |
Selected 9B minus selected 4B was −1.6 points (95% interval −4.2 to +0.8). Full-weight 4B training trailed the matched 1,024-example LoRA adapter by 9.2 points (−13.6 to −5.0). These fixed recipes favor spending further experiments on adapters and smaller models; they do not establish that 9B or full-weight training cannot work.
The 27B teacher trial was canceled before any complete development predictions. Its gate was not evaluated, and there is no 27B accuracy or distillation result.
What this experiment can tell us.
ABCD asks for the next recorded support tool action. Accuracy measures agreement with that action, not successful resolution of an entire customer issue. Public pretraining overlap is unknown.
H2O 4B and JevK5 2B used the same ordered 4,096 examples and native 15/16-option training views as the 4B comparator. Each new model had two epoch candidates. Development accuracy, then lower Brier, then fewer exposures selected H2O epoch 1 and 2B epoch 2. Final evaluation retained all 30 choices. The original 4B/9B search also included 1,024-example runs; the full-weight checkpoint used one fixed epoch and a different learning rate.
One training seed (553), fixed recipes, and different native prompts, calibration, and decision protocols limit generalization. Jev uses a knockout protocol; H2O reads all 30 choices in one pass. JevK5 2B is v0.2; the 4B comparator is v0.3. These are complete-system comparisons, not an isolated test of parameter count.
Paired 95% intervals resample the same 500 conversations 2,000 times, seed 553. They quantify cohort sampling uncertainty, not variation between training seeds, and are not adjusted for multiple comparisons. An interval spanning zero proves neither superiority nor equivalence.
Numerical checks, recovery, and calibration limits
The 9B/4,096 numerical check exceeded the 0.03 probability-error bound on its longest training input: fast BF16 vs FP32 error was 0.04114, and reference BF16 vs FP32 was 0.05380. BF16 cutoff ties changed knockout finalists, although final actions agreed. Replaying identical 16-option prompts reduced fast-vs-FP32 error to 0.01468. The run proceeded under a recorded exception; FP32 equivalence was not established.
H2O serialized both checkpoints before its training process failed a CUDA-memory release assertion. The unchanged weights were evaluated in fresh processes after correcting a PEFT layer-name validation mismatch. Training peak memory and uninterrupted total runtime are unavailable.
New H2O training used native temperature 0.75; the reused 1,024-example adapter was trained at 0.8. Their comparison changes training calibration as well as data exposure. Full controls and original failures are described in the reports.
B300 is the test bench.
All new timing and memory results are B300 measurements. Latency includes prompt preparation, tokenization, and synchronized single-case inference. It excludes loading, warmup, network transport, and queues. There was no new RTX 4090, quantization, concurrent-serving, or production validation of these checkpoints.
An earlier October 9 RTX 4090 ABCD test measured the older H2O adapter at 257 ms median versus 874 ms for Jev, using an older runtime and H2O temperature 0.8. That supports feasibility on the 4090 and the direction of the speed advantage, not exact performance for these new checkpoints.
Other earlier tasks were mixed: H2O tied Jev on Bitcast accuracy and trailed slightly on reply reserve and chess. A support-action benchmark does not identify one default model for every miner job.
The temporary GPU was deleted after verified local retrieval. Cumulative estimated compute cost was $37.95; this is a runtime estimate, not a provider invoice or a price for training a customer model.
Inspect the evidence.
Reports include full scores, paired intervals, development selection, pinned model revisions, and limitations. JSON files preserve aggregate results and source digests. Raw conversations, labels, per-case predictions, and trained weights are not distributed, so these downloads do not reproduce inference.