The measured leaderboard
58 clan models through the canonical CSOAI flywheel on a Kaggle T4 — 4,408 cells, frozen split (45 practice / 31 held-out items per model). Ordered by held-out accuracy. This is a descriptive ordering, not a crown: at n=31 the 95% CI on a single accuracy is roughly ±0.18, so neighbours within that band are statistically indistinguishable. Cross-substrate flags come from the M4-vs-T4 crosscheck (57 joined models, weights-verified join).
sov33-unified — 0.778 practice / 0.645 held-out, 231.8 tokens/correct, cross-substrate consistent (Δ −0.097)
clan-meok-adversarial — 0.710 held-out, but the single most divergent model cross-substrate (Δ +0.355 vs M4). Flagged, not celebrated.
qwen2.5:0.5b — 0.533 / 0.548. All clan-* members are postures of this base (weights-verified join).
| Model | Practice | Held-out | Overfit gap | t/c (held-out) | M4 held-out | Δ T4−M4 | Cross-substrate | Substrate |
|---|---|---|---|---|---|---|---|---|
| Loading measured data… | ||||||||
Click a column header to sort. Overfit gap = practice − held-out (positive = overfit to practice). t/c = tokens per correct answer (cost proxy). M4 columns: weights-verified cross-substrate check; blank = not joined.
Honest notes
- The v7 weight-divergence lesson.
sov33-v7measured 0.581 held-out on T4 but 0.290 on M4 (Δ +0.291). A same-named model on two substrates is not the same model unless the weights join. That is why the crosscheck join is weights-verified, not name-verified. - LOST-WEIGHTS register. Eight custom-weight models (
sov33-dist-c1..c3,sov33-evolved-c1/c3,sov33-evolved,sov33-v6,sov33-evolved-patched) were honestly declared LOST after a blob-store wipe on 2026-08-02. They are absent here rather than silently substituted: a model NAME is not a model. - Salt-split reproducibility. RunPod A4500 control runs v1 (
runpod-v1-control-saltcheck) and v2 (runpod-v2-salt-baseline) share corpus root3729b52e…3ef01d(417 provisions). The anti-Goodhart salt re-splits practice/held-out between runs, so per-split accuracies move while the corpus anchor holds — reproducibility is proven at the anchor level, not by split leakage. - 17 of 57 joined models diverge (|Δ| > 0.15) between T4 and M4 (mean signed Δ 0.096, median 0.097). Substrate is a measurement variable; SwarmBench reports it, always.
Provenance — every number traces to an anchored artifact
| Artifact (estate path, anchored at write time) | sha256 |
|---|
Machine-readable source: /leaderboard/data.json — regenerated by the daily lane; hashes recomputed at generation time.