Clan spread v1 — measured 2026-08-02

The measured leaderboard

58 clan models through the canonical CSOAI flywheel on a Kaggle T4 — 4,408 cells, frozen split (45 practice / 31 held-out items per model). Ordered by held-out accuracy. This is a descriptive ordering, not a crown: at n=31 the 95% CI on a single accuracy is roughly ±0.18, so neighbours within that band are statistically indistinguishable. Cross-substrate flags come from the M4-vs-T4 crosscheck (57 joined models, weights-verified join).

Practice leader
sov33-unified — 0.778 practice / 0.645 held-out, 231.8 tokens/correct, cross-substrate consistent (Δ −0.097)
Held-out leader (T4)
clan-meok-adversarial — 0.710 held-out, but the single most divergent model cross-substrate (Δ +0.355 vs M4). Flagged, not celebrated.
Base model
qwen2.5:0.5b — 0.533 / 0.548. All clan-* members are postures of this base (weights-verified join).
Model Practice Held-out Overfit gap t/c (held-out) M4 held-out Δ T4−M4 Cross-substrate Substrate
Loading measured data…

Click a column header to sort. Overfit gap = practice − held-out (positive = overfit to practice). t/c = tokens per correct answer (cost proxy). M4 columns: weights-verified cross-substrate check; blank = not joined.

Honest notes

Provenance — every number traces to an anchored artifact

Artifact (estate path, anchored at write time)sha256

Machine-readable source: /leaderboard/data.json — regenerated by the daily lane; hashes recomputed at generation time.