Single-model benchmarks are saturated. Group-level measurement is empty. SwarmBench measures MoE models, mixture-of-models teams, swarms, clans and councils as collectives — frozen splits, 95% confidence intervals, published harness, signed attestation for every run.
We retired our own prior positioning after our refutation ledger measured a 33-agent council at an effective ensemble size of n_eff 1.21 of 3 — correlated members voting is one model with extra steps. That retraction is now the product: the same honest accounting, run on your mixture.
Deterministic, frozen-split, reported with 95% CIs and a three-outcome honesty rule (measured / unparseable / UNMEASURED). Status is labelled per lane — measured lanes cite their anchored artifact; open lanes are marked as targets, not results.
Does the group beat its best individual member? The killer metric. 58 member baselines measured on Kaggle T4; collective runs are the next lane.
Did the router pick the right expert? Measured against per-item member outcomes already in the spread cells.
Kill k of N agents, measure verdict drift. Test subject #1: the retracted 33-seat council — n_eff 1.21, published on the refutation ledger.
Tokens per correct answer across configurations, from the same run cells — so uplift is never quoted without its bill.
How much independent signal the votes actually carry. Feeds the Refutation Ledger directly when a group's entropy collapses.
Red-team attacks the swarm; the trust spine scores it. Attack suites are drawn from the estate's adversarial corpus.
Naive Byzantine councils are retracted on our own ledger: correlated members voting adds cost, not signal. The Illusion Index is the public register of measured effective ensemble sizes — starting with ours.
Every SwarmBench run is anchored to a corpus hash and can be sealed with a signed attestation. These endpoints are live: