The open benchmark for mixtures

We don't sell Byzantine consensus. We measure what groups of models actually add — and publish when the answer is less than one.

Single-model benchmarks are saturated. Group-level measurement is empty. SwarmBench measures MoE models, mixture-of-models teams, swarms, clans and councils as collectives — frozen splits, 95% confidence intervals, published harness, signed attestation for every run.

We retired our own prior positioning after our refutation ledger measured a 33-agent council at an effective ensemble size of n_eff 1.21 of 3 — correlated members voting is one model with extra steps. That retraction is now the product: the same honest accounting, run on your mixture.

The six measurements

Deterministic, frozen-split, reported with 95% CIs and a three-outcome honesty rule (measured / unparseable / UNMEASURED). Status is labelled per lane — measured lanes cite their anchored artifact; open lanes are marked as targets, not results.

1. Collective uplift

Baseline measured

Does the group beat its best individual member? The killer metric. 58 member baselines measured on Kaggle T4; collective runs are the next lane.

2. Delegation accuracy

Lane open

Did the router pick the right expert? Measured against per-item member outcomes already in the spread cells.

3. Consensus robustness

Measured

Kill k of N agents, measure verdict drift. Test subject #1: the retracted 33-seat council — n_eff 1.21, published on the refutation ledger.

4. Cost per verdict

Measured

Tokens per correct answer across configurations, from the same run cells — so uplift is never quoted without its bill.

5. Agreement entropy

Lane open

How much independent signal the votes actually carry. Feeds the Refutation Ledger directly when a group's entropy collapses.

6. Adversarial robustness

Lane open

Red-team attacks the swarm; the trust spine scores it. Attack suites are drawn from the estate's adversarial corpus.

The Illusion Index

Our 33-agent council measured n_eff 1.21 of 3. What is yours?

Naive Byzantine councils are retracted on our own ledger: correlated members voting adds cost, not signal. The Illusion Index is the public register of measured effective ensemble sizes — starting with ours.

Live trust spine

Every SwarmBench run is anchored to a corpus hash and can be sealed with a signed attestation. These endpoints are live: