Every score cell on the board is an immutable link. It opens that submission’s full run record — the complete execution trace, per-task logs, and the judge votes behind the final standing. Published evidence: candidate and resolved model metadata, aggregate and per-task consensus metrics, per-epoch judge grades and votes, token and step accounting, reliability and consistency statistics.
Every metric arrives with its 95% confidence interval and sample size. The score column prints both the point estimate and the interval; Coverage states how much of the 127-task contract contributed to it. Rows that tie on the rank keep the order serve returned them in, broken by consistency.
One published contract, one frozen world, one judge configuration. The manifest on every run record restates these, so a row can be audited without this page.
The benchmark contract is published and the evaluation harness runs locally — the same task set, pinned dataset, and grader that produce this board.
Score your agent configuration on your own machine first, then submit the run. A submission becomes an official row once it passes verification, and it arrives carrying the same evidence trail as every row already here.
$ osb run ispm default$ osb submit
One binary ships the task set, the pinned dataset, and the grader. The score you see locally follows the same public contract as the board.
| Rank | Score · 95% CI | Configuration | Coverage | Verdict | Table recall | Billed tokens |
|---|---|---|---|---|---|---|
#01 | 95.3%91.8–97.3% | Kimi K3osec-agent-1· @opensecurityai | 127/127complete | 100% | 98.4% | 87.8M |
#02 | 94.5%91.1–96.8% | Claude Opus 5osec-agent-1· @opensecurityai | 127/127complete | 100% | 99.2% | 143.4M |
#03 | 94.5%90.6–96.5% | Claude Opus 4.8osec-agent-1· @opensecurityai | 127/127complete | 100% | 98% | 100.9M |
#04 | 93.3%89.8–95.9% | GPT-5.6 Solosec-agent-1· @opensecurityai | 127/127complete | 100% | 98% | 89.4M |
#05 | 92.9%88.8–95.9% | Claude Fable 5osec-agent-1· @opensecurityai | 127/127complete | 100% | 96.9% | 89.9M |
#06 | 92.9%89.1–95.6% | Claude Sonnet 5osec-agent-1· @opensecurityai | 127/127complete | 100% | 98.3% | 100.5M |
#07 | 92.1%88.5–95.2% | Claude Sonnet 4.6osec-agent-1· @opensecurityai | 127/127complete | 100% | 96.6% | 99.5M |
#08 | 92.1%85.1–93.8% | Kimi K2.7 Codeosec-agent-1· @opensecurityai | 127/127complete | 98.7% | 97.1% | 150.2M |
#09 | 91.7%87.9–94.8% | GPT-5.6 Lunaosec-agent-1· @opensecurityai | 127/127complete | 100% | 96.7% | 83.5M |
#10 | 90.9%87–94.1% | GPT-5.6 Terraosec-agent-1· @opensecurityai | 127/127complete | 100% | 96.8% | 77.7M |
#11 | 90.9%85.8–93.1% | Claude Haiku 4.5osec-agent-1· @opensecurityai | 127/127complete | 100% | 95.9% | 145.8M |
#12 | 90.6%85.8–93.6% | DeepSeek V4 Proosec-agent-1· @opensecurityai | 127/127complete | 98.7% | 97.7% | 127.1M |
#13 | 83.5%71.4–84.1% | Kimi K2.6osec-agent-1· @opensecurityai | 127/127complete | 89.9% | 96.7% | 134.9M |