Serious engineers discount vendor benchmarks on sight, and they are right to. So the credibility of this suite is not the win rate โ it is the ledger below: what survives a fair fight, what we retracted or corrected, and the contamination we caught benchmarking ourselves. If a result page reads cleanly, it is because the mess was moved here, not hidden.
Run fairly, most categories tie or lose: for raw fact recall a CLAUDE.md or a vector store hands back text and wins on cost. That is the thesis, not a defect โ we are a decision store, not a document store. Two claims survive, and they are the only ones we stand behind:
| result | the claim | why a text store can't close it |
|---|---|---|
| provenance | a file answers "who approved this?" 0/5, structurally | a markdown file has nowhere to put a reviewer |
| contested-canonical | on a rename practice contradicts, raw reasoning ships the superseded term 16/17 | text records what was said, not which decision still stands |
trust-containment used to sit at the top of this list. We retracted it on 2026-07-17 (see the ledger below) โ it graded the opponent against a verdict we never put in the opponent's corpus. Removing it made this list shorter and more honest.
| result | status | what happened |
|---|---|---|
| governed-retrieval | retracted | The prompt told the opponent it could not win โ "the answers are NOT in this repository; they live only in that knowledge base." A static arm holding all three facts scored 0/5 under it ("I could not query the knowledge base" โ it never opened the file) and 5/5 under a neutral question. It also asked for org facts we deliberately don't store. Withdrawn. |
| adherence-at-scale | published loss | We built it to prove selective injection beats context-dumping. It disproved our own thesis: adherence ties 5/5 at N=1/20/60/200 โ a long-context model finds the one governing rule among 200 every time. We published the loss. |
| governed-freshness | corrected | Beat a snapshot; then an ungoverned store with both docs + dates tied us 3ร cheaper on plain supersession. The claim narrowed (pre-registered): governance earns its keep only when the newest document isn't the authoritative one. |
| trust-containment | retracted | The graded answer โ which contradictory decision was accepted โ was never in any document; it lived only in the mla arm's governance verdict, scrubbed from every other arm's corpus. So mla's ceiling was 100% by construction and the opponent's a coin flip. Proof it was a guess: a plain vector-RAG arm on the identical corpus scored 3/3 by picking the rounder number, where "ungoverned" scored 0/3. We also graded honesty (UNRESOLVED) as failure. Withdrawn; a harness gate now refuses the shape. |
| decision-history | retracted | Same defect, caught before announcement. Dated notes let any arm order the values, but which were adopted vs. rejected lived only in the verdict โ so the exclusion the win rested on was impossible for the opponent by construction. |
For months the contamination we caught flattered us; once, one flattered the opponent. But the deepest error wasn't contamination at all โ it was a tautology: on trust-containment and decision-history we graded the opponent against a verdict we never put in its corpus, so the fixture decided the winner before any model ran. We caught it by giving the task a third arm (a plain vector store) that scored 3/3 on the same corpus the "loser" scored 0/3 on. Both are retracted, and a harness gate now refuses the shape. A benchmark you can't lose isn't a benchmark.
An agent under bypassPermissions on a dev machine is filesystem- and loopback-omniscient. Secrecy-by-obscurity loses; only unreachability (deny) or absence (a neutral substrate) wins. A representative slice โ note the direction column, because roughly half flattered us and half flattered the opponent:
| leak | what it would have published | direction |
|---|---|---|
| The prompt forbade the opponent from winning | our flagship 5/5-vs-0/5 โ measuring our question, not the opponent | flattered us |
| Our repo taught the opponent to distrust its own memory | "ungoverned memory makes an agent stuck, not wrong" โ understating the danger | flattered the opponent |
| macOS sleep suspended the process; timers never fired | 3 of 5 "static failures" that were never measurements at all | flattered us |
| Answer smuggled in a filename (-alt suffix) | the ungoverned arm "passing" by reading the path, never its memory | flattered the opponent |
| Empty-queue false-pass (async ingest) | mla graded against a KB that was never populated | against us |
| The mla arm was loading the operator's Gmail (auto-connected MCP) | "governed memory has a 2.1ร fixed overhead" โ invented by the test rig | against us |
| --force merged runs across fixture versions | two different experiments, averaged, reported as one | either |
The harness now aborts rather than measures on each of these: a foreign MCP in any arm, a config swapped mid-run, a corpus whose extraction didn't provably complete, a suspended process, a stale build, a prompt that presupposes a knowledge base, and a fact extracted but never ACCEPTED. A contaminated trial is indistinguishable from a measurement unless something refuses to print it.
We chased governance wins through five confounds on the deprecation benchmark alone, each a tie, each a way we had accidentally handed the opponent the answer โ up to and including giving it the tidy, indexed decision record that governance itself produces. The win only appeared once the opponent got what raw memory actually looks like: a contested mess. The narrow, defensible claims above are what's left when the bench stops flattering anyone.
We don't earn trust by winning the scoreboard. We earn it by publishing the null results, the retractions, and the harness โ so a skeptic can run it and get our numbers, or catch us if they can't.