Adversarial self-benchmarking on coding agents: does the agent act on the decision that is actually in force, or on the one that happened to be in its context? Every run is sealed to a base commit and published with its limitations, including the runs where governed memory tied or lost.
Head to head runs against the tools you already use: a static CLAUDE.md, a verify prompt, Memory Bank, ADRs, spec-driven development, vector RAG. Each page states its N, its base commit, and its caveats. Results that lost are published too.
Ten models, three vendors, 272 external cases. A confident always-loaded summary suppresses the verification search, so the agent writes the stale answer without ever opening the newer notes on disk. Then the harder half: appending the correct values beside that summary corrected nothing when the block was partial or hedged, including the render path we ship, and everything when it was complete and named which value was in force.
A team renamed its core object and the owner ruled on it, but code, tickets, and standups kept using the old name. An agent reading the raw notes ships the superseded name 16 times out of 17. Same notes plus the accept/reject verdict: the in-force name 16 out of 17.
We built this one to prove that context-dumping dilutes at scale. It does not. Adherence ties at 1, 20, 60 and 200 rules, and governed memory costs more. We publish the row that lost, and the cost fix we pre-registered, measured, and rejected.
Ask an agent what was decided, who approved it, and when. A static CLAUDE.md answers the first 5 out of 5 and the other two 0 out of 5, with zero fabrication. Not a model failure: a markdown file has nowhere to put a reviewer.
Two parallel sessions decide incompatible things. All 5 caught in 10 to 15 seconds and delivered to the next turn. Precision is 83%, not 100%: a regional scope carve-out fools it once in three.
On a strong model a written rule ties with a governed one at zero violations. On a weaker model the written rule collapses (4 violations in 5) while the governed rule holds against a red team that tried seven escapes.
A static snapshot ships the stale value 4 times in 5 while governed memory is right 5 out of 5. Corrected after publication: against an ungoverned store holding both dated documents, plain retrieval also scores 5 out of 5. Supersession alone does not need governance.
What each tested approach actually is, and how to set it up.
What we retracted, what we corrected, and the contamination we caught while benchmarking ourselves.