Meetless Research

We benchmarked governed memory against what you already use.

Adversarial self-benchmarking on coding agents: does the agent act on the decision that is actually in force, or on the one that happened to be in its context? Every run is sealed to a base commit and published with its limitations, including the runs where governed memory tied or lost.

Benchmarks

Head to head runs against the tools you already use: a static CLAUDE.md, a verify prompt, Memory Bank, ADRs, spec-driven development, vector RAG. Each page states its N, its base commit, and its caveats. Results that lost are published too.

Delivering the correct value is not the same as the agent using it

31 Jul 2026

Ten models, three vendors, 272 external cases. A confident always-loaded summary suppresses the verification search, so the agent writes the stale answer without ever opening the newer notes on disk. Then the harder half: appending the correct values beside that summary corrected nothing when the block was partial or hedged, including the render path we ship, and everything when it was complete and named which value was in force.

Governed memory vs. a contested rename

16 Jul 2026

A team renamed its core object and the owner ruled on it, but code, tickets, and standups kept using the old name. An agent reading the raw notes ships the superseded name 16 times out of 17. Same notes plus the accept/reject verdict: the in-force name 16 out of 17.

We built this benchmark to prove our thesis. It disproved it.

14 Jul 2026

We built this one to prove that context-dumping dilutes at scale. It does not. Adherence ties at 1, 20, 60 and 200 rules, and governed memory costs more. We publish the row that lost, and the cost fix we pre-registered, measured, and rejected.

A file can hold the decision. It cannot hold who approved it.

13 Jul 2026

Ask an agent what was decided, who approved it, and when. A static CLAUDE.md answers the first 5 out of 5 and the other two 0 out of 5, with zero fabrication. Not a model failure: a markdown file has nowhere to put a reviewer.

Two agents, same hour, incompatible decisions

12 Jul 2026

Two parallel sessions decide incompatible things. All 5 caught in 10 to 15 seconds and delivered to the next turn. Precision is 83%, not 100%: a regional scope carve-out fools it once in three.

A rule in a file is a request

12 Jul 2026

On a strong model a written rule ties with a governed one at zero violations. On a weaker model the written rule collapses (4 violations in 5) while the governed rule holds against a red team that tried seven escapes.

Governed memory vs. a static snapshot

11 Jul 2026

A static snapshot ships the stale value 4 times in 5 while governed memory is right 5 out of 5. Corrected after publication: against an ungoverned store holding both dated documents, plain retrieval also scores 5 out of 5. Supersession alone does not need governance.

Reference

What each tested approach actually is, and how to set it up.

Corrections

What we retracted, what we corrected, and the contamination we caught while benchmarking ourselves.

Built by Meetless. Methodology, raw counts, and corrections live on each page. Questions or a result you cannot reproduce: hi@meetless.ai.