← Exomem

Benchmark fairness contract

How we benchmark memory systems

Almost every benchmark comparing AI memory systems is published by one of the products being compared, and configured by them too. That is not usually fraud. It is that tuning your own system well and a competitor’s system adequately is the natural outcome of knowing one of them properly.

This page is the set of rules Exomem commits to before publishing any comparative number. It is deliberately published ahead of results, because rules announced after the fact are worth very little.

Why this exists

On 31 July 2026 this project published head-to-head findings against Basic Memory. On 8 August an independent adversarial review rejected them, and on 9 August every cross-product figure in that document was withdrawn — not caveated, withdrawn.

The defects ran in both directions, including ones that flattered us:

  • The competitor’s vector index was never built, so every figure recorded for it before that date measured a system that was not working.
  • After that was fixed, the harness was found feeding the competitor hundreds of oracle-normalised fact lines phrased in the query’s own vocabulary — an advantage no real deployment would have.
  • Knowledge time was never transmitted to either system, provenance and abstention were authored by the shared answerer rather than by the products, and the run profile de-tuned Exomem’s own defaults.

Findings about Exomem alone survived that review — the product defects it exposed were real, and were fixed. The comparisons did not survive, and no replacement has been published yet. This contract exists so that the defect class that produced them is structurally impossible rather than merely discouraged.

The one-line rule

Competitor-side configuration is competitor-authored, or it does not run.

What that means in practice

  1. Competitor configuration is competitor-authored. Every configuration value applied to another product traces to that product's own code, by file and line, or to its published documentation, by URL. A setting without provenance refuses to run rather than running on our judgement of what is fair.
  2. We run as a guest, on their harness. Head-to-head rows come from the competitor's own benchmark suite, with Exomem entered as a guest provider. We author only our own integration — the same posture any vendor takes on someone else's public suite.
  3. One reader, one judge, one ledger. Answer and evaluation stages from a competitor's harness are never republished as-is. Retrieval artifacts are exported and re-judged under a single frozen reader, so a difference in scores cannot hide inside a difference in prompts or judges. Where a suite excludes failed questions from its accuracy figure, we count them.
  4. Our own glue is disclosed and measured. Whatever we did write for a row — projectors, drivers, exporters — is accounted for by file, line count and endpoint. If our integration is substantially larger than a competitor's, that asymmetry is itself reported rather than quietly enjoyed.
  5. Harness faults are never contender losses. An unreachable service, a model that failed to load, an index that was never built: these invalidate the row for every product equally. The previous programme published a competitor scoring zero while its embedding model had silently failed to download. That outcome is now structurally forbidden.
  6. Readiness is proven, not assumed. An exit code is not evidence. Each product has a named verification method — vector-chunk counts and a log line, a terminal document status plus a canary, doctor checks that refuse rather than degrade. Where a product's default mode genuinely offers no completion signal, the row is marked unverifiable and disclosed, because invalidating a product's default mode would be its own bias.
  7. The plan is fixed before the run, and changes leave a trail. Scenario families, assertions and acceptance thresholds are ratified before any competitor runs. Ratification leaves the approved bytes unchanged and adds an immutable receipt. Later changes are ordered amendments — mutation, omission, reordering and branch substitution all refuse.
  8. Independent adversarial review before publication. An auto-generated packet of assumptions, confounds and suspicious-win flags goes to a reviewer with no stake in the outcome, and every material objection is either fixed or published beside the claim. A result showing a competitor ahead is a valid, publishable outcome of this programme.

What is not here yet

Results. The comparative programme is being rebuilt under these rules, and publishing numbers before that work is finished would repeat exactly the mistake this page documents. When comparative figures appear, they will arrive with the per-row fairness matrix — configuration provenance, who authored each knob, enumerated asymmetries and the direction each one favours, readiness evidence, version and dataset pins, and any measurement that was blocked, with the reason.

The full contract, including the reviewer’s checklist of what to attack first, is maintained in the open-source repository and is the normative version of this page.

Read the full fairness contract on GitHub ↗

Exomem itself is an open-source, MCP-native memory server over a Markdown vault you own — see the product page, or the honest comparisons with mem0, Letta, Zep and Basic Memory and claude-mem, which are written from published behaviour rather than from benchmark runs.