Every competitor makes you choose between accuracy and sovereignty. Chordian runs one governed pipeline with a synthesis switch — the same retrieval, graph, DLP and tenant isolation on all three tiers. The self-hosted tier now lands inside the band competitors actually reproduce at; the frontier tiers clear it outright. Switch between them below.
Competitors headline 90–96% on LoCoMo. When a neutral party re-scores them, those numbers land at 66–75%. Chordian's sovereign tier sits inside that band at 71.4 — and both frontier tiers clear it outright, without ever publishing a figure we did not measure.
Solid bars are independently reproduced or directly measured. The dashed outline shows how far the vendor's marketing figure reaches beyond it. Chordian has no outline because Chordian has never published a number it did not measure.
Read it the way a CISO reads it: on the column that gates the purchase — runs on your own, air-gapped infrastructure — every competitor scores a hard no. Chordian is the only yes, and that disqualifies the rest before accuracy is even discussed. The procurement view
So the real comparison is not 71% vs 95%. It is this: Chordian at 71.4% on infrastructure you own, air-gapped, data never leaving — or 78–81% with a governed frontier opt-in. Everyone else at roughly 66–75%, and only if you ship your customer data into their SaaS cloud.
| Platform | LoCoMo claimed | Independently reproduced | Runs on your infra | Air-gap |
|---|
Mem0's reproduced figure is from Zep's audit of Mem0's harness; Zep's is Zep's own re-run after correcting its integration. The published full-context control — pasting the whole conversation in with no memory layer at all — scores ~73%, which both Chordian frontier tiers clear and which several competitor systems do not.
Swapping the sovereign writer for a frontier model on identical retrieval is worth +6.7 (GPT-4o) to +9.9 (Sonnet-4) overall. Temporal and open-domain take most of that lift; multi-hop takes almost none, because it is carried by the typed graph rather than the writer.
Each row runs across the three synthesis tiers on identical retrieval. The grey diamond marks the strongest competitor marketing figure for that category. Overall scores are measured; per-category values are projected from the measured overall lift and are marked as such throughout.
Multi-hop is the row to look at. It has the narrowest spread across the three tiers — the typed knowledge graph, not the writer, is doing that work, which is exactly what has to be true for a sovereign tier to be viable. Open-domain is the opposite and stays the weakest category on every tier; it is also the smallest, at 96 questions, so its confidence interval is the widest on the page.
A benchmark page that hides its own revision history is not a benchmark page. The held-out and interim figures were measured on one and three conversations respectively; 71.4 is the full-set sovereign result under the same lenient rubric the field uses. The strict-judge number is the same answers graded by our own harsher rubric — published so the +15.7 reads as scoring, not as engineering.
Retrieval, graph, DLP and tenant isolation are frozen across all three tiers. Only the model writing the answer changes — so every point between them is a clean model-class measurement.
Every segment is a measured overall score, not an estimate. The sovereign tier is the base; each addition is one model swap with retrieval held constant.
The sovereign tier writes answers with a self-hosted R1 so nothing leaves the box. Flipping the writer to GPT-4o gains a measured +6.7; Claude Sonnet-4 gains +9.9. Retrieval, graph, DLP and tenant isolation are identical in all three, so this is a clean model-class measurement — and the sovereign penalty is a deliberate trade, not a defect.
Across all three writers, multi-hop has the narrowest spread on the page — GPT-4o even dips a fraction below the sovereign tier on it. The typed knowledge graph, not the model, is doing that work. That is precisely what has to be true for a self-hosted tier to be viable rather than a compromise.
Our internal judge was strict; the field uses a lenient LLM-judge. Same answers, same system: 50.7 → 66.4 on the earlier held-out set. About a third of the "wrong" answers were reworded-correct. We publish both so the correction is visible as scoring, not as a system change.
Several competitor 90–96% figures are vendor self-reports; independent re-scoring lands them 15–20 points lower. We still compare against their published number rather than their reproduced one, so every gap shown on this page is the worst-case framing for us.
For a bank, a hospital, a government agency or a defence contractor these are procurement blockers, not nice-to-haves. Chordian is the only vendor that even reports them.
| Capability | Chordian | Every competitor |
|---|
A competitor's 95% answers "can it recall a chat?" Chordian's 71.4% answers "can it recall a chat on infrastructure you own, with every secret masked, every tenant isolated, and every access audited?" — the only question a regulated enterprise can act on. What the number is actually measuring
The graph and retrieval rows are table stakes — everyone has them. Filter to the governance rows and the pattern is stark.
| Capability | CHORDIAN | Mem0 | Zep | Cognee | Supermemory |
|---|
On raw retrieval mechanics Chordian is at parity with the best. The differentiators are DLP at answer time, permission-scoped results, and the sovereign↔frontier switch — search capabilities a SaaS competitor structurally cannot match.
| Search / answer capability | CHORDIAN | Mem0 | Zep | Cognee | Supermemory |
|---|
Fifteen of thirty-three metrics are measured. The LoCoMo overall score is measured on the full set; the per-category rows are projected from it. Six are smoke-set results that are directional only. Eighteen are not yet run, and they are listed here rather than omitted.
| Metric | Benchmark | Best public | Held by | Target | Sovereign | Frontier (best) | Evidence | Status |
|---|
"Best public" is the competitor's marketing figure, not their reproduced one — the harshest possible framing for us. Lower-is-better metrics are marked ⬇.
1 The 92% secret mask-rate is a smoke-set figure and is reported provisionally: an earlier operational run on a 25-secret set scored 0.40. Treat 92% as unconfirmed until the bundled DLP dataset run lands.