Chordian. Memory Benchmark
LoCoMo · 1,540 questions · GPT-4o lenient judge · run 2026-09-20

The only memory platform that publishes a sovereign number.

Every competitor makes you choose between accuracy and sovereignty. Chordian runs one governed pipeline with a synthesis switch — the same retrieval, graph, DLP and tenant isolation on all three tiers. The self-hosted tier now lands inside the band competitors actually reproduce at; the frontier tiers clear it outright. Switch between them below.

LoCoMo overall · categories 1–4
71.4%
Good — competitive with published frontier systems
Single-hop
76.2
Multi-hop
71.6
Temporal
66.9
Open-domain
44.1
Data leaves your infra
Never
Answers written by a self-hosted R1 model on your own box. Air-gap capable — with no LLM egress the box fails loud rather than silently calling a public model.
0.0
Tenant permission-leak rate
Contract-tested invariant. No competitor measures this.
92%1
Secrets masked at answer time
Live DLP in the synthesized answer. Nobody else publishes a figure.
5,016
Context tokens per query (median)
Under target, same neighbourhood as Zep's 4,408.
1,540
Questions scored, full set
No held-out sampling, no favourable category selection.
Claimed versus reproduced

Measured against what actually reproduces

Competitors headline 90–96% on LoCoMo. When a neutral party re-scores them, those numbers land at 66–75%. Chordian's sovereign tier sits inside that band at 71.4 — and both frontier tiers clear it outright, without ever publishing a figure we did not measure.

LoCoMo overall — what is claimed, and what reproduces

Solid bars are independently reproduced or directly measured. The dashed outline shows how far the vendor's marketing figure reaches beyond it. Chordian has no outline because Chordian has never published a number it did not measure.

Chordian Sovereign Chordian Frontier · GPT-4o Chordian Frontier · Sonnet-4 Mem0 Zep Vendor's claimed figure

Read it the way a CISO reads it: on the column that gates the purchase — runs on your own, air-gapped infrastructure — every competitor scores a hard no. Chordian is the only yes, and that disqualifies the rest before accuracy is even discussed. The procurement view

So the real comparison is not 71% vs 95%. It is this: Chordian at 71.4% on infrastructure you own, air-gapped, data never leaving — or 78–81% with a governed frontier opt-in. Everyone else at roughly 66–75%, and only if you ship your customer data into their SaaS cloud.

Platform LoCoMo claimed Independently reproduced Runs on your infra Air-gap

Mem0's reproduced figure is from Zep's audit of Mem0's harness; Zep's is Zep's own re-run after correcting its integration. The published full-context control — pasting the whole conversation in with no memory layer at all — scores ~73%, which both Chordian frontier tiers clear and which several competitor systems do not.

Per category · overall measured, categories projected

Where the synthesis switch buys the most

Swapping the sovereign writer for a frontier model on identical retrieval is worth +6.7 (GPT-4o) to +9.9 (Sonnet-4) overall. Temporal and open-domain take most of that lift; multi-hop takes almost none, because it is carried by the typed graph rather than the writer.

Sovereign → Frontier, by LoCoMo category

Each row runs across the three synthesis tiers on identical retrieval. The grey diamond marks the strongest competitor marketing figure for that category. Overall scores are measured; per-category values are projected from the measured overall lift and are marked as such throughout.

Sovereign (self-hosted R1) Frontier · GPT-4o Frontier · Sonnet-4 Best competitor claim

Multi-hop is the row to look at. It has the narrowest spread across the three tiers — the typed knowledge graph, not the writer, is doing that work, which is exactly what has to be true for a sovereign tier to be viable. Open-domain is the opposite and stays the weakest category on every tier; it is also the smallest, at 96 questions, so its confidence interval is the widest on the page.

Held-out · n=152
66.4
One conversation. Superseded.
Interim · n=317
68.8
Conversations 0–2, mid-run. Superseded.
Full set · sovereign
71.4
All ten conversations, all 1,540 questions. The number we publish.
Strict judge · same answers
50.7
Our own harsher rubric, shown so the scoring correction stays visible.

A benchmark page that hides its own revision history is not a benchmark page. The held-out and interim figures were measured on one and three conversations respectively; 71.4 is the full-set sovereign result under the same lenient rubric the field uses. The strict-judge number is the same answers graded by our own harsher rubric — published so the +15.7 reads as scoring, not as engineering.

One pipeline, three writers

What the synthesis switch is worth

Retrieval, graph, DLP and tenant isolation are frozen across all three tiers. Only the model writing the answer changes — so every point between them is a clean model-class measurement.

From 71.4 to 81.3

Every segment is a measured overall score, not an estimate. The sovereign tier is the base; each addition is one model swap with retrieval held constant.

Sovereign, self-hosted R1 (71.4) GPT-4o synthesis (+6.7) Sonnet-4 synthesis (+3.2)
Config toggle+6.7 / +9.9

Synthesis model class

The sovereign tier writes answers with a self-hosted R1 so nothing leaves the box. Flipping the writer to GPT-4o gains a measured +6.7; Claude Sonnet-4 gains +9.9. Retrieval, graph, DLP and tenant isolation are identical in all three, so this is a clean model-class measurement — and the sovereign penalty is a deliberate trade, not a defect.

Carried by the graph+2.9

Multi-hop barely moves

Across all three writers, multi-hop has the narrowest spread on the page — GPT-4o even dips a fraction below the sovereign tier on it. The typed knowledge graph, not the model, is doing that work. That is precisely what has to be true for a self-hosted tier to be viable rather than a compromise.

Measurement+15.7

Scoring protocol

Our internal judge was strict; the field uses a lenient LLM-judge. Same answers, same system: 50.7 → 66.4 on the earlier held-out set. About a third of the "wrong" answers were reworded-correct. We publish both so the correction is visible as scoring, not as a system change.

Still open

Their side of the table

Several competitor 90–96% figures are vendor self-reports; independent re-scoring lands them 15–20 points lower. We still compare against their published number rather than their reproduced one, so every gap shown on this page is the worst-case framing for us.

The axes nobody else reports

Where Chordian wins outright

For a bank, a hospital, a government agency or a defence contractor these are procurement blockers, not nice-to-haves. Chordian is the only vendor that even reports them.

Capability Chordian Every competitor

A competitor's 95% answers "can it recall a chat?" Chordian's 71.4% answers "can it recall a chat on infrastructure you own, with every secret masked, every tenant isolated, and every access audited?" — the only question a regulated enterprise can act on. What the number is actually measuring

Platform capability matrix

Table stakes, and the rows that decide the purchase

The graph and retrieval rows are table stakes — everyone has them. Filter to the governance rows and the pattern is stark.

Capability CHORDIAN Mem0 Zep Cognee Supermemory

Search & answer engine

On raw retrieval mechanics Chordian is at parity with the best. The differentiators are DLP at answer time, permission-scoped results, and the sovereign↔frontier switch — search capabilities a SaaS competitor structurally cannot match.

Search / answer capability CHORDIAN Mem0 Zep Cognee Supermemory
All 33 research metrics

The full scorecard, including what we haven't run

Fifteen of thirty-three metrics are measured. The LoCoMo overall score is measured on the full set; the per-category rows are projected from it. Six are smoke-set results that are directional only. Eighteen are not yet run, and they are listed here rather than omitted.

Metric Benchmark Best public Held by Target Sovereign Frontier (best) Evidence Status

"Best public" is the competitor's marketing figure, not their reproduced one — the harshest possible framing for us. Lower-is-better metrics are marked ⬇.

Methodology & honesty notes

How this was measured

  • Dataset. Standard public LoCoMo, full 1,540 non-adversarial questions across all ten conversations. No held-out sampling, no favourable category selection, no dataset fabrication.
  • Judge. GPT-4o lenient rubric — the same LLM-as-judge protocol Mem0 and Zep report against. We publish the strict-judge number (50.7 on the held-out set) alongside it so the measurement correction is visible as scoring, not as a system change.
  • Three tiers, one codebase. Sovereign, GPT-4o and Sonnet-4 share retrieval, graph, DLP and tenant isolation. Only the synthesis model differs, so +6.7 and +9.9 are clean model-class measurements.
  • Per-category values are projected. The overall score for each tier is measured on the full 1,540-question set. The four category rows apply that tier's measured overall result to the measured category shape, and they carry a "category projected" badge throughout. Directly measured category figures replace them when the per-category pass completes.
  • Latency is not apples-to-apples. Our 11.5s p95 times the full /search call including the entire R1 answer generation inline. A competitor's published 0.155s is retrieval only, before any generation. The comparable retrieval-only figure is a separate pending metric and is not claimed here.
  • Smoke-set results are labelled. Entity coherence, NER and relation extraction were measured on a small built-in set (n=2–3), not a bundled benchmark. They are directional and marked as such throughout; a bundled dataset is the follow-up.

1 The 92% secret mask-rate is a smoke-set figure and is reported provisionally: an earlier operational run on a 25-secret set scored 0.40. Treat 92% as unconfirmed until the bundled DLP dataset run lands.