AI memory compared: mem0, Zep, Letta, Cognee, Supermemory, us
Every comparison of AI memory tools ends in a ranking. And every ranking we found was published by the tool sitting at the top of it.
So we built the table that is missing: where each published number comes from. Which benchmark, which variant, how many questions, which model answered, which model graded. And whether anyone can redo the arithmetic.
Three things stand out at once:
- Two of the highest scores do not say which benchmark variant produced them. LongMemEval ships in several sizes. The hard one buries the answer in ten times more distractor sessions than the easy one. Without that word, nobody can place the number.
- That table has ten rows. Three have every cell filled, and all three are ours. You can check it without reading a sentence: count the em dashes.
- Our result files are public. A
verify.jsrecomputes every figure and fails if one of them does not follow. No other project in the table offers one. Skip to it if that is the only part you want.
One more number, the one we put forward: our retrieval loses 20% when the corpus gets 76 times bigger. In the reference paper, the RAG baselines lose 22 to 29%.
Our interest, stated up front: we build Mnemosyne OS. We are the last rows of both tables. Everything about the other five comes from their sites, their docs and their repositories, read on 13 September 2026. The links are there so you can check.
What these things actually are
These six products are not alternatives to each other. Five are parts a developer puts inside an application. One is an application.
| What it is | For | Data sits | Self-host | Code | Stars | |
|---|---|---|---|---|---|---|
| mem0 | memory infrastructure | developers | their cloud | yes | Apache-2.0 | 65.2k |
| Zep | memory service | enterprise | their cloud | yes, your VPC | see note | 4.9k |
| Letta | agent platform, ex-MemGPT | developers | not stated | yes * | Apache-2.0 | 24.7k |
| Cognee | memory library | developers | your machine, or BYOC | yes | Apache-2.0 | 30.7k |
| Supermemory | memory API, plus plugins | developers | not stated | yes | MIT | 29.7k |
| Mnemosyne OS | desktop app | one person | your machine | no server to host | app closed, SDKs MIT | 13 |
Each one describes itself in its own words. mem0: “drop-in memory infrastructure for AI agents and apps”. Zep: “agent memory, at enterprise scale”. Cognee: “Open Source Memory Platform for Agents”. Supermemory: “the default engine for memory and continual learning for agents”. Those are four names for one job.
“Not stated” means their published docs do not say. * Letta has a Docker self-host guide, but the guide itself says nobody maintains that surface.
The Zep note. Its GitHub repository shows an Apache-2.0 licence and 4.9k
stars. What it holds today: examples, integrations, benchmarks, mcp,
ontology, legacy. The memory service is not in there. Their open graph engine
is a different repository, Graphiti, at
30.9k stars. A licence shown on a repository says nothing about the product.
Keep that in mind before you read any “open source” column, ours included.
The stars column. We have 13. mem0 has 65,200. We are the smallest in the table by a long way, and the column stays.
The published numbers, and what they rest on
Here is what made us drop the ranking.
| Figure | Benchmark | Variant | N | Reader | Judge | Runs | |
|---|---|---|---|---|---|---|---|
| mem0 | 92.5% | LoCoMo | — | — | — | — | — |
| mem0 | 48.6% | BEAM | 10M | — | — | — | — |
| Zep | 94.7% | LoCoMo | not stated | 1,540 | gpt-5.4 | gpt-5.4 | — |
| Zep | 90.2% | LongMemEval | not stated | 500 | gpt-5.4 | gpt-5.4 | — |
| Letta | no figure | — | — | — | — | — | — |
| Cognee | no figure | — | — | — | — | — | — |
| Supermemory | ”state of the art” | 3 named | — | — | — | — | — |
| Mnemosyne OS | 77.1% | LongMemEval | full-haystack, unseen questions | 48 | gemini-3.8-flash | gpt-4o (official) | published |
| Mnemosyne OS | 61.7% | BEAM | 100K | 400 | gemini-2.5-pro | gpt-4.1-mini ** | published |
| Mnemosyne OS | 49.2% | BEAM | 10M | 200 | gemini-2.5-pro | gpt-4.1-mini ** | published |
An em dash means the project does not publish that cell. Letta and Cognee give no
figure on their landing pages. Supermemory claims state of the art on three
benchmarks without giving a single number. At Zep, the judge is the same
gpt-5.4 that wrote the answers. ** On BEAM we did not pick the judge:
gpt-4.1-mini comes with the benchmark.
Read our BEAM row with care. Our 49.2% sits three rows under mem0’s 48.6%, same benchmark, same tier. Do not read it as a ranking. They publish no reader, no judge and no context budget. Nobody can tell whether the two runs did the same work. And 0.6 points fits inside the noise of every judge we have measured.
We answer with files. Both BEAM runs are published, question by question.
A verify.js recomputes 61.7% and 49.2%, and refuses to agree with us if the
numbers do not follow. It is what we would want from mem0.
Read the “variant” column first. A LongMemEval score means nothing until you know which size of benchmark produced it. Two of the highest figures do not say. LoCoMo has the same problem.
Then read each row across. A figure from gpt-5.4 grading gpt-5.4 and a
figure from gemini-2.5-flash grading gemini-2.5-pro come from two different
instruments. We know how much that changes: on another benchmark, the same 400
answers graded by two judges moved 136 verdicts.
mem0 keeps the most complete leaderboard in the category, and writes this on its own page:
None of these numbers were generated using the same model stack, judge model, or retrieval configuration. […] Treat this table as a starting point for further reading, not a settled ranking.
They are right. So no ranking exists for anyone today, us included.
What is weak in our row
Ours is the only row with every cell filled. Filled cells are a publishing habit. They prove nothing about quality. Ours is also the smallest sample in the table. Here are its four weaknesses.
- N = 48. Zep’s figure rests on 500 questions. Ours rests on 48.
- Our sample is easier than the benchmark it comes from, by about 13 points. We measured that and published it next to the number it flatters.
- The judge contradicts itself on about 2.6 verdicts per 48. Any gap under five questions is therefore invisible to this benchmark. That includes our own improvements.
- 77.1% is the official LongMemEval judge, on 48 questions the engine had never seen. The protocol was published before the run. On the 48 questions used to tune the engine, our strict judge gives 85.4%. We publish both and never chain them.
The one figure we put forward
A score in this category cannot be compared with another. A degradation can. It compares a system to itself, and a corpus multiplied by 76 is 76 for everyone.
BEAM is the ICLR 2026 benchmark, and its own official judge graded our answers. Mnemosyne OS goes from 61.7% at the 100K tier to 49.2% at 10M. That is a loss of 20% for 76 times the haystack. Over the same jump, BEAM’s own paper shows its RAG baselines losing 22 to 29%, and the models that read the whole conversation losing 50 to 57%.
Both figures are in the table, next to the one BEAM number another project publishes. Read the two together. BEAM entries are self-published, and no two share a reader, a judge or a context budget. An absolute score ranks nothing.
We also measured something nobody publishes: the context budget actually served. We count it in tokens, the benchmark’s own unit, and we read it from the real system prompt instead of computing it from settings. It fits inside the budget BEAM’s LIGHT method grants itself. A memory score says little without the size of the context spent to get it. That number is missing everywhere.
Check it without trusting us
Everything above asks you to believe a table. This part lets you check it.
Our LongMemEval campaign publishes its per-question results: both strict runs,
the recall files, the holdout, the mechanically extracted ledgers. They sit on a
public provenance page,
with the reading key in
lexical-2026-08.
One command, node verify.js, recomputes every figure and fails on any
mismatch. A quiet rounding in our favour would show up there.
The BEAM runs behind 61.7% and 49.2% are published the same way, in
beam-2026-09,
with their own verify.js. It redoes BEAM’s aggregation instead of approximating
it. The score is a mean per conversation, then per category, then over the ten.
And event_ordering is graded with a Kendall tau, not the judge’s 0/1. A flat
mean gives a different answer.
Both judgings of the 100K run are in there too. So you can recompute the 136 figure above yourself. We checked the tool can go red by breaking the data three ways: one flipped verdict, one deleted row, one masked infrastructure failure.
There are three limits. A re-run redoes our pipeline, not the benchmark: the questions and the gold answers belong to LongMemEval and BEAM, both MIT. The engine that produced the answers is closed, so you audit the scoring and not the retrieval. And the core build behind the two BEAM runs is not stamped. Later runs record which build answered. Those two came before that, and nobody can recover it now.
A score published without its files cannot be recomputed. A score with its files can. Almost every score in this category comes without them.
Which one you actually want
- You are building an agent or an application and want memory inside it: take mem0, Letta, Cognee or Supermemory. We make a desktop application with an SDK. If your memory has to live in your cloud next to your service, we are the wrong shape.
- You are an enterprise team with a VPC, a compliance review and a latency budget: Zep and mem0 are built for that conversation.
- You want a graph you can inspect that runs on local models: Cognee says so and is Apache-2.0.
- You are one person who wants their memory on their own machine, in vaults you can see on disk: that is what we make. You choose which vaults a model may read. It is a different product from the five above.
Remember the second table. A score without a variant, without a judge and without a question count is a score nobody can reproduce. That holds when it favours us too.
More alternatives, on a site we do not run:
Read from each project’s sites, docs and repositories on 13 September 2026. Licences and star counts from the GitHub API the same day. Pages change, so every claim here is dated and linked. None of this judges how well these products work. It is a record of what each one publishes, the only thing checkable from outside. Corrections from the projects named are welcome, and will be applied with their date.