Benchmark · LongMemEval-M · full-haystack
72.9% on the variant nobody cites.
Everyone cites LongMemEval. Almost nobody runs the full-haystack: ~480 distractor sessions per question — the honest variant, the hardest one. We publish the number as a floor, with the protocol, the failures, and everything you need to recompute it yourself.
The number
~480 distractor sessions per question.
The honest variant. The hardest.
There is a more comfortable version of this benchmark. It's the one the industry shows. We ran the other one, and we publish the number as a floor: two categories were never replayed with the full engine, for lack of compute time. Other models still remain to be tested.
No question was picked after the fact. One single configuration for everyone. And a success only counts if it reproduces.
What was measured
| Measure | Before | After | Δ | Conditions |
|---|---|---|---|---|
| Full-haystack, 48 q | 64.6% | 72.9% | +8.3 | all engines |
| Multi-session | 1/8 | 5/8 | +4 | the hardest category |
| Real executeQuery path, 61 q | 78.7% | 83.6% | +4.9 | topK 32 |
| Temporal (A/B inline) | 7/10 | 10/10 | +3 | reproduced ×2 |
| Recall at ~1 M vectors | — | 1.000 | — | MNEMO_ANN |
| 100% local, zero cloud | — | ~50% | floor | Qwen2.5-3B · Vulkan · n=12 · loaded machine |
The "100% local" line deserves its precision: it's a floor, not a ceiling. A three-billion-parameter model, on twelve questions, on a machine busy with something else. The big local reasoning models — we don't yet have the power to test them, so we don't know how high it goes, and we don't claim it.
And the levers that gave nothing, because they count just as much: relationship ledgers (worse), master pass (worse), k-means pre-sorting (useless), reserve sweep (noise). Prime numbers carry no memory signal — falsified six times, buried.
The section nobody writes
What we miss.
On the full-haystack, three questions still resist. They are broken down line by line. Here they are.
×2 — structural to the benchmark
The benchmark stacks several users into a single memory. A stranger's project becomes indistinguishable from yours in the served text. No answer-side filter can decide — verified. Mnemosyne is sovereign and single-human: this failure mode doesn't exist for you. We won't chase those two points.
×1 — our debt, not our design
An imbalance between what we write and what we read back: a corpus ceiling truncates a topic that's too dense. It's a real engineering limit, it has a ticket, it will be fixed. We don't dress it up as a choice.
Don't take our word for it
The grader and the verdicts are public. Recompute.
The published grader and the per-question verdicts let you re-derive the score in one command — no engine, no network. The full methodology, the root-cause analysis and the 16 raw logs are open too.