A multiplier cannot rescue a zero
Note, September 2026. On 2026-09-28 an outside audit found a wrong verdict in the figures below. The erratum corrects the end-to-end result to 35/48 (72.9%) under the strict judge and 37/48 (77.1%) under the flexible one. The current result is a rerun whose protocol was published before it ran: 77.1% (37/48) on questions the engine had never seen, official LongMemEval judge (the benchmark page). This post stays up as it was written.
Ask a memory system “What shift was Admon assigned on Sunday?”. One thing matters: can it find the single session that contains Admon?
Ours could not. The miss came from the design, so it happened every time.
The blind spot
Mnemosyne OS ranks memory by dense-vector similarity: e5 embeddings, cosine, one ranking. That is the right tool for meaning. “The time I complained about the hotel” finds the complaint without sharing a word with it. A rare literal token is another matter. A proper noun, an identifier or a version number is exactly what an embedder maps close to noise. It has never seen that string. It has no meaning to hang on it. The one session you need can sit at rank #800.
We measured this on LongMemEval full-haystack, with the same protocol as the 72.9% campaign: roughly 480 distractor sessions per question, on a 48-question sample. For 6 questions out of 48, not a single evidence session came back. On 35 questions we could track the exact answer-bearing chunk. It was absent from the served context 10 times.
A model cannot answer with a chunk that retrieval never handed it.
The tempting fix, measured at zero
The codebase already contained the obvious answer: a term boost. Extract the distinctive tokens from the question. When a memory chunk contains one verbatim, multiply its cosine by 2.
Two things happened. Both taught us something.
First, fed naively, it was a catastrophe. A prose question capitalises its first word. So “What shift was Admon…” boosted “what”, a token present in nearly every conversational chunk in existence. Evidence recall collapsed from 38/48 to 17/48. One rule survived: a capital letter only evidences a proper noun when the word is not sentence-initial.
Second, once tamed, it measured neutral end-to-end. The arithmetic says why. The sessions we were missing scored near-zero cosine, and 0.30 doubled is still below 0.80. A multiplier amplifies a signal that exists. It cannot create one. The chunks this blind spot loses have no rank at all, so there is nothing to multiply.
A second ranking instead of a bigger multiplier
The fix that worked never consults the vectors at all.
Rank the vault twice. Once by cosine, deeper than the budget (200 instead of
32). Once lexically, with Okapi BM25. BM25 is the forty-year-old
term-weighting scheme that managed “hybrid search” products run under the
hood. Then merge the two rankings with Reciprocal Rank Fusion: each document
scores the sum of 1/(k + rank) across channels.
Ranks, never scores. A cosine lives in [0,1]. A BM25 sum is unbounded. Hybrid search usually goes wrong when someone normalises one against the other. RRF never compares them.
Fusion by rank also carries a safety property. A document that appears in only one ranking can at best tie the other channel’s #1. It can never outrank a document both channels agree on. So the lexical channel cannot seize the budget. It can only displace vector hits ranked so deep that their own reciprocal has decayed. And BM25 returns nothing for a document that shares no query term. So the channels overlap only partially, and that is exactly where fusion has something to add.
The context budget did not move: still 32 sources. The channel changes which chunks are served, never how many.
The numbers
| Measurement (48-question full-haystack sample) | Vector only | + lexical channel |
|---|---|---|
| Questions with all evidence sessions retrieved | 38/48 | 41/48 (+6/−3) |
| Answer-bearing chunk in the served context (35 tracked) | 25/35 | 30/35 (+6/−1) |
| End-to-end, strict judge | 29/48 | 37/48 = 77.1% (paired p = 0.0215) |
Retrieval is measured deterministically: byte-identical inputs, no LLM in the loop, free to reproduce. The end-to-end number goes through an LLM judge, and the judge is noisy. Replaying byte-identical runs through this rig flips about 2.6 verdicts per 48. So this bench cannot resolve a gap under ~5 questions. Retrieval is therefore the instrument, and the answer rate is the confirmation. The +8 clears the floor with room to spare. It reproduced across two independent runs that agreed verdict for verdict.
Audit it. Every number above recomputes from published per-question rows.
The raw runs and the mechanically extracted ledgers live on the
benchmark provenance page.
The raw runs are both strict runs, the deterministic recall files and the
holdout. The campaign’s reading key is in
lexical-2026-08/.
One command, node verify.js, re-derives every cell and fails on any
mismatch.
Two ways to manufacture a result
There are two classic ways to manufacture a result like this. Tune the hyper-parameters on your sample. Report only the sample you developed on.
We did neither.
The parameters are the literature defaults: BM25 k1=1.5, b=0.75, and RRF
k=60, the constant from the original paper. We left them untuned on purpose.
Our 48-question sample runs about 13 points easier than the benchmark it is
drawn from. Fitting to an easy sample manufactures gains that do not transfer.
Then the holdout: 48 fresh questions the channel had never seen, drawn after the design was frozen. The gain transferred: +4 questions with complete evidence, +2 answer-bearing chunks, zero regressions. Not one question retrieved worse. The dominant gain category was the same as on the development sample: temporal reasoning.
What we do not claim
Earlier in August we measured our own retrieval montage augmented with a managed hybrid-search API (Google Vertex AI Search). That Google-augmented montage scored 37/48 on the strict judge. The local channel now reaches the same 37/48, with no API, no account, and nothing leaving your machine.
That sentence does not say we beat anyone. We compared two of our own montages under one protocol we control. One called an external service, the other did not. The managed product was never scored in a way that permits a product-versus-product claim.
The other reservations stay printed next to the number. One benchmark. A 48-question sample that is thinner and easier than its parent set. An answer-side judge with a known noise floor. The holdout removed the objection we feared most: that we had quietly fit our own development set. A second benchmark family is still owed.
The plumbing
- The inverted index is persistent: plain SQLite tables inside the vault database, with a term dictionary keyed on integers. A cold build on a 2,470-chunk vault takes 19.5 s naive (one fsync per posting) and under a second in a single transaction. After that it is incremental: an ingest indexes only what it added.
- Not FTS5, deliberately. The benchmark harness runs on
node:sqlite, which has no FTS5 module. The lever would have been unmeasurable there, and an unmeasured lever is indistinguishable from one that never ran. FTS5’s defaults (k1=1.2, theunicode61tokenizer) would also have silently changed the very ranking we measured. - An incomplete index returns nothing, and the pipeline falls back to the vector ranking, yesterday’s behaviour. A half-built index has wrong document frequencies and produces a confident, silently wrong ranking. A missing feature beats a wrong answer.
- Per-app memory isolation is enforced inside the channel itself, with its own tests. Nothing is delegated to whoever happens to call it.
The channel ships in v1.3.8, on by default. MNEMO_LEXICAL_FUSION=0 restores
the vector-only path byte for byte. Our own reservations demand that escape
hatch.
A proper noun your memory could not hear now finds its session. The gain is local, measured, reproduced twice, and confirmed on 48 questions the channel had never seen.