Skip to content

MemoryResearch by Sarbloh

Language modelscan't say“I never saw that.”

When a fact is missing, a language model reads the silence as “it didn't happen” and answers anyway. Our experiments show that exact, identity-keyed memory removes this failure, replicated on two models, and we publish every experiment, including the ones that failed.

Every number on this page links to a result file. Failed experiments are published too.

Scroll to walk through the evidence

The problem

Hallucination from memory is a classification error between “false” and “not stored”.

Attention always returns a blend of what it holds. It has no output for “no match.” When a model reads a summary or a long context and a fact is missing, it reads the silence as “that did not happen.” The answer sounds confident because nothing in the architecture can express absence.

Pilot · indicative

In a first pilot (one episode, 31 yes/no questions), a model reading only a 48-token summary was confidently wrong 39% of the time. Reading the full 337-token episode: 0%. Every summary error was a confident “no” to a true fact.

One episode, 31 questions, and every absent-fact question had the answer “no”, so the model's lean toward “no” inflates the figure. Indicative only.

Source: history/RESULTS.md (P1)

What we proved

Address memory by identity, exactly, instead of by similarity. Replicated on two models from different families, three seeds each.

Replicated

+41 to +50 pts

Accuracy of an identity-keyed store over chunk retrieval. Identity store: 100% in every run. Stale-belief errors: 0% vs 64–69% for chunk retrieval.

M08 templated +42.4 (Gemma) / +38.9 (Qwen); natural text +51.4 / +48.6.

Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt

Replicated

96.5% at 0.36%

Accuracy of a structural symbol index on a 2,410-file code repository while loading 0.36% of the repository's tokens. Plain BM25 retrieval: 88–89%. Definition-only: 49.6%.

M01b, mean of two models. Gemma 96.7% at 0.39%, Qwen 96.3% at 0.34%.

Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt

0 of 18 seed × model × experiment cells flipped direction. The rewrite check confirmed all facts preserved when synthetic sessions were rewritten into natural chat.

Similarity breaks as memory grows. Exact identity does not.

Share of questions about things that were never stored that receive a confident wrong answer, as the memory grows from 1,000 to 99,000 entities.

Tiny transformer trained from scratch on a synthetic world. The comparison is architecture-level, not a product benchmark.

Even with a perfect key, similarity lookup answers 46–50% of absent questions confidently and wrongly at 99,000 entities. Exact identity answers 0%.

Confident wrong answers on never-stored questions (%)
Memory size (entities)Similarity, perfect key (3 seeds)Exact identity (6 seeds)
1,0000.6%0.0%
10,00020.6%0.0%
99,00048.3%0.0%

Per-seed similarity values at 99k: 49.0 / 49.9 / 45.9.

Source: experiments/M11b_binding/results/eval_seed{0,1,2}.json, eval_seed{1,2}_i.json, eval_seed{3,4,5}_iii.json

In plain words

Imagine finding a book by “which cover looks most like what I want.” With ten books that works. With a hundred thousand, something always looks similar enough, even when your book isn't there. A library card number either exists or it doesn't.

Disclosure

Exact-identity seed 2 shows an accuracy dip at 10k/99k on one question type (true/false “is X's value V”), not seen on seeds 0, 1, 3, 4, 5. Cause unconfirmed.

Honest limits

  • Results are on synthetic worlds and small models: tiny transformers trained from scratch, Qwen3.5-4B, Gemma-4-E4B-it.
  • The identity result is replicated (2 models, 3 seeds each; 6 seeds on the tiny transformer). The in-network fix (M14) is a hypothesis.
  • The comparison of exact identity vs similarity at 99,000 entities comes from one tiny architecture.
  • No test yet on real user data, on models above 4B, or at production scale.
  • Several early pilots are single-seed and are labelled indicative.

What failed, and we say so

Every experiment has a written kill line before it runs. Killed experiments stay in the repository with the reason.

We also retired a broader architecture (“the LLM as a peripheral of a persistent controller”). Its controller and compiler halves have no supporting evidence. The memory result stands on its own.

HYPOTHESIS — NOT YET TESTED

Next: put the structure inside the model

Today the exact store sits beside the model. The next step is an exact lookup inside the network, with the mathematics written down before the run.

We have derived why similarity lookup fails at scale, and what a fix must satisfy. These are derived relations, not tested fixes.

Derived, not tested as a fixNot run
Attention puts weight ≥ 1 − ε on the right key when:
    β · (1 − s) ≥ ln(N / ε)
For random keys in d dimensions, distractor similarity:
    s ≈ √(2 ln N / d)
So the key width must grow with memory size:
    d ≥ 2 ln N / μ²
And the “not stored” threshold must move with N:
    τ_N ≈ √(2 ln N / d) + δ

At d = 64 and N = 99,000 this predicts a distractor similarity near 0.6. We measured 0.53–0.58. A threshold fixed at small scale (0.5) therefore opens at large scale. That is how the earlier network-layer experiment (M11) failed.

  1. M14Not run

    Calibrated read

    Does a size-aware threshold and a wider key remove the 99,000-entity failure?

    Kill line (pre-registered): confident-wrong on absent questions stays above half of M11's rate. TODO(owner): the brief says 45%, the claims register says 54.6%; confirm which

  2. M14Not run

    Exact two-level code

    A bucket × slot identity read through a zero-initialised gate. Target: 0% absent confident-wrong at 99,000 entities, inside the network.

  3. M12Not run

    Integration on a real model

    A frozen Qwen3.5-2B with an exact resolver outside the network and the value injected into the residual stream.

Product-key memory, Titans and nested learning explore memory inside models. Our claim is narrower: exactness and explicit absence that hold as memory grows.

End of the evidence

Back to the notebook: what it costs, how we work, and what to ask us.

07Fund

What funding buys

The bottleneck is compute and time, not ideas. The next results need GPU hours on larger models.

  1. 01

    Scale the replication

    Repeat the identity-vs-similarity results on a model above 4B parameters and on non-synthetic data. This is the largest open risk to the claim.

  2. 02

    Run M14

    Test the derivations. A pass makes the in-network claim publishable. A kill gets published too.

  3. 03

    Integrate and write up

    Run M12 on a real model, then publish with at least three seeds and a second model.

The ask[[OWNER: funding amount, instrument, and how it maps to the three items above.]]

Milestones[[OWNER: public milestone dates, if any.]]

08Method

How we work

  1. 01

    Prediction first

    Each experiment starts with a written hypothesis, a prediction, and a kill criterion.

  2. 02

    Kill criteria are binding

    A killed experiment stays in the repository with its reason. We never round a kill into a pass.

  3. 03

    Evidence rule

    A number without its result file is not a result. Claims for a paper need at least three seeds and a second model.

  4. 04

    Mistakes are logged

    Any bug or wrong assumption that cost more than an hour is written down with its cause.

Each ingredient has prior art (product-key memory, Titans, nested learning, retrieval-augmented generation). Our contribution is narrow: an experimental record that exactness and explicit absence hold as the memory grows, and a derivation of why similarity does not.

Read the experiment ledger

09FAQ

Questions investors ask

Isn't this just RAG / a vector database?

No, and that is the point. RAG and vector stores look up by similarity. Our result is that similarity lookup makes more confident wrong answers as the memory grows, even with a perfect key, while exact identity lookup does not. We also measured that structural identity beats chunk retrieval by 41–50 points on our tasks.

Isn't this already done (Titans, product-key memory, nested learning)?

Parts of the surrounding space are. We do not claim a new memory architecture. We claim an experimental record of where similarity lookup fails and a derivation of why, plus a test of whether an exact lookup can live inside the model.

Does it work on big models?

Not tested yet. Largest model so far is 4B parameters. That is exactly what funding buys first.

Why publish the failures?

Because the failures tell you which ideas to stop funding. Five of our ideas were killed on written criteria. The one that survived did so on replication, not on a single run.

What's the product?

[[OWNER: one sentence, or “Research stage; no product yet.”]]

What are the risks?

(1) Results are on synthetic data and small models. (2) The in-network fix is untested. (3) A larger lab could build the same thing inside a model, since it requires pretraining compute we do not have. (4) The exact-identity approach needs reliable entity extraction from raw text, which our current experiments partly assume.

10Contact

Talk to us

If you fund research, or run evaluations on memory, we would like to hear from you.