Similarity breaks as memory grows. Exact identity does not.
Share of questions about things that were never stored that receive a confident wrong answer, as the memory grows from 1,000 to 99,000 entities.
Tiny transformer trained from scratch on a synthetic world. The comparison is architecture-level, not a product benchmark.
Even with a perfect key, similarity lookup answers 46–50% of absent questions confidently and wrongly at 99,000 entities. Exact identity answers 0%.
Confident wrong answers on never-stored questions (%)| Memory size (entities) | Similarity, perfect key (3 seeds) | Exact identity (6 seeds) |
|---|
| 1,000 | 0.6% | 0.0% |
|---|
| 10,000 | 20.6% | 0.0% |
|---|
| 99,000 | 48.3% | 0.0% |
|---|
Per-seed similarity values at 99k: 49.0 / 49.9 / 45.9.
Source: experiments/M11b_binding/results/eval_seed{0,1,2}.json, eval_seed{1,2}_i.json, eval_seed{3,4,5}_iii.json
In plain words
Imagine finding a book by “which cover looks most like what I want.” With ten books that works. With a hundred thousand, something always looks similar enough, even when your book isn't there. A library card number either exists or it doesn't.
Disclosure
Exact-identity seed 2 shows an accuracy dip at 10k/99k on one question type (true/false “is X's value V”), not seen on seeds 0, 1, 3, 4, 5. Cause unconfirmed.