Experiment ledger
Each experiment starts with a written hypothesis, a prediction and a kill criterion. A killed experiment stays in the repository with its reason. A number without its result file is not a result.
Every experiment, including the ones that failed.
- R01 · M08Replicated
Identity-keyed store vs chunk retrieval
Identity store 100% in every run. Gap over chunk retrieval: templated +42.4 (Gemma) / +38.9 (Qwen), natural text +51.4 / +48.6 points. Stale-belief errors 0% vs 64–69%.
Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt
- R01 · M01bReplicated
Structural symbol index on a code repository
96.5% at 0.36% of a 2,410-file repository's tokens (mean of two models). BM25: 88–89%. Definition-only: 49.6%.
Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt
- M11bReplicated
Exact identity vs similarity as memory grows
Absent questions answered confidently wrong at 99,000 entities: similarity with a perfect key 48.3% (mean of 3 seeds), exact identity 0.0% (6 seeds).
- P1Pilot · indicative
Summary vs full episode
One episode, 31 yes/no questions. 48-token summary: 39% confidently wrong. Full 337-token episode: 0%. Every absent-fact question had the answer “no”, which inflates the figure.
Source: history/RESULTS.md (P1)
- M04Killed
Rewrite facts into joint documents so the model learns multi-hop links
Chance level: 0.045 vs 0.040 for the control. The facts were already known, so there was nothing to learn from.
Source: experiments/M04_cooccurrence_consolidation/results/, history/LEDGER.md
- M11Killed
A learned continuous key inside the network
Accuracy 100% at 10 entities, 92.4% at 10,000, 70.3% at 99,000. The absence gate was calibrated at small scale and opens at large scale.
Source: history/LEDGER.md (M11 row)
- M13Killed
A trained controller that decides “answer / reload / ask the model”
+3.4 and +3.7 points over the model's own choices against a +10 requirement. A hand-written rule scored as high (88.3%).
Source: experiments/M13_calibrated_controller/results/, history/LEDGER.md
- B02Killed
The model compiles repeated answers into programs it no longer needs to ask about
A 4B model could not write the programs: 13.8% of compile attempts were usable, and answer accuracy was 11.9%. Parser, token limit and sandbox were ruled out.
Source: experiments/B02_compilation_curve/results/diag_b02_qwen3.5-4b-hf*.{log,jsonl}
- M07Killed
Write only surprising facts to memory
+1 point overall. Killed on its criterion.
Source: history/LEDGER.md
- M14Not run
Calibrated read: a size-aware threshold and a wider key
Hypothesis — not yet tested. Nothing is claimed.
- M14Not run
Exact two-level code inside the network
Hypothesis — not yet tested. Nothing is claimed.
- M12Not run
Integration on a frozen Qwen3.5-2B
Hypothesis — not yet tested. Nothing is claimed.
- B01Not run
Written, not run. Details in the repository.
Hypothesis — not yet tested. Nothing is claimed.
- B03Not run
Written, not run. Details in the repository.
Hypothesis — not yet tested. Nothing is claimed.