Skip to content
01

Experiment ledger

Each experiment starts with a written hypothesis, a prediction and a kill criterion. A killed experiment stays in the repository with its reason. A number without its result file is not a result.

Every experiment, including the ones that failed.

  1. R01 · M08Replicated

    Identity-keyed store vs chunk retrieval

    Identity store 100% in every run. Gap over chunk retrieval: templated +42.4 (Gemma) / +38.9 (Qwen), natural text +51.4 / +48.6 points. Stale-belief errors 0% vs 64–69%.

    Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt

  2. R01 · M01bReplicated

    Structural symbol index on a code repository

    96.5% at 0.36% of a 2,410-file repository's tokens (mean of two models). BM25: 88–89%. Definition-only: 49.6%.

    Source: experiments/R01_replicate_M08_M01b/results/SUMMARY.txt

  3. M11bReplicated

    Exact identity vs similarity as memory grows

    Absent questions answered confidently wrong at 99,000 entities: similarity with a perfect key 48.3% (mean of 3 seeds), exact identity 0.0% (6 seeds).

    Source: experiments/M11b_binding/results/eval_seed{0,1,2}.json, eval_seed{1,2}_i.json, eval_seed{3,4,5}_iii.json

  4. P1Pilot · indicative

    Summary vs full episode

    One episode, 31 yes/no questions. 48-token summary: 39% confidently wrong. Full 337-token episode: 0%. Every absent-fact question had the answer “no”, which inflates the figure.

    Source: history/RESULTS.md (P1)

  5. M04Killed

    Rewrite facts into joint documents so the model learns multi-hop links

    Chance level: 0.045 vs 0.040 for the control. The facts were already known, so there was nothing to learn from.

    Source: experiments/M04_cooccurrence_consolidation/results/, history/LEDGER.md

  6. M11Killed

    A learned continuous key inside the network

    Accuracy 100% at 10 entities, 92.4% at 10,000, 70.3% at 99,000. The absence gate was calibrated at small scale and opens at large scale.

    Source: history/LEDGER.md (M11 row)

  7. M13Killed

    A trained controller that decides “answer / reload / ask the model”

    +3.4 and +3.7 points over the model's own choices against a +10 requirement. A hand-written rule scored as high (88.3%).

    Source: experiments/M13_calibrated_controller/results/, history/LEDGER.md

  8. B02Killed

    The model compiles repeated answers into programs it no longer needs to ask about

    A 4B model could not write the programs: 13.8% of compile attempts were usable, and answer accuracy was 11.9%. Parser, token limit and sandbox were ruled out.

    Source: experiments/B02_compilation_curve/results/diag_b02_qwen3.5-4b-hf*.{log,jsonl}

  9. M07Killed

    Write only surprising facts to memory

    +1 point overall. Killed on its criterion.

    Source: history/LEDGER.md

  10. M14Not run

    Calibrated read: a size-aware threshold and a wider key

    Hypothesis — not yet tested. Nothing is claimed.

  11. M14Not run

    Exact two-level code inside the network

    Hypothesis — not yet tested. Nothing is claimed.

  12. M12Not run

    Integration on a frozen Qwen3.5-2B

    Hypothesis — not yet tested. Nothing is claimed.

  13. B01Not run

    Written, not run. Details in the repository.

    Hypothesis — not yet tested. Nothing is claimed.

  14. B03Not run

    Written, not run. Details in the repository.

    Hypothesis — not yet tested. Nothing is claimed.