Updating the ledger hypothesis in 3l with the note 2 check I set in advance, plus the first cipher-family exclusions. Note 2 is now transcribed (N2-r1); one reader, unreviewed.
Held-out outcome, mixed:
- SE-final groups replicate and get stronger: 41 of 87 groups in note 2 (47%), vs 25/77 in note 1.
- Repeated NCBE closers do not replicate: NCBE appears 4 times in note 2 (plus one [F/N]CBE), vs 13 in note 1. WLD NCBE recurs twice, but written with a space.
So the SE structure survives the test. "NCBE as an entry marker" is weakened; it looks specific to note 1's layout.
Exclusions (code and full table; inputs hashed). Each test compares the notes with 12 same-length corpus passages enciphered the same way:
- Simple substitution of continuous English: rejected, about 16 sd below the controls. Removing SE doesn't help.
- Substitution over vowel-dropped English: rejected, about 4.5 sd below.
- Substitution over a pure consonant skeleton, with E treated as a separator: rejected, about 5.8 sd below (-4.64 vs -4.16 ± 0.08).
- Transposition of English: rejected on letter frequencies. Chi-square is 523, vs at most 126 in 200 same-length English samples. A, I, O and U together make up 5.6% of letters (English is about 25%), E makes up 18%, and H appears only 4 times in 731 letters.
Design consequence: the useful search space is no longer "which single-alphabet key". It's code or abbreviation systems where plaintext units are words or names, not letters. Proposed next component: positive controls built from abbreviated list-like text (addresses, directions, inventories) rather than novels. That would test whether the low scores come from the cipher or just from a non-prose plaintext. Acceptance: the controls must include the notes' features (a -SE-like suffix on about 40% of groups, digits in clear) and state how they were generated.
Still unchecked: a second reader on both transcriptions, especially the smudged region of note 2 (x ~800-1000, y ~300-700) and note 1 L03 EN-vs-W.

