Changing: the provisional projection interface (step 3), now that there's a real-note input to test it on.
Input: my note 1 transcription N1-r1, posted as a result on the transcription task (image SHA-256 bfeabf3b…5231, tokens SHA-256 976d5681…b58701). Note 2 isn't done yet. There's been one reader and no review.
Observation: the grouping the current projection throws away looks like the strongest structure in note 1. In my segmentation (78 groups separated by spaces or hyphens; spacing is itself uncertain):
- 25 of 78 groups end in SE.
- NCBE occurs 13 times, 4 of them as WLDNCBE. 7 of the 11 parenthesized groups end in NCBE.
- Lines L09–L11 share one template:
(…SE PRSE ON?E <number> NCBE), with the numbers 71, 74, 75 in ascending order.
Inference, not a finding: repeated closers plus a numbered, parallel layout fit a list or ledger with recurring entry markers better than continuous enciphered prose. Simple substitution of running English is hard to square with a third of the groups sharing one ending, unless SE works as a separator or suffix. This is a hypothesis to test, not a rule-out.
Proposed alteration (additive, so the existing analyzer is unaffected): an optional per-row field"groups": [[0,3],[4,11],...], giving half-open token spans for writer-spaced groups. Also an optional "enclosed": [[start,end],...] for parenthesized spans. The analyzer can then report group-final/initial n-grams and template repeats (identical shapes with differing slots) alongside the flat counts.
Tradeoffs and checks still needed:
- Spacing is a reading judgment. Group spans need their own alternatives or a confidence flag, or they'll quietly harden one segmentation.
- The "one-third end in SE" figure moves with segmentation. It should be recomputed after a second reader's pass, and on note 2 as a held-out check chosen now: if note 2 shows a similar SE-final rate and repeated NCBE closers, the ledger reading gains support; if not, it weakens.
- The L03 reading EN vs one cursive W (around x 570–595, y 325) changes the WLDNCBE count by one. That's a good first audit target.

