Everything in this series up to here has been about building an instrument and then pointing it at things. This one is the ledger. Over the last stretch I shipped a date filter, a literal-amount guard, a tax-year selector, two revisions of the grader, and three attempts at rebuilding table rows out of OCR geometry. Some of those moved the numbers. Some of them moved the numbers by changing the ruler. One of them cost thirty-seven hours of re-reading a library and ended up net positive by a margin I would not have predicted in either direction.
The question I want to answer honestly is not "did the system get better", which is the question a changelog answers. It is: for each lever, what did it measurably do, and how much of what it appears to have done is real.
The scoreboard, and why it is hard to read
Here is the difficulty. Between the numbers in post six and the numbers today, three things changed at once: the engine, the grader, and the text the engine reads. Any two of those moving in the same window makes an attribution argument, and all three moved.
- The engine gained a date filter, an amount-literal guard, and a tax-year selector.
- The grader went from v3 to v4, which is stricter in two specific places and therefore re-scores history.
- The text was re-read from the original page images three separate times, at reader versions v2, v3 and v4, because table-row reconstruction needs geometry that was never stored.
That last one is the awkward one, and it is the reason this post has a section about a fourteen-hour null result. A document app's corpus is normally the fixed thing you measure against. Ours was not fixed. It was being rewritten underneath the measurement, by a process whose whole purpose was to make it better.
So the only defensible frame is per-lever, against the baseline that lever was actually run against, and with the ruler named.
Lever one: the date filter. Real, and cheap.
Banking's single money violation for two months was an American Express question. The engine retrieved the right statement, ranked it first by a wide margin, and read the balance off a near-identical sibling from the following month that was sitting in the same eight-document window. The fix was not a better ranker — the ranker had already won. It was removing wrong-dated siblings from the context entirely, so that the model reading six nearly identical AmEx statements is not asked to pick one.
Measured effect: the answer came back correct and correctly cited, and banking's money bar went from FAIL to PASS and has stayed there through every subsequent re-measure, across two grader revisions and three re-reads of the library. Of everything in this ledger it is the lever with the cleanest attribution, and it is also the smallest change. That ordering is not a coincidence and it keeps happening.
Lever two: the amount-literal guard and the tax-year selector. Real, and then undone by something else.
Post six ends on the worst thing I have found in this project: a deterministic extractor reporting $809,111.74 of wages under a label promising an exact extracted value, for a figure that appears nowhere in the library, on the wrong year's form. Two fixes went in. A model-extracted amount must now appear literally in the cited document's own recognised text, and one amount failing that test voids the whole reply into an honest couldn't-read. The form selector now reads the year the form is for, out of the filename stem or a tax-year declaration in the body, rather than the year the envelope was mailed.
Measured effect at the time, same snapshot, nothing else moved: tax went from 15 correct with two money violations to 17 out of 17 with none. That is about as cleanly as a fix ever gets to own its own jump.
Measured effect at the time this section was written: tax reads 13/0/4 with four money violations. (Both the number and the diagnosis below have since been superseded — see lever six in the continuation at the foot of this post. The cause was not recognition variance; it was an unversioned change of render resolution, and tax now reads 15/0/2.)
Nothing in either fix regressed. Both guards still hold and both are still doing their job. What happened is that the pages underneath them were re-read, and a fresh Vision recognition of the same W-2 came back with different words — 3 Socisl security wages in the original, 3 Social securty wages in the re-read, and, more damagingly, the clean alternation of label and figure that the deterministic extractor depends on no longer there. 8,581 documents changed text between the first and second reader versions. The extractor now fails to find its anchors and declines, or reaches for something worse.
I want to be precise about what that means for the ledger, because there is a comfortable reading and it is wrong. The comfortable reading is that the fix worked and an unrelated regression landed on top of it, so the fix still counts. The true reading is that a fix which depends on the exact characters a recogniser emits is a fix with a dependency I had not written down, and the value of that lever is conditional on a thing I do not control and had assumed was stable. The lever moved the number. Something else moved it back. Both facts belong in the same row of the table.
Lever three: the grader. Moved the numbers without touching the engine.
Twice, now, the honest answer to "did the number move" has been "yes, and the engine had nothing to do with it".
The first time was the substring list in the correctness check, fitted by accident to the register of my own documents; widening it lifted the external benchmark by four and a half points with zero engine changes. The second time was the money bar itself, which turned out to have two holes: it only recognised a money question if the key carried cents, so every whole-dollar figure in the library — coverage limits, liability limits, anything a form rounds — was outside the safety bar entirely; and it graded what an answer mentioned rather than what it asserted, so a model that quoted the right figure on its way to the wrong one scored correct.
Measured effect of closing those: money-question denominators grew (tax thirteen to fifteen, insurance twelve to fourteen), several passes became failures, and every historical number had to be re-baselined. Under the old ruler the tax corpus would have looked better than it does. Under the old ruler an answer asserting $16.58 while quoting the $26.41 it had ruled out was a point in my favour.
The thing I would tell anyone keeping a scorecard for longer than a quarter: a grader revision is not a bug fix, it is a change of instrument, and it invalidates your entire history in the direction that flatters you. The only way I have found to keep it honest is to re-run every corpus on the new ruler and keep both sets of numbers in the record, clearly labelled, forever. Every table in this series carries the grader version for that reason.
Lever four: rebuilding table rows. Two not-shippable iterations, then one that shipped.
This is the expensive one, and the only one whose story is worth telling at length.
The defect: Apple's recogniser, handed a two-column billing table, emits all the labels and then all the values. Every token survives and the pairing a human reads is destroyed — ENDING BALANCE sitting flush against 148.65 on a bill whose amount due is 26.41. The fix is to take the word bounding boxes Vision gives us and we were throwing away, band them by vertical position, and weld each band back into a line. Because geometry is not stored, applying it to an existing library means re-running recognition over every document — which is why each attempt costs the better part of a day of wall time and why "just try it" was never available.
Attempt one shipped the rebuild with a gate that did not understand grids. The gate asked whether the rightmost cell of a band looked like a value. A W-2 is not a table but a grid — side-by-side boxes each printing a label above its own value — so banding by vertical position slices across unrelated boxes and produces all-value bands, which satisfy that test for entirely the wrong reason. Tax went from 100% to 76.47% with four money violations. Measured effect: the largest regression in this project's record, from a change that at ingestion demonstrably repaired the three bills it was written for.
Attempt two tightened the gate into a band-shape test — a band must run label-first, value-last, and at least half a page's bands must do so. Reader version bumped, full library re-read, fourteen hours and thirty-nine minutes, matrix re-run. Measured effect: nothing. Every corpus at baseline, one flip since characterised as generation variance. And underneath the null, two findings that mattered more than the matrix:
The W-2 was byte-identical across the two reader versions, so the gate change had never touched it and the tax regression was never the gate's — it belonged to re-OCR itself, per the 8,581-document diff above. Fourteen hours to learn that I had attributed a regression to the wrong cause, when a one-document text diff would have said so beforehand.
And the tightened gate over-declined. Across the 2,462 documents whose text changed, 1,503 lost welded label→value lines against 167 that gained. The ADT bill — the document the whole ticket exists for — regressed to exactly the split text the ticket was opened to repair. A fix that reintroduces its own defect on its own fixture is not a marginal call.
Attempt three kept the grid rejection and dropped the shaped-share requirement, declining a page on the actual grid signature (value-only bands in the majority) rather than on how cleanly its bands pair. And it added the piece that came out of attempt two's failure rather than out of the original design: a re-read replaces stored text only when it welds strictly more lines than what is already there.
Measured effect of that last rule alone, on a third full pass: churn fell from 2,462 changed documents to 976, of which 848 gained welds, adding 4,199 welded lines. Against the pre-rebuild text, the net is 569 documents better, 153 worse — the previous attempt's 1,503-document damage substantially repaired, without anyone having to identify which documents were harmed. The pass ran 22.5 hours with resident memory between 2.3 and 5.5 GB, against a previous pass that had climbed to 22 GB and needed an external watchdog cycling the app to finish at all.
The six-corpus matrix that closed it, on a fresh snapshot under the current grader:
Superseded — left standing as history; the record of note is the reader-v5 / grader-v6 table further down, and in the post next door.
| corpus (2026-09-05, reader v4, grader v4) | correct/partial/wrong | correctness | money questions | zero-wrong-money |
|---|---|---|---|---|
| banking | 17/0/1 | 94.44% | 13 | PASS |
| synthetic-irs | 21/0/2 | 91.30% | 18 | PASS |
| insurance | 14/0/3 | 82.35% | 14 | FAIL (one) |
| openrag (external benchmark) | 35/7/2 | 79.55% | 4 | PASS |
| tax | 13/0/4 | 76.47% | 15 | FAIL (four) |
| utilities | 13/3/3 | 68.42% | 8 | FAIL (one) |
Every corpus at or above the baseline it was handed, no new money violations, two flips both accounted for. Insurance's move from 76.47% is the one that looks like a win, and I am not banking it: the question that flipped is one we have measured flipping in both directions across runs on byte-identical text. Utilities' money bar went PASS to FAIL on the same wrong answer, because the v3 run's phrasing happened to end in a hedge the bar credited and this one did not — a hollow pass turning into an honest failure, which is the metric working rather than a regression.
So: did the levers move the numbers?
Superseded — six rows, left standing as history. The current ledger is the fourteen-row table at the foot of this post.
| lever | measured effect | how much of it is real |
|---|---|---|
| date filter | banking money bar FAIL → PASS, held through everything since | fully real, cleanest attribution in the ledger |
| amount-literal guard + tax-year selector | tax 15/2 violations → 17/17/none, same snapshot | real, then buried by re-OCR variance; both guards still hold |
| grader v3 → v4 | denominators grew, several passes became failures | real as measurement; zero engine change |
| row rebuild, attempt 1 | tax 100% → 76.47%, four money violations | real regression; ingestion wins real but outweighed |
| row rebuild, attempt 2 | matrix flat; 1,503 documents lost welds | net negative; its value was the refutation, not the code |
| row rebuild, attempt 3 (variance guard) | 848 documents gained welds; net 569 better / 153 worse vs pre-rebuild | real, and the guard is the transferable part |
Two of the six rows are levers that failed, and one is a lever that improved the ruler rather than the engine. That is a worse-looking ledger than the one I would have written from memory, and the reason it is worth keeping is the shape of the two failures.
Both not-shippable attempts failed the same way: they were correct about the diagnosis and wrong about the instrument. The grid diagnosis was right both times. The rightmost-cell test and then the fifty-percent shaped-share threshold were both the wrong way to express it, and both were only findable by running the change across the whole library rather than across the two documents that motivated it. Every one of the three fixes post six describes as measured-and-discarded failed in exactly that register too — repairing the fixture, breaking the corpus.
And the single most useful thing to come out of thirty-seven hours of re-reading is not a gate at all. It is the rule that a migration must earn its overwrite: propose new text, and accept it only if it is measurably better than what it replaces, on the axis the migration exists to improve. Both bad attempts would have been mostly harmless under that rule. It is a small rule, it came out of a failure rather than a design, and it is the part of this I expect to still be using when none of the rest is left.
The last word belongs to the document that started it. The ADT bill's re-read did not weld more than the text it would have replaced, so the guard declined it, correctly by its own rule, and the bill still reads with its ending balance on a separate line from its label. But the answer is wrong for a different reason: the engine cites a sibling statement from later that spring. Two full library re-reads, a repaired gate, a new acceptance rule — and the headline failure that motivated all of it turns out to be a retrieval problem in an OCR costume. That is the next lever, and I do not yet know whether it will move the number either.
The ledger, continued
That last paragraph asked whether the retrieval lever would move the number. It did, and then eight more levers went through the same process: three moved numbers, one moved a number by making the ruler honest, one changed real product behaviour the scoreboard cannot credit, and three were measured and thrown away. Same rules as above: per lever, against the baseline it was actually run against, with the ruler named.
Lever five: sibling date parsing. Real, and it was a two-line regex.
The ADT bill was retrieving a May 2012 statement when asked about February. The date filter from lever one was supposed to make that impossible, and the reason it did not is not in the filter at all. Filenames in this corpus write dates unpadded — 5_7_2012 — and the parser wanted zero-padded components, so those files parsed to no date at all in 697 documents and to the wrong date in 97 more. A document with no date cannot be excluded by a date filter; it is invisible to it. The siblings were not out-ranking the right bill, they were walking past the gate.
Measured effect: util-adt-amount answers $26.41 citing the right statement, and utilities went 13/3/3 68.42% to 14/3/2 73.68% with the money bar at PASS. The headline failure of the entire row-rebuild arc — three attempts, thirty-seven hours of re-reading — was closed by fixing a regex that parses filenames, which is either a satisfying ending or a humiliating one depending on how attached you are to the previous six weeks.
Lever six: best-of-two DPI. Real, and it repaired the thing three re-reads could not.
The four tax money violations in the table above were attributed, in this post, to re-OCR recognition variance. That attribution was wrong, and the correction is the most useful measurement in this stretch.
Eight fresh ingests of the same W-2 into an empty library produce byte-identical text and byte-identical structured output, eight times out of eight. Vision is deterministic at a fixed DPI on a fixed OS. What was not fixed was the DPI: a commit in late July had moved page rendering from 200 to 300 without bumping the reader version, and the 300-DPI read of that W-2 recognises individual glyphs slightly better while losing the sidecar table that carries 1 Wages, tips, other compensation\n953974.49 and 2 Federal Income tax withheld\n241582.46. The deterministic extractor needs those cells; without them it declines. Lexical quality barely registers the loss — 0.629 dictionary-word rate at 200 DPI against 0.594 at 300 — so any word-level quality score would have called it a tie or preferred the worse read.
The library-wide census says there is no right answer to pick globally: 308 documents have a sidecar table only at 200 DPI, 772 only at 300, and 391 of the 6,194 with tables at both resolutions have fewer at 300. So reader v5 renders a challenger at 200 DPI when the 300-DPI read shows loss signals and keeps whichever read wins an answer-key-free ordering — tables, then welded lines, then adjacency, ties to the incumbent.
Measured effect, full pass, durable counters: 15,392 of 15,398 documents stamped, 2,908 re-reads accepted, 8,484 rejected, twenty-seven hours wall, resident memory 0.7–3.7 GB. Tax 13/0/4 with three money violations → 15/0/2 88.24% with one (the grader-v4 table above counts four; by the time this pass ran, one had become an honest abstention rather than a stated figure, on an unrelated extractor fix), both W-2 questions answering as exact extracted values; insurance 14/0/3 → 15/0/2 with its standing violation cleared. Two corpora paid a point in the same matrix, and both are accounted for elsewhere in this ledger: banking on the grader's $35-versus-35.00 format hole, closed by lever seven, and utilities on a tied window, closed by lever eight. Note the rejection count against lever four's variance guard: three refusals for every acceptance, on a pass that still produced the largest single recovery in the ledger. And note what the guard needed in order not to sabotage it — the acceptance metric had to be taught to count tables before welds, because the W-2's better read wins on tables and only ties on welds, and without that change the migration would have rejected precisely the 308 documents it existed to rescue.
Lever seven: grader v5 and v6. The ruler again, twice, in the unflattering direction.
v5 fixed a format hole: "a late fee of up to $35" against a key of 35.00 graded WRONG on a substring test, and the money bar counted that asserted-correct figure as a violation. Money needles that name only an amount now compare as amounts in integer cents. Offline re-grade of the standing reports with answers held identical: exactly one flip in 183 runs, and it is the filed defect. Banking back to 17/0/1 94.44%, money PASS.
v6 fixed a stance hole, and this one costs points rather than granting them. The bar had been grading what an answer mentioned rather than what it finally asserted: a wrong figure followed by a hedge read as an honest decline, and a right figure derived inside a hypothetical that concludes "cannot be determined" satisfied its needle out of the retracted reasoning. The regression set was built before the mechanism, mined from 1,314 runs on disk, with hedged-but-correct answers pinned as an explicit false-positive budget. Measured over that history: three grade flips, all of them the filed defect, and money-bar violations +19 / −0, every new one landing on an answer already graded wrong. One extension — applying the same decline logic to the negative questions — was measured, found to flip three honest declines, and rejected on the evidence.
Measured effect on the matrix: zero grade changes, one new money violation — a utilities answer stating $54.96 against a key of $133.89 and then hedging about a due date, which under v5 the hedge had excused. Same character as the hollow ADT pass described earlier in this post: not a corpus getting worse, an instrument getting honest.
Lever eight: deterministic tie resolution. Zero measured points, and I would ship it again.
Several questions in this project have been flapping for months — correct on one run, declining on the next, on text that had not changed — and every one of them was filed under generation nondeterminism.
The decoder was never involved. Eval generation is greedy argmax at temperature zero. Thirteen unpinned runs against six runs with Swift's hash seed pinned: unpinned, one question produced seven distinct answer texts in thirteen runs and another flipped its grade on two of thirteen; pinned, every suspect collapses to one byte-identical text, and the model-free retrieval reports diverge across unpinned runs too, which puts the cause upstream of generation entirely.
The mechanism: retrieval sums a term vector over a Set and a dot product over a Dictionary, Swift seeds hash order per process, float addition is not associative, and chunks meant to tie exactly — one question has a six-way exact tie at 0.478 among six near-identical gas bills — end up an ulp apart, so the (documentID, chunkIndex) tie-break written for that case never fires and the ordering is decided by float noise. Different chunk order, different prompt, different answer from a deterministic decoder. The dice were in the arithmetic.
Fix: accumulate in sorted token order, so scores are a pure function of content and genuine ties are bit-equal. Cost: 62.5 ms per query sorted against 68.5 ms unsorted, best of three over a thousand chunks and twenty queries — inside the noise, and slightly favourable, because sorted iteration improves dictionary locality.
Measured effect on the matrix: zero. All 138 answer texts byte-identical. That is not evidence the change does nothing, and it must not be reported as a win either: pinning the seed had already frozen one arbitrary draw for the harness, and this draw and the sorted-order draw happen to coincide. The evidence lives in unit tests (tied scores bit-equal, a six-way tie resolving to the tie-break under every insertion order, non-tied rankings keeping both order and bit-exact scores). What it buys is not points. It is that the shipping app, which runs unpinned, stops resolving ties by the process's hash seed, so the same question over the same library gives the same answer twice.
Lever nine: value-chunk swap inside the per-document quota. Real, one point, one flip.
The last fixable money violation was the NW Natural bill: right document, rank one, wrong answer. The window dump shows why. The per-document cap gives that bill three slots, spent in pure score order on the P.O. box page, a utility-commission notice and a "gas meters record the volume" explainer; the chunk carrying PLEASE PAY THIS AMOUNT … 1/03/2012 $133.89 sits fourth, 0.0029 below the third, because every chunk of one bill shares that bill's vocabulary and the lexical cosine spreads all seven of them over about 0.005. Noise was choosing the window.
Raising the cap is the obvious move and it is wrong — the cap exists to stop one document crowding out its rivals. The change is about which chunks spend the quota: on a money question, if a document's slots contain no chunk carrying a money figure and a lower-ranked chunk of the same document does, its last slot is swapped for the highest-ranked qualifying one. Single slot, positive-evidence only, and explicitly not a reorder: on a bill, "carries a money figure" is satisfied by the previous balance and the late-payment charge as readily as by the amount due, so a value-first reorder would evict the context that disambiguates which figure is the answer and produce a confident wrong number — the failure the whole ticket exists to prevent.
Measured effect against the pinned baseline, 138 question-runs: 134 answer texts byte-identical, three text-only changes, one grade flip. The flip is the NW Natural question, WRONG → CORRECT at $133.89 with the right citation; utilities 13/3/3 68.42% → 14/3/2 73.68%, money bar FAIL → PASS. One text-only change is the documented over-trigger behaving as designed: an external-benchmark maths question whose LaTeX $\psi(x)$ reads as a money cue, moving one chunk slot inside a document already in the window and leaving the answer identical.
Three levers that did not pay, and were not shipped
Region-scoped row rebuild. The idea was to assemble columnar regions rather than accept or decline whole pages, which is a better-shaped mechanism than an all-or-nothing gate and was designed off the ADT class. Benched on a 52-document cohort before any library touch: 2 wins against 3 losses, on a no-change noise floor of 1/1, firing on 2 of 107 eligible pages, and producing nothing at all on the class it was designed from. Closed without shipping. The diagnosis it came from is still true; this repair for it is not.
Query-entity refinement. Entities are already extracted and cached on the document side, so using them as a gazetteer to refine which documents a query's entity prior boosts costs nothing at query time. Measured pinned, six corpora, boost-only and fail-open: banking 17/0/1 → 17/0/1, tax 15/0/2 → 15/0/2 with nine of fifteen questions' retrieval moving and zero grade flips, insurance 14/0/3 → 14/0/3 with two offsetting flips, utilities 13/3/3 → 13/3/3 with five of sixteen moving and zero flips, and the two ingested benchmark corpora inert by construction. It moves retrieval a great deal and grades not at all, and the one insurance flip in its favour is a known flapper on byte-identical retrieval while the flip against it is genuinely attributable. Reverted in full, no commit. A lever that churns the window without moving the score is a lever that is spending risk for nothing.
The VLM-OCR spike. A vision language model as an arbiter or fallback for the recogniser, run on the field bench against Vision page for page. The expected disqualifying finding did not appear: across ten pages GLM-OCR invented zero money figures, and on the one page where the two engines disagreed about money it was Vision that was wrong — the VLM read all seven figures correctly and Vision got one of seven, fabricating two amounts that are not on the page. The actual disqualifier is worse and lives where the designed safeguard cannot see it. Asked to read the W-2 as a table — the mode an arbiter would use — it attached 108178.23, which is Box 3, to boxes 7, 10 and 11, which are blank on the form, and scored a perfect money-token diff while doing so, because every fabricated value is a verbatim token from elsewhere on the same page. "Values must appear verbatim in the Vision read" does not verify a pairing. It also silently omits: 4% of the W-2's characters, 2% of one receipt, and all three ADT statements' money tables skipped entirely while the marketing prose came through. And it costs 18.6 seconds a page against Vision's 0.57 — a 33× multiplier. Arbiter-worthy in a narrow band, not viable as a fallback, nothing shipped.
The ledger as it stands
Where all of that leaves the scoreboard — reader v5, grader v6, snapshot 20260907-144447, pinned (SWIFT_DETERMINISTIC_HASHING=1), --mode topDoc3 --datefilter 1 --semweight 0, serial, app quit before the run:
| corpus (2026-09-07/08, reader v5, grader v6) | correct/partial/wrong | correctness | money questions | zero-wrong-money |
|---|---|---|---|---|
| banking | 17/0/1 | 94.44% | 14 | PASS |
| tax | 16/0/1 | 94.12% | 15 | PASS |
| synthetic-irs | 21/0/2 | 91.30% | 18 | PASS |
| insurance | 14/0/3 | 82.35% | 14 | FAIL (one) |
| openrag (external benchmark) | 35/7/2 | 79.55% | 4 | PASS |
| utilities | 14/3/2 | 73.68% | 8 | PASS |
And the per-lever ledger, which is the point of this post:
| lever | measured effect | how much of it is real |
|---|---|---|
| date filter | banking money bar FAIL → PASS, held through everything since | fully real, cleanest attribution in the ledger |
| amount-literal guard + tax-year selector | tax 15/2 violations → 17/17/none, same snapshot | real, then buried by a raster change; both guards still hold |
| grader v3 → v4 | denominators grew, several passes became failures | real as measurement; zero engine change |
| row rebuild, attempt 1 | tax 100% → 76.47%, four money violations | real regression; ingestion wins real but outweighed |
| row rebuild, attempt 2 | matrix flat; 1,503 documents lost welds | net negative; its value was the refutation, not the code |
| row rebuild, attempt 3 (variance guard) | 848 documents gained welds; net 569 better / 153 worse | real, and the guard is the transferable part |
| sibling date parsing | utilities 13/3/3 → 14/3/2, money FAIL → PASS; 794 documents had been dateless or misdated | fully real, and it closed the arc's headline failure |
| best-of-two DPI (reader v5) | tax 76.47% → 88.24%, money FAIL(3) → FAIL(1); insurance +1, its standing violation cleared | fully real; the largest recovery in the ledger, at the actual cause |
| grader v5 → v6 | +1 correct on the format hole; zero grade changes and one hidden violation surfaced on the stance hole | real as measurement; the v6 half costs points on purpose |
| deterministic tie resolution | zero matrix deltas; 138/138 texts unchanged | real as product behaviour, unmeasurable as score; evidenced by tests, not the table |
| value-chunk swap | utilities 14/3/2, money FAIL → PASS; 134/138 texts byte-identical | real, one point, and the design's over-trigger cost exactly what it was budgeted to |
| region-scoped rebuild | 2 wins / 3 losses on a 1/1 noise floor; fired on 2 of 107 pages | refuted by its own bench; never shipped |
| query-entity refinement | retrieval moved on a large share of tax, insurance and utilities questions; zero net grades | measured not paying; reverted, no commit |
| VLM-OCR arbiter/fallback | zero invented figures, but a fabricated label→value pairing the verification design cannot detect; 33× slower | refuted as designed; nothing shipped |
| reconciliation-row extraction | tax 15/0/2 → 16/0/1, money FAIL(1) → PASS; one grade flip in 138 runs | fully real; a brokerage summary's pre-adjustment column was being summed as if reportable — the row now has to prove itself (three figures, A − B = C) before its final column is read |
Fifteen rows. Seven moved a score in the direction I wanted, one moved it decisively the wrong way, two moved the ruler rather than the engine, four failed outright and were never shipped, and one — the tie fix — is real product behaviour the scoreboard is structurally unable to credit, which is its own kind of honest. The reconciliation-row fix closed the last tax money violation, which leaves exactly one red cell on the board: an insurance question that flips on a knife-edge the stance grader now at least calls honestly.
The pattern across the failures did not change: correct about the diagnosis, wrong about the instrument. Region scoping had the right read of the ADT class and a detector that fires on two pages in a hundred. The VLM spike had the right worry about hallucination and a check aimed at the wrong half of it — digits, when the danger was pairings. Query-entity had a real signal and no evidence it converts. All three were only findable by running the change over the whole corpus rather than over the documents that motivated it, which is the same sentence this post has now had to write five times.
And the correction I would put at the top if I were rewriting the whole ledger: for months this project attributed a structural regression to "recognition variance" and a set of flapping answers to "generation nondeterminism", and both were deterministic mechanisms with names. One was a render resolution that changed without a version bump. The other was float addition over a hash-ordered set. Neither was a model doing something inscrutable, and in both cases the experiment that settled it — ingest the same page eight times; run the same question thirteen times with the hash seed pinned — costs an afternoon and was available the entire time. The expensive part was never the measurement. It was how long I was willing to keep an explanation that did not require one.