← board

The fifteen reversed rows are the real goal-1 backlog, not the emulator

This ticket exists because its predecessor answered its own question and the answer moved the work somewhere else. frankuser's framing, which is the sentence to keep: upgrading borg's emulator "does not deliver a green tier — it converts 'never green, and the verdict carries no information' into 'sometimes green, and a red means something'. The deliverable is INFORMATION, not GREEN."

What is already established, so nobody re-derives it

The two jobs, in order

1. Decide whether there are fifteen rows. Re-run tools/tstate_toolchain_reversals.py after borg's qemu moves. Strike every row that clears, with the date; keep every row that stays red, because that one is in our own tree. Do not start fixing rows before this — nine of the eleven checkable ones already pass on a second 10.2.1 host, so some fraction of this list is a fact about a retired machine.

2. Install the corpora on the release-candidate host. skip_holes == 0 is the goal-1 bar and a 40-job hole is not it. Note the recursion, which is worth a grin and a check: tools/install_lib_candidates.sh is itself one of the skipped jobs, so the tool that would close the hole is inside it.

The trap, named in advance

A row on this list that reddens after the upgrade will look like a new regression, because the upgrade and any compiler work will land in the same week. The predecessor ticket pre-registers that expectation; read it before attributing anything, and prefer git log -S'<the row>' -- <the skip or expected file> over an inference from timing.

REJECTED WITHIN THE HOUR, BY ITS OWN AUTHOR'S NEXT MEASUREMENT — THE PREMISE IS FALSE

rejected/ and not low-prio/, because the report is WRONG rather than unimportant: there is no backlog of fifteen rows. Every one of the fourteen reversed rows that had any reds had already ENDED before seven's last report — none was red in its final 20 reports.

Why the premise looked true. I computed each row's red rate over a time window and treated it as a per-run probability. It is not. Those reds are single past EPISODES:

row reds longest consecutive run last red
size_canary.py 189 158 in a row 2026-09-10
test_libwriteln_parity.pas 86 86 in a row, one day 2026-09-06
test_emit_obj.pas 65 13 2026-09-06

A row that is red for four days and then fixed reads as "52.4% red" over a seven-day window, and 52.4% invites an inference about the next run that the data cannot support. So P(all nine pass) ≈ 13% was not a weak result, it was a meaningless one — there was no per-run probability for it to be about.

AND THIS IS THE SECOND TIME TONIGHT THE SAME QUESTION WENT UNASKED. frankuser asked me hours earlier, of a different table, "are the 16 greens CLUSTERED or SPREAD?", and the answer dissolved a 2.6x I had published. I had the lesson in hand, wrote a new table, and did not apply it. Promoted to the playbook for that reason and not for the arithmetic.

What survives and where it went: the corpora half, to task-t-a-release-grade-full-green-needs-the-corpora-installed-skip-holes-is-forty. What this changes for the owner escalation: the predicted cost of upgrading borg's qemu from these rows is essentially zero, so the trade earlier reported has no cost side. frankuser guessed the cost would collapse; it collapsed further than either of us measured, and for a different reason than the one I checked.