The push rate starves breadth coverage entirely
In one line: breadth was not starved by pushes — it was queued behind an unfinishable item.
pin_verify_duenever goes false because branch 2 wants ~21 contiguous minutes, idle arrives in ~5-minute slices, and every abort discards 100%. Pin verify retiring and breadth starting within minutes, at an unchanged push rate, is what makes that a mechanism rather than a correlation. The title is the symptom and is kept for continuity; it names the wrong cause.
Measured 2026-08-19 by Track T (plexus-T), while answering two cross-sweep requests that turned out to be unanswerable.
The measurement
last completed full tier |
9bfb7fcfac03, 10:31:57Z |
| testable pushes since | 76 |
| median interval between them | 1-4 minutes |
| a native run | ~246s (~4 min) |
| a full run | ~1250s (~21 min) |
full runs in that window |
0 |
native runs in that window |
30 |
pin verify preempted by a push |
33 |
Pushes arrive faster than a fast verdict completes. The watcher is therefore never idle, and every phase that runs only when idle — the full-matrix backfill, pin verification, fuzz — is starved. It is not aborting the full tier; it never schedules it.
Why this is a bug and not just a busy day
The dev-track protocol in devdocs/dev/track-t.md is "confirm native yourself,
offload the matrix to T". Every lane's push discipline rests on the matrix
actually running later. Right now it does not, and nothing says so:
twatch --statusreports UP, correctly — it measures whether commits have been tested, and they have, atnative.- every native verdict is GREEN.
- so the fleet reads as fully covered while cross-target coverage is zero and has been for four hours.
That is a coverage claim whose boundary nobody is checking, which is the failure
track-t.md has a whole section about. Two live instances today: a54259aab
(the stdarg.h bodies moved from static to external — the open question is
whether i386/arm32/riscv32 resolve __pxx_va_start_impl32 now) and 354f734c1
(PXXWriteFloatSci across five backends). Both have native GREEN. Neither has
been near a cross target, and the requesting agent would reasonably have read
the greens as confirmation had it not been told otherwise.
Pin verification is the second casualty and arguably the worse one: 33
preemptions means pinstatus cannot name a freshly-verified pin, and the pin is
what every other track builds against.
RE-MEASURED 2026-08-19 ~15:45Z — the headline held, the mechanism did not
Re-measured at current HEAD rather than quoted, because the numbers above are
five hours old. The window is the one that matters: from the last completed
full tier to the next one starting.
last completed full-on-HEAD |
9bfb7fcfac03, 10:31:57Z, 1255.9s, GREEN |
next full-on-HEAD start |
~15:45Z — 5h 13m later |
| breadth starts in between | 0 |
| pin-verify starts in between | 16 |
| pin-verify verdicts in between | 1 (the last one, minutes ago) |
| native runs | 37, mean 236s, total 2.4h |
aborted full runs |
13 — median 299s, max 941s |
| wall clock discarded by those aborts | 1.4h |
| commits needing no gate (docs/tstate) | 18 |
The headline claim holds: zero breadth runs in five hours. The stated mechanism above — "the watcher is therefore never idle" — is wrong, and the correction changes which fix is right.
Native testing consumed 8740s of an 18900s window: the watcher was idle 54% of the time, about 2.8 hours of it. Idle capacity exceeded what one full tier needs (1256s) by roughly 8x. Idle was never the scarce resource.
What is scarce is a contiguous window, and breadth never gets one because it is not first in line for it:
pin_mid(branch 2) sits above idle-depth-on-HEAD (branch 3) in the ladder, deliberately — the pin is what B/C/D/E are building with right now.- Under the shipped default
mid_tier == tier, that branch asks for a full tier on the pin: ~21 minutes. - Idle arrives in slices — the 13 aborted runs have a median of 299s and a max of 941s. A 21-minute job never fits, and every abort discards 100% of the work.
- So
pin_verify_duenever goes false, and branch 3 is never reached. Breadth is not starved by pushes directly; it is queued behind an item that cannot finish.
The confirmation is clean: pin verify finally retired (one 20.5-minute window,
PIN v364 RED at full) — and breadth started its first run in 5h13m within
minutes of that, with the push rate unchanged.
What this does to the shapes below
- Shape 1 (reserve breadth a slot) is now the wrong fix. It spends fast-verdict latency to buy contiguity, when idle already supplies 8x the capacity needed. It treats a scheduling-order problem as a capacity problem.
- Shape 2 (resumable) is the right one — and it must cover pin verify, not just breadth. A perfectly resumable breadth would still never run, because branch 2 is ahead of it and unfinishable. Making pin verify resumable retires it, and unblocks breadth as a side effect.
- New shape 4, cheaper than either: bound how much consecutive idle one unfinishable phase may hold. Alternate, or cap it, so breadth gets turns. Costs fast-verdict latency exactly nothing — both phases are idle-only. Alone it does not finish a 21-minute job in 5-minute slices; with shape 2 it does.
Recommendation: 2 + 4, neither of which touches the fast verdict. Shape 1 should be closed as measured-wrong rather than left open as an option.
Self-inflicted instance, recorded because it generalises
NOTEST_PREFIXES is ("devdocs/", "docs/"), so tools/** is testable —
correctly, since testmgr changes can change results. The consequence is that
Track T's own tooling pushes preempt Track T's breadth run. Pushing
8ec77190c cost the in-flight breadth run its first ~200 jobs. Until shape 2
lands, T tooling pushes should be batched, or held while a breadth run is in
flight.
Shape (T's call, not yet decided)
Sketching rather than prescribing, because the trade is real — the fast verdict is load-bearing and its latency IS the dev loop's latency:
-
Reserve breadth a slot.CLOSED 2026-08-19 — measured wrong. The proposal was to run the full tier after N fast verdicts and let pushes queue behind it, accepting fast-verdict latency as the price. The re-measurement above shows there is nothing to buy: the watcher was idle 54% of the window, ~2.8 hours, about 8x the 1256s one full tier needs. Capacity was never the constraint; contiguity was. This shape spends the one genuinely scarce resource (fast-verdict latency, which IS the dev loop's latency) to buy capacity already present eightfold.Left visible rather than de-ranked, because the next person to notice zero breadth will propose it again — it is the obvious fix, and it is wrong for a reason nobody can see without the idle number.
-
Make the backfill resumable rather than all-or-nothing, so a 21-minute run can complete across several idle slices instead of needing one contiguous window that never comes. More work; does not need to steal latency.
-
Say so out loud. Whatever else,
--statusand the tstate report should carry "no full-tier verdict for N hours" — the cheap half, and it converts a silent hole into a visible one. Worth doing even if 1 and 2 are declined.
(3) is separable and small; recommend it regardless.
Note
Not caused by the two-phase design being wrong — the design says the fast verdict wins and a backfill is discardable, which is right. What has changed is the arrival rate crossing the point where "idle" stops occurring at all. A rate threshold nobody set explicitly is worth naming before it is tuned.
2026-08-19 — shape 3 SHIPPED (the visibility half). 1 and 2 remain open.
The scheduling trade is untouched and still undecided; what is fixed is that the hole is no longer silent.
--status now prints a breadth line whenever a host has ever run a full
tier — always, not only past the threshold, because the age IS the boundary of
the claim and a boundary nobody can see does not get checked:
tstate: host plexus last 6fba42d69830 RED (native, ...); full through 9bfb7fcfac03 GREEN
tstate: breadth — newest full tier is 4h old, 39 testable commit(s) behind
Past BREADTH_STALE_SECS (6h) it adds
[STALE — no cross-target verdict on this tree; native GREEN does NOT cover i386/arm32/riscv32/aarch64].
"Testable commits behind" counts what a gate OWES (needs_test), not raw
commits: on this repo the watcher's own publishes are most of the log, so a raw
count would read as alarming on a quiet day and bury the signal on a busy one.
Published reports carry it too, which is the half that matters days later —
a native report is what a reader reaches for, and GREEN on it says nothing
about the cross targets:
BREADTH IS 7h STALE. The newest
fulltier on this host is 7h old, so no cross target has seen this tree. Anativeverdict covers x86-64 only — do not read it as matrix coverage.
A host that has never completed a full tier gets its own wording, because an
undefined age must not render as fine. A full report gets no banner — it is
the breadth run.
The framing worth keeping, and why this was worth doing before the hard part
--status said UP and was correct: it measures whether commits were tested,
and they were. Every verdict said GREEN and each was true of the tier that
produced it. The defect was never a wrong answer — it was a true answer to a
narrower question than the reader believes it is answering, and this is that
shape at its most load-bearing, because CLAUDE.md tells every lane to run
quick + self-host and offload the matrix, and for four hours there was no
matrix.
The only thing standing between that and a wrong conclusion was one agent warning another by hand. A correctness property that depends on somebody remembering is a habit, not a property, and habits do not survive a context clear.
Gate: tools/twatch_breadth_visibility_devtest.py (new, in
make tools-devtest) — the stale banner and its age, the never-ran wording, the
fresh case, the full-tier case, and that secs_since returns None rather than 0
on a malformed timestamp (0 would render "0h old" and mean the opposite).
Still open: shape 1 (reserve breadth a slot) and shape 2 (resumable backfill). Both trade against fast-verdict latency, which is the dev loop's latency, and neither should be decided from one busy afternoon.
2026-08-19 — shapes 4 and 2 SHIPPED. Every open shape is now closed or landed.
Approved by the coordinator on the reasoning rather than the authority ("neither shape touches fast-verdict latency"), which is the right test: step 1 of the ladder is the dev loop's latency and nothing here may spend it.
Shape 4 — an unfinishable idle phase yields the slot instead of holding it
(546771cbe). After IDLE_YIELD_AFTER (3) consecutive preemptions on one
target, a phase yields a single turn to the phase below it.
- Three, not one. Yielding every other turn would invert a priority that exists for a real reason — pin verify is the binary B/C/D/E are building with right now, while HEAD is a sha nobody has adopted yet.
- Keyed on target as well as phase, so a new pin starts with a full budget rather than inheriting the previous pin's exhaustion.
- The yield is spent BEFORE the run it enables. Clearing it afterwards would leave it standing through an abort — the likely outcome, since the same push rate is what earned it — and hand breadth the next slot too, turning a one-turn loan into the inversion this is careful to avoid.
Shape 2 — an aborted run costs the work it had LEFT, not the work it had
done. The expensive half turned out to be already built: testmgr handles
SIGINT, tears its jobs down, and still writes its report, verdict
INTERRUPTED, listing every job it had finished. run_gate returned before
reading it. So the fix is a carry-over, not a checkpointing engine:
- twatch saves that report as a partial keyed on
(sha, tier)(.testmgr/resume.json, untracked); - the next slice of the same work passes it to testmgr as
--resume; - testmgr skips the jobs it already decided and merges their verdicts into the report it finally publishes.
Three ways that could be wrong, each pinned by a check:
- Results not attributable to this binary. A partial is only valid against
a byte-identical compiler, so the partial carries the compiler's
sha256and testmgr — the only thing that knows the bytes before any job runs — discards on mismatch. The invariant is real (the compiler is built through the self-host fixedpoint at the tested sha) but holds at the default-Olevel, which is the only level anything compilescompiler.pasat today; a tier that ever built it at another level would break it quietly, and comparing bytes is what turns that into a discarded partial instead of a wrong verdict. That assumption is stated at the check site, not just here. - Carrying a job whose artifacts are gone. The aborted slice ran in its own
RUN_TMPand dropped it at exit, so anything a still-to-run job depends on must be re-run — transitively. This would not have failed loudly: it would have run a dependent against a missing artifact and reported an ordinary red. - Laundering a carried RED into a GREEN. The failure was decided in an
earlier process, so this run's
rcknows nothing about it.carried_red()is what stops the report carrying a red job while announcing green.
The verdict is never partial. An aborted run still publishes nothing at all
— run_gate still returns (None, "aborted") and the caller still records no
verdict. A partial is an input to the next run, not an output to the fleet.
Counted, because the failure mode here is silence
A resume that always discards is indistinguishable from one that works: both
eventually produce full coverage, and the broken one merely costs what today
already costs. So .testmgr/resume-stats.json counts partials saved, runs
carried, testmgr-side discards, aborts that left no report (SIGKILL after the
30s grace), and supersedes; resume_health() prints the rates, not the events.
Same fix as the breadth banner, applied to its own machinery — a property that
holds only because somebody remembered is a habit, not a property.
What shape 4 must not be credited with
It makes the lower branches reachable, not fast. It does not finish a 21-minute job in 5-minute slices; shape 2 is what makes the turns add up. A metric will move when the daemon restarts and it will be tempting to read that as the fix landing. If breadth's first post-restart run is still an abort, that is expected, not a regression — the push rate has not changed.
Gate: tools/twatch_idle_yield_devtest.py (shape 4; pins both opposite
failure directions — yielding too eagerly and yielding permanently) and
tools/twatch_resume_devtest.py (shape 2; the three wrong-carry modes above
plus the counting), both in make tools-devtest.
Status: shapes 1/3/4 done; shape 2 shipped, then measurably did nothing.
1 closed as measured-wrong, 3 shipped (visibility), 4 shipped and working
(idle_yield counts on the live tstate). Shape 2 shipped as e2449adc5 and
delivered nothing at all — see the correction below.
Correction, 2026-08-19: shape 2 had a 100% loss rate
The line above once read "all four shapes resolved". That was written from the code landing, not from the numbers, and the numbers say otherwise. From the watcher clone at 20:23:12Z, over the feature's entire life:
saved_partials 9 · saved_jobs 1420 · carried_runs 0 · superseded 9
Nine saves, nine supersedes, zero carries. Root cause: .testmgr/resume.json
was a single slot every gate run claimed, and resume_arg() deleted a partial
belonging to other work rather than declining it — so the push-driven native
verdict that ends an idle slice destroyed the pin-verify partial saved when
that slice was preempted. Filed and fixed as
bug-t-a-saved-partial-is-evicted-by-the-next-run-of-different-work (keyed
partial store, PARTIAL_CAP, read-only resume_arg).
So this ticket's own conclusion — "shape 2 is what makes the turns add up" — was never in force. A 21-minute pin verify still could not complete in 5-minute slices, for the whole period the fix was believed to be deployed.
This ticket stays open on a NUMBER, not on code. It closes when
carried_runs on the live clone leaves zero and a full tier lands off resumed
slices. Watching the stats file is the remaining work — which is what the stats
file was built for, and this correction is the first time it earned its keep.
Measurement, 2026-08-20: the first half is met; the second half cannot be, for breadth
carried_runs has left zero. From the live clone at 04:24Z:
saved_partials 16 · saved_jobs 3282 · carried_runs 1 · carried_jobs 15
last_note: "resume: partial accepted — 15 job(s) already decided against
this exact binary (1479b663dd15)"
So the keyed store works and the eviction bug is genuinely fixed. But the one carry was a pin verify, and that is not luck. Breadth cannot carry, and the proof is short enough to state in full:
- Breadth targets
tested = st["last"].sha— the newest sha the fast tier has verdicted (twatch.py:5047, used at:5134). - Breadth aborts through
make_preempted(clone, tested)(:3899), which returns True only when a commit betweentestedand head passesneeds_test()— i.e. touches anything outsidedevdocs//docs/(:3892,NOTEST_PREFIXESat:3602). - That is the same predicate that makes
do_testfire on the next cycle (:5051), which runs the fast tier at head and movestestedto it. - The partial is keyed
(sha, tier)(:1887) andresume_arg()opens only that one file (:755,:666).
⇒ the abort that creates a breadth partial is caused by the very sha move that
makes its key unmatchable. Not a race, not a tuning problem: a breadth partial
is dead at the instant it is written. The only breadth path that can carry is
STOP — a daemon restart, where tested does not move. Pin verify carries
because it targets a frozen pin sha across slices.
Live, from the log, in one 40-minute window:
testing e0f9cb893fea (full) -> aborted -> 64 jobs kept -> key dead
testing 011e3904e52e (full) -> aborted -> 537 jobs kept -> key dead
testing ef3ea948003d (full) -> aborted -> 15 jobs kept -> key dead
616 jobs saved, 0 carried, and the newest full tier is 49a511e43271 at
2026-08-19T23:23:18Z — 5h12m stale at the time of measurement, with five
pushes in the interval. The starvation this ticket names is live right now.
The key is NOT too strict — do not "fix" it by keying on the binary
The obvious next move is to drop the sha from the key, since testmgr already
validates the partial against the compiler's sha256 (testmgr.py:1474) and
that is the invariant that actually matters. That would be unsound. The
compiler sha256 answers "same binary"; it says nothing about the corpus or the
harness. tested only ever moves when a commit touches something outside
devdocs//docs/ — which means compiler/, test/, tools/, or lib/. The
sha256 covers the first. The git sha is what covers the other three, and a
carried verdict for test_foo measured against an old expected-output file is
exactly the wrong-answer-with-confidence this repo keeps paying for.
So the two-part key is correct as designed, and shape 2's near-zero carry rate for breadth is the mechanism refusing to carry work it cannot attribute — not a defect in it.
What that means for this ticket
The close condition as written — "a full tier lands off resumed slices" — is unachievable for breadth by construction, so it cannot be the bar. Shape 2 is doing its job on the path where carrying is sound (pin verify); breadth's remaining lever is the one shape 1 already identified and this measurement re-confirms: contiguity, not capacity. Idle time is abundant; a contiguous 21-minute window is not, because a push arrives every ~8 minutes and every push restarts the ladder from zero.
Revised close condition: a full tier lands within a working day of the sha it tests, sustained over a week. That is the outcome the ticket actually wants, and unlike the old one it is reachable by more than one mechanism.
Found on the way: a killed job was being published as a failed one
Reading the partials to write the above turned up
ef3ea948003d-full.json holding selfhost-fixedpoint#00 as a 26.7s "fail"
after an eight-second run — an artifact of teardown(), not a verdict. Filed
and fixed separately as [[bug-t-a-killed-job-is-published-as-a-failed-job]],
because it is a phantom red on the published path too, not only on this
ticket's resume path.
2026-08-25 — the premise expired, and a reservation now exists
This ticket's own conclusion — "the fix is resumability plus bounding consecutive idle, NOT reserving a slot" — was right for its measurement and is wrong for today's. Recorded here rather than edited away: on 2026-08-19 the box was idle 54% of the window, so breadth did not need a slot reserved, it needed the unfinishable item ahead of it (pin verify) to stop holding the queue.
What changed is the box, not the queue. Plexus became a workstation on 08-20 and the watcher was capped at 6 of 12 cores. Measured tonight:
| 08-19 | 08-25 | |
|---|---|---|
| a native run | ~246s | ~490s |
| testable commits landing during one | — | ~9 |
cycles where do_test is true |
some | essentially all |
| idle ladder reached | 54% of the window | never |
So breadth no longer merely runs rarely — it cannot start. do_test
outranks the ladder unconditionally, and a new commit is always pending by the
time a native run ends. That also made the commitment point (572524c7c, which
lets a running breadth run survive a push) inert: nothing was ever running.
What landed
572524c7c— a breadth run pastfull_commit_secs(300s) finishes instead of being discarded. Native stays preemptive at every moment.b61125007— the global deadline scales with the core budget. At 6 of 12 cores the matrix (~24000 cpu-seconds) cannot fit an hour at any packing, so the old fixed 3600s ceiling was unsatisfiable by construction and every breadth run would have been torn down at the same minute.breadth_overdue()— the reservation, deliberately narrow: it fires only when breadth is already stale pastBREADTH_STALE_SECS(the threshold the reports have been printing all along), takes one run, and is guarded by aBREADTH_RETRY_SECSbackoff.
The backoff is the part that matters more than the reservation. A breadth run
that does not LAND leaves last_full untouched, so the staleness test stays
true forever — without a backoff the reservation fires every cycle and the fast
verdict, which is what every lane actually steers by, stops entirely. That
failure would be strictly worse than no breadth. twatch_breadth_slot_devtest.py
fences it, along with the "a landed run releases the slot immediately, backoff or
not" case.
What this ticket still owes
carried_runs leaving zero, per the summary line — the resume path is
untouched by tonight's work. The reservation makes a full run possible;
resumability is what makes a preempted one cheap. Still open.
RESOLUTION 2026-08-26 — breadth recovered, and the fix was the shape this
ticket explicitly ruled out
Breadth is healthy. Measured over 3,828 run records:
| window | runs | full-to-full gap |
|---|---|---|
| last 24h | 48 (42 native, 4 full, 1 slow, 1 opt) | median 1.3h, worst 3.1h |
| last 7d | 306 (254 native, 28 full) | median 1.2h, worst 40.1h |
Current staleness at the time of writing: 3.3h since the last full. The ticket was filed on 0 full runs in 5h13m.
The headline says the wrong thing, and the dates prove it
The summary line reads: "Fix is resumability plus bounding consecutive idle, NOT reserving a slot." The evidence says the opposite — the thing it ruled out is the thing that worked.
| landed | what | breadth afterwards |
|---|---|---|
2026-08-19 17:56 546771cbe |
shape 4 — an unfinishable idle phase yields the slot | — |
2026-08-19 18:16 e2449adc5 |
shape 2 — resumable partials | gaps got worse: 12.8h, 9.4h, 21.6h, 19.2h, 31.5h, 40.1h |
2026-08-25 20:44 572524c7c |
commit_after — a run 5 min in finishes instead of being discarded |
last full before: 08-25 20:32 |
2026-08-25 21:25 7457f6aee |
breadth_overdue — breadth reserves a slot when stale, commit_after=0 |
next gaps: 1.1h, 3.1h, 1.3h |
Shapes 2 and 4 shipped six days before recovery and breadth degraded steadily across the whole interval. The reservation landed and the next three gaps are 1-3h.
The confound, stated rather than hidden: 2026-08-20 is also when borg's PSU failed and this box became the human's workstation, so load rose at almost the same moment shapes 2/4 landed. That is real, and it is why the fair reading is shapes 2 and 4 were not harmful, they were insufficient against the new load. But it does not rescue the headline, because the confound never went away — plexus is still a shared workstation, Track T is still throttled to ~50% of the box, pushes still arrive on the same cadence — and breadth recovered anyway. Load held constant, mechanism changed, outcome changed.
And the resume ledger says the same thing numerically
saved_partials 83 saved_jobs 22280 carried_runs 3 carried_jobs 73
superseded 70 empty_partials 2 no_report_on_abort 4
73 of 22,280 saved job-results were ever reused: 0.33%. Not because the
mechanism is broken — superseded: 70 is the explanation and it is a correct
one. A partial is keyed on (sha, tier), and when the run is aborted the
watcher re-targets to the new HEAD, so the partial it just saved is for a sha
nobody will ask about again. Resumability can only pay when the same (sha,
tier) is retried, and the push-driven ladder almost never retries one. That is
a structural ceiling, not a bug, and it should be recorded as such rather than
chased.
The one place it does pay is the phase that does retry the same sha: pin
verify, which is bounded to one pin. last_note confirms it — "partial
accepted — 56 job(s) already decided against this exact binary" — and the live
log carries kept 432 decided job(s) from the aborted full run.
What this ticket is worth to the next person
Its close condition was carried_runs leaving zero. That is satisfied (3), but
the boolean was the wrong instrument and would have closed this ticket six
days early, on a mechanism recovering 0.33% of what it saves, while breadth was
in the middle of a 40-hour gap. resume_health()'s own docstring says it:
"One line of RATES, not events." The condition that actually answers "is
breadth starved" is breadth staleness, which is now measured above and is the
number to re-check.
Resolved on the outcome, not on the bit.
Log
- 2026-08-26 — measured; headline corrected; closed on breadth staleness.
- 2026-08-26 — resolved, commit 18054e41c.