Three defects found by watching v0.23.5 un-stick a real PR (cliftonc/drizzle-cube#1016) in production.
The gate ran twice, and the second run could not change anything
Run 49c101aa: the agent pushed at 11:03:31 having already run the repo’s whole suite itself. GitHub’s checks went fully green at 11:06:25. The harness spent 11:04:00–11:10:48 in a fresh container re-running the same suite a third time before agreeing — a third of the run, spent re-proving something the real CI had already answered.
until: is evaluated before until_bash and short-circuits it, so outcome=pushed now ends the loop. Once the commit is on the branch, GitHub is the strictly better authority: the real CI environment rather than a sandbox approximation, warm rather than a cold container, covering matrix legs the sandbox cannot reproduce.
It short-circuits on pushed only. no-change and gave-up still pay for the gate — nothing was pushed, so there is no new commit, no new check run and no external authority at all; the local gate is the only evidence that exists and its RED verdict is what earns the agent its next iteration.
What this gives up, stated plainly: after a push the harness no longer independently checks the agent’s self-reported gate=green. That gate ran after the push and therefore never gated it — the self-report was already the only thing between a bad fix and the branch, on every run this workflow has ever done. What actually catches a bad fix is untouched: red checks re-dispatch the fix family, bounded by fix.maxAttempts / fix.maxCostUsd.
The gate the agent wrote was a CI clone, because we asked for one
The fixing skill said the gate is “whatever CI runs” and both fix prompts said “mirror CI”, with nothing bounding cost or scope. So for a merge conflict — fixed by merging main and regenerating a lockfile — the agent faithfully wrote five builds, two full test suites, and three docker branches that can never run (there is no docker in the sandbox).
The instruction now asks for the narrowest command that would have failed before the fix and passes after it, targeting under two minutes, with explicit exclusions: not the whole suite, not a check already watched passing this session, nothing that starts a service, nothing that mutates git state.
Writing no gate stays wrong — it burns the remaining iterations and forces gate=skipped, which is RED and never authorises a push — so the nothing-to-verify case gets an ordered fallback: the coherence check the repair implies, else an honest one-line exit 0 saying why.
A third of the run was invisible
The until_bash window was not a phase, had no ledger row, no pipeline node and no timer — and the iteration’s own phase entry was withheld pending its verdict, so currentPhase read diagnose (twelve minutes stale) while every surface showed a run that looked finished but stuck.
Iterations now persist when their work finishes, and the check gets its own <phase>_iter_N_check row: open with a start time while in flight, closed with a duration and a condition_met / condition_not_met verdict. success records whether the check ran, not what it said — a red gate is the loop working as designed, and renders muted rather than failed.
Also
- A paused
explore socratic round used to be recorded nowhere at all (the interactive gate returned above the persist call). It now lands in history before the pause.
- A push followed by a red or exhausted gate no longer posts “Couldn’t auto-fix — leaving it for a human” on a PR it had in fact just fixed.
Full changelog: https://github.com/nearform/lastlight/compare/v0.23.5...v0.23.6