v0.38.2 @libsql/client 0.18

Latest View on GitHub

Highlights

  • @libsql/client 0.17 → 0.18 (#433). 0.18 replaces the local client’s connection relay with a real connection pool, and two state-layer pieces are adapted to it:
    • SQLite busy timeout now covers every connection. The write lock’s PRAGMA busy_timeout re-arm reached one pooled connection of many, so a write on another failed SQLITE_BUSY at once instead of waiting out another process (lastlight state, a backup). The timeout is now set as the client’s timeout option, which libsql applies to every connection the pool opens.
    • The legacy messaging_sessions rebuild runs as one client.migrate(), with its foreign-key check enforced inside the transaction. A violation now rolls the rebuild back instead of being detected after COMMIT.

No schema or config changes.

Packages

  • lastlight 0.38.2 · lastlight-core 0.38.2
  • lastlight-evals 0.18.2
  • lastlight-shared unchanged (0.11.1) · lastlight-code-facts unchanged (0.8.1) · lastlight-workflow-engine unchanged (0.11.0) · agentic-pi unchanged (0.7.0)

Full changelog: https://github.com/nearform/lastlight/compare/v0.38.1...v0.38.2

v0.38.1 dependency refresh

View on GitHub

Highlights

  • Dependency refresh (#432). Folds the open Dependabot PRs (#411–#419) into one change with a single lockfile regeneration.
    • vitest 5 in every package (the CLI was already on it).
    • Runtime: zod 4.6, hono 4.13.9, OpenTelemetry 0.222 / 2.11, react 19.3, astro 7.3, @clack/prompts 1.8, drizzle-orm 0.45.3, @slack/bolt 5.1, @ff-labs/pi-fff 0.11, plus patch bumps.
    • Tooling: vite 8.3, wrangler 4.139, biome 2.5.14, drizzle-kit 0.31.11, pglite 0.5.8, tsx, @types/node.
  • Held back: @libsql/client 0.18. It changes connection-release behaviour under the #421 SQLite write lock, so it will ship in its own PR after a closer look.

No behaviour changes intended.

Packages

  • lastlight 0.38.1 · lastlight-core 0.38.1
  • lastlight-code-facts 0.8.1 · lastlight-shared 0.11.1
  • lastlight-evals 0.18.1
  • lastlight-workflow-engine unchanged (0.11.0; dev-only bump) · agentic-pi npm unchanged (0.7.0; the sandbox image vendors the workspace copy, so it picks up the new deps anyway)

Full changelog: https://github.com/nearform/lastlight/compare/v0.38.0...v0.38.1

v0.38.0 re-reviews converge

View on GitHub

Highlights

  • Re-reviews converge (#429, #431). With review.analysis on, pr-review now remembers what it reviewed and re-reviews only what changed.
    • A review ledger per PR. Each review records the units it covered, per-file line hashes, and every finding (posted or withheld). The next review is dispatched with it. Finding status (open / withheld / addressed / resolved) comes only from structured signals — the quoted code gone at the new head, our thread resolved — never from reply text. No table or migration: the ledger rides the run’s scratch.
    • Delta-scoped sites. Units get a stable identity and a delta against the last review (new / changed / affected / unchanged). Rows of unchanged units are carried, not investigated, so a small push gets a few sites and a push that changed nothing the last review covered plans none.
    • A per-line convergence gate. A re-review finding whose lines were all already there at the last review is a late discovery: it is withheld as converged unless it is must-fix on non-low-risk code, which posts labelled “Missed in an earlier review”. A re-found open finding is withheld as already-raised, and the summary opens with what was addressed and what is still open.
  • Risk tiers (review.risk.rules). First-match globs — a repo’s .lastlight/ rules, then the operator’s, then built-ins (docs/tests/generated low; migrations/schema/auth/CI high) — raised one tier by a security/state obligation or high fan-in. They weigh site ranking, gate strictness and coverage; they never decide whether a changed unit is surveyed.
  • Review coverage. Each run writes per-unit surveyed vs investigated, risk-weighted, and the in-scope units nobody investigated. The run detail page gets a Review tab showing coverage and the ledger.
  • Evals: chained re-review cases. A pr-review case can list rounds: [{ head_commit }] to replay a PR’s review history through the real workflow, carrying the ledger forward. Each round is diffed from its own merge base, and seeded human discussion can carry from_round so an early round never reads the review a later one is graded on. The scorecard and dashboard show late discoveries, converged / already-raised withholds, cumulative gold recall and per-round coverage and cost.

Measured on real skillspro review chains: a push that changed nothing reviewed costs about a third of a first review (~35 s, no investigators); a 1-of-34-unit push investigates one file; across both repeats 1 of 8 later-round comments landed on already-reviewed lines.

Packages

  • lastlight 0.38.0 · lastlight-core 0.38.0
  • lastlight-code-facts 0.8.0 · lastlight-shared 0.11.0
  • lastlight-evals 0.18.0
  • lastlight-workflow-engine unchanged (0.11.0) · agentic-pi unchanged (0.7.0)

Full changelog: https://github.com/nearform/lastlight/compare/v0.37.0...v0.38.0

v0.37.0 dynamic fan-out

View on GitHub

Highlights

  • Fan-outs can take their branch list from run-time data (#423). A type: fanout phase can declare branches_from: { file, max } and a branch: template instead of a fixed branches: list.

    • The template is rendered once per item of a JSON manifest an earlier phase wrote into the workspace, with each item available as {{item.*}}.
    • max caps the spend.
    • An empty manifest is a no-op.
    • A malformed manifest fails that phase, loudly, rather than the run.
  • pr-review runs one site investigator per real site (#423). site-review no longer declares sixteen fixed slots and pre-closes the unused ones. site-plan writes the list of sites it actually formed, so a PR with two sites runs two investigators. With models.review-site-pair set, each site’s second investigator is site-00N-b.

  • Compact fan-out view in the dashboard (#423). A fan-out with five or more branches is drawn as one block:

    • status counts and one chip per branch, including chips for branches that haven’t started yet;
    • clicking a chip opens that branch;
    • an expand toggle brings back the full cards.

    The workflow diagram labels dynamic fan-outs fan-out · dynamic ≤ N.

  • Evals record the right model for every fan-out branch (#423). Branch results now carry the model template they ran under, so pair investigators are scored on review-site-pair. Stored runs in the old sixteen-slot format still replay.

Packages

  • lastlight 0.37.0 · lastlight-core 0.37.0
  • lastlight-workflow-engine 0.11.0 · lastlight-code-facts 0.7.0 · lastlight-shared 0.10.3
  • lastlight-evals 0.17.0
  • agentic-pi unchanged (0.7.0)

Full changelog: https://github.com/nearform/lastlight/compare/v0.36.0...v0.37.0

v0.36.0 superseded reviews stop cleanly, kinder re-reviews, range-anchored comments

View on GitHub

Highlights

  • A superseded review stops cleanly (#426). When a new push supersedes an in-flight pr-review, the cancel is now final. Before, the killed phase’s failure could flip the run from cancelled to failed, and the dead run carried on through its remaining phases. A fan-out also stops launching branches, gates and gate re-runs into a killed sandbox, and the phase shows superseded: … instead of a misleading Sandbox agent failed (exit 137).
  • No model call for an empty selection (#426). When the site investigators find nothing, select is skipped: its only correct answer is fixed. This removes the “empty completion” failure it used to hit. Workflow skip_if gains an anchored startsWith(...) expression.
  • Brief but kind review summaries (#427). A clean first review reads “Looks good — no issues to raise.” A clean re-review thanks the author and says it’s good to merge, but only when nothing from an earlier review is still open: an open ledger point, an already-raised finding, or an unresolved inline thread of ours. An @bot review of an unchanged head isn’t treated as a re-review. The default soul gains a “brief, not curt” line.
  • Review comments can highlight a line range (#428). Site investigators can give an optional startLine, so a comment highlights the defect’s own stretch of code rather than a single line with unrelated context above it. A range that doesn’t fit the diff falls back to exactly the placement it would have had without one, so asking for a range never costs a finding its inline comment.
  • micro-select eval (#424). A new lastlight-evals script replays pr-review’s select phase over preserved runs, to compare selection models on identical input, with gold judging and a dashboard view.

Packages

  • lastlight 0.36.0 · lastlight-core 0.36.0
  • lastlight-workflow-engine 0.10.0 · lastlight-code-facts 0.6.0 · lastlight-shared 0.10.2
  • lastlight-evals 0.16.0
  • agentic-pi unchanged (0.7.0)

Full changelog: https://github.com/nearform/lastlight/compare/v0.35.2...v0.36.0

v0.35.2 resumed runs keep their context

View on GitHub

Highlights

  • A resumed run keeps its dispatch’s template context (#425). resumeSimpleRun rebuilt only the base template fields and dropped everything the dispatch had rendered onto the run: the PR snapshot, analysisEnabled, triageEnabled, and a comment’s commentBody. Every re-entry goes through resume: boot recovery after a restart, admission of a run queued at the concurrency cap, and Retry. So a pr-review on a busy instance, or one in flight during a deploy, skipped all its analysis phases and silently posted the light single-pass review, or failed at post-review with no findings. Resume now restores the persisted context underneath the fields it owns.

Packages

  • lastlight 0.35.2, lastlight-core 0.35.2, lastlight-evals 0.15.2

Full changelog: https://github.com/nearform/lastlight/compare/v0.35.1...v0.35.2

v0.35.1 SQLite write lock, review supersede, sweep grace window

View on GitHub

Highlights

  • SQLite write lock (#421). Every SQLite write now holds one in-process lock. Before, a plain write racing an open transaction could fail with SQLITE_BUSY: database is locked, and the next transaction then failed to commit with cannot commit transaction - SQL statements in progress. The worst case was the boot resume sweep re-dispatching several orphaned runs at once, which failed every one of them.
  • A newer commit replaces an in-flight review (#422). When a review of an older commit is still running and the PR’s current commit is due a review, the old run is cancelled: its containers are killed, and the new review waits for it to stop before using the shared workspace. A review of the same commit is never replaced, so an @bot review sent mid-flight gets the “already working” reply.
  • The review sweep no longer reviews mid-CI (#422). New review.sweepPendingGraceMinutes (default 60): the 30-minute sweep reviews a PR whose checks are still pending only once they have been pending that long. It is still the safety net for checks that never finish, but it no longer fires a minute after a push.
  • review.placeholderCheck (#422). Set it to false (with postsCheck: true) and the last-light/review check appears only once a review actually starts, with no queued placeholder while CI runs.
  • Workspace guard (#422). A run now refuses to reset a per-PR workspace that another still-running run owns (WorkspaceBusyError). If two runs collide, one fails instead of both.

Packages

  • lastlight 0.35.1, lastlight-core 0.35.1, lastlight-shared 0.10.1, lastlight-evals 0.15.1

Full changelog: https://github.com/nearform/lastlight/compare/v0.35.0...v0.35.1

v0.35.0 pr-review: units + sites, paired investigators

View on GitHub

Highlights

  • One pr-review analysis path: units + sites. With review.analysis.enabled, the review now runs facts → seed → a unit survey (one bounded model call per changed function or module region, in parallel, cached) → sites → reconcile → post-review. The survey’s rows are treated as a signal of where to look: rows cluster into sites, and each top site gets one investigator that reads the code, runs probes, and writes 1–3 grounded findings, or a none it has to back with an executed probe. One select pass then merges duplicates, sets importance and writes the comments. The five-branch agent survey, the adjudicator, dossier and jev-classify are gone.
  • Paired investigators. Set models.review-site-pair and every site gets a second investigator on a different model. Two models miss different things, so the pair finds more than either one, and select merges what both report.
  • Coverage and context. review.analysis.siteTop (1–8, default 5) sets how many sites get an investigator. Test-file sites now fill slots the code sites leave free instead of being skipped. Investigators read the PR’s title, description and closed issues as a claim to check. select reads the PR’s earlier reviews, threads and comments, so a point someone already raised is recorded rather than posted again.

On the Martian held-out set (18 PRs across 5 repos, 54 known defects, two repeats), judged on the posted review: sites with gpt-6-luna investigators matched 24/19 defects at $0.37/PR, deepseek-v4-flash 26/28 at $0.61, and the luna + deepseek pair 28 in both repeats at $0.70. The earlier Haiku-investigated arm matched 13 of 44. The full record of what was tried is in docs/plans/pr-review-units-sites.md.

Fixes

  • A single fan-out branch that failed (for example a provider’s transient 404) no longer marks a posted review as failed. A failed run left the head unassessed, so the review sweep would re-dispatch it and post again. Branches that hit a provider error are also retried once.
  • A light re-review now runs review and posts, instead of posting nothing.
  • On the in-process backends, a timed-out bash phase now kills its whole process group, and bash phases no longer block the event loop.

Deploy notes

Bump deploy.version to v0.35.0 in the overlay. This changes what pr-review runs when review.analysis.enabled is on:

  • Pin the investigator model. models.review-site is unset by default and falls back to models.review-survey. Measured, cheap models beat Haiku here: openai/gpt-6-luna (with variants.site-review: medium) or opencode/deepseek-v4-flash. models.review-select falls back to models.review, so pin it too if review points at a model you don’t want writing the comments.
  • Removed keys are ignored with a warning: review.analysis.surveyEngine, reviewEngine, independentReview, adjudicate, jevModel, admit, jevTimeoutSeconds, surveyPasses, and models.review-adjudicate. Delete them from overlays.
  • surveyConcurrency is now siteConcurrency (the old name is still read). With the pair, a PR runs up to 10 investigators, and the docker backend runs at most 6 at once.
  • probes is inert for now. Falsify is not yet attached to the sites engine; the investigators run their own probes.
  • Kubernetes: review.analysis.enabled is refused at startup on the kubernetes backend.

Packages

lastlight 0.35.0 · lastlight-core 0.35.0 · lastlight-evals 0.15.0 · lastlight-code-facts 0.5.0 · lastlight-shared 0.10.0 · lastlight-workflow-engine 0.9.0 · agentic-pi 0.7.0 (unchanged)

Full changelog: https://github.com/nearform/lastlight/compare/v0.34.1...v0.35.0

v0.34.1 sandbox CPU and memory per phase

View on GitHub

Highlights

  • Sandbox CPU and memory, per phase and per workflow. Docker and Kubernetes sandboxes now record the CPU time and memory high-water mark of every phase, read from the sandbox’s own cgroup. The home page gains a CPU card and a Sandbox CPU time chart, with the day’s largest memory peak as a line. Recent Workflows shows each run’s CPU (and its peak memory in the tooltip), and phase detail shows CPU Time and Peak Memory against the limit. A fan-out’s shared container is recorded once rather than split across its branches. The in-process backends (gondolin / none) report nothing (#409).
  • Token stats count cache writes. The Tokens total and token chart now include cache-write tokens, shown as a cache write bar. Anthropic reports the uncached prompt as a cache write while OpenAI-compatible providers (OpenCode Zen) report it as input, so without them a switch to open models looked like a jump in input tokens. Historical totals rise by the Anthropic cache-write volume. That’s a correction, not new usage (#409).

Fixes

  • PR review no longer puts prose in “Apply suggestion” blocks (prompt-level). The adjudicator had lost the definition of suggestion and was filling it with a description of the fix, which pressing Apply would have committed into the source file. The field is now defined everywhere it’s written as exact replacement code, or omitted (#410).

Deploy notes

Bump deploy.version to v0.34.1 in the overlay. The release adds three nullable executions columns (additive migration on SQLite and Postgres, applied at boot). No config changes.

On Kubernetes, the usage reading travels on the pod’s log stream. A pod that dies early (OOM-kill, deadline) records none, and it’s a best-effort metric rather than a security boundary. See spec/09-sandbox.md → Resource usage.

Packages

lastlight 0.34.1 · lastlight-core 0.34.1 · lastlight-evals 0.14.1 · lastlight-shared 0.9.1 · lastlight-workflow-engine 0.8.1 · lastlight-code-facts 0.4.0 · agentic-pi 0.7.0

Full changelog: https://github.com/nearform/lastlight/compare/v0.34.0...v0.34.1

v0.34.0 PR review on open models, command policy, fewer and better review comments

View on GitHub

Highlights

  • PR review on open models. pi 0.87 and OpenCode Zen (opencode/… models, OPENCODE_API_KEY): the whole pr-review pipeline can now run without an Anthropic key (#402).
  • Command policy per phase. Workflow phases can block or log installs, test runs, and (new) any bash call that reaches outside the agent’s workspace (host: whole-disk find, ~/.nvm, global node_modules, PATH pointing outside). pr-review blocks installs and test runs in the phases that must not run code, and blocks host on survey, review and adjudicate (#403, #404).
  • Fewer, better review comments (#405):
    • Impact rule: style, locale, “no tests”, dead-code and convention findings are recorded but no longer posted. A finding adjudicate calls a defect is never demoted this way.
    • Verdict hygiene: falsify’s reproduced now means the scenario was executed; a grep or file read is corroborated.
    • Derived severity: severity is computed from the evidence, not the adjudicator’s label, and the posting caps rank on it, with a tie-break computed in code within each severity.
    • Summary after the caps: the review’s opening summary is written after the caps decide what posts, so it never mentions a finding that was held back.
  • Calibration harness for pr-review findings, and the two fixes it found (#398).

On the Martian cal.com set (3 cases × 3 repeats, all-open stack): posted comments per case are down about 20%, precision .220 → .246, cost at or below the previous release.

Fixes

  • {{artifactUrl}} links no longer 400 on docker/kubernetes with buildAssets.location: server (#400).
  • The survey gate rewrites hypothesis ids to their canonical form, which removes a mismatch that sent adjudicate into a second full pass in about half of cases.

Deploy notes

Bump deploy.version to v0.34.0 in the overlay. Config changes that can matter:

  • review.analysis.maxInlineComments default 10 → 5. Combined with maxBodyComments: 0, a review now posts at most 5 comments. Pin maxInlineComments in the overlay to keep the old cap.
  • review.analysis.thresholds and internalFloor are removed. They’re accepted and ignored with a warning; delete them from overlays.
  • review.analysis.probes takes three values (off | static | full). true still reads as static, and false or unset as off.
  • New model key models.review-summary (the post-cap summary; one tool-free call). Unset, it uses models.default. Pin it if default points at a provider this deployment has no key for.
  • New review.analysis keys: adjudicate (legacy | dossier | jev, default legacy), jevModel, falsifyTimeoutSeconds (600), jevTimeoutSeconds (120). All have safe defaults.
  • To run pr-review on open models, set OPENCODE_API_KEY in the host .env, then recreate the agent container (lastlight server start agent); a restart doesn’t pick up a changed secret.

Packages

lastlight 0.34.0 · lastlight-core 0.34.0 · lastlight-evals 0.14.0 · lastlight-code-facts 0.4.0 · lastlight-shared 0.9.0 · lastlight-workflow-engine 0.8.0 · agentic-pi 0.7.0

Full changelog: https://github.com/nearform/lastlight/compare/v0.33.2...v0.34.0

v0.33.2 workflow definition crash + feedback config banner

View on GitHub

Fixes

  • Workflow definitions: clicking a phase no longer crashes the page. Phases whose timeout_seconds or generic_loop.max_iterations is a config reference ({ from: gate.phaseTimeoutSeconds }, e.g. build’s Executor) threw React error #31. They now render as text.
  • Feedback page shows when GitHub feedback is off. A banner appears when feedback.github is disabled (only Slack reactions are recorded), or when feedback.enabled turns collection off entirely.

Deploy note: bump deploy.version to 0.33.2 in the overlay.

Full changelog: https://github.com/nearform/lastlight/compare/v0.33.1...v0.33.2

v0.33.1 autonomy WIP limits + issues-only board

View on GitHub

Fixes

  • Runs paused at an approval gate stay paused. A queued build started later by admission (and boot recovery / Retry) was marked succeeded when it stopped at post_architect, which moved its issue to ready-for-human with the approval still pending — and approving it could not resume the build. It now stays paused; the run store refuses a paused → succeeded flip; and on boot, runs already left succeeded at waiting_approval with a pending approval are restored to paused.
  • autonomy.budget.maxConcurrentBuilds ignores runs waiting for approval. Only queued/running autonomous builds count, so plans awaiting review no longer block the pipeline.
  • The build limit can’t be overshot in one sweep. The dispatch gate now counts one issue at a time and holds a reservation for each approved dispatch until its run exists (also applied to maxBuildsPerRepoPerDay). Previously a sweep could start 7 builds against a limit of 1.
  • The pipeline board shows issues only. Pull requests are no longer cards; each issue links the PRs that close it (open / draft / merged / closed), fetched in the same GraphQL query — no extra requests per issue.

Deploy note: bump deploy.version to 0.33.1 in the overlay. Issue labels already moved to ready-for-human by the old bug are not moved back.

Full changelog: https://github.com/nearform/lastlight/compare/v0.33.0...v0.33.1

v0.33.0 actionable pipeline board

View on GitHub

Highlights

  • Actionable pipeline board: dispatch builds by dragging cards, live updates, and failure reasons on cards (#394).
  • Quieter quota deferrals: error_quota phase outcomes now log at warn instead of error (#370).

Packages

  • lastlight-workflow-engine 0.7.0
  • lastlight-shared 0.8.1
  • lastlight-core 0.33.0
  • lastlight 0.33.0
  • lastlight-evals 0.13.3

Full changelog: https://github.com/nearform/lastlight/compare/v0.32.2...v0.33.0

v0.32.2 recovered tool failures, fast hermetic test suite

View on GitHub

Patch release.

Fixes

  • Empty completion after a recovered tool failure is no longer a hard error_tool (#387, #389). agentic-pi’s RunResult gains lastToolErrored (the last tool’s state, not sticky); mapStopReason keys error_tool off it, so such a run maps to the soft, retryable unknown, and a stale tool error no longer leaks into its message.
  • Octokit throttling is disabled for loopback GitHub base URLs (#388, #390). The throttle’s module-global notifications limiter spaced review POSTs 3s apart even against a local mock — the evals harness and the test suite. Real GitHub URLs are unaffected.

Test suite (#388, #390)

Full turbo run test in a 2-CPU / 6 GB container: 263s → 90s (67s with VITEST_MAX_WORKERS=2). Shared per-file PGlite with reset (fixes a 3.8 GB worker peak), scanners disabled in code-facts tests unless opted in, .env no longer leaks into core tests, agentic-pi test runs unit tests only, and the repo’s guardrails gate stopgap is removed.

Packages

  • lastlight 0.32.2, lastlight-core 0.32.2, lastlight-evals 0.13.2
  • agentic-pi 0.6.1 (published separately via agentic-pi-v0.6.1)

Full changelog: https://github.com/nearform/lastlight/compare/v0.32.1...v0.32.2

v0.32.1 opengrep locale fix for pr-review patterns

View on GitHub

Fixes

opengrep failed in the sandbox, so pr-review’s pattern scan found nothing

  • The sandbox image set no locale. opengrep, a bundled Python program, decodes its rules file with the process locale, so it read rules/review.yaml as ASCII and crashed on the first non-ASCII byte: an em dash in a comment. It exited 2 with empty stdout, so every pr-review patterns run recorded a degraded entry and produced zero pattern findings.
  • Fixed in two places:
    • sandbox-base now sets LANG=C.UTF-8 LC_ALL=C.UTF-8, inherited by sandbox and sandbox-qa.
    • extractPatterns spawns opengrep with a forced UTF-8 locale (withUtf8Locale), so host runs (--sandbox none, dev boxes with LANG=C) are covered too.

lastlight-code-facts tests that only failed inside Last Light’s own sandbox

Found by the new v0.32.0 guardrails gate on a build for this repo. It ran the full suite once, and these tests failed there while passing on CI:

  • The opengrep rules tests only run where opengrep is installed, so they had never run on CI. In the sandbox they hit the locale crash above.
  • The toolchain resolution test asserted “nothing resolves” against the real /opt/lastlight/bin, which exists in the sandbox image. resolveToolBin / resolveFactsBin now take an optional baked directory, and the tests pass one that cannot exist.
  • New regression test: opengrep loads a rules file containing non-ASCII text when spawned from an environment with no locale.

Packages

  • lastlight-core 0.32.1, lastlight 0.32.1
  • lastlight-code-facts 0.3.1
  • lastlight-evals 0.13.1

Full Changelog: https://github.com/nearform/lastlight/compare/v0.32.0...v0.32.1

v0.32.0 config-only timeouts and an exit-code guardrails gate

View on GitHub

Highlights

Config-only timeouts and an exit-code guardrails gate (#386, closes #385)

  • Guardrails proves the full suite green, by exit code. The build workflow’s guardrails is now two phases: the agent sets up (install, typecheck, lint — fail fast) and writes .git/lastlight-gate.sh with the repo’s test commands; then a new harness-run guardrails_gate phase runs the full suite once under gate.timeoutSeconds. Exit 0 → READY; failure → BLOCKED with the exit code and log tail; timeout → BLOCKED “raise gate.timeoutSeconds”. The model can no longer mark a suite it never saw pass as READY.
  • Every workload timeout lives in config. config/default.yaml is the only source; there are no numeric fallbacks in code or workflow YAML, and a missing key fails config load naming it. New keys: sandbox.agentTimeoutSeconds / commandTimeoutSeconds / untilBashTimeoutSeconds and a gate: block (timeoutSeconds, maxTimeoutSeconds, phaseTimeoutSeconds).
  • Repos can size their own gate. gate.timeoutSeconds is settable in .lastlight/lastlight.yml, clamped to the operator’s gate.maxTimeoutSeconds. fix.gateTimeoutSeconds still works as a deprecated overlay alias.
  • Agents stop guessing bash timeouts. agentic-pi --gate-timeout (always passed by core) tells the model to give installs, builds and full suites the gate timeout, judge them by exit code, and never retry a timed-out gate; too-short timeouts on recognised build/test commands are raised to the gate value.
  • Per-phase agent timeouts are enforced. An agent phase’s timeout_seconds now bounds its run on docker, smol and k8s.
  • on_output (BLOCKED + bootstrap bypass, requires_marker) now works on type: bash phases.
  • Core’s test run is quiet by default (LOG_LEVEL=fatal under vitest).
  • New docs page: Making your repo agent-friendly.

Upgrade notes

  • No overlay changes are required: the new timeout keys ship in default.yaml.
  • If an overlay sets fix.gateTimeoutSeconds, rename it to gate.timeoutSeconds (it is mapped across with a warning for now).
  • Overlay-forked workflow YAML using { from: …, default: N } still works (with a warning); packaged workflows now use { from: … } only.

Packages

  • lastlight-core 0.32.0, lastlight 0.32.0
  • lastlight-workflow-engine 0.6.0, lastlight-shared 0.8.0
  • agentic-pi 0.6.0 (published separately via agentic-pi-v0.6.0)
  • lastlight-evals 0.13.0

Full Changelog: https://github.com/nearform/lastlight/compare/v0.31.0...v0.32.0

v0.31.0 Slack digests post as threads

View on GitHub

Highlights

  • Slack repo digests now post as threads (#384, closes #383). The channel gets one short top-level message with the repo header and narrative; the content lists, stats and date range move into a threaded reply. Feedback still anchors on the top-level message. Falls back to an unthreaded post if Slack returns no ts.
  • CLI: <cmd> help no longer fires a real workflow dispatch (#380, fixes #361).
  • dependabot-ci-fix: the repair commit is typed so it can’t cut a release (#382).

Packages

  • lastlight-core 0.31.0
  • lastlight 0.31.0
  • lastlight-evals 0.12.5

Full Changelog: https://github.com/nearform/lastlight/compare/v0.30.0...v0.31.0

v0.30.0 tiered PR review

View on GitHub

Tiered PR review (#378, #379)

A 73-file pull request was reviewed, a Merge branch 'main' landed eleven minutes later, and the whole pipeline ran again over a push that changed no line the author wrote. This release prices a re-review by how much the change actually moved.

A merge that adds nothing gets no review. review.skipUnchangedDiff (on by default) skips a re-review whose own three-dot diff base...head is byte-identical to the one already reviewed. That covers a merge from the base, a rebase that preserves the tree and an empty force-push, and it correctly does not fire when the base touched a file the pull request also touches. No path-based rule could catch this: on a merge commit the delta since the last review is every file the base brought in, which is exactly what the existing generated-only gate must refuse to suppress. The prior verdict is carried onto the new head as a completed last-light/review check, so a required check is never left missing.

A small delta gets a light review. review.triage (on by default) adds one cheap pass at the head of pr-review.yaml, on a re-review only. It reads the diff since the last posted review and answers full or light. On light the seven evidence-pipeline phases skip and the reviewer makes a single focused pass over the delta, naming the prior review’s SHA and verdict. The survey fan-out is roughly 75% of a review’s spend and 90% of its branch-seconds, so this is where the money is.

Every failure direction lands on a full review: a missing or unrecognised marker, a failed or skipped triage phase, a degraded diff read, a truncated compare. A first review of a pull request is untouched, and an explicit @last-light review, the request label and the check’s Re-run button all still force the full one.

Operator notes

  • models.review-triage ships pinned to anthropic/claude-haiku-4-5-20251001, like review-survey. A non-Anthropic deployment should override it, or the phase reaches for a provider you have no key for.
  • Turn either gate off with review.skipUnchangedDiff: false / review.triage.enabled: false.
  • Repo clamps: skipUnchangedDiff is downward-only (a repo may buy itself more review, never less); triage is operator-only, like analysis.

Versions

lastlight-shared  0.6.0  -> 0.7.0
lastlight-core    0.29.0 -> 0.30.0
lastlight         0.29.0 -> 0.30.0
lastlight-evals   0.12.3 -> 0.12.4

Full changelog: https://github.com/nearform/lastlight/compare/v0.29.0...v0.30.0

v0.29.0 point your models at your own LLM gateway

View on GitHub

Model endpoints are configuration now (#373)

Provider endpoints were compile-time constants, so running Last Light through a self-hosted or corporate LLM gateway meant patching source. Now it’s a providers: block in your overlay config.yaml:

providers:
  # A gateway speaking Anthropic's dialect: only the URL moves. Models, request
  # shape and the key env var are inherited.
  anthropic:
    baseUrl: https://gateway.internal/anthropic

  # A provider the registry has never heard of — self-hosted, Azure/Bedrock-fronted.
  acme:
    baseUrl: https://llm.corp.example/v1
    api: openai-completions        # or anthropic-messages
    envKey: ACME_API_KEY

That’s how a deployment gets central spend accounting, key custody, rate limiting and audit — and it’s the only way to reach a model that isn’t one of the ~18 registry entries.

One key reaches every model path. The registry is resolved once at config load, and the cheap screener/classifier helpers, the sandboxed workflow phases, in-process chat, and the sandbox egress allowlist all read the resolved value — so the firewall follows your gateway instead of denying it, with no separate allowlist edit.

Also supported: re-pointing envKey on a built-in (a gateway holding its own credential — leave it alone and pi keeps resolving the key itself, so OAuth subscription logins still work); ANTHROPIC_BASE_URL / OPENAI_BASE_URL-style env vars; and LASTLIGHT_PROVIDERS as a JSON map. lastlight server setup offers “Self-hosted / gateway endpoint” as a provider choice and writes the block for you.

Fails loudly, on purpose. A malformed URL, a custom provider with no baseUrl, or plaintext http:// off-loopback refuses to boot rather than warning — quietly ignoring an override would send your prompts and your API key to the vendor instead. http://localhost:… needs no opt-in; LASTLIGHT_ALLOW_INSECURE_PROVIDER_URLS=1 covers anywhere else.

Docs: Configuration → Provider endpoints and spec/02-configuration.

Packages

Package Version
agentic-pi 0.4.4 → 0.5.0 (--providers / AGENTIC_PI_PROVIDERS)
lastlight-shared 0.5.0 → 0.6.0
lastlight-core 0.28.2 → 0.29.0
lastlight 0.28.2 → 0.29.0
lastlight-evals 0.12.2 → 0.12.3

Full changelog: https://github.com/nearform/lastlight/compare/v0.28.2...v0.29.0

v0.28.2 the workflow kill switch is read before the first side effect

View on GitHub

Fix: disabling a workflow now really disables it

A workflow switched off in the admin dashboard still reacted 👀 on the triggering event, and still posted the last-light/review check. isWorkflowEnabled was read in exactly one place — runSimpleWorkflow, the bottom of the dispatch stack — and by the time it said “disabled”, dispatch had already acked the event and, for pr-review under review.trigger: after-checks, taken the skip path that posts the queued “Waiting for CI” placeholder.

That placeholder was then stranded permanently: the run that concludes a check is the run the switch drops, and the 30-minute sweep cannot post a superseding one (postReviewCheckForSkip is limited to the attention route). A deployment that requires last-light/review had its PRs left unmergeable by switching the reviewer off.

The switch is now read at each choke point, earliest first:

  • dispatch — before the 👀 ack, covering the webhook and Slack routes.
  • dispatchWorkflow — before the PR state machine spends its GitHub reads, covering cron, /api/run and resume.
  • runSimpleWorkflow — last, keeping the invariant stated where workflow_runs rows are created.

The five in-process handlers (chat, chat-reset, status-report, explore-reply, approval-response) are exempt — they have no definition to toggle. GitHub stays silent, which is the point of the switch; a messaging turn gets one line back, because answering a direct request with nothing reads as a broken bot rather than as a setting.

Full changelog: https://github.com/nearform/lastlight/compare/v0.28.1...v0.28.2

v0.28.1 the review talks about the change, and the pipeline view fits on screen

View on GitHub

Patch on v0.28.0. Two fixes, both found by looking at real output.

The review no longer names its own machinery

A posted review on cliftonc/drizzle-cube#891 opened: “This adjudication keeps those findings reconciled as not applicable and adds the hypothesis ledger…” — three internal terms in one sentence, on a pull request about a search box. Two separate sources:

  • Ours, written verbatim by code. _Reported here rather than inline by the adjudicating pass._ and _Below the reporting confidence bar for their family._ are literal strings in review-poster.ts naming a phase and an obligation family. They now say what happened in ordinary review language.
  • The model’s. adjudicate-pass gains a “who reads what you write” section, restated at the point of output in review-adjudicate.md and review.md: the maintainer did not build this pipeline, so summary / title / body are about their change in their vocabulary. The bookkeeping fields (family, obligation, confidence, hypotheses) are machine-read and never rendered — that is where it belongs.

An instruction is not a mechanism, so internalJargon is a pure predicate in the poster plus one warn in post-review. It reports and never rewrites: silently editing a review’s wording changes a claim nobody re-read, and the failure it catches is a prompt drifting rather than a one-off.

Fan-outs and loops are containers, not stacks

The v0.28.0 pipeline view drew a fan-out as a card with its branches stacked below it in the same column, each branch followed by its own gate card — on a real pr-review that is a ten-deep ladder running off the canvas and pushing the rest of the row out of frame. Loops had the same shape.

Both are now containers the children sit inside (React Flow parentId). The _check / _retry gate cards are gone and folded into the colour of the row they judge — a red until_bash tints its branch unmet rather than spending a whole card on a tone. The one real difference is preserved: a loop chains its iterations, a fan-out draws no internal edges at all, because a line between concurrent branches asserts an order the run does not have.

Cards are also wider (110 → 150px, so sentence-length labels stop wrapping to different heights) and start time + duration share one line.

Full changelog: https://github.com/nearform/lastlight/compare/v0.28.0...v0.28.1

v0.28.0 role-scoped review skills, and the ceiling made measurable

View on GitHub

Review pipeline

Each pass of the pr-review fan-out now reads one role-scoped skill, or none. survey reads survey-pass, adjudicate reads the new adjudicate-pass, and falsify stages nothing (the workspace layout is inline in its prompt). review keeps [pr-review, code-review] — it produces a review from a diff, which is where the precision gate is meant to fire — and code-review is unchanged on build’s branch-diff reviewer and dependabot-pr-merge.

For adjudicate this is a correctness fix, not a token one: three of pr-review’s contracts contradict what that phase must do — its output schema prescribes content fields only where adjudicate must carry tier/family/obligation/confidence/hypotheses/evidence plus dropped[]; its §1 instructs {"skip": true} on an already-reviewed head, which there discards five surveys and an oracle; and its confidence gate permits dropping a finding against code where adjudicate may delete only against a probe transcript.

Critical now needs a trust boundary, not a category. A finding must name the boundary the input crosses and a capability its supplier does not already have. Severity feeds inline ordering, so an inflated one spends a maintainer’s top slot on a hazard nobody can reach. The adjudicator demotes rather than drops.

The truncation notice renders again. It matched a prose substring and a rename had silently killed it, so a branch handed twelve of fifty-nine obligations was told nothing about the forty-seven it did not get. It now reads the structured families[].minted/cap, and obligations.json carries those fields.

Measurement

cap-sweep and ceiling-pressure (in apps/evals/scripts/) replay stored envelopes at any ceiling with no model calls. Across the eight gate cases the shipped ceiling names 21 of 25 gold and uncapped names the same 21 — so the per-family ceiling is not what is costing recall, and the four it misses are a seeding-shape gap rather than a budget one.

The skill subtraction does not validate as a recall win and cannot at this n: micro-recall 0.320/0.440 against a baseline band of 0.240/0.280/0.400, McNemar p >= 0.500 on all six pairings. What it does establish is that discovery is unchanged — hypotheses per repeat 290 -> 277, obligations identical at 187.

Also

  • Dashboard: tool I/O view, parallel fan-out branches, per-phase outcomes, run failure reasons
  • pr-review skips a PR the bot authored, and says so when asked; one place decides head-SHA dedup and explicit re-reviews post
  • lastlight CLI sends prNumber on PR-scoped triggers

Full changelog: https://github.com/nearform/lastlight/compare/v0.27.0...v0.28.0

v0.27.0 the review evidence pipeline

View on GitHub

The review evidence pipeline — Last Light’s PR reviewer stops guessing and starts proving.

The review evidence pipeline (#360)

A code review is now a multi-phase workflow (pr-review.yaml) rather than one agent pass:

  • lastlight-code-facts — a new published package (lastlight facts / lastlight-facts) that does deterministic PR analysis with ts-morph + ast-grep. It computes the obligations a diff creates — contracts, coverage, dependency edges, discharge — and hands the reviewer a seeded, fact-grounded starting point instead of a blank diff. It ships in the CLI, not just the sandbox image, so the eval harness can run it on the host under --sandbox none.
  • Seeded surveys — six parallel survey phases (contract, enforcement, security, spec, state, tests), each with its own prompt, fan out over the seeded facts.
  • A probe oracle and a falsification pass — candidate findings must survive review-falsify before they can be posted.
  • An adjudicated attention boundary — review-adjudicate decides what actually reaches the reviewer, with an anchor cascade so comments land on the right lines.
  • Polarity discipline — a correctness conclusion is a claim, not a measurement; the findings gate carries nullable advisory fields and an internal[] id-list.

Supporting work: a fanout workflow handler and phase-ref/schema support in lastlight-workflow-engine, a rebuilt review-poster, and the code-review 2.4.0 rubric.

Evals

The harness gained review-specific metrics, pipeline stats, repeat/band analysis for run-to-run variance, a match-v2 judge, per-phase model overrides, and a much richer dashboard (repeat view, phase sidebar, micro panel, session modal).

Also in this release

  • Dependency PRs blocked on a required review now escalate instead of stalling silently (#357).
  • TypeScript 7 across the workspace (#354) — dependency-cruiser was replaced by scripts/lint-import-boundaries.mjs, because dep-cruiser refuses to parse TS ≥ 7 and exits 0 anyway, leaving the boundary gate green while seeing nothing.
  • Batched Dependabot majors: tailwind 4, daisyUI 5, vite 8 (#353), plus production and dev dependency groups (#359, #350).

Versions

Package Version
lastlight (CLI) 0.27.0
lastlight-core 0.27.0
lastlight-code-facts 0.2.0
lastlight-workflow-engine 0.5.0
lastlight-shared 0.5.0
lastlight-evals 0.11.0
agentic-pi 0.4.4 (published from agentic-pi-v0.4.4)

Full changelog: https://github.com/nearform/lastlight/compare/v0.26.0...v0.27.0

v0.26.0 Drizzle state layer + sandbox dependency services

View on GitHub

The state layer moves from better-sqlite3 to Drizzle ORM with an async store API and dual-dialect support, and workflow phases can now run against repo-declared dependency services (a test postgres/redis/mssql in the sandbox).

This is a five-package release — lastlight-workflow-engine’s ports changed shape, so the bump cascades across the whole graph.

Drizzle state layer (#351)

  • better-sqlite3 is gone. The driver is @libsql/client + drizzle-orm/libsql — natively async, so the SQLite and Postgres code paths share one shape, and its prebuilt bindings let the agent image drop python3 make g++ from both build stages.
  • Every store method returns a Promise. new StateDb(path) → await StateDb.open(pathOrUrl). RunStore / ExecutionLedger / PhaseReporter in lastlight-workflow-engine are now Promise<T> — breaking for direct consumers of that package.
  • Migrations are journaled. The 530-line boot-time migrate.ts is deleted; __drizzle_migrations records what has been applied, so the data backfills that used to re-run on every boot now run once. Closes two of the three weaknesses in #345 (silent failed migrations; backfills idempotent only by convention). The downgrade direction remains unguarded.
  • Postgres runs the entire test suite via PGlite, with a pgTable mirror using real jsonb. Production stays on SQLite — StateDb.open() throws on a postgres:// URL and no PG driver is a runtime dependency.
  • New database.url config slot (DATABASE_URL env → overlay → default → file: + dbPath). Setting none of them is byte-identical to previous behaviour.

Verified against a real 41 MB production snapshot: migrations no-op’d in 96 ms, every row count held, nothing dropped, and the second open was idempotent.

Sandbox dependency services (#347)

A repo can declare services: in its .lastlight/ layer so a phase’s test suite gets a real database. Implemented for both the docker and kubernetes backends (native sidecars), with an operator image allowlist that permits nothing by default.

Upgrading

No action for a normal deployment — bump your overlay’s deploy.version to v0.26.0. The baseline migration is schema-neutral on an existing database (every CREATE no-ops; the only addition is __drizzle_migrations, which older code ignores), so rollback to the previous image reads the same file unchanged.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.9...v0.26.0

v0.25.9 the digest reports what happened

View on GitHub

The weekly repo digest now says what actually happened

It reported arithmetic — “6 PRs merged, 2 opened, 3 issues closed” — plus one narrative sentence forbidden from naming anything. It now lists the week’s merged pull requests, issues opened and issues closed, each linked, and spends the same one cheap model call on a themed summary over that text instead of over the counts.

The content comes from one GraphQL request (three aliased search queries carrying each item’s bodyText), separate from the counts and allowed to fail — if it does, the digest posts exactly what it posted before.

Three details that were easy to get wrong:

  • closingIssuesReferences reports links, not causes. GitHub returns every issue linked to a merged PR whether or not the merge closed it — on this repo, #344 lists #341, closed by hand three days earlier. Attribution now requires the issue to have closed within seconds of the merge, drops cross-repo refs, and counts only what it actually folded away.
  • The enrichment can never fail the tick. A failed repo fails the tick on purpose so a revoked token surfaces; that rule is right for facts and wrong for decoration. Octokit throws on partial GraphQL responses too.
  • callLlm defaults to maxTokens: 256. On a reasoning model that covers thinking as well as output, so a paragraph-length summary would have come back empty — silently, forever.

Also: untrusted titles and model output are escaped before any link markdown is added (<!channel> in an issue title is not markup a formatter strips); bot PRs fold to a count so a week of Dependabot can’t bury the human work; and Block Kit sections are clamped to Slack’s limits.

New config: digest.listItems (8) caps each content list, digest.detailItems (25, clamped) bounds what the summariser reads. digest.maxItems keeps its old meaning and default.

Also in this release

  • A fire-grain cron_runs ledger (#344) — every cron fire, scheduled or manual, handler or workflow, writes one row keyed on the cron’s name, so a cron that dispatches nothing is no longer invisible to failure alerting.
  • Dependency bumps across 3 directories (#348).

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.8...v0.25.9

v0.25.8 digest previews, k8s retry, CLI repo scoping

View on GitHub

Four fixes.

Slack digest: no more preview cards

Linking the PR numbers in v0.25.7 made the digest actionable and then unreadable — Slack expanded each citation into a preview card, so six lines of summary arrived buried under GitHub chrome. The digest now posts with unfurling off (both unfurl_links and unfurl_media, which govern different target types).

Opt-out per message, not a connector default: a conversational reply where someone shares a link still gets its preview. (#342)

lastlight security / health targeted the wrong field

Both CLI commands sent repos where the dispatch path reads repo, so a repo-scoped run never received its target. (#340)

A failed phase now logs at error level, with its cause

It was logged at info with the cause dropped, so the one line you’d search for after a failure was both invisible to a level filter and missing the reason. (#337)

Kubernetes: a retry can reclaim the previous attempt’s pod

A retried phase collided with the pod its earlier attempt left behind and could not proceed. (#338)

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.7...v0.25.8

v0.25.7 digest PR links

View on GitHub

The weekly repo digest printed #294 as plain text, so acting on “2 PRs waiting on a human: #294, #282” meant retyping the number into GitHub. Every number it prints is now a link.

• Oldest unreviewed: #46 Update postgres Docker tag to v18 — open 321 days
• ⚠️ 2 PRs waiting on a human: #294, #282

…where each #N now opens the pull request. Both the notification fallback text and the message blocks carry the links, since they render from the same lines.

The owner/repo heading is unchanged and stays plain text — Slack header blocks are plain_text only and cannot carry a link.

No configuration change and no action on upgrade.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.6...v0.25.7

v0.25.6 the weekly Slack repo digest

View on GitHub

The weekly repo digest

A per-repo Slack summary of what happened and what Last Light did about it, posted Monday 09:00 (repo-digest cron).

*acme/widgets — week to 9 Aug*

A steady week: seven PRs landed and the review queue shrank.

*Repo*
• 7 PRs merged, 3 opened
• 5 issues closed, 4 opened
• 9 PRs open (2 awaiting review)
• Oldest unreviewed: #412 Refactor the loader — open 9 days

*Last Light*
• 14 runs — 12 ok, 2 failed
• reviewed 6 · fixed CI on 3
• $4.12 across 40 phases
• ⚠️ 2 PRs waiting on a human: #401, #408

The numbers are computed in code, from the GitHub API and the harness’s own SQLite — not written by a model. One cheap model call turns those settled facts into the opening sentence, and a failure there drops the sentence rather than the digest (digest.narrative: false skips it entirely).

It is inert until you name a channel, most specific first:

  1. the repo’s own .lastlight/lastlight.yml → notifications.slack.channel
  2. the operator’s slack.repoChannels: { "owner/repo": "C…" } (overlay config.yaml)
  3. SLACK_DELIVERY_CHANNEL

No channel means no post, no GitHub request and no model call — so upgrading changes nothing until you configure one. Remember to invite the bot to the channel; Slack refuses delivery to a channel the app was never added to.

handler: — host-side crons

A cron YAML now declares exactly one of workflow: (dispatch an agent workflow) or handler: (run host-side code). The digest needs the latter: half its content lives in the harness’s own database, which a sandboxed phase cannot reach, and it posts to Slack, for which there is no agent tool. Handler crons are first-class — dashboard toggle, schedule override, per-repo participation, and lastlight cron trigger repo-digest.

Every handler tick writes an executions row, so consecutive-failure alerting and the dashboard’s failure count work for a cron that dispatches nothing.

Removed

MessageDeliveryService and SlackConnector.sendToDeliveryChannel were registered at boot and never called from anywhere — no call site in the codebase, none in git history — while four documents described SLACK_DELIVERY_CHANNEL as the channel cron reports go to. The digest makes that sentence true; the dead wiring is gone. slack.deliveryChannel itself stays, and is now finally what it always claimed to be.

Also removed: the deliverSlackSummary flag and two skills’ claims that the harness routed their reports to a channel.

No action needed on upgrade unless you were relying on SLACK_DELIVERY_CHANNEL — in which case it never worked, and now does.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.5...v0.25.6

v0.25.5 stop the review sweep re-reviewing bot PRs

View on GitHub

Fix: the review sweep’s unbounded spend loop

check-prs-awaiting-review re-dispatched the same bot-authored PRs every 30 minutes indefinitely. On the nearform instance that was 1260 pr-review:review executions with 0 reviews ever posted, at roughly $1.30/hour. Two independent gaps compounded:

Discovery’s bot filter was narrower than the webhook’s. review-discovery.ts dropped only authorLogin === botLogin, while the webhook route drops any author ending [bot]. An instance running as nearform-lastlight[bot] on repos carrying last-light[bot] PRs matched nothing — so the webhook dropped those PRs and the sweep offered them. Discovery now uses the webhook’s predicate verbatim.

The per-head dedup could never latch. resolveReviewTrigger keyed “already reviewed” on botReviewAtHead — a posted review. A run that completes without posting one left no trace, so the sweep re-dispatched the same SHA forever. This is general, not specific to self-review: any pr-review run that posts nothing hit it. The resolver now also skips on assessedHeadShaByWorkflow["pr-review"] === headSha.

@bot review, the request label and the check’s Re-run button still force a review — both gates sit below the explicit-request branch. Only succeeded runs count as an assessment, so a crashed run is still retried, and a push clears it.

Also includes four dependabot merges (dompurify, concurrently, daisyui, typescript).

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.4...v0.25.5

v0.25.4 pi-ai 0.84.1

View on GitHub

Patch release. Dependency bump only.

  • chore(deps): bump pi-ai + pi-coding-agent to 0.84.1 — moves off the ^0.82.1 range (a 0.x caret pins to <0.83.0). pi-ai 0.83 made the AbortSignal required on the provider-facing auth surface (ProviderAuthInteraction = AuthInteraction & { signal: AbortSignal }, and refresh(credential, signal)), so three call sites in lastlight-shared’s oauth.ts now pass a never-aborted controller — preserving the previous no-cancellation behaviour rather than synthesising a timeout.

Versions: agentic-pi 0.4.3, lastlight-shared 0.3.6, lastlight-core 0.25.4, lastlight 0.25.4, lastlight-evals 0.9.19. lastlight-workflow-engine (0.3.4) is unchanged.

Note: this does not change which models resolve. openai/gpt-5.6-{sol,terra,luna} already resolved under 0.82.1, and bare openai/gpt-5.6 resolves under neither — that is an invalid id in a deployment’s own config, not a stale pi.

Caveat: the migrated code is the subscription-login path (Claude Pro/Max, ChatGPT Codex, GitHub Copilot). Types check and the suite is green, but no test exercises a real OAuth handshake.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.3...v0.25.4

v0.25.3 signed-commit publish, triage routing, observable sandbox sweep

View on GitHub

Patch release. Three PRs since v0.25.2.

  • fix(router): triage needs an issue, and needs to be told what was asked (#290, #291) — triage no longer dispatches without an issue in context, and the ask is carried through to the run.
  • feat(cron): make sandbox-sweep observable (#282) — the sweep reports what it reclaimed instead of running silently.
  • feat(github): publish signed commits via createCommitOnBranch (#292) — the agentic-pi GitHub extension publishes through the GraphQL commit API, so agent commits are verified/signed; adds worktree-diff.

Versions: agentic-pi 0.4.2, lastlight-core 0.25.3, lastlight 0.25.3, lastlight-evals 0.9.18. lastlight-workflow-engine (0.3.4) and lastlight-shared (0.3.5) are unchanged.

Deploy: bump the overlay’s deploy.version to v0.25.3.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.2...v0.25.3

v0.25.2 chat says what this deployment can actually do

View on GitHub

Asked “what do you do?” in Slack, the bot recited a hand-written constant listing five triggers. It had drifted, and structurally could not do otherwise: CHAT_SYSTEM_SUFFIX was a literal concatenated into the system prompt once at boot, with no reference to the workflow registry.

verify and qa-test shipped chat-routable and unadvertised for several releases. An overlay that adds a workflow got a routable intent for free (#164) but no way to be mentioned in chat. And nothing stopped the bot naming a workflow an operator had switched off — typing that trigger no-ops silently at dispatch.

The chat prompt is now composed from the enabled workflow set, the way the classifier prompt has been since #164: a forkable base template (workflows/prompts/chat-system.md) plus one new chat: block per workflow, filtered through the dashboard kill switch and re-derived per turn.

Highlights

  • New chat: block on the workflow schema (trigger? / summary / deflect? / reply?). An explicit opt-in, not a derivation from classification: — the classifier tagging an intent is not the same as a human being told to type it, and the two diverge in both directions.
  • Composed per turn, so the dashboard’s per-workflow kill switch takes effect without a restart.
  • Three duplicate lists removed — skills/chat/SKILL.md’s copy of the triggers; triggers.ts’s five-workflow Slack list (the dashboard showed no Slack trigger for verify / qa-test / demo / answer); and CHAT_SKILL_NAMES, now chat: true frontmatter resolved through the asset layer stack, so an overlay can finally add a chat skill or override a built-in.
  • demo’s dead Slack route fixed — it shipped with a classification block and a routes.slack.demo entry but no branch in the Slack switch, so every demo-classified message fell through to plain chat against a configured route.
  • The dependency workflows no longer half-dispatch from Slack — both are pr_scoped and reach handlePrFix via context.prNumber, which no Slack branch sets. A new GITHUB_ONLY_INTENTS set, consulted only by the message case, routes them to chat instead. Deliberately not a WELL_KNOWN_INTENTS entry: the GitHub comment ladder reaches them through that same fallback.
  • A fifth copy of the routes table pinned — defaultRouteConfig() had drifted from config/default.yaml (missing verify / qa_test / demo), which silently removed those workflows’ @bot mention rows from the trigger table. Now enforced equal by test.

Packages

lastlight-workflow-engine 0.3.4 · lastlight-shared 0.3.5 · lastlight-core 0.25.2 · lastlight 0.25.2 · lastlight-evals 0.9.17. agentic-pi is untouched at 0.4.1.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.1...v0.25.2

v0.25.1 the repo-scope filter finds its own repos, and its own escape hatch

View on GitHub

A patch on v0.25.0, fixing two things about the optional “my repos” dashboard filter (teamVisibility). Both surfaced while turning it on for the first time on a live instance.

The control hid itself in the one state where it was needed (#287)

It rendered only once real team grants resolved. So the ordinary path in — enable the feature, see no toggle, go create a team and grant it repos — still showed nothing, because the “no teams” answer is cached for ttlMinutes (60 by default). And visibleRepos() is stale-while-revalidate, so even waiting out the hour serves the old answer once more: it took two page loads to see a change.

Meanwhile POST /admin/api/me/repos/resync and the SPA hook’s own resync() both existed, and nothing in the UI ever called either. The escape hatch was unreachable from the state it was built for.

The control now renders whenever teamVisibility.enabled is on, and the unresolved states are explanatory and retryable — the button names why (no teams, no GitHub identity, over budget, GitHub error) and clicking it re-syncs. enabled: false remains the one state it isn’t drawn in, and is the one a re-sync provably can’t change, since resync() short-circuits on that flag before reaching GitHub.

The filter now counts the repos you OWN (#287)

Teams are an org concept, so a repo under a personal account can never be granted by a team — and a strictly team-derived filter therefore hid every personal repo an instance managed. On the deployment this was found on, that was nine of eleven repos. The gap was an artefact, not the deliberate approximation the rest of the feature makes: team grants stand in for involvement, and nobody is more involved in a repo than its owner.

The test is owner === login, not “the owner is not an organization” — the second picks out the same set on a single-owner instance and is a disclosure bug on any other, putting everybody else’s personal repos into your filter. It costs no extra API call and is derived per request rather than cached, so it re-intersects with the live managed list and can’t go stale.

An incomplete team answer (truncated / error) is deliberately not rescued by it: “your own repos plus an unknown fraction of your teams’” would confidently hide the org repos the failed half would have contributed, so those still fail open.

The label follows the set — “my teams” → “my repos” — in the UI, config/default.yaml, the spec and the docs site.

⚠️ Nothing here narrows anything that didn’t narrow before. The filter is still opt-in, still off by default, still remembered per browser, and every failure path still fails open. The ownership union only ever adds to the allow-list.

Versions

lastlight-core 0.25.1 · lastlight 0.25.1 · lastlight-evals 0.9.16. lastlight-workflow-engine (0.3.3), lastlight-shared (0.3.4) and agentic-pi (0.4.1) are unchanged.

Upgrading

Bump your overlay’s deploy.version to v0.25.1 and push.

Full changelog: https://github.com/nearform/lastlight/compare/v0.25.0...v0.25.1

v0.25.0 repo-identifier convergence, honest execution reads, resilient CI reads

View on GitHub

Four changes, three of them fixes for reads that quietly returned less than they claimed to.

Highlights

One repo-identifier rule (#279). Storage converges on owner + a bare repo across workflow_runs, executions, feedback_anchors and feedback_signals; everything user-facing speaks the qualified owner/repo, and src/state/repo-ref.ts is the only place that join is expressed. Previously the same column meant two things depending on who wrote it, and six read sites each re-derived “may be bare or qualified” and disagreed — which is how #278 shipped a filter that compared lastlight against {nearform/lastlight} and hid rows rather than showing them. An idempotent backfill in migrate() converges both tables on boot.

ExecutionRecord reads now populate their fields (#285). The executions table is snake_case and the record is camelCase, so SELECT * cast to ExecutionRecord[] type-checked while leaving issueNumber, startedAt and workflowRunId undefined. The Slack status report rendered (started undefined) on every status request, and the admin cancel loop filtered on a workflowRunId that could never match a row. All four record-returning reads now share one aliased column list and one row mapper.

⚠️ GET /admin/api/executions changes its wire shape from snake_case to camelCase as a result. No dashboard component or CLI command reads that endpoint today (the CLI uses /workflow-runs/:id/executions), so nothing renders differently — but note it if you consume the admin API directly.

A missing Commit statuses: read no longer blinds the whole CI read (#277). getChecksSummary reads check runs and the combined commit status together, needing two different App permissions — and Checks: read does not imply Commit statuses: read. Under Promise.all a 403 on the status leg discarded the check-runs result the App was permitted to read, so an App missing only statuses lost its entire CI signal while runs still recorded success = true. Now Promise.allSettled: the status leg degrades to “no status contexts” and the ref is judged on check runs alone.

👉 Grant your App Commit statuses: read. It was missing from the setup docs entirely, so an App built to the documented table was broken by construction. This release adds it to the setup page, the landing page, the README and the lastlight-server skill. It is optional — without it CI is judged on check runs alone — but you want it if any of your CI reports via classic statuses (CircleCI, Jenkins).

Dashboard: optional “my teams’ repos” filter (#169, #278). A GitHub-authenticated admin can scope the dashboard to the managed repos their org teams can reach. Dormant unless you grant the org Members: read permission, subscribe to team/membership/organization, and set teamVisibility.enabled: true — and it fails open at every read path, so without those it is exactly today’s behaviour.

Versions

lastlight-core 0.25.0 · lastlight 0.25.0 · lastlight-workflow-engine 0.3.3 · lastlight-shared 0.3.4 · lastlight-evals 0.9.15. agentic-pi unchanged at 0.4.1.

Minor rather than patch because of the dashboard capability above. No breaking changes beyond the admin-API shape note.

Upgrading

Bump your overlay’s deploy.version to v0.25.0 and push — the Deploy Action pins the host CLI and pulls the images for you.

Full changelog: https://github.com/nearform/lastlight/compare/v0.24.5...v0.25.0

v0.24.5 three ways a review stopped stranding at `queued`

View on GitHub

Three independent defects, each of which could leave a PR with a last-light/review check queued forever — and on a repo that makes the check required, an unmergeable PR.

Superseded check re-runs (#280)

checks.listForRef?filter=latest de-dupes per check suite, not per check name, so a job re-run in a fresh suite came back alongside the red attempt it replaced — and the aggregate stayed failing for the life of the SHA, whatever the re-run said. Every check read now collapses to the latest run of each (app, name). Latest-wins is what branch protection does anyway.

getCiFailureReport gets the same collapse, so the fix agent is no longer handed a red job that has since gone green.

Settle on either colour (#280)

The red and green check_suite.completed branches each ran their own aggregate test, and the green one demanded passing. A PR with one failing job whose sibling suite finished green afterwards was dropped by both. Which event fires is now decided by the aggregate alone; the suite’s own conclusion only decides whether an aggregate read is worth paying for. Fix-outranks-review precedence is preserved.

Fork PRs (#283)

GitHub fills check_suite.pull_requests[] only for same-repo PRs, so a fork PR’s checks could never find their PR. Invisible until after-checks became the default in v0.24.4 — under eager the review fired from pr.opened, which always carries the PR. An empty array now falls back to asking the base repo which open PR the commit heads, filtered to head.sha matches that are open. Both Re-run buttons get it too.

No new policy: reviewing fork PRs is what the harness already did under eager, and GitHub’s maintainer-approval gate on fork Actions is inherited for free — no approval means no checks, so nothing ever settles.

Full changelog: https://github.com/nearform/lastlight/compare/v0.24.4...v0.24.5

v0.24.4

View on GitHub

Patch on top of v0.24.3.

Fixed

  • Feedback chart had two zeros (#276). The score chart carries two Y scales — counts on the left, average score on the right — but the reader sees one horizontal zero. The count axis had no explicit domain, so recharts auto-fitted it ([-1, 3] for one 👍 and one 👎), putting its zero near the bottom of the plot while the score axis’s symmetric [-2, 2] put its zero in the middle. Bars were drawn correctly from their zero but appeared to float below the visible gridline. The count axis is now pinned symmetric, so both zeros are the same pixel by construction, with an explicit ReferenceLine drawing it.

Dashboard-only; no server behaviour changed.

v0.24.3

View on GitHub

Feedback signals from emoji reactions (#255)

A 👍 or 👎 on something Last Light wrote now becomes a scored eval signal against the specific workflow run that wrote it, so the effect of a prompt or skill change on quality is measurable instead of felt. Analytical only — nothing reads a signal back into the agent’s behaviour.

Score GitHub Slack also accepts
+2 🎉 · 🚀 · ❤️ tada, heart_eyes
+1 👍 · 😄 smile, smiley, grinning
0 👀 (recorded, unscored) eyes
-1 👎 thumbsdown
-2 😕 disappointed, cry, sob

👀 is deliberately unscored and excluded from every average — it’s Last Light’s own ack emoji, so counting it as criticism would poison the dataset.

Surfaces: a Feedback tab (score over time, per-workflow leaderboard, raw feed), a per-run badge, four admin read endpoints, and — with OTel enabled — a lastlight.feedback.signal span parented on the original run’s span, so the score lands on the trace it grades rather than a disconnected one.

Upgrading

Slack needs a re-consent. Add the reactions:read bot scope and the reaction_added / reaction_removed event subscriptions (both are in deploy/slack/slack-manifest.json), then reinstall the app. Until you do, Slack delivers no reaction events and the feature is silently dormant — nothing else is affected.

GitHub reaction polling ships OFF (feedback.github: false). GitHub sends no webhook for reactions, so that half is a poller and is opt-in. When you do enable it, the cost is bounded by the data rather than the schedule: individual bot comments (never issues), retired after windowDays, refreshed through one batched GraphQL query per 100 anchors — measured at one rate-limit point per request.

No action needed beyond bumping deploy.version; the two new tables are created by the additive migration on first boot.

Also

  • Chat replies are thumbable too, attributed to their messaging session and reported under chat.
  • WorkflowRunStore’s terminal observer became a list — it was a single slot owned by the last-light/review check, and a second registration would have silently unhooked it.

Full notes: #275. Closes #255.

v0.24.2 Slack threads keep their context; questions stop burning a sandbox

View on GitHub

Slack threads keep their context across workflow-answered turns

A Slack thread is one conversation however each message was handled, but only the chat runner ever wrote to messaging_messages. A message the classifier routed to a workflow — answer, build, explore, a router refusal — left the thread’s history untouched, so the next chat turn in that same thread rehydrated nothing and replied “there’s no prior context in our conversation… this appears to be the start of our session”.

The transcript is now written for every messaging path the chat runner doesn’t own, wired once at the dispatch choke point so it covers every Slack-answering workflow (including ones an overlay introduces). Two further bugs fixed alongside it:

  • Session staleness — touchSession was likewise chat-only, so a thread carried by workflow turns lapsed past the 30-minute window and silently re-keyed to a fresh session mid-conversation.
  • getHistory returned the oldest 50, not the newest — a long thread permanently rehydrated its own opening and never saw what was just said.

QUESTION is now a capability test, not a seriousness one

The answer workflow provisions a sandbox; chat replies in-process in seconds — and chat already reads repos, issues and comments, PRs and their diffs, file contents, commit history, and searches code. The classifier category now admits only what chat genuinely cannot do: the web (comparisons with other tools, upstream docs, current external facts) and real exploration of a checkout. Repo-scoped questions like “what does this workflow do?” or “what changed in PR #273?” are answered by chat.

A newly-opened GitHub issue that asks a question still always runs answer — an issue has no chat surface, so downgrading there would file questions as work items.

Full changelog: https://github.com/nearform/lastlight/compare/v0.24.1...v0.24.2

v0.24.1 link App installations from the admin console

View on GitHub

Follow-up to v0.24.0. The Config → Managed repos pane listed each App installation by account and id, and left you to construct the GitHub URL by hand — which is the next click from every question the pane raises.

  • Each installation now deep-links to its GitHub settings page (repo grant / suspend / uninstall). The path shape depends on the account type and guessing wrong 404s — an org install is /organizations/<login>/settings/installations/<id>, a personal one a viewer-scoped /settings/installations/<id> — so installationSettingsUrl() builds it server-side from account.type and returns nothing when that type isn’t known yet, rendering plain text instead of a link that may not resolve.
  • Webhook-learned installs are linkable too. The directory now records account.type off the payload and backfills it onto a record seeded by an earlier event that lacked it.
  • The uninstalledOwners warning offers the fix, not just the diagnosis: it links the App’s install page via a new appInstallUrl (derived from botName, which is the App slug).

GET /admin/api/managed-repos gains installations[].htmlUrl and appInstallUrl.

Full changelog: https://github.com/nearform/lastlight/compare/v0.24.0...v0.24.1

v0.24.0 multi-installation GitHub App support

View on GitHub

Multi-installation GitHub App support

A GitHub App is installed per account, and each installation mints its own tokens. Last Light threaded one statically-configured GITHUB_APP_INSTALLATION_ID into every mint and every App-authenticated Octokit — so an App installed on a second account failed on every path: token mints returned 422 There is at least one repository that does not exist or is not accessible to the parent installation, and every harness-side call (comments, reactions, check runs, .lastlight/ fetches, post-review) 404’d.

One instance now serves every account its App is installed on, resolved per repository owner.

What changed

  • InstallationDirectory (apps/server/src/engine/github/installations.ts) is the single owner→installation authority. Fed by every webhook’s payload.installation and by GET /app/installations under an App JWT. Concurrent misses share one request and negatives are cached briefly, so a cron fan-out over N repos costs one call, not N. A suspended installation is withheld from resolution (it 403s every mint) but stays listed and flagged.
  • GitHubClient resolves its Octokit per owner, memoized per installation. Every method already took owner first, so DispatchDeps, the router, dispatcher, pr-state, review-check and repo-config are unchanged. The read-only chat tools get the same treatment; the two search tools read the account from the query’s repo: / org: / user: qualifier.
  • The per-run token mint resolves from githubAccess.owner. An owner with no usable installation fails the phase immediately, naming the account — no sandbox, no API call.
  • Installation repo discovery is keyed by installation id. Those webhooks are per-account: against one flat set, a second org’s created reset the managed list to just that org and its deleted cleared it entirely. suspend / unsuspend are handled too — they previously fell through silently.
  • GET /admin/api/managed-repos reports every installation (account, id, repo count, selection, suspended) plus uninstalledOwners. Config → Managed repos renders them and warns when a managedRepos owner has no installation — visible before it becomes a failed run.

Upgrade notes

GITHUB_APP_INSTALLATION_ID is now optional. Installations are discovered from the App JWT; the env var carries no account, so it is used only when that lookup itself fails (network, revoked PEM). An existing single-installation deployment keeps its exact behaviour with no .env edit — nothing to do on upgrade.

To manage repos across several accounts, install the App on each one and list their repos in managedRepos. There is nothing else to configure.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.8...v0.24.0

v0.23.8 severity is a field, not a substring

View on GitHub

Three contributions from @robinbowes, plus the release plumbing that makes forks usable.

Structured JSON logging with explicit levels (#258). Operational logging in lastlight-core was ~320 raw console.* calls emitting unstructured text, so a log pipeline had to guess severity from a substring. All of them — plus the nine in the workflow engine — now go through a leveled logger that emits one JSON object per line to stderr: level, component, msg, err ({message,stack}), time, and trace_id/span_id stamped from the active OpenTelemetry span so Loki and Tempo correlate. The seam is a dependency-free LoggerPort in lastlight-workflow-engine — the deepest leaf of the graph — injected through the existing EnginePorts and defaulting to a noopLogger, so the port is optional and nothing breaks by omission. The concrete pino adapter is confined to apps/server, which is what keeps pino out of the published lastlight CLI’s dependency tree; the existing dep-cruiser boundary enforces it. Levels were triaged rather than mapped mechanically — per-turn and hot-loop tracing dropped to debug, once-per-trigger stayed info.

Operator note: logs are now JSON on stderr, where they were previously mixed plaintext across stdout and stderr. docker logs still shows everything, so the admin log endpoints and lastlight logs are unaffected — but a plaintext grep pipeline should move to a JSON-aware view. LOG_LEVEL (default info) and LOG_FORMAT (json | pretty) are read straight from the environment; the format defaults to pretty on a TTY and json otherwise, so containers get JSON with no configuration. stdout is now reserved exclusively for the sandbox NDJSON protocol.

publish.yml is fork- and branch-agnostic (#260). The images job hardcoded ghcr.io/nearform, so a fork had to patch the workflow to build anything. It now targets ghcr.io/${{ github.repository_owner }} — identical on upstream, the fork’s own namespace on a fork — and the tag input became optional: a workflow_dispatch on a branch with no input builds that branch and tags the image <branch>-<short-sha>. Release behaviour is unchanged, and the npm job stays gated to release events, so a fork dispatch can never publish to npm.

Raw NUL bytes removed from three source files (#263). interventionKey() and the test-support ledgerKey() joined composite keys with a literal 0x00 separator, and a parser test fed a raw NUL+SOH pair as adversarial input. Embedding the byte made git treat those files as binary: no text diff, ripgrep skips them silently, and a three-way merge is refused outright — so a rebase manufactures a conflict on files whose changes never overlapped. That is not hypothetical; merging #258 and #263 together hit exactly that refusal on fakes.ts. Both keys are only ever compared for equality in memory, so JSON.stringify of the field tuple is unambiguous without any separator-collision assumption, and the test input is byte-identical written as \x00\x01. No behaviour change.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.7...v0.23.8

v0.23.7 never satisfy a check by disabling it

View on GitHub

The fix loop could satisfy its own gate by suppressing the thing the gate verifies — ignoreBuildErrors: true, a shimmed NODE_OPTIONS, a deleted test — and then honestly report green, because the check no longer existed. Two changes close it, one norm and one gate.

A hard rule in agent-context/rules.md. Never satisfy a check by weakening, suppressing, narrowing, bypassing or removing the check itself. It lives there rather than in the fix prompts because that file is concatenated into AGENTS.md for every agent session — so it covers build, pr-fix, dependabot-ci-fix and security-feedback uniformly — and a repo’s own agent-context/*.md is additive-only, so no managed repo can shadow it. Mechanisms are named across ecosystems, not just JS. It carries the test to apply (would the original failure still be caught?), the honest exit (outcome=gave-up, flag for a human — stopping cleanly is a correct outcome), and a carve-out: repairing a genuinely-wrong check is legitimate — an env-mismatch repair aligning CI to reality is the fixing skill’s own advice — but it must be declared prominently, never done quietly.

dependabot-pr-merge becomes the backstop. Suppression turns the checks GREEN, so pr.checks_failed never fires and the fix family never sees the PR again — it routes to the merge workflow, the single owner of the merge decision. STEP 1e scans the file list STEP 1a already fetched for CI/pipeline definitions, test files, type/lint/build config, hooks and manifest scripts, then opens at most the one matching file’s patch: filenames, not diffs, so it survives that prompt’s “never pull giant diffs” constraint. STEP 2 makes “nothing weakens how the repo verifies itself” a conjunct of the TRIVIAL test — the one thing between a PR and github_enable_auto_merge. It holds however green the PR is (the green may be a product of the change) and whoever authored the commit, including Last Light’s own fix commits. The verdict is FUNCTIONAL — routed to a human — not a refusal.

Deliberately not shipped: a suppression detector. A regex table over the diff enumerates known cheats, so it is open-ended and biased toward whichever ecosystem wrote it — the same arms race 09-state-machine.md §S1 already declined for the gate script, moved to the diff. Both halves here are prompt-level: a model that ignores STEP 1e reports a clean assessment and the harness reads it as truth, exactly as autoMergeMaxImpact is already policy-the-agent-is-asked-to-honour rather than a code-enforced ceiling. build and pr-fix land code by other routes, where rules.md is the only protection.

Closes #264.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.6...v0.23.7

v0.23.6 the fix loop stops re-running CI

View on GitHub

Three defects found by watching v0.23.5 un-stick a real PR (cliftonc/drizzle-cube#1016) in production.

The gate ran twice, and the second run could not change anything

Run 49c101aa: the agent pushed at 11:03:31 having already run the repo’s whole suite itself. GitHub’s checks went fully green at 11:06:25. The harness spent 11:04:00–11:10:48 in a fresh container re-running the same suite a third time before agreeing — a third of the run, spent re-proving something the real CI had already answered.

until: is evaluated before until_bash and short-circuits it, so outcome=pushed now ends the loop. Once the commit is on the branch, GitHub is the strictly better authority: the real CI environment rather than a sandbox approximation, warm rather than a cold container, covering matrix legs the sandbox cannot reproduce.

It short-circuits on pushed only. no-change and gave-up still pay for the gate — nothing was pushed, so there is no new commit, no new check run and no external authority at all; the local gate is the only evidence that exists and its RED verdict is what earns the agent its next iteration.

What this gives up, stated plainly: after a push the harness no longer independently checks the agent’s self-reported gate=green. That gate ran after the push and therefore never gated it — the self-report was already the only thing between a bad fix and the branch, on every run this workflow has ever done. What actually catches a bad fix is untouched: red checks re-dispatch the fix family, bounded by fix.maxAttempts / fix.maxCostUsd.

The gate the agent wrote was a CI clone, because we asked for one

The fixing skill said the gate is “whatever CI runs” and both fix prompts said “mirror CI”, with nothing bounding cost or scope. So for a merge conflict — fixed by merging main and regenerating a lockfile — the agent faithfully wrote five builds, two full test suites, and three docker branches that can never run (there is no docker in the sandbox).

The instruction now asks for the narrowest command that would have failed before the fix and passes after it, targeting under two minutes, with explicit exclusions: not the whole suite, not a check already watched passing this session, nothing that starts a service, nothing that mutates git state.

Writing no gate stays wrong — it burns the remaining iterations and forces gate=skipped, which is RED and never authorises a push — so the nothing-to-verify case gets an ordered fallback: the coherence check the repair implies, else an honest one-line exit 0 saying why.

A third of the run was invisible

The until_bash window was not a phase, had no ledger row, no pipeline node and no timer — and the iteration’s own phase entry was withheld pending its verdict, so currentPhase read diagnose (twelve minutes stale) while every surface showed a run that looked finished but stuck.

Iterations now persist when their work finishes, and the check gets its own <phase>_iter_N_check row: open with a start time while in flight, closed with a duration and a condition_met / condition_not_met verdict. success records whether the check ran, not what it said — a red gate is the loop working as designed, and renders muted rather than failed.

Also

  • A paused explore socratic round used to be recorded nowhere at all (the interactive gate returned above the persist call). It now lands in history before the pause.
  • A push followed by a red or exhausted gate no longer posts “Couldn’t auto-fix — leaving it for a human” on a PR it had in fact just fixed.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.5...v0.23.6

v0.23.5 stuck-PR recovery

View on GitHub

Makes a stuck pull request recoverable by a maintainer, and makes it get stuck far less often in the first place.

The base a fix merges is no longer stale

prePopulateWorkspace refreshed origin/<base> on the fresh-clone and per-PR-reuse paths, but returned early on the same-run path — so the base ref was frozen from the run’s first phase and a fix phase merged a base tens of minutes old. On a repo taking several dependency bumps a day the merge lands a superseded base, leaving the PR dirty. GitHub cannot compute a merge ref for a conflicting PR, so no pull_request workflow runs at all — not failed, absent — and checksState then reads green off whatever commit-status app is left.

It is now refreshed on every provisioning path, on both the host and the k8s init-clone script (which mirrored the same early exit). Safe on the preserve path because it writes remote-tracking refs only — never HEAD, the index, or the working tree. The base is also added to remote.origin.fetch, so the fix prompt’s own git fetch origin <base> stops being a silent no-op.

A new hold label: lastlight-ignore

Apply it to any issue or pull request and Last Light stops acting on that subject entirely. It outranks every guard except “we could not read the PR at all”, including an explicit @<bot> request — which gets exactly one reply naming the label, so a direct instruction is refused rather than ignored. Remove it and the bot resumes, with no record to clear.

Operator-configurable as hold.label / LASTLIGHT_HOLD_LABEL.

Behaviour change: requires-human no longer suppresses dispatch. It is now purely a notification the bot writes, and nothing in the code reads it — which is what its own module header always claimed. If you have been hand-applying requires-human to keep the bot off a PR, use lastlight-ignore instead. The old behaviour only ever worked on a PR the bot had never touched; on any PR it had worked, a hand-applied label already read as the bot’s own escalation.

Four ways to un-stick an escalated PR

Previously exactly one worked — pushing a non-bot commit — and two of the other three posted a duplicate escalation comment. All four now re-arm the attempt counter and the cost budget identically:

  • push a commit (unchanged)
  • comment @<bot> retry [reason] — the reason is recorded and handed to the next attempt
  • remove requires-human
  • lastlight pr retry <owner/repo#N> [reason], over the new POST /admin/api/prs/:owner/:repo/:number/retry

The agent’s journal now survives a retry (with a seam line marking the boundary) while a push still wipes it — a push changed the code, a retry changed only patience. Model escalation moves off attempt to priorAttempts.length, so a retry no longer downgrades the model on a PR that has already failed three times.

escalatePr gained a same-head dedup, so no bypass can produce a second escalation comment under any ordering, and the escalation comment now names only exits that actually work.

Note for operators

Retries are unbounded and re-arm the full window by design, so fix.maxCostUsd is now a futility guard rather than a spend ceiling. A server-level spend cap is tracked in #261 and is the only remaining backstop.

Plan, locked decisions and execution notes: docs/plans/stuck-pr-recovery/

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.4...v0.23.5

v0.23.4 the fix loop's local gate actually runs now

View on GitHub

Two independent, systemic bugs that between them meant the fix loop’s push gate has never once been enforced by the harness. Both affect pr-fix and dependabot-ci-fix.

The gate never ran

until_bash invoked the gate as sh .git/lastlight-verify.sh. /bin/sh in the sandbox image is dash, and sh <script> discards the script’s shebang — so the set -euo pipefail the fixing skill has the agent open the gate with was rejected on line 2 and the command exited 2 in milliseconds.

The gate was therefore a constant red. Every loop burned all its iterations, no iteration could satisfy the condition, and the harness’s half of “no green gate ⇒ no push” never fired — leaving the agent’s own self-reported gate=green (it runs the script directly, so its shebang is honoured) as the only thing gating a push. The loop still iterated, so it looked alive.

Now invoked as bash <script>. Applies to every backend running this image, kubernetes included.

The harvest never found it

prNotesRepoDir resolved the checkout from context.repo — a field no real run row carries. dispatchWorkflow reads a qualified owner/repo off the context, splits it into the row’s owner + bare repo columns, and the persisted context keeps only owner. It resolved to "" → null, and both readers were skipped: scratch.fixMarkers.verifyScript was null and notes was [] on every fix run ever recorded, while the files sat in the workspace.

That is the admin panel’s “No .lastlight-verify.sh was recorded” message, and the reason the PR journal has never carried a note. Now read off the run row’s column. Still inert on kubernetes, where the harness has no filesystem access to the PVC.

Operator note

The gate now actually runs, so each loop iteration pays a real install + test + lint + typecheck in a fresh sandbox — bounded by fix.gateTimeoutSeconds, with package downloads served by the shared /cache volume. Fix runs will be materially slower than they have been.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.3...v0.23.4

v0.23.3 dependency PRs merge again after we fix them

View on GitHub

Two fixes to the dependency-PR pipeline, both found on drizby.

The merge handoff was blocked by our own fix commit

resolveMergeDisposition inherited the fix route’s headSha === escalatedAtSha || headIsOurs test for “we escalated and nothing has changed since”. On the fix route the headIsOurs disjunct is load-bearing — our own retry push must not re-arm the attempt counter. On the merge route it is inverted: our commit at the head is the resolution, and the whole dependabot-ci-fix → pr.checks_passed → dependabot-pr-merge handoff ends with it.

So every PR we had ever escalated became structurally unmergeable — ci-fix repaired the branch, CI went green, and the merge route skipped because we were the one who fixed it. On drizby that was 27 of 37 green dependency PRs.

The route now compares headSha === escalatedAtSha only. It stays bounded: a merge run that declines re-stamps escalatedAtSha at the head it assessed, and one that succeeds populates assessedHeadShaByWorkflow — so the per-head already-assessed dedup allows one assessment per head SHA, the same bound every never-escalated PR already lives under.

dependabot-ci-fix now repairs merge conflicts instead of misdiagnosing them

The sweep dispatches this workflow for dirty / behind / blocked PRs as well as red ones, and the fix prompt already merged the base in and regenerated conflicted lockfiles. But diagnose ran first, and its taxonomy is CI-failure-shaped: shown green checks and no failing job, it honestly answered infra-dependent — one of the three stopping classes in the fix phase’s skip_if. The fix was skipped and the conflict left in place, on a succeeded run.

diagnose is now skipped outright for those three reasons; there is nothing to classify when no check is failing. Both prompts are also hardened for the paths where diagnose still runs on a merge-blocked PR.

Full changelog: https://github.com/nearform/lastlight/compare/v0.23.2...v0.23.3

v0.23.2

View on GitHub

Two production fixes found while debugging why a @bot comment on a dependency PR did the wrong thing.

The classifier was silently dead

models.classifier and models.screener were unpinned, and defaultFastModel deliberately ignores models.default and walks the provider table in order — so adding a KIMI_API_KEY silently repointed both at a reasoning model. On a reasoning model max_tokens bounds reasoning and visible output together, so the classifier’s 128-token budget (screener: 64) was consumed by hidden reasoning and the model returned "". Both parses then fell back to their safe defaults — intent: "chat" and flagged: false — which are indistinguishable from real answers.

Live effect: PR comments fell through to pr-comment, issue mentions routed to ignore, and the prompt-injection screener flagged nothing. No error, no metric, no log line.

  • One shared HELPER_MAX_TOKENS = 2048 for all three cheap helpers. The old budgets were sized for the visible reply, which is exactly the assumption a reasoning model breaks. The classifier drives the number: its system prompt is composed from every workflow’s classification: block (11 workflows, ~34 examples, ~3.7k tokens today) and grows with each workflow an operator adds — more intents to weigh means more reasoning before the first visible token. Raising it is close to free: a cap is not an allocation, billing is on tokens generated, and the only exposure is latency on a runaway, which the cap bounds.
  • The fallback now warns unconditionally. The diagnostic explaining all of this already existed, but was gated behind explain, which is never set in production.

Measured against the real composed prompt: gpt-5.4-nano 7/8 at 2048 and wrong on the reported comment at 128; gpt-5.4-mini 4/8 even at 2048; gpt-5-nano/gpt-5-mini empty at ≤512. The budget and the model pin are both load-bearing — neither alone fixes it. Operators should pin models.classifier / models.screener explicitly.

The Live tab could hide running work

The runs list paginated on ORDER BY started_at DESC alone. The Live filter asks for queued|running|paused; a cron fan-out enqueues a batch newer than anything executing; and queued rows are hidden client-side after pagination. So the batch filled the first page, running work sat on page 2, and the tab rendered “no workflow runs” mid-run. Ordering is now active-first (running < paused < queued < terminal) ahead of the date tiebreak. Terminal runs stay purely chronological, so the day/week ranges are unchanged.

Packages

  • lastlight-core 0.23.1 → 0.23.2
  • lastlight 0.23.1 → 0.23.2
  • lastlight-evals 0.9.1 → 0.9.2

agentic-pi (0.4.1), lastlight-shared and lastlight-workflow-engine (0.3.0) are unchanged.

v0.23.1

View on GitHub

Patch release: agentic-pi’s GitHub read tools now project and page their payloads instead of passing Octokit’s REST response through verbatim.

Why

Octokit’s payload is written for API clients, not for a context window — a dozen *_url fields per object, a ~1 kB user object per actor, a complete repository object under a PR’s head and base*, and then the genuinely large text: a Renovate changelog body, a lockfile patch`, a long review thread. An agent re-sends every one of those on each subsequent step of its loop, so the cost is the payload times the turns that follow it.

Measured raw against real repos, then through the new projections:

call raw projected
listPullRequests (5 open) 189,582 2,422 78×
listIssues (23) 267,880 9,666 28×
listPullRequestFiles (7) 26,679 987 27×
getPullRequest (one Renovate) 78,760 4,571 17×
listCommits (30) 160,905 12,334 13×

What changed

Three rules, applied uniformly so an agent learns one shape:

  1. Project to the fields prompts branch on — URLs, nested actor/repo objects and reaction counts dropped. List entries carry no body.
  2. Cap prose, and name the lift. Bodies/reviews/commit messages at 4000 chars, a file’s patch at 2000, each truncation notice naming the flag that returns it whole (full_body, full_bodies, full_messages, full_patch).
  3. Page, don’t cut. Every list returns { items, page, per_page, has_more, next_page }, so a long comment thread is paged rather than silently truncated. Search adds the API’s real total_count.

github_list_pull_request_files now omits each file’s patch unless include_patch is passed — it was 86% of a measured 7-file payload, and two prompts already forbade reading it in prose while the tool handed it over unbidden.

Prompt side: the agent was opening with an unprompted github_list_pull_requests to “find” a PR it had been handed. pr-review’s skill already guarded against that; the guard now also sits in dependabot-pr-merge, pr-comment and fixing (shared by pr-fix + dependabot-ci-fix).

Packages

  • agentic-pi 0.4.0 → 0.4.1
  • lastlight-core 0.23.0 → 0.23.1
  • lastlight 0.23.0 → 0.23.1
  • lastlight-evals 0.9.0 → 0.9.1

lastlight-shared and lastlight-workflow-engine are unchanged at 0.3.0.

v0.23.0 dependency-PR resilience

View on GitHub

Dependency-PR resilience — bounded fix retries, impact-classified merges, and one PR state machine. Closes #251 and #252 (PR #257).

What changes

The fix loop diagnoses before it retries, and is bounded. pr-fix and dependabot-ci-fix run a cheap diagnose phase that classifies the failure as reproducible / env-mismatch / flaky / infra-dependent / upstream-broken; only the classes another attempt can help with reach the fix phase. Attempts are counted across runs per PR, capped by fix.maxAttempts and a cumulative dollar ceiling, and a PR the bot gives up on gets a label and an explanatory comment instead of silence.

Major bumps are classified by impact, not semver magnitude. A major is judged against a new dependency-impact rubric and auto-merges at or below dependencies.autoMergeMaxImpact.

pr-review gets a trigger policy. Once per settled head SHA rather than per push, and it can cite what CI actually said.

PR state is resolved once, at the dispatchWorkflow choke point, into one PrState snapshot every policy question is a pure function over. That restructure closes three latent defects and replaces a per-PR concurrency guard that had never matched a row — so nothing had ever stopped two agents cloning and pushing the same branch at once.

Action required

  1. Re-consent the GitHub App for Actions: read (new optional permission — every existing installation must accept). Nothing hard-fails without it; diagnosis quality is capped at check-run annotations until granted, and the degradation is now stated in the prompt rather than silently substituting worse evidence.
  2. Audit your overlay’s forks. One fails every run (a forked fix prompt missing the CI_FIX_COMPLETE: marker instruction); the rest fail silently. Full list in docs/plans/dependency-pr-resilience/RELEASE-NOTES.md → “Audit your overlay’s forks”.
  3. Otherwise nothing: no migration, no schema change, every new key ships with a default.

Recommended rollout is fix.maxAttempts: 2 in the overlay, measure cost per attempt from the executions.cost_usd rollups, then raise.

Packages

Package Version
lastlight / lastlight-core 0.23.0
agentic-pi 0.4.0 (published separately from agentic-pi-v0.4.0)
lastlight-workflow-engine / lastlight-shared 0.3.0
lastlight-evals 0.9.0

Full changelog: https://github.com/nearform/lastlight/compare/v0.22.0...v0.23.0

v0.22.0 per-repository configuration

View on GitHub

Per-repository configuration

A managed repo can now commit a .lastlight/ directory and get its own configuration — models, prompts, skills, agent context, cron participation and approval gates — instead of everything being global to the instance. See the docs.

.lastlight/
├── lastlight.yml            # models, variants, disabled.workflows, approval (add-only), crons
├── workflows/prompts/*.md   # prompts only — a repo may never contribute workflow YAML
├── skills/<name>/SKILL.md
└── agent-context/*.md       # additive only

.lastlight/ is always read from the repo’s default branch, never a PR head — a PR cannot reconfigure the agent reviewing it. Agent context is additive only (a repo cannot shadow the operator’s security.md/rules.md), and approval is add-only (a repo may add a gate, never remove one). Operators bound the whole surface with a new repoConfig: block (enabled, allowKeys, allowedModels, allowAssets); precedence is default → overlay → env → repo.

A malformed or out-of-bounds lastlight.yml warns and drops the offending keys — it never fails a run.

Crons gain enable / disable at every layer

crons:
  enable:  [security-scan]
  disable: [repo-health]

Valid in config/default.yaml, an instance overlay, and a repo’s lastlight.yml. Legacy disabled.crons still works and is unioned in, so existing deployments are unaffected.

Note the semantic change at the operator layer: crons.disable now means off by default rather than structurally removed, so a repo can opt back in. A globally-disabled cron keeps its scheduler tick and resolves participation at tick time. The operator’s un-overridable kill switch is dropping crons from repoConfig.allowKeys.

New surfaces

  • CLI — lastlight repo fork, lastlight repo config validate (offline), lastlight repo config show <owner/repo>.
  • Dashboard — a config sub-tab on the repo page showing effective config with per-leaf provenance; repo-sourced values are visually distinct.
  • Evals — a repo-config tier plus a fixture seam on the fake GitHub, so a repo’s committed layer can be exercised end-to-end.

Fixes

  • Dependabot cron discovery bypassed per-repo participation entirely.
  • The dashboard cron toggle unregistered the tick, so a repo opt-in silently did nothing until a restart.
  • Three admin cron routes rebuilt contexts without _cronName, ignoring opt-outs.
  • Cron control keys (_cronName, _cronGloballyEnabled) were overridable from a cron YAML’s own context: block.
  • The repo-config byte cap admitted blobs whose tree entry reported no size.
  • redactPublic/SENSITIVE_KEY_RE and the config layer merge were duplicated.
  • Substantial spec/02-configuration.md drift, a stale k8s agent-context description in spec/09-sandbox.md, and an incorrect public FAQ answer about managing multiple repos.

Upgrade notes

No action required — the feature is inert until a repo commits a .lastlight/. Roll out by bumping deploy.version: v0.22.0 in your overlay repo.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.14...v0.22.0

v0.21.14 per-run GitHub credentials

View on GitHub

Patch release. Bumps agentic-pi to 0.3.1, lastlight-core/lastlight to 0.21.14, lastlight-evals to 0.7.30.

Fixes

  • Per-run GitHub credentials are threaded explicitly, no longer spliced through process.env (#215). Concurrent runs could race on the shared process environment; the executor now passes each run’s credentials down the call chain, and in-process sandbox runs no longer mutate process.env at all. agentic-pi gained the matching explicit auth-env plumbing.
  • post-review skips the local git diff when there is no checkout (#249, thanks @yo61) — fixes the k8s sandbox path where the workspace isn’t on the host.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.13...v0.21.14

v0.21.13 major dependency bumps + dependabot workflow fixes

View on GitHub

Dependency majors

Took the five open major dependabot bumps in one pass (each had been opened twice, once for the workspace root and once for the package dir):

  • better-sqlite3 12 → 13 — N-API rewrite (node-addon-api), drops the deprecated prebuild-install. No API surface change.
  • @slack/web-api 7 → 8 — axios → fetch, errors redesigned around Error subclasses. We only use new WebClient(token), chat.postMessage/chat.update, reactions.add, assistant.threads.setStatus, users.info and the KnownBlock type, and every catch site goes through err instanceof Error, so none of the removed options reach us. @slack/bolt 5 already resolves web-api 8.
  • chalk 5 → 6 (cli, shared, evals) — only raises the Node floor to 22, already every package’s declared minimum.
  • astro 6 → 7 (lastlight.dev) — Astro 7 flips the compressHTML default from true to 'jsx', which strips whitespace between inline elements and silently ate the spaces around inline links and <code> spans across the docs site. astro.config.mjs pins compressHTML: true; with that, all 59 generated .md files are byte-identical to the Astro 6 build.
  • recharts 2 → 3 (dashboard) — Tooltip formatter now receives ValueType | undefined, so the two HomePage formatters coerce with Number(v ?? 0).

Workflow fixes

  • Request a rebase on a conflicted dependency PR regardless of verdict (#245).
  • Retract the enqueue ack when a queued run leaves the queue (#244).

Versions

lastlight / lastlight-core 0.21.13, lastlight-shared 0.1.7, lastlight-evals 0.7.29. lastlight-workflow-engine and agentic-pi are unchanged.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.12...v0.21.13

v0.21.12 PR comments post again

View on GitHub

Fixed

pr-comment (and any comment workflow on a PR) can post again — #243, fixes #239.

GitHub resolves POST /repos/:owner/:repo/issues/:n/comments against the target’s type: an issue is checked against the issues permission, a pull request against pull_requests. The issues-write profile minted pull_requests: read, so it could comment on issues and 403’d (Resource not accessible by integration) on every PR — while the run still reported succeeded, leaving the answer only in the transcript.

issues-write now carries pull_requests: write, unblocking pr-comment plus verify / qa-test / demo / issue-comment whenever the router points them at a PR. Tool sets are unchanged: github_create_pull_request / github_create_pull_request_review remain review-write+, and that registration gate — not the token scope — is what keeps a comment workflow from submitting a formal review.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.11...v0.21.12

v0.21.11 OpenInference span tree + metrics toggle

View on GitHub

OpenTelemetry: OpenInference span tree (#224)

Traces now render as a proper agent tree in OpenInference-aware backends (e.g. Arize Phoenix) instead of a flat two-span shape:

lastlight.workflow.run     (CHAIN)
└─ lastlight.workflow.phase (CHAIN)
   └─ lastlight.agent.execute (AGENT)  — model, total tokens, total cost
      ├─ turn N               (LLM)    — per-turn tokens + cost
      │  └─ <tool>            (TOOL)   — tool.name, is_error, args/result (gated)
  • Spans carry OpenInference attributes (openinference.span.kind, llm.model_name/llm.system, llm.token_count.*, llm.cost.total, tool.*). Content (input.value/output.value/tool args+results) stays gated behind LASTLIGHT_OTEL_INCLUDE_CONTENT.
  • Built from the same pi event stream; the flat pi.* span events are kept as a fallback.

Optional: disable OTLP metrics

New LASTLIGHT_OTEL_METRICS_ENABLED=false (overlay otel.metrics: false) exports traces only — for a backend that rejects the metrics signal (Phoenix 404s the metrics endpoint).

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.10...v0.21.11

v0.21.10 OTLP protobuf export (Phoenix fix)

View on GitHub

Fixes

  • OpenTelemetry: default OTLP export to http/protobuf. The harness exported traces/metrics as OTLP/JSON (@opentelemetry/exporter-*-otlp-http), which protobuf-only backends — Arize Phoenix among them — reject with HTTP 415, so traces silently never landed. Export now defaults to http/protobuf (the OTLP spec default) and honors the standard OTEL_EXPORTER_OTLP_PROTOCOL (+ per-signal OTEL_EXPORTER_OTLP_{TRACES,METRICS}_PROTOCOL) to opt back into http/json.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.9...v0.21.10

v0.21.9 k8s/QEMU deploy example, live-pipeline fix, PR-fix build gate

View on GitHub

Highlights

  • feat(images): publish lastlight-agent-qemu variant for gondolin/k8s, plus a usable-as-is k8s deploy example (apps/server/deploy/k8s/) and docs.
  • fix(dashboard): stop the live workflow pipeline blanking mid-run.
  • feat(#218) / fix(pr-fix): gate the pr-fix flow so it no longer omits the build step; harden failed-checks handling in the GitHub engine (+ tests).
  • Docs + www run-it page updates.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.8...v0.21.9

v0.21.8 event router playground, PR-review base-branch diff, workspace reaping

View on GitHub

Highlights

  • feat(dashboard): Event Router Playground — visualize + dry-run event routing (e150a16)
  • feat(dependabot-ci-fix): teach the agent to run tests efficiently under a time budget (108e38f)
  • fix(pr-review): fetch base + deepen to merge-base so reviews anchor inline — git diff origin/<base>...HEAD now works in the workspace (c8cd1d8)
  • feat(sandbox): reap workspaces in-harness (#106) (a1ac7c2)

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.7...v0.21.8

v0.21.7 deploy downtime fix (redundant recursive chown)

View on GitHub

Deploy performance

Root-caused from a nearform deploy that was offline ~3 minutes:

  • entrypoint.sh no longer runs a full chown -R over the state volume every boot. On a busy instance that walked >1.5M files (23 GB, mostly cloned repos under sandboxes/) on an IOPS-limited disk — ~3 min of blocking before the harness starts, despite every file already being uid 10001. It’s now gated on the volume not already being lastlight-owned, so steady-state boots skip it. Boot drops from ~3 min to ~10 s.
  • server update pulls the agent image and recreates first, then pulls the sandbox images with the agent already back online (sandbox images are build-only profiles used at workflow time, never started by up -d).

Net: the nearform deploy goes from ~5 min (3 min offline) to well under a minute.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.6...v0.21.7

v0.21.6 dashboard: CLI-login cleanup, actor chip, pipeline hardening

View on GitHub

Dashboard fixes

  • CLI-login handoff no longer traps the tab. The sessionStorage cli_login stash is now cleared on the Authorize (success) path, not just Cancel — returning to /admin in the same tab after a lastlight login handoff no longer re-shows the “Authorize CLI login” screen.
  • Actor chip consistency. When an avatar resolves, the per-type channel icon (CLI / cron / Slack / …) now renders ahead of the avatar in the run detail panel, matching the list. GitHub actors are exempt.
  • Run list. The current phase shows as an italic sub-state right after the workflow name while running; the repo/issue/actor line flex-wraps so the actor chip drops cleanly to its own row on long repo/user names.
  • Pipeline hardening. A malformed/legacy phase_history row (missing phase) no longer crashes the pipeline view.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.5...v0.21.6

v0.21.5 gate dependabot-ci-fix on real dependency PRs

View on GitHub

Fixes

  • dependabot-ci-fix no longer misfires on human PRs. The red-CI webhook path (pr.checks_failed) now applies the same deterministic dependency-PR gate the green pr.checks_passed path already uses — commit author dependabot[bot]/renovate[bot] or branch prefix dependabot//renovate/. Previously any PR whose checks settled red was handed to the LLM classifier alone, which could route a human’s failing PR onto the dependabot-ci-fix repo-write workflow. A human’s red PR now fires nothing and never even makes the getChecksConclusion call. (95b90e1)

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.4...v0.21.5

v0.21.4 fix dashboard GitHub/Slack OAuth login

View on GitHub

Fix

Dashboard GitHub & Slack OAuth login could fail with ArcticFetchError / invalid content-length header. Once any in-process agent run replaced Node’s global undici dispatcher (a non-default build that rejects a manually-set Content-Length), arctic’s token exchange — which pre-sets that header — threw, and users were locked out of GitHub/Slack login (password login still worked). Replaced the two arctic token exchanges with a small shared exchangeOAuth2Code helper that performs the same confidential-client exchange without setting Content-Length (undici computes it). Adds a regression test.

Not caused by v0.21.3 — the OAuth code and lockfile were unchanged; the v0.21.3 deploy merely forced a re-login that exposed the pre-existing break.

Packages

lastlight 0.21.4 · lastlight-core 0.21.4 · lastlight-evals 0.7.20.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.3...v0.21.4

v0.21.3 red sweep picks up behind/dirty/blocked dependency PRs

View on GitHub

Highlights

  • fix(cron): red dependency-PR sweep now covers green-CI-but-unmergeable PRs. The daily red sweep (dependabot-ci-fix) previously only picked up settled-red CI; a PR that was green but behind / dirty / blocked matched neither sweep and sat forever (the #270-class blind spot). It now also enqueues those states (threading a reason), and the ci-fix prompt flags requires-human on a blocked PR it can’t unblock so it stops looping. unknown is left for a later tick by design.
  • fix(dashboard): Recent Workflows shows only finished runs.
  • fix(agentic-pi): surface swallowed provider errors on the synthesized terminal (shipped as agentic-pi@0.3.0, incl. @octokit/rest v22).

Packages

lastlight 0.21.3 · lastlight-core 0.21.3 · lastlight-evals 0.7.19 · agentic-pi 0.3.0 (own tag agentic-pi-v0.3.0).

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.2...v0.21.3

v0.21.2 cron fan-out unqueued, quieter dependabot + OAuth logs, queued-workflow dashboard

View on GitHub

Fixes

  • Cron fan-out no longer throttles. The dependabot/cron fan-out fired runs in slices of 3, serializing dispatch. It now fires all discovered runs at once and lets the global concurrency cap (maxWorkflows) + admission queue be the sole limiter — over-cap runs park as cheap queued rows and drain as slots free.
  • No more spurious OAuth warnings. The executor stopped logging Model '…' needs an OAuth login for 'anthropic' on every run when the provider authenticates fine via an API key (ANTHROPIC_API_KEY). It warns only for oauth-only providers or when no credential is usable.
  • Dependabot PRs stop getting re-commented. On a trivial PR where auto-merge is disabled on the repo, the workflow now flags requires-human and comments once, instead of re-posting “a maintainer should merge” on every check-pass and daily cron.

Dashboard

  • Queued workflow runs are hidden by default in the workflow list, with an “N queued workflows” show/hide toggle.
  • Live Activity gains a Queued Workflows count; “Active Workflows” is now running+paused only (no longer double-counting queued).

CLI

  • lastlight skills install defaults to the remote nearform/lastlight marketplace so installed skills auto-update (claude plugin update) instead of being pinned to the CLI. Adds --local to force the bundled offline copy; migrates a mismatched existing registration.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.1...v0.21.2

v0.21.1 streaming server-log viewer

View on GitHub

Highlights

  • Admin console Logs tab — a streaming docker logs viewer for any lastlight-* container, riding the existing /server/logs[/stream] endpoints (previously CLI-only). Live SSE follow with reconnect + ring-buffer cap, time-windowed snapshot mode (--since), container picker, max-rows, client-side text filter, and dependency-free syntax highlighting in a small monospace body. Closes the gap where the dependency-merge cron’s code-based discovery/skip decisions log only to stdout, not the DB.

Full changelog: https://github.com/nearform/lastlight/compare/v0.21.0...v0.21.1

v0.21.0 admin resource stats + dependabot dedup/recovery

View on GitHub

Highlights

Admin dashboard (home page):

  • Host-level CPU + memory added to the Resource Usage panel (true host figures via the agent container’s non-namespaced /proc; memory uses MemAvailable so reclaimable cache isn’t counted as used).
  • Containers are classified agent / sandbox / infra — infra sidecars (caddy, coredns, nginx-egress, otel) are hidden by default behind a toggle, with per-kind icons.
  • Recent Workflows moved under Stats and now shows per-run token and cost totals (aggregated over executions).
  • Resource Usage panel gained a proper loading state.

Engine:

  • Dependabot PR event de-duplication + orphaned-workflow recovery after a harness restart.

Dependencies: major bumps merged this cycle — @slack/bolt 5, better-sqlite3 12, @hono/node-server 2, @octokit/rest 22, plus grouped production/dev updates.

Full changelog: https://github.com/nearform/lastlight/compare/v0.20.1...v0.21.0

v0.20.1 bot git-identity fix

View on GitHub

Fixes

Agent git commits are now authored as the configured bot login (e.g. nearform-lastlight[bot]) instead of the default last-light[bot], on the docker and smol sandbox backends (#167).

  • sandbox: forward git-identity env to the agent via docker exec -e / smolvm exec -e instead of agentic-pi --sandbox-env, which is a no-op under --sandbox none — so GIT_AUTHOR_*/GIT_COMMITTER_* (and the github.com auth header) actually reach the agent’s git.
  • agentic-pi: github_clone_repo no longer hardcodes a repo-local last-light[bot] identity; the clone inherits whatever the environment configures.
  • sandbox: thread git identity into command/script phases too, so their commits are attributed (closes the command-phase gap in #167).

Full changelog: https://github.com/nearform/lastlight/compare/v0.20.0...v0.20.1

v0.20.0 workflow-run owner column + dashboard fixes

View on GitHub

Highlights

  • fix(state): store workflow-run owner as its own column
  • fix(dashboard): render skipped phases as skipped, not failed
  • fix(dashboard): link workflow-run repos/issues via qualified owner/repo
  • fix(ci): drop agentic-pi from publish.yml npm loop; publish it via agentic-pi-npm.yml

Bumped: lastlight-core & lastlight → 0.20.0, lastlight-evals → 0.7.14. agentic-pi/shared/workflow-engine unchanged.

Full changelog: https://github.com/nearform/lastlight/compare/v0.19.0...v0.20.0

v0.19.0 git-auth fix, dashboard GitHub links, safer dependabot/pr-merge

View on GitHub

Highlights

  • sandbox/agentic-pi (0.2.19): git auth via GIT_CONFIG_* http.extraheader; finish failed ledger rows.
  • server: dependabot-ci-fix is now fix-only; single-owner merge in pr-merge.
  • dashboard: link repos + issue/PR numbers to GitHub; fix stale repo selection.
  • workflow-engine (0.1.2): phase-executor / scheduler fixes.

Packages: lastlight/lastlight-core 0.19.0 · agentic-pi 0.2.19 · lastlight-shared 0.1.2 · lastlight-workflow-engine 0.1.2 · lastlight-evals 0.7.13.

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.6...v0.19.0

v0.18.6 cron CLI/dashboard controls + dependabot backstop + sandbox fixes

View on GitHub

Patch release.

Highlights

  • feat(cli): lastlight cron list/trigger/enable/disable + dashboard “Run now”
  • feat(server): dependabot red-PR cron backstop, label state machine + mention routing
  • feat(server): dependabot-pr-merge rebases behind/dirty PRs; repo-write gains workflows scope
  • fix(sandbox): bake pnpm via corepack + let the non-root agent fnm-install Node

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.5...v0.18.6

v0.18.5 dependabot-pr-merge no-op guards + per-PR cron fan-out

View on GitHub

Supersedes v0.18.4, which failed npm provenance validation the first time lastlight-workflow-engine / lastlight-shared were auto-published (their package.json was missing repository). This release adds that metadata; no code change vs v0.18.4.

Highlights

The daily dependency-merge sweep could report succeeded while merging nothing — a repo’s scan agent buried itself in every open PR’s lockfile churn until its context overflowed (or returned an empty completion), and the harness recorded the empty run green. It also occasionally direct-merged a RED PR.

  • on_output.requires_marker (workflow-engine) — a phase fails if its final output lacks a required completion marker. dependabot-pr-merge must now emit ASSESSMENT_COMPLETE, so a silent no-op is recorded RED, not green.
  • reclassifySuccess (executor) — a terminal agent_end carrying no final answer (an empty completion, including agentic-pi’s synthesized backstop) is demoted from success to the soft unknown outcome. Generic across all workflows.
  • RED-PR merge gate — the prompt confirms mergeable_state === "clean" before a direct merge (a repo with no required checks reports a failing PR mergeable).
  • Per-PR cron fan-out — the daily backstop now discovers green dependency PRs in code and fans out one bounded single-PR run per PR, retiring the mode: scan whole-repo agent sweep. One PR per run makes context overflow structurally impossible.

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.3...v0.18.5

v0.18.3 dependabot-pr-merge direct-merge fallback

View on GitHub

Fixes

  • dependabot-pr-merge now merges PRs GitHub refuses auto-merge on. github_enable_auto_merge returns { ok: false } for several distinct reasons; the workflow treated them all as “auto-merge disabled” and left a maintainer comment. GitHub actually refuses with reason "Pull request is in clean status" when the PR is already mergeable with nothing to wait for — typically a repo with no required checks — where the correct action is a direct squash merge. The workflow now branches on reason: clean/mergeable → direct squash merge; “not allowed for this repository” → maintainer comment; conflict/other → comment.

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.2...v0.18.3

v0.18.2 dependabot-pr-merge context fix

View on GitHub

Fixes

  • dependabot-pr-merge no longer overflows the model context. Scan-mode assessment read every open dependency PR’s full diff via github_get_pull_request_diff; lockfile churn (package-lock.json, pnpm-lock.yaml, …) is tens of thousands of lines, so repos with several open bumps blew past the context window and the run died mid-assessment. It now inspects via github_list_pull_request_files (file list + line counts), skips lockfile diffs, only pulls the full diff for a small non-lockfile source change, and caps the scan at 10 PRs/run.
  • A run that only reads files with no verdict/action is now treated as a failure, not a false success.

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.1...v0.18.2

v0.18.1 dependabot cron hardening

View on GitHub

Highlights

  • fix(server): harden the dependabot cron against stale managedRepos. A managedRepos entry whose repo was deleted, transferred to another org, or had App access revoked made every repo-scoped token mint 422 — which silently disabled agentic-pi’s github_* tools, so dependabot-pr-merge (which, unlike pr-fix, has no checkout to fall back on) ran with no tools at all. Now:
    • getAccessibleManagedRepos() intersects configured managedRepos with the discovered installation set; the cron fan-out uses it so a stale entry never spawns a doomed scan (falls back to the raw list before boot discovery populates).
    • The executor fails fast with an actionable error_fatal result when an expected token mint fails, instead of running a toolless agent.
    • cron-dependabot-merge schedule */30 → 0 14 * * * (daily backstop; the pr.checks_passed webhook handles real-time merges).
  • evals: fallback per-token cost for flat-rate models.

Full changelog: https://github.com/nearform/lastlight/compare/v0.18.0...v0.18.1

v0.18.0 sandbox drift fix + 40% slimmer images

View on GitHub

Highlights

  • Fix — production sandbox outage. Every sandbox agent crashed at boot after a routine rebuild (does not provide an export named 'AuthStorage'): upstream pi-coding-agent shipped a breaking change within 0.80.x, and the sandbox’s npm install -g resolved past the lockfile-tested version. Ported agentic-pi to the 0.80.10 ModelRuntime API.
  • Build — the sandbox now vendors agentic-pi from the workspace. Instead of installing the published npm tarball (which ignored the lockfile), sandbox*.Dockerfile builds a pnpm deploy bundle from source, pinning the whole dependency tree to exactly what CI tested. Removes the agentic-pi.pin file, its regen script, and the drift-guard — the npm publish is now for external consumers only. The transitive-drift class that caused the outage is gone.
  • Perf — sandbox images ~40% smaller. Base moved to node:24-slim, dropped the pre-baked fnm Node 22+24 (fnm stays for on-demand via nodejs.org), and purged semgrep caches. Lean sandbox 2.92 GB → 1.75 GB; shared base 2.22 → 1.49 GB (sandbox-qa drops ~730 MB too).
  • Dashboard — Repos tab (Artifacts folded under it).

Ships agentic-pi@0.2.18 (published separately) as the vendored coding-agent harness.

Full changelog: https://github.com/nearform/lastlight/compare/v0.17.1...v0.18.0

v0.17.1 Artifacts tab at scale

View on GitHub

Artifacts tab: repo discovery, search & time window at scale

The admin dashboard’s Artifacts tab now scales to large stores and stops showing an empty repo picker.

  • Repos actually list. New `GET /artifact-repos` scans the store for repos that have artifacts (server-side search + pagination, newest-first), replacing the empty-by-default `managedRepos` picker.
  • Time window + search on the issues list. `GET /artifacts` gains `q`/`since`/`limit`/`offset` and returns `{ keys, total }`; the per-repo run list respects the shared header time range (issues list only, so the repo list never goes empty under the 24h default).
  • Built for scale. Async `fs/promises` listing with a 5s TTL cache (invalidated on write/harvest). Repo index is O(repos), not O(issues): 1000 repos scan in ~74ms (0ms cached); a 2000-issue repo in ~146ms (0ms cached). Both lists page in with load-more, each row showing its age.

Full changelog: https://github.com/nearform/lastlight/compare/v0.17.0...v0.17.1

v0.17.0 react to failing PR checks (Dependabot auto-fix)

View on GitHub

Highlights

  • New dependabot-ci-fix workflow — Last Light now reacts to a PR whose CI goes red (check_suite.completed → new pr.checks_failed event), routed through the intent classifier so workflows self-register. For a red dependency-update PR it diagnoses + pushes a minimal fix, then assesses the diff and, only when trivial, enables GitHub auto-merge (merge-once-green).
  • agentic-pi github_enable_auto_merge tool (repo-write) — GraphQL enablePullRequestAutoMerge; published independently as agentic-pi@0.2.17 and baked into the sandbox image via the pin.
  • Docs-sync skill + hook repointed at the monorepo layout.

Operator note: subscribe the GitHub App to Check suite events (Checks: read) or failing-CI PRs won’t reach the connector.

Full diff: https://github.com/nearform/lastlight/compare/v0.16.2...v0.17.0

v0.16.2 server-mode handoff docs staged outside the repo tree

View on GitHub

Fix

Server-mode build handoff docs (.lastlight/<issueKey>/) no longer leak into the target repo’s PRs. For pre-cloned workflows (build, pr-*) on whole-workspace backends (docker/none/smol), the docs are staged at the sandbox workspace root — a sibling of the checkout, reached via {{issueDir}} = ../.lastlight/<key> — so the executor’s git add -A can’t sweep them into the feature commit. gondolin (mounts only cwd) keeps the in-repo path + commit gate. The decision is computed once (artifactIssueDir) and carried on ExecutorConfig.buildAssetsRelocated.

Also

  • Removed the legacy tracked .lastlight/issue-* docs and gitignored .lastlight/.
  • CI/Docker: build the vendored agentic-pi in the agent image, and fix the automated npm publish loop (pnpm pack invocation + skip already-published versions) — both latent since the monorepo vendoring, first exercised by this release.
  • Docs synced (spec 02/07, CLAUDE.md); new artifactIssueDir unit tests.

Packages: lastlight / lastlight-core 0.16.2, lastlight-evals 0.7.3.

Full diff: https://github.com/nearform/lastlight/compare/v0.16.1...v0.16.2

v0.16.1 first post-migration release

View on GitHub

First release cut from the consolidated monorepo.

Packages

  • @lastlight/core → 0.16.1, lastlight (CLI) → 0.16.1, lastlight-evals → 0.7.2
  • @lastlight/workflow-engine and @lastlight/shared publish for the first time at 0.1.0
  • Claude plugin manifest → 0.16.1

Changes

  • feat(evals): vendor a sample run per tier so the evals dashboard deploys populated instead of an empty shell (no eval runs in the build)
  • ci(deploy): release-gate the www + evals Cloudflare deploys (fire on Release + workflow_dispatch, not every push to main)
  • ci(publish): fix the monorepo buildx bake — FS entitlement + invoke from apps/server so the parent-dir context resolves
  • docs(readme): lead CLI examples with the global lastlight command

This Release fires the GHCR image build (publish.yml, moving :latest) and the release-gated www + evals deploys. npm publishing is manual (see docs/RELEASING.md).

Full diff: https://github.com/nearform/lastlight/compare/v0.16.0...v0.16.1

v0.16.0 Slack rich formatting

View on GitHub

Slack rich formatting

Upgrades Slack output from a single text-only mrkdwn path to richer Block Kit rendering, plus interactive approvals.

  • Hardened tables — per-column width cap + total-width budget, over-wide cell truncation, long-table elision, and a *label*: value fallback for wide 2-column tables.
  • Block Kit progress — workflow progress renders as a header + context meta + divider + sectioned checklist (renderProgressBlocks). The notifier transport gains publish(markdown, model?) so Slack renders blocks while GitHub keeps markdown — one content source, no drift.
  • Interactive approvals — approval gates post Approve/Reject buttons; new POST /webhooks/slack/interactions (signature-verified, trigger_id-deduped) routes clicks into the same approval-response path as /approve, then rewrites the prompt to a resolved state. Socket mode uses Bolt action listeners.
  • Inline images — ![alt](url) is auto-promoted to Block Kit image blocks, with a plain-text fallback if Slack rejects them.

Also adds scripts/slack-format-demo.ts and updates spec/03-integrations.md.

Deploy note: the interactions endpoint requires Interactivity enabled in the Slack app (Request URL → <public-url>/webhooks/slack/interactions).

Full changelog: https://github.com/nearform/lastlight/compare/v0.15.0...v0.16.0

v0.15.0 forkable, workflow-driven intent classifier

View on GitHub

Highlights

Fork the classifier prompt + workflow-driven intents (#164, #165). The intent classifier — which routes free-text GitHub comments and Slack messages to a workflow — is no longer a single hardcoded prompt string. It’s now composed at runtime and forkable per-deployment:

  • Forkable base template — workflows/prompts/classifier.md, overridable via the overlay and lastlight fork classifier.
  • Per-workflow categories — each workflow YAML carries its own classification: block (intent + description + examples).
  • Add a routable intent by adding a workflow — an overlay incident.yaml with classification.intent: incident teaches the classifier the category, the parser the token, and the router the route via getWorkflowByIntent, with no core change. Well-known intents keep their bespoke, context-dependent routing.
  • Retired the separate new-issue question classifier (folded into the main QUESTION intent); the re-triage gate is now the forkable classify-adds-info.md.

Classifier/screener model is now config-driven. defaultFastModel(taskType) reads the config.yaml models: map (models.classifier / models.screener) before the env OPENCODE_MODELS map and the provider fast-model fallback — no more env-only pinning. Only an explicit per-task entry counts (never models.default), so the cheap helpers stay cheap.

New CLI: lastlight fork classifier forks the base classifier prompts.

Docs & skills: updated spec/, CLAUDE.md, and the lastlight-overlay Claude Code plugin skill (forkable classifier + an explicit model-setup section).

Full compare: https://github.com/nearform/lastlight/compare/v0.14.0...v0.15.0

v0.14.0 managed repos from the App installation

View on GitHub

Highlights

  • Managed-repo list sourced from the GitHub App installation. When the overlay’s managedRepos is left empty, the effective list now falls back to the repos the App installation can access — discovered at boot (apps.listReposAccessibleToInstallation) and kept live by installation / installation_repositories webhooks — instead of managing nothing. A non-empty configured managedRepos still wins (back-compat / further restriction). An org install that already limits the App to a subset of repos need not maintain a second copy in config.
  • New admin /managed-repos endpoint and a Config → Managed repos dashboard pane showing the configured / installation / effective lists and their source.

Notes

  • Cron repo-wide scans enumerate the effective list, so with managedRepos empty they now cover every discovered repo — set an explicit list to scope tighter.
  • For a repository_selection: "all" install, a newly-created org repo is picked up on the next boot fetch (GitHub fires no webhook); the selected case is fully covered by webhooks.

Full changelog: https://github.com/nearform/lastlight/compare/v0.13.0...v0.14.0

v0.13.0 Retry a failed workflow run

View on GitHub

Highlights

Retry a failed workflow run from the phase that failed — a new Retry action resumes a failed run from the exact phase that failed, keeping the same context, taskId and workspace. It reuses the ledger-driven resume path that already powers approval-gate resumption (the failed phase’s executions row is success=0, so it re-runs while already-succeeded phases skip).

  • Dashboard: Retry button on failed runs (list + detail panel).
  • CLI: lastlight workflow retry <id>.
  • API: POST /admin/api/workflow-runs/:id/retry (authed; 404/400/503 guards).
  • Works for GitHub-issue- and Slack-thread-scoped runs (e.g. an explore started from Slack), by reconstructing context from the stored run row rather than a lossy owner/repo/issue rebuild.
  • restartRun is a compare-and-set (failed→running), so a double-click no-ops.

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.8...v0.13.0

v0.12.8 agentic-pi 0.2.16

View on GitHub

Changes

  • Bump agentic-pi to 0.2.16 — sandbox install pin (sandbox/agentic-pi.pin) and the harness dependency in lockstep. Build + full test suite green.

This is also the first release to build its Docker images against the registry-backed build cache added in v0.12.7 — the agentic-pi bump only invalidates the thin tail below the Chromium layers, so sandbox images should restore Chromium from cache and rebuild only the agentic-pi/agent-context tail.

Full Changelog: https://github.com/nearform/lastlight/compare/v0.12.7...v0.12.8

v0.12.7 registry-backed Docker build cache

View on GitHub

CI / infra

Docker image builds now use a registry-backed cache that survives across releases.

`docker-publish.yml` only runs on a release (a tag ref), and GitHub Actions cache is ref-scoped — a run can only restore caches written by its own ref or the default branch. So each release wrote a `type=gha` cache under its own tag scope that the next release could never read, and every release built cold (v0.12.6 restored 1 layer and rebuilt 23, including sandbox-qa’s ~300 MB Chromium download).

Each image now caches to/from a per-image GHCR manifest (`ghcr.io/nearform/lastlight-:buildcache`, `mode=max`), which any run can read — so unchanged layers survive across releases and a dashboard-only release re-tags near-identical sandbox images in seconds.

This release seeds the new registry cache (its own build is still cold); the speed-up shows on the next release that reads it.

No runtime/CLI behavior change.

Full Changelog: https://github.com/nearform/lastlight/compare/v0.12.6...v0.12.7

v0.12.6 explore survives an empty socratic iteration

View on GitHub

Fixes

Explore no longer dies with a bare `unknown` on a degenerate socratic turn.

A generic-loop iteration where the agent exits cleanly but produces no final text and no `agent_end` event (surfaced by `mapStopReason` as `unknown`) used to hard-fail the entire workflow, cascading to skip `synthesize`/`publish` — on `explore` this discarded every accumulated socratic Q&A round and never wrote the spec.

  • New declarative `generic_loop.on_soft_failure: { retries, then }` policy — a soft iteration (stop reason `unknown`/`error_truncated`, as opposed to a hard crash: terminated / `error_fatal` / `error_tool` / `error_exit_*`) retries up to N times under a distinct `_iter_n_retry` ledger label, then either advances the loop as complete or hard-fails. Absent ⇒ prior behavior.
  • Extracted a generic `isSoftOutcome()` classifier, shared with the reviewer loop’s fallback recovery (removes the duplicated inline check).
  • `explore.yaml`’s socratic phase opts in (`retries: 1`, `then: complete`) — a degenerate turn now retries once, then advances to synthesis with the Q&A gathered so far.
  • A bare `unknown` error now reads `agent produced no output (no final text, no agent_end)`.

Full Changelog: https://github.com/nearform/lastlight/compare/v0.12.5...v0.12.6

v0.12.5 reviewer verdict-file recovery

View on GitHub

Fixes

  • Reviewer marked failed on a clean review: some models (e.g. gpt-5-codex) end their final turn on the reviewer-verdict.md write with no trailing stdout, so the VERDICT: marker was absent and the run classified a clean APPROVED review as a failure — which also ran a needless fix cycle and suppressed the PR-link comment (posting is gated on run success). The reviewer loop now recovers the verdict from reviewer-verdict.md when the stdout marker is missing (gated to soft outcomes so a real sandbox failure can’t be masked), and reviewer.md now requires the reviewer to end with a final text VERDICT + summary rather than a tool call.

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.4...v0.12.5

v0.12.4 guardrails: absent lint/typecheck is fine

View on GitHub

Fixes

  • Guardrails: clarified the contract to present-must-pass, absent-is-fine. A configured lint/typecheck that fails still blocks, but a repo that simply has no lint (or no typecheck) command no longer blocks the build when tests pass. Previously a missing lint script could BLOCK a build whose tests + typecheck were green.

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.3...v0.12.4

v0.12.3 PR base branch fix

View on GitHub

Fixes

  • PR creation: the build’s pr phase hardcoded base: main, so on a master-default repo GitHub rejected the PR with 422 base invalid. It now opens against {{baseBranch}} (the repo’s real default branch, resolved at dispatch). Completes the default-branch fix started in v0.12.2 (reviewer scoping).

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.2...v0.12.3

v0.12.2 reviewer default-branch fix

View on GitHub

Fixes

  • Reviewer scoping: the reviewer phase diffed against a hardcoded main, so on a repo whose default branch isn’t main (e.g. master) every scope command threw fatal: ambiguous argument 'main..HEAD' and derailed the review. It now scopes against the repo’s real default branch ({{baseBranch}}), resolved at dispatch for build/issue runs.

CI

  • Docker image publish now builds with a per-image buildx GitHub Actions cache (this is the first release to exercise it) — unchanged layers (notably sandbox-qa’s Chromium) restore from cache instead of rebuilding cold.

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.1...v0.12.2

v0.12.1 header version label

View on GitHub

Highlights

  • Dashboard: the full-width “Pinned to vX.Y.Z” strip under the header is gone. The pinned core version now shows as a small muted label tucked left of the theme toggle, and UpdateBanner is warnings-only (Update available / Redeploy needed).

Full changelog: https://github.com/nearform/lastlight/compare/v0.12.0...v0.12.1

v0.12.0 prebuilt Docker images (pull instead of build)

View on GitHub

Highlights

  • lastlight server update now pulls prebuilt Docker images from GHCR by default instead of building them on the deploy host — a pull is seconds where a build was minutes.
    • A GitHub Release now also publishes ghcr.io/nearform/lastlight-{agent,sandbox-base,sandbox,sandbox-qa} (new .github/workflows/docker-publish.yml, amd64, public), tagged vX.Y.Z and :latest.
    • server update pulls the tag from the overlay’s deploy.version pin (else :latest) and re-tags each to its fixed local name, so docker-compose.yml and the harness find them unchanged — no compose or runtime change.
    • --local (on server update, also honoured by server setup) reverts to building from source. server build stays the always-local build.
  • Stock sidecars (coredns/nginx/otel-collector/caddy) were already remote images and egress-init reuses the agent image, so only four images are published.

Deploy note: to roll out to a pinned instance, cut the release (this builds+pushes the images), then bump the overlay’s deploy.version to the tag and run lastlight server update. Upgrade the host’s global CLI (npm i -g lastlight@0.12.0) first so it knows the pull path.

Full changelog: https://github.com/nearform/lastlight/compare/v0.11.1...v0.12.0

v0.11.1 Fix annotated pin-tag drift comparison

View on GitHub

Fix: pinned drift comparison for annotated tags

Follow-up to v0.11.0’s core-version pinning. The drift check resolved a pinned tag with git ls-remote origin refs/tags/<pin>, which for an annotated tag returns the tag-object SHA rather than the commit. Since git checkout <pin> lands HEAD on the commit, the two never matched — producing a permanent false “redeploy needed” in lastlight server status and the dashboard banner for every pinned instance (this repo tags releases annotated).

Fixed by resolving with the glob form refs/tags/<pin>* (the only pattern that surfaces the peeled …^{} commit row) and centralizing the parse in pickTagCommit() — shared by the host CLI and the in-container banner.

The pin/checkout behavior itself was always correct; this only affected drift reporting.

Full changelog: https://github.com/nearform/lastlight/compare/v0.11.0...v0.11.1

v0.11.0 Overlay-driven core version pinning

View on GitHub

Overlay-driven core version pinning

The deployment overlay can now declare which core version an instance runs, making the overlay repo the declarative source of truth for deploys.

What’s new

  • deploy.version in the overlay config.yaml — a git tag/ref (e.g. v0.10.6) that pins the core version. lastlight server update pulls the overlay first, then checks core out at the pinned tag (detached HEAD) instead of tracking main. lastlight server setup applies the same pin before its first build. Unset — or the sentinels main/latest — tracks main as before.
  • LASTLIGHT_CORE_VERSION env override — pin without editing config.yaml (handy for CI).
  • Pin-aware drift surfaces — when pinned, lastlight server status and the dashboard “update available” banner compare the running image against the pinned tag rather than main: “behind” now means “pin bumped, redeploy needed”, and an on-pin instance shows a quiet “pinned vX.Y.Z” label instead of a nudge.

Workflow

Bump deploy.version in the overlay repo, commit, and run lastlight server update on the host — a CI/CD job can now drive the deploy declaratively.

Full changelog: https://github.com/nearform/lastlight/compare/v0.10.6...v0.11.0

v0.10.6 configurable bot name / @mention handle

View on GitHub

Highlights

  • Configurable bot name / @mention handle (#148). A new botName config field (default last-light) is the single source of truth for the bot’s identity — set it via the overlay config.yaml (botName: nearform-lastlight) or the GITHUB_APP_BOT_NAME env var. It derives:
    • the incoming @mention handle the router triggers on (configured handle only — the default stays last-light, so existing deployments are unaffected),
    • botLogin (<botName>[bot], still overridable with BOT_LOGIN) for self-comment / self-review filtering,
    • the git commit author for agent commits (host + sandbox), keyed off the resolved botLogin so a BOT_LOGIN override is honoured.

Full changelog: https://github.com/nearform/lastlight/compare/v0.10.5...v0.10.6

v0.10.5 OAuth subscription logins

View on GitHub

Highlights

  • OAuth subscription logins (#147) — new host-local lastlight oauth commands (login / list / status / test / logout) to authenticate the model provider via ChatGPT/Codex, Claude Pro, or Copilot. Credentials are stored as auth.json under $STATE_DIR and consumed by the agent runtime via agentic-pi’s new authFile run option.
  • deps: bump agentic-pi → 0.2.14 (adds authFile support).

Full changelog: https://github.com/nearform/lastlight/compare/v0.10.4...v0.10.5

v0.10.4 repo moved to nearform org

View on GitHub

Maintenance release — repo home moved from cliftonc to the nearform GitHub org. Canonical repository metadata, plugin manifest, and documentation links updated to nearform. No functional changes.

https://github.com/nearform/lastlight/compare/v0.10.3...v0.10.4

v0.10.3 setup wizard provider picker

View on GitHub

Highlights

  • Setup wizard provider picker (community PR #143 by @servatj) — the install wizard’s step-4 now surfaces pi-ai’s full wizard-able provider set (17 providers) instead of a hardcoded 3-provider shortlist, with per-provider API-key prefix validation and a “custom” escape hatch for any provider/model string.
  • New provider registry (src/providers.ts) — a single source of truth wired end-to-end:
    • agent-executor.ts forwards every registered provider’s env key into the sandbox
    • egress-allowlist.ts derives PROVIDER_HOSTS from the registry
    • llm.ts (screener/classifier cheap helper) routes via two generic families (anthropic-messages + openai-completions), covering all wizard-able providers
  • Backward-compatible: existing OPENAI / ANTHROPIC / OPENROUTER_API_KEY setups keep working unchanged.

Providers: Anthropic, OpenAI, OpenRouter, Google (Gemini), Mistral, Groq, Cerebras, xAI, Hugging Face, Moonshot, NVIDIA, Fireworks, Together, DeepSeek, Z.AI, Kimi for Coding, MiniMax.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.10.2...v0.10.3

v0.10.2 Fix PR-review posting for webhook-triggered PRs

View on GitHub

Fixes

  • pr-review: carry the PR number into the run context so post-review can actually post. A pr.opened/synchronize/reopened webhook routes with only prNumber (the router drops the issueNumber mirror), so the runtime context was built with issueNumber: 0 and no prNumber — post-review then failed with “no PR number in run context; cannot post review”. The review ran and produced findings, but was never posted. This broke every pr-review triggered by a real PR webhook. Verified in production against a live reopened-PR webhook.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.10.1...v0.10.2

v0.10.1 project-independent agent image + explicit server build

View on GitHub

Fixes

  • Deployment no longer requires the checkout dir to be named lastlight. The agent service is now pinned to image: lastlight-agent, so its built tag is independent of the compose project name (which defaulted to the working-directory basename). Previously a serverHome like ~/work/lltest made docker compose up try to pull lastlight-agent and fail with “pull access denied”.
  • New lastlight server build — the explicit image-build step (agent + sandbox + sandbox-qa; no git pulls, no up). First-run is now server build → server start; server update still folds build + up together.
  • server start pre-checks the image. If lastlight-agent isn’t built, it points at server build instead of failing on an opaque docker pull error.
  • Setup papercuts. server setup now prints the absolute overlay path (the message read “in instance/” but writes to <home>/instance), and warns loudly when the working dir came from a saved serverHome rather than the folder you’re in.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.10.0...v0.10.1

v0.10.0 GitHub-free mode + PAT fallback

View on GitHub

Highlights

Last Light can now run without a GitHub App — easier local setup.

  • Always-on HTTP server. The HTTP server was owned by the GitHub webhook connector, so with no GitHub App nothing listened on the port and the lastlight CLI + dashboard couldn’t connect. It’s now owned by main(); the webhook connector just registers /webhooks/github. Admin, /api/*, /health, and the Slack webhook mount are always available.
  • Chat-only mode. With only a provider API key, the CLI chat + dashboard work with zero GitHub configured.
  • PAT fallback (GITHUB_TOKEN). A read-only Personal Access Token (far easier than a full App — no webhook, no PEM, no installation) enables read-only GitHub tools in chat plus CLI-driven read-only workflows (triage / review / health). App always wins when both are set; the PAT is static, so a read-only fine-grained PAT is the safe default and repo-write profiles warn.
  • Three-way setup wizard. lastlight setup now offers Full GitHub App / PAT / Chat-only.
  • Chat prompt only advertises GitHub tools when auth is present; cron jobs are skipped without a GitHub client (no chat-only no-op dispatch noise).

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.9.0...v0.10.0

v0.9.0 lastlight-evals-loop score-improvement skill

View on GitHub

Highlights

  • New skill: lastlight-evals-loop — a disciplined, anti-gaming loop that drives a pr-review eval toward a target F1. Diagnoses on a TRAIN split, validates on a BLIND held-out split, proposes ONE generic overlay fix per iteration, and keeps it only if the held-out split doesn’t regress. Prefers generic prompt/skill/persona edits (auto, auditor-gated) and stops for human sign-off before any per-repo context or gold/dataset edit. Ships references/{levers,guardrails,journal-format}.
  • Discoverability updates across the plugin README, lastlight-guide (quick-menu + overlay↔evals bridge), and the lastlight-evals skill, plus docs for the repo-context injection lever (--no-inject-context, repo-context/ + context/<id>/).

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.8.0...v0.9.0

v0.8.0 lastlight-guide + PR-review dataset authoring

View on GitHub

New: lastlight-guide skill + PR-review dataset authoring

  • lastlight-guide — a plugin-wide orientation/router skill over the server / client / overlay / evals skills. Agent-invocable, with an explicit trigger boundary so concrete asks still go straight to the owning skill (or /lastlight-guide).
  • lastlight-evals — a “Start here” front door + an interactive “build a PR-review dataset from gold PRs” playbook (prompt for one/many URLs → curate → smoke-run), a new authoring-pr-review.md, --code-fix added to the code-fix authoring docs (--pr now needs an explicit kind), and pr/review_gold documented in instance-schema.md.

Pairs with lastlight-evals v0.5.0 (add-case --pr <url> --review).

v0.7.8 gondolin skills + pr-review inline anchors

View on GitHub

Fixes

  • fix(sandbox): copy skill bundle for gondolin, not symlink. gondolin mounts only the agent cwd into the guest, so a symlinked skill bundle dangled (its target lives in the install tree, outside the mount) and the agent couldn’t read SKILL.md. Now copies the dereferenced tree into the mounted cwd, exactly as docker does; none keeps the cheaper symlink. Adds a regression test.
  • fix(pr-review): deepen base history on no-merge-base. When a PR forked far behind the branch tip, the --depth 50 base fetch left the merge-base beyond the shallow boundary and the three-dot diff died with “no merge base” — demoting every finding to the PR body. Now --unshallowes and retries once so the review still anchors inline.

Docs

  • docs(release): release when the lastlight/evals barrel surface changes too. Harness changes on the workflow-execution path that lastlight-evals drives via src/evals-api.ts now warrant an npm release even when the CLI is untouched — which is exactly why this release exists.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.7...v0.7.8

v0.7.7 first-class PR-review posting

View on GitHub

Note: no CLI-surface change in this release — the fixes below reach production via `lastlight server update`, not npm. Cut to mark the milestone.

Highlights

  • PR reviews now post via a first-class, in-process action (`type: post-review` → `PhaseExecutor.runPostReview`), replacing the ~150-line in-sandbox script that depended on the AI hand-writing `pr_number`/`base_ref`/`head_sha` into findings.json and silently no-op’d on any mismatch. The reviewer agent now writes content only (`{ skip?, summary, event, findings[] }`); the harness supplies the PR number, base ref, head SHA and diff itself.
  • Genuine posting failures now fail the phase visibly (checklist + dashboard pipeline) instead of masquerading as success; a legitimate skip succeeds without posting. Idempotent on resume.
  • New tested `review-poster` module + `GitHubClient.createPullRequestReview`/`getPullRequestDiff`; skill contract (SKILL.md v7) and spec updated.
  • Dashboard: opening a run via `?run=…` deeplink now auto-selects the run’s first substantive phase (shows logs immediately); teaches the UI the `bash`/`script`/`post-review` phase types.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.6...v0.7.7

v0.7.6 30-day tokens, sliding auto-refresh, login fixes

View on GitHub

Highlights

  • 30-day admin tokens (up from 7) with a 7-day grace window — a briefly-lapsed session can still renew instead of forcing a full re-login.
  • CLI auto-refresh — the saved token silently renews once past its half-life (or within the grace window), so an active session slides forward indefinitely and never hands the CLI a dead-on-arrival token.
  • Login hang fixed — the browser-handoff loopback no longer keeps the process alive (Connection: close + explicit exit).
  • Expired-token guard — lastlight login refuses to persist an already-expired token (e.g. from a stale dashboard) and points you at --password.

Notes

  • Server + dashboard changes reach prod via lastlight server update (already deployed).
  • Update the CLI with npm i -g lastlight@0.7.6.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.5...v0.7.6

v0.7.5 pr-review is a pure code review

View on GitHub

Highlights

  • pr-review is now a pure code review. The workflow no longer installs, builds, or runs the repo — it reads the diff and reasons statically, and CI validates that the change actually works. Dropped the building skill from the pr-review phase (skills: [pr-review, code-review]), skill → v5.0.0.
  • Precision-first reviews. code-review v2.0.0 + pr-review post only Critical / Important findings (Suggestions/Nits dropped as noise), each with a concrete-impact line, past a self-refutation confidence gate. Tuned against Martian’s Code Review Bench (F0.5).

Also included since v0.7.4 (previously unreleased):

  • demo: human-paced recordings with a visible cursor + slow typing
  • dashboard: skill-badge label fixes + plural skills[] in the workflow browser
  • guardrails: install deps before running check commands

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.4...v0.7.5

v0.7.4 sandbox repoSubdir + evals docs

View on GitHub

Highlights

  • sandbox: repoSubdir — lets a caller that pre-seeds a <workspace>/<repo>/ checkout (the evals harness in static-token mode) nest the agent cwd the same way production does, without the token-gated clone path. Workspace root stays the home for AGENTS.md / .lastlight-skills/ (siblings outside the repo’s git tree).
  • evals skill docs — document suite vs hold-out grading modes, lastlight-evals --version, and TAP reporter setup for vitest/mocha/jest.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.3...v0.7.4

v0.7.3 lastlight version / --version

View on GitHub

Highlights

  • CLI version output — lastlight version, lastlight --version, and lastlight -v now print the CLI version (with --json support). The bare lastlight help banner also shows the version in its header.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.2...v0.7.3

v0.7.2 evals add-case skill docs

View on GitHub

Docs

  • lastlight-evals skill → v1.2.0: documents add-case. Adds guidance for the lastlight-evals add-case --pr/--issue subcommand — author a code-fix eval case from a real GitHub PR (git-source: base/head commits, held-out test_patch, auto-detected FAIL_TO_PASS/PASS_TO_PASS) or a triage case from an issue (problem statement, applied labels, reviewer comments). Includes the new git-source case flavor (test_cmd/setup_cmd/head_commit) in instance-schema.md and a new authoring-from-pr.md how-to. (aa88f5d)

These ship in the npm package + via lastlight skills install, so lastlight-evals consumers pick them up on update.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.1...v0.7.2

v0.7.1 PR-targeted workflow branch fixes

View on GitHub

Fixes

  • qa-test / verify now test the PR’s real head ref. These workflows synthesized a lastlight/<prNumber>-<title-slug> branch that never matched a PR’s actual head ref (named after the originating issue), so the sandbox fell back to cloning the default branch — QA’ing the base without the PR’s changes and reporting features missing (a false-negative QA). They now resolve pr.head.ref and pre-clone the actual PR code. (436558a)
  • pr-fix bails fast on fork PRs. A fork PR’s head branch lives on a repo we can’t push to and isn’t on this repo’s origin, so pr-fix would clone the wrong branch and fix the wrong code. It now detects cross-repo (and deleted-fork) PRs before provisioning a sandbox, posts a short notice, and skips. (abea8dc)

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.7.0...v0.7.1

v0.7.0 fork from CLI-bundled assets + `fork all`

View on GitHub

Highlights

  • lastlight fork no longer needs a core checkout. It reads the built-in workflows / skills / agent-context it forks from out of the assets bundled in the npm package, so it works from a deployment overlay or a Last Light Evals workspace with the CLI installed and no git checkout anywhere. A colocated checkout (or a server home that is one) is still preferred when present, so local unpublished asset edits get forked.
  • New lastlight fork all — fork every workflow (plus the prompts & skills each references) and all agent-context in one pass, deduping shared assets. The “fully fork the defaults” shortcut for evals/overlay tuning.
  • Evals-workspace aware — running fork from a workspace root that contains an overlay (instance/ + evals/) targets that instance/.
  • Overlay skill docs updated to drop the old “needs a checkout” framing and cover the evals case + fork all.

Internal

  • Finishes the src/ regroup begun in the prior refactor: removes the duplicate old-location files that were left tracked, points the package bin at dist/cli/cli.js, adds a typecheck:test script + tsconfig.test.json, and fixes import paths and spec docs to match.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.6.2...v0.7.0

v0.6.2 evals skill: separate/plain layouts + auto-detect

View on GitHub

What’s changed

Refines the lastlight-evals Claude Code skill to match the current lastlight-evals CLI:

  • Separate layout (recommended): lastlight-evals init <dir> --clone <org/overlay-repo> clones an existing deployment overlay into instance/ as its own checkout; the overlay (./instance) and datasets (./evals/datasets) are auto-detected, so a bare lastlight-evals run works with no --overlay flag.
  • Plain layout: the self-contained overlay+evals workspace, run with --overlay ..
  • Notes that init is non-interactive without a TTY (agent/CI/piped), and corrects dataset authoring paths to the workspace’s evals/ dir.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.6.1...v0.6.2

v0.6.1 Claude Code skills marketplace + smolvm sandbox backend

View on GitHub

Highlights

Claude Code skills marketplace (#138)

Last Light now ships a Claude Code plugin (marketplace at the repo root, one lastlight plugin under plugins/) that teaches Claude Code to install, configure and operate Last Light + Last Light Evals. Four skills:

  • lastlight-server — install & configure a server (agent + docker stack)
  • lastlight-client — point the CLI at a server and log in
  • lastlight-overlay — create a deployment overlay and fork workflows/prompts/skills/persona
  • lastlight-evals — scaffold & run a Last Light Evals workspace

New host-local CLI command installs them into a local Claude Code:

lastlight skills install [--scope user|project] [--no-marketplace]
lastlight skills list
lastlight skills uninstall

It prefers claude plugin marketplace add of the bundled, version-matched path (works offline, even on a non-git npm install) and falls back to copying the skills into ~/.claude/skills. Assets ship in the npm package.

smolvm micro-VM sandbox backend

New smol sandbox backend — a structural peer of the docker backend that runs agent tasks inside local micro-VMs (LASTLIGHT_SANDBOX=smol). Also splits agent-executor for the shared path.

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.5.1...v0.6.1

v0.5.1 extract eval harness; add lastlight/evals public API

View on GitHub

Extracts the eval harness into a standalone repo (lastlight-evals) and exposes the public API it depends on.

Highlights

  • New lastlight/evals public API barrel — getWorkflow, runWorkflow, ExecutorConfig, TemplateContext, plus the overlay bootstrap helpers (detectGh, bootstrapOverlayRepo, scaffoldOverlayFiles).
  • exports map added to package.json with a ./dist/* back-compat wildcard, so existing deep lastlight/dist/... importers keep working.
  • Eval harness removed from core (evals/, the eval/eval:compare scripts). It now lives in its own repo and consumes core via lastlight/evals.
  • Slim seam guard retained (src/engine/agent-executor.seam.test.ts) proving core still forwards ExecutorConfig.githubApiBaseUrl into agentic-pi.

Why

This is the deliberate release that publishes the ./evals barrel so the new lastlight-evals package can depend on lastlight@^0.5.0.

No runtime/behaviour change to the harness, workflows, or dashboard.

v0.4.0 lastlight fork + overlay override visibility

View on GitHub

Highlights

lastlight fork — copy built-in assets into the instance/ overlay so a deployment can customise them (the overlay wins by logical name at startup), instead of hand-copying files.

  • lastlight fork <workflow> — copies the workflow YAML plus every prompt and skill its phases reference.
  • lastlight fork agent-context [file] — copies the persona files (soul.md / rules.md / security.md), all or one. Forked only via the explicit target, never inferred from a bare filename.
  • lastlight fork — lists forkable targets and marks what’s already forked.
  • Resolution: explicit --home wins, else the cwd if it’s an overlay/checkout, else the server home. Skip-existing by default; --force overwrites.

Override visibility — forked assets are surfaced everywhere, each tagged shadows default vs added, via a shared enumerateOverlayAssets enumerator:

  • lastlight server status gains an Overrides section.
  • GET /admin/api/overrides exposes them.
  • The dashboard Config tab gains an Overrides pane.

setup wizard now delegates its docker build/launch to the canonical server update flow (live progress, builds agent + sandbox + sandbox-qa, restarts sidecars, health-checks) instead of re-implementing compose.

Upgrade the CLI: npm i -g lastlight@0.4.0

Full changelog: https://github.com/cliftonc/lastlight/compare/v0.3.0...v0.4.0

v0.2.3 server update builds sandbox-qa; retire deploy.sh

View on GitHub

Changes

  • lastlight server update now also builds the browser-QA sandbox image (docker compose build sandbox-qa, non-fatal, after the base sandbox) — so it fully reproduces the old production deploy.sh.
  • deploy.sh is retired: lastlight server update is now the canonical deploy path. Docs recentered on the CLI.

Full Changelog: https://github.com/cliftonc/lastlight/compare/v0.2.2...v0.2.3

v0.2.2 auth required with OAuth (security fix)

View on GitHub

Security fix

Dashboard + /api/* auth was gated only on ADMIN_PASSWORD. Clearing the password left both surfaces fully open even when Slack/GitHub OAuth was configured.

Now auth is required when any login method is set — a password or a working OAuth provider (Slack needs client id + secret; GitHub also needs GITHUB_ALLOWED_ORG). Both the dashboard and the trigger API share one authIsEnabled() gate. POST /login no longer hands out an open-access token when auth is on but no password is set. The login screen hides the password box for an OAuth-only gate.

The dashboard is only fully open when no login method is configured.

Full Changelog: https://github.com/cliftonc/lastlight/compare/v0.2.1...v0.2.2

v0.2.1 fix executable bin

View on GitHub

Fix

The published dist/cli.js was packed with mode 644, so on some installs (notably nvm + npm under Node 24) the global lastlight binary wasn’t executable — bash: …/bin/lastlight: Permission denied. The build now sets the executable bit so the published tarball carries it.

If you hit this on 0.2.0, either reinstall (npm i -g lastlight@latest) or chmod u+x $(which lastlight).

Full Changelog: https://github.com/cliftonc/lastlight/compare/v0.2.0...v0.2.1

v0.2.0 server lifecycle CLI + drift banner

View on GitHub

Highlights

lastlight server lifecycle commands — the CLI is now the control plane for the docker stack, run host-local against a working directory (repo checkout + instance/ overlay + compose override symlink):

  • server setup — scaffold/adopt the working dir (clones core + overlay)
  • server start|stop|restart [service]
  • server update — deploy.sh equivalent: pull core + overlay → build (stamping GIT_SHA) → up -d --remove-orphans → restart egress sidecars → health-check, with live progress
  • server status — compose state + core/overlay version drift

Working dir resolves from --home → LASTLIGHT_HOME → saved serverHome → ~/lastlight. The CLI runs on the host, so it survives the agent container recreating itself during an update.

lastlight setup now asks client vs server first (--client / --server skip the prompt).

Version-drift detection — GET /admin/api/server/info compares the baked LASTLIGHT_GIT_SHA (new Dockerfile ARG) and overlay HEAD against their remotes; a dashboard UpdateBanner nudges to run server update when behind. behind is only true when both SHAs are known, so an unreachable remote never false-flags.

Notes

  • New env vars: LASTLIGHT_HOME, LASTLIGHT_GIT_SHA / LASTLIGHT_BUILD_DATE.
  • Docs synced (CLAUDE.md, spec/02-configuration.md); the lastlight-www site is updated separately.

Full Changelog: https://github.com/cliftonc/lastlight/compare/v0.1.15...v0.2.0

v0.1.14

View on GitHub