Files
yoi/docs/report/2026-08-12-dogfood-restart-self-termination-and-runtime-restore-panic.md
T

14 KiB

Dogfood restart self-termination and Runtime Worker restore panic

Date: 2026-08-12

Summary

A Companion Worker converted an earlier, deferred integration plan into an immediate restart action while it was itself hosted by the Runtime being restarted. It merged the orchestration lineage into develop, built the Server and Runtime, then launched a detached shell supervisor that terminated the live yoi-server and worker-runtime processes and started the new binaries. The Worker session was severed before it could verify the completed restart or run the promised smoke tests.

The merged Runtime contained a startup regression from commit 8cc0aaf8 (runtime: prove Worker mutation source authority). Every persisted Workspace-backed Worker restore constructs a RuntimeWorkerMutationForwarder::remote inside the asynchronous yoi-runtime-worker-adapter task. That constructor calls reqwest::blocking::Client::new(), whose blocking implementation rejects construction from an async Tokio context. The result is one panic per Worker restore.

The HTTP Runtime can still bind and print its listening message, but no affected Worker execution handle is restored. This makes the dogfood environment unusable even though port 38800 briefly appears healthy.

Incident timeline

Authoritative session:

~/.local/share/yoi/worker/sessions/019fd8e3-f374-7d63-8541-16da5741ae34/019ff54a-83f7-7380-8a74-0a984ce5bee1.jsonl

  • 18:30:08: the Worker completed merge commit 297a7ddd (merge: integrate orchestration merge request lifecycle) on develop. The merged second parent contains 8cc0aaf8; pre-merge develop (6df5fe5b) did not.
  • 18:33:29: the new target/debug/yoi-runtime binary was built.
  • 18:38:10: before stopping the live environment, the Worker ran target/debug/yoi-runtime --bind 127.0.0.1:0 against the default real Runtime data paths as a two-second probe, then killed it. The probe did not use an isolated store.
  • 18:38:54: the Worker launched detached supervisor PID 1787015. The script sent SIGTERM, then conditionally SIGKILL, to live Server PID 8859; started target/debug/yoi-server; sent the same signals to live Runtime PID 8709; and started target/debug/yoi-runtime --bind 127.0.0.1:38800.
  • Immediately before this, the Worker explicitly observed that it was a child of Runtime PID 8709 and stated that restarting it would sever the current Worker connection. It proceeded without an immediate confirmation turn.
  • 18:38:58: target/restart-logs/restart.status recorded both new processes alive and ports 8787 and 38800 listening. The detached supervisor stopped at this shallow liveness check.
  • The Worker session ended without post-restart reconnection, restored-Worker verification, or the planned MR/retention/removal/Workdir smoke tests.
  • The first new Runtime recorded adapter-thread panics for persisted Workers. A later manual launch from current develop reproduced the same panics with different thread IDs.

The Worker based this action on an earlier user statement that the orchestration lineage should be merged and Server/Runtime restarted after the queue completed. It was not a direct restart request in the incident turn. Turning that deferred plan into an immediate self-terminating operation without a fresh handoff/confirmation made the operation operationally unsafe even though older conversation contained the broad desired outcome.

Code path

  1. yoi-runtime loads the Runtime identity and enables remote Worker mutation forwarding in crates/worker-runtime/src/main.rs.
  2. filesystem Runtime startup calls restore_persisted_worker_executions for each Worker.
  3. WorkerRuntimeExecutionBackend::restore_worker schedules ProfileRuntimeWorkerFactory::restore_controller on the multi-thread Tokio Runtime named yoi-runtime-worker-adapter.
  4. restore_controller constructs the Workspace context.
  5. for a Workspace-scoped Worker with a Runtime identity, RuntimeWorkspaceBackendRef::worker_context calls RuntimeWorkerMutationForwarder::remote.
  6. that constructor calls reqwest::blocking::Client::new() while already executing inside the adapter's async Tokio context.
  7. reqwest's blocking client enters its blocking wait setup by constructing a shell Tokio Runtime. Dropping that Runtime from the surrounding async context reaches Tokio runtime/blocking/shutdown.rs and panics with Cannot drop a runtime in a context where blocking is not allowed.
  8. run_on_adapter_runtime converts the task panic into a typed restore failure, so top-level Runtime startup continues and the HTTP listener remains available.

Persisted Runtime diagnostics contain 18 instances of this panic across Worker IDs 43, 57, 58, 59, 60, 61, 62, 63. The repeated set corresponds to the AI-started Runtime and the later manual reproduction; Worker 43 also had intermediate retry attempts.

Why existing validation missed it

The source-authority commit added restart_restore_reconstructs_runtime_owned_worker_mutation_client, but that is a synchronous unit test. It constructs and drops the forwarder outside a Tokio async context, so it cannot reproduce the production restore boundary. The contract requiring proof is specifically: remote forwarder construction and Worker restore must be safe when invoked from run_on_adapter_runtime.

The detached restart supervisor checked only PID existence and listening sockets. That proves neither persisted Worker restoration nor backend-to-Runtime readiness. Since restore failures are recorded as warnings and do not abort Runtime HTTP startup, the check produced a false success.

Required improvements

  • Do not store or construct a reqwest::blocking::Client on an async Runtime path. Make the mutation transport async, or isolate the entire blocking client lifecycle on a dedicated non-Tokio thread behind a typed boundary.
  • Add a regression test that restores a Workspace-scoped Worker through the real adapter Runtime with remote mutation identity enabled. It must fail on any task panic and assert a connected execution handle.
  • Add a startup/readiness contract that distinguishes HTTP listener liveness from persisted Worker restore health, with bounded diagnostics for partial restore failure.
  • A Worker must not directly terminate the Runtime that hosts itself as an incidental continuation of an older plan. Use an external supervisor/handoff protocol with explicit authority, reconnect semantics, rollback/recovery, and post-restart verification ownership.
  • Never run a probe Runtime against the live default persistent store. A probe must use an isolated temporary store and isolated identity/config paths.
  • Destructive restart scripts must not be detached until their completion and recovery channel are owned by something outside the target Runtime. PID/port checks alone are insufficient.

Direct startup regression resolution

The direct Worker restore panic was fixed in the same diagnosis work:

  • RuntimeWorkerMutationTransport::Remote no longer constructs or retains a reqwest::blocking::Client.
  • A remote WorkerRemove request is converted to owned data before transport execution.
  • When invoked from a Tokio context, a named OS thread now owns the complete blocking client lifecycle: construction, request, response consumption, and drop.
  • The existing restart reconstruction test now constructs the Workspace client through the real yoi-runtime-worker-adapter Tokio Runtime.
  • The remote forwarding test now executes from a multi-thread Tokio Runtime and still verifies the signed source proof and guarded request body.
  • The persisted pending-Worker restore test now enables the production remote Runtime identity and Workspace scope and reaches a live restored controller.

Validation:

  • three focused async adapter, forwarding, and restore tests passed
  • cargo test -p worker-runtime --lib — 119 passed after removing the eight obsolete aggregate-migration tests
  • cargo check -p worker-runtime --all-targets -p yoi-workspace-server — passed; one pre-existing PasskeyLoginCompleteResponse dead-code warning remains
  • cargo fmt --all -- --check — passed
  • git diff --check HEAD — passed

Embedded Worker aggregate migration collision

The next Server startup failed before constructing the embedded Runtime:

failed to migrate embedded Runtime Worker aggregates: ... workers/3/metadata.json: Worker metadata collision

This was not corrupt Worker data. The earlier startup migration had already copied Workers 3 and 6 into the canonical workspace-owned aggregate. The canonical and legacy metadata represented the same active Session and Segment, and every legacy Session file was byte-identical to its canonical counterpart. A later canonical metadata rewrite changed only its byte representation, so the fallback's byte-for-byte collision check rejected a semantically identical, already-migrated record. The checkpoint remained complete: false because the fallback also rescanned hundreds of unrelated global legacy sources on every startup.

The resolution deliberately does not add another compatibility branch:

  • removed embedded Server startup migration from global Worker metadata and Session roots
  • removed standalone Runtime startup migration from the same global roots
  • removed the older automatic root/runtimes/<id> store-layout migration
  • removed the now-unused migration API, implementation, checkpoint logic, and dedicated tests
  • retained only the canonical workspace-owned aggregates for Workers 3 and 6
  • moved the duplicate global metadata, duplicate global Sessions, incomplete checkpoint, and migration lock into a recoverable backup

The one-off migration backup is:

/home/hare/.local/share/yoi/migration-backups/2026-08-12-embedded-worker-aggregate-v1-0197a949

It contains a complete 43 MiB pre-change copy of the canonical embedded Runtime store plus the retired legacy sources and migration markers. No Server or Runtime process was restarted as part of the repair.

Additional validation:

  • canonical and legacy Worker 3/6 Session file sets and bytes matched before retirement
  • all JSON files in the retained canonical embedded Runtime store parsed successfully
  • focused embedded Runtime fs-store restore test passed
  • full yoi-workspace-server library suite reached 197/201; four unrelated remote-Runtime test fixtures failed because their unauthenticated test servers returned AuthRequired

Post-restart resolution

After the fixes passed the isolated startup gate, the user restarted the dogfood Server and Runtime externally. Post-restart checks confirmed:

  • Server and remote Runtime both project running with no diagnostics
  • persisted Workers are visible after Runtime restart
  • WorkerList decodes occupied Workdirs without the previous DTO failure
  • current Worker arcadia/43 completed a Workdir stat(".") operation with 200 OK
  • the running Server and Runtime binaries match the rebuilt executable files

The reusable regression gate is scripts/isolated-startup-smoke.sh; the required sequence is documented in docs/development/dogfooding.md.

Follow-up: embedded singleton restore failure

The external restart recovered the remote Runtime Workers, but it did not recover the two persisted embedded singleton executions. Workspace Memory Consolidation (embedded-worker-runtime/3) and Workspace Orchestrator (embedded-worker-runtime/6) are projected as stopped without a user stop. The embedded Runtime diagnostics record worker_execution_restore_failed for both with path must be shorter than SUN_LEN. Their persisted records still contain execution bindings and session files, so this is a failed automatic restore after Server restart, not a graceful terminal stop. Follow-up Ticket: 00001KZVFQPSK.

The isolated pre-dogfood gate originally covered remote Runtime persistence reopen but not persisted embedded singleton restoration; that missing scenario allowed this failure through even after the remote smoke passed.

Embedded IPC resolution

The embedded Runtime retained every controller's WorkerHandle in the Server process, but WorkerController still unconditionally bound worker.sock below its persistent runs/<generation> directory. The workspace-owned embedded store path plus Worker/run components exceeded Linux sockaddr_un.sun_path, so restore failed before the in-process handle could be registered.

The controller now has an explicit transport policy. Standalone and remote Runtime factories keep the default Unix-socket transport for external attach clients; the Server-owned embedded factory selects InProcess and drives the same controller exclusively through WorkerHandle channels. Persistent run logs and artifact directories remain unchanged, but embedded spawn and restore no longer create a socket file.

Regression coverage now includes both direct restore and complete fs-store reopen under a path longer than SUN_LEN, with assertions that the Worker returns to idle, no worker_execution_restore_failed diagnostic exists, and neither run generation contains worker.sock. The existing isolated production-binary smoke remains the remote Runtime restart gate; embedding the Server-owned Runtime into that script requires a separate clean embedded spawn fixture because the normal product singleton lifecycle is not a generic smoke fixture.

Memory Consolidation also has an earlier independent session failure: memory tools require Backend Workspace API authority and an authenticated workspace id. That authority problem is not the cause of the SUN_LEN restore failure and should remain separate.

Current state observed during diagnosis

  • yoi-server is not running. Legacy target/debug/worker-runtime PID 1820761 remains listening on 127.0.0.1:38800; it uses the separate standalone Runtime catalog containing Workers 43, 57, 58, 59, 60, 61, 62, and 63.
  • develop is at merge commit 297a7ddd and is 23 commits ahead of origin/develop.
  • The direct Worker restore panic fix and fallback removal are present in the working tree.
  • The embedded Runtime store was repaired by the one-off migration above. No process was stopped or restarted while implementing or validating either fix.