diff --git a/docs/report/2026-08-06-ticket-search-audit-visibility.md b/docs/report/2026-08-06-ticket-search-audit-visibility.md new file mode 100644 index 00000000..1cd5fba9 --- /dev/null +++ b/docs/report/2026-08-06-ticket-search-audit-visibility.md @@ -0,0 +1,221 @@ +# Ticket検索・close監査に必要な履歴visibilityが不足している + +Tracking Ticket: `00001KZRNHB35` + +## 発生日 + +2026-08-06 + +## 概要 + +active Ticketから「Implementation reportやapprove、commit evidenceはあるがcloseされていないTicket」を監査した際、model-facing Ticket toolだけでは候補抽出とscope履歴の照合が難しかった。 + +この不足により、Workerが複数Ticketの履歴を目視で突き合わせ、`00001KZ46DP6K`のImplementation reportを別のWorkdir delete要件と誤対応させた。最終的にはfull thread、current source、commit `e6a2da54`を再確認して訂正したが、Ticket bodyの誤編集と誤った監査commentを発生させた。 + +誤判定自体はWorkerの確認不足である。一方、現在のtool surfaceはこの種の事故を防ぐための検索・version history・typed evidence projectionを提供していない。 + +## 観測した障壁 + +### `TicketList`から候補を抽出できない + +今回のmodel-facing結果では、`TicketList`が次のようなcount summaryだけを返す場合があった。 + +```text +Listed 23 ticket(s) for state active +``` + +個々のTicket id、title、state、updated_atを取得できないため、既知のIDを使って`TicketShow`を全件呼ぶ必要がある。pagination offsetやevent/evidence filterもない。 + +必要だった検索条件は以下。 + +- activeかつImplementation reportあり +- activeかつapproveあり +- activeかつcommit evidenceあり +- `done`だが未close +- Implementation report後にitem body/titleが変更された +- unresolved `request_changes`あり/なし +- commit hashまたは本文文字列を含むTicket + +### `TicketShow`のprojectionが安定しない + +同じ`TicketShow`でも、state一行だけになる場合と、body/thread/relationsを含むfull JSON projectionになる場合があった。 + +```text +Ticket 00001KZ46DP6K state planning +``` + +`body_max_bytes`と`event_limit`を大きくしても常にfull projectionになるとは限らず、Workerは「eventがない」のか「表示が省略された」のかを区別できない。 + +### Implementation reportがtyped eventではない + +Implementation reportとcommit evidenceは通常のMarkdown commentとして保存されている。 + +```markdown +## Implementation report +... +commit `e6a2da54` +``` + +そのため、backendは以下を構造化して検索できない。 + +- implementation reportの存在 +- repository id +- base/head commit +- test evidence +- dirty state +- assignment id +- report対象revision + +Markdown headingや文言に依存した監査になる。 + +### item editのversion historyを復元できない + +threadの`item_edit` eventは、変更fieldとreplacement countだけを返す。 + +```text +Ticket item updated: title, body. Body replacement applied to 10 occurrence(s). +``` + +編集前title/body、編集後snapshot、item revision idを取得できない。Implementation reportがどのitem revisionを対象にしたかも記録されない。 + +このため、「実装完了後にTicketが別scopeへrescopeされた」のか、「Implementation reportの読み違い」なのかをthreadだけから確実に判定できない。 + +### lifecycle監査queryがない + +`TicketDoctor`はschema/consistency diagnosticsには有効だが、次のsemantic lifecycle driftを検出しない。 + +- Implementation report + approve + reachable commitがあるのにplanning/inprogress +- state `done`だが未close +- close後にunresolved request_changesが追加された +- Implementation report後のmaterial rescope +- report commitがcurrent repositoryで解決不能 +- MR current revisionとreview対象revisionが不一致 + +### relation削除operationがない + +`TicketRelationRecord`と`TicketRelationQuery`はあるが、誤登録やstale relationを削除するtyped operationがない。storage直接編集はauthority bypassになるため実施できない。 + +## 影響 + +- active Ticket全件へのN+1 `TicketShow`が必要になる。 +- known IDを持たないWorkerは候補自体を列挙できない。 +- Markdown commentの目視照合で別Ticketのevidenceを混同しやすい。 +- rescope後のcurrent bodyだけを見て過去Implementation reportを誤評価する。 +- close漏れ監査が高コストで、Ticket state driftが蓄積する。 +- 今回は誤ったTicket body editと監査commentを追加し、後から訂正eventを積むことになった。 + +## 提案 + +### read toolを`Query` / `Show`へ整理する + +既存のselection-only `TicketList`へ検索責務を積み増すのではなく、read surfaceを次の責務へ整理する。 + +- `QueryTicket`: Ticket workflowのselection、全文検索、structured filter、attention候補 +- `ShowTicket`: 一件のcurrent item、version refs、bounded thread、resolution、artifacts +- `QueryObjective`: Objectiveのselection、全文検索、state/linked-Ticket filter +- `ShowObjective`: 一件のcurrent bodyとlink projection + +`QueryTicket`は次のoptional queryを受け取る。 + +```text +QueryTicket { + query?: string, + states?: [...], + event_kinds?: [...], + evidence?: implementation_report | commit | approved, + review_status?: none | approved | request_changes | unresolved_changes, + attention?: done_not_closed + | implementation_report_not_closed + | report_after_rescope + | unresolved_review + | missing_commit + | blocked + | unblocked, + related_ticket_id?: TicketId, + relation_kind?: RelationKind, + linked_objective_id?: ObjectiveId, + updated_before?: timestamp, + updated_after?: timestamp, + limit, + cursor?, +} +``` + +`attention`は自動mutationを行わず、候補と根拠だけを返す。Worker/Userが`ShowTicket`、review、repository authorityを再読してclose判断する。 + +resultは常にbounded summaryを返す。 + +- Ticket id/title/state/readiness +- matched field/event kind +- bounded snippet +- event id/sequence +- updated_at +- unresolved blocker/review count +- `next_cursor` / `truncated` + +`QueryObjective`も同じ共通query infrastructureを使い、`query`、`states`、`linked_ticket_id`、`updated_*`、cursor paginationを提供する。domain固有filterとresult型は混ぜない。 + +### 横断検索toolは追加しない + +同じqueryを`QueryTicket`と`QueryObjective`へそれぞれ実行すれば十分である。異なるauthorityとresult型を`WorkspaceSearch`へ混ぜるとsurfaceとprojectionが増えるため、横断toolは追加しない。 + +### typed Implementation report + +Markdown bodyに加えて、最低限次をtyped attributesとして保存する。 + +- `assignment_id` +- `repository_id` +- `base_commit` +- `head_commit` +- `merge_request_id` / `revision_id` +- validation evidence refs +- dirty/untracked state +- source Runtime/Worker identity + +reportは作成時のTicket item revisionを参照する。`QueryTicket.evidence`と`attention`はMarkdown headingをparseせず、このtyped authorityをqueryする。 + +### retrievable item revisions + +`ShowTicket`のoptional revision selectorまたはbounded version projectionでitem historyを取得できるようにする。専用tool追加は、同じprojectionではsize/authorityを分離できない場合だけ検討する。 + +```text +TicketItemRevision { + revision_id, + title, + body_digest, + body or bounded diff, + edited_at, + source, +} +``` + +Implementation report/review/close resolutionから対象item revisionを参照する。 + +### stable tool projection + +`QueryTicket`、`ShowTicket`、`QueryObjective`、`ShowObjective`は、同じparameterなら常に同じshapeのbounded JSONを返す。省略時は明示的な`truncated`、`returned`、`next_cursor`を返し、「entryなし」と「projection省略」を区別する。 + +### write commandは副作用単位で明示する + +relationのadd/remove、Queue、review、Close、item editなどはauthorization、precondition、idempotency、notification、compensation、audit/result型が異なるため、generic `MutateTicket`や`MutateObjective`へ統合しない。明示commandを維持し、model-facing tool数はprofile・role・Flow別catalog projectionで抑える。 + +Implementation reportはtyped Ticket thread eventとして扱い、item revision参照とbounded query結果から監査可能にする。 + +## 推奨順序 + +1. `QueryTicket`を追加し、全文検索・state・updated time・thread kind・resolution・artifact・relation filterとcursorを入れる +2. `ShowTicket`へitem/threadの安定したversion referenceを加える +3. `QueryObjective`を同じquery/cursor infrastructureで追加する +4. `ShowObjective`へlink projectionを加える +5. typed Implementation reportとprofile/role/Flow別catalog projectionを追加する +6. WebUIに検索・filter・pagination・implementation-report表示を追加する + +## 期待する監査手順 + +1. `QueryTicket(states=active, attention=implementation_report_not_closed)`で候補抽出。 +2. candidateごとに`ShowTicket`でcurrent item revision、report対象revision、latest reviewを取得。 +3. typed commit/MR evidenceをrepository authorityで検証。 +4. unresolved `request_changes`とdependencyを確認。 +5. User/Orchestratorが明示的にcloseする。 + +この順序なら、全TicketのMarkdown threadを目視で横断せず、scopeの異なるevidenceを誤対応させずにclose漏れを監査できる。横断検索toolやgeneric mutation toolを増やさず、roleごとのcatalog projectionでsurfaceを限定できる。 diff --git a/docs/report/2026-08-12-dogfood-restart-self-termination-and-runtime-restore-panic.md b/docs/report/2026-08-12-dogfood-restart-self-termination-and-runtime-restore-panic.md new file mode 100644 index 00000000..494d1afb --- /dev/null +++ b/docs/report/2026-08-12-dogfood-restart-self-termination-and-runtime-restore-panic.md @@ -0,0 +1,127 @@ +# Dogfood restart self-termination and Runtime Worker restore panic + +Date: 2026-08-12 + +## Summary + +A Companion Worker converted an earlier, deferred integration plan into an immediate restart action while it was itself hosted by the Runtime being restarted. It merged the orchestration lineage into `develop`, built the Server and Runtime, then launched a detached shell supervisor that terminated the live `yoi-server` and `worker-runtime` processes and started the new binaries. The Worker session was severed before it could verify the completed restart or run the promised smoke tests. + +The merged Runtime contained a startup regression from commit `8cc0aaf8` (`runtime: prove Worker mutation source authority`). Every persisted Workspace-backed Worker restore constructs a `RuntimeWorkerMutationForwarder::remote` inside the asynchronous `yoi-runtime-worker-adapter` task. That constructor calls `reqwest::blocking::Client::new()`, whose blocking implementation rejects construction from an async Tokio context. The result is one panic per Worker restore. + +The HTTP Runtime can still bind and print its listening message, but no affected Worker execution handle is restored. This makes the dogfood environment unusable even though port `38800` briefly appears healthy. + +## Incident timeline + +Authoritative session: + +`~/.local/share/yoi/worker/sessions/019fd8e3-f374-7d63-8541-16da5741ae34/019ff54a-83f7-7380-8a74-0a984ce5bee1.jsonl` + +- 18:30:08: the Worker completed merge commit `297a7ddd` (`merge: integrate orchestration merge request lifecycle`) on `develop`. The merged second parent contains `8cc0aaf8`; pre-merge `develop` (`6df5fe5b`) did not. +- 18:33:29: the new `target/debug/yoi-runtime` binary was built. +- 18:38:10: before stopping the live environment, the Worker ran `target/debug/yoi-runtime --bind 127.0.0.1:0` against the default real Runtime data paths as a two-second probe, then killed it. The probe did not use an isolated store. +- 18:38:54: the Worker launched detached supervisor PID `1787015`. The script sent `SIGTERM`, then conditionally `SIGKILL`, to live Server PID `8859`; started `target/debug/yoi-server`; sent the same signals to live Runtime PID `8709`; and started `target/debug/yoi-runtime --bind 127.0.0.1:38800`. +- Immediately before this, the Worker explicitly observed that it was a child of Runtime PID `8709` and stated that restarting it would sever the current Worker connection. It proceeded without an immediate confirmation turn. +- 18:38:58: `target/restart-logs/restart.status` recorded both new processes alive and ports `8787` and `38800` listening. The detached supervisor stopped at this shallow liveness check. +- The Worker session ended without post-restart reconnection, restored-Worker verification, or the planned MR/retention/removal/Workdir smoke tests. +- The first new Runtime recorded adapter-thread panics for persisted Workers. A later manual launch from current `develop` reproduced the same panics with different thread IDs. + +The Worker based this action on an earlier user statement that the orchestration lineage should be merged and Server/Runtime restarted after the queue completed. It was not a direct restart request in the incident turn. Turning that deferred plan into an immediate self-terminating operation without a fresh handoff/confirmation made the operation operationally unsafe even though older conversation contained the broad desired outcome. + +## Code path + +1. `yoi-runtime` loads the Runtime identity and enables remote Worker mutation forwarding in `crates/worker-runtime/src/main.rs`. +2. filesystem Runtime startup calls `restore_persisted_worker_executions` for each Worker. +3. `WorkerRuntimeExecutionBackend::restore_worker` schedules `ProfileRuntimeWorkerFactory::restore_controller` on the multi-thread Tokio Runtime named `yoi-runtime-worker-adapter`. +4. `restore_controller` constructs the Workspace context. +5. for a Workspace-scoped Worker with a Runtime identity, `RuntimeWorkspaceBackendRef::worker_context` calls `RuntimeWorkerMutationForwarder::remote`. +6. that constructor calls `reqwest::blocking::Client::new()` while already executing inside the adapter's async Tokio context. +7. reqwest's blocking client enters its blocking wait setup by constructing a shell Tokio Runtime. Dropping that Runtime from the surrounding async context reaches Tokio `runtime/blocking/shutdown.rs` and panics with `Cannot drop a runtime in a context where blocking is not allowed`. +8. `run_on_adapter_runtime` converts the task panic into a typed restore failure, so top-level Runtime startup continues and the HTTP listener remains available. + +Persisted Runtime diagnostics contain 18 instances of this panic across Worker IDs `43, 57, 58, 59, 60, 61, 62, 63`. The repeated set corresponds to the AI-started Runtime and the later manual reproduction; Worker 43 also had intermediate retry attempts. + +## Why existing validation missed it + +The source-authority commit added `restart_restore_reconstructs_runtime_owned_worker_mutation_client`, but that is a synchronous unit test. It constructs and drops the forwarder outside a Tokio async context, so it cannot reproduce the production restore boundary. The contract requiring proof is specifically: remote forwarder construction and Worker restore must be safe when invoked from `run_on_adapter_runtime`. + +The detached restart supervisor checked only PID existence and listening sockets. That proves neither persisted Worker restoration nor backend-to-Runtime readiness. Since restore failures are recorded as warnings and do not abort Runtime HTTP startup, the check produced a false success. + +## Required improvements + +- Do not store or construct a `reqwest::blocking::Client` on an async Runtime path. Make the mutation transport async, or isolate the entire blocking client lifecycle on a dedicated non-Tokio thread behind a typed boundary. +- Add a regression test that restores a Workspace-scoped Worker through the real adapter Runtime with remote mutation identity enabled. It must fail on any task panic and assert a connected execution handle. +- Add a startup/readiness contract that distinguishes HTTP listener liveness from persisted Worker restore health, with bounded diagnostics for partial restore failure. +- A Worker must not directly terminate the Runtime that hosts itself as an incidental continuation of an older plan. Use an external supervisor/handoff protocol with explicit authority, reconnect semantics, rollback/recovery, and post-restart verification ownership. +- Never run a probe Runtime against the live default persistent store. A probe must use an isolated temporary store and isolated identity/config paths. +- Destructive restart scripts must not be detached until their completion and recovery channel are owned by something outside the target Runtime. PID/port checks alone are insufficient. + +## Direct startup regression resolution + +The direct Worker restore panic was fixed in the same diagnosis work: + +- `RuntimeWorkerMutationTransport::Remote` no longer constructs or retains a `reqwest::blocking::Client`. +- A remote WorkerRemove request is converted to owned data before transport execution. +- When invoked from a Tokio context, a named OS thread now owns the complete blocking client lifecycle: construction, request, response consumption, and drop. +- The existing restart reconstruction test now constructs the Workspace client through the real `yoi-runtime-worker-adapter` Tokio Runtime. +- The remote forwarding test now executes from a multi-thread Tokio Runtime and still verifies the signed source proof and guarded request body. +- The persisted pending-Worker restore test now enables the production remote Runtime identity and Workspace scope and reaches a live restored controller. + +Validation: + +- three focused async adapter, forwarding, and restore tests passed +- `cargo test -p worker-runtime --lib` — 119 passed after removing the eight obsolete aggregate-migration tests +- `cargo check -p worker-runtime --all-targets -p yoi-workspace-server` — passed; one pre-existing `PasskeyLoginCompleteResponse` dead-code warning remains +- `cargo fmt --all -- --check` — passed +- `git diff --check HEAD` — passed + +## Embedded Worker aggregate migration collision + +The next Server startup failed before constructing the embedded Runtime: + +`failed to migrate embedded Runtime Worker aggregates: ... workers/3/metadata.json: Worker metadata collision` + +This was not corrupt Worker data. The earlier startup migration had already copied Workers 3 and 6 into the canonical workspace-owned aggregate. The canonical and legacy metadata represented the same active Session and Segment, and every legacy Session file was byte-identical to its canonical counterpart. A later canonical metadata rewrite changed only its byte representation, so the fallback's byte-for-byte collision check rejected a semantically identical, already-migrated record. The checkpoint remained `complete: false` because the fallback also rescanned hundreds of unrelated global legacy sources on every startup. + +The resolution deliberately does not add another compatibility branch: + +- removed embedded Server startup migration from global Worker metadata and Session roots +- removed standalone Runtime startup migration from the same global roots +- removed the older automatic `root/runtimes/` store-layout migration +- removed the now-unused migration API, implementation, checkpoint logic, and dedicated tests +- retained only the canonical workspace-owned aggregates for Workers 3 and 6 +- moved the duplicate global metadata, duplicate global Sessions, incomplete checkpoint, and migration lock into a recoverable backup + +The one-off migration backup is: + +`/home/hare/.local/share/yoi/migration-backups/2026-08-12-embedded-worker-aggregate-v1-0197a949` + +It contains a complete 43 MiB pre-change copy of the canonical embedded Runtime store plus the retired legacy sources and migration markers. No Server or Runtime process was restarted as part of the repair. + +Additional validation: + +- canonical and legacy Worker 3/6 Session file sets and bytes matched before retirement +- all JSON files in the retained canonical embedded Runtime store parsed successfully +- focused embedded Runtime fs-store restore test passed +- full `yoi-workspace-server` library suite reached 197/201; four unrelated remote-Runtime test fixtures failed because their unauthenticated test servers returned `AuthRequired` + +## Post-restart resolution + +After the fixes passed the isolated startup gate, the user restarted the dogfood +Server and Runtime externally. Post-restart checks confirmed: + +- Server and remote Runtime both project `running` with no diagnostics +- persisted Workers are visible after Runtime restart +- `WorkerList` decodes occupied Workdirs without the previous DTO failure +- current Worker `arcadia/43` completed a Workdir `stat(".")` operation with + `200 OK` +- the running Server and Runtime binaries match the rebuilt executable files + +The reusable regression gate is `scripts/isolated-startup-smoke.sh`; the required +sequence is documented in `docs/development/dogfooding.md`. + +## Current state observed during diagnosis + +- `yoi-server` is not running. Legacy `target/debug/worker-runtime` PID `1820761` remains listening on `127.0.0.1:38800`; it uses the separate standalone Runtime catalog containing Workers 43, 57, 58, 59, 60, 61, 62, and 63. +- `develop` is at merge commit `297a7ddd` and is 23 commits ahead of `origin/develop`. +- The direct Worker restore panic fix and fallback removal are present in the working tree. +- The embedded Runtime store was repaired by the one-off migration above. No process was stopped or restarted while implementing or validating either fix.