Files
yoi/docs/design/durable-operations.md
Hare a595af133c feat: merge durable Workdir updates from develop
# Conflicts:
#	docs/README.md
#	docs/design/durable-operations.md
#	web/workspace/deno.json
2026-09-03 02:47:50 +09:00

15 KiB

Durable operation design

Durable operation records make retries converge on the same authorized intent and preserve the facts needed to explain an externally visible result. They are not execution traces and must not mirror every Rust function or implementation step as persisted state.

This document defines the cross-domain rules for operation identity, state, checkpoints, child operations, failure evidence, and terminal disposition. Domain code may use different records where its atomicity boundary differs, but it should classify the operation before choosing a schema.

Core rule

Persist authority and non-reconstructable facts, not control flow.

A value belongs in durable operation state only when at least one of the following is true:

  • it identifies the caller's stable intent and detects conflicting reuse;
  • it freezes authority or configuration that an exact retry must continue to use;
  • it binds a preallocated or created resource to the operation;
  • it records an externally visible side effect that cannot be safely derived or repeated;
  • it records the final domain result or disposition;
  • it provides bounded evidence needed for retry, reconciliation, or audit.

A local step does not become durable merely because it occurs before or after another function call. If current authority can be reread or the step can be repeated safely, derive or repeat it instead of adding a stage.

Resilience boundary

These rules cover normal product retry and recovery boundaries: duplicate requests, returned errors, timeouts, known partial-completion outcomes, process restart from the last committed facts, and retries of provider operations with an explicit idempotency or observation contract.

They do not require the system to survive an unexpected stop at every point while Rust code is executing. A panic, abort, process kill, machine loss, or power failure may occur between an external side effect and its next durable checkpoint. Yoi does not attempt to close every such instruction-level crash window by writing a stage before and after every await, function, or provider call.

Consequently:

  • do not claim general exactly-once execution;
  • do not introduce write-ahead stages solely to model arbitrary Rust control-flow interruption;
  • rely on SQLite transaction atomicity for work inside one database transaction;
  • prefer provider idempotency, compare-and-swap, stable resource identity, and authoritative observation for external effects;
  • when an uncovered crash window cannot be reconciled automatically, retain the last committed facts and surface an unknown or attention-required disposition rather than guessing that the side effect did or did not happen.

A domain may require a stronger crash-consistency contract for a specific destructive or security-sensitive effect. That requirement must be explicit and must define the provider protocol, checkpoint ordering, replay behavior, and reconciliation evidence. It is not implied by calling a record a durable operation.

Classify the record before adding state

Not every record containing an operation_id is a state machine. Use one of the following shapes.

Atomic idempotency ledger

Use an idempotency ledger when all authoritative mutations and result recording commit in one database transaction.

The record normally contains:

  • Workspace and operation identity;
  • a fingerprint of stable caller intent;
  • the created resource or result identity;
  • the committed revision and timestamp where relevant.

It does not need pending, executing, or intermediate stages. An exact retry returns the recorded result. Reusing the same operation identity with a different fingerprint fails.

Repository secret mutation results, Workspace resource creation results, and transactionally appended domain events are examples of this shape.

Reservation

Use a reservation when an identity or exclusive right must exist before a later binding can complete.

Persist factual transitions such as:

  • the reserved resource identity;
  • the immutable request fingerprint and authority snapshot;
  • the concrete resource or assignment bound to the reservation;
  • reservation expiry or release evidence when the contract requires it.

Do not model internal dispatch, validation, construction, or callback steps as reservation states. A nullable result binding or a small reserved | created state can be sufficient when those values correspond to real authority facts.

Durable side-effect operation

Use a durable side-effect operation when work crosses a database/provider boundary and a retry needs durable intent or result evidence.

The default lifecycle is deliberately small:

pending -> completed
pending -> failed
failed  -> pending     # only when the domain explicitly permits retry

Existing code may use succeeded for the successful terminal value; new naming should prefer completed. Do not rewrite applied migrations or historical audit text only to normalize that word.

The operation should contain:

  • stable operation identity and request fingerprint;
  • immutable resolved authority needed by an exact retry;
  • preallocated resource identity where it prevents duplicate creation;
  • only the necessary irreversible checkpoints;
  • bounded failure evidence;
  • the final result and domain disposition.

pending means that the intent remains open and current authority must be reread before progress. It does not identify which Rust function should execute next. failed records the latest terminal attempt outcome; retryability is an explicit domain rule, not something inferred from the word. completed means the operation's required result and evidence are durably committed.

Parent workflow

A parent workflow coordinates domain operations but does not duplicate their lifecycle.

Persist:

  • the parent intent and fencing authority;
  • stable child operation identities;
  • the final workflow result or disposition;
  • bounded attention or decision evidence.

Read child state from the child authority. Do not copy child states, provider stages, Worker status, attachment status, or Workdir status into a second parent state machine. A parent cleanup workflow will often need only pending | completed; child failure remains on the child operation and appears in the parent as current attention metadata.

Creating or binding a child must itself be idempotent. Prefer a deterministic child operation identity or persist the child reference atomically with the parent decision so a retry cannot create siblings for one intent.

Checkpoint rules

A checkpoint records a fact that changes retry semantics. It is not a progress notification.

Add a checkpoint only when all of the following hold:

  1. A side effect may already have occurred outside the current transaction.
  2. Current authority cannot derive the fact reliably enough for safe retry, or repeating the effect is not safe under the provider contract.
  3. The retry algorithm changes after the fact is committed.
  4. Tests can exercise behavior before and after the checkpoint.

Prefer factual fields over stage names:

  • provider_deleted_at is evidence that provider deletion succeeded;
  • child_operation_id binds delegated work;
  • result_revision identifies the committed result;
  • target_ref_after records verified merge evidence.

Avoid fields such as validating, closing_session, detaching, deleting_registry, or finalizing. Those names describe code location, not durable authority. If those steps are safe to rerun or their result can be read from Worker, attachment, Workdir, repository, or provider authority, they are not checkpoints.

A checkpoint must never claim more than the authority that produced it. For example, sending a provider request is not proof that provider deletion completed, and receiving a Worker notification is not proof that a Ticket or cleanup workflow completed.

State, failure, blockers, and disposition are separate

Do not overload one enum with unrelated dimensions.

  • Operation state says whether the intent is open, completed, or has a recorded failed attempt.
  • Failure evidence records a bounded category, timestamp, and safe diagnostic detail for the latest failure.
  • Blockers and eligibility are normally derived by rereading current authority. Persist them only as audit or attention evidence, not as a substitute for live validation.
  • Disposition records what the domain decided to retain, delete, release, tombstone, abandon, or leave unknown.
  • Attention metadata explains why automated progress currently cannot continue and what authority must change.

Values such as blocked, executing, stale, dirty, retained, and deleted therefore do not all belong in one operation-state enum. Some are derived conditions, some describe transient execution, and some are domain results.

Before every retry or side effect, reread live authority and revalidate its fence. A previously recorded blocker does not prove that the operation remains blocked, and a previously unblocked operation does not retain permission after assignment, ownership, revision, or attachment authority changes.

Identity and fingerprinting

Every externally retryable operation has a stable identity in its owning Workspace or authority scope. The operation fingerprint represents stable caller intent, not generated results or mutable observations.

Include inputs whose change would mean a different requested operation. Exclude:

  • generated resource IDs when the Server allocates and persists them as the result;
  • timestamps assigned by the Server;
  • retry counters and diagnostics;
  • current provider observations that are expected to change;
  • secret bytes and credential material.

Resolved authority snapshots may be stored separately from the caller fingerprint. An exact retry uses the persisted snapshot where replay convergence requires it; a new operation resolves current authority. Unknown, foreign, or conflicting operation identity fails closed.

Transactions and external providers

Keep database work in one transaction whenever the owning authority and result live in the same database. Do not create a durable operation merely to split a transaction that can remain atomic.

When an external provider is involved:

  1. reserve stable intent and identity if retry needs them;
  2. invoke the provider with the strongest available idempotency, expected-old revision, or stable resource key;
  3. verify the provider result through authoritative response or observation;
  4. commit only the checkpoint or result evidence that changes retry behavior;
  5. on retry, reread both the operation and current domain/provider authority before acting.

Compensation is a domain operation, not an invisible finally block. If compensation has its own external side effects or retry lifecycle, give it a stable child operation identity rather than expanding the parent into a list of cleanup stages.

Workdir removal application

Workdir removal is one durable side-effect operation in the Workspace Server DB. It binds the Workspace, Workdir, owning Runtime, Repository/materialization identity, source actor, stable intent fingerprint, lifecycle, retry metadata, and bounded result. Runtime URL, provider handle, host path, credentials, and caller-selected Runtime are not operation inputs.

A durable one-pending-operation constraint plus an atomic attempt claim prevents concurrent callers from entering the provider side effect for the same Workdir; the in-process resource lock is an additional serialization layer, not the sole authority. Each active attempt persists the Server process ID and process-start marker. Recovery reclaims only an owner proven missing or replaced; a live or unobservable owner is never stolen. The reclaim transaction compare-and-set checks the exact proved owner snapshot and attempt count so stale orphan proof cannot overwrite a newer live claim.

Each attempt:

  1. resolves or revalidates the persisted same-Workspace Workdir, Runtime, Repository, and materialization identity;
  2. checks current attachments, attachment reservations, current assignment occupancy, retention/cleanup holds, and pending materialization authority; a failed Workdir-create retry must atomically return to pending before provider work and is rejected while removal is pending;
  3. retains dirty, occupied, blocked, or otherwise unknown Workdirs without detaching a Worker or forcing deletion;
  4. observes the owning Runtime/provider and calls its existing Workdir cleanup only for an eligible clean Workdir;
  5. treats only successful provider cleanup or exact working_directory_not_found as removal evidence;
  6. deletes the Backend Workdir registry row and commits the operation's completed/removed result in one SQLite transaction.

A provider error leaves the registry intact and records a bounded attention_required result with explicit retryability. Startup recovery lists pending and retryable failed operations, then executes this same path after rereading live authority. WorkdirDelete, Workspace REST removal, Runtime cleanup execution, and recovery must not maintain separate inline provider-delete paths.

The public request contains only working_directory_id plus a bounded reason. The public result contains only the Workdir ID, removed | retained | attention_required, retryability, and an optional bounded failure category. Internal operation identifiers, checkpoints, provider paths, and credentials are not public DTO fields.

Diagnostics and audit

Persist bounded error categories and identifiers needed to investigate or retry. Do not persist credentials, provider handles, raw command output, raw prompts, full session transcripts, or host paths in ordinary operation diagnostics.

Attempt counts, last-attempt timestamps, and safe provider categories may help operations, but they are telemetry and evidence rather than lifecycle authority. Logs may describe detailed execution stages; the durable record should remain centered on intent, checkpoints, result, and disposition.

Applying this rule

For a new or materially changed operation:

  1. identify the owning authority and transaction boundary;
  2. classify it as an atomic ledger, reservation, durable side-effect operation, or parent workflow;
  3. define stable identity, fingerprint, and exact-retry behavior;
  4. list external side effects and decide which are idempotent or authoritatively observable;
  5. add only checkpoints that change retry behavior;
  6. keep child operation state in the child authority;
  7. separate failure, blocker, attention, and disposition from lifecycle state;
  8. state the unsupported crash windows honestly;
  9. test fingerprint conflict, exact retry, authority revalidation, checkpoint replay, and result/disposition projection as applicable.

Existing operation schemas need not be rewritten solely for vocabulary consistency. When an operation is changed for functional reasons, use this classification to remove derived or control-flow stages rather than adding another special-case lifecycle.