# Task26 Recovery Runbook, candidate 2e0894ea (v2)

Version: v2 (2026-08-15, supersedes the v1 recovery runbook listed below)
Scope: failure handling for the Task26 same-actor live rehearsal bound to candidate `2e0894eac92bc396cc4723bf1f18ebc653b95018dd41574df435941c235da925`, wheel `af4a9d0a1ffffb6eb7551c1d6dc2b32853ca6d024332a4f8f5702bbf992f141b`, amended plan `7ace03c6dad33d2fc3ef223621cbca68a150fde8429932138e252fb8498ac582`, actor `8527916639`.

Supersession (append-only):

- Supersedes `task26-runbook-rebind-2e0894ea-st_01a0054e/task26-recovery-runbook-2e0894ea.md` sha256 `f4f2d2347b4b68a080b9c437a813accbef8c4bcc471cbf21d67ffff31e56f89b` (v1) and its golden-path sibling sha256 `1f55b8f967c6564113574d41bf1c15ce977d35c38b5687f1623ba37f235a7bdb`; the v1 pair had a G8/Section 2 lock deadlock, a wrong ledger authority path, and a generic reset interface, all fixed in v2.
- Still supersedes `.omo/evidence/dualcoach-recovery-runbook.md` sha256 `ae2f5f9046c06f8f0b42693be5aa0d9c31cada8024ae1e5ab8ba19a3cf10f6fc` and its golden-path sibling sha256 `32a379d855c6e5af978bd9886e3bf49c616c7c20f1c1d5f100c19d8adfc5eb4a` (stale candidate `19ed0e6047f0a4d7650c47a7413ee0298a243f768a63bcb8fb9cffe44d140a1a`). No prior file is modified.

Hard rule: there are no recovery shortcuts. Every recovery path ends in one of two states: return to the golden path at a named gate with new evidence, or abort per the matrix in Section 4. The plan's forbidden list (archive restore/prepopulation, direct durable-state mutation, raw Telegram listener or cursor edits, service replay, forced publication, recovery shortcuts, changes to other profiles) applies to every procedure below.

## 1. Supported commands for recovery (exhaustive; anything else is invented and forbidden)

Lifecycle and inspection CLIs (run from `/home/cube/.hermes/profiles/dualcoachtest/.venv` unless noted):

- `./bin/python -m checkin_cli.customer_admin --registry <registry> <list|show|create|activate|disable|reset|set-next-checkin|delete> ...` (lifecycle CLI; `create` is forbidden in this run; `reset`/`delete` only per R9 with explicit human approval)
- `./bin/python -m hermes_cli.nutrition_readiness --customer <key> --profile-root ... --data-root ... --json`
- `./bin/python -m hermes_cli.nutrition_review_service --customer <key> ... <request-generation|...> --json`
- `./bin/python -m hermes_cli.dualcoach_admin --profile .. provider-auth check [--json] [--allow-billable-active-probe --probe-all]` (the only existing `dualcoach_admin` subcommand; the billable probe requires prior written human authorization)
- `./bin/python -m hermes_cli.profile_cli list --profiles-root /home/cube/.hermes/profiles --json` (read-only inventory)

Candidate verifier (run from the repository root): `python3 ".omo/evidence/task26/task26-repaired-archive-successor-2e0894eac92bc396cc4723bf1f18ebc653b95018dd41574df435941c235da925/verify_candidate.py" ".omo/evidence/task26/task26-repaired-archive-successor-2e0894eac92bc396cc4723bf1f18ebc653b95018dd41574df435941c235da925/verifier-input.json"` (exit `0` and `status: PASS` means the candidate still verifies; the positional argument is the sealed verifier-input file, not the candidate root).

Retained-archive verifier (read-only; run from the repository root): `python3 .omo/evidence/task26/verify_retained_archives.py .omo/evidence/task26/retained-archive-schema-inventory.json --permission-seal .omo/evidence/task26/retained-archive-verifier-permission-seal.json`. Verifier sha256 `12b97aa72e2df86719dbb7b80b477237d590aaaea81a9f812f321a1566936528`, inventory sha256 `29db87f2855c4804dbbc69244e68d9fb9f51808f7f7759fb592ed68f5f090185`, seal sha256 `28ff72fdb7925cbdcd8e7dd7e8c058bcca42cab3377c34078164766197ed6968`; sealed PASS receipt sha256 `597a37e4912d729b3b5f6022bccff6dc12a73b024128805980f5ab1ade4fb83d`; sealed test file sha256 `1b92ac7250d1af013204d8074176b971307211eddee91606d4cdd0fcd05efd79` with 14/14 tests passing. Without `--receipt` it prints to stdout and writes nothing.

Reset controller (sealed, blocker B2 CLOSED): `python3 '.omo/evidence/task26/reset-controller-st_01a0054d/reset_controller.py' <dry-run|execute|verify> --profile '/home/cube/.hermes/profiles/dualcoachtest' --other-profile '/home/cube/.hermes/profiles/physique-coach' --archive-root '/home/cube/.hermes/profiles/dualcoachtest/data/profile-reset-archives' --contract '.omo/evidence/task26/reset-controller-st_01a0054d/schema-contract.json' --permission '.omo/evidence/task26/reset-controller-st_01a0054d/permission-receipt.json' --candidate '2e0894eac92bc396cc4723bf1f18ebc653b95018dd41574df435941c235da925' --wheel-sha256 'af4a9d0a1ffffb6eb7551c1d6dc2b32853ca6d024332a4f8f5702bbf992f141b' --plan-sha256 '7ace03c6dad33d2fc3ef223621cbca68a150fde8429932138e252fb8498ac582' --approval 'TASK26_ARCHIVE_FIRST_PROFILE_RESET_APPROVED' --run-id '<unused run id>'` (run from the repository root; line breaks display only). Pins: controller sha256 `bd051dda8666ce9e014ec79c58acd4d7b8df2c27fad01a7a78c76e46c3388baf`, contract sha256 `4128cece087f3f4eb84c3917207299fae6bd1ac9971d7e4e8676549106bb7c20`, permission sha256 `8a290b11cac5b6957c772366abe875c7f635b8a3e7956a665471ffaa90b6c495`, sealed dry-run receipt sha256 `2fc323f8b00c18f21598f717fb6a19a7e342b7786909e59adcf4bc3955bcef99`, readiness receipt sha256 `776e90301a63dce8912bf8a4110cd037f833c3a0e875a3ddd1289c0c9ee2ac9e`. Mode order is dry-run, then execute, then verify; verify is post-execute only. The pre-reset consumes run id `task26-live-reset-2e0894ea`; the cleanup reset uses a fresh unused run id. `--approval` takes the literal phrase shown, presented by the human. `--archive-root` is where the NEW archive lands (`<archive-root>/<run-id>` via atomic rename from `.pending-<run-id>`); the controller reads prior archives from both `data/rehearsal-reset-archives` and the archive root and never writes to them.

Observer CLIs (from the trusted workspace source checkout): the three `scripts/telegram_nutrition_onboarding_e2e.py` subcommands `subscribe-outbox`, `audit-tail`, `watch-deliveries`.

Handset actions (actor `8527916639`): invite link open, `/start`, consent review, onboarding answers, owner review card actions, `readiness_check` with photo upload, check-in answers, and the review card's single `Send to customer` action.

systemd: `systemctl --user <start|stop|status|is-active> hermes-gateway-dualcoachtest.service`, `journalctl --user -u hermes-gateway-dualcoachtest.service`, `systemctl --user list-timers`.

Planned but NON-EXISTENT interfaces (any runbook or habit referencing them is void): `dualcoach_admin invite prepare|expire`, `dualcoach_admin bootstrap prepare-expire`, `dualcoach_admin delivery watch|status|deadline`, `dualcoach_admin recovery audit`, `dualcoach_admin read-model`. Only `provider-auth check` exists. The old runbook's `--dry-run` invite flag is void because the command itself does not exist.

Bounded deadlines: every subscription watch window is 120 seconds and re-arms with an advanced cursor; service start to gateway-connected is 90 seconds; claim after `/start` is 120 seconds; consent-to-owner-review visibility is 300 seconds; generation to staff-review publish is 600 seconds; send decision to exactly one DM is 300 seconds; readiness/activation CLIs should exit within 60 seconds. An exceeded bound is a STOP to this runbook, never a silent retry.

## 2. Global abort criteria (any one is sufficient)

- Any digest drift on candidate, wheel, plan, config, unit file, sealed reset artifacts (controller `bd051dda`, contract `4128cece`, permission `8a290b11`), or sealed archive-verifier artifacts (verifier `12b97aa7`, inventory `29db87f2`, seal `28ff72fd`).
- Any sign the running code is not the exact candidate bytes (G5/G3-proof mismatch at any recheck).
- Any unexpected customer DM, any duplicate delivery, any implicit send, any extra Telegram message beyond the script.
- Any unauthorized write to `physique-coach` or `quarantine`, or a digest change in their pinned files.
- Loss of Telegram auth, provider auth failure that persists across one recheck, or a handset account mismatch.
- Any need for a forbidden action (state hand-edit, restore, replay, forced publication, lock hand-removal, bootstrap-ledger edit) to proceed.
- A second invite becoming necessary while the first is still claimable.
- Owner rejection that the candidate's canonical card actions cannot express.
- A bounded deadline exceeded with no matching audit/outbox event.

## 3. Procedures

R1, invite link dead or wrong bot: confirm the link came from this run's sealed preparation receipt and the bot username is `dual_coach_pilot_test_bot`. If the link was mistyped or stale, do not retry with edits; check the invite expiry deadline from the receipt. If expired unbound, follow R3. If the link is valid but the bot does not respond, go to R6.

R2, stale `gateway.lock`: pre-reset, the sealed reset disposes of the lock (`gateway.lock` is inside the contract's `approved_clear_scopes`); the only pre-reset requirement is that no process holds the lock (`flock -n <lock> true` succeeds). At authoring the lock file (sha256 `28420d4aa8fc1ee1298c9fe93a69e47aa187c0d2777ca672fb603cfbebc569cf`, stale pid `4091167`) has no holder. Post-reset, `gateway.lock` must be absent (gate G15); the controller's `verify` fails closed if it is not. If the lock exists AND a live process holds it at any point, a gateway is already running: stop it through systemd (`systemctl --user stop hermes-gateway-dualcoachtest.service`, confirm inactive) and recheck. Never remove the lock by hand, and never run `flock -u`. The unit's `ExecStartPre` `/bin/rm -f gateway.lock` is only acceptable when no live process exists, and it is not an operator step.

R3, invite expired before `/start`: the session is `PREPARED` past expiry. Canonical cleanup is `RoomBootstrapStore.expire_unbound` through a sealed harness (precedent: session `cb_rmDnfrqA6gkmjoQEdwxu0g`, receipt in `task26-fresh-canonical-invite-20260815T040558Z/expiry-cleanup-receipt.json`), then exactly one new sealed preparation. Link both receipts in the evidence directory. Never reuse the expired link, never hand-edit the bootstrap ledger, never send a bare `/start`.

R4, `/start` sent without invite payload, or double `/start`: expected candidate behavior is no claim (no bootstrap session exists for a bare payload) or a single `role_swap_blocked` outcome. Verify through the audit subscription that no claim succeeded and no `role_conflict` message was sent. If a claim DID succeed from a bare `/start`, that is a candidate defect: abort rehearsal, capture evidence.

R5, consent card stalls (no progress within 120 seconds): confirm the service is active and the audit subscription cursor is current; re-arm the subscription once with `--after-seq` advanced. Do not resend `/start`. If the card is present but buttons do nothing, capture a screenshot and the audit tail, then abort: no state edits exist for this and none may be invented. Note for later: leftover nonterminal bootstrap sessions with recovery attempts or claims block any future reset (the sealed contract's terminal-authority check), which is handled by R16 at cleanup time.

R6, service fails to start or crashes: `journalctl --user -u hermes-gateway-dualcoachtest.service -n 100 --no-pager`. If the log shows a candidate import or config error, verify the loaded bytes (golden path Section 3, step 4). A loaded-byte mismatch means the deployment is not the candidate: abort; do not reinstall from any other source. Lock-related startup failure goes to R2. `Restart=always` with `RestartSec=5` will retry; if the unit enters a failed restart loop, stop it and abort.

R7, onboarding stuck in `reconciling` or `safety_hold`: `safety_hold` is a canonical state, not a failure: use the owner card's `safety_hold` action, then `review_questions`. Stuck `reconciling` beyond 120 seconds with no audit events: capture the audit tail and the onboarding projection file digest, then abort; the reconciler owns projection writes and there is no operator repair.

R8, owner review card missing or wrong topic: expected route is review group `-1004484458766`, topic `59` (staff review) or `219` (preview), with the operator card in owner DM `8693203710`. If no card arrives within 300 seconds of consent completion, check the outbox subscription for the publication event. Missing publication event: abort with evidence. Card in a wrong topic is a routing defect: abort with evidence. Do not republish by hand.

R9, customer lifecycle CLI errors (activation/disable/reset): all lifecycle CLIs fail closed with `error`, `error_en`, `suggestion` fields and never mutate on failure. Read the suggestion; if it maps to a golden-path step order problem, correct the order and retry once. `activate` is one-use: a failed activation consumes nothing, a succeeded one cannot be repeated. `reset` and `delete` are forbidden in this run except to retire the synthetic customer AFTER the rehearsal if the human explicitly approves; they are never a recovery shortcut mid-run.

R10, provider auth failure: `provider-auth check --json` distinguishes cheap-check failure from active-probe failure. A cheap-check failure aborts the run (no provider, no rehearsal). The billable probe (`--allow-billable-active-probe --probe-all`) requires the written human authorization recorded as blocker B3; without it, use cheap-check results only and treat missing billable confirmation as an OPEN gate, not a failure by itself.

R11, generation produces no staff-review card within 600 seconds: check the audit subscription for `request-generation` acceptance and the review service CLI output. If generation was accepted but no card published, capture the outbox transcript and abort; do not request a second generation for the same check-in (the idempotency token is single-use by design).

R12, delivery anomalies: stale card (card state older than the current generation): do not act on it; use the newest card only. Duplicate click: the second click is a no-op by candidate design; verify exactly one `adaptive_plan_delivered` audit event exists. Failed send (Telegram error in the audit trail): do not resend by hand; capture the audit event and abort to the human for a decision, because a manual resend would violate the exactly-once requirement.

R13, service restart between send decision and delivery: systemd restarts the unit within 5 seconds (`Restart=always`, `RestartSec=5`). The candidate's delivery path is crash-safe: the outbox record survives and delivery completes after restart. Proof obligation: exactly one DM and exactly one `adaptive_plan_delivered` event across the restart boundary; the delivery watcher transcript must show no duplicate. If a duplicate appears, that is a candidate defect: abort with evidence.

R14, owner rejection paths: edit answers/targets through the card actions and re-approve, or use the candidate's canonical decline/regenerate card action if one is visible on the current card. If the card offers no canonical rejection path for the needed case, STOP: do not invent CLI or state edits; escalate to the human.

R15, handset or account problems: if the wrong Telegram account is used at any step, stop immediately; only actor `8527916639` is authorized. If the handset loses the DM thread, recover by reopening the bot DM from the contact, never by fresh `/start` spam; one additional `/start` is acceptable only if no claim state exists yet (verify via audit tail first).

R16, cleanup reset blocked by nonterminal bootstrap authority (blocker B6): symptoms are controller errors `customer-bootstrap ledger contains nonterminal authority` or `required current authority missing` at cleanup dry-run/execute. Meaning: after a successful lifecycle the rehearsal session is `ACTIVE` with a retained `role_claims` entry, and candidate `2e0894ea` provides no transition from `ACTIVE` to any contract terminal state (`EXPIRED`, `CANCELLED`, `FAILED`, `COMPLETED`); `transition()` allows only REGISTERING to AWAITING_CONSENT to AWAITING_ACTIVATION to ACTIVE, and no code path assigns CANCELLED or FAILED. Allowed actions: stop, keep the disabled rehearsal customer in place, retain all evidence, escalate to the plan owner for a superseding sealed contract/permission revision or a candidate-supported terminal transition. Forbidden: bootstrap-ledger edits, forced transitions, local contract edits, archive restores. The rehearsal result stands as rehearsal-complete, cleanup-blocked until the superseding artifact exists.

## 4. Rollback and abort matrix

| Trigger | Allowed actions | Forbidden | End state |
|---|---|---|---|
| Preflight gate FAIL/UNKNOWN | Fix the environment cause (mode bits, service state), rerun the gate | Weakening a gate, skipping a gate, editing sealed artifacts | Return to golden path Section 1, or abort |
| Candidate/wheel/config/unit digest drift | Recompute, identify source, restore the pinned artifact from its sealed location | Rebuilding, repinning, proceeding on drift | Abort unless drift is explained as external and pins re-verify |
| Pre-reset dry-run FAIL | Diagnose against the contract; rerun after the cause is fixed | Hand-clearing state, editing the contract or permission | Return to Section 2 Step A, or abort |
| Pre-reset execute FAIL | Controller auto-rolls back to the pre-execute state (rollback boundary: archive directory rename); capture stderr and the staging directory state | Manual cleanup of `.pending-*` while investigating; any state hand-edit | Abort; profile is back at pre-execute state by controller rollback |
| Post-reset verify FAIL or lock present (G15) | Capture verify output; do not start the service | Deleting leftover state by hand | Abort; archives intact for forensics |
| Deployment/loaded-byte mismatch | Redeploy once from the identical wheel path; recheck digests | Installing from any other source or rebuilding | Return to Section 3, or abort |
| Service start failure | R6 diagnosis; one restart via systemd | Lock hand-removal, config edits | Return to Section 4, or abort |
| Invite expired unbound | R3 canonical expiry plus one new sealed preparation | Reusing the link, ledger edits | Return to Section 5 |
| Claim/consent/onboarding stall or defect | R4/R5/R7 evidence capture | State edits, extra `/start`, free-form goal text | Abort |
| Owner rejection inexpressible | R14 stop and escalate | CLI/state workarounds | Abort pending human decision |
| Provider auth failure (cheap check) | R10 recheck once | Billable probe without authorization (B3) | Abort |
| Delivery anomaly (duplicate, implicit, none) | R12/R13 evidence capture | Manual resend, forced publication | Abort with exactly-once evidence |
| Post-rehearsal lifecycle CLI failure | R9 single guided retry | `reset`/`delete` without explicit approval | Abort pending human decision |
| Cleanup dry-run/execute FAIL (nonterminal authority, B6) | R16: stop, retain, escalate for superseding sealed contract/permission | Ledger edits, forced transitions, restore/prepopulate | Rehearsal-complete, cleanup-blocked (B6 OPEN) |
| Other-profile drift detected at any point | Stop everything; capture both profiles' digests | Any write to the other profiles | Abort; incident evidence |
| Human abort at any point | Stop service if safe, preserve evidence | Any cleanup beyond evidence preservation | Aborted by decision |

## 5. Other-profile non-touch verification

- `physique-coach` pins (from evidence, recompute at run time): `config.yaml` `e04e3a0a3c40a77a5f5876d453354763d44f9eaa6994a07d0a25484be2ba33f4`, `secrets/telegram-bot-token` `30b0e026c8b78c40a88ae276ba296807696134365ee62dce2c76a610d14e82e1`, `secrets/gateway-credentials.json` `e233d558ec022144f3f09094c6dea48adb1497c7cbb6b4c4d5175ac90814be56`. Additionally the sealed reset controller records the whole-tree digest (`435d8b83e12ede00ca1494e254e146bcdcda06210779cb08fb38b3ecb4a897d8` at dry-run time) into the reset manifest and aborts if it changes before commit.
- `quarantine` pins: `config.yaml` `9ec0b0f34940f7e10d92e96653ce6f50d22a7cc420df8c45be3f3c7d2b87b1c3`, `secrets/telegram-bot-token` `4b4a8c5a2b759d1b74a04ff8395f446a4d597ac2ff9b4f78cfa31309b8e56a52`.
- The cleanup reset passes `--other-profile /home/cube/.hermes/profiles/physique-coach` exactly as sealed. Any drift: stop, do not touch, escalate.

## 6. Secrets and redaction

Never copy into receipts or logs: bot tokens, gateway credential file contents, `.gateway-auth-key`, `auth_39664143.session*`, provider OAuth material under `dualcoach-provider-auth/`, invite nonces beyond their first 8 hex chars, and full chat contents. Redact in shared evidence as `sha256:<digest>` or `<redacted:kind>`. Evidence directories stay mode 700, files mode 600. Telegram message quotes in evidence are limited to the minimal JSON fields needed to prove the event (kind, seq, ids, timestamp).

## 7. Receipt rules during recovery

Every procedure invocation appends a receipt line to the run's evidence directory: procedure id, UTC time, trigger, commands run (exact argv), outcomes, digests of any artifact touched or produced, and the resulting decision (return-to-gate or abort). No receipt is edited after writing; corrections are new receipts that reference the earlier one.
