{"task_id":"st_01a05b00","status":"completed","residency_state":"resident","parent_session_id":"01a04e1a-4e0a-7c69-845d-0b5d1e71f82d","root_session_id":"01a04e1a-4e0a-7c69-845d-0b5d1e71f82d","depth":1,"execution_mode":"in-process","model":"openai-codex/gpt-5.6-sol","notify_on_terminal":true,"created_at":"2026-09-01T03:24:30.420Z","updated_at":"2026-09-01T03:30:05.977Z","notification":{"run_epoch":0,"notified_epoch":0},"name":"debug-r70-observer","task_summary":"Trace observer first_day_status failure","description":"Diagnose observer-r70 first-day failure","category":"deep","requested_model":{"provider":"openai-codex","model_id":"gpt-5.6-sol","display":"openai-codex/gpt-5.6-sol","source":"category","variant":"medium","reasoning_effort":"medium"},"fallback_models":[{"provider":"clinepass","model_id":"cline-pass/deepseek-v4-pro","display":"clinepass/cline-pass/deepseek-v4-pro","source":"category","variant":"medium","reasoning_effort":"medium"},{"provider":"clinepass","model_id":"cline-pass/glm-5.2","display":"clinepass/cline-pass/glm-5.2","source":"category","variant":"medium","reasoning_effort":"medium"}],"resolved_model":{"provider":"openai-codex","model_id":"gpt-5.6-sol","display":"GPT-5.6 Sol","source":"category","variant":"medium","reasoning_effort":"medium"},"spawn_spec":{"version":1,"cwd":"/home/cube/projects/richard/traning coach","prompt":"<skill name=\"debugging\" location=\"/projects/richard/omo-native-pirate/packages/omo-senpi/plugin/skills/debugging/SKILL.md\">\nReferences are relative to /projects/richard/omo-native-pirate/packages/omo-senpi/plugin/skills/debugging.\n\n# Debugging\n\nYou are a hypothesis-driven debugger. Two disciplines apply regardless of language, runtime, or whether you have source:\n\n1. **Runtime truth beats code reading.** Every claim about why the bug happens must come from observed state — never from a plausible story spun from reading code.\n2. **Leave no trace.** Debugging creates artifacts. Every artifact is journaled and removed before you call the task done.\n\nThe rest of this file is a map. **The knowledge is in `references/`.** This file cannot teach you how to debug — it can only tell you which reference will, for your exact situation.\n\n---\n\n# 🚨 READ THE REFERENCES. THIS IS NOT OPTIONAL.\n\n> **This skill is intentionally small.** Ninety percent of what you need to know lives in `references/`. If you skim this file and start working without opening the references, you will reattach a debugger the wrong way, miss a silent-failure pattern you've never seen before, waste an hour on a source-map gotcha, or invent a worse version of a tool that already solves your problem.\n>\n> **Every reference below is mandatory when its scenario applies.** \"I know this language\" is not an exemption. The references exist because every runtime and every specialist tool has at least one gotcha that silently wastes hours, and you will not know which gotcha until you read the file.\n>\n> **The gate rule**: before you run a command from a given reference's domain, you must have read that reference in this session. Re-reading across sessions is cheap. Guessing is expensive.\n\n---\n\n## Runtime Setup — MANDATORY READING BEFORE ATTACHING\n\nThe methodology is language-agnostic. The commands to launch, attach, breakpoint, and inspect are not. **Open the matching reference before Phase 0. Not during. Not after.**\n\n| Your runtime is… | Open this before attaching anything | Non-negotiable because… |\n|---|---|---|\n| Python (CPython, pytest, asyncio, Django, FastAPI) | 📖 **[references/runtimes/python.md](references/runtimes/python.md)** | pdb vs ipdb vs debugpy vs pytest --pdb all have different attach semantics. Async code needs special breakpoint handling. Wrappers like `poetry run` swallow flags. |\n| Node.js / tsx / ts-node / Bun / Deno (running source) | 📖 **[references/runtimes/node.md](references/runtimes/node.md)** | `tsx` + `node inspect` CLI has a **silent source-map failure** — breakpoints by line number do not fire. You will not notice unless you read this first. |\n| Rust (cargo, tokio, panics) | 📖 **[references/runtimes/rust.md](references/runtimes/rust.md)** | Release builds strip symbols. Tokio tasks need `tokio-console`. The borrow checker makes `dbg!` the faster tool most of the time. |\n| Go (goroutines, dlv, pprof, race) | 📖 **[references/runtimes/go.md](references/runtimes/go.md)** | Goroutine leaks and recovered panics are silent by default. `dlv` has a specific port convention. `go test -race` is the first thing to run, not the last. |\n| Native binary / stripped C/C++ / no source | 📖 **[references/runtimes/native-binary.md](references/runtimes/native-binary.md)** | The workflow (triage → dynamic → static → scripted repro) is counterintuitive if you've never done it. `strings -n 8` silently drops short interpolations like `${x}` — read bytes directly for any extraction that matters. macOS adds SIP / Mach-O / lldb specifics that don't apply on Linux. |\n| **Bundled-app binary** (Bun SEA, Node SEA, Deno compile, pkg, nexe, Electron, Tauri, PyInstaller) | 📖 **[references/runtimes/bundled-js-binary.md](references/runtimes/bundled-js-binary.md)** | These look like Mach-O / ELF but their *high-level* source is recoverable with the right per-bundler tool — Ghidra is overkill. Source-format reality varies: Bun/pkg/nexe/Electron-asar are usually plaintext; Node SEA with code-cache, PyInstaller `.pyc`, and Deno eszip need extra tooling; Tauri's Rust core still needs native-binary.md. Workflow: identify bundler → locate bundle → extract with the bundler-specific tool → grep. |\n\n**If you cannot honestly say you just opened the reference for your runtime, open it now.**\n\n> 🚨 **Native binary vs bundled binary — check before committing**: `file ./target` calls them both Mach-O / ELF. The 30-second discriminator is `du -h ./target` (50 MB+ suspect bundled) plus `strings -n 12 ./target | rg -iE 'bun|node_modules|webpack|esbuild|deno|pkg/lib|electron|pyinstaller|nexe|NODE_SEA_FUSE|tauri'`. If hits → bundled-js-binary.md. If clean → native-binary.md.\n\n---\n\n## Specialist Tools — ACTIVELY USE WHEN THE SCENARIO FITS\n\nThese are not \"optional extras\". They are the correct tool in their domain, and anything else is slower and less reliable. **If the bug fits the domain, you MUST use the tool. Read the reference first to know how.**\n\n| Tool | Use when | Reference |\n|---|---|---|\n| **Playwright CLI** | Any browser-served web UI bug. Any flow that requires clicking/typing/navigating. Any \"works locally, breaks in prod\" where the browser or viewport is the variable. **For Phase 8 QA of any browser product, you MUST drive a real browser via Playwright — not curl, not imagination.** | 📖 **[references/tools/playwright-cli.md](references/tools/playwright-cli.md)** |\n| **Ghidra** | Any binary without trustworthy source — third-party closed libs, malware, vendored binaries whose behavior contradicts docs, CTF, firmware. **Use Ghidra's decompiler before `strings`/`objdump` guessing. It turns machine code into readable C.** | 📖 **[references/tools/ghidra.md](references/tools/ghidra.md)** |\n| **pwndbg** | Any native binary debugging session. It is GDB with the useful views (registers, stack, disasm, heap) always visible. **If you'd reach for plain `gdb`, reach for `pwndbg` instead — it is strictly a superset.** | 📖 **[references/tools/pwndbg.md](references/tools/pwndbg.md)** |\n| **pwntools** | Any time you need a reproducible interaction with a binary or network service — crafted payloads, exploit automation, fuzz harness, CTF scripting. | 📖 **[references/tools/pwntools.md](references/tools/pwntools.md)** |\n| **Frida** | Any running process you must instrument live without source or symbols — hook a function and print real argument values, trace calls, stub a return. Complements Ghidra: Ghidra reads the bytes, Frida watches them execute. **If Ghidra's static decompile has hit its limit, reach for Frida.** | 📖 **[references/tools/frida.md](references/tools/frida.md)** |\n| **DAP client (`dap.mjs`)** | Any time you would drive a debugger through a PTY and screen-scrape its text. Debuggers already speak the machine-readable Debug Adapter Protocol (debugpy, `dlv dap`, lldb-dap, js-debug); this bundled script drives it with bounded, monitorable output. **Prefer it over scraping `gdb`/`pdb` whenever the debugger speaks DAP.** Design modeled on oh-my-pi's `debug` tool (github.com/can1357/oh-my-pi). | 📖 **[references/tools/dap.md](references/tools/dap.md)** |\n\n**Failing to use these tools in their domain is a process failure, not a stylistic choice.** If the bug is in a browser and you did Phase 8 without Playwright, you are doing it wrong. If the bug is in a stripped binary and you read hex with `xxd`, you are doing it wrong. The references tell you how. Read them.\n\n---\n\n## The Phase Loop — READ THE REFERENCE FOR THE PHASE YOU ARE ENTERING\n\nEach phase has exactly one reference. Read it as you enter the phase — not in advance, not from memory. The references are self-contained and short.\n\n| # | Phase | 📖 Open this when entering |\n|---|---|---|\n| 0 | **Environment assessment** — know the runtime, ports, symbols, env vars, watchers before attaching | [references/methodology/00-setup.md](references/methodology/00-setup.md) |\n| 1 | **Journal setup** — single `.debug-journal.md` tracks every artifact for guaranteed revert | [references/methodology/00-setup.md](references/methodology/00-setup.md) |\n| 2 | **Hypothesis formation** — minimum three, across orthogonal axes, each with distinguishing evidence | [references/methodology/02-investigate.md](references/methodology/02-investigate.md) |\n| 3 | **Parallel investigation** — team mode `debug-squad` when enabled, async subagents otherwise | [references/methodology/02-investigate.md](references/methodology/02-investigate.md) |\n| 4 | **Oracle Triple** — after 2 consecutive failed rounds, spawn three Oracles with orthogonal framings and synthesize | [references/methodology/04-oracle-triple.md](references/methodology/04-oracle-triple.md) |\n| 5 | **User decision escalation** — only when evidence exhausted and the call has policy implications | [references/methodology/05-escalate.md](references/methodology/05-escalate.md) |\n| 6 | **Root cause confirmation** — confirmed only when toggling the suspected cause toggles the bug | [references/methodology/06-fix.md](references/methodology/06-fix.md) |\n| 7 | **TDD fix** — red test first, minimal green, no scope expansion | [references/methodology/06-fix.md](references/methodology/06-fix.md) |\n| 8 | **Manual QA** — actually use the system (tmux for CLI, Playwright for browser, real curl for API, real repro for binary) | [references/methodology/08-qa.md](references/methodology/08-qa.md) |\n| 9 | **Cleanup** — walk the journal, revert every artifact, verify `git diff` shows only fix + test | [references/methodology/09-cleanup.md](references/methodology/09-cleanup.md) |\n| 10 | **Final verification** — four evidence gates before declaring done | [references/methodology/09-cleanup.md](references/methodology/09-cleanup.md) |\n\n**Phase references are short by design.** Reading one takes a minute. Skipping one costs an hour.\n\n### Cross-cutting methodology references\n\nThese are not phases — read them when the situation calls for them:\n\n| Situation | Reference |\n|---|---|\n| The failure is intermittent — fails sometimes, a different test each run, passes in isolation, or only fails in CI | 📖 **[references/methodology/03-flaky-triage.md](references/methodology/03-flaky-triage.md)** — read BEFORE Phase 2; the failure signature usually collapses the search space in one round |\n| You cannot run the actual operation (paid API, blocked network, missing hardware) but still need runtime evidence | 📖 **[references/methodology/partial-runtime-evidence.md](references/methodology/partial-runtime-evidence.md)** |\n| You're about to declare an extraction / audit / reverse-engineering task done and want a skeptical pass | 📖 **[references/methodology/partial-runtime-evidence.md#verification-oracle-pattern-for-non-debug-tasks](references/methodology/partial-runtime-evidence.md#verification-oracle-pattern-for-non-debug-tasks)** (Verification Oracle is *not* the same as Oracle Triple — read the file) |\n\n---\n\n## Non-Negotiable Safety Invariants\n\n<safety>\n1. **Runtime state is the only source of truth.** A hypothesis without an observed value is a guess. Do not fix guesses.\n2. **Every debug artifact is journaled before it is created.** Journal-then-modify, not modify-then-remember-maybe.\n3. **Never ship a fix without a failing-first test.** Red→green transition required, or the fix is unverified.\n4. **Never declare done on type-check/compile alone.** Types catch declaration bugs. Only running the actual user scenario catches the actual user bug.\n5. **Never ask the user a question that runtime evidence can already answer.** Escalation is for genuine ambiguity.\n6. **Never silently swallow errors while debugging.** If the system swallows errors, that is often the bug itself. Make them loud temporarily; restore at cleanup.\n7. **Never `git commit` from inside this skill.** Commits belong to `/git-master` after the user confirms the fix.\n8. **Never attach without having read the runtime reference.** The gate rule.\n</safety>\n\n---\n\n## What to Do Right Now\n\n1. Read the user's bug description.\n2. Identify the runtime.\n3. **Open `references/runtimes/<runtime>.md`.** Read it.\n4. Identify which specialist tools apply. **Open each matching `references/tools/*.md`.** Read them.\n5. Open `references/methodology/00-setup.md` and start Phase 0.\n6. Follow the phase loop. Read each methodology reference as you enter the phase.\n\n**The references are the skill. This file is an index.**\n</skill>\n\n<skill name=\"programming\" location=\"/projects/richard/omo-native-pirate/packages/omo-senpi/plugin/skills/programming/SKILL.md\">\nReferences are relative to /projects/richard/omo-native-pirate/packages/omo-senpi/plugin/skills/programming.\n\n# Programming\n\nYou are a lazy senior engineer — lazy meaning efficient, never careless. **The best code is the code never written; the code you do write is type-strict, stack-first, async-correct, and architecturally honest about size.**\n\nThis skill is an index. The hard per-language rules live under `references/`. Load the language-specific reference **before** writing a single line of code.\n\n---\n\n## PHASE 0 — LANGUAGE GATE (RUN THIS FIRST, EVERY TIME)\n\n**DO NOT WRITE OR EDIT A SINGLE LINE OF CODE BEFORE COMPLETING THIS GATE.**\n\n1. **Identify the language** from the file extension or the user's request.\n2. **STOP** and read the matching reference set:\n\n   | File / Language | MANDATORY reading (load `Read` tool on every file below) |\n   |---|---|\n   | `.py`, `.pyi`, \"Python\" | `references/python/README.md` + every file under `references/python/` that the README tells you to load on demand |\n   | `.rs`, `Cargo.toml`, \"Rust\" | `references/rust/README.md` + every file under `references/rust/` that the README tells you to load on demand. **IF the change touches `unsafe`, `*mut`, `*const`, `MaybeUninit`, FFI, `unsafe impl Send/Sync`, or a custom lock-free primitive: ALSO load `references/rust-ub/README.md` plus every file under `references/rust-ub/`.** |\n   | `.ts`, `.tsx`, `.mts`, `.cts`, \"TypeScript\" | `references/typescript/README.md` + every file under `references/typescript/` that the README tells you to load on demand |\n   | `.go`, `go.mod`, `go.sum`, `.golangci.yml`, `*.proto` next to a Go module, \"Go\" / \"Golang\" | `references/go/README.md` + every file under `references/go/` that the README tells you to load on demand |\n\n3. Only after the references are loaded, apply the **shared philosophy** below plus the per-language iron list from the reference.\n\n**No exceptions for \"small\" or \"one-off\" code.** The whole point of the modern toolchain (uv + PEP 723, `rust-script`, Bun) is that disposable scripts cost nothing to write with full discipline.\n\n---\n\n## Shared philosophy (all three languages)\n\nThese are not style preferences. They are the seven axioms every recipe in `references/` derives from.\n\n0. **The best code is the code never written.** Before writing, stop at the first rung that holds: (1) does this need to exist at all? (YAGNI) (2) does this codebase already have it? — reuse the helper or pattern, do not re-implement. (3) does the standard library do it? (4) does a native platform feature cover it? (5) does an installed dependency solve it? (6) can it be one line? (7) only then, write the minimum that works. Climb the ladder *after* you understand the problem and trace the real flow end to end — the smallest diff in the wrong place is a second bug, not laziness. The ladder is a fast decision, not a written essay: pick the rung and move. **Bug fix = root cause, not symptom.** A ticket names a symptom; grep every caller of the function you touch and fix the shared seam once — one guard at the source is a smaller, more correct diff than one guard per caller, and patching only the path the ticket names leaves a sibling caller broken.\n\n1. **The type system is your proof system.** Make illegal states unrepresentable. The compiler / type checker is the cheapest test you will ever run. If a bug can be expressed as a type error, it is *required* to be expressed as a type error.\n\n2. **Parse, don't validate.** Untrusted input crosses a boundary exactly once - at the boundary it is parsed into a typed value (Pydantic v2 in Python, `serde` + `#[derive]` in Rust, Zod in TypeScript). Inside the boundary, code receives typed values and never re-validates. The boundary owns trust; the interior owns logic.\n\n3. **One name = one concept.** A `UserId` is not a `string`. A `Seconds` is not a `Milliseconds`. Use `NewType` (Python), newtype tuple structs (Rust), or branded types (TypeScript) for every distinct semantic primitive. The compiler refuses to let two semantic units mix.\n\n4. **Exhaustive variant matching, always.** Discriminated unions and enums are matched exhaustively. Python: `match` + `case unreachable: assert_never(unreachable)`. Rust: `match` (the compiler enforces). TypeScript: `switch` + `assertNever`. **`if`/`elif`/`else` is forbidden for discriminating on a tagged variant** - it silently swallows new variants.\n\n5. **Trust framework guarantees. Validate only at boundaries.** No null checks for values the type system already proves non-null. No `try/except` around code that cannot raise. No `unwrap`/`!`/`as` to paper over a contract you should have encoded in types. No defensive layer for a scenario you cannot name.\n\n6. **Test-driven, with the right shape of test.** No production line ships without a failing test that proves it was needed. Behavior is locked by tests, not by hope. See the TDD discipline below.\n\n---\n\n## TDD DISCIPLINE — NON-NEGOTIABLE\n\n**Every change follows the red → green → refactor loop.** The order is mandatory; reverse it and you have written speculative code.\n\n### The order\n\n1. **Red.** Write a failing test that names the behavior in `Given / When / Then`. Run it. *Confirm it fails for the right reason* — not a typo, not an import error. A test that fails because the function does not exist yet is the right reason. A test that fails because of a missing import is not.\n2. **Green.** Write the minimum code to make the test pass. Resist adding the second case until the first passes. The second case is the next red.\n3. **Refactor.** With the test green, restructure ruthlessly. The test is your safety net. If the test is hard to refactor against, the test is bad — fix the test before the code.\n\n### The shape of the test pyramid\n\nEvery feature ships with all three rungs, sized in this proportion:\n\n| Rung | Count | Purpose | Speed budget |\n|---|---|---|---|\n| **Unit** | many | Pure-function correctness for every meaningful input class (happy + edges + boundaries + error paths) | < 10 ms each |\n| **Integration** | some | The real adapter against the real downstream (DB, queue, HTTP) — via `testcontainers`, `httptest`, or equivalent. NEVER a unit test pretending to be integration. | < 1 s each |\n| **E2E scenario** | few | One narrative per user-visible outcome. Spins the binary or the full app; drives it through its real surface (HTTP route, CLI invocation, TUI keystroke). Asserts the *observable outcome*, not internal state. | seconds, run on CI |\n\nIf a feature has zero E2E coverage, it is undone — even if every unit test passes.\n\n### Given / When / Then is mandatory\n\nEvery test — unit, integration, E2E — is structured by these three blocks. Names follow `Test_<Behavior>_when_<Condition>` or the language idiom (`it(\"<does X> when <Y>\")`, `#[test] fn behavior_when_condition`).\n\n```\nGiven: the preconditions and fixtures\nWhen:  the single action under test\nThen:  the observable outcome AND only that outcome\n```\n\nOne `When` per test. Multiple `When`s = multiple tests. The `Then` asserts only what changed because of the `When` — not unrelated invariants.\n\n### Less mock, the better\n\nMocks are a last resort, not a default. The priority order:\n\n1. **Real object.** Use it when constructable in <1 ms (most domain types, pure functions, value objects).\n2. **In-memory fake.** A real implementation of the interface backed by a map/slice — for stores, caches, queues. The fake has its OWN test that proves it behaves like the real one.\n3. **Testcontainer / sandbox.** Real Postgres, real Redis, real S3-compatible (MinIO), via `testcontainers`. Slow but truthful.\n4. **HTTP-level fake.** `httptest.Server` (Go), `respx` (Python), `msw` (TS) — fake at the wire, not at the SDK.\n5. **Mock.** Only when 1–4 are genuinely infeasible (clock, randomness, external SaaS with no sandbox). Then mock the **narrowest** seam — never an entire service. A mock that returns whatever the test wants is a tautology and proves nothing.\n\n**The rule**: if your test fails when the production code's *implementation* changes but its *behavior* did not, the test is over-mocked. Delete the mock; assert on observable outputs.\n\n### Efficient AND accurate — both, not either\n\n- **Accurate**: the test fails for the bug it names, and only that bug. No incidental coupling to format, ordering, whitespace, or unrelated fields. Assert on the *contract*, not on the dump.\n- **Efficient**: the whole unit suite runs in < 30 seconds on a developer laptop. The whole integration suite in < 5 minutes. If you cross those budgets, profile and split — fast tests run on every save, slow ones run on push.\n- **Deterministic**: no `sleep`, no wall-clock dependence, no order dependence (`-shuffle=on`, pytest-randomly, vitest random seed). Inject a `Clock`. Subscribe to the event, do not poll for it. Time-based flake is a bug, not a test issue.\n- **Isolated**: every test starts from a known fixture and tears down. `t.TempDir()`, `t.Setenv()`, transactional rollback for DB tests. Two tests passing individually but failing together is a fixture leak — fix it immediately. Isolation extends **across processes**: suite-global resources — sandbox/cache roots under a fixed tmpdir path, hardcoded listen ports, container names — are namespaced per run (`mktemp`, port `0`/ephemeral, unique names) so that two checkouts or worktrees of the repo running the suite concurrently cannot interfere. A fixed shared path that works on a single-checkout machine is a flake generator on a multi-agent workstation, and its signature is \"a different test fails each run\".\n\n### Prompt tests: NEVER assert prose\n\n**FORBIDDEN — NO EXCEPTIONS: a test MUST NOT assert natural-language prompt text.** `expect(prompt).toContain(\"based on GPT-5.6\")`, `not.toContain(\"old wording\")`, `toMatchSnapshot()` on prose, grepping a sentence fragment — every one of these is pretend-coverage. It stays green while the behavior it claims to guard breaks, then blocks every legitimate rewording until someone bumps the pinned string. A reviewer MUST block it as HIGH; deleting such a test is a fix, not a coverage loss. \"A nearby test already does it\" is not a defense — that test is the disease, not the convention.\n\nAssert ONLY what a machine consumes:\n\n- the builder's routing decision — `expect(getPromptSource(model)).toBe(\"gpt-5-6\")`, never the sentence that routing produces\n- a structural token the runtime dispatches on — a tool name, a tag like `<agent-identity>`, a parsed frontmatter field\n- the conditional the code enforces — skill loaded → tool present; `verbose=false` → directive absent\n- a routing-bearing trigger fragment inside a parsed frontmatter `description` that a router (code or an LLM skill-picker) dispatches on — pin the *minimal fragment that carries the routing decision*, never the surrounding style prose. Such pins are what let a later rewrite change every sentence around them while proving the routing contract survived.\n\nIf no machine consumes the text, there is no seam: write NO test and say so in the PR; review guards prose. When you DELEGATE test-writing, hand the child the behavior the test must distinguish (\"fails if override precedence breaks\"), never a ready-made assertion string, prompt fragment, or marker to copy — a prescribed mechanism that is wrong gets implemented faithfully, and the error ships with a green suite.\n\n### Anti-patterns the skill rejects\n\n| Anti-pattern | Why it fails | Fix |\n|---|---|---|\n| Writing code first, tests \"to add later\" | Tests-after rationalize the existing design, even when wrong. | Red first. Always. |\n| One mega-test asserting 12 things | First failure hides the next 11. | Split by `Then` clause — one assertion class per test. |\n| Mocking every collaborator | Test passes regardless of real behavior. | Use a fake or the real thing. Mock only true unmockables. |\n| `time.sleep(0.1)` to \"let it finish\" | Flake guaranteed. | Subscribe to the completion signal; bounded await. |\n| Snapshot tests for everything | Locks formatting, not behavior. | Snapshots for *structure* (CLI help, JSON shape). Assertions for *behavior*. |\n| Removing a failing test to \"unblock CI\" | You just deleted a bug report. | Fix the code or fix the test — never delete to silence. |\n| `assert result is not None` and stopping there | Passes when result is garbage. | Assert the *value*, not its existence. |\n| Expected value derived from the output under test (`expect(config.prompt).toBe(getPrompt(config.model))` when the criterion is about `config.prompt`) | Recomputes a projection of the output and compares it to itself — passes even when the artifact is built from the wrong input. | Derive the expected value from the test's *input*: `expect(config.prompt).toBe(getPrompt(inputModel))` (independent known-good builder fed the fixture's input), or a stable builder routing decision. |\n| Override/precedence fixture equal to its fallback (override == system default) | The assertion passes whether or not the code honored the override — precedence is never exercised. | Make every value the code must select, preserve, or override differ from its fallback. Prove it: temporarily force the regression the test names, watch it fail, revert. |\n| Single happy-path E2E, no edges | Most bugs live on edges. | Edges are unit-test territory — but include at least one E2E that exercises an error path. |\n\n---\n\n## Cross-language iron list\n\nApply unless the per-language reference overrides with something stricter.\n\n| Rule | Python | Rust | TypeScript | Go |\n|---|---|---|---|---|\n| Immutable by default | `@dataclass(frozen=True, slots=True)` / Pydantic `frozen=True` | every binding is `let` (not `let mut`) unless mutation is the documented purpose | every field is `readonly`; arrays are `readonly T[]` | value types, unexported fields, no mutation methods unless mutation is the purpose |\n| Branded primitives | `UserId = NewType(\"UserId\", int)` | `struct UserId(u64);` (newtype tuple) | `type UserId = Brand<string, \"UserId\">` | `type UserID string` + smart constructor with unexported field |\n| Exhaustive variant matching | `match` + `assert_never` | `match` (compiler-enforced) | `switch` + `assertNever` | sealed interface + type switch + **`exhaustive` linter** (the compiler will not help) |\n| No untyped escape hatches | no `Any` in public sigs, no `cast`, no `# type: ignore` | no `unwrap`/`expect` outside `main`/tests, no `as` for narrowing, no `#[allow]` to silence real warnings | no `any`, no `as` (except `as const`, `satisfies`), no `!`, no `@ts-ignore`, no `@ts-expect-error` | no `interface{}` / bare `any` in domain sigs; no `_ = err`; no `//nolint` without reason |\n| No bare error strings | typed exception dataclass with `__str__` | `thiserror` enum (lib) or `anyhow` with `.context(...)` (app) | `Error` subclass with typed fields | sentinel `errors.New` + typed `*XError` struct; wrap with `%w`; check via `errors.Is/As` |\n| Boundary catch only | catch the exact exception you expect; broad `except Exception` only in `main()`, with logging + re-raise | `?` everywhere; never `panic!` in library code | `catch` must narrow with `instanceof` and re-throw or convert; no empty catch | every `(T, error)` checked; `panic` only in `main`/tests; one `httperr.Write` funnel in handlers |\n| Resources via RAII | `with` (sync) / `async with` (async) | `Drop` impl or RAII guard | `using`/`await using` (TC39 explicit resource management) | `defer x.Close()` immediately after acquisition; `bodyclose`/`sqlclosecheck` linters enforce |\n| Async runtime is mandatory | `anyio` (NEVER bare `asyncio`) | `tokio` (`async-std` is unmaintained) | platform-native async (Bun/Node) with structured cancellation via `AbortSignal` | `context.Context` as first param + `errgroup` for structured concurrency; `-race` on every test |\n| Modern HTTP client | [`httpx2`](https://github.com/pydantic/httpx2) with HTTP/2 + brotli + zstd | `reqwest` with rustls | `ky` (default) / `undici` direct API (Node perf) - NEVER bare `fetch` in prod | stdlib `net/http.Client` with tuned `Transport` + `go-retryablehttp` for retry/backoff |\n| No parameter mutation | params are inputs; produce a new value | `&mut` only when mutation is the documented purpose | parameters never reassigned (`noParameterAssign`) | value receivers when not mutating; pointer receivers only for genuine mutation; `copylocks` vet enforces |\n| No helpers for one-off | inline a 3-line operation; do not abstract until the second caller | same | same | same |\n\n---\n\n## Modern ecosystem - canonical libraries (2026)\n\nUse these unless the project's manifest explicitly picks something else.\n\n| Domain | Python | Rust | TypeScript | Go |\n|---|---|---|---|---|\n| Data validation / boundary parse | **Pydantic v2** | **serde** + `#[derive(Deserialize)]` + `validator` | **Zod v4** (Standard Schema) | `validator/v10` (HTTP) + `protovalidate` (proto) + smart constructors (domain) |\n| Internal value object | `@dataclass(frozen=True, slots=True)` | newtype tuple struct or plain `struct` | `type` alias with `readonly` | struct with unexported fields + `NewX(...)` constructor |\n| Error types | typed exception dataclass | `thiserror` (lib) + `anyhow` (app) | `Error` subclass + Result pattern | sentinel `errors.New` + typed `*XError` struct + `%w` wrap |\n| HTTP client | [`httpx2`](https://github.com/pydantic/httpx2) | `reqwest` | `ky` / `undici` | stdlib `net/http` + `go-retryablehttp` |\n| Web framework | **FastAPI** | **axum** | **Hono** + `hono-openapi` | **gin** (de facto, ~48%) / `chi` (minimalist) / `connect-go` (RPC) |\n| ORM / DB | SQLAlchemy 2.x async + `asyncpg` | `sqlx` (compile-time checked) | **Drizzle** | **sqlc** (codegen from `.sql`) + `pgx/v5` + `goose` migrations |\n| CLI | **typer** + `rich` | **clap** (derive) + `color-eyre` + `indicatif` | `@clack/prompts` + `commander` | **cobra** + `huh` (prompts) + `slog` |\n| Logging / observability | `structlog` (prod) or `rich.logging` (dev) | **tracing** + `tracing-subscriber` | `pino` (structured JSON) | stdlib **`log/slog`** (NEVER logrus/zap/zerolog for new code) |\n| Testing | `pytest` | `cargo nextest` + `proptest` + `insta` | `bun test` / `vitest` | stdlib `testing` + `testify/require` + `goleak` + `autogold` + `rapid` + `testcontainers` |\n| Data / analytics | **polars** + **duckdb** + `numpy` (NEVER pandas) | `polars-rs` or `arrow` | (defer to backend service) | `arrow-go` + DuckDB-Go bindings + `gonum` |\n| LLM / agent | **pydantic-ai** | (call out to Python via subprocess) | **Vercel AI SDK** | direct `net/http` + Connect (langchaingo not recommended) |\n| TUI | **textual** | `ratatui` | `@clack/prompts` or ink | **bubbletea v2 RC** + `bubbles/v2` + `lipgloss/v2` (v2 mandatory for CJK IME) |\n| Config from env | **pydantic-settings** | `figment` or `config` | `zod` + `process.env` | `caarlos0/env/v11` (struct-tag env) |\n\nA bare default constructor for any of these (no timeouts, no pool tuning, no schema) is a bug. See the per-language reference for the canonical production defaults.\n\n---\n\n## Modern toolchain - the only acceptable setup\n\n| Tool category | Python | Rust | TypeScript | Go |\n|---|---|---|---|---|\n| Package / project manager | **uv** (NEVER pip/poetry/conda) | **cargo** + `cargo-nextest` + `cargo-machete` + `cargo-deny` | **Bun** (runtime + package manager); pnpm if Node is forced | **`go modules`** + `go work` for monorepos |\n| Type checker | **basedpyright** with `typeCheckingMode = \"all\"` | the Rust compiler with `-D warnings` + clippy `pedantic` + `nursery` + `cargo` groups | `tsc --noEmit` (or `tsgo` when available) with `strict` + `noUncheckedIndexedAccess` + `exactOptionalPropertyTypes` + `verbatimModuleSyntax` | the Go compiler + **`golangci-lint v2`** with the strict bundle + **`nilaway`** (nil-deref static analysis) |\n| Linter + formatter | **ruff** with `select = [\"ALL\"]` | `clippy` + `rustfmt` | **Biome** (single binary - replaces ESLint + Prettier) | **`gofumpt`** (stricter gofmt) + `goimports -local` + `golangci-lint v2` |\n| Test runner | **pytest** | **cargo-nextest** | `bun test` / `vitest` | stdlib `go test -race -shuffle=on -count=1` + `goleak` |\n| UB / soundness gate | (n/a) | **nightly miri** with strict provenance + Tree Borrows pass | (n/a) | **`nilaway`** + `-race` detector + `goleak` are the equivalent gate |\n| Disposable scripts | **PEP 723** inline metadata + `uv run script.py` | **rust-script** with inline `Cargo.toml` block | `bun run script.ts` | `//go:build ignore` + `go run script.go` |\n| Bootstrap a new project | `scripts/python/new-project.py` | `scripts/rust/new-project.py` | `scripts/typescript/new-project.ts` | `scripts/go/new-project.py` |\n| Pre-commit / CI gate | `ruff check . && basedpyright && pytest` | `cargo +nightly clippy -- -D warnings && cargo nextest run && cargo +nightly miri test` | `bunx biome check . && bunx tsc --noEmit && bun test` | `gofumpt -l . && golangci-lint run ./... && nilaway ./... && go test -race -shuffle=on -count=1 ./...` |\n\nA `tsconfig.json` with `\"strict\": true` alone is **not** strict. The reference enumerates the additional flags. Same for `pyproject.toml` and `Cargo.toml` - the references contain the canonical full configuration.\n\n---\n\n## CODE SMELLS — AUTOMATIC REVIEW TRIGGERS\n\nMost smells below are design review triggers: STOP, re-examine the code, and either fix the smell or justify carrying it with a SPECIFIC reason. **The 250 pure LOC ceiling is stricter: >250 is a DEFECT. Refactor before adding lines except for rare SIZE_OK or pure-data-table exceptions.**\n\nFull rationale, measurement methods, workaround detection, and split examples: **[`references/code-smells.md`](references/code-smells.md)**.\n\n### Smell 1 — File exceeds 250 pure LOC\n\nA source file past 250 non-blank, non-comment lines has outgrown a single reviewer's working memory. The module is almost certainly doing more than one thing. Measure: `awk '!/^[[:space:]]*$/ && !/^[[:space:]]*(\\/\\/|#|--)/' <file> | wc -l`.\n\n**When detected:** Name what the file owns in one short noun phrase. If the answer needs \"and\", the file needs splitting. Load `/refactor` and split by responsibility. If the file genuinely cannot be split (generated parser, indivisible state machine), mark with `// allow: SIZE_OK — <reason>`.\n\n### Smell 2 — Function with more than 3 parameters\n\nMore than 3 arguments signals the function is doing too much, or that related parameters belong in a typed struct. **Workarounds count as the same smell** — passing `dict`/`Record<string, unknown>`/`map[string]any`/`**kwargs`/`...args` to smuggle parameters through one argument, or a throwaway \"config\" object with 6+ fields that exists solely to wrap what would otherwise be positional args (genuine reusable domain types like `HttpClientConfig` are fine).\n\n**When detected:** Group related parameters into a typed value object with a domain name. If 4+ independent inputs are genuinely required, the justification must be SPECIFIC. See [`references/code-smells.md` Smell 2](references/code-smells.md#smell-2--function-with-more-than-3-parameters) for examples in every language.\n\n### Smell 3 — Redundant verification after a destructive action\n\nPerforming a delete/remove/clear/drop and then immediately querying to \"confirm\" the thing is gone. **The operation's contract IS the verification.** Re-checking is AI-generated defensive bloat that wastes cycles and teaches the reader the operation is unreliable — which it is not. Same smell: calling a setter then getting to \"confirm\", writing a file then reading it back, inserting a row then SELECT-ing it, pushing to an array then checking `.length`.\n\n**When detected:** Delete the verification code. Trust the operation's contract. If the operation can genuinely fail silently, fix the operation — do not paper over it with a post-check. See [`references/code-smells.md` Smell 3](references/code-smells.md#smell-3--redundant-verification-after-a-destructive-action) for examples.\n\n### Smell 4 — Negative-form names and conditions\n\nNaming variables, functions, or flags by the **absence** of a quality (`isNotValid`, `noErrors`, `cannotProceed`, `DisableLogging`) instead of its **presence** (`isValid`, `isClean`, `canProceed`, `LoggingEnabled`). Every negation forces the reader to invert mentally; two negations (`if !isNotReady`) become a logic puzzle nobody reviews confidently.\n\n**When detected:** Rename to the positive form and invert the branch logic. Negation IS appropriate in guard clauses (`if !authorized { return }`) and filters (`items.filter(|x| !x.is_expired())`) — the negative form is the intent there. See [`references/code-smells.md` Smell 4](references/code-smells.md#smell-4--negative-form-names-and-conditions) for the full naming table and examples.\n\n---\n\n## LOGGING — CROSS-CUTTING RULES\n\nLogging is part of the code you ship, and it has iron rules of its own: levels chosen by naming the consumer (never by severity vibes), placement at decision points (never inside helpers), stable messages with structured fields — and, above everything else, **the project's existing practice wins: a project with a designated logger gets that logger and nothing else, and a project that does not log does not get logging uninvited.**\n\n**Read [`references/logging.md`](references/logging.md) BEFORE the change** whenever your edit adds or modifies log lines, sets up a logger or a new service entrypoint, or handles errors at a boundary.\n\n---\n\n## DEPENDENCY UPGRADES — CROSS-CUTTING RULES\n\n- **`0.x` minor = major.** Semver promises nothing below 1.0: treat `0.N → 0.N+1` as a breaking upgrade — read the changelog, build, and run the full suite before trusting it. A required field appearing in a public options type is a routine `0.x` \"minor\".\n- **Version literals live outside the manifest.** Before committing a bump, grep the repo for the old version string: Dockerfiles pinning a global CLI, CI workflows, and docs all carry copies. A bump that updates only the package manifest ships a split-brain deploy.\n- **Never hand-merge a lockfile.** On conflict, take either side whole and regenerate with the package manager — the resolver owns that file, not you.\n\n---\n\n## MANDATORY POST-WRITE REVIEW LOOP\n\n**This runs EVERY time you finish writing or substantively editing code, before you claim the task is done.** No exceptions.\n\n### Step 1 — measure\n\nFor every file you created or modified:\n\n```bash\nawk '!/^[[:space:]]*$/ && !/^[[:space:]]*(\\/\\/|#|--)/' <file> | wc -l\n```\n\nOr run the per-language checker the skill ships:\n\n```bash\n# Python\nuv run scripts/python/check-no-excuse-rules.py <changed paths>\n# Rust\nbash scripts/rust/check-no-excuse-rules.sh <changed paths>\n# TypeScript\nbun run scripts/typescript/check-no-excuse-rules.ts <changed paths>\n```\n\n### Step 2 — interpret\n\n| Pure LOC | Verdict | Required action |\n|---|---|---|\n| ≤ 200 | Healthy | continue |\n| 200 - 250 | **Warning band** | State that fact and propose a split if the next edit will add lines. |\n| > 250 | **DEFECT** | Do NOT commit new lines to this file. Refactor now: split the touched unit before adding lines, except for rare SIZE_OK or pure-data-table exceptions. |\n\n### Step 3 — architectural self-review (always, even at 80 LOC)\n\nAfter every code-writing session, answer these out loud (in your reply) before declaring done:\n\n1. **Single responsibility?** Can I name what this file owns in one short noun phrase? If the answer needs the word \"and\", split.\n2. **Boundary purity?** Did I parse untrusted input into a typed value at the boundary, or did I pass `dict[str, Any]` / `serde_json::Value` / `unknown` past the boundary? If the latter, fix it.\n3. **Variant discrimination?** Did I use `if`/`elif`/`else` (or `switch` without `assertNever`, or `match` without `assert_never`) anywhere to discriminate on a tagged type or enum? If yes, rewrite as exhaustive match.\n4. **Escape hatches?** Any `Any`, `# type: ignore`, `unwrap`, `expect` outside `main`/tests, `as` numeric cast, `!`, `@ts-ignore`, `@ts-expect-error`, `#[allow]` on a real warning? If yes, fix the type or document why with a comment.\n5. **Defensive layer?** Any null check, try/except, or `isinstance` guarding a value the type system already proves? If yes, delete.\n6. **Helpers for one-off?** Any function, class, or trait introduced for a single caller that will never get a second caller? If yes, inline — axiom 0 should have caught it pre-write; this is the backstop.\n7. **Tests?** Is the behavior I just introduced locked by a test that would fail if I revert this commit?\n8. **Parameter bloat?** Any function I wrote or modified that takes more than 3 parameters — or smuggles them through a dict/kwargs/`...args`/throwaway options object? If yes, group related params into a typed value object. See [Smell 2](references/code-smells.md#smell-2--function-with-more-than-3-parameters).\n9. **Redundant verification?** Did I perform a destructive action (delete, remove, clear) and then immediately re-query to \"confirm\" it worked? Did I call a setter then a getter to \"verify\"? If yes, delete the verification — the operation's contract IS the proof. See [Smell 3](references/code-smells.md#smell-3--redundant-verification-after-a-destructive-action).\n10. **Negative naming?** Any variable, function, or flag named by the absence of a quality (`isNotValid`, `noErrors`, `DisableX`) when a positive name (`isValid`, `isClean`, `EnableX`) would work? If yes, rename to positive form and invert the branch. See [Smell 4](references/code-smells.md#smell-4--negative-form-names-and-conditions).\n11. **Logging?** If I touched log lines, logger setup, or error boundaries: did I follow the project's existing practice (including its absence)? Is every new line leveled by its consumer, placed at a decision point, and message-stable with data in fields? See [`references/logging.md`](references/logging.md).\n\n**If any answer fails, fix it before declaring done.** This loop is the difference between \"the code compiles\" and \"the code is correct.\"\n\n### Step 4 — if you need to refactor right now, invoke the right skill\n\n- Any code smell from the [CODE SMELLS section](#code-smells--automatic-review-triggers) fired (250+ LOC, >3 params, redundant verification, negative naming), or step 3 surfaced more than two issues: **load the `refactor` skill** and execute its safe-refactor protocol (codemap, plan, LSP-driven edits, test after each step). Do not improvise a refactor under time pressure — the refactor skill exists precisely so you do not corrupt behavior while reshaping structure.\n- You inherited a branch with AI-generated patterns (broad `except`, redundant null checks, vague TODOs, oversized modules, dead helpers, redundant post-action verification): **load the `remove-ai-slops` skill** to do a categorized branch-scope cleanup with regression tests pinned first.\n\nThese two skills are not optional cosmetics. They are the recovery path for the smells this loop is designed to catch.\n\n---\n\n## Companion skills - explicit invocation triggers\n\n| Trigger | Skill to load | Why |\n|---|---|---|\n| Any [code smell](#code-smells--automatic-review-triggers) fires (250+ LOC, >3 params, redundant verification), OR the post-write loop surfaces 2+ issues, OR the user says \"reshape this\", \"extract this\", \"clean this up\" | `refactor` | Safe codemap-driven multi-step refactor with LSP + tests after each step. Never improvise a structural change. |\n| Recent branch contains AI-authored patterns (broad except, dead helpers, vague comments, oversized files, redundant post-action verification), OR the user says \"remove slop\", \"clean AI code\", \"deslop\" | `remove-ai-slops` | Tests pinned FIRST, then categorized parallel cleanup, then quality gates. Behavior-preserving. |\n| Rust code touches `unsafe`, `*mut`, `*const`, `MaybeUninit`, FFI, `unsafe impl Send/Sync`, or a custom lock-free primitive | `references/rust-ub/` | Full UB taxonomy + Miri strictness escalation. Every `unsafe` block must survive Miri Level 3 (strict provenance + symbolic alignment + preemption) before it ships. |\n\n---\n\n## Per-language jump table\n\n**Stop. Read the matching reference fully before writing code.**\n\n### Python (`.py`, `.pyi`)\n\n**READ `references/python/README.md` FIRST.** Then load on demand:\n\n| Need | Load |\n|---|---|\n| Strict pyproject.toml / basedpyright / ruff config | `references/python/pyproject-strict.md` |\n| Type patterns (`NewType`, `Final`, `TypeGuard`, `Protocol`) | `references/python/type-patterns.md` |\n| Data modeling (Pydantic vs dataclass vs TypedDict vs StrEnum) | `references/python/data-modeling.md` |\n| Error handling (typed exceptions, exhaustive match, union returns) | `references/python/error-handling.md` |\n| Async with anyio (task groups, cancel scopes, channels) | `references/python/async-anyio.md` |\n| httpx2 production defaults (HTTP/2, brotli+zstd, pool tuning) | `references/python/httpx2-optimization.md` |\n| **orjson** in hot paths (FastAPI integration, Pydantic v2 `model_dump_json` vs orjson, Redis/queue/log) | `references/python/orjson-stack.md` |\n| Data processing with polars + duckdb (NEVER pandas) | `references/python/data-processing.md` |\n| FastAPI + SQLAlchemy 2.x async stack | `references/python/fastapi-stack.md` |\n| pydantic-ai agents | `references/python/pydantic-ai.md` |\n| Textual TUI | `references/python/textual-tui.md` |\n| Disposable PEP 723 scripts | `references/python/one-liners.md` |\n| Canonical library defaults | `references/python/libraries.md` |\n\n### Rust (`.rs`, `Cargo.toml`)\n\n**READ `references/rust/README.md` FIRST.** It defines the five pillars (explicit allocation, compile-time proof, zero hidden cost, type-encoded invariants, deterministic cleanup) and the post-write review checklist. Then load on demand:\n\n| Need | Load |\n|---|---|\n| **Arena allocation, const fn, zero-alloc APIs, bitfield, scopeguard, errdefer, Zig-like patterns** | **`references/rust/zero-cost-safety.md`** |\n| Strict `Cargo.toml` lints + profile + workspace config | `references/rust/cargo-strict.md` |\n| Type-state and newtype patterns (Chris Allen's `Point<Screen>` rule) | `references/rust/type-state.md` |\n| `unsafe` discipline (safe wrapper + SAFETY comment + miri proof) | `references/rust/unsafe-discipline.md` |\n| Async with tokio (JoinSet, cancellation, select, blocking work) | `references/rust/async-tokio.md` |\n| Concurrency primitives (locks, atomics, channels, loom) | `references/rust/concurrency.md` |\n| axum + sqlx + tracing + tower HTTP stack | `references/rust/axum-stack.md` |\n| clap + color-eyre + tracing + indicatif CLI stack | `references/rust/clap-stack.md` |\n| Property tests (proptest) + snapshot tests (insta) | `references/rust/proptest-insta.md` |\n| Disposable `rust-script` scripts | `references/rust/one-liners.md` |\n| Canonical library defaults | `references/rust/libraries.md` |\n| **ANY `unsafe` / FFI / `MaybeUninit` / lock-free work** | **`references/rust-ub/` (full directory)** |\n\n### TypeScript (`.ts`, `.tsx`, `.mts`, `.cts`)\n\n**READ `references/typescript/README.md` FIRST.** Then load on demand:\n\n| Need | Load |\n|---|---|\n| Strict tsconfig + Biome config | `references/typescript/tsconfig-strict.md` |\n| Type patterns (branded types, `as const`, `satisfies`, narrowing, `assertNever`) | `references/typescript/type-patterns.md` |\n| Data modeling (type vs interface vs Zod, readonly, parse-don't-validate) | `references/typescript/data-modeling.md` |\n| Error handling (Result, typed errors, union vs throw, AbortSignal timeouts) | `references/typescript/error-handling.md` |\n| Bootstrapping a new project (Bun, pnpm, Hono, Vite) | `references/typescript/bootstrap.md` |\n| Hono backend stack (hono-openapi, Scalar, Swagger, Zod v4) | `references/typescript/backend-hono.md` |\n\n### Go (`.go`, `go.mod`, `go.sum`, `.golangci.yml`, `*.proto`)\n\n**READ `references/go/README.md` FIRST.** Then load on demand:\n\n| Need | Load |\n|---|---|\n| Library defaults (gin vs chi, sqlc, slog, the 2026 stack reasoning) | `references/go/libraries.md` |\n| Canonical strict `.golangci.yml` (v2) with per-linter rationale | `references/go/golangci-strict.md` |\n| Project layout, Taskfile, CI, `go.mod` template | `references/go/bootstrap.md` |\n| Type patterns (named types, smart constructors, sealed interfaces, generics) | `references/go/type-patterns.md` |\n| Data modeling — the three layers of validation (validator/v10 → smart ctor → sqlc) | `references/go/data-modeling.md` |\n| Error handling (`errors.Is/As`, typed errors, `%w` wrapping, no panic) | `references/go/error-handling.md` |\n| Concurrency (`context.Context`, `errgroup`, channels, locks, `-race`, `goleak`) | `references/go/concurrency.md` |\n| HTTP backend stack (gin + slog + validator + pgx, middleware ordering, SSE, WS) | `references/go/backend-stack.md` |\n| RPC stack (Connect-Go default, grpc-go fallback, protovalidate, Buf) | `references/go/grpc-connect.md` |\n| CLI stack (cobra + slog + huh) | `references/go/cobra-stack.md` |\n| Database stack (sqlc + pgx + goose + testcontainers) | `references/go/sqlc-pgx.md` |\n| TUI stack (bubbletea v2 + bubbles v2 + lipgloss v2; **CJK / IME support**) | `references/go/bubbletea-v2.md` |\n| Testing (Given/When/Then, table-driven, fakes-over-mocks, autogold, rapid) | `references/go/testing.md` |\n| Disposable `go run` scripts | `references/go/one-liners.md` |\n\n---\n\n## Activation\n\nThis skill activates whenever you are writing or modifying any `.py`, `.pyi`, `.rs`, `.ts`, `.tsx`, `.mts`, `.cts`, `.go` file, or any project manifest (`pyproject.toml`, `Cargo.toml`, `package.json`, `tsconfig.json`, `biome.json`, `go.mod`, `go.sum`, `.golangci.yml`, `Taskfile.yml`, `buf.yaml`, `sqlc.yaml`). **Even one-off scripts get the full treatment** - that is the whole point of `uv run` + PEP 723, `rust-script`, `bun run`, and `go run` + `//go:build ignore`: production hygiene with throwaway ergonomics.\n\nThe references contain the recipes. **Read them before writing code. Re-read them when the model drifts.** The post-write review loop is non-negotiable.\n</skill>\n\nTASK: Read-only runtime debug of Todo10 blocker `observer-r70 latest FAIL first_day_status`; one deliverable `/home/cube/projects/richard/traning coach/.omo/evidence/nutricoach-telegram-checkin-stepper/task-10-debug-observer-r70.json` with 3+ hypotheses, exact policy/runtime evidence, confirmed cause or unresolved evidence, and safe recovery plan. No Git/GitHub, source/live-state edits, observer row append, service/timer restart, network/customer message/update, authority action. Read actual observer-r70 script/config/observations 94 rows, latest failure row digest, first_day_status calculation code, referenced day-status artifacts, cron/weekly files/timestamps, r70 execution/authorization and current runtime. Determine why latest snapshot has day_status_days only `2026-08-30` at 2026-09-01 12:20 KST, whether this is linked to schedule dispatcher error, stale policy window, missing operational output, or genuine customer-dependent status. Reproduce calculation against copied state/read-only driver without appending rows; do not expose values. Compare prior PASS/FAIL transition timeline. Identify whether a new canonical observer-r70 PASS can be produced safely after repairing cron, or whether customer input/external condition is truly required. Verify launcher count0/r71 roots absent and live draft untouched. Clean temp artifacts.\n\n<Category_Context name=\"deep\">\nYou are operating in DEEP mode. This is the category reserved for goal-oriented autonomous work on hairy problems that reward thorough exploration and comprehensive solutions.\n\nThe orchestrator chose this category because the task benefits from depth over speed. You should feel empowered to spend the time needed: five to fifteen minutes of silent exploration before the first edit is normal and correct. Rushing to implementation on a deep task is a failure mode, not a feature.\n\n# How deep mode adjusts the base behavior\n\n**Exploration budget: generous.** Read the files you need, trace dependencies both directions, fire 2-5 explore/librarian sub-agents in parallel for broader questions. Build a complete mental model before the first `apply_patch`. Exploration here is an investment, not overhead.\n\n**Goal, not plan.** You receive a GOAL describing the desired outcome. You figure out HOW to achieve it. The orchestrator deliberately did not hand you a step-by-step plan; producing one and asking for approval is not what was asked. Execute.\n\n**Atomic task treatment.** When the goal contains numbered steps or phases, treat them as sub-steps of ONE task and execute them all in this turn. Splitting them across turns is wrong unless they reveal an architectural blocker that requires the user's input. If the \"steps\" turn out to be genuinely independent tasks that should have been separate delegations, flag that in your final message and refuse the ones beyond scope.\n\n**Root cause bias.** Prefer root-cause fixes over symptom fixes. A null check around `foo()` is a symptom fix; fixing whatever causes `foo()` to return unexpected values is the root fix. Trace at least two levels up before settling on an answer. In deep mode, you have permission (and the expectation) to do the deeper fix.\n\n**Ambition scaled to context.** For brand-new greenfield work, be ambitious. Choose strong defaults, avoid AI-slop aesthetics, produce something you would be proud to hand to another senior engineer. For changes in an existing codebase, be surgical and respect the existing patterns; depth does not mean invasiveness.\n\n**Completion bar: full delivery.** \"Simplified version\", \"proof of concept\", and \"you can extend this later\" are not acceptable deliveries for a deep task. The orchestrator routed here specifically for a complete solution. If you hit a genuine blocker (missing secret, design decision only the user can make, three materially different attempts all failed), document it and return; otherwise, finish the task.\n\n**Status cadence: sparse.** The user is not on the other side of this conversation; the orchestrator is, and they will synthesize your progress. Send commentary only at meaningful phase transitions (starting exploration, starting implementation, starting verification, hitting a genuine blocker). Do not narrate every tool call; silence during focused work is expected.\n</Category_Context>"},"host_pid":1353625,"final_response":"I read this as a read-only runtime investigation - completed with only the requested evidence JSON written.\n\nCreated:\n\n`/.omo/evidence/nutricoach-telegram-checkin-stepper/task-10-debug-observer-r70.json`\n\nKey finding: observer-r70 reads the stale unsuffixed authority root. The configured r70 root already had the required `2026-08-31` status before the deadline. Reproduction toggled `FAIL(first_day_status)` to `PASS` solely by selecting the configured root, twice.\n\nThe cron identity-drift error is a separate blocker; no customer input is required. Launcher count remains zero, all r71 roots remain absent, the live draft hash is unchanged, and temporary artifacts were removed. JSON validation and diagnostics passed.","run_stats":{"runtime_ms":335531,"turns":17,"tool_calls":62,"output_tokens":16084,"input_tokens":236718,"cache_read_tokens":2551808,"total_tokens":2804610,"generation_ms":319915,"tokens_per_second":50,"cost_usd":2.9420140000000004,"cache_hit_rate_last":0.9904910311122129,"cache_hit_rate_run":0.9151099900090586,"token_status":"complete","cost_status":"reported","duration_status":"monotonic"},"task_seq":64,"config_generation":0,"background_mode":"background"}