# Consensus plan — 자유 윤문 + 강한 검증·실패 시 원문 복귀

## Decision, boundary, and research basis

**User reconciliation decision (binding):** use a real, open-ended Korean *connective-prose* postprocessor: `자유 윤문 + 강한 검증·실패 시 원문 복귀`. It is not a finite variant table, fixed style lexicon, or model-selected template. Every value, date, predicate/decision, action, safety state, approval/delivery state, heading, button, envelope, and layout boundary remains a locally expanded immutable placeholder or locked canonical byte.

Stage-09 and stage-10 artifacts are research only. Their concrete gaps are closed below: shadow proof goes through the coordinator and all four real surface paths; shadow has an explicit no-selected-output contract; config migration has an owner/API/command; the live marker is registered; and adaptive presentation is committed atomically with current authority.

This is presentation-only, disabled by default. It does not activate a customer, alter consent, proposal calculation, adaptive epoch, routing, Gate-D, approval, or delivery authority. One configured direct OpenAI-compatible provider is allowed; Kimi, auxiliary/provider routers, alternate providers, retries, and repairs are out of scope. Any ineligibility, configuration/provider/journal/validation/commit error selects the already rendered canonical bytes.

## Surface contracts and insertion points

| Surface | Eligible canonical slot | Immutable/local material | Selected-byte owner / replay |
|---|---|---|---|
| `daily.v1` | one status paragraph between daily facts and `오늘 할 일` in `TelegramAdapter._nutrition_daily_text` | title, seven facts, headings, actions/order, all values and safety state | no selected persistence; initial edit only; every replay is canonical/no-call |
| `weekly.v1` | typed interpretation and rationale paragraphs in `TelegramAdapter._nutrition_report_text`, only on fully typed branches | report title, facts/ranges/rates, judgment, action bullets, explicit source text, layout | existing prepared scheduled-delivery `body`; postprocess before reservation only |
| `adaptive.dualcoachtest.card.v1` | full `현재 판단:` value plus code-mapped `검토 필요` reason lines in the current dual card grammar | envelope, title/customer/date/facts, recommendations, notes, delivery line, buttons | existing adaptive `card_payload`; recovery republishes exact bytes/no call |
| `adaptive.physique.card.v1` | new code-owned `검토 안내` sentence immediately before the exact final status in the physique card | every existing structured card line, final status, envelope/buttons | existing adaptive `card_payload`; recovery republishes exact bytes/no call |

`gateway/platforms/telegram.py` currently renders daily feedback through `_saved_physique_coaching_feedback`, `_generate_physique_coaching_feedback`, `_request_physique_coaching_feedback`, `_nutrition_daily_interpretation(feedback)`, and `_render_physique_feedback_replay`; remove this finalized-render route rather than retain an alias. Refactor `_nutrition_daily_text(snapshot)` to canonical complete/partial status first. Add `PhysiqueCheckinBridge.finalized_coaching_artifact(session_id) -> FinalizedCoachingArtifact | None` in `gateway/platforms/physique_checkin.py`; it loads one finalized session once and returns the trusted snapshot and internal finalized event id together. The id only derives the opaque artifact key and is never sent/logged. First-save may call the postprocessor; replay is canonical only.

For weekly, change `_send_nutrition_coaching_tick` so a new weekly task renders canonical, obtains selected-or-canonical presentation, computes its digest, then calls `reserve_customer_task_delivery`. An existing `prepared` weekly receipt first calls `load_prepared_customer_task_delivery`; it must not rerender or call the postprocessor. Add `PinnedScheduledDelivery` and `load_prepared_customer_task_delivery(profile_root, receipt)` to both profile copies of `checkin_cli/customer_schedule.py` and export it from both `checkin_cli/__init__.py`. Under the existing schedule lock, it validates ready fence, complete hash chain/tombstones, receipt identity/state, body/destination digests, and registry/config pins, then returns only the exact durable body/destination. Read failure hard-stops; no generic lookup, row/schema change, resend, or model call is permitted.

## Open Korean contract and validation boundary

Create `/home/cube/projects/richard/hermes-agent/gateway/platforms/korean_expression.py` as the sole feature module. It owns frozen/slotted `KoreanExpressionConfig`, `CanonicalSurface`, `ProseSlot`, `Placeholder`, `PresentationCandidate`, `KoreanExpressionCoordinator`, strict JSON parser, surface adapters, direct client, journal, and config-cutover preflight receipt/API. It imports neither profile-local `checkin_cli` nor any generic/auxiliary provider resolver.

The provider receives a fixed instruction plus a bounded JSON capsule containing only schema version, surface, HMAC artifact key, each slot's issued semantic/action IDs, issued literal placeholders, and limits. It receives no canonical body, expanded placeholder text, customer/session/proposal/schedule IDs, name/address, notes/history, measurements, dates, raw facts, raw action text, destination, or telemetry. It returns one JSON object with exactly `schema_version`, `surface`, `artifact_key`, and `slots`; every slot has exactly `slot_id`, `semantic_ids`, `action_ids`, `placeholder_tokens`, and `generated_ko`. Arrays must exactly echo the issued IDs/tokens in order.

Each slot's `generated_ko` is genuine unrestricted Korean connective prose around the issued placeholders: there is **no** finite allowed-word list, local variant id, selector, or enumerated output catalog. The only structural requirements are NFC UTF-8, exact issued placeholders once each in declared order, no other placeholder/identifier syntax, and per-slot/document byte/character limits. All truth-bearing content is issued as local placeholders and expanded only after validation; slots never carry raw values or semantic expansions.

Validation is fail-closed and document-wide (one invalid slot rejects the whole response):

1. Reject invalid UTF-8, BOM/control/newline where the slot forbids it, non-NFC text, JSON over 16 KiB, duplicate JSON keys at every depth, non-exact keys/types, wrong schema/surface/HMAC key, or missing/reordered/duplicate/extra slots, IDs, or placeholders.
2. Mask only exact issued placeholders, then reject every new numeric/date/unit token (ASCII/full-width digits, signed/percent/range forms, ISO/Korean date forms, and number-word-plus-time/unit forms), URLs, Latin identifiers, new bracket/tag syntax, and unknown semantic/action/placeholder identifiers.
3. Reject `GLOBAL_FORBIDDEN_CONNECTIVE_TERMS` and surface-specific normalized Korean stem/phrase sets outside placeholders, including medical/diagnostic/prescription, risk/safety/emergency, evidential/factual, quality/adherence, decision/adjustment/action, and approval/hold/delivery claims. The static lists are code-reviewed and include inflected stem matches; a placeholder expansion is masked before this scan. Add adversarial tests for synonyms/inflections, but do not misrepresent the list as a complete Korean semantic theorem.
4. Enforce per-slot maximum generated characters, maximum final slot expansion, and final-body change budget; expand local placeholders and assert byte identity outside declared slots, exact headings/blank lines/envelope/buttons, Telegram UTF-16 bounds, and no newly introduced numeric/date token outside protected canonical spans.

Prompt policy says prose may only connect the supplied placeholders and may not assert a fact, cause, recommendation, safety conclusion, approval, or delivery event. These deterministic gates make semantic safety strong best effort, **not mathematically complete**: Korean paraphrase can evade a finite lexical detector. The ADR and operator guide must state this limit honestly; the canonical semantic anchors, strict bans, bounded change, shadow evidence, default-off rollout, and immediate canonical fallback are the risk controls—not a proof that every free Korean sentence is harmless.

Per-surface manifests retain the prior typed branches: daily has exactly `P_DAILY_COMPLETE` or `P_DAILY_PARTIAL`; weekly has only local typed interpretation/rationale branch IDs; dual maps only known decision/reason enums; physique has only `A_OPERATOR_REVIEW_DECISION`. All values, decisions, actions, safety, and approval/delivery language are their exact local expansions. Unknown/untyped branch, explicit unstructured source text, malformed card grammar, menu/view/recovery/terminal payload, or future schema is canonical/no-call.

## One direct provider, outcome journal, and exact shadow contract

`KoreanExpressionConfig.from_extra` is parsed at `TelegramAdapter.__init__`. Missing config is disabled. A positive configuration is exactly one direct OpenAI-compatible identity: `provider_id=openai-direct`, `transport=chat_completions`, `endpoint=https://api.openai.com/v1`, `model=gpt-4.1-mini`; all unknown/normalized/override fields fail closed. Only `OPENAI_API_KEY` is read at direct-client construction. Construct `openai.OpenAI(api_key=..., base_url="https://api.openai.com/v1", max_retries=0, timeout=3.0)` and issue at most one non-streaming `chat.completions.create` call (`n=1`, literal model, JSON object response, `store=False`, no tools). Worker deadline is four seconds. Missing key, 429, 5xx, cancellation, timeout, malformed/late response, or any validation failure makes zero further requests and selects canonical. No Kimi/Moonshot, auxiliary client, OAuth, proxy, pool, model default, fallback provider, or application retry may be reached.

Keep only metadata in `data/korean-expression-v1/`: 0700 root, 0600 no-follow UID-owned identity key/lock/daily JSONL journal, revalidated before use. No feature-owned text store exists. Artifact identity is HMAC-SHA256 over versioned length-prefixed internal surface identity; raw identity inputs never leave process.

A canonical journal row has exactly `schema_version, sequence, artifact_key, surface, mode, attempt_id, state, outcome, started_at_utc, deadline_at_utc, finalized_at_utc, previous_row_digest, row_digest`, with canonical UTF-8 JSON, contiguous per-segment sequence, zero predecessor for the first row, and SHA-256 hash chaining. Start is `attempt_started`; terminals are `selected`, `shadow_selected`, `credential_unavailable`, `provider_error`, `timeout`, `invalid_response`, `validation_rejected`, `commit_failed`, `publication_failed`, or `claim_expired_canonical`. It never contains prompt/response/body, placeholder or semantic IDs, raw identity, binding digest, exception text, provider/model, destination, or lengths. Retain complete validated segments for 35 completed UTC days plus current day.

**Shadow is normative, not a near-enabled mode:** for an eligible artifact, the coordinator may make its one provider attempt and fully validate it. On acceptance it appends terminal `state=finalized, outcome=shadow_selected`, then immediately discards selected text. It returns only the canonical `PresentationCandidate`; it exposes no selected body to the adapter/caller, journals no selected text, and permits no selected text in Telegram edits, schedule reservation/ledger, adaptive pending/published cards, logs, or telemetry. Normal daily edit, weekly reservation/send, and adaptive pending publication proceed using exact canonical bytes; `shadow_selected` is validation evidence, not a publication result. Disabled/invalid config creates no row or client. An unexpired competing claim returns canonical/no-call; post-expiry appends `claim_expired_canonical` and still makes no second attempt. In enabled mode only, terminal `selected` follows the existing durable destination write (daily edit, weekly reservation, adaptive atomic pending write); any failure is terminal canonical, never repair/retry.

## Adaptive atomic presentation boundary

Replace fresh-card gateway use of public `AdaptiveOperatorService.mark_publish_pending` with:

```python
def validate_and_mark_presented_card(
    self,
    callback_data: object,
    *,
    canonical_payload: Mapping[str, object],
    presentation: PresentationCandidate,
    origin_message_id: object = "",
) -> Mapping[str, object]: ...
```

in `gateway/platforms/nutrition_coaching.py`. Extract the current durable append as private `_mark_publish_pending_unlocked`. `PresentationCandidate` is internal-only and holds rendered selected-or-canonical payload, surface version, attempt ID, opaque artifact key, and a binding digest over canonical callback/session action, proposal digest/revision, canonical enveloped text/buttons, origin provenance, and source schema version.

Under the existing `_authority_session_lock()` and `_publication_ledger_lock()`, the method reloads the latest session; revalidates exact callback action, expiry, operator/config/registry/consent/activation/source/epoch/owner pins, proposal digest/revision, provenance, and fresh `status == "card"`; reparses the canonical payload with the appropriate adapter; recomputes and matches binding digest; reruns slot/locked-byte validation; then appends that exact payload through `_mark_publish_pending_unlocked` before returning it. Telegram edits only this returned payload, then use existing `mark_published`. Shadow supplies a canonical candidate, so its durable card is canonical. Menu stays on current canonical persistence; terminal/view/operator-input/recovery never call the postprocessor. Stale/unequal/racing candidates append no pending row and edit nothing; same candidate is idempotent; recovery reads the sole persisted payload and never regenerates.

## Explicit legacy-config preflight and ordered migration

The migration owner is `gateway/platforms/korean_expression.py::audit_physique_checkin_cutover(extra: object) -> PhysiqueCheckinCutoverReceipt`. It returns bounded, redacted enum-only evidence: schema version, `ready`, checks `legacy_feedback_removed`, `physique_bridge_valid`, `postprocessor_disabled_or_valid`, config SHA-256, and reason codes from `legacy_coaching_feedback_enabled`, `physique_checkin_invalid`, `postprocessor_config_invalid`, `config_shape_invalid`. No config values, profile paths, IDs, or secrets appear in output.

Add `/home/cube/projects/richard/hermes-agent/scripts/audit_physique_config_cutover.py` as the sole release invocation. It reads each passed YAML directly with `yaml.safe_load` (no default/fallback loader), extracts `platforms.telegram.extra`, calls the API, emits one redacted receipt per input, and exits nonzero unless every receipt is ready. The mandatory pre-cutover and promotion invocation, from the Hermes root, is:

```text
python scripts/audit_physique_config_cutover.py \
  --config /home/cube/.hermes/profiles/physique-coach/config.yaml \
  --config /home/cube/.hermes/profiles/dualcoachtest/config.yaml
```

First remove only `coaching_feedback_enabled` from both real config files and add an exact disabled `korean_expression_postprocessor` block. Run this command before parser removal and before every promotion. It must prove the active physique profile still yields a valid enabled `PhysiqueCheckinConfig`; the disabled dualcoachtest shape is also checked, so it cannot hide the legacy key behind an early disabled return. Then remove the dataclass field and all old finalized-feedback calls. After cutover, an enabled raw `physique_checkin` containing the legacy key produces the existing `_physique_checkin_error` startup diagnostic with the explicit legacy-key reason, never a silent disabled bridge or compatibility route. Rollback is config-first: set postprocessor mode/surfaces disabled, stop/restart after claim deadline, preserve journal/ledger/card evidence, and never restore the legacy key.

## Exact file and symbol plan

1. **New** `/home/cube/projects/richard/hermes-agent/gateway/platforms/korean_expression.py`: all contracts above, `KoreanExpressionConfig.from_extra`, `KoreanExpressionCoordinator`, strict response validator, four pure adapters, direct client/journal, `audit_physique_checkin_cutover`, and `PhysiqueCheckinCutoverReceipt`.
2. `/home/cube/projects/richard/hermes-agent/gateway/platforms/physique_checkin.py`: `FinalizedCoachingArtifact` and `finalized_coaching_artifact`; preserve active/conversation behavior.
3. `/home/cube/projects/richard/hermes-agent/gateway/platforms/physique_checkin_config.py`: remove `coaching_feedback_enabled` from `PhysiqueCheckinConfig`/`from_extra` and reject the supplied obsolete key.
4. `/home/cube/projects/richard/hermes-agent/gateway/platforms/telegram.py`: change `_nutrition_daily_text`, `_render_physique_callback_prompt`, `_render_physique_feedback_replay`, `_send_nutrition_coaching_tick`, and `_handle_adaptive_review_callback`; delete the obsolete finalized-feedback methods/route; inject only the coordinator at the new slots.
5. `/home/cube/projects/richard/hermes-agent/gateway/platforms/nutrition_coaching.py`: add `AdaptiveOperatorService.validate_and_mark_presented_card` and private unlocked persistence extraction; do not alter proposal/event/delivery schemas or recovery semantics.
6. Both `/home/cube/.hermes/profiles/{physique-coach,dualcoachtest}/workspace/checkin_cli/checkin_cli/customer_schedule.py` and `__init__.py`: `PinnedScheduledDelivery` and `load_prepared_customer_task_delivery` only.
7. Both `/home/cube/.hermes/profiles/{physique-coach,dualcoachtest}/config.yaml`: ordered key removal plus disabled feature block.
8. **New** `/home/cube/projects/richard/hermes-agent/scripts/audit_physique_config_cutover.py`; `/home/cube/projects/richard/hermes-agent/pyproject.toml`: register `live_korean_expression: opt-in real OpenAI Korean-expression shadow proof` (and mark the live test `integration` so default `not integration` still excludes it).
9. **New** `/home/cube/projects/richard/hermes-agent/tests/gateway/test_korean_expression.py` and `test_korean_expression_live_shadow.py`; extend `test_telegram_physique_checkin.py`, `test_nutrition_coaching.py`, `test_adaptive_nutrition.py`, and `test_telegram_group_gating.py`. Extend both profile `tests/test_customer_schedule.py` and `tests/test_adaptive_nutrition.py`.
10. `/home/cube/projects/richard/traning coach/듀얼코치_사용설명서.md`: disabled default, user reconciliation, explicit best-effort limit, provider/no-retry/privacy/shadow/migration/rollback guidance. **New** `/home/cube/projects/richard/hermes-agent/docs/adr/ADR-021-korean-expression-postprocessor.md`: the binding decision and its limits.

## Verification and release gates

Hermetic tests cover: exact canonical and selected goldens for daily complete/partial, all typed weekly branches, and both adaptive grammars; full request/response/local-render vectors; open prose accepted when it uses novel benign Korean connective language (proving no finite style list); every malformed schema/duplicate key/NFC/control/identifier/ID/token/order/number/date/unit/medical/safety/approval/delivery/prohibited-vocabulary/length/locked-byte attack falling back to canonical; one-call/no-retry/direct-config/credential isolation; journal hashing/permissions/torn/claim/late-worker/retention traces; and no content in journal/log/telemetry.

The migration tests exercise both current legacy shapes and assert enum `legacy_coaching_feedback_enabled`, nonzero script exit, no filesystem mutation, active physique bridge continuity after removal, disabled dual profile handling, missing/disabled/malformed feature config, and zero requests. A marker/collection test asserts `pytest --markers` contains `live_korean_expression` and that the explicit live node cannot silently select zero tests.

Schedule tests prove prepared weekly success, pin/digest/config mismatch, corrupt ledger hard-stop, exact post-crash send, and zero re-render/model call. Adaptive tests use a barrier after candidate creation: owner/config/epoch/revision change causes zero pending rows/edits; same candidate commits once; unequal candidates conflict; shadow persists canonical bytes only; recovery republish makes zero calls. Both profile suites capture their own `render_operator_card` grammar before adapter tests.

The live proof is coordinator-driven, not direct-client-only. `test_korean_expression_live_shadow.py` runs all four product paths using temporary journal/profile roots, synthetic redacted manifests, injected direct OpenAI client, a fake Telegram editor, real temporary weekly ledger seam, and temporary adaptive service. It asserts for each accepted shadow result: one provider call, terminal `shadow_selected`, canonical visible daily bytes, canonical reserved weekly body, canonical adaptive pending/published payload, no selected output in any durable artifact/log/telemetry, and no selected candidate returned to the adapter. Four cases × 12 distinct synthetic capsules = exactly 48 calls in a 24-hour release-candidate window; no timeout/provider/validation/credential outcome is allowed.

Run from `/home/cube/projects/richard/hermes-agent`:

```text
pytest -q tests/gateway/test_korean_expression.py tests/gateway/test_telegram_physique_checkin.py tests/gateway/test_nutrition_coaching.py tests/gateway/test_adaptive_nutrition.py tests/gateway/test_telegram_group_gating.py
pytest -q /home/cube/.hermes/profiles/physique-coach/workspace/checkin_cli/tests/test_adaptive_nutrition.py /home/cube/.hermes/profiles/physique-coach/workspace/checkin_cli/tests/test_customer_schedule.py
pytest -q /home/cube/.hermes/profiles/dualcoachtest/workspace/checkin_cli/tests/test_adaptive_nutrition.py /home/cube/.hermes/profiles/dualcoachtest/workspace/checkin_cli/tests/test_customer_schedule.py
python scripts/audit_physique_config_cutover.py --config /home/cube/.hermes/profiles/physique-coach/config.yaml --config /home/cube/.hermes/profiles/dualcoachtest/config.yaml
RUN_KOREAN_EXPRESSION_LIVE=1 pytest -q -m live_korean_expression tests/gateway/test_korean_expression_live_shadow.py
```

Promotion is human-approved and surface-by-surface: default disabled → all hermetic/preflight gates plus 48/48 controlled proof → test-profile shadow for 14 days with at least 10/10 accepted observations and zero failures independently for daily, weekly, dual adaptive, and physique adaptive → enabled daily → enabled weekly → shared adaptive only after both adaptive evidence sets pass. Any failure disables only that surface and restores canonical output; this is not real-customer activation, Gate-D completion, delivery approval, or permission to resend.

## ADR and attribution

ADR-021 records the rejected alternatives: finite local style composition fails the user's free-prose decision; unrestricted whole-body rewrite violates immutable semantic boundaries. The accepted design is free connective prose around local semantic placeholders, with best-effort lexical/structural validation and canonical fallback. Consequence: measurable residual semantic risk remains and must not be described as a formal guarantee.

The guide and ADR preserve the non-copy attribution: `https://github.com/Gaeduck-0908/im-not-ai-kiro/tree/901e378ae7f77035d17a491067343ac3f83d0214` and `https://github.com/epoko77-ai/im-not-ai/tree/53e24e8f92cf344efcb812103f7c2b203e7efffc` were consulted only for a high-level preservation goal. No source code, prompts, rules, workflow, or artifacts were copied; no license notice is reproduced because no material is copied.
