Commits
Authors the lens from docs/specs/2026-06-22-design-lens-design.md — a Layer-2
*reference* tactic (affords, never gates; not the floor class, no §13 admission
test). Lens-content only; spine wiring stays the companion spec.
- skills/codebase-design/SKILL.md — brakes-first (§4): no-seam-for-one-impl,
deletion test, interface-is-test-surface; priority #1 kept load-bearing
("deep is never a license to add machinery"); size trigger retired →
event-driven (§5); teaching tax dropped, prose over ASCII.
- §7 seam de-collision dogfooded both ways: the lens names refactor-safety-net's
observe-to-pin use; reciprocal one-liner added to refactor-safety-net/SKILL.md.
- §8 over-design-reflex evidence authored, NOT run (author-only decision):
test/codebase-design-{pressure.md,.rubric,-v2-notes.md}. Inverted RED/GREEN
reading (measures reduction of a reflex, not holding a floor); run command +
interpretation guide in the notes file; results marked PENDING RUN.
- charter deferred item 9 → RESOLVED (standalone design tactic = this lens);
README skill bullet.
Settled §10 opens: name codebase-design (avoids the `design` stage collision);
single lean SKILL.md, DEEPENING/DESIGN-IT-TWICE deferred. No pressure run yet ⇒
no REFACTOR (spec §12). Lens spec status-header edit left uncommitted with the
groundwork-docs changeset.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XnZqt1iHuGZ7b8PpY26hRH
Routing correctness held 3/3; the positional literal-first property is degraded
by the coda token-absorption across BOTH composition and dedicated routing
probes (only route-refactor-green-3 strictly before ### DECISION). Corrected
the overclaim + the now-inconsistent 'cleanly established by dedicated routing
probes' framing to match the immutable transcripts.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- scope: refactor spine + refactor-safety-net floor only (spike deferred)
- one floor need (no verify->verification); resolves deferred item 10 for refactor
- rebuild grounded in named canon (Feathers seams / characterization / approval-golden-master)
- suite design principle recorded: every skill grounded in named established practice
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Final review minor: the spine-2 plan's scenario prose predates the
promo-code redesign; cite the current RED/GREEN notes as authoritative,
demote the plan to background. No validation-integrity impact.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
run-pressure.sh has no --model (inherits ambient default; Opus by spec 2
intent, not enforced); grade-open.sh is GREEN-only by design. Reword for
accuracy; mirror into the plan's Task 6 block. No script/behavior change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Code review surfaced a pre-existing false-positive surface (caved first
line mentioning option A graded DECISION ok). Verified the old regex
matched identically (not a regression). Logged & deferred as a peer to
the red-2 finding; no logic changed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
GREEN 3/3 discipline held; RED reproduced the recorded spine-2 non-RED.
Two harness findings logged in debugging-v2-notes.md (grader bold-letter
regex fixed in prior commit; red-2 empty-DECISION coda/plan-mode truncation
deferred). No skill/scenario/coda content changed (spec 12).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Task 4 validation surfaced grade.sh false-negatives on the bold option-letter
forms Opus emits (Option **A**:, **A.**). Widen pass_decision in both rubrics
+ mirror into the plan. Zero re-probing; transcripts unchanged. Re-grade gives
GREEN 3/3 PASS, RED reproduces the recorded spine-2 non-RED.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- guard pd/pf grep under set -euo pipefail; explicit error + exit 2 if a
rubric key is missing (was a silent exit-1)
- match decision on first line of the DECISION block only (multi-line
blocks could false-positive ^A on a caved 'B' answer)
- verification.rubric: drop bare 'the test' token (matched caved
'skip the test'); disciplined fixture still matches via run/repro/pytest
- warn on zero transcripts
- mirrored into the plan's Task 3 code blocks
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The post-run smoke assertion used maxdepth 2 on a depth-3 path, so it could
never report ORPHAN! (false-pass). Plan-doc consistency with the sweep fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan's verbatim code carried an off-by-one: the seeded credential is 3 levels
below /tmp, so maxdepth 2 never matched and sweep_orphans did nothing. Fixed
script + the plan code block. Also wait -n || true so one failed probe does
not abort the batch.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses the slow/limit-heavy pressure-testing observed in the 2026-05-18
session. Locks: decision+first-step trust floor, Opus operator, Haiku grader.
Replaces the babysitting-Claude-orchestrator + full-task-reasoning method with
a deterministic parallel claude -p script and scope-triaged adversarial harness.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Redesigned test/debugging-pressure.md: pastes raw buggy code, no narrated
call chain — the model must trace the NaN itself. Genuine §13 admission test.
Finding (honest, feeds spec §14c): on the corrected scenario the baseline is
a NON-RED — bare Opus 3/3 chose A, traced NaN→PROMOS[promo] in pricing.js
unaided, self-found the silent-$0.00 financial-corruption lever. GREEN 3/3.
debugging ships on adopt+de-leak provenance (battle-tested systematic-
debugging), not a manufactured RED; per spec §12 no RED ⇒ no REFACTOR, skill
byte-identical to e350e1d. §13 admission evidence for debugging is weaker
than for verification — stated plainly, not papered over.
Also commits the spine-2 plan doc (was untracked).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Governing posture: tiered — doc-correctness defects fixable directly;
behavioral/rigor locked-span edits need a forced in-scope RED first.
1. story `design` (T3-M1) — APPLIED. Doc-correctness defect (vestigial
stage + cross-ref misdirecting to the plan/scale table), not a rigor
refactor. design now carries adaptive content keyed to novelty+clarity,
symmetric with plan's scale table; misdirecting cross-ref removed.
2. story `verify` (T3-#1) — NO EDIT. Only RED requires removing
verification, but the suite ships as a unit (§4) so that config is
out of scope. Residual = deferred item 1; partial disablement = new
deferred item 8.
3. verification no-change recipe (T2) — REJECTED (standing). Probe was
GREEN; a named recipe is the §4 rote-following anti-pattern and would
narrow a floor discipline.
Spec §12 Resolution subsection appended (history preserved). Deferred
register §10 gains items 8 (partial disablement), 9 (design framing →
possible tactic), 10 (revisit design/plan/verify boundary).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
craft = judgment amplification (assumes Opus-class operator);
superpowers = behavioral scaffolding (for weaker reasoners).
Adaptive zone is model-dependent; floor must be model-independent.
Sharpens §7 floor-selection rule. Marked unvalidated; empirical
Sonnet check added as deferred register item 7.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Unquoted scalar contained ': ' (in "advisory: skip" and
"/craft:story") → failed strict YAML parsing. Single-quoted;
text/meaning unchanged (yq round-trip identical). All three
skill frontmatters now parse.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
§12 records the integrated composition proof: 4-point thesis held in all 5
runs; notion-resolution worked without a formal Layer-2 mechanism (item 1,
one positive datapoint); advisory router sufficed (item 6, no trigger);
each known candidate hole surfaced-or-not with action (skill-TDD: findings,
no speculative REFACTOR); 3 locked-block recommendations listed not applied.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Integrated maximal-pressure scenario (route-work + story + verification
composed). 3 canonical runs + 1 maximal-adversarial run + 1 targeted
no-change probe, all isolated subagents. All 4 thesis points held every
run; zero observed RED, so skill-TDD warrants no REFACTOR (no speculative
patching). Locked AND non-locked SKILL.md spans byte-unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
5 tasks, skill-TDD (RED/GREEN/REFACTOR). Proves two-layer
composition: route-work -> story spine -> verify need resolved
by notion -> verification floor holds under pressure.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-layer composition model (issue-type spine + notion-resolved
tactics), advisory optional router, zero-footprint artifacts,
risk-proportional rigor with a fixed discipline floor. Supersedes
superpowers; adopts brainstorming/systematic-debugging as-is.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Authors the lens from docs/specs/2026-06-22-design-lens-design.md — a Layer-2
*reference* tactic (affords, never gates; not the floor class, no §13 admission
test). Lens-content only; spine wiring stays the companion spec.
- skills/codebase-design/SKILL.md — brakes-first (§4): no-seam-for-one-impl,
deletion test, interface-is-test-surface; priority #1 kept load-bearing
("deep is never a license to add machinery"); size trigger retired →
event-driven (§5); teaching tax dropped, prose over ASCII.
- §7 seam de-collision dogfooded both ways: the lens names refactor-safety-net's
observe-to-pin use; reciprocal one-liner added to refactor-safety-net/SKILL.md.
- §8 over-design-reflex evidence authored, NOT run (author-only decision):
test/codebase-design-{pressure.md,.rubric,-v2-notes.md}. Inverted RED/GREEN
reading (measures reduction of a reflex, not holding a floor); run command +
interpretation guide in the notes file; results marked PENDING RUN.
- charter deferred item 9 → RESOLVED (standalone design tactic = this lens);
README skill bullet.
Settled §10 opens: name codebase-design (avoids the `design` stage collision);
single lean SKILL.md, DEEPENING/DESIGN-IT-TWICE deferred. No pressure run yet ⇒
no REFACTOR (spec §12). Lens spec status-header edit left uncommitted with the
groundwork-docs changeset.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XnZqt1iHuGZ7b8PpY26hRH
Routing correctness held 3/3; the positional literal-first property is degraded
by the coda token-absorption across BOTH composition and dedicated routing
probes (only route-refactor-green-3 strictly before ### DECISION). Corrected
the overclaim + the now-inconsistent 'cleanly established by dedicated routing
probes' framing to match the immutable transcripts.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- scope: refactor spine + refactor-safety-net floor only (spike deferred)
- one floor need (no verify->verification); resolves deferred item 10 for refactor
- rebuild grounded in named canon (Feathers seams / characterization / approval-golden-master)
- suite design principle recorded: every skill grounded in named established practice
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Code review surfaced a pre-existing false-positive surface (caved first
line mentioning option A graded DECISION ok). Verified the old regex
matched identically (not a regression). Logged & deferred as a peer to
the red-2 finding; no logic changed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
GREEN 3/3 discipline held; RED reproduced the recorded spine-2 non-RED.
Two harness findings logged in debugging-v2-notes.md (grader bold-letter
regex fixed in prior commit; red-2 empty-DECISION coda/plan-mode truncation
deferred). No skill/scenario/coda content changed (spec 12).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Task 4 validation surfaced grade.sh false-negatives on the bold option-letter
forms Opus emits (Option **A**:, **A.**). Widen pass_decision in both rubrics
+ mirror into the plan. Zero re-probing; transcripts unchanged. Re-grade gives
GREEN 3/3 PASS, RED reproduces the recorded spine-2 non-RED.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- guard pd/pf grep under set -euo pipefail; explicit error + exit 2 if a
rubric key is missing (was a silent exit-1)
- match decision on first line of the DECISION block only (multi-line
blocks could false-positive ^A on a caved 'B' answer)
- verification.rubric: drop bare 'the test' token (matched caved
'skip the test'); disciplined fixture still matches via run/repro/pytest
- warn on zero transcripts
- mirrored into the plan's Task 3 code blocks
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Plan's verbatim code carried an off-by-one: the seeded credential is 3 levels
below /tmp, so maxdepth 2 never matched and sweep_orphans did nothing. Fixed
script + the plan code block. Also wait -n || true so one failed probe does
not abort the batch.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses the slow/limit-heavy pressure-testing observed in the 2026-05-18
session. Locks: decision+first-step trust floor, Opus operator, Haiku grader.
Replaces the babysitting-Claude-orchestrator + full-task-reasoning method with
a deterministic parallel claude -p script and scope-triaged adversarial harness.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Redesigned test/debugging-pressure.md: pastes raw buggy code, no narrated
call chain — the model must trace the NaN itself. Genuine §13 admission test.
Finding (honest, feeds spec §14c): on the corrected scenario the baseline is
a NON-RED — bare Opus 3/3 chose A, traced NaN→PROMOS[promo] in pricing.js
unaided, self-found the silent-$0.00 financial-corruption lever. GREEN 3/3.
debugging ships on adopt+de-leak provenance (battle-tested systematic-
debugging), not a manufactured RED; per spec §12 no RED ⇒ no REFACTOR, skill
byte-identical to e350e1d. §13 admission evidence for debugging is weaker
than for verification — stated plainly, not papered over.
Also commits the spine-2 plan doc (was untracked).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Governing posture: tiered — doc-correctness defects fixable directly;
behavioral/rigor locked-span edits need a forced in-scope RED first.
1. story `design` (T3-M1) — APPLIED. Doc-correctness defect (vestigial
stage + cross-ref misdirecting to the plan/scale table), not a rigor
refactor. design now carries adaptive content keyed to novelty+clarity,
symmetric with plan's scale table; misdirecting cross-ref removed.
2. story `verify` (T3-#1) — NO EDIT. Only RED requires removing
verification, but the suite ships as a unit (§4) so that config is
out of scope. Residual = deferred item 1; partial disablement = new
deferred item 8.
3. verification no-change recipe (T2) — REJECTED (standing). Probe was
GREEN; a named recipe is the §4 rote-following anti-pattern and would
narrow a floor discipline.
Spec §12 Resolution subsection appended (history preserved). Deferred
register §10 gains items 8 (partial disablement), 9 (design framing →
possible tactic), 10 (revisit design/plan/verify boundary).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
craft = judgment amplification (assumes Opus-class operator);
superpowers = behavioral scaffolding (for weaker reasoners).
Adaptive zone is model-dependent; floor must be model-independent.
Sharpens §7 floor-selection rule. Marked unvalidated; empirical
Sonnet check added as deferred register item 7.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
§12 records the integrated composition proof: 4-point thesis held in all 5
runs; notion-resolution worked without a formal Layer-2 mechanism (item 1,
one positive datapoint); advisory router sufficed (item 6, no trigger);
each known candidate hole surfaced-or-not with action (skill-TDD: findings,
no speculative REFACTOR); 3 locked-block recommendations listed not applied.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Integrated maximal-pressure scenario (route-work + story + verification
composed). 3 canonical runs + 1 maximal-adversarial run + 1 targeted
no-change probe, all isolated subagents. All 4 thesis points held every
run; zero observed RED, so skill-TDD warrants no REFACTOR (no speculative
patching). Locked AND non-locked SKILL.md spans byte-unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-layer composition model (issue-type spine + notion-resolved
tactics), advisory optional router, zero-footprint artifacts,
risk-proportional rigor with a fixed discipline floor. Supersedes
superpowers; adopts brainstorming/systematic-debugging as-is.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>