Decision 060: Testing is planned and verified as two nodes; the plan is the record, and a skip is not a decision
- Status: proposed
- Date: 2026-08-06
- Deciders: @MadaraUchiha-314 (issue #163)
- Work item: issue-163
- Spec:
docs/specs/issue-163/ - Refines: decision-041 (the PDLC is an executable graph) and decision-014 (Gherkin scenario docstrings and contract-first APIs). Nothing in either is reversed; this adds the step that decides which kinds of testing a work item gets, and the step that proves they ran.
Context
Issue #163. the-loop can only claim a work item is done if it can verify it, and verification was the least declared part of the loop:
design.mdcarried a one-paragraph Testing strategy.tasks.mdcarried a_Test:_line per task.- The graph's
implementationnode ranverify-tests— a hook that is a no-op unless the node declares a command, and no shipped node declares one. - Everything after implementation (
self-review…reviewer-briefing) is about opinion, not about proof. Theevidencenode gates on an execution-log section, but the section is written by the same agent that decided what to test, at the moment it is writing the PR.
So which kinds of testing applied, whether they ran, and what the evidence was were all left to judgement at review time. That is the same shape as decision-045's defect (a gate reporting success without running) and issue-148's (prose describing a process the graph did not execute).
The ticket also names two things the-loop must not do: own the complexity of testing a multi-repository system, and pretend a single test command covers every work item.
Decision
Two nodes, one new artifact, no new runtime concepts.
design → test-planning → design-approval → tasks-breakdown → implementation → verification → self-reviewtest-planningproducestesting-plan.md, locked, carrying Test matrix, Verification environment, Evidence plan and Verification results.verificationre-declares the same artifact and gates oncheckmarks: completeplus a non-empty Verification results.
| Sub-decision | What was chosen | Why |
|---|---|---|
| D1 — plan before tasks | test-planning is locked before tasks-breakdown | Each task's _Test:_ names a matrix row. Planning after the DAG would reverse-engineer the plan from the tasks it is meant to constrain. |
| D2 — no approval node of its own; the design gate covers it | test-planning sits between design and design-approval, so the one human gate approves design.md and the plan derived from it, and feedback is recorded into both | Owner's call on PR #166, and the better answer than the original "no gate at all": the plan is significant enough to want human approval, but not a sixth stop. changes-requested returns to design, which re-derives the plan — the plan can never be approved against a design that moved under it. |
| D3 — the plan is the record | verification re-gates testing-plan.md rather than minting a verification-report.md | One artifact, one diff: a reviewer reads intent beside outcome. A second artifact would need its own template, manifest entry and parity coverage, and would duplicate both the plan and the execution log's Final validation evidence. It is the shape implementation already uses to re-gate tasks.md. |
| D4 — real phases | Both nodes carry a phase:, so both get loop: labels | A node the ticket cannot show is not a node in the PDLC. The nodes that share needs-review do so because they are review rounds on one state; test-planning and verification are distinct states. |
| D5 — catalogue, not enum | The testing types live in the bundled template and reference/testing.md; the schema gains nothing | An enum would have to be exhaustive to be useful, and adding "chaos testing" would become a schema migration. The gate checks the section exists and is non-empty; the reviewer judges the content — the same footing as "no new attack surface is written and justified". |
| D6 — declare, don't manage | The plan states repositories, services, fixtures, credentials-by-reference and the project's own commands | the-loop facilitates verification and owns no runner, orchestrator or environment manager. Where an operator has written the setup down, the plan links the registered customInstructions doc instead of restating it. |
The matrix rule
Rows are candidate testing types — unit, integration, contract, end-to-end, UI/visual, snapshot, performance, security/abuse-case, accessibility, migration/upgrade, manual exploratory, plus whatever the work item needs. No type is mandatory in itself; the matrix is work-item dependent. What is mandatory is the decision: a type that does not apply is n/a with a written reason. An unexplained blank is not a decision.
Evidence
Evidence is committed under docs/specs/<id>/evidence/. A link to a CI run that expires is not evidence. UI verification captures screenshots of each verified state, and an animated capture (GIF or equivalent) when the behaviour under test is a flow. Because the directory is as public as the repository, captures are redacted before they are committed; one that cannot be redacted is not committed, and the results row says so. Credentials appear in the plan by reference only — a literal secret in a committed plan is a leaked secret, to be rotated rather than edited out.
A skip is not a decision (the defect found while implementing this)
run_chain short-circuited on the first result that was not pass — including skip. Two consequences, one cause:
- Hooks after a skipping one never ran.
design's chain isvalidate-artifacts, enforces-boundaries-from, lint-artifacts, and the middle hook skips whenever the upstream artifact is absent — taking the lint gate down with it. - A chain ending in a skip routed on the outcome
"skip", for which no edge is declared.implementation's chain ends inverify-tests, which skips whenever no command is bound — soimplementationparked atno_edgeand escalated instead of advancing. Theimplementation → verificationedge this decision introduces would have been unreachable.
A hook that declines to run has said nothing about the node. The chain now continues past a skip and, if nothing objects, the node passes on the outcome pass. Blocking and waiting are unchanged, and NodeReport.satisfied already treated skip as satisfied — this makes the chain agree with it.
This is also why verification re-declares produces: validate-artifacts returns skipped for a node that declares no artifacts, so a verification node without it would have been a gate that reports success without ever running.
Alternatives considered
- Extend
design.md's Testing strategy instead of adding an artifact — cheapest, no new node, no new template. Rejected: the artifact has a second life atverification, and a file edited after implementation cannot also be a design artifact locked before it. The two would drift, and the gate would have nothing to re-read. - A
verification-report.mdproduced by the verification node — clean separation of plan from result. Rejected as D3: it splits intent from outcome across two files and two diffs, and buys a template, a manifest entry and parity coverage for the privilege. - Bind a project's test command to
verify-testson the verification node — makes the gate literally run the tests. Rejected for now: it re-introduces the-loop as a runner by the back door (whose command? which environment? which of eleven testing types?) and the hook is left in the chain as the declared seam for a future graph revision or a user-authored graph. - A
test-plan-approvalhuman node — consistent with requirements and design. Rejected as D2: the design gate covers the plan instead, which buys the same human approval without a sixth stop. - Produce
testing-plan.mdfrom thedesignnode itself (produces: [design.md, testing-plan.md], notest-planningnode at all) — the owner's literal suggestion on PR #166, and it would also put the plan under the design gate. Rejected in favour of keeping the node one step earlier in the chain: a node is what gives the plan aphase, aloop:test-planninglabel and a shape gate of its own, and folding two artifacts into one node means onevalidate-artifactscall whosesections:list cannot distinguish which file is missing what. Ordering the node before the gate gets the shared approval the owner asked for while keeping the state visible on the ticket — which is D4's whole argument. - Make specific testing types mandatory by risk tier (e.g. tier ≥ 4 requires performance testing) — rejected. It would force work items to run kinds of testing their change cannot exercise, and the-loop's existing answer to "prove you considered it" is a written justification, not a forced activity.
- Fold verification into
implementation— no new phase, no new label. Rejected: it is precisely the conflation issue-109 measured, where six distinct states hid behind oneneeds-reviewlabel and the drift piled up in the part nobody could see.