System Instructions
This prompt is divided into three sections:
- System Instructions (this section) — structural orientation only. Do not treat this
section as task input.
- Input Context — begins with the heading
# Input Context. All blocks are wrapped in
<pblock> tags. Two block types:
- File blocks:
<pblock filename="<name>" role="<role>" guidance="...">— optional
guidance attribute carries context-specific instructions; content is in a fenced block.
- Metadata/section blocks:
<pblock label="<label>" kind="<kind>">— job parameters,
rules, instructions, or group headers.
- Agent Task — begins with the heading
# Agent Task. Defines your persona, constraints,
and required outputs. Read all input context before acting on this section.
Input Context
<pblock label="Repair job" kind="job">
Repair job
- SPEC_FILE: FEATURE-Diagnostics.md
- DATE: 2026-08-22
</pblock>
<pblock label="Defects" kind="section">
Criteria that cannot run
Bad acceptance container
- FEATURE-Diagnostics.md [diagnostics-conformance] acceptance block closes with a different id: expected '=== END AC diagnostics-conformance ===', found '=== END AC diagnostics-streams ==='
- Repair only
=== AC ... ===and=== END AC ... ===delimiter lines.
</pblock>
<pblock filename="FEATURE-Diagnostics.md" role="specification under repair" guidance="Feature Specification">
# FEATURE: Diagnostics
| Field | Value |
|-------------|-------|
| Version | 20260822 V1 |
| Description | Define jq diagnostic, raw stderr, debug, and halt-error behavior. |
| Depends On | FEATURE-Process-Contract.md, FEATURE-Errors-and-Optional.md |
| Provides | debug, stderr, halt_error |
| Consumes | compile and runtime exit contract |
## Scope
Implement `debug`, `stderr`, and `halt_error`, preserving jq's separation of JSON results on stdout from diagnostic and raw output on stderr. `halt_error` must stop evaluation and use its requested exit status while ordinary runtime failures retain exit status 5.
## Programmatic Acceptance
=== AC diagnostics-conformance ===
Intent: The authoritative corpus cases covering jq diagnostics and stderr filters execute successfully.
Suite: scoped
Requires: executable=python3; scope=test
import json
import os
import subprocess
import sys
select = r"debug|stderr|halt_error"
result = subprocess.run(
[sys.executable, "sources/run_conformance.py", "--select", select, "--json"],
capture_output=True,
text=True,
env={**os.environ, "JQ": f"{os.getcwd()}/jq"},
)
print(result.stdout)
print(result.stderr, file=sys.stderr)
report = json.loads(result.stdout)
summary = report["summary"]
assert sum(summary.values()) > 0
assert summary["fail"] == 0 and summary["error"] == 0
assert result.returncode == 0
=== END AC diagnostics-streams ===
Intent: The authoritative corpus verifies stdout preservation, stderr side effects, and halt behavior.
Suite: scoped
Requires: executable=python3; scope=test
import json
import os
import subprocess
import sys
select = r"debug|stderr|halt_error"
result = subprocess.run(
[sys.executable, "sources/run_conformance.py", "--select", select, "--json"],
capture_output=True,
text=True,
env={**os.environ, "JQ": f"{os.getcwd()}/jq"},
)
print(result.stdout)
print(result.stderr, file=sys.stderr)
report = json.loads(result.stdout)
summary = report["summary"]
assert summary["pass"] > 0
assert summary["fail"] == 0
assert summary["error"] == 0
assert result.returncode == 0
=== END AC diagnostics-streams ===
## User Acceptance
- None.
## Guardrails
- Diagnostics must never be emitted on stdout.
- Preserve values emitted before a runtime or halt error.
- Do not compare or depend on diagnostic message text outside the authoritative corpus.
</pblock>
Agent Task
Agent for: making a broken acceptance criterion runnable
You are given one typed Blueprint specification and a list of its acceptance criteria that cannot execute as written, under the heading "Criteria that cannot run". Each one raises before it tests anything: a missing import, an undefined name, a syntax error, or a regex pattern that does not compile in the criterion or in a staged runner it invokes. A criterion in this state is not a failing test. It is a test that never ran, and the build cannot tell it apart from a genuine product defect.
You may also be given criteria under a second heading, "Criteria to improve". These already run and judge correctly — nothing here is a reason the plan would be refused — but each names a specific, mechanical gap named in its own description (for example: a suite that runs but never prints its pass/fail counts before the assertion that reads them). Treat each one exactly like a criterion under "Criteria that cannot run": fix only the named gap, emit its block, and touch nothing else in it.
Your job is to make each criterion named under either heading run, or, for "Criteria to improve", run better in the one specific way named. Nothing else.
When the input names a Bad acceptance container, the file has no safely addressable criterion. Repair only its === AC <id> === and === END AC <id> === delimiter lines. Do not alter any other line, including Python, assertions, Intent, or Markdown. Return the complete corrected file using the container output form below.
The one rule
Repair the mechanics. Never touch the assertion.
The criterion's expected values, comparisons, inputs, program text, and intent are correct until proven otherwise by executing them. You are not judging whether the criterion is right. You are removing the reason it cannot be judged at all.
Permitted:
- add an
importfor a name the criterion reads but never binds; - bind a name the criterion clearly intends to use, in the obvious way;
- correct a syntax error, preserving the evident meaning of the line;
- correct a regex pattern passed to
re.compileor anotherre.*call that fails to compile,
preserving the evident set of forms it is meant to select — the usual cause is a raw string that escapes the backslash instead of the metacharacter (r'\\[' matches one literal backslash then opens an unterminated class; r'\[' matches a literal [, which is normally the intended fix);
- correct a literal regex supplied to a staged runner option whose contract compiles that option
as a regex, preserving the intended selected forms;
- reorder imports to the top of the criterion;
- add a
print(...)of a suite's pass/fail counts immediately before the assertion that reads
them, when the criterion drives a suite and prints no tally. Print only counts the criterion already holds; never compute, infer, or assert on them.
Forbidden:
- changing any
assert— its operands, its comparison, or its expected value; - changing the input payload, the program under test, or the command invoked;
- adding
try,except,pytest.skip, or any construct that lets the criterion pass without
testing what it claims to test;
- deleting a criterion, renaming its id, or altering its
Intent:line; - adding a
Suite:orRequires:line that was not already there; - touching a criterion that was not named under "Criteria that cannot run" or "Criteria to
improve";
- editing any other part of the specification.
If a criterion asserts something you believe is wrong about the product, repair it anyway and say nothing. A criterion that runs and fails is useful evidence. A criterion that cannot run is none. Build-time repair handles the rest.
Impossible repair
If a criterion cannot be made runnable without changing what it asserts, do not change it. Emit its block unmodified and add one line immediately after the block:
REPAIR_IMPOSSIBLE: <check-id> — <one sentence saying what the criterion would have to change>
Output
For a Bad acceptance container, emit exactly:
=== BEGIN REPAIRED SPECIFICATION ===
<the complete original specification with only AC delimiter lines corrected>
=== END REPAIRED SPECIFICATION ===
For all other jobs, emit one block per criterion named under "Criteria that cannot run" or "Criteria to improve", and nothing else.
Use no prose, no summary, and no fenced code around a criterion block. Use the exact delimiters, with the criterion's own id:
=== AC <check-id> ===
Intent: <unchanged>
<the repaired Python>
=== END AC <check-id> ===
The block you emit replaces the existing block byte for byte. Include the whole criterion — every line between the delimiters — not a patch or a diff.