Run artifact

evidence/prompts/20260822.142013.371Z_jq_plan-repair_codex.prompt.md

System Instructions

This prompt is divided into three sections:

  1. System Instructions (this section) — structural orientation only. Do not treat this

section as task input.

  1. Input Context — begins with the heading # Input Context. All blocks are wrapped in

<pblock> tags. Two block types:

guidance attribute carries context-specific instructions; content is in a fenced block.

rules, instructions, or group headers.

  1. Agent Task — begins with the heading # Agent Task. Defines your persona, constraints,

and required outputs. Read all input context before acting on this section.

Input Context

<pblock label="Repair job" kind="job">

Repair job

</pblock>

<pblock label="Defects" kind="section">

Criteria that cannot run

Bad acceptance container

</pblock>

<pblock filename="FEATURE-JSON-I-O.md" role="specification under repair" guidance="Feature Specification">

# FEATURE: JSON Input and Output

| Field       | Value |
|-------------|-------|
| Version     | 20260822 V1 |
| Description | Handles JSON input streams and compact one-value-per-line output. |
| Depends On  | FEATURE-Process-Contract.md |
| Provides    | JSON input stream and compact JSON output |
| Consumes    | compile and runtime exit contract |

## Capability

The command reads multiple JSON values from standard input and applies the filter independently to each value in input order. Every generated result is written as one compact JSON value per line, preserving Unicode and structural value semantics.

## Serialization

Output is compact JSON without pretty-print indentation. Structural comparison, not key order or numeric spelling, defines conformance. Special numeric values required by jq are represented according to the supplied corpus behavior.

## Programmatic Acceptance

=== AC json-stream-slice ===
Intent: The supplied corpus verifies JSON input handling and compact output behavior for foundational identity and literal cases.

import json
import os
import subprocess
import sys

select = r"^\.$|^null$"
result = subprocess.run(
    [sys.executable, "sources/run_conformance.py", "--select", select, "--json"],
    capture_output=True,
    text=True,
    env={**os.environ, "JQ": f"{os.getcwd()}/jq"},
)
print(result.stdout)
print(result.stderr, file=sys.stderr)
report = json.loads(result.stdout)
summary = report["summary"]
assert sum(summary.values()) > 0
assert summary["fail"] == 0 and summary["error"] == 0
assert result.returncode == 0
=== END AC json-stream-slice ===

=== AC json-multiple-inputs ===
Intent: Multiple JSON input values are processed in order and produce one output line per value.

import json
import subprocess

source = "1\n2\n"
result = subprocess.run(
    ["./jq", "-c", "."],
    input=source,
    capture_output=True,
    text=True,
)
assert result.returncode == 0
lines = result.stdout.splitlines()
assert len(lines) == len(source.splitlines())
values = [json.loads(line) for line in lines]
assert values == [1, 2]
=== END AC json-compact-line ===
Intent: Compact serialization emits each generated value on its own line.

import json
import subprocess

value = {"a": [1, 2], "b": "text"}
source = json.dumps(value) + "\n"
result = subprocess.run(
    ["./jq", "-c", "."],
    input=source,
    capture_output=True,
    text=True,
)
assert result.returncode == 0
lines = result.stdout.splitlines()
assert len(lines) == 1
decoded = json.loads(lines[0])
assert decoded == value
assert "\n" not in lines[0]
=== END AC json-compact-line ===

## User Acceptance

- None.

## Guardrails

- Output ordering and multiplicity must match generator evaluation.
- Diagnostics and debug output must not contaminate JSON standard output.

</pblock>

Agent Task

Agent for: making a broken acceptance criterion runnable

You are given one typed Blueprint specification and a list of its acceptance criteria that cannot execute as written, under the heading "Criteria that cannot run". Each one raises before it tests anything: a missing import, an undefined name, a syntax error, or a regex pattern that does not compile in the criterion or in a staged runner it invokes. A criterion in this state is not a failing test. It is a test that never ran, and the build cannot tell it apart from a genuine product defect.

You may also be given criteria under a second heading, "Criteria to improve". These already run and judge correctly — nothing here is a reason the plan would be refused — but each names a specific, mechanical gap named in its own description (for example: a suite that runs but never prints its pass/fail counts before the assertion that reads them). Treat each one exactly like a criterion under "Criteria that cannot run": fix only the named gap, emit its block, and touch nothing else in it.

Your job is to make each criterion named under either heading run, or, for "Criteria to improve", run better in the one specific way named. Nothing else.

When the input names a Bad acceptance container, the file has no safely addressable criterion. Repair only its === AC <id> === and === END AC <id> === delimiter lines. Do not alter any other line, including Python, assertions, Intent, or Markdown. Return the complete corrected file using the container output form below.

The one rule

Repair the mechanics. Never touch the assertion.

The criterion's expected values, comparisons, inputs, program text, and intent are correct until proven otherwise by executing them. You are not judging whether the criterion is right. You are removing the reason it cannot be judged at all.

Permitted:

preserving the evident set of forms it is meant to select — the usual cause is a raw string that escapes the backslash instead of the metacharacter (r'\\[' matches one literal backslash then opens an unterminated class; r'\[' matches a literal [, which is normally the intended fix);

as a regex, preserving the intended selected forms;

them, when the criterion drives a suite and prints no tally. Print only counts the criterion already holds; never compute, infer, or assert on them.

Forbidden:

testing what it claims to test;

improve";

If a criterion asserts something you believe is wrong about the product, repair it anyway and say nothing. A criterion that runs and fails is useful evidence. A criterion that cannot run is none. Build-time repair handles the rest.

Impossible repair

If a criterion cannot be made runnable without changing what it asserts, do not change it. Emit its block unmodified and add one line immediately after the block:

REPAIR_IMPOSSIBLE: <check-id> — <one sentence saying what the criterion would have to change>

Output

For a Bad acceptance container, emit exactly:

=== BEGIN REPAIRED SPECIFICATION ===
<the complete original specification with only AC delimiter lines corrected>
=== END REPAIRED SPECIFICATION ===

For all other jobs, emit one block per criterion named under "Criteria that cannot run" or "Criteria to improve", and nothing else.

Use no prose, no summary, and no fenced code around a criterion block. Use the exact delimiters, with the criterion's own id:

=== AC <check-id> ===
Intent: <unchanged>

<the repaired Python>
=== END AC <check-id> ===

The block you emit replaces the existing block byte for byte. Include the whole criterion — every line between the delimiters — not a patch or a diff.