Build Instructions: A jq Interpreter
Objective
Build an interpreter for the jq language as described in sources/jq-manual.txt. Correctness is measured by the upstream jq conformance corpus, sources/jq.test, taken verbatim from jq 1.8.2. The goal is to pass every case the corpus supplies, with none failed and none errored. The suite's size is a property of the pinned corpus; never assert a case count.
The implementation language is Python, fixed by this Target's TECHNOLOGY_STACK.md and governed by stack/python.md.
jq is a small language with a large semantic core. Almost every filter is a generator: it takes one input and produces a stream of zero, one, or many outputs, and downstream filters run once per upstream output. Backtracking through that stream is not an optimisation, it is the evaluation model, and reduce, foreach, limit, first, label/break, and the ?// destructuring alternative are all defined in terms of it. An implementation that treats a filter as a function returning one value will pass the early cases and then stall permanently. Decide the evaluation model before writing builtins.
Run Harness
sources/full_test.sh is the single scoring entry point. It is supplied, not authored: it is staged verbatim into the build directory alongside the other imported assets, and drydock uat runs sh sources/full_test.sh from the completed application root and takes its exit code and output as the score. It reads:
#!/bin/sh
# full_test.sh — scoring entry point. Do not filter, skip, or reinterpret.
set -eu
if [ ! -x ./jq ]; then
echo "error: no executable ./jq at the application root." >&2
echo "The deliverable is an executable named jq that reads JSON on stdin." >&2
exit 1
fi
JQ="$PWD/jq" exec python3 sources/run_conformance.py
Before relying on any path above, run ls sources/ in the application directory and correct the paths in the harness against what is actually on disk. Correcting a path is the only edit permitted to this script. Do not add flags, filters, skips, or a redirection of the exit code.
The interface check is deliberately separate from the conformance run so that a missing program and a genuine conformance failure are distinguishable in the evidence. JQ is the harness's only knowledge of the implementation language; the harness itself is language-neutral.
Read-only scoring assets
These four files are the exam. They are hash-verified against the import and restored before grading, so a modification is reported as tampering rather than honoured:
sources/full_test.sh— the scoring entry pointsources/run_conformance.py— the scoring instrumentsources/exclusions.txt— the declared skipssources/jq.test— the conformance corpus
Do not write to them. Build ./jq so that the supplied entry point succeeds; changing the entry point is not a repair, and a repair pass spent editing one of these files is wasted.
Interface contract
The program is a filter: an executable file named jq at the application root, invoked as
./jq -c '<program>'
with JSON on stdin. It writes each value the program produces to stdout as one compact JSON value per line, and exits 0.
-c is the only option exercised. The program need not implement any other jq command-line option, and the manual's "Invoking jq" section is omitted from sources/jq-manual.txt for that reason.
Exit codes follow jq's own, and the distinction is load-bearing because the harness grades on it:
| Exit | Meaning |
|---|---|
0 | the program compiled and ran to completion |
3 | the program did not compile — a syntax or static error |
5 | the program compiled but raised at run time |
A case may legitimately emit several values and then raise; the harness compares the values produced before the raise, so exit 5 is not by itself a failure. Exit 3 on a valid program is always a failure. Diagnostics go to stderr and are never compared.
Any implementation shape that satisfies this contract is acceptable. A #!/usr/bin/env python3 script named jq that imports the real work from a package alongside it is the obvious one; main should parse arguments and delegate.
Test / verification process
The imported source files are placed in a sources/ subdirectory of the application directory. The only tools required are python3 and a POSIX sh, both already present. No installation step, no package download, and no network access are required at any point.
JQ="$PWD/jq" python3 sources/run_conformance.py # the scored run
JQ="$PWD/jq" python3 sources/run_conformance.py -v # list passing cases too
JQ="$PWD/jq" python3 sources/run_conformance.py --json # machine-readable
JQ="$PWD/jq" python3 sources/run_conformance.py --list # print cases, run nothing
JQ="$PWD/jq" python3 sources/run_conformance.py --select 'reduce' # one construct at a time
JQ="$PWD/jq" python3 sources/run_conformance.py --select 'reduce' --list # size that slice
During development, sh sources/full_test.sh does the interface check and the conformance run together, and is the same command the score is taken from.
--select takes a regular expression matched against the case's program text. It is how an intermediate story runs its own slice of the corpus, and it is required, not optional.
The harness's own --help text for --select calls it a development aid and states that the acceptance gate always runs the whole corpus. That text predates this rule and does not govern. It is part of a read-only scoring asset and cannot be corrected in place, so it is corrected here: where the harness's help text and this document disagree, this document is authoritative.
Exactly one acceptance check runs the whole corpus. One terminal story — the last one — runs sh sources/full_test.sh, asserts only result.returncode == 0, prints the captured stdout and stderr so a failure can be diagnosed from the evidence, and carries the Sea Trial. No other acceptance check may run the corpus unscoped, and none may invoke full_test.sh at all. A partial implementation fails most of the corpus by construction, so an unscoped mid-build run reports the schedule rather than a defect, and it is slow in exact proportion to how incomplete the code is.
Every other story runs its own slice, and the slice executes cases. Each story's acceptance invokes sources/run_conformance.py with a --select expression scoped to the construct that story implements, supplies JQ, and asserts result.returncode == 0. Those are exactly the cases the story's code is supposed to pass, so the run is neither slow nor red by construction, and the corpus goes green in the order the plan builds it.
--list prints the matching cases and runs nothing. Use it while planning, to size and inspect a selector before committing to it. An acceptance check that invokes it executes nothing, passes before the story's code exists, and steers no repair; such a check is a defect. The one exception is the story whose only obligation is that the corpus parses, the exclusion list applies, and the harness starts — it implements none of the behaviour under test, so listing is the correct thing for it to assert.
A selector matching no case is a defect too: it buys the story no coverage and reports success. Check the count with --list while planning and widen the expression until it covers the story's construct. Together the slices cover the corpus; a case no slice reaches is first executed by the terminal gate, where a failure arrives with the whole build already spent and no story to attribute it to.
No acceptance check may assert that an imported or staged file merely exists — a file-presence check is not acceptance.
The summary line is:
jq conformance: NNN passed, N failed, N errored, N skipped (corpus jq.test @ jq-1.8.2)
The harness reserves exit 2 for its own faults — a missing corpus, an unset JQ, a stale exclusion. Exit 2 never means the interpreter is wrong.
The corpus format
sources/jq.test documents its own format in its header. Cases are separated by blank lines; blank lines and # lines are ignored. A case is a program line, an input line, and then the expected output values, one per line. A case preceded by %%FAIL is a program that must be rejected at compile time: the following lines are upstream jq's diagnostic, which this harness records but never compares. Reproducing jq's exact error text is reverse-engineering a C implementation, not conforming to a specification, so a %%FAIL case passes on exit 3 alone.
Values are compared structurally, not textually. 1 and 1.0 are the same jq value; so are two objects whose keys are printed in a different order. Formatting of output is therefore not under test, but the number and order of values is.
Declared exclusions
sources/exclusions.txt names the corpus cases this kit cannot run, with the reason. They are the module-loader cases: import and include resolved against a search path of fixture files that this kit's flat source import cannot carry. They are reported as skipped and are not part of the score.
The module grammar cases are not excluded and must pass. module (.+1); 0, module []; 0, include "a" (.+1); 0, include "a" []; 0, include "\ "; 0, include "\(a)"; 0, and %::wat are all %%FAIL cases: the front end must parse the module syntax far enough to reject them, without ever touching the filesystem.
Source Roles
Record this table in the Analysis so every asset is staged onto disk in the build directory. sources/jq-manual.txt, sources/jq.test, sources/parser.y, and sources/lexer.l are large and must be readable from disk during implementation rather than carried in prompt text.
| Source | Role | Plan disposition | Build disposition |
|---|---|---|---|
jq-manual.txt | normative specification | context | stage |
jq.test | conformance test suite | context | stage |
parser.y | normative specification | context | stage |
lexer.l | normative specification | context | stage |
builtin.jq | reference implementation | context | stage |
run_conformance.py | conformance harness | context | stage |
full_test.sh | conformance harness | context | stage |
exclusions.txt | conformance harness | context | stage |
INSTRUCTIONS.md | author intent | context | prompt-only |
What each staged file is for:
sources/jq-manual.txt— the jq language manual at 1.8.2, rendered to plain text. The
primary specification, and the normative description of every builtin.
sources/jq.test— the conformance corpus. Also the most precise available statement of
the semantics, especially for generators and backtracking.
sources/parser.y— upstream's yacc grammar. The authority on operator precedence,
associativity, and the shape of every syntactic form.
sources/lexer.l— upstream's lexer. The authority on tokens, string interpolation, and
escape handling.
sources/builtin.jq— the subset of jq's builtins that upstream defines in jq itself.
Read it as a specification of those builtins' semantics.
sources/run_conformance.py,sources/full_test.sh,sources/exclusions.txt— the
scoring instruments, read-only as stated above.
Suggested implementation order
The difficulty is concentrated in one place — the evaluation model — and not spread evenly across the corpus. Build the core correctly before reaching for coverage.
- Lexer and parser. Follow
sources/lexer.landsources/parser.ydirectly. Produce
an AST. Precedence, ? suffixes, string interpolation, and the def forms are all settled there. Reject invalid programs with exit 3.
- The generator core. Evaluate a filter as something that yields a stream of values:
., literals, |, ,, field access, iteration, arithmetic, comparison, and empty. Every later feature is expressed in terms of this. Get [.[] | f], cartesian products over multi-output arguments, and short-circuiting right.
- Paths and assignment.
path(f),getpath,setpath,delpaths,del, and then
=, |=, +=, and friends. Assignment is defined over path expressions, so this cannot precede step 2.
- Control flow.
if/then/elif/else/end,try/catchand?,//,
reduce, foreach, label/break, limit, first, last, until, while, recurse. This is where backtracking is tested hardest.
- Functions, variables, and destructuring.
defwith arity and closures,as
bindings, object and array patterns, and the ?// alternative operator.
- Builtins. Work outward from
sources/builtin.jqand the manual: strings, arrays,
objects, sort_by/group_by/unique_by, @base64/@uri/@csv/@tsv/@sh formats, the date functions, tostream, input/inputs, $__loc__, debug.
- Numbers and edge cases.
nan,infinite, integer/float equality, large literals,
and the have_decnum builtin — return false from it and the corpus takes its non-decNumber branch, which native floats satisfy.
The manual is normative and the corpus is precise. Follow both directly rather than inferring behaviour from jq's printed output.
Definition of Done
sh sources/full_test.shruns cleanly with zero errors and exits zero.- The program satisfies the
./jq -c→ stdin → one-JSON-value-per-line → exit-code
contract.
- Every corpus case that runs passes: the failed and errored counts are both zero, and the
skipped count matches the declared exclusions. The harness exit status is the verdict and the whole verdict — assert returncode == 0 and stop there. Do not assert on the text of the summary line at all: the case totals belong to the pinned corpus, and a check that reads a runner's printed output is measuring the runner rather than the interpreter.
- Do not create acceptance checks asserting that imported or staged files merely exist.
- The interpreter is written from the specification. **Every third-party jq
implementation or binding is forbidden** — jq.py, pyjq, jqlang, gojq, jaq, and any other — as is shelling out to a system jq binary. A wrapper around real jq scores perfectly and makes the exercise meaningless.
- The project declares no third-party runtime dependency. The standard library is
sufficient: json, decimal, math, re, datetime, time, base64, unicodedata, itertools, functools, dataclasses, argparse, sys.
- No network access at any point, including at test time. No package is installed, and no
tool beyond python3 and POSIX sh is invoked.
- Deliver a concise project
README.mddocumenting the stdin/stdout interface, the exit
codes, and the sh sources/full_test.sh command.