# System Instructions

This prompt is divided into three sections:

1. **System Instructions** (this section) — structural orientation only. Do not treat this
   section as task input.

2. **Input Context** — begins with the heading `# Input Context`. All blocks are wrapped in
   `<pblock>` tags. Two block types:
   - File blocks: `<pblock filename="<name>" role="<role>" guidance="...">` — optional
     `guidance` attribute carries context-specific instructions; content is in a fenced block.
   - Metadata/section blocks: `<pblock label="<label>" kind="<kind>">` — job parameters,
     rules, instructions, or group headers.

3. **Agent Task** — begins with the heading `# Agent Task`. Defines your persona, constraints,
   and required outputs. Read all input context before acting on this section.


# Input Context

<pblock label="Build block job" kind="job">
## Build block job
- TARGET: jq
- BUILD_DIRECTORY: /mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/build/jq
- WORKING_DIRECTORY: /mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/build/jq
- BUILD_BLOCK: Block 1 · Foundational (block-1)
- STORIES: Define the standalone interpreter architecture and module boundaries. (architecture)
- DATE: 2026-08-22
- BUILD_SCOPE: exactly one MANIFEST.md build block

</pblock>

<pblock label="Stories in this block" kind="section">
## Stories in this block
- Define the standalone interpreter architecture and module boundaries. (architecture) [story]

</pblock>

<pblock label="Files on disk" kind="section">
## Files on disk in the build directory
- sources/builtin.jq
- sources/exclusions.txt
- sources/full_test.sh
- sources/jq-manual.txt
- sources/jq.test
- sources/lexer.l
- sources/parser.y
- sources/run_conformance.py

These are the only imported files present on disk. Every other file named in this prompt is supplied as prompt context and is not on disk; read it here and do not report it as a missing input.

</pblock>


## COMPASS - Target Orientation

<pblock filename="COMPASS.md" role="compass" path="/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/COMPASS.md" guidance="Important: This is the core project intent, constraints, and guardrails. It should have presecedence in conflicts.">
````markdown
# COMPASS: jq

## Compass

Build a standalone interpreter for the jq language as described in `sources/jq-manual.txt`. The
product is an executable file named `jq` at the application root. It reads JSON from standard
input, evaluates a jq filter as an ordered generator, and writes each value the filter produces to
standard output as one compact JSON value per line.

Correctness is measured by the upstream jq conformance corpus `sources/jq.test`, taken verbatim
from jq 1.8.2, minus the cases named in `sources/exclusions.txt`. The goal is every case passing,
none failed and none errored.

## Constraints

- Implement in Python using only the standard library.
- Provide an executable named `jq` at the application root, invoked as `./jq -c '<program>'`.
- `-c` is the only option exercised. No other command-line option is required.
- Run without network access, package installation, or external runtime dependencies.
- Exit `0` when the program compiled and ran to completion, `3` when it did not compile, and `5`
  when it compiled and raised at run time. The harness grades on this distinction.
- Diagnostics go to standard error and are never compared.

## Guardrails

- Do not shell out to a system `jq` executable.
- Do not use a third-party jq implementation or binding.
- Do not modify, rewrite, trim, regenerate, or substitute any file under `sources/`. Those assets
  are restored before grading and an edit is reported as tampering.
- Preserve generator ordering, multiplicity, backtracking, and partial-output runtime behavior.
- Keep compile failures distinct from runtime failures using exit codes 3 and 5.

## Verification Protocol

This section is normative. It governs which story may invoke the supplied harness, and how.

### Invoking the harness

`sources/run_conformance.py` **requires** the environment variable `JQ`, the command that runs the
candidate implementation. Without it the harness exits `2` on its own usage code, which is a
harness fault and never a verdict about the interpreter. Every invocation, in every acceptance
criterion and every developer command, supplies it:

```bash
JQ="$PWD/jq" python3 sources/run_conformance.py            # whole corpus, the scored run
JQ="$PWD/jq" python3 sources/run_conformance.py --select 'reduce'   # run one construct for real
```

Those two commands are the only ways this build runs the harness. They are specified verbatim
below under *The two harness invocations*, together with the flag this build forbids.

`sources/` is read only. No story edits, patches, or regenerates `sources/run_conformance.py`,
`sources/jq.test`, or `sources/exclusions.txt`; a harness defect is reported, not repaired in
place. A story that needs to experiment with the harness works on a copy outside `sources/`, and
every acceptance criterion invokes the original `sources/run_conformance.py`.

An acceptance criterion written in Python supplies it by **extending** the inherited environment,
never by replacing it:

```python
env={**os.environ, "JQ": str(build_dir / "jq")}
```

`env={"JQ": ...}` alone leaves the child with no `PATH`, so nothing it invokes resolves and the
criterion is false at every level of implementation quality.

`sources/full_test.sh` sets `JQ` itself for the runner it wraps and therefore takes no environment
from its caller.

The harness reserves exit `2` for its own faults — a missing corpus, an unset `JQ`, a stale
exclusion list. Exit `2` never means the interpreter is wrong.

The summary line is:

```
jq conformance: NNN passed, N failed, N errored, N skipped (corpus jq.test @ jq-1.8.2)
```

### The two harness invocations

An acceptance criterion that runs `sources/run_conformance.py` uses one of these two commands. No
criterion in this build passes any other flag to the harness.

| Story kind | Command | Executes cases? | Asserts |
|---|---|---|---|
| Every behavioral story | `--select <regex> --json` | Yes, the selected slice | exit `0`, zero `fail`, zero `error`, non-zero case count |
| Terminal story (once, last) | `sh sources/full_test.sh` | Yes, all of them | exit `0` |

The staging story does not appear in this table. It does not run the harness at all; see *The
staging story* below.

#### `--list` is never run

`sources/run_conformance.py` accepts a flag, spelled `--list`, that prints the names of the
matching cases and then exits without executing any of them. `sources/INSTRUCTIONS.md`, the file
header, and `--help` all document it.

**This build never runs it. Not in an acceptance criterion, not in a story, not in a script, not
in a command typed by a build agent, not while developing and not while verifying. The string
`--list` does not appear anywhere in this project's output. If you have written it, that line is
wrong — delete it and use one of the two commands above.**

A Drydock build is headless. There is no one watching the output, so a mode whose entire purpose
is to print something for a person to read has no reader and no reason to run.

The flag returns `0` at the top of the run — before the harness reads `JQ`, before it resolves the
candidate command, before it executes a single case. A criterion built on it passes when `jq` is
an empty file, when `jq` does not exist, and when the story it gates was never written. It is not
a weak proof, not a partial proof, and not an acceptable proof for staging, for scaffolding, or
for an early story whose implementation is incomplete. It is not a proof. Thirty-six criteria in
one earlier plan of this project used it, every one of them reported green, and it cost three days.

If you are writing a criterion and reaching for that flag, the reason is always the same: the
story's code does not exist yet and you want a command that will not fail. That is the definition
of a criterion that proves nothing. Write the `--select ... --json` form instead and let it be red
until the story makes it green. **A criterion is supposed to fail before its story is built.**

The same prohibition covers any other flag whose effect is to not execute the cases — enumeration,
dry-run, validation, or help. If a flag's documented purpose is "run nothing", it has no place in
an acceptance criterion.

#### Behavioral criterion — copy this, changing only `SELECT`

```python
import json
import os
import subprocess
import sys

SELECT = r"reduce"

result = subprocess.run(
    [sys.executable, "sources/run_conformance.py", "--select", SELECT, "--json"],
    capture_output=True,
    text=True,
    env={**os.environ, "JQ": f"{os.getcwd()}/jq"},
)
print(result.stdout)
print(result.stderr, file=sys.stderr)
report = json.loads(result.stdout)
tally = report["summary"]
assert sum(tally.values()) > 0, f"selector matched no case: {SELECT}"
assert tally["fail"] == 0 and tally["error"] == 0, tally
assert result.returncode == 0, result.returncode
```

Three assertions, and all three are required.

1. **The selector matched something.** `--select` is a regular expression matched against the
   program text of each case. A selector that matches nothing yields zero cases, zero failures,
   and exit `0` — green, and worth nothing. Alternations naming ideas rather than syntax
   (`closure`, `recursive`, `optional`) match no jq program and are the common way to write one
   by accident. Select on syntax the corpus actually contains: `reduce`, `foreach`, `def `,
   ` as \$`, `try `, `//`, `path(`.
2. **No case failed or errored.** Read off the parsed JSON tally, not off any printed line.
3. **The exit status is `0`.** The harness returns `0` only when `fail` and `error` are both zero,
   and reserves `2` for its own faults — a missing corpus, an unset `JQ`, a stale exclusion list.
   Exit `2` is never a verdict about the interpreter.

`--json` writes the report and nothing else to stdout, so `json.loads(result.stdout)` is total. Do
not assert against the human summary line, and do not grep stdout for `passed` or `failed`.

### The terminal story

The **terminal story** is the last story in the build order: the one on which every other story is
a transitive dependency, and after which no further story runs. It is a verification story. Its
job is not to add capability but to prove that the capability every preceding story delivered is
present, together, at the end of the build.

The terminal story of this project runs `sh sources/full_test.sh`, asserts `returncode == 0`,
prints the captured stdout and stderr so a failure is diagnosable from the evidence alone, and
carries the Sea Trial. It is the only story permitted to run the whole corpus.

A story is not terminal because its name contains "verify", because it is a test harness, or
because it stages the test assets. Staging the corpus is foundational work that happens early;
running the corpus is terminal work that happens last. Do not place a whole-corpus gate on a
story that cannot yet run it — it fails vacuously and teaches nothing.

### Scope of every other story

Every non-terminal story is gated on its own declared behavior only, through `--select` against
the constructs that story implements, and the criterion asserts the selected slice passes. A
non-terminal story never invokes `sources/full_test.sh` and never runs the corpus unfiltered: a
partial interpreter fails most of an authoritative corpus by construction, and its unimplemented
cases exhaust the harness's per-case timeout rather than returning, so the unscoped run costs the
most exactly where it teaches the least.

Regression across stories is not the responsibility of any story's criteria. Drydock re-runs every
previously proven criterion after each block and attributes a criterion that was green and is now
red to the block that broke it, so a criterion proven at story 2 and broken at story 6 fails story
6. Do not author a mid-build story whose purpose is to re-run earlier stories' checks.

### The staging story

The story that stages the conformance assets is gated on the assets being present, complete, and
mutually consistent — not on a bare file-existence assertion, and not on the corpus running. It
proves that in process, by importing the harness and calling its parsers directly. It never
launches the harness, so the question of which flags to pass does not arise:

```python
import sys

sys.path.insert(0, "sources")
import run_conformance as harness

EXPECTED_CASES = 550
EXPECTED_EXCLUSIONS = 13

cases = harness.parse_corpus(harness.CORPUS.read_text(encoding="utf-8"))
excluded = harness.apply_exclusions(cases, harness.parse_exclusions(harness.EXCLUSIONS))
assert len(cases) == EXPECTED_CASES, len(cases)
assert len(excluded) == EXPECTED_EXCLUSIONS, len(excluded)
```

This reads state rather than output: the harness module imports, the corpus parses into the
expected number of cases, and every exclusion still matches a case — `apply_exclusions` raises on
a stale entry, so a corpus and an exclusion list that have drifted apart fail here rather than
silently skipping cases later.

It claims nothing about the interpreter, because at this point in the build there is nothing to
claim. Every story that claims a construct works runs that construct through
`--select ... --json`.

## The corpus

`sources/jq.test` documents its own format in its header. Cases are separated by blank lines;
blank lines and `#` lines are ignored. A case is a program line, an input line, and then the
expected output values, one per line. A case preceded by `%%FAIL` is a program that must be
rejected at compile time; the following lines are upstream jq's diagnostic, which the harness
records but never compares. Reproducing jq's exact error text is reverse-engineering a C
implementation rather than conforming to a specification, so a `%%FAIL` case passes on exit `3`
alone.

Values are compared structurally, not textually. `1` and `1.0` are the same jq value, and so are
two objects whose keys print in a different order. Output formatting is therefore not under test,
but the number and order of values is.

`sources/exclusions.txt` names the corpus cases this kit cannot run, with the reason. They are the
module-loader cases, whose `import` and `include` resolve against a search path of fixture files a
flat source import cannot carry. They are reported as `skipped` and are not part of the score.

The module *grammar* cases are not excluded and must pass. `module (.+1); 0`, `module []; 0`,
`include "a" (.+1); 0`, `include "a" []; 0`, `include "\ "; 0`, `include "\(a)"; 0`, and `%::wat`
are all `%%FAIL` cases: the front end parses the module syntax far enough to reject them, without
ever touching the filesystem.

<!-- drydock:build-write-guardrail:start -->
## Build Write Guardrail

- Authorized build directory: `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/build/jq`
- Authorized Target directory: `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq`
- Build agents have permission to create, modify, and remove files required by the active build block inside these authorized directories.
- No path outside these authorized directories may be modified.
- Protected Drydock artifacts:
  - `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/blueprint/`
  - `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/MANIFEST.md`
  - `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/COMPASS.md`
  - `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/QuarterDeck/`
  - `/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/evidence/`
<!-- drydock:build-write-guardrail:end -->
````
</pblock>


## STACK - Technology HOW

<pblock filename="common.md" role="stack" path="/mnt/c/Users/barlo/projects/drydock/Rigging/stack/common.md">
````
# Common Best Practices

**Version:** 20260320 V1  
**Category:** Technologies
**Description:** Common development patterns shared across all stack configurations

Always included regardless of technology stack. Covers project structure conventions, shell scripts, metadata files, git hygiene, and development workflow. This file does not change between projects.

---

## 1. Project Directory Layout

**Rule**: Every project follows a predictable directory structure.

```
project-name/
├── bin/                # Operation scripts (see Shell Scripts section)
├── data/               # Runtime data (DB, logs, backups) — gitignored
│   ├── logs/           # Script and process output logs
│   └── backups/        # Database and file backups
├── docs/               # Project documentation
├── tests/              # Test suite
├── PROJECT/            # Build specification (optional, for specification-driven projects)
├── .env                # Environment config — gitignored
├── .env.example        # Template with placeholder values — committed
├── .gitignore
├── CLAUDE.md           # AI agent instructions
└── Links.md            # External links
```

Additional directories depend on the stack (e.g., `templates/`, `static/`, `migrations/`).

**Why**: Consistent layout lets any developer or AI agent locate files instantly.

---

## 2. Shell Scripts (bin/ Directory)

**Rule**: All user-facing operations live in `bin/` as bash scripts with standardized headers, logging, and error handling.

### Script Template

```bash
#!/bin/bash
# CommandCenter Operation
# Name: Human Readable Name
# Type: daemon|batch
# Port: 8000

# --- Standard Preamble ---
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PROJECT_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
LOG_DIR="$PROJECT_DIR/data/logs"
mkdir -p "$LOG_DIR"

TIMESTAMP=$(date '+%Y-%m-%d_%H%M%S')
SCRIPT_NAME=$(basename "$0" .sh)
LOG_FILE="$LOG_DIR/${SCRIPT_NAME}_${TIMESTAMP}.log"

echo "=== $SCRIPT_NAME started at $(date '+%Y-%m-%d %H:%M:%S') ===" | tee "$LOG_FILE"
echo "Arguments: $*" | tee -a "$LOG_FILE"
echo "Working dir: $PROJECT_DIR" | tee -a "$LOG_FILE"
echo "---" | tee -a "$LOG_FILE"

cd "$PROJECT_DIR"

# --- Your Commands Here ---
# All output goes to both console and log file via tee
your_command 2>&1 | tee -a "$LOG_FILE"

echo "=== $SCRIPT_NAME finished at $(date '+%Y-%m-%d %H:%M:%S') ===" | tee -a "$LOG_FILE"
```

### Header Fields

The first comment block is parsed by Command Center's scanner for auto-discovery:

| Field | Required | Values | Description |
|-------|----------|--------|-------------|
| `# CommandCenter Operation` | Yes | literal | Marks script as discoverable |
| `# Name:` | Yes | free text | Display name in UI |
| `# Type:` | No | `daemon` or `batch` | Default: `batch`. Daemons stay running. |
| `# Port:` | No | integer | Port number for daemon services |

Scripts without the `# CommandCenter Operation` header are still valid scripts but won't appear in Command Center's UI.

### Standard Scripts

Every project should have these scripts (where applicable):

| Script | Type | Purpose |
|--------|------|---------|
| `bin/start.sh` | daemon | Start the dev server |
| `bin/stop.sh` | batch | Stop the dev server |
| `bin/test.sh` | batch | Run test suite |
| `bin/build.sh` | batch | Build/compile the project |
| `bin/deploy.sh` | batch | Deploy to production |
| `bin/backup.sh` | batch | Backup data/database |

### Logging Pattern

- All stdout and stderr captured via `tee` to `data/logs/`
- Log filename includes script name and timestamp: `start_2026-03-07_143022.log`
- First lines of log always record: timestamp, arguments, working directory
- These logs are viewable from Command Center's Process Monitor

**Why**: Standardized scripts make every project operable the same way. Log capture enables monitoring, alerting, and post-mortem analysis.

---

## 3. External Links (Links.md)

**Rule**: Every project maintains a `Links.md` file at its root with a markdown table of relevant URLs.

```markdown
| Label | URL |
|-------|-----|
| Local Dev | http://localhost:5001 |
| Production | https://example.com |
| Docs | https://docs.example.com |
| GitHub | https://github.com/user/repo |
```

Rules:
- One table, two columns: Label and URL
- Labels are short and descriptive
- Command Center's scanner reads this on startup and stores links in the project's `extra` JSON
- Links appear in the project's configuration page

**Why**: Centralizes all project URLs in one discoverable, parseable location. Both AI agents and Command Center consume it.

---

## 4. CLAUDE.md Convention

**Rule**: Every project has a `CLAUDE.md` at its root following this section structure:

1. `## Project Overview` — What the project does, key features
2. `## Architecture` — Tech stack, key files, patterns
3. `## Dev Commands` — Bash commands to run the project (in a code block)
4. `## Service Endpoints` — URLs: `- Label: https://url`
5. `## Bookmarks` — Grouped links: `### Group` then `- [Title](URL)`

Section rename rules — always use the standard name:
- `## Commands` / `## Development Commands` / `## Build Commands` → `## Dev Commands`
- `## Overview` / `## Project Purpose` → `## Project Overview`
- `## Stack` → `## Architecture`

**Why**: Consistent structure lets AI agents parse project context reliably.

---

## 5. Git Hygiene

**Rule**: Maintain a comprehensive `.gitignore`. Never commit secrets, generated files, or runtime data.

```gitignore
# Runtime
data/
*.db
*.log

# Environment
.env
venv/
node_modules/

# Python
__pycache__/
*.pyc
*.egg-info/
dist/
build/

# OS
.DS_Store
Thumbs.db
```

Rules:
- `data/` — runtime databases, logs, backups, uploads
- `.env` — secrets and local config
- Commit `.env.example` with placeholder values
- Write imperative commit messages: "Add health endpoint" not "Added health endpoint"

**Why**: Clean repos are cloneable and runnable. No secrets in history.

---

## 6. Development Workflow

**Rule**: Follow these rules when working in any git-managed project.

1. **Always commit changes immediately** after completing a task if the task has no errors
2. **Commit messages** should have descriptive text (no AI/tool mentions)
3. **DO NOT push** — only commit to local git
4. **NO co-authored-by lines** in commits
5. **Always end code change responses with a restart notice** for any project that runs a web server:
   - If only templates/CSS/static files changed: "No restart needed — browser refresh is enough."
   - If any Python/JS server files changed: "Restart required — run the start script or equivalent."

**Why**: Consistent workflow prevents accidental pushes, keeps commit history clean, and ensures developers know when to restart.

---

## Summary Checklist

- [ ] Standard directory layout with `bin/`, `data/`, `docs/`, `tests/`
- [ ] Shell scripts in `bin/` with CommandCenter headers and tee logging
- [ ] `Links.md` for external URLs
- [ ] `CLAUDE.md` following section convention
- [ ] `.env.example` committed, `.env` gitignored
- [ ] Comprehensive `.gitignore`
- [ ] Commit immediately, don't push, no AI mentions
````
</pblock>

<pblock filename="python.md" role="stack" path="/mnt/c/Users/barlo/projects/drydock/Rigging/stack/python.md">
````
# Python Best Practices

**Version:** 20260716 V3  
**Category:** Technologies
**Description:** Python language conventions and patterns for specification-driven projects

Technology reference for Python development. Framework-agnostic — applies to any Python project. This file does not change between projects.

Prerequisite: `stack/common.md`

---

## 1. Configuration Management

**Rule**: All environment access goes through one typed `Config` class (see `stack/persistence.md`) — never read `os.environ` elsewhere. Every field is inherited from the environment; there are no `Dev`/`Prod`/`Test` subclasses. Never hardcode secrets, ports, or paths.

```python
# config.py
import os
from dataclasses import dataclass
from dotenv import load_dotenv
load_dotenv()

@dataclass(frozen=True)
class Config:
    secret_key: str
    database_path: str
    port: int
    debug: bool = False

    @classmethod
    def load(cls) -> "Config":
        try:
            return cls(
                secret_key=os.environ["SECRET_KEY"],
                database_path=os.environ.get("DATABASE_PATH", "data/app.db"),
                port=int(os.environ.get("APP_PORT", "5001")),
                debug=os.environ.get("APP_DEBUG") == "1",
            )
        except KeyError as e:
            raise RuntimeError(f"Missing required env var: {e}") from e
```

Commit `.env.example` listing every variable `Config` reads, with dummy values. The build creates the local `.env` from it; `.env` itself is never committed.

```bash
# .env.example — committed; every variable Config reads, dummy values only
SECRET_KEY=change-me
DATABASE_PATH=data/app.db
APP_PORT=5001
APP_DEBUG=0
```

Rules:
- `python-dotenv` is a runtime dependency in `pyproject.toml` — not a dev dependency. Nothing in the toolchain loads `.env` on its own: `python run.py`, `uv run`, and `pytest` all leave the environment untouched. The application loads its own configuration or it does not start.
- `load_dotenv()` runs at import of `config.py`, before any `Config.load()` call, so every entry point — `run.py`, `bin/start.sh`, a management command, a worker — reads the same `.env`.
- `.env.example` is committed and lists every variable `Config` reads, no more and no fewer. A variable missing from it never reaches the running application.
- The startup error names the missing variable (`Missing required env var: 'SECRET_KEY'`). A generic message such as `Invalid application configuration` tells the operator nothing and is a defect.
- The entry point must start on a clean shell with nothing exported by hand: `.env` plus the defaults in `Config` are the whole configuration.

**Why**: A single typed `Config` is the only env reader. Typed fields crash on a missing or malformed variable at startup, not at first use. The environment (`.env`) selects configuration — not a Python subclass. A `Config` that reads `os.environ` without loading `.env` passes every test that constructs the app with overrides, then fails on the operator's first real run.

---

## 2. Code Style and Understandability

**Rule**: Code must be understandable on its own through naming, structure, small focused units, explicit types, clear interfaces, appropriate abstractions, and tests. If a reader needs a comment to follow the mechanics, improve the code instead.

Rules:
- **Naming** — names state intent: `load_active_users()`, not `get_data()`; `retry_limit`, not `n`. No abbreviations a new reader must decode.
- **Structure** — modules have one responsibility each (see §10 layout); related code lives together; call depth stays shallow.
- **Small focused units** — functions do one thing at one level of abstraction; a function that needs a section comment ("# now validate…") is two functions.
- **Explicit types** — public interfaces are fully typed (§3); the signature answers "what goes in, what comes out" without reading the body.
- **Clear interfaces** — few parameters, typed returns, no boolean flags that change what a function fundamentally does, no output-by-mutation surprises.
- **Appropriate abstractions** — introduce a layer only to remove real duplication or isolate a boundary (DB, cloud, external API). No speculative generality.
- **Tests** — tests are the executable specification of behavior (§6); a behavior worth keeping is a behavior worth a test.
- Comments state constraints the code cannot express (invariants, external quirks, why-not-the-obvious-way) — never restate what the next line does.

**Why**: Code is read far more often than written. Every hour invested in clarity is repaid at each future read, debug, and review — including by the author six months later.

---

## 3. Type Hints and Static Typing

**Rule**: Use modern type hints on all public interfaces. Model data that crosses module or process boundaries with typed structures — dataclasses, `TypedDict`, or Pydantic models — never bare dicts or tuples of implicit shape. Run a static type checker when practical.

```python
from dataclasses import dataclass

@dataclass(frozen=True)
class User:
    id: int
    email: str
    roles: list[str]

# Public interfaces are fully typed; modern syntax only
def load_users(db: Database, limit: int | None = None) -> list[User]: ...

# Boundary shapes are explicit types, never dict[str, Any]
def to_response(user: User) -> UserResponse: ...
```

Rules:
- Type every public function, method, and class attribute. Module-private helpers may omit hints when the types are obvious.
- Use modern syntax: built-in generics (`list[str]`, `dict[str, int]`) and unions (`X | None`) — never `typing.List`, `Optional`, or `Union`.
- Schemas, serializers, services, and data structures are typed classes where appropriate: frozen dataclasses for internal data, Pydantic models or `TypedDict` at serialization and validation boundaries.
- Never rely on implicit or ambiguous shapes — no bare `dict`, positional tuples, or `Any` crossing a module boundary. If a shape matters, give it a name and a type.
- Run a suitable static type checker when practical: `uv add --dev mypy` then `uv run mypy .` (pyright is an acceptable alternative). Run it alongside ruff and pytest in CI.

**Why**: Typed interfaces make wrong calls fail at check time instead of runtime, and named shapes document intent where docstrings drift. The type checker is the cheapest reviewer the project has.

---

## 4. Logging

**Rule**: Use Python's `logging` module with named loggers, never `print()`. Configure formatters and handlers at startup.

```python
import logging
import os

def setup_logging(level=None):
    level = level or ('DEBUG' if os.getenv('APP_DEBUG') else 'INFO')

    formatter = logging.Formatter(
        '%(asctime)s %(name)s %(levelname)s %(message)s',
        datefmt='%Y-%m-%d %H:%M:%S'
    )

    console = logging.StreamHandler()
    console.setFormatter(formatter)

    root = logging.getLogger()
    root.setLevel(level)
    root.addHandler(console)

    # File handler
    os.makedirs('data/logs', exist_ok=True)
    file_handler = logging.FileHandler('data/logs/app.log')
    file_handler.setFormatter(formatter)
    root.addHandler(file_handler)
```

```python
# In any module
import logging
logger = logging.getLogger(__name__)

logger.info('Server starting on port %s', port)
logger.error('Failed to connect: %s', err)
```

**Why**: Named loggers trace messages to source modules. Structured format enables log parsing.

---

## 5. Environment Separation

**Rule**: Maintain distinct `.env` files per environment; the same typed `Config` reads whichever `.env` is present. Never run debug mode in production.

| Setting | Dev | Test | Prod |
|---------|-----|------|------|
| APP_DEBUG | 1 | 0 | 0 |
| DATABASE_PATH | data/app.db | :memory: | data/app.db |
| SECRET_KEY | .env value | .env value | .env value (required) |
| LOGGING | DEBUG | WARNING | INFO |

**Why**: Environment separation lives in `.env` values, not Python config subclasses, so the same code path runs everywhere. This prevents dev shortcuts from reaching production.

---

## 6. Testing

**Rule**: Use `pytest` with fixtures. Isolate each test with a fresh database. Test at the boundary, not internals. Every Python project build must include a complete pytest suite regardless of whether the specification mentions tests — a project without tests does not satisfy ACTIVE conformity.

### Required test files

**`tests/conftest.py`** — fixtures shared across all test modules:
- `app` fixture: `create_app(TestConfig)` — in-memory DB, `TESTING=True`
- `client` fixture: `app.test_client()`
- `db` fixture: fresh `init_db(':memory:')` per test, yielded inside `app.app_context()`

```python
# tests/conftest.py
import pytest

@pytest.fixture
def app(monkeypatch):
    monkeypatch.setenv("SECRET_KEY", "test")
    monkeypatch.setenv("DATABASE_PATH", ":memory:")
    from app import create_app
    from config import Config
    yield create_app(Config.load())

@pytest.fixture
def client(app):
    return app.test_client()

@pytest.fixture
def db(app):
    from db import Database
    with app.app_context():
        yield Database(':memory:')
```

**`tests/test_smoke.py`** — liveness checks:
- App factory returns a Flask app without error
- `GET /health` returns 200 and `{"status": "ok"}`
- Root route `GET /` returns 200

**`tests/test_routes.py`** — one test per registered route:
- Every `GET` page route: `assert response.status_code == 200`
- Every `POST` API route: assert status in `{200, 201, 204}` with a minimal valid payload
- HTMX routes: include `HX-Request: true` header; assert 200 and non-empty `response.data`
- Routes with `{id}` params: use a fixture-created record for the ID

**`tests/test_db.py`** — only if project has a DATABASE.md:
- Schema test: after `Database(path)` init, all expected tables exist (`SELECT name FROM sqlite_master WHERE type='table'`)
- Round-trip per major table: insert a minimal valid row, read it back, assert field values match
- FK enforcement: inserting a row with an invalid FK raises `IntegrityError` (requires `PRAGMA foreign_keys=ON`)

### Configuration

**`pytest.ini`** at project root:
```ini
[pytest]
testpaths = tests
addopts = -v
```

Add `pytest` to `pyproject.toml` dev dependencies (see §8).

### What not to test
- Third-party library internals (Flask, SQLite, HTMX)
- Configuration loading — tested implicitly by fixture startup
- Private helper functions — test through the public interface that uses them

**Why**: Fixtures ensure clean state per test. In-memory DB makes tests fast.

---

## 7. Security Basics

**Rule**: Validate all user input. Use parameterized queries exclusively. Never trust client data.

Checklist:
- Parameterized queries for all DB operations (`?` placeholders, never f-strings)
- `secure_filename()` for any file path from user input
- Length and type validation on inputs
- Secret key loaded from environment, not hardcoded in prod
- Never expose stack traces to end users

**Why**: These basics prevent the most common attack vectors with minimal effort.

---

## 8. Dependency Management (uv)

**Rule**: Use `uv` for venv creation and dependency management. `pyproject.toml` is the required manifest; `uv.lock` is committed.

```bash
uv venv                         # creates .venv/
uv add flask python-dotenv      # add runtime deps → updates pyproject.toml + uv.lock
uv add --dev pytest ruff mypy   # add dev deps
uv sync                         # install from uv.lock (standard clone setup)
uv sync --frozen                # strict install (CI — fail if lock is stale)
```

```toml
# pyproject.toml
[project]
name = "my-project"
version = "0.1.0"
requires-python = ">=3.11"
dependencies = [
    "flask>=3.1",
    "python-dotenv>=1.0",
]

[project.optional-dependencies]
dev = ["pytest>=8.0", "ruff>=0.4", "mypy>=1.10"]

[tool.ruff]
line-length = 88

[tool.ruff.lint]
select = ["E", "F", "I", "UP", "B"]

[tool.ruff.format]
quote-style = "double"
```

Rules:
- Use `uv add` / `uv pip install` — never bare `pip install`
- Use `uv venv` — never `python -m venv`
- Commit `pyproject.toml` and `uv.lock`; `.venv/` is gitignored
- Keep runtime dependencies minimal; dev deps in `[project.optional-dependencies].dev`
- When migrating an existing project: `uv venv`, `uv pip install -r requirements.txt`, `uv lock`, commit `uv.lock`

**Why**: uv resolves and locks dependencies deterministically, eliminating "works on my machine" drift. `uv sync --frozen` in CI guarantees the exact locked versions are installed.

Full toolchain conventions (uv workflow, ruff rulesets, local/CI gates): see `stack/uv_ruff.md`.

---

## 9. Health Check and Startup Validation

**Rule**: Validate required config and DB connectivity at startup. Crash early on misconfiguration.

```python
def validate_startup(config: Config, db: Database):
    """Crash early on misconfiguration. Config.load() already validates required
    env vars; here we confirm the database is reachable."""
    try:
        db.healthcheck()          # runs SELECT 1 inside the Database class
    except Exception as e:
        raise RuntimeError(f'Database not accessible: {e}')

    logger.info('Startup validation passed')
```

**Why**: Required env vars are validated when `Config.load()` constructs the typed config, so startup validation only needs to confirm connectivity. Catches misconfigurations immediately rather than at first user request.

---

## 10. Project Directory Layout (Python-specific)

Python web projects extend the common layout:

```
project-name/
├── app.py              # Entry point / app factory
├── routes.py           # Route handlers
├── models.py           # Data models and type registries
├── db.py               # Database class: typed tables (row dataclass + CRUD), connection, schema, migrations
├── ops.py              # Business logic and operations
├── config.py           # typed Config class — the only env reader (stack/persistence.md)
├── templates/          # Jinja2 or Django templates
│   ├── base.html
│   └── types/          # Type-specific partials
├── static/
│   ├── css/
│   └── js/
├── tests/
│   ├── conftest.py
│   └── test_*.py
├── bin/                # (from common.md)
├── data/               # (from common.md)
├── pyproject.toml      # preferred dependency manifest
├── uv.lock             # committed — reproducible install record
├── .env
├── .gitignore
└── CLAUDE.md           # endpoints/bookmarks live in AGENTS.md — no Links.md
```

---

## Summary Checklist

- [ ] One typed `Config` class is the only env reader; no hardcoded secrets, no Dev/Prod/Test subclasses; `.env.example` maintained (`stack/env_variables_and_secrets.md`)
- [ ] Code understandable through naming, structure, small units, explicit types, clear interfaces, appropriate abstractions, and tests
- [ ] All persistence/services through typed classes (`stack/persistence.md`) — no raw SQL/`os.environ`/`open()`/SDK in app code
- [ ] Modern type hints on all public interfaces; typed schemas/serializers/services; no ambiguous shapes across boundaries; type checker run when practical
- [ ] Logging with named loggers, not `print()`
- [ ] Distinct dev/test/prod configs
- [ ] pytest with fixtures and isolated test DB
- [ ] Input validation, parameterized queries
- [ ] `uv` for venv + deps; `pyproject.toml` + `uv.lock` committed; `.venv/` gitignored
- [ ] Startup validation for required config
````
</pblock>


## IMPLEMENTS - Authoritative Step Specifications

<pblock label="Implementation recency anchor" kind="section">
The files in this section are the load-bearing specifications for this build block.
Build these files exactly. Treat earlier sections as constraints and context.

</pblock>

<pblock filename="ARCHITECTURE.md" role="implements" path="/mnt/c/Users/barlo/projects/drydock/uat/jq/runs/20260822.044627/workspace/targets/jq/blueprint/ARCHITECTURE.md" guidance="Foundational Architecture Specification">
```markdown
# ARCHITECTURE: jq Interpreter

| Field       | Value |
|-------------|-------|
| Version     | 20260822 V1 |
| Description | Defines the modular standard-library Python architecture for the standalone jq interpreter. |
| Depends On  | — |
| Provides    | interpreter module boundaries, executable boundary |
| Consumes    | — |

## Intent

The interpreter is a standalone executable named `jq`. It accepts `./jq -c '<program>'`, reads JSON values from standard input, evaluates jq filters as ordered generators, and writes compact JSON values to standard output.

## Technology Stack

- Python 3.11 or newer, using only the standard library.
- POSIX `sh` for the supplied scoring entry point.
- No third-party runtime dependency, network access, package installation, jq binding, or shell-out to another jq executable.

## Modules and Boundaries

| Boundary | Responsibility |
|---|---|
| `jq` executable | Parse command-line arguments, read standard input, invoke the interpreter, serialize outputs, and map failures to exit codes. |
| Lexer | Tokenize jq literals, identifiers, fields, bindings, keywords, operators, delimiters, comments, formats, and interpolated strings. |
| Parser | Produce an AST or equivalent intermediate representation, enforce precedence and static validity, and reject invalid programs. |
| Evaluator | Execute filters as ordered streams, preserving multiplicity, backtracking, Cartesian products, and partial output. |
| Runtime values | Represent JSON values, literal-aware numbers, NaN, infinities, arrays, objects, and immutable transformations. |
| Builtins | Implement jq primitives and standard-library filters over the evaluator's stream model. |
| Paths and assignment | Discover paths and apply immutable reads, writes, updates, and deletions. |
| Diagnostics | Keep compile failures, runtime failures, stderr output, and successful completion distinct. |

The executable boundary must not depend on a system jq command or an external implementation. Parser and evaluator interfaces remain internal to the Python implementation; the only public interface is the executable process contract.

## Runtime Contracts

- Compile failure exits `3`.
- Runtime failure exits `5`.
- Successful completion exits `0`.
- Diagnostics are written to stderr.
- Each produced value is serialized as one compact JSON value per line.
- Values emitted before a runtime failure remain on stdout.
- Generator ordering and multiplicity are observable and must be preserved.

## Module Ownership

| Concern | Owning boundary | Allowed dependencies |
|---|---|---|
| Process and CLI | executable boundary | parser, evaluator, serializer, diagnostics |
| Syntax | lexer and parser | standard-library text handling, AST definitions |
| Evaluation | evaluator | AST, runtime values, builtins, paths |
| Persistence/configuration | none | jq has no persistent store or application configuration |
| File store | none | module loading is excluded by the fixed interface |
| External services | none | network and external runtimes are forbidden |

## Numeric Decision

Use a literal-aware standard-library numeric model. Preserve source/input number spelling where jq semantics require it, perform arithmetic using standard-library numeric operations, and support special floating-point values without adding dependencies.

## Source Role Context

The implementation uses the staged lexer, parser, manual, builtin reference, corpus, and conformance harness as read-only context. The staged harness remains external to the implementation and is never modified.

## Programmatic Acceptance

- None. This specification defines architecture and boundaries; executable behavior is verified by the implementing feature specifications and the terminal conformance story.

## User Acceptance

- None.

## Guardrails

- Do not add third-party dependencies.
- Do not modify files under `sources/`.
- Do not shell out to a system jq executable.
- Do not collapse generator streams into single return values.
- Do not conflate compile exit `3` with runtime exit `5`.
```
</pblock>

<pblock label="Build instructions" kind="instructions">
### Build instructions for this block

#### Define the standalone interpreter architecture and module boundaries. (architecture)
Establish the Python standard-library architecture, executable boundary, lexer/parser/evaluator
ownership, immutable value and path-operation boundaries, and diagnostic flow. Do not add
third-party dependencies or shell out to another jq implementation.

</pblock>

<pblock label="Reusable compact request" kind="section">
## Reusable compacts

The Blueprint sources below are consumed as context by later Manifest blocks.
In this same response, extract their consumer-facing contract surface. Preserve
interfaces, schemas, constraints, configuration, and cross-file obligations; drop
implementation narrative and repetition. Do not write these files yourself.

Before the required RESULT block, emit one optional payload per source exactly as:
<reusable-compact filename="SOURCE.md">
compact content
</reusable-compact>

Emit no payload when a source has no useful technical surface. These payloads are
advisory and do not change the required build result or file-change report.

Sources eligible for reusable compaction:
- ARCHITECTURE.md

</pblock>


# Agent Task

You are a Drydock build agent implementing exactly one build step of a larger plan.
The build job block below names the target, the build working directory, and the
step. Everything you need is stacked into this prompt under role headings:

- `compass` — the Target's COMPASS.md orientation.
- `implements` — the Typed Specification files this step builds. These are
  authoritative; implement them exactly.
- `context` — read-only support specifications. Do not reimplement them.
- `stack` — enterprise stack and technology rules. Honor them.
- `rules` — governance and branding rules. Honor them.

Operating contract:

1. Follow the write authorization and protected paths in the stacked `COMPASS.md` exactly.
   That persisted guardrail is the sole authority for paths this build may modify.
2. Start by inspecting the build working directory. Preserve existing application
   files unless this step's specifications require a change.
   Its `sources/` subdirectory holds staged build assets — imported test corpora,
   conformance harnesses, and fixtures — placed there for you. They are read-only
   inputs: run them, import them, and write code against them, but never create,
   rewrite, trim, regenerate, or substitute one, even to make a check pass. A step
   that modifies a staged asset fails and the asset is restored. If an asset you
   expect is absent, report that; do not author a replacement.
3. Implement only this step. Use `context`, `stack`, and `rules` as constraints,
   not as additional work to perform.
4. Follow the stack and rules for languages, structure, naming, and branding.
5. The programmatic acceptance assertions in the `implements` specifications are
   this step's **Definition of Done** — human-owned, declared before the build,
   and fixed. Build the story and, in this same step, write the deterministic
   tests that prove each declared assertion, as a TDD master would; add finer
   tests for coverage. Every test you write follows the same rule the acceptance
   assertions do: **act on the system, read the state back, compare to expected.**
   The oracle is a return value, parsed JSON, a status code, a stored row, file
   contents read back, or an exit status — never a substring of captured stdout or
   stderr, a test-runner tally, or a log line. Write tests in the project's own
   language using that language's libraries; an in-language HTTP client yields a
   status code and a parsed body, where `curl` yields text to scrape. Round-trip
   anything that stores state: act, then read back through the public interface.
   Assert declared failure signals on negative paths, never message wording. You may add tests but must never remove, soften, or weaken
   a declared acceptance assertion. A `Suite: full` conformance check gates on the
   entire imported test suite: the step is done only when it passes in full, never on a
   representative subset — reproduce the standard exactly rather than wrapping a
   third-party library that approximates it. For a suite, the runner's exit status is the
   verdict and the whole verdict: print its captured output for diagnosis, never assert on the
   text of its summary. When an assertion is a static or filesystem
   scan (import boundary, "X never appears outside Y," grep/AST gate), honor the
   scope the specification states and never widen it: scan production source only,
   exclude `.venv/`, `site-packages`, and vendored or generated code, and do not
   flag test doubles or fixtures that use the guarded dependency.
   Run every declared acceptance assertion before returning. For a conformance suite,
   use its section or example filters to diagnose coherent root-cause clusters, but
   rerun the full declared scope before reporting the result. Treat failing examples as
   a work queue for fixing general behavior; never add example-specific exceptions.
6. Grow the project's own test suite as you write the code, and treat it as the project's
   real coverage. The acceptance assertions in `implements` are gates: few, fixed, and
   written before any code existed, so every expectation in them is a prediction. The tests
   you write are written *beside* the finished code, so their expected values are observed
   rather than predicted — which is why exhaustive coverage belongs here and not there.
   Extend the suite in the project's established location and runner, keep it runnable by the
   project's declared test command, and leave it green when you return.
   Cover, at minimum: every public entry point and every verb it declares, including declared
   error paths; the boundaries — empty, exactly one, many, absent optional fields, declared
   maxima; declared idempotence, applied twice; one behavior per test, named for the behavior;
   and isolation — each test arranges its own data, with a fresh store or explicit teardown, so
   a run leaves no residue behind in the build directory.
   Where a staged authoritative suite already covers a surface, that suite is the coverage:
   run it, and do not restate its cases. Run it **only through the invocation this step's
   acceptance criteria declare**. A criterion marked `Suite: scoped` names the whole of this
   step's obligation to that suite; running the suite's unscoped entry point instead is not
   extra rigor, it is a different step's gate executed early. A partial capability fails most
   of an authoritative corpus by construction and its unimplemented cases exhaust the runner's
   per-case timeout rather than returning, so the unscoped run costs the most where it teaches
   the least, and interrupting it forfeits the step. If no criterion in this step invokes the
   staged suite, do not invoke it. Report the pass/fail counts of the invocations you did run
   in your `SUMMARY` so a reader can see coverage moving across steps.
7. Treat `User Acceptance` entries as review evidence requirements. Implement
   the supporting behavior, but do not claim to have performed human judgment.
8. The `implements` section is authoritative and intentionally stacked late in
   the prompt as the recency anchor. Build that WHAT exactly; do not substitute
   generic framework defaults.
9. Before adding or installing Python dependencies, verify each package name
   against the declared registry. Do not invent package names. If a needed
   package cannot be verified or appears newly published, fail explicitly
   instead of installing it.
10. Use the stack's required package manager workflow for dependency changes.
   When the stack requires `uv`, update manifests through `uv` conventions
   rather than bare `pip install`.
11. Do not claim success unless you actually created or modified project files in
   the build working directory. If you cannot write files or cannot complete the
   step, report failure explicitly.
12. Do not run `git add`, `git commit`, create branches, create tags, rewrite
   history, or otherwise mutate Git history. Drydock owns the final build
   directory commit after you return.
13. End your response with this exact closing structure:

```text
RESULT: SUCCESS | FAILED

FILES CHANGED:
- relative/path

SUMMARY:
<brief reviewable summary>

BLOCKERS:
- <only if any>
```
   Before `RESULT`, you may emit one optional JSON payload when implementation required a bounded
   choice not already settled by the owning specification. This records what you did; it does not
   ask permission, create a questionnaire, or excuse incomplete work:

```text
<blueprint-decisions>
[{"spec":"FEATURE-Example.md","severity":"Material","subject":"Chosen behavior","decision":"Options A and B were available. I implemented B because ... Is that acceptable, or should this change on replan?"}]
</blueprint-decisions>
```

   Name only a specification implemented by this build block. Use `Low` or `Material`; Build never
   emits a Blocking decision. Omit the payload when no implementation decision was necessary.

14. `FILES CHANGED` must list only files actually written in the build working
   directory. If no files were written, use `RESULT: FAILED`.
15. On `RESULT: FAILED`, append two additional lines so the failure is actionable
   without opening logs. `FAILURE_SUMMARY` is one line naming the cause;
   `FAILURE_DETAIL` states what happened, why, and what to change before a rerun.
   Name concrete conditions when they apply: token or context limit exceeded,
   could not execute commands in this environment, a required input was missing,
   or a specific tool or command failed.

```text
FAILURE_SUMMARY: <one line naming the cause>
FAILURE_DETAIL: <what happened, why, and what to change before rerunning>
```

16. When a declared acceptance criterion cannot pass no matter how the code is written,
   say so with this exact token. You may not edit the criterion — it is staged and
   restored before grading:

```text
AC_BROKEN: <check-id>[, <check-id>]
```

   This is a report, not a verdict, and it stops nothing. A criterion reaches you only
   when its expected value is one its author could not have invented — a status code, a
   staged suite's exit status, a value the criterion itself supplied as input — so your
   claim that the criterion rather than the code is at fault is the less likely
   explanation, and the budget is spent as it would be for any other failure. A
   criterion whose expectation *was* hand-typed already settles `DISPUTED` on its own,
   without you naming it. Emit the token only after running the criterion and confirming
   the underlying command succeeded while the assertion still failed. Name the affected
   check ids, emit it alongside your normal `RESULT` line, state the reasoning in
   `FAILURE_DETAIL`, and emit it even when `RESULT: SUCCESS`. Do not use it for a
   criterion you merely failed to satisfy.

