PROP-055 — ChatGPT-only campaign execution pool
01Status: owner-ruled operating contract for the central OpenAI ChatGPT/Codex coordinator while the lifecycle/extensions campaign is active. It consolidates worker routing, launcher use and test-panel cadence. It does not change VibeVM product semantics, Claude Code behavior, or a delegated packet.
02Owner ruling 2026-09-08: the subscription-backed external GLM launcher is retired and MUST NOT be invoked, probed, retried, deployed, or selected as a worker lane. The host root dependency, installed projection, and generated boot contribution are removed. The canonical package source remains unchanged as historical product evidence; it is not current execution authority. Dated observations and legacy anchor names below describe completed runs only and never authorize another command. Decision: remove the host root edge, materialized slot, and boot contribution, and forbid the launcher family as a current lane. Why: the subscription is cancelled, so the transport cannot execute work while its boot still consumes every central session's context. Considered and rejected: keeping a dormant installed copy still pays that context cost; deleting or rewriting the canonical package source would destroy deliberately preserved product evidence and violates the owner's explicit perimeter. When to revisit: only after the owner restores a compatible subscription and explicitly asks to reinstall the package into this host project.
03The portable
direct-GLM contract was extracted into
tool:org.vibevm.world/zai-glm-claude and remains preserved in its canonical
package source. This host no longer consumes its boot or launcher artifacts.
The rest of this PROP retains the live ChatGPT campaign controls that are
independent of that retired transport.
1. Applicability — a harness gate, not a repository-wide mode
- 04This document applies
only when the current owner-facing central coordinator runs under an
OpenAI ChatGPT/Codex harness and is executing
campaigns/packages-2026-09/TZ-LIFECYCLE-EXTENSIONS-v0.1.md(or the owner explicitly reactivates this contract by name). All conditions are required. - Claude Code, a
Gemini and every other non-OpenAI harness MUST NOT read
or apply this document. It is deliberately absent from the shared
STATIC.xml/INDEX.mdboot lane. The byte-identicalCLAUDE.md/AGENTS.md/GEMINI.mdtriplet carries one conditional pointer: an applicable OpenAI coordinator follows it; every other harness skips it. - A task carrying
##subagent-quiet-clause, a worker packet or an explicit reviewer packet is a subagent even when its executable is Codex. It follows that packet and MUST NOT promote itself to this central contract. Native workers,codexrunnerand plainclaudeworkers are all excluded by this rule. - When the lifecycle campaign completes, pauses, or the owner selects another campaign, this contract becomes inactive without changing global project practice. A successor campaign needs an explicit owner ruling before it inherits any item.
2. The central coordinator keeps judgment
- 05The central ChatGPT/Codex session owns architecture, cross-wave planning, owner rulings, destructive-action decisions, review verdicts, integration, staging allow-lists, commits, mirrors and final gates. A worker may propose or implement; it never self-accepts, moves authoritative status, commits, pushes or declares the campaign complete.
- Delegate a bounded task when independent verification is cheaper than generating it centrally. Keep a cross-cutting edit in one thread when its write perimeter intersects live atoms or correctness depends on an unresolved architectural choice.
- For a
substantive implementation, prefer review by a different model family or
fresh independent thread. Diversity counts only with exact file/line
evidence, a root rerun of the decisive test and applicable ordinary
negative-case or invariant checks; another model's confidence is not
evidence. Counterfactual mutation follows the package's
MUTATION-DIAGNOSTIC-LAWand is never required by diversity alone. - Retired
transport tombstone. No execution or correction command is authorized by
this anchor. The enduring rule is
CENTRAL-OWNS-JUDGMENT: worker claims have no acceptance authority. - Owner ruling 2026-08-28: do not commission native collaboration subagents that inherit Sol Ultra for acceptance or review. Every such spawn pays the repository's enormous boot-lane prefix again before it can inspect the diff, making the review disproportionately expensive. Acceptance stays in the already-booted central Ultra context; an independent second family may use an explicitly available Opus or lower-cost external lane, never a native Sol Ultra reviewer. This does not forbid a narrowly justified native execution task; it removes native Sol Ultra from the acceptance route.
3. Worker pool and routing
06The pool is a set of interchangeable execution lanes with separate limits. Availability is probed on live work; no lane is assumed healthy from an old ping.
- 07Native collaboration
workers are the current low-latency execution lane for a narrow
implementation or inventory task that benefits from current conversation
context. Select model and reasoning effort explicitly when the task needs
an override. Use disjoint write
perimeters in parallel; retain packet/report/no-commit discipline. A Sol
Ultra native worker is never used for acceptance, per
NO-NATIVE-SOL-ULTRA-ACCEPTANCE. codexrunner—gpt-5.6-sol, effortxhigh— is the second external lane for hard, well-specified coding, broad code reading and evidence-backed repair when a separate Codex process or model diversity is useful. It runs in a dedicated worktree. Its yolo transport makes root diff review and a closed write allow-list mandatory; worker claims about review, commits or panels have no authority.- Retired transport tombstone. The subscription is cancelled; no health probe, fan-out, continuation, or fallback through this lane is permitted.
- Plain Claude Code with
Opus and max effort is invoked explicitly as
claude --model opus --effort max; follow-up repair usesclaude -cin the exact same dedicated cwd. Use it for ambiguous code, adversarial architecture/code review, or difficult correction. Read run metadata to learn the resolved model; the alias is not proof of a version. - Never select Fable. Do not name it, configure it as fallback, or accept automatic fallback to it. If Opus is unavailable, route to Sol, a native worker or the central coordinator.
3.1 Default routing calculus
- 08Hard but explicit code with a cheap oracle: use a native collaboration worker first; Sol xhigh or an explicitly available external family may provide independent execution when diversity is worth the spend. Use both concurrently only for disjoint writes. Ambiguous high-impact code remains centrally designed and independently reviewed.
- Security, wire, transaction and lifecycle reviews: central Ultra in its already-paid context and, when independence is valuable, a fresh external model family that did not author the change. Never pay the whole boot lane to a native Sol Ultra acceptance subagent. Read-only external reviews may run beside one code worker.
- Mechanical edits, fixtures, corpus additions and exact test matrices: use a native worker or Sol according to context transfer and verification cost. The packet supplies the skeleton; the worker fills named refinement points rather than redrawing the architecture.
- A deterministic
micro-edit remains central when its complete implementation and oracle
are cheaper than provisioning, reporting and reviewing an external
worker. A 2026-08-26 external-worker measurement took about 150 seconds
plus packet/report review for one prescribed
specmark::scope!line. That confirms deterministic micro-edits remain central. This is the verification-cheaper-than-generation law. - Architecture, provider boundaries, public grammar, final trade-offs, commit grouping and integration stay central. External architecture work is adversarial review of a central draft, not delegation of the final decision.
4. Worker transport and resource limits
- 09Every writing worker
gets a dedicated worktree/cwd and packet file. Continuation (
-c) is used only from that exact cwd. Large context travels in the packet file, never a command-line argument. - Owner
ruling 2026-08-26: authorized native and external workers may use the
normal task-capable tool surface. Do not spend orchestration effort
starving them of PowerShell/Bash, Read/Glob/Grep,
repository history, formatters, generators and exact tests. Launch an
autonomous worker with an explicit non-interactive permission mode and
a normal task-capable tool surface; a restricted
--toolsexperiment is justified only when the restriction itself is being measured. Claude Code 2.1.220 distinguishes--tools(the actual universe) from--allowedTools(auto-approval), but this is an observability fact, not a mandate to micromanage a strong worker. A resumed thread still gets an explicit mode so stale session permissions are not accidental. - The protected boundary is the handoff representation, not the worker's access to the machine. A writing worker leaves its result as ordinary unstaged working-tree files plus the durable report in its dedicated cwd. It MUST NOT package or hide the state in the index, a commit/branch, stash, private worktree, patch-only deliverable, ignored/cache-only product copy, reset/checkout state or an external service. Read-only git history, blame, status and diffs are encouraged. Root can then inspect one visible diff, rerun the oracles, stage, commit and integrate without decoding a worker-private state machine.
- Parallelism is routed on write intersections, not read intersections. Read-only reviewers may overlap. Intersecting or many-place writes are serialized under one worker or the central coordinator.
- The box ceiling is two
simultaneous cold Cargo workers. Prefer one cargo-heavy implementation
plus one cargo-heavy review/test at most; text/read-only workers may use
otherwise free lanes. Cargo runs use
CARGO_BUILD_JOBS=4where the repository rule requires it. - A packet-level “do not run
Cargo” rule does not prevent Claude Code's Rust language server from
starting
rust-analyzer/flycheck after the first Rust edit. Measured on 2026-08-27: one patch-only worktree grew a 2,225,158,486-byte localtarget/; a second had already grown 297,254,505 bytes when its process was restarted. Root therefore setsCARGO_TARGET_DIRto the one active warm integration target before launching every patch-only Rust worker. The sibling launched with that environment before its first edit created no local target. Setting it only on a continuation is too late because an LSP spawned by the original process may outlive that process. Orphaned rust-analyzer processes are stopped only after their worker terminates; the exact worktree is then reclaimed by##ROLLING-WORKTREE-GC. - A process census is
observation, not exclusion: on 2026-08-26 three independent workers all
observed zero Cargo processes and then started three cold builds in the
same race window. Every ChatGPT-commissioned worker therefore runs each
repository Cargo command through
python campaigns/packages-2026-09/tasks/chatgpt-cargo.py -- cargo …. The wrapper holds one of two OS file locks for the complete child lifetime, emits a waiting heartbeat, setsCARGO_BUILD_JOBS=4unless already set, and releases on normal exit or process death. A packet may bypass it only for a central coordinator's explicitly serialized command. The wrapper is campaign harness infrastructure, never a VibeVM product command or standing instruction for Claude. - When this
campaign changes in-workspace
file://package sources and the central coordinator refreshes the local lock/materialisation, runvibe install --offline --assume-yes. The required source bytes are local and the selected dependency closure is already cached; an online resolver adds latency and an unnecessary external failure mode. Drop--offlineonly when the task explicitly needs a remote package version not present locally. If this campaign has changed install/materialiser code, drive that refresh through the current workspace CLI (chatgpt-cargo.py -- cargo run -p vibe-cli -- install --offline --assume-yes) or a separately verified current build, never an assumed PATH binary: on 2026-08-28 the stale system launcher rewrote current.vibe-slot.tomlrecords back to legacy.vibe-derived.toml; the workspace binary immediately restored the intended diff. - Campaign worktrees
are reclaimed continuously, not at campaign close (owner ruling
2026-08-26). As soon as an atom is accepted and integrated, root proves
that its product diff is landed, its report/decisions have a durable home,
its worker process is closed, and no unique untracked artifact remains;
then it immediately removes that exact worktree with
git worktree removeand runsgit worktree prune. A clean obsolete worktree already proved superseded is treated the same way.--forceis used only after the same evidence proves every remaining untracked file disposable; a live or unintegrated dirty worktree is never reclaimed for space. Root measures and reports reclaimed bytes. Retaining hundreds of gigabytes of rebuildabletarget/trees until the epic ends is itself a process defect. - Every packet
carries
##subagent-quiet-clause, an expected write perimeter, exact acceptance, narrow self-verify commands, a durableWORKER-REPORT-<id>.md, the reviewable-worktree law, a worker-side full-panel ban and finalTASK-DONE. A reasoned perimeter expansion is disclosed in the report (and is forbidden only when it would intersect a parallel writer). Screen prose is discarded; reports and visible diffs are the deliverable. - When the central
Codex harness drives a long external CLI worker through unified exec, use
ordinary text output plus the durable worker report by
default.
--output-format stream-json --verboseemits a JSON event for each thinking-token update: the 2026-08-28 A4b read-only audit produced more than 2.7 MiB of transport output for about 10.4k thinking tokens while its complete report was about 12 KiB. Feeding that stream back through the tool wastes context and can fill the process pipe. Stream JSON remains legitimate for a short transport/model-metadata probe or when redirected to a file the model does not ingest; long work is observed through process liveness, the visible diff and the bounded report. - Default a worker report to at most 150 lines: verdict, findings, changed-file set, decisive evidence, commands and deviations. A longer report is justified only by a large finding matrix or protocol review. Do not ask the worker to restate every source it read; exact citations replace narrative. The 2026-08-26 probes showed that a good Opus review can consume 47k output tokens and a GLM review 30k when no report budget is stated.
4.1. Accepted evidence and terminal lane state
- 10An untracked
worker report is evidence, never the sole durable home of an accepted
fact. Before root accepts an atom or permits its report/worktree/cache to
become disposable, root reviews every non-obvious implementation,
architecture, diagnostic and operational finding and gives it a home by
genre: normative product law → PROP/FEAT or named spec-debt; accepted
campaign architecture/status → the implementation ledger and TASKS;
real but deliberately unowned gap → BACKLOG or the tracked retrospective
harvest; harness-only launcher/panel fact → this PROP-055; local invariant
→ code contract/test plus specmap edge where applicable. Rejected,
duplicated or non-durable observations may remain report-only, but root
makes that disposition consciously rather than losing them by omission.
Later waves drain the tracked harvest/backlog, not untracked
cache/archaeology. This classification is part of acceptance, not a session-end courtesy, and applies equally to findings made by root during review. - An explicitly
available
claudelane runs against its own subscription quota. Provider JSON may expose a nominalcostUSD; record it as comparative telemetry, never treat it as the owner's marginal bill or set--max-budget-usd. Five-hour quota availability and review throughput are the practical limits. Deep-research agents are a separate, genuinely expensive class and are commissioned deliberately; they are not implied by ordinary Opus/GLM delegation. - The central ChatGPT/Codex lane preserves a 15% pause reserve (owner ruling). Before each major campaign atom, after a long gate and at least once per working hour, root runs
python campaigns/packages-2026-09/tasks/chatgpt-usage.py --pause-at-remaining 15. The helper speaks the signed-in local Codex app-server's read-onlyaccount/rateLimits/read, selects the mainlimit_id = codexbucket (Spark's separate bucket is irrelevant unless root actually switches to it), and takes the minimum remaining percentage across primary, secondary and individual windows. It prints only safe plan/window metadata; no auth, account id, prompt or response content. Exit 75 / remaining ≤15 is a temporary resource pause, never campaign completion, cancellation or blockage: start no new work or worker, preserve every working-tree/report/cache/handoff artifact, finish only the current non-destructive safe checkpoint tail when needed to avoid ambiguity, and yield to the owner. Do not delete or compact away campaign state and do not mark the plan blocked/complete. The owner then either obtains more usage and resumes this same session, or requests a durable handoff to continue with a new model/resource set. A read failure gets one bounded retry; if it remains unavailable, do not guess enough reserve exists before another large atom. Live commissioning measured 65% used / 35% remaining in the main 10,080-minute Pro window; a threshold-equal negative probe exited 75. - Workers never receive token contents or a task that reads a token file. Secret-bearing live smoke stays central and prints only a safe verdict. Launchers load their own credentials without exposing them in packet, argv, log or report.
- A quota/429/limit failure is transport state, not a code verdict. Confirm it once, keep partial artifacts, then re-route a fresh packet with durable context. Do not spin retries, kill by broad process name, or discard a useful diff because the runner exited non-zero.
- Root acceptance compares actual changed files with the report, reads the diff as a PR, re-runs the decisive essential commands and applicable ordinary negative/invariant checks, and issues PASS, repair, re-commission or discard. A counterfactual mutation runs only for a concrete suspicious result under @spec://org.vibevm.world/multi-user-planning/flows/multi-user-planning/campaign-execution#MUTATION-DIAGNOSTIC-LAW. Cosmetic tails are fixed centrally; wrong judgment or behavior uses same-cwd continuation when economical.
5. Fast development gates and the full-panel cadence
11Owner ruling, 2026-08-26: the approximately twenty-minute full repository panel must not run after every small atom. During this ChatGPT campaign, exact tests are the development gate; the full panel is an integration/batch gate. This is the scoped successor to older campaign-TZ wording and changes no other project or harness.
- 12Before accepting an atom, run formatting for each touched workspace,
git diff --check, exact unit/integration tests, checks/clippy for touched crates, applicable ordinary negative-case/invariant tests, and the task oracle: codegen+wire corpus for schemas; compiler byte identity; isolated homes. Counterfactual mutation is not an atom ritual and runs only under the package's suspicion-and-economy law. A Rust edit in aconform.toml-gated crate also runscargo xtask conform check --scope <crate-path>; split test-support carries truthful per-item#[cfg(test)]because the syntax frontend reads it without its parent. New/movedspecmark::scope!or#[specmark::spec(...)]edges runcargo xtask specmapand review the derived map in the same atom. Expand only for another named consumer. This closes the 2026-08-28 R3.4 findings: exact tests/clippy missed 52 conform findings; four later lawful edges stopped on a stale map. cargo xtask check-codegenfinishes by comparing the worktree againstHEAD; in the mandated ordinary-unstaged worker handoff, an intended generated product diff therefore makes it fail by construction. A schema/codegen worker runs the generator, captures a stable hash of the relevant visible diff, runs the generator again and proves the hash unchanged (plus focused wire/corpus tests). Root runs the realcheck-codegenafter staging the atom in its human-authored commit/integrated tree. A worker's dirty-tree exit is not codegen drift evidence by itself.- A gate transcript
is evidence only when it carries the exit status of the GATE process.
Never spell it as
command | tail; echo $?: that status belongs totail, and an A6A codegen run on 2026-08-28 printed zero over an actualos error 5publication failure. Run the gate unpiped, or redirect its output, capture the child status immediately ($?/$LASTEXITCODE/PIPESTATUS[0]as appropriate), and inspect the saved output afterward. PowerShell cmdlet/conversion errors are non-terminating by default, so a gate such as XML parsing first sets$ErrorActionPreference = 'Stop'(or uses-ErrorAction Stop); otherwise it can print an error, continue into a hand-written “clean” line and even commit invalid XML — observed and repaired on 2026-08-28. A green-looking suffix with a masked parent failure is no verdict. - A worker never runs the full panel. It runs the narrow self-verify named by its packet. The central coordinator does not run the full panel merely because a worker finished.
- A coherent batch may contain several atomic local Conventional Commits, each accepted by its exact gate. They remain unmirrored until the batch integration panel is green. If that panel finds a defect, repair forward in the same batch and rerun newly relevant narrow tests before the next panel.
- Run the full
CARGO_BUILD_JOBS=4 tools/self-check.shpanel when: (1) a cross-crate integration batch is ready to leave local state; (2) a lifecycle wave or owner scenario is declared green; (3) before mirror/push of accumulated code; (4) before final epic acceptance; (5) a high-risk change has no cheaper complete affected set. One panel may satisfy several triggers for the same unchanged tree. - A panel result belongs to the exact tree measured. A later product change invalidates it. A report/docs-only change runs its document gates and may ride the next integration panel, but cannot rewrite an earlier code verdict.
- Mirror only after the accumulated batch panel is green and the staging allow-list is reviewed. This preserves atomic human-authored commits while amortising the slow panel over useful integration work.
6. Live and historical lane observations
13All
dated observations whose anchor or prose names the retired subscription GLM
transport are historical evidence only. Their commands, retry advice,
quota assumptions, and availability conclusions are non-operative after
SUBSCRIPTION-GLM-LANE-RETIRED. Observations about other live lanes retain
their ordinary force.
- 14Record only a live task's exit, durable report, resolved-model metadata and reviewed quality. A ping proves transport, not coding quality; an old successful run does not prove a five-hour quota remains open.
- Historical external-lane probe. Its durable transport evidence was extracted to the preserved package source; the current rules retain only the conclusions that a ping is not a quality verdict and root acceptance remains mandatory.
- Retired transport budget tombstone. No current worker budget or continuation rule is derived from this anchor.
- Retired account-quota tombstone. No current retry, wait, or fan-out rule is derived from this anchor.
- Retired heartbeat tombstone. No availability probe, scheduled run, retry, or repository work may target the cancelled transport. No matching automation exists on the current Codex host as of 2026-09-08.
- Historical recovery tombstone. A past successful run does not establish present availability.
- Historical R4 transport tombstone. The accepted product evidence remains in commits and gates; this anchor authorizes no runner.
- A
later R4.1 T1/T3 pair passed its worker gates yet central review found four
identity-boundary defects: five-digit values entered a four-digit TOML
date, host provider plus document path was unconstructible, legacy plan
paid a temporary collection, and
CompiledSelector::Eqtreated OR-set order/duplicates as semantic while the promised digest did not. Two same-cwd worker repairs plus an ordinary Opus/max byte-schedule audit closed them before T2; conform also forced the enlarged selector RED into a focused sub-600-line cell. Durable packet/acceptance rule: whenever a canonical identity lands, pair its independent byte vector with Rust equality, absent-versus-empty, typed constructor composition and legacy allocation/order checks. A green digest golden alone does not prove one semantic identity. - Two
disjoint R4.1 writers proved that write-perimeter independence does not
make cross-crate gates temporally independent. The registry writer's
downstream compile twice observed the workspace writer mid-move and went
red on transient duplicate/missing symbols; after that writer reached a
stable file boundary, the same gate passed unchanged. Cargo-slot
serialisation prevents concurrent cargo processes, not compilation of a
peer's half-written source. Schedule downstream gates after peer writers
are stable or classify the exact foreign path before repairing anything.
The same run exposed a second trap:
cargo test -p vibe-install lifecycleexited 0 while running zero tests; root replaced it with the full package suite. Acceptance records test counts, never exit alone. Finally, one split continuation usedgit restoredespite a no-mutating-git packet to heal its own transient newline collapse; the desired bytes survived, but text prohibition is not enforcement — root still verifies the final ordinary diff and ownership perimeter. - Retired multi-lane quota tombstone. No current availability or failover rule is derived from this anchor.
- Retired
initial-stall tombstone. Generic transport failures remain evidence, not
code verdicts, under
LIMIT-FAILOVER. - A
plain fresh
codexrunner execloaded repositoryAGENTS.mdbefore its named worker packet and began the forbidden full boot despite that packet's##subagent-quiet-clause; root stopped it before product edits. The strict live repair is--strict-config --ignore-user-config -c project_doc_max_bytes=0, withmodel_reasoning_effort="xhigh"repeated in the later exec layer and sharedCARGO_TARGET_DIRset before launch. A no-tool probe returned exactQUIET_OK; the real retry resolvedgpt-5.6-sol/xhigh, read the packet first and loaded only its six named standing files. A packet's quiet clause controls reading only when the harness has not already injected project documentation ahead of it. tools/self-check.shlexically countsrun_steplines while seven calls execute dynamically, so final R3.4 printed[54/47]and still endedall green/ exit 0. The denominator is an observability bug, never coverage evidence: trust the ordered gates + green tail. A later tooling atom derives the runtime total or removes the false denominator; this docs-only finding does not rewrite the green verdict.
7. Canonical worker packet template
15The central coordinator copies this skeleton into every substantial worker packet and fills every bracketed field. Do not point a worker at PROP-055 itself: this spec is ChatGPT-only. Task-specific instructions override a template default only when the packet says so explicitly.
16# PACKET-<task-id> — <one bounded outcome>
## Role and authority
You are an EXECUTION worker, not the architect, reviewer, integrator or owner.
The central coordinator owns architecture, acceptance, staging, commits and
final gates. Your `PASS` is advisory. Do not claim that root reviewed/accepted
anything. A prior report is a lead, never authority.
On a continuation or follow-up, discard your previous verdict where the
correction packet contradicts it. Re-read the correction packet first and
repair the named failure; do not defend the old answer.
## Task
<exact behavior to implement or question to review>
## Required inputs
- <files/spec anchors/reports to read>
- <facts supplied by root; confirm or refute each with exact evidence>
If an ambiguity would change public grammar, architecture, security posture or
scope, STOP and write the exact blocker. Do not silently choose a new design.
## Write perimeter
- <expected files/directories to change; mark any path intersecting a parallel writer as closed>
- `WORKER-REPORT-<task-id>.md`
Leave every product change as an ordinary **unstaged** working-tree diff in
this cwd. Preserve pre-existing/user changes. If solving the task genuinely
requires another non-conflicting path, change it and explain the expansion in
the report. Do not create/package state in a branch, commit, index, stash,
private worktree, patch-only deliverable or ignored/cache-only product copy;
root already provisioned the cwd. Do not spawn subagents unless the packet
explicitly delegates a nested swarm.
## Acceptance
1. <criterion with observable result>
2. <negative/security/failure criterion>
3. <compatibility/idempotence criterion>
Name the essential ordinary negative-case or invariant check for each
load-bearing branch. If a concrete observation makes a result suspicious, name
one smallest counterfactual diagnostic under `MUTATION-DIAGNOSTIC-LAW`; never
commission an open-ended mutation search. An empty grep/test output is not proof
until a known positive control demonstrates the instrument can fire.
## Constraints
- Use the machine normally: Read/Glob/Grep, PowerShell/Bash, read-only git
archaeology, formatters, generators and the exact tests are available when
they help. `git log/show/blame/diff/status` are normal evidence sources.
- Keep the handoff reviewable: NO git add/commit/push/merge/rebase,
checkout/restore/reset, stash, clean, worktree, tag, config mutation or
history rewrite; do not leave the only deliverable in generated scratch,
cache, a patch file or an external service.
- NO full repository panel; run only the exact commands listed below.
- NO specs/status/SPEC-DEBT/WAL/CONTINUE/TASKS/BACKLOG/AGENTS/campaign edits.
- Work inside the assigned task. Network, user-home or external writes are
allowed when the task needs them, but publishing/pushing or another
irreversible external action still requires the packet to grant it. Never
expose a token value in argv, logs or the report.
- Run every repository Cargo command through
`python campaigns/packages-2026-09/tasks/chatgpt-cargo.py -- cargo …`; the
wrapper owns the two-slot ceiling and supplies `CARGO_BUILD_JOBS=4`. Do not
replace it with a process census, which races across workers.
- Preserve the actual exit of every gate. Do not pipe a gate into `tail`,
`head`, `grep` or a formatter and then report the pipeline's last process as
the gate result; redirect first and capture the gate child status directly.
- For an intended generated-code diff, follow `##DIRTY-CODEGEN-GATE`: prove
two emissions identical in the worker and reserve `check-codegen`'s
HEAD-clean verdict for root after the generated atom is committed.
- NO new `unsafe`, FFI, handwritten OS ABI or home-grown security/locking
primitive unless the packet explicitly grants that architecture. Stop and
report the missing primitive instead; do not invent a mini-runtime.
- Every changed/new file remains at or below 600 physical lines. Check while
editing; split early, not after fmt.
- Do not weaken, delete or rewrite an oracle merely to make it green.
- Treat a non-zero runner exit as transport state; still write durable evidence
when possible.
## Allowed self-verify
```text
<exact fmt/test/check/clippy/codegen commands; no broader substitutes>
```
## Deliverable report — draft within 140 lines
Do not spend turns repeatedly trimming a long report. Create its concise
skeleton after the first product slice, then update decisions/evidence
incrementally before long parent gates (140-line target leaves headroom under a
150-line hard cap). Never defer the first durable report write to the final two
turns: a turn cap, network loss, or transport failure can end the next call,
and a worker without the report has failed its deliverable even when useful
edits remain.
```markdown
# WORKER-REPORT — <task-id>
## Verdict
PASS | FAIL | BLOCKED (advisory only)
## Changed files
- exact path — why
## Acceptance evidence
- criterion → result, exact file:line/test output
## Suspicion-triggered diagnostics
- suspicious signal → bounded diagnostic and result, or none
## Commands
- command → exit + decisive lines
## Decisions and why
- every choice within granted latitude, or none
## Perimeter expansions / external effects
- disclose every extra changed path, external write, broad test or state effect, or none
## Leftovers
- exact remaining work, or none
```
## subagent-quiet-clause
Your screen text is discarded. Output only mandated `PROGRESS:` heartbeats,
the durable report, and final `TASK-DONE`. No greeting, task restatement,
intermediate chat reasoning or final screen summary.
- 17Launch with an
explicit autonomous non-interactive permission mode and the normal tool
surface; do not construct a tiny deny-list as a substitute for review.
Read Glob Grep, PowerShell/Bash, read-only git history, formatters and exact tests are legitimate worker instruments. The short pointer prompt repeats only the handoff invariants: keep product work as an unstaged visible diff, do not add/commit/stash/branch/reset/push, do not run the full panel, and write the report before finishing. Root review, not tool starvation, is the safety and acceptance mechanism. Do not set a maximum dollar budget for Opus/Claude runs (owner ruling 2026-08-26); record cost post-facto when JSON telemetry is available. - After
TASK-DONE, root mechanically compares actual changed files to the report, reads the diff, checks Decisions/Deviations, reruns decisive essential commands and only those diagnostics whose suspicion trigger actually fired, then accepts, sends exact-crepair notes, re-commissions or discards. The packet/template reduces review cost; it never replaces review.
8. Opus 5 max — CLI lane and accumulated lessons
18Opus/max uses the canonical packet above. Root still supplies the idea, accepted architecture, task boundaries and essential verification; Opus implements or independently audits, leaves a reviewable worktree, and receives exact acceptance repairs from root.
8.1 Fresh run and continuation
19Run the installed
Claude Code launcher, never a hand-written direct API client. A first
turn is fresh claude -p; -c means continue the most recent session in
that cwd and is therefore used only for root's repair/acceptance follow-up.
Every concurrent Opus thread gets its own cwd so -c cannot resume a
neighbour. Do not use --no-session-persistence when a repair loop may
be needed.
20# Fresh task, from the worker's dedicated cwd
claude -p $pointer --model opus --effort max `
--permission-mode bypassPermissions --output-format json
# Same-cwd root correction / continuation
claude -c -p $repair --model opus --effort max `
--permission-mode bypassPermissions --output-format json
21Omit a fallback
model and inspect the JSON result's canonicalModel; never accept Fable.
The 2026-08-26 probes resolved --model opus to
claude-opus-5 with a 1,000,000-token context. Omit
--max-budget-usd by owner ruling: subscription quota, not nominal API
accounting, is this lane's resource model.
8.2 Task shape and parallelism
- 22Prefer Opus/max for wide code archaeology, adversarial architecture/transaction review, ambiguous repairs and implementation whose verification needs sustained cross-file judgment. It may code as freely as any authorized writing worker; root retains architecture and acceptance. Deterministic micro-edits still fail the verification-cheaper-than-generation test.
- Multiple Opus processes may run concurrently in separate cwd/session pairs. Parallel readers may overlap; writers are separated by intersecting product paths. Its quota is independent of Codex, so using both families at once can expand throughput and supply review diversity.
- Give Opus the normal machine toolset. The first two audits were unnecessarily launched with a read-only tool profile; one sensible PowerShell line-count command was denied. Both still completed, but the owner corrected the policy: PowerShell/Bash, Grep/Glob/Read, read-only git history and exact tests are ordinary reasoning instruments. Protect only the reviewable worktree handoff.
8.3 Live measurements — 2026-08-26
23Two Opus/max
read-only audits ran concurrently through Claude Code 2.1.220. R3
verifier/trace: terminal success, 91 turns, about 9.2 minutes wall,
37,491 output tokens, nominal telemetry $6.57, 215-line report. R4
staged transforms: terminal success, 94 turns, about 9.8 minutes wall,
41,137 output tokens, nominal telemetry $11.01, 236-line report. Both
wrote only their report and produced concrete file/line matrices, red
proofs and commit slices. The nominal dollar fields do not change the
subscription-lane ruling.
24A broad Opus audit genuinely benefits from a 200–240-line cap; implementation reports stay near the canonical 140-line target. Require the report before final completion because screen output is not the handoff. JSON telemetry is kept for model/turn/time observation; it does not replace the durable report.
25An A5a Opus/max implementation reached a green exact test, wrote the complete product diff, then produced no non-telemetry session event for about fifteen minutes while its final max-reasoning turn remained open. Root verified liveness from the Claude session JSONL, stopped only that exact marked claude.exe PID (never by name or tree), and resumed with claude -c in the same cwd. The session retained the diff and prior tool result; the continuation applied the acceptance repair and finished 13 focused + 42 parent tests. Therefore a killed/stalled final turn is not a code verdict and does not require regeneration: inspect the durable diff/report and last non-telemetry event, preserve them, then prefer same-cwd continuation. The exact-PID/no-tree rule applies to ordinary Opus too, even though its process family is not Codex's GUI family.
26The A5b implementation ignored its exact-gate budget at the tail and ran full vibe-cli sweeps through cargo test | grep/head/awk; the first spent roughly nine minutes inside unrelated cli_spec_format local install/submodule work. The rerun eventually printed 635/0, but the pipeline had discarded Cargo's exit status, so that number remained telemetry rather than root acceptance. Root used only unpiped, terminating exact e2e/golden/unit commands plus strict touched-crate clippy/workspace check/conform for the verdict. A worker package sweep is neither a substitute for the central full-panel trigger nor permission to mask the child status; when an overbroad sweep enters an unrelated slow test, stop its exact process chain and return to the packet's named tests.
9. History
27Authored from the owner's 2026-08-26 rulings: use the former external GLM lane and Opus/max alongside Sol/native workers because quotas are separate; never select Fable; consolidate the strategy in one ChatGPT-only document; keep the agent-instruction triplet byte-identical; replace per-atom full panels with exact gates and batch panels; trust both launchers with normal machine tools; preserve their result as an unstaged, centrally reviewable worktree rather than a worker commit/stash/index; and never cap subscription runs by nominal dollar budget.
28The owner removed native Sol Ultra subagents from the acceptance pool after observing that every reviewer repays the enormous boot-lane prefix. Central Ultra now accepts in its existing context; independent review is routed through subscription launchers or another explicitly cheaper external lane.