VibeVM
Contents
On this page
en
Publisher
org.vibevm.core
Version
1.0.0latest
Audiences
Reading time
23 min
Rendered
Read aloud
never

PROP-055 — ChatGPT-only campaign execution pool

01Status: owner-ruled operating contract for the central OpenAI ChatGPT/Codex coordinator while the lifecycle/extensions campaign is active. It consolidates worker routing, launcher use and test-panel cadence. It does not change VibeVM product semantics, Claude Code behavior, or a delegated packet.

02Owner ruling 2026-09-08: the subscription-backed external GLM launcher is retired and MUST NOT be invoked, probed, retried, deployed, or selected as a worker lane. The host root dependency, installed projection, and generated boot contribution are removed. The canonical package source remains unchanged as historical product evidence; it is not current execution authority. Dated observations and legacy anchor names below describe completed runs only and never authorize another command. Decision: remove the host root edge, materialized slot, and boot contribution, and forbid the launcher family as a current lane. Why: the subscription is cancelled, so the transport cannot execute work while its boot still consumes every central session's context. Considered and rejected: keeping a dormant installed copy still pays that context cost; deleting or rewriting the canonical package source would destroy deliberately preserved product evidence and violates the owner's explicit perimeter. When to revisit: only after the owner restores a compatible subscription and explicitly asks to reinstall the package into this host project.

03The portable direct-GLM contract was extracted into tool:org.vibevm.world/zai-glm-claude and remains preserved in its canonical package source. This host no longer consumes its boot or launcher artifacts. The rest of this PROP retains the live ChatGPT campaign controls that are independent of that retired transport.

1. Applicability — a harness gate, not a repository-wide mode

  • 04This document applies only when the current owner-facing central coordinator runs under an OpenAI ChatGPT/Codex harness and is executing campaigns/packages-2026-09/TZ-LIFECYCLE-EXTENSIONS-v0.1.md (or the owner explicitly reactivates this contract by name). All conditions are required.
  • Claude Code, a Gemini and every other non-OpenAI harness MUST NOT read or apply this document. It is deliberately absent from the shared STATIC.xml / INDEX.md boot lane. The byte-identical CLAUDE.md/AGENTS.md/GEMINI.md triplet carries one conditional pointer: an applicable OpenAI coordinator follows it; every other harness skips it.
  • A task carrying ##subagent-quiet-clause, a worker packet or an explicit reviewer packet is a subagent even when its executable is Codex. It follows that packet and MUST NOT promote itself to this central contract. Native workers, codexrunner and plain claude workers are all excluded by this rule.
  • When the lifecycle campaign completes, pauses, or the owner selects another campaign, this contract becomes inactive without changing global project practice. A successor campaign needs an explicit owner ruling before it inherits any item.

2. The central coordinator keeps judgment

  • 05The central ChatGPT/Codex session owns architecture, cross-wave planning, owner rulings, destructive-action decisions, review verdicts, integration, staging allow-lists, commits, mirrors and final gates. A worker may propose or implement; it never self-accepts, moves authoritative status, commits, pushes or declares the campaign complete.
  • Delegate a bounded task when independent verification is cheaper than generating it centrally. Keep a cross-cutting edit in one thread when its write perimeter intersects live atoms or correctness depends on an unresolved architectural choice.
  • For a substantive implementation, prefer review by a different model family or fresh independent thread. Diversity counts only with exact file/line evidence, a root rerun of the decisive test and applicable ordinary negative-case or invariant checks; another model's confidence is not evidence. Counterfactual mutation follows the package's MUTATION-DIAGNOSTIC-LAW and is never required by diversity alone.
  • Retired transport tombstone. No execution or correction command is authorized by this anchor. The enduring rule is CENTRAL-OWNS-JUDGMENT: worker claims have no acceptance authority.
  • Owner ruling 2026-08-28: do not commission native collaboration subagents that inherit Sol Ultra for acceptance or review. Every such spawn pays the repository's enormous boot-lane prefix again before it can inspect the diff, making the review disproportionately expensive. Acceptance stays in the already-booted central Ultra context; an independent second family may use an explicitly available Opus or lower-cost external lane, never a native Sol Ultra reviewer. This does not forbid a narrowly justified native execution task; it removes native Sol Ultra from the acceptance route.

3. Worker pool and routing

06The pool is a set of interchangeable execution lanes with separate limits. Availability is probed on live work; no lane is assumed healthy from an old ping.

  • 07Native collaboration workers are the current low-latency execution lane for a narrow implementation or inventory task that benefits from current conversation context. Select model and reasoning effort explicitly when the task needs an override. Use disjoint write perimeters in parallel; retain packet/report/no-commit discipline. A Sol Ultra native worker is never used for acceptance, per NO-NATIVE-SOL-ULTRA-ACCEPTANCE.
  • codexrunnergpt-5.6-sol, effort xhigh — is the second external lane for hard, well-specified coding, broad code reading and evidence-backed repair when a separate Codex process or model diversity is useful. It runs in a dedicated worktree. Its yolo transport makes root diff review and a closed write allow-list mandatory; worker claims about review, commits or panels have no authority.
  • Retired transport tombstone. The subscription is cancelled; no health probe, fan-out, continuation, or fallback through this lane is permitted.
  • Plain Claude Code with Opus and max effort is invoked explicitly as claude --model opus --effort max; follow-up repair uses claude -c in the exact same dedicated cwd. Use it for ambiguous code, adversarial architecture/code review, or difficult correction. Read run metadata to learn the resolved model; the alias is not proof of a version.
  • Never select Fable. Do not name it, configure it as fallback, or accept automatic fallback to it. If Opus is unavailable, route to Sol, a native worker or the central coordinator.

3.1 Default routing calculus

  • 08Hard but explicit code with a cheap oracle: use a native collaboration worker first; Sol xhigh or an explicitly available external family may provide independent execution when diversity is worth the spend. Use both concurrently only for disjoint writes. Ambiguous high-impact code remains centrally designed and independently reviewed.
  • Security, wire, transaction and lifecycle reviews: central Ultra in its already-paid context and, when independence is valuable, a fresh external model family that did not author the change. Never pay the whole boot lane to a native Sol Ultra acceptance subagent. Read-only external reviews may run beside one code worker.
  • Mechanical edits, fixtures, corpus additions and exact test matrices: use a native worker or Sol according to context transfer and verification cost. The packet supplies the skeleton; the worker fills named refinement points rather than redrawing the architecture.
  • A deterministic micro-edit remains central when its complete implementation and oracle are cheaper than provisioning, reporting and reviewing an external worker. A 2026-08-26 external-worker measurement took about 150 seconds plus packet/report review for one prescribed specmark::scope! line. That confirms deterministic micro-edits remain central. This is the verification-cheaper-than-generation law.
  • Architecture, provider boundaries, public grammar, final trade-offs, commit grouping and integration stay central. External architecture work is adversarial review of a central draft, not delegation of the final decision.

4. Worker transport and resource limits

  • 09Every writing worker gets a dedicated worktree/cwd and packet file. Continuation (-c) is used only from that exact cwd. Large context travels in the packet file, never a command-line argument.
  • Owner ruling 2026-08-26: authorized native and external workers may use the normal task-capable tool surface. Do not spend orchestration effort starving them of PowerShell/Bash, Read/Glob/Grep, repository history, formatters, generators and exact tests. Launch an autonomous worker with an explicit non-interactive permission mode and a normal task-capable tool surface; a restricted --tools experiment is justified only when the restriction itself is being measured. Claude Code 2.1.220 distinguishes --tools (the actual universe) from --allowedTools (auto-approval), but this is an observability fact, not a mandate to micromanage a strong worker. A resumed thread still gets an explicit mode so stale session permissions are not accidental.
  • The protected boundary is the handoff representation, not the worker's access to the machine. A writing worker leaves its result as ordinary unstaged working-tree files plus the durable report in its dedicated cwd. It MUST NOT package or hide the state in the index, a commit/branch, stash, private worktree, patch-only deliverable, ignored/cache-only product copy, reset/checkout state or an external service. Read-only git history, blame, status and diffs are encouraged. Root can then inspect one visible diff, rerun the oracles, stage, commit and integrate without decoding a worker-private state machine.
  • Parallelism is routed on write intersections, not read intersections. Read-only reviewers may overlap. Intersecting or many-place writes are serialized under one worker or the central coordinator.
  • The box ceiling is two simultaneous cold Cargo workers. Prefer one cargo-heavy implementation plus one cargo-heavy review/test at most; text/read-only workers may use otherwise free lanes. Cargo runs use CARGO_BUILD_JOBS=4 where the repository rule requires it.
  • A packet-level “do not run Cargo” rule does not prevent Claude Code's Rust language server from starting rust-analyzer/flycheck after the first Rust edit. Measured on 2026-08-27: one patch-only worktree grew a 2,225,158,486-byte local target/; a second had already grown 297,254,505 bytes when its process was restarted. Root therefore sets CARGO_TARGET_DIR to the one active warm integration target before launching every patch-only Rust worker. The sibling launched with that environment before its first edit created no local target. Setting it only on a continuation is too late because an LSP spawned by the original process may outlive that process. Orphaned rust-analyzer processes are stopped only after their worker terminates; the exact worktree is then reclaimed by ##ROLLING-WORKTREE-GC.
  • A process census is observation, not exclusion: on 2026-08-26 three independent workers all observed zero Cargo processes and then started three cold builds in the same race window. Every ChatGPT-commissioned worker therefore runs each repository Cargo command through python campaigns/packages-2026-09/tasks/chatgpt-cargo.py -- cargo …. The wrapper holds one of two OS file locks for the complete child lifetime, emits a waiting heartbeat, sets CARGO_BUILD_JOBS=4 unless already set, and releases on normal exit or process death. A packet may bypass it only for a central coordinator's explicitly serialized command. The wrapper is campaign harness infrastructure, never a VibeVM product command or standing instruction for Claude.
  • When this campaign changes in-workspace file:// package sources and the central coordinator refreshes the local lock/materialisation, run vibe install --offline --assume-yes. The required source bytes are local and the selected dependency closure is already cached; an online resolver adds latency and an unnecessary external failure mode. Drop --offline only when the task explicitly needs a remote package version not present locally. If this campaign has changed install/materialiser code, drive that refresh through the current workspace CLI (chatgpt-cargo.py -- cargo run -p vibe-cli -- install --offline --assume-yes) or a separately verified current build, never an assumed PATH binary: on 2026-08-28 the stale system launcher rewrote current .vibe-slot.toml records back to legacy .vibe-derived.toml; the workspace binary immediately restored the intended diff.
  • Campaign worktrees are reclaimed continuously, not at campaign close (owner ruling 2026-08-26). As soon as an atom is accepted and integrated, root proves that its product diff is landed, its report/decisions have a durable home, its worker process is closed, and no unique untracked artifact remains; then it immediately removes that exact worktree with git worktree remove and runs git worktree prune. A clean obsolete worktree already proved superseded is treated the same way. --force is used only after the same evidence proves every remaining untracked file disposable; a live or unintegrated dirty worktree is never reclaimed for space. Root measures and reports reclaimed bytes. Retaining hundreds of gigabytes of rebuildable target/ trees until the epic ends is itself a process defect.
  • Every packet carries ##subagent-quiet-clause, an expected write perimeter, exact acceptance, narrow self-verify commands, a durable WORKER-REPORT-<id>.md, the reviewable-worktree law, a worker-side full-panel ban and final TASK-DONE. A reasoned perimeter expansion is disclosed in the report (and is forbidden only when it would intersect a parallel writer). Screen prose is discarded; reports and visible diffs are the deliverable.
  • When the central Codex harness drives a long external CLI worker through unified exec, use ordinary text output plus the durable worker report by default. --output-format stream-json --verbose emits a JSON event for each thinking-token update: the 2026-08-28 A4b read-only audit produced more than 2.7 MiB of transport output for about 10.4k thinking tokens while its complete report was about 12 KiB. Feeding that stream back through the tool wastes context and can fill the process pipe. Stream JSON remains legitimate for a short transport/model-metadata probe or when redirected to a file the model does not ingest; long work is observed through process liveness, the visible diff and the bounded report.
  • Default a worker report to at most 150 lines: verdict, findings, changed-file set, decisive evidence, commands and deviations. A longer report is justified only by a large finding matrix or protocol review. Do not ask the worker to restate every source it read; exact citations replace narrative. The 2026-08-26 probes showed that a good Opus review can consume 47k output tokens and a GLM review 30k when no report budget is stated.

4.1. Accepted evidence and terminal lane state

  • 10An untracked worker report is evidence, never the sole durable home of an accepted fact. Before root accepts an atom or permits its report/worktree/cache to become disposable, root reviews every non-obvious implementation, architecture, diagnostic and operational finding and gives it a home by genre: normative product law → PROP/FEAT or named spec-debt; accepted campaign architecture/status → the implementation ledger and TASKS; real but deliberately unowned gap → BACKLOG or the tracked retrospective harvest; harness-only launcher/panel fact → this PROP-055; local invariant → code contract/test plus specmap edge where applicable. Rejected, duplicated or non-durable observations may remain report-only, but root makes that disposition consciously rather than losing them by omission. Later waves drain the tracked harvest/backlog, not untracked cache/ archaeology. This classification is part of acceptance, not a session-end courtesy, and applies equally to findings made by root during review.
  • An explicitly available claude lane runs against its own subscription quota. Provider JSON may expose a nominal costUSD; record it as comparative telemetry, never treat it as the owner's marginal bill or set --max-budget-usd. Five-hour quota availability and review throughput are the practical limits. Deep-research agents are a separate, genuinely expensive class and are commissioned deliberately; they are not implied by ordinary Opus/GLM delegation.
  • The central ChatGPT/Codex lane preserves a 15% pause reserve (owner ruling). Before each major campaign atom, after a long gate and at least once per working hour, root runs python campaigns/packages-2026-09/tasks/chatgpt-usage.py --pause-at-remaining 15. The helper speaks the signed-in local Codex app-server's read-only account/rateLimits/read, selects the main limit_id = codex bucket (Spark's separate bucket is irrelevant unless root actually switches to it), and takes the minimum remaining percentage across primary, secondary and individual windows. It prints only safe plan/window metadata; no auth, account id, prompt or response content. Exit 75 / remaining ≤15 is a temporary resource pause, never campaign completion, cancellation or blockage: start no new work or worker, preserve every working-tree/report/cache/handoff artifact, finish only the current non-destructive safe checkpoint tail when needed to avoid ambiguity, and yield to the owner. Do not delete or compact away campaign state and do not mark the plan blocked/complete. The owner then either obtains more usage and resumes this same session, or requests a durable handoff to continue with a new model/resource set. A read failure gets one bounded retry; if it remains unavailable, do not guess enough reserve exists before another large atom. Live commissioning measured 65% used / 35% remaining in the main 10,080-minute Pro window; a threshold-equal negative probe exited 75.
  • Workers never receive token contents or a task that reads a token file. Secret-bearing live smoke stays central and prints only a safe verdict. Launchers load their own credentials without exposing them in packet, argv, log or report.
  • A quota/429/limit failure is transport state, not a code verdict. Confirm it once, keep partial artifacts, then re-route a fresh packet with durable context. Do not spin retries, kill by broad process name, or discard a useful diff because the runner exited non-zero.
  • Root acceptance compares actual changed files with the report, reads the diff as a PR, re-runs the decisive essential commands and applicable ordinary negative/invariant checks, and issues PASS, repair, re-commission or discard. A counterfactual mutation runs only for a concrete suspicious result under @spec://org.vibevm.world/multi-user-planning/flows/multi-user-planning/campaign-execution#MUTATION-DIAGNOSTIC-LAW. Cosmetic tails are fixed centrally; wrong judgment or behavior uses same-cwd continuation when economical.

5. Fast development gates and the full-panel cadence

11Owner ruling, 2026-08-26: the approximately twenty-minute full repository panel must not run after every small atom. During this ChatGPT campaign, exact tests are the development gate; the full panel is an integration/batch gate. This is the scoped successor to older campaign-TZ wording and changes no other project or harness.

  • 12Before accepting an atom, run formatting for each touched workspace, git diff --check, exact unit/integration tests, checks/clippy for touched crates, applicable ordinary negative-case/invariant tests, and the task oracle: codegen+wire corpus for schemas; compiler byte identity; isolated homes. Counterfactual mutation is not an atom ritual and runs only under the package's suspicion-and-economy law. A Rust edit in a conform.toml-gated crate also runs cargo xtask conform check --scope <crate-path>; split test-support carries truthful per-item #[cfg(test)] because the syntax frontend reads it without its parent. New/moved specmark::scope! or #[specmark::spec(...)] edges run cargo xtask specmap and review the derived map in the same atom. Expand only for another named consumer. This closes the 2026-08-28 R3.4 findings: exact tests/clippy missed 52 conform findings; four later lawful edges stopped on a stale map.
  • cargo xtask check-codegen finishes by comparing the worktree against HEAD; in the mandated ordinary-unstaged worker handoff, an intended generated product diff therefore makes it fail by construction. A schema/codegen worker runs the generator, captures a stable hash of the relevant visible diff, runs the generator again and proves the hash unchanged (plus focused wire/corpus tests). Root runs the real check-codegen after staging the atom in its human-authored commit/integrated tree. A worker's dirty-tree exit is not codegen drift evidence by itself.
  • A gate transcript is evidence only when it carries the exit status of the GATE process. Never spell it as command | tail; echo $?: that status belongs to tail, and an A6A codegen run on 2026-08-28 printed zero over an actual os error 5 publication failure. Run the gate unpiped, or redirect its output, capture the child status immediately ($? / $LASTEXITCODE / PIPESTATUS[0] as appropriate), and inspect the saved output afterward. PowerShell cmdlet/conversion errors are non-terminating by default, so a gate such as XML parsing first sets $ErrorActionPreference = 'Stop' (or uses -ErrorAction Stop); otherwise it can print an error, continue into a hand-written “clean” line and even commit invalid XML — observed and repaired on 2026-08-28. A green-looking suffix with a masked parent failure is no verdict.
  • A worker never runs the full panel. It runs the narrow self-verify named by its packet. The central coordinator does not run the full panel merely because a worker finished.
  • A coherent batch may contain several atomic local Conventional Commits, each accepted by its exact gate. They remain unmirrored until the batch integration panel is green. If that panel finds a defect, repair forward in the same batch and rerun newly relevant narrow tests before the next panel.
  • Run the full CARGO_BUILD_JOBS=4 tools/self-check.sh panel when: (1) a cross-crate integration batch is ready to leave local state; (2) a lifecycle wave or owner scenario is declared green; (3) before mirror/push of accumulated code; (4) before final epic acceptance; (5) a high-risk change has no cheaper complete affected set. One panel may satisfy several triggers for the same unchanged tree.
  • A panel result belongs to the exact tree measured. A later product change invalidates it. A report/docs-only change runs its document gates and may ride the next integration panel, but cannot rewrite an earlier code verdict.
  • Mirror only after the accumulated batch panel is green and the staging allow-list is reviewed. This preserves atomic human-authored commits while amortising the slow panel over useful integration work.

6. Live and historical lane observations

13All dated observations whose anchor or prose names the retired subscription GLM transport are historical evidence only. Their commands, retry advice, quota assumptions, and availability conclusions are non-operative after SUBSCRIPTION-GLM-LANE-RETIRED. Observations about other live lanes retain their ordinary force.

  • 14Record only a live task's exit, durable report, resolved-model metadata and reviewed quality. A ping proves transport, not coding quality; an old successful run does not prove a five-hour quota remains open.
  • Historical external-lane probe. Its durable transport evidence was extracted to the preserved package source; the current rules retain only the conclusions that a ping is not a quality verdict and root acceptance remains mandatory.
  • Retired transport budget tombstone. No current worker budget or continuation rule is derived from this anchor.
  • Retired account-quota tombstone. No current retry, wait, or fan-out rule is derived from this anchor.
  • Retired heartbeat tombstone. No availability probe, scheduled run, retry, or repository work may target the cancelled transport. No matching automation exists on the current Codex host as of 2026-09-08.
  • Historical recovery tombstone. A past successful run does not establish present availability.
  • Historical R4 transport tombstone. The accepted product evidence remains in commits and gates; this anchor authorizes no runner.
  • A later R4.1 T1/T3 pair passed its worker gates yet central review found four identity-boundary defects: five-digit values entered a four-digit TOML date, host provider plus document path was unconstructible, legacy plan paid a temporary collection, and CompiledSelector::Eq treated OR-set order/duplicates as semantic while the promised digest did not. Two same-cwd worker repairs plus an ordinary Opus/max byte-schedule audit closed them before T2; conform also forced the enlarged selector RED into a focused sub-600-line cell. Durable packet/acceptance rule: whenever a canonical identity lands, pair its independent byte vector with Rust equality, absent-versus-empty, typed constructor composition and legacy allocation/order checks. A green digest golden alone does not prove one semantic identity.
  • Two disjoint R4.1 writers proved that write-perimeter independence does not make cross-crate gates temporally independent. The registry writer's downstream compile twice observed the workspace writer mid-move and went red on transient duplicate/missing symbols; after that writer reached a stable file boundary, the same gate passed unchanged. Cargo-slot serialisation prevents concurrent cargo processes, not compilation of a peer's half-written source. Schedule downstream gates after peer writers are stable or classify the exact foreign path before repairing anything. The same run exposed a second trap: cargo test -p vibe-install lifecycle exited 0 while running zero tests; root replaced it with the full package suite. Acceptance records test counts, never exit alone. Finally, one split continuation used git restore despite a no-mutating-git packet to heal its own transient newline collapse; the desired bytes survived, but text prohibition is not enforcement — root still verifies the final ordinary diff and ownership perimeter.
  • Retired multi-lane quota tombstone. No current availability or failover rule is derived from this anchor.
  • Retired initial-stall tombstone. Generic transport failures remain evidence, not code verdicts, under LIMIT-FAILOVER.
  • A plain fresh codexrunner exec loaded repository AGENTS.md before its named worker packet and began the forbidden full boot despite that packet's ##subagent-quiet-clause; root stopped it before product edits. The strict live repair is --strict-config --ignore-user-config -c project_doc_max_bytes=0, with model_reasoning_effort="xhigh" repeated in the later exec layer and shared CARGO_TARGET_DIR set before launch. A no-tool probe returned exact QUIET_OK; the real retry resolved gpt-5.6-sol / xhigh, read the packet first and loaded only its six named standing files. A packet's quiet clause controls reading only when the harness has not already injected project documentation ahead of it.
  • tools/self-check.sh lexically counts run_step lines while seven calls execute dynamically, so final R3.4 printed [54/47] and still ended all green / exit 0. The denominator is an observability bug, never coverage evidence: trust the ordered gates + green tail. A later tooling atom derives the runtime total or removes the false denominator; this docs-only finding does not rewrite the green verdict.

7. Canonical worker packet template

15The central coordinator copies this skeleton into every substantial worker packet and fills every bracketed field. Do not point a worker at PROP-055 itself: this spec is ChatGPT-only. Task-specific instructions override a template default only when the packet says so explicitly.

16# PACKET-<task-id> — <one bounded outcome>

## Role and authority

You are an EXECUTION worker, not the architect, reviewer, integrator or owner.
The central coordinator owns architecture, acceptance, staging, commits and
final gates. Your `PASS` is advisory. Do not claim that root reviewed/accepted
anything. A prior report is a lead, never authority.

On a continuation or follow-up, discard your previous verdict where the
correction packet contradicts it. Re-read the correction packet first and
repair the named failure; do not defend the old answer.

## Task

<exact behavior to implement or question to review>

## Required inputs

- <files/spec anchors/reports to read>
- <facts supplied by root; confirm or refute each with exact evidence>

If an ambiguity would change public grammar, architecture, security posture or
scope, STOP and write the exact blocker. Do not silently choose a new design.

## Write perimeter

- <expected files/directories to change; mark any path intersecting a parallel writer as closed>
- `WORKER-REPORT-<task-id>.md`

Leave every product change as an ordinary **unstaged** working-tree diff in
this cwd. Preserve pre-existing/user changes. If solving the task genuinely
requires another non-conflicting path, change it and explain the expansion in
the report. Do not create/package state in a branch, commit, index, stash,
private worktree, patch-only deliverable or ignored/cache-only product copy;
root already provisioned the cwd. Do not spawn subagents unless the packet
explicitly delegates a nested swarm.

## Acceptance

1. <criterion with observable result>
2. <negative/security/failure criterion>
3. <compatibility/idempotence criterion>

Name the essential ordinary negative-case or invariant check for each
load-bearing branch. If a concrete observation makes a result suspicious, name
one smallest counterfactual diagnostic under `MUTATION-DIAGNOSTIC-LAW`; never
commission an open-ended mutation search. An empty grep/test output is not proof
until a known positive control demonstrates the instrument can fire.

## Constraints

- Use the machine normally: Read/Glob/Grep, PowerShell/Bash, read-only git
  archaeology, formatters, generators and the exact tests are available when
  they help. `git log/show/blame/diff/status` are normal evidence sources.
- Keep the handoff reviewable: NO git add/commit/push/merge/rebase,
  checkout/restore/reset, stash, clean, worktree, tag, config mutation or
  history rewrite; do not leave the only deliverable in generated scratch,
  cache, a patch file or an external service.
- NO full repository panel; run only the exact commands listed below.
- NO specs/status/SPEC-DEBT/WAL/CONTINUE/TASKS/BACKLOG/AGENTS/campaign edits.
- Work inside the assigned task. Network, user-home or external writes are
  allowed when the task needs them, but publishing/pushing or another
  irreversible external action still requires the packet to grant it. Never
  expose a token value in argv, logs or the report.
- Run every repository Cargo command through
  `python campaigns/packages-2026-09/tasks/chatgpt-cargo.py -- cargo …`; the
  wrapper owns the two-slot ceiling and supplies `CARGO_BUILD_JOBS=4`. Do not
  replace it with a process census, which races across workers.
- Preserve the actual exit of every gate. Do not pipe a gate into `tail`,
  `head`, `grep` or a formatter and then report the pipeline's last process as
  the gate result; redirect first and capture the gate child status directly.
- For an intended generated-code diff, follow `##DIRTY-CODEGEN-GATE`: prove
  two emissions identical in the worker and reserve `check-codegen`'s
  HEAD-clean verdict for root after the generated atom is committed.
- NO new `unsafe`, FFI, handwritten OS ABI or home-grown security/locking
  primitive unless the packet explicitly grants that architecture. Stop and
  report the missing primitive instead; do not invent a mini-runtime.
- Every changed/new file remains at or below 600 physical lines. Check while
  editing; split early, not after fmt.
- Do not weaken, delete or rewrite an oracle merely to make it green.
- Treat a non-zero runner exit as transport state; still write durable evidence
  when possible.

## Allowed self-verify

```text
<exact fmt/test/check/clippy/codegen commands; no broader substitutes>
```

## Deliverable report — draft within 140 lines

Do not spend turns repeatedly trimming a long report. Create its concise
skeleton after the first product slice, then update decisions/evidence
incrementally before long parent gates (140-line target leaves headroom under a
150-line hard cap). Never defer the first durable report write to the final two
turns: a turn cap, network loss, or transport failure can end the next call,
and a worker without the report has failed its deliverable even when useful
edits remain.

```markdown
# WORKER-REPORT — <task-id>
## Verdict
PASS | FAIL | BLOCKED (advisory only)
## Changed files
- exact path — why
## Acceptance evidence
- criterion → result, exact file:line/test output
## Suspicion-triggered diagnostics
- suspicious signal → bounded diagnostic and result, or none
## Commands
- command → exit + decisive lines
## Decisions and why
- every choice within granted latitude, or none
## Perimeter expansions / external effects
- disclose every extra changed path, external write, broad test or state effect, or none
## Leftovers
- exact remaining work, or none
```

## subagent-quiet-clause

Your screen text is discarded. Output only mandated `PROGRESS:` heartbeats,
the durable report, and final `TASK-DONE`. No greeting, task restatement,
intermediate chat reasoning or final screen summary.
  • 17Launch with an explicit autonomous non-interactive permission mode and the normal tool surface; do not construct a tiny deny-list as a substitute for review. Read Glob Grep, PowerShell/Bash, read-only git history, formatters and exact tests are legitimate worker instruments. The short pointer prompt repeats only the handoff invariants: keep product work as an unstaged visible diff, do not add/commit/stash/branch/reset/push, do not run the full panel, and write the report before finishing. Root review, not tool starvation, is the safety and acceptance mechanism. Do not set a maximum dollar budget for Opus/Claude runs (owner ruling 2026-08-26); record cost post-facto when JSON telemetry is available.
  • After TASK-DONE, root mechanically compares actual changed files to the report, reads the diff, checks Decisions/Deviations, reruns decisive essential commands and only those diagnostics whose suspicion trigger actually fired, then accepts, sends exact -c repair notes, re-commissions or discards. The packet/template reduces review cost; it never replaces review.

8. Opus 5 max — CLI lane and accumulated lessons

18Opus/max uses the canonical packet above. Root still supplies the idea, accepted architecture, task boundaries and essential verification; Opus implements or independently audits, leaves a reviewable worktree, and receives exact acceptance repairs from root.

8.1 Fresh run and continuation

19Run the installed Claude Code launcher, never a hand-written direct API client. A first turn is fresh claude -p; -c means continue the most recent session in that cwd and is therefore used only for root's repair/acceptance follow-up. Every concurrent Opus thread gets its own cwd so -c cannot resume a neighbour. Do not use --no-session-persistence when a repair loop may be needed.

20# Fresh task, from the worker's dedicated cwd
claude -p $pointer --model opus --effort max `
  --permission-mode bypassPermissions --output-format json

# Same-cwd root correction / continuation
claude -c -p $repair --model opus --effort max `
  --permission-mode bypassPermissions --output-format json

21Omit a fallback model and inspect the JSON result's canonicalModel; never accept Fable. The 2026-08-26 probes resolved --model opus to claude-opus-5 with a 1,000,000-token context. Omit --max-budget-usd by owner ruling: subscription quota, not nominal API accounting, is this lane's resource model.

8.2 Task shape and parallelism

  • 22Prefer Opus/max for wide code archaeology, adversarial architecture/transaction review, ambiguous repairs and implementation whose verification needs sustained cross-file judgment. It may code as freely as any authorized writing worker; root retains architecture and acceptance. Deterministic micro-edits still fail the verification-cheaper-than-generation test.
  • Multiple Opus processes may run concurrently in separate cwd/session pairs. Parallel readers may overlap; writers are separated by intersecting product paths. Its quota is independent of Codex, so using both families at once can expand throughput and supply review diversity.
  • Give Opus the normal machine toolset. The first two audits were unnecessarily launched with a read-only tool profile; one sensible PowerShell line-count command was denied. Both still completed, but the owner corrected the policy: PowerShell/Bash, Grep/Glob/Read, read-only git history and exact tests are ordinary reasoning instruments. Protect only the reviewable worktree handoff.

8.3 Live measurements — 2026-08-26

23Two Opus/max read-only audits ran concurrently through Claude Code 2.1.220. R3 verifier/trace: terminal success, 91 turns, about 9.2 minutes wall, 37,491 output tokens, nominal telemetry $6.57, 215-line report. R4 staged transforms: terminal success, 94 turns, about 9.8 minutes wall, 41,137 output tokens, nominal telemetry $11.01, 236-line report. Both wrote only their report and produced concrete file/line matrices, red proofs and commit slices. The nominal dollar fields do not change the subscription-lane ruling.

24A broad Opus audit genuinely benefits from a 200–240-line cap; implementation reports stay near the canonical 140-line target. Require the report before final completion because screen output is not the handoff. JSON telemetry is kept for model/turn/time observation; it does not replace the durable report.

25An A5a Opus/max implementation reached a green exact test, wrote the complete product diff, then produced no non-telemetry session event for about fifteen minutes while its final max-reasoning turn remained open. Root verified liveness from the Claude session JSONL, stopped only that exact marked claude.exe PID (never by name or tree), and resumed with claude -c in the same cwd. The session retained the diff and prior tool result; the continuation applied the acceptance repair and finished 13 focused + 42 parent tests. Therefore a killed/stalled final turn is not a code verdict and does not require regeneration: inspect the durable diff/report and last non-telemetry event, preserve them, then prefer same-cwd continuation. The exact-PID/no-tree rule applies to ordinary Opus too, even though its process family is not Codex's GUI family.

26The A5b implementation ignored its exact-gate budget at the tail and ran full vibe-cli sweeps through cargo test | grep/head/awk; the first spent roughly nine minutes inside unrelated cli_spec_format local install/submodule work. The rerun eventually printed 635/0, but the pipeline had discarded Cargo's exit status, so that number remained telemetry rather than root acceptance. Root used only unpiped, terminating exact e2e/golden/unit commands plus strict touched-crate clippy/workspace check/conform for the verdict. A worker package sweep is neither a substitute for the central full-panel trigger nor permission to mask the child status; when an overbroad sweep enters an unrelated slow test, stop its exact process chain and return to the packet's named tests.

9. History

27Authored from the owner's 2026-08-26 rulings: use the former external GLM lane and Opus/max alongside Sol/native workers because quotas are separate; never select Fable; consolidate the strategy in one ChatGPT-only document; keep the agent-instruction triplet byte-identical; replace per-atom full panels with exact gates and batch panels; trust both launchers with normal machine tools; preserve their result as an unstaged, centrally reviewable worktree rather than a worker commit/stash/index; and never cap subscription runs by nominal dollar budget.

28The owner removed native Sol Ultra subagents from the acceptance pool after observing that every reviewer repays the enormous boot-lane prefix. Central Ultra now accepts in its existing context; independent review is routed through subscription launchers or another explicitly cheaper external lane.

For an agent

This page has a machine mirror. The citation carries the version rather than latest, so what an agent quotes does not move under it.

spec://org.vibevm.core/vibevm@1.0.0/common/PROP-055-chatgpt-campaign-execution

.md.xmlllms.txt