Orchestrate Vol. 1 · Agent Orchestration · Est. 2026

An Agent Skills package · MIT Licensed

Orchestrate
the whole crew.

One coordinator. Many agents. Orchestrate splits a job across live-verified coding-agent runtimes and in-session subagents — routing each task by capability and risk, isolating parallel writers in their own git worktrees, capturing every command, and refusing to call the work done until deterministic checks pass and — for anything above read-only or scoped-write work — an independent arbiter agrees.

claude code
› /orchestrate "research three caching strategies and compare them"
› /orchestrate "refactor the session API" --internal
› /orchestrate plans/jobs.yaml --yes
› /orchestrate --resume plans/reports/orchestrate-260723-1442

  → live runtime probe · 4 candidates verified
  → 6 jobs · 2 stages · concurrency 2 · 3 worktrees
  → arbiter: independent C3 route (illustrative; family verified live)
01

Parallel agents are easy.
Trustworthy ones are not.

Fanning out work to several agents takes one loop. What it costs you is everything that makes the result believable: two agents editing the same file, a runtime that silently vanished, a model name that stopped existing last week, a permission prompt swallowed by a headless process, and a cheerful summary claiming success nobody verified.

Orchestrate treats those as the actual problem. Every route is resolved from live evidence at execution time — never from a catalog written into a file. Every parallel writer gets its own worktree. Every job records its command, its exit state, and its artifacts. And nothing is reported as finished until deterministic checks pass. Read-only and scoped-write work may be accepted on a calibrated micro-arbiter verdict; everything above that goes to a separate arbiter, on an independently selected route, to check the claims against the evidence.

"Runtime and model catalogs drift. Resolve every route from live execution-time evidence. Never treat a runtime, provider, model, alias, flag, or agent seen in this file as currently available."
02

The seven-stage pipeline

Each stage has an owner, a gate, and an artifact. A stage that cannot prove its evidence blocks the run rather than guessing past it.

01

Brainstorm & intake

Clarify the outcome, constraints, non-goals, and acceptance evidence. Refuse any plan that would put secrets or private data into prompts. If orchestration adds no real parallelism or arbiter value, say so and recommend a single agent instead.

02

Build the job graph

Turn the outcome into jobs with explicit task, cwd, timeout, expected output, and file ownership. depends_on forms the stages. Same-stage jobs run concurrently only when their ownership and outputs do not overlap.

03

Discover, profile, route, onboard

Build a live runtime inventory, profile each candidate's permissions, isolation, capture and budgets, then resolve a route by capability and risk tier. An optional System-1 layer may supply scored signals, and it may only raise a floor, never lower one. A missing, unauthenticated, or under-controlled candidate cannot satisfy a route. A pinned runtime that is missing is onboarded as a visible setup step.

04

Apply the safety gate

Confirm every cwd, writable root, and expected side effect. Least privilege by default; a permission bypass is never enabled — a job that needs more privilege gets scoped permissions with explicit approval, a stronger external boundary, or it is blocked. Destructive, deployment, release, or credentialed work needs approval for that exact scope.

05

Dispatch, observe, verify

Create worktrees before dispatch, start independent jobs up to the concurrency limit, and update state.json on every transition. Capture the redacted command, bounded stdout/stderr, exit status, wall time, and artifacts, and observe every attempt until it settles.

06

Arbiter review

A separate judgment route — preferably a different model family — compares each escalated result against its expected output, runs the declared checks, and flags contradictions, unsupported claims, and missing artifacts. R2, R3, any job whose artifact is a verdict on another job's work, and any attempt whose tier a semantic signal raised always escalate here. Read-only and scoped-write work may clear on a calibrated micro-arbiter gate only when no risk-floor raise applied. With no different-family route the verdict is labeled not-independent and the job is blocked, unless a fresh, independently configured context substitutes.

07

Report

One report.md with per-job status, resolved runtime and model, tiers, artifacts, errors, the arbiter verdict, reproduction commands, pending worktree diffs, and the unresolved questions — listed plainly, last.

03

Routed by capability, gated by risk

Two independent axes. Capability asks how much judgment the task needs; risk asks how much damage a wrong answer can do. A job runs only where both floors are met — otherwise it is marked blocked, never quietly downgraded.

Capability tiers

C1Throughput. Accurate search, extraction, summarisation, bounded repetitive changes — scout, docs, mechanical fan-out.
C2Delivery. Multi-file implementation judgment, test design, failure-path handling — normal implementation and tests.
C3Judgment. Deep trade-off analysis, conflict resolution, security reasoning, independent arbitration — architecture, review, audit, arbiter.

Risk tiers

R0Observe. Read and report only. Explicit cwd, bounded timeout, captured result, no unnecessary write or shell grant.
R1Scoped write. Reversible edits in owned files. Scoped write boundary, tool restrictions, diff capture, no permission bypass.
R2Isolated write. Parallel, high-impact, untrusted, or hard-to-revert changes. Separate worktree or stronger isolation, enforced sandbox where available, explicit checks and arbiter review.
R3External / destructive. Deploy, release, delete, credentialed, or other external side effect. Explicit approval, rollback plan, strongest verified controls — blocked when controls are unavailable.
04

A probabilistic layer that may only tighten the gate

A small classifier can watch traces, triage failures and pre-score an attempt before the arbiter. It is optional, provider-neutral, and it is allowed to make exactly one kind of change: raise a floor.

What it may do

01Watch traces continuously and flag a stalled or drifting attempt.
02Triage a failure into the failure taxonomy.
03Supply three scored signals to the pre-arbiter gate.
04Raise a floor, recorded as capabilityFloorDelta or riskFloorDelta.

What it may never do

01Decide eligibility, a tier, or a control. The hard filter and the safety gate own those.
02Lower a floor, restore a rejected candidate, or widen the no-C3 path.
03Authorize its own egress. With no recorded user authorization scope, the plane is disabled for the run.
04Accept an attempt. Only the escalation matrix does that.

A riskFloorDelta changes the recorded tier, adds controls, and always escalates to C3 — so a classifier can shrink the fast path, never extend it.

The no-C3 path needs 30 C3-audited outcomes, a measured agreement of at least 0.95, and a threshold of at least 0.90. Those are policy constants, not job inputs, so a spec cannot set its own bar. With no valid record the plane observes and logs only — it never decides.

The same discipline applies to runtimes. A file under runtimes/ is a note, not a support claim. Only a live probe recorded in runtimes.json makes a runtime selectable, and only a verified available state makes it eligible for routing.

Jev (TypeSafe), the reference model

01Typed decisions, not text. Jev is TypeSafe's System One model: it answers a declared set of typed questions about a piece of state and returns a probability for each answer, so no generated text has to be parsed back into a policy field.
02A distribution, not a verdict. A choice comes back with the runner-up probabilities beside it, so the plane can return a low-confidence signal instead of a confident wrong answer — the case the escalation matrix exists to catch.
03Reference, never a requirement. The contract was written against Jev, and Jev is one optional provider behind it. With no eligible classifier, the plane is disabled and deterministic policy decides.

Configuring the key

01The variable is TYPESAFE_API_KEY, created in TypeSafe's own console. It is read from the four locations documented in decision-plane.md, which owns the order.
02A key alone is not enough: a call also needs a recorded egress authorization naming the provider and the credential source. With no key, or no such authorization, the plane is disabled for the run — never silently re-routed.
03The value is never printed, prompted for or committed, and a shadowed source is reported. Nothing on this page, and nothing in a run, shows it.
05

Defaults that assume something will go wrong

Worktree isolation

One worktree per isolated job, branched from the accepted base ref. Never shared, never silently reused. Integration is coordinator-owned and happens after the arbiter pass — a worktree defers collisions, it does not resolve them.

External timeouts

Every CLI process gets a timeout enforced from outside. An internal timeout is treated as accounting only, unless the harness has proven it actually cancels the work.

Resumable runs

An interrupted run reloads its spec and state, keeps successful outputs, preserves prior attempts, revalidates every live route, and redispatches only the jobs that were interrupted.

Fail loud, never guess

A timeout, a permission prompt, an unknown flag or model, or a failed check is a failure. No silent retry, no substituted model name, no relabelled output.

Redacted capture

Secrets, tokens, keys and dotenv values are refused at intake and redacted from prompts, commands, logs and reports. Capture stays inside the run directory.

Advisory metrics

Each finished job appends one JSON line of history. Metrics can suggest a policy change with evidence — they can never authorise one automatically.

06

A job graph, in plain YAML

Write it once and rerun it. Free-form requests get the same treatment — the coordinator writes the spec first, then dispatches from it. Placeholders are resolved from live evidence at run time, never from memory.

jobs.yaml
version: 1
concurrency: 2
jobs:
  - id: scout-session-api
    runtime: internal
    task: scout
    cwd: <workspace-root>
    prompt: "Inspect the session API
      and report extension points."
    timeout: 10m
    expected_output: "Files read +
      recommended seams."

  - id: independent-review
    runtime: <verified-cli-runtime>
    fallback_runtime: [<verified-fallback>]
    task: review
    depends_on: [scout-session-api]
    importance: high
    isolation: worktree
    timeout: 10m
    expected_output: "Verdict with checks
      and unresolved risks."
plans/reports/orchestrate-<timestamp>/
orchestrate-260723-1442/
├── jobs.yaml          # private resolved input
├── state.json         # attempts + acceptance
├── metrics.jsonl      # per-attempt outcomes
├── runtimes.json      # live probe evidence
├── decisions.jsonl    # decision traces
├── calibration.json   # arbiter calibration
├── trace.jsonl        # correlated record
├── report.md          # the deliverable
├── worktrees/
│   └── <job-id>/      # isolated writes
├── supervisor/
│   └── <run-id>/
│       ├── events.jsonl # normalized journal
│       ├── graph.json   # private invocation
│       └── output-<job-id>.log # bounded, redacted
└── <job-id>/
    ├── command.txt     # redacted, CLI jobs
    ├── stdout.txt      # bounded
    ├── stderr.txt
    ├── result.md       # internal + native jobs
    ├── session.json    # Pi session handle
    ├── status.json     # resolved route
    ├── native-<attempt-id>.json  # dispatch receipt
    ├── artifacts/
    └── attempt-<n>/    # preserved failures

A summary of the normative run-directory tree in output-layout.md, which owns the layout.

07

What this release adds

Four behaviours, each with the bound that keeps it honest.

Benchmark-ranked routing

Routing ranks the candidates that already passed the hard filter, by measured success rate, cost per task and task duration at a given reasoning effort. Evidence is cached between runs. Benchmarks rank and never gate: they cannot set eligibility, a floor, a tier, a control or an approval.

Promotion and fail-safe

A quota limit, an outage or a crash promotes the job to the next candidate, after the retry budget and honouring a declared order first, ending in a logged blocked state if none is left. A failed check never promotes, and neither does a permission stop.

Trace and logs

Every routing decision, promotion, gate outcome and verdict carries one correlation identity, and fetches are traced like decisions. The trace is redacted on write, excluded from exports unless reviewed, and a run whose trace is incomplete says so.

Credentials

The provider key is read from four documented locations in a fixed order, passed only through the inherited child environment, and never printed or prompted for. A missing key disables the decision plane rather than failing the run.

08

How a job actually moves

These are the routing hops of the same pipeline listed above as stages, so the count differs from seven.

RequestuserDecision planeapproverRoutingorchestratorTerminalsystemRequestScored needsServiceProbeProcessFetchProcessHard filterDecisionRankProcessRouteProcessNo candidateServiceStartFailureDecisionProcess
Routing half: probe, fetch, filter, rank and route, steered by the decision plane, with a no-candidate exit.
Safety gateapproverExecutionworkerReviewapproverTerminalsystemGateDecisionDispatchProcessObserveProcessCheckDecisionPromoteWaitingRe-enter routingProcessVerdictDecisionReportFail-safeSuccessFailureWaitingDecisionProcess
Gate, execution and review: the safety gate authorises, a settled check reaches a verdict, infrastructure failure promotes, and an exhausted budget fails safe.
  1. Probe: live inventory of the runtimes on this machine.
  2. Fetch: benchmark evidence for each surviving candidate and effort level.
  3. Filter: eligibility, capability floors and risk floors.
  4. Rank: order the survivors by measured outcome.
  5. Route: select the runtime, model and reasoning effort.
  6. Gate: tier, controls, approval and egress authority.
  7. Dispatch: run the job in its isolated worktree.
  8. Observe: capture events, output and settlement.
  9. Check: deterministic verification of the work product.
  10. Verdict: independent arbiter review where the tier requires it.
09

Install

Orchestrate runs on any harness that has Agent Skills installed. The primary path is the npx skills CLI; the Claude Code marketplace path and the manual copy both still work.

Install with the skills CLI

npx skills add bestagentkits/orchestrate installs for the current project; add -g for a global install or -a <harness> for one named harness.

Add the marketplace

Point Claude Code at this repository once.

Install the plugin

Pull orchestrate from the marketplace you just added.

Restart, then run it

Restart Claude Code so the skill registers, then invoke /orchestrate with a job spec or a plain-language request.

Or copy the skill directly

No plugin needed — drop the skill folder into .claude/skills/ in your project, or ~/.claude/skills/ for every project.

install
# 1 · add the marketplace
› /plugin marketplace add bestagentkits/orchestrate

# 2 · install
› /plugin install orchestrate@orchestrate

# 3 · restart Claude Code, then
› /orchestrate "your request"

# — or, without the plugin system —
$ git clone \
    https://github.com/bestagentkits/orchestrate.git
$ cp -R orchestrate/plugins/orchestrate\
    /skills/orchestrate ~/.claude/skills/
10

What it is not

Stated plainly, because an orchestrator that oversells its guarantees is worse than no orchestrator.

Not a daemon. No scheduler, no dashboard, no account pool, no provider adapter. It coordinates runtimes that already exist on your machine.

Not a sandbox. A git worktree prevents edit collisions between agents. It does not isolate processes, the network, or the filesystem.

Not shared memory. Jobs share nothing implicitly. Anything a downstream job needs must travel through an explicit dependency.

Not a stable catalog. CLI commands, models, auth and safety behaviour drift constantly. Every run revalidates them from scratch — that is the point.