Multi-agent harness

A role does one job in the pipeline, and each role points at a runner: a CLI, a model and its permission flags. Stages are assigned to roles, never to vendors. Assign no stages and nothing changes: the main agent does everything itself.

Why three layers

stage  →  role       →  runner
testing   tester        gemini (CLI + model + permission flags)

Every model is strong at something different: breadth of research, prose, code, review. A role is how you say “this job needs that strength” without hard-wiring a vendor’s name into the pipeline.

Why “resume the latest session” is never needed

kanban-flow keeps its state in files: .works/, the artifacts and .kfw.json. It never relies on what a conversation remembers. A worker for a stage is a fresh headless session whose context is that stage’s skill plus the artifacts on disk. A session is kept per work item and per role, so the repair loop from FAIL back to implement continues in the session that did the work, and that id is either minted by kf or fetched deterministically, never guessed as “the latest”. Two roles sharing one runner still get two separate sessions, so coder and reviewer on the same CLI never see each other’s context.

Config

The harness block in .kf/config.json, which kf harness displays. This example is a hand-tuned one, spreading roles across several runners. What kf init seeds is deliberately plainer: all six roles on the single agent you chose, and "stages": {} — so nothing is routed to a worker until you assign it:

"harness": {
  "main": "architect",
  "roles": {
    "architect":  "claude-opus",
    "researcher": { "runner": "codex",  "brief": "Explores breadth: prior art, libraries, comparable features.", "output": "research.md" },
    "writer":     { "runner": "gemini", "brief": "Turns agreed decisions into precise prose." },
    "coder":      "claude-opus",
    "tester":     { "runner": "devin",  "brief": "Runs the real suite and records exact commands and exit codes." },
    "reviewer":   "codex"
  },
  "stages": {
    "brainstorm":     ["researcher", "writer"],
    "implementation": "coder",
    "testing":        "tester",
    "review":         "reviewer"
  },
  "runners": {
    "claude-opus": {
      "start":  ["claude", "-p", "{prompt}", "--session-id", "{session}", "--permission-mode", "acceptEdits", "--output-format", "json"],
      "resume": ["claude", "-p", "{prompt}", "-r", "{session}", "--permission-mode", "acceptEdits", "--output-format", "json"],
      "session": "provided",
      "usage": "json"
    },
    "codex": {
      "start":  ["codex", "exec", "--json", "{prompt}"],
      "resume": ["codex", "exec", "resume", "{session}", "{prompt}"],
      "session": { "stdout": "\"thread_id\":\"([^\"]+)\"" },
      "usage": "json"
    },
    "devin": {
      "start":  ["devin", "-p", "{prompt}", "--permission-mode", "accept-edits"],
      "resume": ["devin", "-p", "{prompt}", "-r", "{session}", "--permission-mode", "accept-edits"],
      "session": { "command": ["devin", "list", "--format", "json"], "idField": "id", "matchField": "title" }
    },
    "gemini":   { "start": ["gemini", "-p", "{prompt}", "--approval-mode", "auto_edit"] },
    "opencode": { "start": ["opencode", "run", "{prompt}"] }
  }
}

Adding a model means adding a runner and pointing a role at it. Two models on one CLI, such as claude-opus and claude-haiku, are two runners, assigned to the expensive role and the cheap one.

Runner fields

Field What it means
start The argv for a fresh session. {prompt} is mandatory. {session} is allowed only under session: "provided", where kf mints a UUID before running
resume The argv to continue a stored session; it must contain {session}. Without it the role always starts fresh
session "provided", { "stdout": "<regex, group 1>" }, or { "command", "idField", "matchField" }, which runs a command returning a JSON array and picks the entry whose matchField contains the kf-run:<id> marker kf puts at the head of the prompt
usage "json" parses usage.input_tokens and usage.output_tokens, plus total_cost_usd when present, from JSON on stdout
skillsDir The runner’s skills directory. Defaults follow kf install: .claude/skills, .agents/skills (codex), .gemini/skills, .kiro/skills, .cursor/skills, .opencode/skills, and any unfamiliar name falls back to .agents/skills
resumeFailure A regex that recognises a failed resume; the default is session\|not found\|no such\|unknown\|does not exist

Presets, as verified on 2026-09-19

Runner Session What was verified
claude kf mints the UUID via --session-id, resumes with -r; usage and cost come from the JSON Run for real: start and resume both work
codex thread_id from the JSONL of exec --json, resumes with exec resume <id>; usage from turn.completed Run for real: start and resume both work
devin devin list --format json matched on title against the marker; resumes with -r <id> Run for real: start, list and resume all work. Devin refuses an untrusted directory, so open devin interactively once inside the repo
gemini no resume Could not log in on the verification machine. -r accepts latest or an index; whether it accepts a UUID is unknown
opencode no resume How to obtain the id is unverified; run -s <id> does exist

Flow

  1. Before each stage, the main agent, running the kanban-flow skill, reads kf status --change <f> --json. When assignedRoles is present it runs kf run <f>, adding --detach for a long stage such as implementation and then polling kf runs <f>.
  2. kf run resolves the stage’s role chain and runs it in order. Each role’s prompt carries: the role and its brief, the work item, the stage, the folder, the path to that runner’s phase SKILL.md, pointers to kf status and kf instruct, the requirement to write output when the role has one, a “Previous step” section naming the previous role, its output file and its log from the second step onward, and the contract: do only this stage’s work, never run kf stage, kf approve, kf archive or kf run, never edit an approved contract, never commit, and finish with STATUS: DONE | DONE_WITH_CONCERNS | BLOCKED | NEEDS_CONTEXT followed by Summary:.
  3. Each role is its own run: its own process group, its own runs/<id>.log, and one entry in .kfw.json.runs[] carrying both role and runner.
  4. A role that does not end in DONE or DONE_WITH_CONCERNS stops the chain. The roles after it do not run, kf run exits 1, and it names which role stopped and which were skipped.
  5. The main agent reads STATUS and Summary from the tail of the log, without swallowing the transcript, runs kf validate and decides on the transition. The gate does not take the worker’s word for anything.
  6. On the repair loop, the next kf run resumes that role’s session for that work item, and the prompt gains the path to the current FAIL report. --fresh forces a new session.

kf run --role <name> runs exactly one role from the chain. --dry-run prints the plan for the whole chain, one argv and prompt block per role. The timeout is 30 minutes by default (--timeout <minutes>, 0 for none); on expiry the whole process group is killed and the run is marked timeout. There is no automatic retry: a failed resume resets the session exactly once and starts fresh. A work item has at most one running run at a time.

With --detach, kf writes the chain plan, spawns a detached kf run --supervise <id> as the supervisor and returns immediately. The supervisor runs the whole chain and finds the folder again by work item name before writing results; when the folder is gone it writes .works/harness/orphan-<id>.json. kf runs reports failed (supervisor lost) when a record is still running but its pid is dead, and it never repairs the metadata on its own.

Observability

Typing a runner name into stages gets a message that says so, rather than “unknown role”:

Invalid project config: .kf/config.json — harness.stages.testing "gemini" is a runner,
not a role; declare a role in harness.roles that points at it

Token cost

Limits

Terms of use

kf only spawns each vendor’s official CLI using the headless flags that vendor documents (claude -p, codex exec, gemini -p, devin -p, opencode run), under the account already logged in on the machine running it. kf never reads, stores or forwards any CLI’s credentials or tokens, never calls a vendor API directly and never retries in bursts. Each user remains responsible for the terms of the plan they are on, covering commercial use, account sharing and rate limits. This document is not legal advice.