Workflow gates for AI coding agents

Your agent says the tests passed.

Sometimes it is right, and you cannot tell which time from the transcript — because the transcript is written by the same thing you are checking. kanban-flow moves the proof into files on disk: the gates do not read the agent's summary, they read .works/.

$ npm install -g @phuthuycoding/kanban-flow

brainstorm planning implementation testing review dones

A gate is a file that exists, or does not

Every work item walks one pipeline, and only the CLI moves it between stages. Most stages owe an artifact, and a transition is refused when that artifact is missing, empty, still full of template placeholders, or carrying a real secret.

Refusals, not warnings

A testing report that does not say PASS does not reach review. A review report that does not say PASS does not archive.

Three decisions stay human

Confirming the requirement, approving the execution contract, and choosing to start now or hold in backlog. The tool never takes those.

Approval bound to the bytes

Approval carries a SHA-256 fingerprint of the contract. Edit the plan and it is invalid until planning approves it again.

An override leaves a mark

--force works, and is written into .kfw.json where a reviewer will find it. Nothing is skipped quietly.

Evidence that expires

Most of the mechanics here exist elsewhere. One I have not found anywhere else: every entry into testing mints a fresh execution id, and a report is accepted only when its execution: field matches the current one.

Fix a failure, go round again, and yesterday's green report is inert. It is not deleted, argued with, or trusted less — it simply stops matching, so passing again costs exactly one honest test run.

# first pass$ kf stage x testing → execution run-1$ kf stage x review ✓ report says run-1 # review came back FAIL, code changed$ kf stage x implementation$ kf stage x testing → execution run-2$ kf stage x review ✗ refused report still says run-1

One stage, several agents

Stages route to roles, and roles point at runners. A runner is one way to invoke a CLI; a role is a job such as researcher or tester. Pointing a role at a different model is a one-line edit, and the stage never changes.

Chains, in order

A stage runs its roles in sequence, and a role that declares an output file hands it to the next.

Sessions never cross

Kept per work item and per role, so two roles on one CLI do not share context and a repair loop resumes the session that did the work.

Presets included

claude, codex, devin, gemini, opencode. Assign no stages and the harness stays out of the way entirely.

Start in one command

Needs Node 20 or newer. kf init seeds .kf/, installs the agent skills, and leaves your existing AGENTS.md alone.

npm install -g @phuthuycoding/kanban-flow

cd your-project
kf init                             # seeds .kf/, installs the skills
kf new user-login --context auth    # opens Phase 1 in brainstorm
kf status --change user-login       # the checklist, and what is blocking
kf doctor                           # is the project itself healthy?