I built my own agentic harness. Here is what's inside.
A while back I asked a coding agent to add a small feature to a service. It wrote the code, ran the tests, and told me everything passed. Then I read the diff. Two assertions in the test file had been loosened to match what the code did. The agent had not fixed a bug. It had moved the goalposts.
The model did what the setup rewarded. One agent held every role. It wrote the spec in its head, wrote the code, wrote the tests, and graded its own work. Nobody in that session had a reason to disagree with anybody.
That session is why I built Pebble, the agentic harness I now use for my software work. This post covers what a harness is, why I wrote one instead of adopting one, and how the eight agents inside Pebble hand off work to each other. I will stay at the architecture level. No code, no prompts.
What a harness is
A harness is everything around the model that decides what it sees, what it may touch, and what happens after it answers. The model is the engine. The harness is the chassis, the brakes, and the steering.
It decides which agent runs and when. It decides what each agent can read and write. It decides hoA while back I asked a coding agent to add a small feature to a service. It wrote the code, ran the tests, and told me everything passed. Then I read the diff. Two assertions in the test file had been loosened to match what the code did. The agent had not fixed a bug. It had moved the goalposts.How much context each one carries. It decides which steps need a model at all. And it decides when a human has to say yes.
A coding tool on its own gives you a capable agent in a loop. A harness gives you a process. The same model behaves very differently inside one.
Why a harness is worth having
A single agent session fails in predictable ways, and each failure has a structural fix.
The first is self-grading. When the author of the code is also the judge of the code, a pass means little. The fix is separation of duties: the agent that writes tests is not the agent that writes the implementation, and neither one signs off on its own work.
The second is context decay. A long session fills up with stack traces, test output, and file dumps. The original request sinks under the noise, and the agent starts answering the most recent thing it saw instead of the thing you asked for. The fix is giving noisy work its own context window.
The third is scope drift. You ask for a login fix and get a refactor, a new dependency, and a renamed folder. The fix is permissions that match the role. An agent that only writes specs has no write access to source code.
The fourth is cost with no ceiling. An agent that cannot make a failing test pass will keep trying, and every attempt is billed. The fix is a hard cap on retries and a deliberate stop.
The fifth is that nothing repeats. The same request takes a different path every time. The fix is a fixed pipeline, so that when a result is bad, you know which stage to blame.
I wrote about this discipline for agents you ship in Your agent works. That's not the same as production-ready. A harness applies the same thinking to the agents that build your software.
Why I built my own
I looked at what was available before I wrote anything. There is no shortage of harnesses, and some of them are good. Four constraints kept pushing me toward building.
Token budget had to be a design input. Most setups treat cost as something you read on the invoice. I wanted every phase to justify its tokens up front, which meant deciding early which steps get a model and which get a shell command.
Separation of duties had to be enforced by permissions. You can write "do not edit the tests" in a prompt. Under pressure, with a red test blocking the finish line, the agent will find a reason. If the agent has no write access to the test directory, the argument never starts.
The process had to survive a change of tools. Models and coding tools change every few months. I did not want my way of working married to one of them. Pebble defines the roles, the order, and the gates once, and each host gets a thin adapter.
Humans had to own the irreversible actions. The agents prepare. I commit.
Most of what I evaluated optimizes for how much an agent can do unattended. I wanted the other dial: how predictable the result is when I come back to it. And when something misbehaves, I want to know which part to open.
One idea in Pebble is not mine. Interleaving fixed deterministic steps with agent steps in a set blueprint comes from Stripe's write-up on how their internal coding agents work. I adapted it for a much smaller scale.
Pebble at a glance: the eight agents
Pebble is a pipeline of eight agents. Each is named after a stone, which began as a memory aid and stuck. Each owns one phase of the work, from a vague requirement to updated docs, and each has one thing it is not allowed to do.

The order is fixed: Flint, Slate, Mica, Jade, Onyx, Quartz, then Coral and Chalk. Before any of it runs, the request gets classified. A typo, a refactor, an infra change, a bug fix, and a new feature each get a different slice of the pipeline. Only a real feature runs all of it.
The agents are tuned for their role. The ones that judge and verify run with the least randomness I can give them. The one that interprets vague requirements runs a little warmer.
How the agents hand off work to each other
The agents never talk to each other. They read what the previous agent left behind. Flint leaves a spec. Slate reads it and leaves a blueprint. Mica reads both and leaves failing tests. Jade reads all three and leaves code. No agent sees another agent's reasoning, only its work product.
This gives me an audit trail for free, because every handoff is a document. It also means I can rerun or replace one agent without disturbing the rest.
Two handoffs are gated by me. I approve the spec, and I approve the blueprint. After that, the pipeline runs without stopping until the final handover. The placement is deliberate. A wrong line of code gets caught by a test. A wrong requirement does not, because every later agent will faithfully build the wrong thing. The spec is the most expensive place to be mistaken, so that is where a human looks.
The first four steps follow test-driven development on purpose. Mica writes tests against a blueprint for code that does not exist yet, then confirms they fail for the right reason: a missing implementation, not a typo or a broken import. Jade gets a fixed target. Since Jade has no license to edit those tests, getting stuck produces a report that says blocked, not a quietly weakened assertion. That is the failure from the opening of this post, closed off at the permission level.
Onyx runs security first and quality second. If the security pass finds something severe, comments about naming are moot, so the quality review never starts. Findings go back to Jade with exact fixes. I merged security and review into one agent because two agents meant two cold starts and two passes over the same files.
Quartz is the last gate, and it takes nobody's word. It does not trust the claim that tests pass. It runs the full suite twice, and a test that passes once and fails once is a blocker. It scans the output for runtime errors that tests tolerate. Where a project has browser tests, it runs those too, and any failure blocks the handover.
Coral and Chalk are the only pair that run in parallel. Both depend on the finished code, and neither depends on the other. Coral only runs when the change adds a deployable service, new configuration, or a deploy change. Otherwise, it is skipped and only Chalk runs.
When the orchestrator briefs any agent, it passes four things: why the task matters, which artifacts to read, what done looks like, and what is off limits. A vague handoff produces vague work.
Context: what each agent carries
Every agent in Pebble works in its own context window. This is the biggest single lever on both quality and cost.
Six of the eight run in isolation: Slate, Mica, Jade, Quartz, Coral, and Chalk. Each of them is noisy. Slate reads half the codebase to find the existing patterns. Mica and Jade produce test output and stack traces by the page. Quartz runs the suite twice and captures all of it. Chalk reads the whole repository to write accurate docs. If that output landed in one shared session, the conversation would be mostly noise by the time review started.
The main session, the orchestrator, stays small. It holds the original request, a status line or two per phase, the paths of files created or changed, and the gate decisions. Each isolated agent does its work, returns a short structured report, and its working memory is thrown away.
Two agents stay in the main session. Flint stays because it has to ask me questions when a requirement is ambiguous, and a conversation does not work through a closed door. Onyx stays because its output is a verdict, and I want the verdict next to the decisions it affects.
Hosts name this differently. One calls it a forked context; another calls it a subagent mode. Pebble's role definitions say which agents need isolation, and each adapter translates that into whatever the host supports.
One rule keeps the structure honest: agents cannot spawn other agents. Nested delegation makes cost and failure paths impossible to reason about. There is one orchestrator and one level of delegation.
Deterministic nodes and token cost
Not every step needs a model. Linting, formatting, type checking, listing files, reading git state, and checking that a frontend actually builds are predictable commands with known outputs. The orchestrator runs them directly. No agent, no tokens. My estimate is 15 to 20 percent saved per feature, and the results are more reliable because a linter does not have moods.
The frontend check deserves a sentence. Before any review agent spends tokens, the harness builds the app, serves it, and requests the pages. A blank page or a crashed build gets caught for the price of a shell command.
The rule of thumb is simple. If a step has a predictable command and needs no judgment, it is not an agent's job. That is the same split I wrote about in Neuro-Symbolic AI: When the LLM Doesn't Get the Final Say: the model handles judgment and deterministic code handles everything else.
A few smaller choices add up. Agents load only the standards that match the language they are touching, so a Go change never pays for Python conventions. Examples and shared patterns load on demand instead of living in every prompt. Security and review share one agent. The agents that analyze run cold.
The pipeline also classifies work before it starts. Spending the full eight-agent run on a typo would be a waste, so small changes take a short path that still ends in verification and review.
Where Pebble stops on purpose
The loop of implement, review, and test runs twice at most. If the work is still failing after the second round, the pipeline stops and marks it as needing a human. I get the failing tests, the build logs, and the state of the branch. Over the past two rounds, the agent is rarely converging. It is circling, and each lap costs money.
The agents never commit, push, merge, or rewrite history. They can read git state, and they stash my uncommitted work before a run starts so nothing gets mixed. The orchestrator can prepare a commit message and show it to me. I read the diff and run the commit myself. A harness that can push on its own is one I would not leave running.
This is the same blast radius thinking I described in Your Agent Has More Access Than Your Junior Developer. That's a Problem. Every agent gets the least access its job needs, and the irreversible actions stay with a person.
One harness, five hosts
Pebble runs on Claude Code, OpenCode, Antigravity, and Codex. The roles, the order, the gates, and the handoff artifacts are identical on all of them. What changes is a thin layer: how an agent is declared, how isolation is requested, how a command is invoked, and how permissions are written.
The OpenCode version started as a port of the Claude Code one. Putting the two side by side showed that they differ in naming and configuration and almost nowhere else. That comparison is what convinced me the process belongs in the harness and not in any single tool.
Models will keep changing, and so will the tools wrapped around them. The sequence of spec, blueprint, tests, code, review, verification and docs does not depend on either.
Where it falls short
Pebble is heavy for small changes. The classification step and the short path exist because I got tired of paying for eight agents on a one-line edit, and the short path is still a compromise.
The ceiling is the quality of the spec. I gate it because it matters most, but if I approve a sloppy spec, the pipeline builds it with great discipline. A harness makes a process repeatable. It does not make the input good.
The sequential chain costs wall-clock time. The dependencies between stages are real, so there is little to run in parallel beyond the final docs and infra pair. I accepted slower runs in exchange for results I can trust without rereading every line.
And it depends on the repository. Mica needs a test setup to write into, and Quartz needs one to run. Where a project has none, the first feature through the pipeline pays for building it.
Found this useful? I do 1:1 sessions on AI architecture and strategy. → Book a session
Found this useful? I do 1:1 sessions on AI architecture and strategy. → Book a session