▲ Agent-based development
What a harness is and how it makes an AI agent trustworthy.
By César Soto · 6 min read
▲ In this article
▲ In short
- A harness is everything around the model in an AI agent: specs that tell it how to decide, automated reviewers, human gates and tests.
- Developers' distrust of a loose agent is reasonable. A harness turns it into a process the team can see, question and improve.
- After adding an agent that verifies the acceptance criteria of every requirement, the requirements QA used to send back as incomplete dropped to practically zero on our team.
When an engineering team says AI is not good enough yet, they are usually talking about a loose agent with nothing around it checking its work. This article explains what a harness is, what it is made of, and what changed in our process when we added a third automated reviewer.
What is a harness in agent-based development?
Birgitta Böckeler, of Thoughtworks, uses a similar definition in Harness engineering for coding agent users: the harness is everything in an agent except the model. She tells guides, which steer the agent before it acts, apart from sensors, which observe what it did so it can correct itself. She also separates deterministic controls, such as tests, from the ones that use another AI to review the work.
Put simply, the harness is the agent's working system. It defines what the agent knows before it starts, who reviews what it produces, and when a person steps in.
Why do developers distrust AI?
I recently spoke with a CEO whose developers were telling him that AI is not good enough yet. He saw other startups, mostly in the United States, shipping faster with agents, and he was looking for ways to make his own team more efficient.
I think his developers are right to distrust a loose agent, because an agent without review makes mistakes. The useful question is what system surrounds the model and what happens when it fails. If the team can see that system, question it and change it, they have a process to lean on.
What pieces does a harness have?
The harness in our method has four pieces, and they can be built with different tools.
- Specs in the repository. They tell the agent how to decide, and they are updated in the same PR that changes the code.
- Automated reviewers. Other agents audit security, quality and acceptance criteria before a PR is opened.
- Human gates proportional to risk. Product validates what gets built, a person reviews high-risk changes, and QA approves behavior before release.
- Operations that feed back into the flow. Production errors are diagnosed and return to the process as a ticket or a fix PR.
What changes when agents read the code?
For years developers were measured by the code they write, and now agents write a good part of it. That changes who the code is written for: the first reader of specs, conventions and documentation is an agent, which uses them to decide.
That is why these materials stop being paperwork and become part of the work. If a spec is out of date, the agent decides with the wrong instruction, which is why specs are updated in the same PR that changes the code.
How does the acceptance agent work?
In our process, every PR went through two automated reviewers: one for security and one for quality. Even so, QA sent back several requirements because they arrived incomplete, without meeting all of their acceptance criteria. Neither reviewer asked that question, so we added a third.
- Security
- Quality
- Acceptancenew
Repeats until all criteria are met.
The acceptance agent reads the criteria of every requirement included in the PR and looks for evidence of each one, in the code or by running the application. If one is not met, it sends it back to the main agent with which ones are missing and why. The two repeat that loop until all criteria are met, and only then does the human process continue.
| Reviewer | Question it answers | What it returns to the main agent |
|---|---|---|
| Security | Does the change introduce vulnerabilities? | The security findings that need fixing |
| Quality | Does it follow the project's conventions and have tests? | The deviations and the missing tests |
| Acceptance | Does it meet every acceptance criterion of the requirement? | The criteria that are not met, and why |
“Code can be well written and secure and still not do what was asked.”
In Böckeler's terms, the acceptance agent is a behavior sensor. She notes that the behavior harness, the one that checks that a change does what was asked, is the least solved of the three kinds she describes, and in our case that is where the gap showed up.
What was the result?
This is the observation of one team and one product, not a controlled study, so I do not present it as a promise. What it does show is where the gap was: nobody was asking the PR whether it did what the requirement asked.
To measure it in your team you need a starting point. The three indicators I propose in the diagnostic are deliveries per week, time from idea to production, and requirements QA sends back.
Where do you start?
- Pick a bounded flow. One kind of requirement that repeats, not the whole product at once.
- Write verifiable acceptance criteria. Each criterion must be checkable with evidence, not with an opinion.
- Add a reviewer that checks those criteria. It sends back to the building agent what is missing and why, and both repeat until the criteria are met.
- Define what a person reviews. Human review is proportional to risk: low-impact changes do not need it and high-risk ones do.
- Measure before and after. Deliveries per week, time from idea to production, and requirements QA sends back.
If you want to see how each piece fits in the full process, PeakSyn's method describes the four stages and the three human gates with their artifacts.
Frequently asked questions
Does a harness replace developers?
No. The goal is for the team you already have to ship more. Developers run the system and decide at the gates: what gets built, which high-risk changes are merged, and which behavior is approved.
Does a harness depend on a specific model or tool?
The pieces of a harness (specs, reviewers, gates and tests) can be built with different tools and stay useful when the model changes. The technique changes every quarter, which is why the harness is kept up to date instead of being rebuilt.
How long until the benefit shows?
It depends on each team's flow, so I do not promise timelines. What we saw is that the benefits appear when the model is adopted with a good harness and integrated into the development process; in our case it showed in the requirements QA sent back.
What is an acceptance agent?
It is an automated reviewer that reads the acceptance criteria of a requirement and checks, one by one, that the change meets them. If one is missing, it sends it back to the building agent with which ones are missing and why.
Does it work for small teams?
Yes. I work with startups and scale-ups that already have a product in the market and an engineering team. If you are still validating an MVP from scratch, that is not what I do.
References
Want to apply it in your team? Let's talk for 30 minutes.
Free, on video, with César Soto.
Book a conversation