A better AI model can write a better answer. It cannot, by itself, know which customer record is current, when to ask a person for approval, or whether a task actually finished. The harness is the operating setup around the model that makes those decisions explicit and testable.
This is why the next useful question for a business is not only “Which model should we use?” It is “What will the agent be allowed to see, do, remember, and prove?” OpenAI's harness-engineering write-up describes how tools, environment, feedback, and observability made its coding agents more useful. Anthropic's agent-building guide similarly distinguishes simple workflows from agents that choose actions dynamically. Those examples come from software development; the same design questions apply to business workflows, but not every small business needs a coding agent.
Visual: The model reasons; the harness supplies the approved context and tools, checks actions, and records outcomes.
What is an AI harness?
Think of the model as an engine. The harness is the rest of the vehicle: steering, brakes, dashboard, and rules for where it may go. In practice it has six parts:
- A defined job: one outcome, start trigger, success condition, and point at which the agent stops.
- Relevant context: current policies, product information, and the specific customer data needed for the task, with a way to identify stale information.
- Tools and integrations: calendar, CRM, email, ticketing, or document access through narrow, auditable connections.
- Permissions and approvals: what the agent may read, draft, or send; which actions need a person; and what happens when the request is outside scope.
- State and handoffs: what has already happened, what remains open, and where a human can resume without guessing.
- Testing and monitoring: realistic test cases, errors, cost and latency, customer outcomes, and a way to pause or roll back the workflow.
The exact software varies. The point is to make behavior visible and recoverable. Anthropic's long-running-agent research illustrates how an agent can lose track of progress across sessions unless the environment leaves a clear record. A business cannot treat “the model sounded confident” as proof that an appointment was booked or a customer was contacted.
A small-business example: after-hours leads
Say an HVAC company wants to respond to after-hours requests. A model-only demo might read a message and draft a polite reply. A working harness needs more:
- It reads the approved service area and hours, not a random old PDF.
- It checks an authorized calendar or dispatch system before suggesting a slot.
- It flags emergency language and hands the request to the on-call person.
- It drafts a reply and sends it only under the business's chosen approval rule.
- It logs whether the lead received a response and whether a booking actually happened.
That is a narrower, more useful outcome than “deploy an AI employee.” If the calendar connection fails, the system should say it cannot confirm availability rather than invent a time.
When is an agent the wrong choice?
If the steps never vary, a conventional automation may be cheaper and easier to audit. If the task is only to summarize a document, a single model call with review may be enough. Use an agent when the work genuinely requires choosing among tools or paths, and when the benefit justifies the additional controls. Anthropic's agent guidance makes the same distinction: start with the simplest approach that meets the need.
The harness is not a permission slip for broad access. Start with read-only or draft-only capabilities, then add actions after the workflow passes tests. Our existing AI agent safety guide covers the security side in more detail.
Five questions to ask before buying or building
- Which exact task will the agent complete, and what result will we measure?
- Which systems and fields does it need, and who owns those permissions?
- Which actions can it take without approval? What gets escalated?
- How will we detect a wrong answer, failed tool call, or duplicate action?
- Can a person inspect the record, correct the outcome, and stop the agent?
If the vendor cannot answer those questions, a polished demo is not a production plan. Talk to us about an AI-agent workflow if you want to scope one job, its approvals, and a testable first pilot.
Ready to Automate Your Business?
Book a free workflow review and discover which processes you should automate first.