Harness engineering for coding agents

11 min read

Engineers reviewing software architecture on a large monitor

Agents need a harness, not just better prompts

The first few times I used coding agents seriously, I treated the prompt like a long ticket description. I explained the feature, pasted a few constraints, and expected the agent to infer the rest of my engineering process.

That works for demos. It breaks down on real teams.

The problem is not that agents cannot write code. The problem is that most teams give them an undefined job. One prompt asks for architecture, implementation, testing, product judgment, file ownership, and release planning all at once. Then everyone is surprised when the agent edits too much, misses a constraint, or optimizes for a different definition of "done."

Harness engineering is the work around the agent. It is the prompt contract, role boundary, review loop, and handoff format that turns a flexible coding model into a predictable engineering participant.

I think about it the same way I think about CI. The compiler is only one part of the system. The useful part is the harness around it: inputs, checks, logs, failure modes, and a clear pass/fail contract.

Start with the role

Before I write a task prompt, I decide what role the agent is allowed to play.

The role matters because "build this" is too broad. A frontend implementer, backend implementer, test writer, reviewer, and debugging partner should not have the same instructions. They should not own the same files. They should not make the same decisions.

For coding work, these are the roles I reach for most often:

  • Planner: turns a request into implementation steps, risks, and acceptance criteria.
  • Implementer: changes code inside a bounded ownership area.
  • Tester: adds or updates tests based on the intended behavior.
  • Reviewer: looks for defects, regressions, missing tests, and unclear contracts.
  • Debugger: reproduces a failure, isolates root cause, and proposes the smallest fix.
  • Integrator: checks whether independently changed pieces still work together.

Those roles can be handled by people, agents, or both. The important part is that the role is explicit.

Here is the difference in practice:

Weak role: You are a senior engineer. Implement this feature. Useful role: You are the frontend implementer for this task. Your job is to update only the owned UI files, preserve the existing component patterns, add focused tests before production changes, and write a handoff for review. Do not alter backend contracts, package files, or unrelated pages.

The second prompt is less glamorous, but it is much easier to trust.

Write prompt contracts, not essays

A good task prompt is a contract. It tells the agent what job it has, what context it must read, what it may change, what it must not change, and what evidence proves completion.

I try to keep the structure boring:

CONTEXT - Repository, branch, ticket, product goal. - Why the change exists. - Current constraints or known risks. ROLE - The exact engineering role the agent is playing. - What decisions it owns. - What decisions it must escalate. INPUTS - Files, docs, tickets, API contracts, design links, logs. - The order they should be read in. SCOPE - Files or directories the agent may edit. - Files or directories it must not edit. - Any behavior that must remain unchanged. TASK - The concrete change to make. - Acceptance criteria. - Required edge states. WORKFLOW - Test-first or reproduce-first expectations. - Review and verification commands. - Handoff format.

The shape is more important than the exact headings. I want the prompt to remove ambiguity before the agent starts writing code.

The biggest improvement usually comes from adding negative space: "do not touch this," "do not invent this contract," "do not solve adjacent problems." Agents are good at continuing patterns. If the prompt leaves the boundary open, they will often continue past the part I wanted.

Separate task prompts from standing instructions

I do not want every prompt to repeat the team's entire engineering culture. That makes prompts long, inconsistent, and easy to drift.

Instead, I split instructions into two layers:

  • Standing instructions: how the team works in general.
  • Task instructions: what this specific change needs.

Standing instructions cover things like:

  • test expectations
  • review posture
  • accessibility standards
  • file editing rules
  • dependency rules
  • branch and commit conventions
  • security and privacy constraints
  • how to report blockers

Task instructions cover the local problem:

  • the feature or bug
  • owned files
  • acceptance criteria
  • relevant contracts
  • product or design constraints
  • validation commands for this task

This split matters because the task prompt should be small enough for a lead to review carefully. If every prompt is a wall of repeated process text, the important task-specific constraints get buried.

Give agents file ownership

The simplest way to reduce damage is to give every implementation agent a file ownership boundary.

I prefer prompts that say this plainly:

You may edit: - app/settings/profile-form.tsx - app/settings/profile-form.test.tsx You may not edit: - API routes - database schema - shared design system components - package files - existing unrelated pages

This does not replace code review. It just gives the agent a small workspace where its decisions are expected.

Ownership also improves collaboration. If a backend change and a frontend change can happen independently, I want two separate agents or two separate human tasks with clear contracts between them. If the frontend agent needs a field that the backend does not provide, the right output is not to invent the field. The right output is to report the contract gap.

The more shared the file, the tighter the prompt should be. A change to a widely used component needs stronger constraints than a change to a one-off page.

Define what "done" means

Agents will happily stop when code compiles if the prompt lets them. I rarely want that.

For implementation tasks, my done contract usually includes:

  • the requested behavior is implemented
  • existing behavior is preserved
  • relevant tests were added or updated
  • validation commands were run fresh
  • known failures are called out with evidence
  • the handoff lists changed files and review notes

For review tasks, done is different:

  • findings are ordered by severity
  • every finding has a file and line reference
  • speculative issues are labeled as assumptions
  • no code is changed
  • test gaps are called out separately

For debugging tasks, done is different again:

  • the failure is reproduced
  • root cause is isolated
  • the fix is minimal
  • the original reproduction now passes
  • nearby regressions are checked

That is the point of role-based prompts. "Done" changes with the job.

Make escalation explicit

One of the worst prompt mistakes is asking the agent to be autonomous without explaining when autonomy should stop.

I want agents to proceed through routine ambiguity, but I do not want them guessing on product, data, security, or contract decisions. So I include escalation rules.

Examples:

Escalate if the task requires changing the public API response. Escalate if the design cannot fit the existing component system. Escalate if validation requires credentials that are not available. Escalate if the requested behavior conflicts with an existing test.

This is especially useful for technical leads because it makes the agent's uncertainty visible. A good blocker report is not a failure. It is a control point.

Ask for evidence, not confidence

I do not care if an agent says a change is "complete." I care what it ran and what happened.

A useful handoff includes concrete validation:

Validation: - npm test -- profile-form.test.tsx: pass, 8 tests - npm run typecheck: pass - npm run lint: pass - npm run build: skipped, outside task scope Notes: - Did not change API contract. - Empty state and server error state are covered by tests. - Needs reviewer attention on keyboard focus order.

This format is deliberately plain. It gives the reviewer something to verify quickly and makes skipped checks visible.

If a command fails because of an environment issue, the agent should say that. "Build failed because DATABASE_URL is missing" is useful. "Build should pass in CI" is not.

Keep examples small and executable

Prompt examples should be short enough that a busy engineer will actually use them.

Here is a role prompt I would give to a frontend implementation agent:

You are the frontend implementer for this task. Read the ticket, then inspect the existing component and tests before editing. Implement only the requested UI behavior in the owned files. Follow the existing styling and state management patterns. Handle loading, empty, error, and success states where the UI can enter them. Write or update focused tests before production changes. Run the task-specific test, type check, and lint command if available. Do not change API contracts, package files, shared utilities, or unrelated components. If the backend data shape does not support the UI, stop and report the contract gap. Return a handoff with summary, files changed, validation, and review risks.

And here is one for a reviewer:

You are reviewing this change for correctness. Prioritize bugs, regressions, missing tests, accessibility problems, and unclear contracts. Do not rewrite the code. Do not comment on style unless it affects maintainability or user behavior. List findings first, ordered by severity. Every finding must include a file and line reference. If there are no findings, say that and list remaining risks or unverified areas.

These prompts are not magic. They are guardrails. The value comes from using them consistently enough that the team can compare outputs across tasks.

Design the loop around the agent

The harness is not just one prompt. It is the loop.

For most code changes, I like this sequence:

  1. A planner turns the request into scoped work and acceptance criteria.
  2. An implementer changes the owned files and records validation.
  3. A reviewer checks the diff from a defect-first posture.
  4. The implementer addresses accepted findings.
  5. An integrator runs broader validation if multiple parts changed.

Small tasks can collapse some of those steps into one person or one agent session. Large tasks should not. The point is to avoid asking one role to be its own unchecked reviewer.

This also keeps prompts smaller. The implementer does not need to be told how to do final release review. The reviewer does not need to be told how to build the whole feature. Each role gets a job it can actually complete well.

Watch for prompt drift

Prompts age like code. They pick up exceptions, one-off warnings, and old assumptions.

I review agent prompts when I see these symptoms:

  • agents repeatedly edit outside their scope
  • reviewers keep finding the same missing state
  • validation reports are vague
  • agents invent contracts instead of reporting gaps
  • tasks need long manual cleanup after every run
  • prompts contain stale file paths or outdated commands

The fix is usually not "write a longer prompt." It is to sharpen the contract.

For example, this:

Make sure the UI is accessible and well tested.

is weaker than this:

Use semantic controls, preserve keyboard navigation, expose loading and error states to assistive technology, and add tests for the submit success and failure paths.

Specific instructions are easier to follow and easier to review.

What I would standardize first

If a team is starting from scratch, I would standardize only a few things:

  • role templates for planner, implementer, tester, reviewer, debugger, and integrator
  • a task prompt format with context, scope, task, workflow, and output
  • file ownership rules
  • escalation rules
  • validation and handoff format
  • a short list of things agents must never do without approval

That is enough to make agent work less random.

I would not start by building a complicated internal framework. Most teams need better contracts before they need more tooling. Once the contracts are stable, tooling becomes obvious: prompt templates, checklists, validation scripts, and dashboards can grow from real repeated behavior.

The practical test

The test I use is simple: could another engineer review the prompt and predict the shape of the output?

If yes, the harness is probably good enough.

If no, the agent is being asked to infer too much. It may still produce something impressive, but the team will spend the saved time on cleanup, review churn, and coordination.

Coding agents are most useful when they are treated like powerful but bounded collaborators. Give them a role. Give them a contract. Give them a narrow surface to change. Ask for evidence. Then improve the harness the same way you improve any other engineering system: by watching where it fails and tightening the loop.

Share:TwitterLinkedIn