Skip to main content

Agentic software development is 90% harness. How good is yours?

Your developers already use AI. Is that happening systematically, or by accident? Now that unspeakable nonsense like tokenmaxxing seems to be behind us, we show you how to work productively with the limitations of LLMs in software development, even in larger teams. To that end, we bring an SDLC harness into your existing development world, adapt it for your company, and stay through the change until your team carries it on its own.


At first, the rules of good engineering have nothing to do with AI

AI tools have arrived, but discipline hasn't kept up. One developer uses AI for every line, the next wouldn't even know how, a third doesn't trust LLMs at all, and none of the three ways of working gets measured. Quality varies, and nobody can quite say why. And productivity turns into more free time for individuals, which we begrudge nobody, but that too could be organised better. Data protection and compliance... well, you know the answer.

The usual blanket measures, things like "everyone must use AI", "everyone presents something AI made", "we're measuring token usage!", are a desperate attempt to control a situation that remains, at its core, misunderstood. We ensure three things for you:

  1. 1. A strategy for using the right tools for your situation
  2. 2. Putting that strategy into operation, including best practices and change management
  3. 3. Building an agentic harness specific to your use cases, tooling, and other requirements

The model accounts for about 10 percent of it

Most agent failures, examined honestly, are configuration failures, not model failures. The model's share of agent behaviour runs at roughly 10 percent. The remaining 90 percent is decided by harness configuration: what instructions the agent gets, which tools it can reach, what environment it runs in, how it's orchestrated, controlled, and observed. That makes the harness the actual engineering task, not the choice of model.

On proprietary enterprise platforms, that ratio tips even further away from the model (more on that on the blog).

Model 10% Harness 90%

A harness has six components:

  • Instructions

    what the agent knows and how it's directed

  • Tools

    what it can access, and what it can't

  • Sandboxes

    the environment it executes in

  • Orchestration

    how multiple steps or agents work together

  • Guardrails

    which actions get blocked or flagged automatically

  • Observability

    what's measurable and traceable after the agent has acted

Where a dev organisation sits on the spectrum from vibe coding to agentic engineering shows up in these six components.


Copying has its limits. Your delivery capability depends on how well you maintain your harness.

No. You can't buy a finished harness off the shelf. LinkedIn, TikTok, YouTube will tell you the five things you need. That's always exactly as wrong as it's right. The technology's strength lies in the inherent necessity of applying it specifically. Without your context, without your problems, without your people's ways of working, copied-in approaches produce results far short of what's possible.

That's why we build the harness in your environment, docked to your existing CI/CD pipeline. If you don't have one, we know where to start too. Once the change has taken hold, the harness is yours: configuration, documentation, operation.

For every phase of the software development lifecycle - from ideation, through requirements detailing, solution engineering, implementation, to quality assurance and documentation - we bring best practices, meaning starting points, which then get adapted for you.


Start small, stay cancellable

  1. 1

    Assess (2 weeks). Where does your dev organisation sit on the spectrum from vibe coding to agentic engineering? We look at your existing pipeline, your current tool use, and the total cost of ownership math behind it: a read you can argue with internally, not an off-the-shelf audit.

  2. 2

    Pilot (4 weeks). One harness, one real project, not a sandbox. We configure the six components for a concrete use case in your pipeline and measure the effect before anyone decides on a rollout.

  3. 3

    Scale. Only once the pilot holds does the rollout reach the wider team: with support until your developers run the harness themselves, without us.

No boiling the ocean. The pilot is deliberately small, and cancellable at any point.


One conversation. No pitch.

If section 2 sounded familiar, a conversation about the assess step is worth having, no obligation, no sales pressure.

Looking for an AI solution for your business processes rather than your dev organisation? That's /ai.

Would rather help build a harness like this than read a page about it: About.