Blog
The Harness Is Everything
A new study found an 88.5% performance lift across 18 AI models with the model itself unchanged. The gains came from the harness, not the model.
There’s a research paper that should change how MSPs and internal IT teams evaluate AI agents. An 88.5% average performance lift across 18 AI models, with the models frozen: improvements came from the harness, not the model.
- New research found an 88.5% average performance lift across 18 AI models without changing the model itself, just the environment around it, what researchers call the “harness.”
- The harness is everything surrounding the model: what it’s allowed to do, how information gets presented to it, and how mistakes get caught before they execute.
- IT support is a harness-heavy problem. Agents need live context, a curated toolset, and validation checks to act safely on a customer’s machine, none of which a smarter model solves on its own.
- SuperIT’s harness runs on four layers: live context from the Digital Twin, a curated set of tools with defined boundaries, a three-layer safety pipeline for validation, and context management across long-running, multi-day tickets.
- The right vendor questions now are: how does the system get context before it acts, what stops it doing something it shouldn’t, and how does it improve over time. If the honest answer to any of those is “we trust the model,” that vendor hasn’t built a harness yet.
Here’s a research finding that didn’t get the attention it should have.
A paper called Life-Harness ran the same task across 126 different setups, using 18 different AI models, and measured how much the performance improved when the researchers changed only the environment the model was operating in, not the model, the prompts, or the fine-tuning: the tools the model had access to, the way information was presented to it, the way its output was checked and fed back.
The mean improvement was 88.5%. 116 out of the 126 setups got better. The models were frozen. None of them changed.
That should stop you for a second.
For two years the AI industry has been competing on model quality. The headlines have all been about which model is biggest, which is fastest, which scores highest on benchmarks. Vendors selling AI tools talk about the model they use as if the model is the product. But this research, and a steadily growing body of similar findings, is making it harder and harder to argue that the model is the thing that matters most.
The thing that matters most is what surrounds the model. The research community has a word for it: the harness. And it’s the single biggest difference between an AI agent that resolves real tickets in production and one that demos beautifully and breaks the moment a real customer environment touches it.
This piece explains what a harness is, why two AI agents using identical models can have completely different real-world behaviour, and what it means for anyone making a buying decision about AI in IT support over the next twelve months.
What “harness” actually means
The harness is everything the model isn’t.
When you imagine an AI agent solving a problem, you tend to picture the model in the centre and a thin layer of code around it that calls the model and reads the answer. That mental model is wrong, in a way that turns out to matter enormously. The actual structure of a production agent looks more like this:
The model is in the middle. Around it, in successive layers, sits the harness: the tools the model is allowed to call (read a file, query a database, run a shell command), the format in which information is presented to it (one summary line vs ten thousand grep matches), the way its history is compressed as the conversation grows long, the guardrails that catch obvious mistakes before they execute, the way work is broken up into manageable chunks, the way state is preserved across the agent’s many decisions, and the validation that happens between the agent saying “I’m done” and any human believing it.
Each of those layers is engineering. Each of them affects what the model can and can’t do well. Most of them are invisible to the user. All of them are where the actual difference between a good agent and a bad one lives.
The Princeton NLP group made this argument formally in a 2024 paper that became one of the foundational documents of AI agent engineering. They built two agents that used the same model (GPT-4 Turbo), gave them the same task (fix bugs in real GitHub repositories from the SWE-bench Lite dataset), and changed only one thing between the two: the interface the agent used to interact with the codebase.
The first agent used a standard Linux shell. The second agent used a purpose-built set of tools: file viewers with line numbers, search commands that capped output, file editors with built-in syntax checking. The model was the same. The compute budget was the same. The prompts were comparable.
The first agent resolved 11% of the issues. The second agent resolved 18%.
That’s a 64% relative improvement, and it came entirely from changing what surrounded the model. The model itself did exactly the same thing in both setups. The harness was the difference between a tool that mostly fails and a tool that mostly works.
Why the harness compounds in IT support
There’s a reason the Princeton group used software engineering tasks as their benchmark. Coding agents have been the canary in the coal mine for harness design, because writing code surfaces every weakness in an AI agent’s environment quickly and visibly. If the agent can’t navigate a codebase, it produces broken patches. If its output isn’t checked, it confidently breaks the build. If its context fills up with noise, it loses track of what it was doing.
IT support is exactly the same shape of problem, with higher stakes.
A service desk agent needs to navigate a customer environment it didn’t write. It has to figure out which user is reporting the issue, which device they’re on, what software is installed, what state the relevant systems are in. It has to query that information without flooding its own context with noise. It has to take actions, reset a password, restart a service, push a configuration, without breaking anything. It has to verify that the action worked. It has to hand off cleanly when it can’t proceed alone. And it has to do this in minutes, in environments where a bad action has real cost: a locked-out user, a broken production system, a customer who lost confidence.
Every one of those is a harness problem. None of them are solved by picking a smarter model.
The Life-Harness research found that the gains from harness improvements were broadly consistent across model quality. A better harness around a small model produced better results than a worse harness around a large model in many setups. This is the finding that should change how MSPs and internal IT teams evaluate AI tools.
The four harness layers that matter most for service desk work
Speaking from how we’ve built SuperIT, and what we observe matters in production, there are four harness layers that account for most of the difference between an agent that resolves tickets and one that pretends to.
The context layer. This is what the agent knows about the environment it’s operating in when it starts work, and it comes from two places. The Digital Twin supplies the live state: who this user is, what devices they have, what software is installed, what state the relevant systems are in. The knowledge base supplies the accumulated procedural knowledge: how this specific desk has handled this kind of issue before, what’s tribal knowledge here that isn’t tribal knowledge anywhere else. That knowledge base isn’t static either; it’s continuously curated from the desk’s own ticket history, kept current rather than assembled once and left to go stale. A generic AI assistant has neither of these and has to ask. The agent doesn’t ask for information the harness can already supply, because asking wastes turns and burns context window space. The right context arrives unprompted, and it’s specific to this desk, not generic IT knowledge.
The tool layer. This is what the agent is allowed to do, and how it’s allowed to do it, and it’s where owning the endpoint agent matters most. We don’t give the model a raw shell prompt on a customer’s machine or hand action off to someone else’s API. Our own endpoint agent is the controlled channel to the device, running our workflow feature and able to craft the exact PowerShell command for the specific situation in front of it, rather than being limited to a fixed script. Every one of those commands still passes through the same verification and command-policy pipeline before it reaches a device: a “reset a Microsoft 365 password” action is not a “run any cmdlet you like” action. Both could in theory solve the same problem. The first one has a knowable failure surface. The second one will eventually break something nobody asked it to break. Owning the endpoint agent is what makes that distinction enforceable rather than aspirational.
The validation layer. This is what catches mistakes before they cascade. The single most cited finding from the Princeton work was that running a syntax check immediately after every code edit eliminated a class of failure that would otherwise compound across the rest of the session. The equivalent in IT support is that every action the agent proposes goes through a three-layer safety pipeline before it executes. The first layer parses the command structurally and rejects high-risk constructs. The second layer applies policy (allow, needs approval, deny, with the default being needs approval, so the unknown is always escalated). The third layer governs execution itself, including running every command as the signed-in user rather than as SYSTEM, so the agent operates inside the same permission boundary the user would. None of these layers are AI. All of them are checks the harness performs around the model.
The context-management layer. This is what keeps the agent coherent over long sessions. An agent that handles a ticket over a five-minute conversation is easy. An agent that handles a ticket where the user drops out for an hour, comes back, asks something tangentially related, and then circles back to the original issue, is hard. The harness has to maintain a coherent thread across all of that: compressing old context that’s no longer relevant, surfacing context that’s becoming relevant again, and never losing the durable state that matters (who this user is, what they’re trying to achieve, what’s been tried, what worked, what didn’t). This is one of the most-engineered and least-talked-about parts of any serious agent system.
Each of these layers can be improved without changing the model. Each of them, when missing or weak, will sink an otherwise capable agent. And each of them is the kind of work that doesn’t show up in a benchmark or a demo. It shows up in whether the agent still performs on the thousandth ticket the way it did on the first.
The implication for buyers
When you’re evaluating an AI tool for your service desk, the conversation has historically focused on what the AI can do: feature lists, demo videos, model claims. The Life-Harness research, the Princeton ACI work, and the broader shift in how serious AI engineering teams talk about their own systems all point to a different evaluation framework.
The questions that matter, in our view:
How does the system get context before it acts? If the answer is “it asks the user,” that’s an agent without a context layer. It will do fine in a demo. It will be exhausting in production.
What’s the agent allowed to do, and what stops it from doing things it shouldn’t? If the answer is “we trust the model not to do bad things,” that’s an agent without a validation layer. It will be fine most of the time. The failures, when they come, will be the kind nobody wants to debug at 2am.
How does the agent stay coherent across a multi-hour or multi-day conversation? If the answer is “we keep the whole conversation in one prompt,” that’s an agent without a context-management layer. It will work brilliantly for short tickets and lose the plot on anything that requires waiting for a user response.
How does the system improve over time? If the answer is “we’ll release a new version that’s smarter,” that’s an agent whose improvement plan depends entirely on the model vendor’s roadmap. Compare that to a system where the harness itself learns from every resolved ticket, accumulating skills, sharpening tools, and tightening validation. That’s improvement that compounds in your own environment, not in someone else’s release cycle.
None of these questions are about the model. All of them are about the engineering around the model. And all of them are the kind of question that, in our experience, separates the vendors who are genuinely engineering AI systems from the vendors who are reselling someone else’s API with a logo on it.
What we’re really building
Most of what we’ve spent the last few years engineering at SuperIT is the harness. The model behind SuperIT matters, we evaluate it constantly, and we’ll move when the frontier moves, but it’s not where the work is. The work is in the four layers described above, plus the curation cycle that keeps them improving over time.
None of our resolution capability comes from the model alone. The model is necessary; without it none of this works. But the model is also commoditising fast, and the gap between the best and the merely-good is narrowing every quarter. What does not commoditise is the harness: the engineered environment that turns a capable model into a system you can trust on a customer’s machine.
The Life-Harness paper used the phrase “the harness is everything” as a slight overstatement, and we’d agree that’s directionally right but worth tempering. The harness isn’t everything: the model still has to be capable, the knowledge base still has to be curated, the Digital Twin has to stay maintained, the safety pipeline auditable, and the handoff to humans clean.
But the model is the easy part. Everyone has access to the same models. The harness is where the actual engineering happens, and it’s where the real differences between AI products in our industry will show up over the next twelve months.
The vendors who have built one will be fine. The vendors who haven’t will keep talking about which model they use, and hope no one notices the rest.
Want to see the harness at work on your own tickets? Book a demo and we’ll show you.
Common questions
What is an AI "harness"?
The harness is everything surrounding the model: the tools it's allowed to call, how information is presented to it, how its output is checked, and how it stays coherent across a long conversation. Two agents can use the identical model and behave completely differently depending on the harness around it.
What questions should I ask an AI vendor about their harness?
How does the system get context before it acts? What's the agent allowed to do, and what stops it doing things it shouldn't? How does it stay coherent across a multi-hour or multi-day conversation? How does it improve over time? If the honest answer to any of those is "we trust the model," the vendor hasn't built a harness yet.
About the author
The SuperIT Team. Ex-MSP operators and engineers, writing about what we're building and what we're seeing across MSP and internal IT service desks.
The ideas here are ours; we use AI to help draft, edit and publish these posts.
Related reading
Blog
The Context Bottleneck: Why the AI Conversation Just Shifted, and What It Means for IT Support
The smartest models are already smart enough. What's missing is company-specific context, and IT support has the worst version of that problem.
Blog
The Loop Behind SuperIT: How an AI Agent Actually Resolves an IT Ticket
SuperIT isn't a flowchart or a scripted automation. It runs a loop that observes, decides, acts and checks until a ticket is actually resolved.
See how SuperIT works in your environment
The best evaluation is your own queue. Book a demo and we will walk through the ticket types that matter to you.