AI Harness Wars
Which harnesses do you use? How do you compare harnesses? What power do they give you? What do they lack? What will harnesses look like a year from now? Join us to share insights with your peers!
Two central themes came out of the discussion: first, what do these terms even mean? And second, who is using the harness?
Context #
“Context” is a very overloaded term. The uncertainty around “What is context?” was a central theme during this talking point.
The conversation largely settled at the level of company-wide context, rather than model context windows or repository-level context. The term “context lake” came up several times as shorthand for this broader scope: the collection of business knowledge that needs to be stored, maintained, and made available across the company. Other levels of context were also discussed, but it was often unclear which level someone meant when they used the term “context,” particularly when contrasted with “context lake.”
The main questions around company-wide context were:
- How do I store it?
- How do I keep it up to date?
- How do I make it available to everyone in the company?
The clearest example discussed was using a Markdown-based wiki, with some feedback that this can be a dangerous approach if done incorrectly. Keeping the wiki up to date is not a trivial task, and controlling when updates take place was a point of discussion.
One advantage of Markdown files is that models are heavily trained on Markdown.
Several other topics came up but did not have strong alignment among the group:
- Who does the context serve?
- What is considered context?
- What specifically do people mean when they talk about context?
There was a lot of discussion around these questions, with no clear resolution on what exactly “context” meant.
Effectiveness #
The conversation here focused on how to measure harness effectiveness.
Several KPIs were proposed:
- Tokens
- Dollars
- Wall time
Some third-party evaluation tools were also mentioned:
- Arena.ai
- Databricks internal evaluation
A key consideration was that the developer time spent operating the harness is also a cost.
Ultimately, there was no clear resolution on how to gauge the effectiveness of a harness.
Token Usage #
The question of how to optimize token usage was discussed.
One of the first comments was:
Why are you trying to optimize token usage? #VCMoneyForTheWin
But this is a real question. For some people, token usage is something that needs to be considered; for others, it is not.
Several observations and strategies came out of the discussion:
- Gateways, such as OpenRouter, can optimize spend without requiring you to worry directly about tokens. Harnesses generally do not perform this routing.
- Cached tokens and new tokens are different. Cached tokens are cheaper, but there is a fair amount of nuance around them:
- Cached tokens expire and become new tokens after enough time away from a session.
- Cached tokens are order- and session-dependent.
- One strategy is to use a chat for roughly three iterations and then hand off to a new chat to control context size.
- Tool output can heavily bloat context.
- Information in the middle of a large context window can lose weight and be overlooked by the model.
- Claude’s recommended best practices were also mentioned:
- Files of around 200 lines are recommended.
- Use agents to decompose large files into smaller ones.
Buy vs. Build #
Should you buy or build a harness?
This was the topic with the strongest alignment across the group. The general opinion expressed was that models work better with the harnesses they are shipped with:
- OpenAI models with Codex
- Claude models with Claude Code
What You Get When You Buy #
- System prompts
- Model-specific optimizations
What You Get When You Build #
- Control flows implemented in the harness rather than encoded into context
How Do You Evaluate Harnesses? #
UI responsiveness and usability were major considerations for some people when evaluating harnesses from companies outside the frontier labs.
For others, lock-in and business continuity mattered greatly.
These concerns did not necessarily contradict the general alignment around harness performance, but they do feed into the trade-offs involved in buy vs. build.
The KPIs discussed earlier did not re-enter the conversation here.
Multimodal Harnesses #
Should harnesses be multimodal? Should voice, video, and other non-language inputs and outputs be built into these harnesses?
This was the final topic discussed. Much of the conversation centered on clarifying what a multimodal harness actually meant.
The discussion largely converged on one point: the frontier reasoning models are primarily language-based. Voice-to-voice models are not currently as advanced in reasoning.
Using a speech-to-text tool to feed a language model, then a text-to-speech tool for output, appears to be a common approach. However, that was not what was meant by a truly multimodal harness in this discussion.
Nolan Aguirre is a principal engineer working on AI devops tooling.