AI Harness Wars

Which harnesses do you use? How do you compare harnesses? What power do they give you? What do they lack? What will harnesses look like a year from now? Join us to share insights with your peers!

Two central themes came out of the discussion: first, what do these terms even mean? And second, who is using the harness?

Context #

“Context” is a very overloaded term. The uncertainty around “What is context?” was a central theme during this talking point.

The conversation largely settled at the level of company-wide context, rather than model context windows or repository-level context. The term “context lake” came up several times as shorthand for this broader scope: the collection of business knowledge that needs to be stored, maintained, and made available across the company. Other levels of context were also discussed, but it was often unclear which level someone meant when they used the term “context,” particularly when contrasted with “context lake.”

The main questions around company-wide context were:

The clearest example discussed was using a Markdown-based wiki, with some feedback that this can be a dangerous approach if done incorrectly. Keeping the wiki up to date is not a trivial task, and controlling when updates take place was a point of discussion.

One advantage of Markdown files is that models are heavily trained on Markdown.

Several other topics came up but did not have strong alignment among the group:

There was a lot of discussion around these questions, with no clear resolution on what exactly “context” meant.

Effectiveness #

The conversation here focused on how to measure harness effectiveness.

Several KPIs were proposed:

Some third-party evaluation tools were also mentioned:

A key consideration was that the developer time spent operating the harness is also a cost.

Ultimately, there was no clear resolution on how to gauge the effectiveness of a harness.

Token Usage #

The question of how to optimize token usage was discussed.

One of the first comments was:

Why are you trying to optimize token usage? #VCMoneyForTheWin

But this is a real question. For some people, token usage is something that needs to be considered; for others, it is not.

Several observations and strategies came out of the discussion:

Buy vs. Build #

Should you buy or build a harness?

This was the topic with the strongest alignment across the group. The general opinion expressed was that models work better with the harnesses they are shipped with:

What You Get When You Buy #

What You Get When You Build #

How Do You Evaluate Harnesses? #

UI responsiveness and usability were major considerations for some people when evaluating harnesses from companies outside the frontier labs.

For others, lock-in and business continuity mattered greatly.

These concerns did not necessarily contradict the general alignment around harness performance, but they do feed into the trade-offs involved in buy vs. build.

The KPIs discussed earlier did not re-enter the conversation here.

Multimodal Harnesses #

Should harnesses be multimodal? Should voice, video, and other non-language inputs and outputs be built into these harnesses?

This was the final topic discussed. Much of the conversation centered on clarifying what a multimodal harness actually meant.

The discussion largely converged on one point: the frontier reasoning models are primarily language-based. Voice-to-voice models are not currently as advanced in reasoning.

Using a speech-to-text tool to feed a language model, then a text-to-speech tool for output, appears to be a common approach. However, that was not what was meant by a truly multimodal harness in this discussion.


Nolan Aguirre is a principal engineer working on AI devops tooling.


Image
Image
Image
Image
Image
Image
Image
Image