Brief
Harness choice shows no resolved average task advantage in paired agentic coding tests
A preprint reports paired same-model contrasts on a private contamination-controlled suite of 256 repository and post-cutoff contest tasks. It finds no resolved average harness advantage, with cost and completion caveats.
The study ran the same 80 tasks under claude-agent-sdk and deepagents on claude-opus-4-8, and under openai-codex SDK and deepagents on gpt-5.5. Neither contrast resolved an average advantage: -1.25 pp for Opus 4.8 and +1.25 pp for GPT-5.5. The Opus average combines opposite strata: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks, a partition chosen after seeing the data and needing a designed replication. Also, 22 of 81 runs cancelled at the wall-clock ceiling had produced a passing patch. Re-priced from raw per-turn usage at frozen list prices, the neutral harness cost 1.3 to 1.6 times as much per solved task on Opus 4.8 and 1.2 times on GPT-5.5, though billed ordering is unresolved because 58 Anthropic runs left no usage record.
Our reading
Our reading is that teams should not assume vendor-native harnesses improve task success and should measure cost per solved task and completion behavior per task type.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- The same 80 tasks ran under claude-agent-sdk and under deepagents on claude-opus-4-8, and under the openai-codex SDK and deepagents on gpt-5.5.
- Neither contrast resolves an average advantage for either harness: -1.25 pp for Opus 4.8 (48.8% vs 50.0%) and +1.25 pp for GPT-5.5 (55.6% vs 54.4%).
- The Opus average combines opposite strata: the native harness trails by 9.0 pp on the 61 repository tasks and leads by 23.7 pp on the 19 contest tasks (label-permutation p = 0.003).
- 22 of 81 runs cancelled at the wall-clock ceiling had produced a passing patch.
- Re-priced from raw per-turn usage at frozen list prices, the neutral harness cost 1.3 to 1.6 times as much per solved task on Opus 4.8 and 1.2 times on GPT-5.5.
Sources
- arXivText stored 14 September 2026
How this story was checked. Written from the 1 page listed above, stored 14 September 2026; claims checked against that stored text on 14 September 2026.
What that means
- 5 of 5 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.