Standard HarnessAI systems / Singapore
Research in progress

Published evidence / Ongoing synthesis

Better harnesses. Measured in real work.

The harness around a coding agent shapes what it sees, how it acts, and how the result is checked. Codegraff has published evaluations of those choices. This page brings the measured outcomes together and keeps their conditions in view.

Recorded results

Completion and cost, by task suite.

In the published Graff FrontierHarness report, Graff with Grok 4.6 completed 20 of 21 selected terminal tasks; Graff with Kimi K3 completed 17. Both recorded one pass in the separate nine-task DeepSWE patch suite. These are repository-reported runs, with the task and cost details available in the evaluation explorer.

Selected evaluation slices

What these runs completed

Each mark counts one task: filled for a pass, outlined for a nonpass. Mark positions do not identify the same task across runs.

Terminal · 21 tasks

Most terminal tasks passed

Completed tasks in the recorded terminal slice.

Graff with Grok 4.6

20 / 2195.2% passed

Recorded total spend$6.2335

Spend per pass~$0.31

Recorded summary

Graff with Kimi K3

17 / 2181.0% passed

Recorded total spend$3.6463

Spend per pass~$0.21

Recorded summary

DeepSWE · 9 tasks

One patch passed per run

Patch outcomes in the published report; evidence status is shown per run.

Graff with Grok 4.6

1 / 911.1% passed

Recorded total spend$25.5616

Spend per pass~$25.56

Verifier artifact

Graff with Kimi K3

Reported result

1 / 911.1% passed

Recorded total spend$21.9148

Spend per pass~$21.91

Reported result

No Kimi task-level verifier file is included in the pinned source.

These are selected stored runs, not a new rerun or a full leaderboard. The terminal runs used additional evaluation instructions and local Docker setup. Spend uses historical estimated prices and includes failed attempts; spend per pass divides the recorded total by passes. Read the results and suite definitions and evaluation protocol.

Evaluation path

What the verifier sees.

01 / TaskIssue and starting state
02 / ConfigurationModel, harness, instructions, runtime
03 / ExecutionTools operate on the task
Grade by task type
Terminal-Bench sliceFinal environment → public tests
DeepSWE sliceSubmitted patch → apply + project verifier
Verified outcome recorded spend
Schematic of the two distinct grading paths in the published comparison; see the pinned evaluation protocol for execution and scoring details.

Method and limits

Keep the conditions attached.

The selected terminal tasks grade a final environment with public tests. DeepSWE checks whether a submitted patch applies and passes the project verifier. A missing pass remains a miss in the published results. The Graff terminal runs used additional evaluation instructions and a local Docker runtime; the wider published field used other conditions. The recorded differences therefore describe full configurations and do not isolate a harness-only effect.

Costs include failed attempts and use the evaluator’s historical list-price estimates. The pinned run summary and protocol make the setup inspectable. Grok’s DeepSWE pass also has a committed verifier record. Kimi’s one pass is reported in the published summary, but its task-level verifier file is not present at that pinned revision. We have not rerun the suites for this page.

Published work

Read the source reports.

For the practical methods behind this work, read our notes on evaluating agents, repository context, and routing.