Field notes / methods
Research for AI that gets work done.
Methods and evidence behind better AI systems: checking results, finding the right information, and choosing a practical route for each task.
Research in progress
Explore the evidence Published evidence / Featured research
Better harnesses. Measured in real work.
Explore published Graff task outcomes, costs, and runtime experiments, with the evaluation conditions attached to every result.
Published reports
CodegraffMethods notes
01—0301
Evaluation ·
02Evaluating coding agents by verified work
A practical method for comparing coding agents across completion, cost, latency, and reproducible execution conditions.
Context ·
03Repository context as a navigation problem
How to give a coding agent the structural evidence it needs without flooding its context window.
Routing ·
Routing coding agents by task and budget
A framework for choosing an agent configuration using task type, verified completion, runtime, and total spend.
