Measure coding-agent results beyond tokens and lines of code
Track accepted changes, review effort, cycle time, and regressions without confusing activity with impact.

Diffs got cheap, which makes diff size a useless measure of progress. An agent can produce a large one while moving a ticket no closer to completion. It can also delete a few lines and fix a bug that has annoyed customers for months.
Tokens and lines changed describe activity. To understand whether the workflow helps your team, follow the work through acceptance and the time people spend around it.
Start with one question
Choose the decision the measurement should support. Perhaps you want to know whether agents reduce the effort spent on dependency updates. Perhaps the question is whether a background agent clears routine bugs without increasing review time.
Keep the scope narrow enough to compare similar work. Combining framework migrations with copy fixes can make almost any average look plausible. Tag task type and rough scope when the work enters the sample, before the outcome influences your judgment.
Write an explicit definition of accepted work. A merged PR is a useful event, but it may not be the end of the story. Decide how you will record rollbacks and follow-up fixes during the period you care about.
Record every attempt
Use one row per task, with links to all associated attempts and PRs. Record when the task became ready, when an agent started, when a reviewer first inspected it, and when the team accepted or rejected the result.
Those timestamps answer different questions. Time waiting for review may dominate elapsed time even when agent execution is quick. A long-running background task may consume little human attention. Don't turn both into the same productivity number.
Keep briefing minutes and active review minutes as separate fields. Record correction rounds and the reason a task was abandoned. A failed attempt can reveal an environment problem that would otherwise recur for the entire team.
Collection is usually the part that fails. When runs happen on individual laptops, the record depends on who remembered to write an attempt down. If you don't want to keep that record by hand, a Hoplite thread keeps the prompt, commands, approvals, and PR for every run, including the failed ones. Your sheet then holds the judgments and a link.
The measurement definitions still belong to your team. An activity log does not decide whether a task was useful.
Choose the right denominator
Report accepted tasks divided by attempted tasks, alongside the number of eligible tasks that never started. Report first-submission acceptance separately from acceptance after correction.
For review effort, show a typical value and inspect the expensive cases. An average can hide one change that consumed a day of senior-engineer time. Keep task types visible so readers can tell whether the result improved because the workflow improved or because the work became easier.
Track defects against the changes that introduced them. Define the follow-up window in advance and avoid treating a lack of observed incidents as proof that every accepted change was correct.
Compare against a baseline
If your question is time saved, collect a comparable baseline. Historical tasks can help, though they may differ in difficulty or in the people doing them. A matched set from the same team is more informative than comparing agent runtime with a rough guess about manual work.
Document changes in setup during the evaluation. Better fixtures, clearer instructions, and repaired test scripts can improve both human and agent work. That is a useful result, but attributing all of it to a model would hide what the team should keep investing in.
Combine the outcome record with per-task usage attribution when you need a cost view. Keep external subscriptions and human time in the calculation. Workspace token charges alone cannot answer whether the adopted workflow pays for itself.
Bring one rejected task and one accepted task to the next retrospective. Read the instructions, the interventions, and the final diff together. Choose one change to the process that the evidence supports, then see whether it helps the next batch.