Evaluate a coding agent on 30 real backlog tickets
A practical study design for measuring accepted work, failed attempts and reviewer effort across a real backlog.

An agent can look useful on a ticket you chose because you already knew it would be easy. The harder question is what happens to the rest of the queue.
A 30-ticket evaluation is a manageable way to look beyond the best screenshot. It is large enough to include different kinds of work and small enough that a team can inspect every result. This article describes an evaluation you can run, for instance by delegating the tickets one by one to Hoplite cloud agents. It does not report a completed Hoplite experiment or a measured success rate.
Choose the tickets before you see the results
Take a defined slice of your backlog. You could use the next 30 eligible maintenance tickets, ordered by the team's existing priority. Record the selection rule so you cannot quietly drop a failure from the sample.
Define eligibility first. A ticket needs a repository, an intended behavior and a way to assess the result. Some tasks will fail that check because the product decision is still open. Exclude those, record the reason, and report them with the study. Ambiguity is part of the cost of delegation.
Keep a mix of jobs your team genuinely wants done. If you exclude every task involving dependencies, UI behavior or unfamiliar code, the result says very little about ordinary engineering work.
Freeze the brief and the starting state
Write the acceptance criteria before assigning a ticket. Record its starting commit and any setup required to run the relevant checks. Keep the initial brief with the eventual result.
Hoplite's task delegation guidance is a useful starting point for choosing bounded work. For the study, be more disciplined than in day-to-day work. A follow-up instruction counts as intervention even when typing it takes only a moment.
Use separate threads for independent tickets. Keep dependent changes in order. If ticket B only works after ticket A changes a shared API, document that dependency. Do not count them as two independent parallel attempts.
Decide how outcomes will be classified
Give every selected ticket a final state. Useful categories are accepted on the first submission, accepted after intervention, rejected and unfinished when the study closes.
Keep an additional field for what happened during review. A task may satisfy its acceptance criteria yet remain unmerged because the team changes priorities. A merged PR may later need a correction. A single success flag misses both cases.
Ask reviewers to record time spent reviewing, how many rounds of corrections they requested, and the reason for rejection. Keep setup failures visible. If a run cannot install dependencies, record the time spent repairing the environment and whether a later attempt succeeded.
Measure the work around the agent
Capture agent usage from the thread and workspace usage records. Record human time separately. Include the time spent preparing instructions and checking the final behavior, not just the time looking at the diff.
Agree on an observation window for regressions before starting. A change that appears complete on Friday may reveal a missed case after release. Add that outcome to the original ticket instead of reporting the success and the later failure separately.
Do not subtract agent runtime from an engineer's workday and call the difference time saved. Background tasks can overlap with other work. A fairer comparison uses similar historical tickets or a separate matched sample and still acknowledges differences in difficulty.
Publish the awkward cases
A useful write-up includes the selection rule, every outcome and examples of where the agent needed help. Show a first-pass success and a failure that changed how you briefed the next task.
Report fractions with their denominators. If 18 of 30 tickets were accepted, that is an illustrative calculation of 60 percent, not a Hoplite result. Explain how many needed intervention and how much reviewer time they took before you call the evaluation worthwhile.
Start by exporting the candidate ticket list and recording why each belongs in the sample. Keep that file unchanged while the runs proceed. Your team will trust the results more.