All stories
3 min read

AI code review needs a better scorecard

GitHub's new ReviewBench puts AI reviewers under the microscope. The useful question is how much good feedback survives the noise.

AIDevelopmentEngineering
A cream and blue robot selects orange review cards beside a retro computer, with lavender comment bubbles behind it.

Picture a pull request with twelve automated comments. Ten are easy to dismiss. One needs a small tidy-up. One catches a bug that would have reached production.

Was that a good review? The answer depends partly on how much work it took to find the useful comment. Counting comments alone cannot tell us.

That question has a timely new reference point. On 5 October, GitHub introduced ReviewBench, an open benchmark for AI code review. Its research preview contains 219 pull requests from 187 public repositories across 19 languages.

GitHub used the distribution of 103.9 million pull requests to help shape that sample, with a deliberate adjustment towards more substantial changes. The benchmark itself runs on 219 PRs. Keeping those two numbers separate is a useful start when reading the announcement.

Useful comments and missed problems

Review quality has at least two sides. How much of the feedback is useful? How much of the important work does the reviewer miss?

Imagine a change with ten known bugs. A reviewer identifies five correctly and adds five false alarms. In that simplified example, half its comments are correct, and it finds half the known bugs. Those are the basic ideas behind precision and recall.

Choosing between reviewers means deciding which failures matter most. A quiet tool that catches a serious issue occasionally might fit a busy team's workflow. A security review may justify more investigation if it catches problems that would otherwise escape. Neither preference removes the need to check the findings.

ReviewBench makes those choices visible through severity and category filters. It also separates findings that match its existing reference set from additional findings assessed by a model. The published methodology contains an important wrinkle: grounded precision excludes unmatched comments. Augmented precision includes them after classification. Those scores answer different questions.

Someone still has to judge the judge

Building the reference set is difficult too. A reviewer might discover a genuine issue that the benchmark's creators missed. Automatically treating every unfamiliar finding as wrong would punish that discovery.

ReviewBench uses a model to assess unmatched findings, but that adds another source of uncertainty. Its methodology also warns against ranking tools solely by augmented recall: each tool's newly accepted findings change its denominator, so the comparison no longer uses a common yardstick.

The human audit is encouraging, with reported agreement of 96.6% on true-or-false-positive labels. It also corrected 47 findings the initial classifier had wrongly accepted. Both details matter. Human checking improved the dataset; a high agreement rate does not make the remaining judgments infallible.

There are other ways to examine the problem. Martian's Code Review Bench combines a fixed dataset with an online benchmark that compares bot suggestions with developers' subsequent fixes. That brings real responses into the picture, although a code change is still a proxy for usefulness. An accepted suggestion can be mistaken; a useful suggestion can be deferred.

Try the question on your own code

A sensible team trial could start with a small, varied set of changes: an ordinary feature, a bug fix and something that crosses an awkward boundary between components.

Before comparing tools, agree what a worthwhile comment looks like. It should identify a concrete problem, explain the consequence and give the author enough context to investigate. Keep a separate record of important issues found later by people or tests.

Then track the work around the comments. How long did verification take? Which suggestions led to worthwhile changes? What did the reviewer miss? Include waiting time and cost, and keep some examples aside rather than tuning repeatedly against everything in the trial.

GitHub's own responsible-use guidance recommends verifying Copilot's feedback and retaining careful human review. A benchmark result does not remove that responsibility.

A useful scorecard should help a team spend its attention well. The interesting outcome is a problem understood and fixed, with enough confidence to ship the change. The comment count can stay in the background.

Cover: original AI-generated editorial illustration.

Adam Spice

Developer. Curious human.

More stories