In everyday words
If one AI is asked to decide which of two code answers is correct, it may sound confident even when it has no real proof. The paper shows a multi-step “check each claim” approach can still fail for code, because it may not get different evidence for each option. A practical fix is for the judge to sometimes say, “I can’t tell from the evidence,” based on signals...
What you need to know
- Who is affected
- Software teams using AI to compare or review code, Managers relying on AI-generated code review decisions, Teams building workflows where AI checks other AI outputs
- What changed
- Researchers studied when an AI system that judges code is actually supported by evidence. They argue the evidence must be independent from the code being judged and must differ between the two code options. In code judging, that second condition can fail. Testing a published multi-step judge (MARCH) on two code-judging comparisons, they found it often rated both options equally good.
- Why it matters
- At work, teams may use AI to review or compare code changes. This paper suggests some multi-step “verification” methods can still produce confident-sounding judgments without a real basis. The authors’ main contribution is a way to detect, without extra labels, when the judge lacks support and should refuse to decide. That can reduce misplaced confidence in automated code review.
- What to watch next
- Whether teams adopting AI-based code review add a “refuse to decide without evidence” step, especially for head-to-head comparisons.
Four useful details
- In tests, the multi-step judge often said both code options were equally good.
- A log-based check let the judge skip comparisons it couldn’t support.
- The paper’s focus is detecting “no basis,” not claiming a better judge overall.