If one AI is asked to decide which of two code answers is correct, it may sound confident even when it has no real proof. The paper shows a multi-step “check each claim” approach can still fail for code, because it may not get different evidence for each option. A practical fix is for the judge to sometimes say, “I can’t tell from the evidence,” based on signals...
ForSoftware teams using AI to compare or review code · Managers relying on AI-generated code review decisions