They measure consistency well. Every candidate gets the same questions, in the same order, evaluated against the same criteria, and the fiftieth candidate is assessed with the same attention as the first. Human screeners cannot do that, and their bar drifts across a long day.
They measure clarity of spoken answers reasonably well, which is genuinely relevant for many roles.
They measure poorly anything requiring follow-up. A human interviewer notices hesitation and digs into it; an AI interviewer moves to the next question. They also disadvantage candidates with poor connectivity, non-standard accents, or anxiety about being recorded, none of which correlates with job performance.
So a well-designed process treats an AI interview as a first filter, not a final decision.