Case Study

How an AI Interviewer Evaluates Candidates: Rubrics, Scoring, and the Human in the Loop

Grady GardnerNovember 5, 202510 min read
How an AI Interviewer Evaluates Candidates: Rubrics, Scoring, and the Human in the Loop

An AI interviewer evaluates candidates by scoring each answer against a rubric written before the interview, then handing a recruiter the score, the reasoning behind it, and the recording. It does not form an overall impression, and it does not decide who gets hired. This page is about that evaluation step, which is the part buyers ask about last and get burned by first.

For the category definition and how the tooling compares with adjacent systems, start with what AI interview software is. What follows assumes you have that and want the scoring mechanism, the failure modes, and the evidence on whether any of it works.

The short answer

  • The mechanism. Capture the answer, transcribe it, score each response against rubric criteria, attach the text that produced each score, then hand the scorecard to a recruiter.
  • The rubric is the product. A score with no written anchors and no traceable rationale cannot be reviewed, only accepted.
  • Does it work. In a randomized field experiment across roughly 70,000 applicants, candidates interviewed by AI voice were 12 percent more likely to receive an offer.
  • The catch. Public opinion is not with it. Pew found 71 percent of US adults oppose AI making final hiring decisions, which is exactly why the human review step is not optional.

How an AI interviewer works, stage by stage

Scroll to see all columns

StageWhat happensWho controls it
1. Role configurationThe competencies, question set, and scoring rubric are defined for the roleThe recruiter or hiring manager, before any candidate sees it
2. InvitationThe candidate receives a link, usually within minutes of applying, and starts whenever they chooseAutomated from the applicant tracking system
3. ConversationThe agent asks, listens, and generates a contextual follow-up based on the answer givenThe model, inside the boundaries of the configured question set
4. TranscriptionSpeech is converted to text by an automatic speech recognition modelThe model, and this is where accent-related error enters
5. ScoringEach response is mapped to rubric criteria and scored, with the supporting text attachedA separate scoring pass, ideally isolated from the conversational model
6. ScorecardA ranked, filterable report with scores, rationale, transcript, and recordingDelivered to the recruiter
7. DecisionA person reads the evidence and decides who advancesThe employer, always

Scroll to see all columns

Stage three is what separates a real AI interviewer from a one-way video tool. A one-way tool asks a fixed question, records an answer, and moves on. An adaptive agent hears a vague answer and asks the question a good human interviewer would ask next. That is covered in more depth in our explainer on adaptive interviewing.

How an AI interviewer evaluates candidates

An AI interviewer evaluates candidates by scoring each answer against pre-defined rubric criteria, not by forming an overall impression. This is the part most explanations skip, and it is the part that determines whether the output is defensible.

A rubric criterion, written out

Take a customer-facing role and the competency "de-escalation under pressure." A usable rubric does not say "rate 1 to 5." It defines what each level looks like:

Scroll to see all columns

LevelWhat the answer contains
1No specific incident. Describes intention rather than action.
2A real incident, but the candidate's own actions are vague or passive.
3A specific incident with named actions taken by the candidate.
4Specific actions plus the reasoning behind them, and what the candidate weighed at the time.
5All of the above, plus the outcome and what the candidate changed afterwards.

Scroll to see all columns

Scores against anchors like these are reviewable. A recruiter can open the transcript, see the sentence that produced a 4, and disagree with it. A single composite number with no anchors and no trace cannot be reviewed, only accepted.

What to demand from any scoring engine

  • Per-criterion rationale. Every score links to the specific response text behind it.
  • Separation of concerns. The model that conducts the conversation should not be the model that grades it, and scoring should return a fixed schema rather than free prose.
  • Explicit non-assessment. The report should state what was not evaluated, so nobody infers a judgment the system never made.
  • A human review trigger. Defined rules for when a case must go to a person, such as a technical failure, a very short interview, or a candidate-flagged issue.
  • No auto-rejection. The system should not be able to end a candidacy on its own.

"Communication is multi-modal: presence, energy, articulation, and confidence are legitimate signals that hiring managers weigh in every final-round interview," writes Adam Jackson, Braintrust founder and CEO, in a comparison of AI voice and video interviews.

How it compares with the alternatives

Scroll to see all columns

FormatConsistencyFollow-up depthCost per candidateCandidate reaction
Recruiter phone screenLow, varies by interviewer and time of dayHigh when the recruiter is strongHighestGenerally positive
One-way recorded videoHighNone, the question set is fixedLowConsistently poor
Structured human interviewHighHighHigh, and hard to sustain at volumePositive
AI interviewerHigh, identical rubric for every candidateModerate to high, through adaptive follow-upLow, and roughly flat with volumeMixed, and improves when a human review step is disclosed

Scroll to see all columns

The structural advantage is consistency at volume. Structured interviews carry an operational validity of .42, against .19 for unstructured interviews. Human recruiters are capable of running structured interviews. They are rarely resourced to run them for every applicant.

What an AI interviewer does not evaluate

  • Truthfulness. No system detects lying reliably. Claims of deception detection should be treated as a reason to disqualify the vendor.
  • Personality from a face. Inferring traits from facial expression has weak scientific support and carries direct legal exposure.
  • Credentials. A license, degree, or certification is verified with the issuing body, not in an interview.
  • Culture fit. An unstructured impression, which is what a rubric exists to remove.
  • Hands-on execution. Talking about a skill is not the same as demonstrating it, which is why technical roles pair the interview with a work sample.

Does a human ever see it

They should, and in most jurisdictions the design assumption is that they do. Public expectation on this is unambiguous. Pew Research found 71 percent of US adults oppose AI making a final hiring decision, against 7 percent in favor, and 66 percent said they would not want to apply to a job that used AI in hiring decisions.

Candidate trust is similarly low. Gartner found 26 percent of applicants trust AI to evaluate them fairly, and 62 percent said they were more likely to apply when an employer requires in-person interviews at some stage.

There is a measured brand cost too. A meta-analysis of technology-mediated interviews found an small negative effect on how attractive the employer looked, at a correlation of -.17, while the interview ratings themselves showed no significant difference at -.01. Read carefully, that is a useful result. The format does not appear to change who scores well. It changes how the employer is perceived. The mitigation is disclosure, a fast human follow-up, and a clearly communicated review step.

Fairness, accessibility, and the law

An AI interviewer is a selection procedure, and selection procedures are regulated whoever or whatever administers them.

Scroll to see all columns

RequirementWhat it means in practice
NYC Local Law 144An independent bias audit within the past year, published results, and 10 business days' notice to candidates
Illinois AI Video Interview ActNotice, explanation, and consent before AI evaluation, plus deletion within 30 days on request
EU AI Act, Annex III 4(a)Candidate evaluation is high risk, bringing documentation, oversight, and conformity duties
29 CFR 1607.4(D)Adverse impact where a group's selection rate falls below 80 percent of the highest group's rate
29 CFR 1630.11A test must measure the skill, not an impaired sensory, manual, or speaking ability, which makes an accommodation path mandatory rather than optional

Scroll to see all columns

Accessibility deserves its own line because almost nobody writing about AI interviews addresses it. Any voice-based interview inherits the error profile of its speech recognition layer, and that profile is uneven. Across five commercial systems, researchers in 2020 measured an average word error rate of 0.35 for Black speakers against 0.19 for white speakers. Candidates with speech differences, strong regional accents, or non-native fluency face the same mechanism. An employer running these interviews needs a documented alternative route, offered before the interview and not only on request.

Compliance in the market is thin. New York State's Comptroller reported that enforcement of Local Law 144 produced two complaints and 17 instances of potential non-compliance among 32 firms reviewed over a two-year audit period. Braintrust publishes an independent third-party bias audit and maps its obligations by jurisdiction. Ask any vendor, Braintrust included, for the impact ratios by group behind the headline audit result.

Does it actually work

Until recently the only evidence was vendor evidence. That changed with a randomized field experiment covering roughly 70,000 job applicants, in which candidates routed to an AI voice interview were 12 percent more likely to receive a job offer, with higher job starts and higher retention than the human-recruiter control group. Coverage of the same study reports that, when given the choice, 78 percent of applicants chose the AI interviewer over a human recruiter, and that about 7 percent of AI-led interviews hit a technical difficulty.

Two caveats belong with that result. It is one study in one high-volume context, and the 7 percent technical failure rate is a real operational cost that argues for a human fallback path.

Where the return shows up

For high-volume employers, the economics change at the top of the funnel rather than at the offer stage. One enterprise retail client running seasonal surges previously hired temporary contractors to absorb call volume. After automating the first interview, cost per hire fell by more than 60 percent and time to hire moved from 24 days to 6, according to Braintrust client data.

The full model, including which line items survive scrutiny from a finance partner and which do not, is set out in our quantitative ROI analysis.

Braintrust AIR runs live, adaptive interviews, scores against a configured rubric, and returns a ranked scorecard with the recording attached. Candidates preparing for one will find our candidate guide useful, and the product specifics are covered in the AIR FAQ. To experience one from the candidate's side, run a live interview, or book a demo.

Frequently Asked Questions

How does an AI interviewer score an answer

It maps the answer to each rubric criterion defined for the role, assigns a level against written anchors, and attaches the specific response text that produced the score. The recruiter reads the score, the rationale, and the recording together. For the category definition itself, see what AI interview software is.

How does an AI interview agent evaluate candidates

It maps each answer to pre-defined rubric criteria and assigns a score with the supporting response text attached. Good systems publish per-criterion rationale, keep scoring separate from the conversational model, and state what was not assessed.

How is an AI interviewer different from a one-way video interview

A one-way video tool asks fixed questions and records answers with no interaction. An AI interviewer listens and generates follow-up questions in real time, which is what produces depth on vague answers and makes rehearsed scripts harder to sustain.

Does a human review the AI interview

With a properly configured system, yes. Braintrust AIR is built so that it cannot auto-accept or auto-reject, and a recruiter reviews the scorecard and recording before any decision. Candidates in several jurisdictions also have notice and explanation rights.

Are AI interviews fair

Fairness depends on the rubric, the audit, and the accommodation path rather than on the model. Speech recognition error rates differ measurably across speaker groups, so an employer needs impact ratios by group and an alternative route for candidates who need one.

Can an AI interviewer detect lying or cheating

It cannot reliably detect deception, and any vendor claiming otherwise should be removed from your shortlist. Adaptive follow-up questions do make rehearsed or model-generated answers harder to sustain, because they probe specifics a script does not contain.

Do AI interviews use facial recognition

Some systems have historically analyzed facial expression, a practice with weak scientific support and significant legal exposure. Ask any vendor directly whether facial analysis contributes to the score, and get the answer in writing.

What happens to my interview recording

That depends on the employer's configuration and the vendor's policy. Ask about retention period, who inside the employer can view it, whether it is used to train models, and where it is stored. Illinois grants a deletion right within 30 days on request.

What if I have an accent or a speech difference

Ask the employer for an alternative assessment route before the interview. Automatic speech recognition error rates vary by speaker group, and under 29 CFR 1630.11 a test must measure the skill in question rather than an impaired speaking ability.

Do AI interviewers replace recruiters

They replace the first-round screen, not the recruiter. Sourcing strategy, candidate relationships, offer negotiation, and every advance or reject decision stay with people.

ROIHigh VolumeRecruiting
Grady Gardner
Grady Gardner

GM and CRO

See how Braintrust can help

Book a demo to explore AI-powered recruiting, talent marketplace, and workforce automation.

Book a Demo