An AI interviewer evaluates candidates by scoring each answer against a rubric written before the interview, then handing a recruiter the score, the reasoning behind it, and the recording. It does not form an overall impression, and it does not decide who gets hired. This page is about that evaluation step, which is the part buyers ask about last and get burned by first.
For the category definition and how the tooling compares with adjacent systems, start with what AI interview software is. What follows assumes you have that and want the scoring mechanism, the failure modes, and the evidence on whether any of it works.
The short answer
- The mechanism. Capture the answer, transcribe it, score each response against rubric criteria, attach the text that produced each score, then hand the scorecard to a recruiter.
- The rubric is the product. A score with no written anchors and no traceable rationale cannot be reviewed, only accepted.
- Does it work. In a randomized field experiment across roughly 70,000 applicants, candidates interviewed by AI voice were 12 percent more likely to receive an offer.
- The catch. Public opinion is not with it. Pew found 71 percent of US adults oppose AI making final hiring decisions, which is exactly why the human review step is not optional.
How an AI interviewer works, stage by stage
Scroll to see all columns
| Stage | What happens | Who controls it |
|---|---|---|
| 1. Role configuration | The competencies, question set, and scoring rubric are defined for the role | The recruiter or hiring manager, before any candidate sees it |
| 2. Invitation | The candidate receives a link, usually within minutes of applying, and starts whenever they choose | Automated from the applicant tracking system |
| 3. Conversation | The agent asks, listens, and generates a contextual follow-up based on the answer given | The model, inside the boundaries of the configured question set |
| 4. Transcription | Speech is converted to text by an automatic speech recognition model | The model, and this is where accent-related error enters |
| 5. Scoring | Each response is mapped to rubric criteria and scored, with the supporting text attached | A separate scoring pass, ideally isolated from the conversational model |
| 6. Scorecard | A ranked, filterable report with scores, rationale, transcript, and recording | Delivered to the recruiter |
| 7. Decision | A person reads the evidence and decides who advances | The employer, always |
Scroll to see all columns
Stage three is what separates a real AI interviewer from a one-way video tool. A one-way tool asks a fixed question, records an answer, and moves on. An adaptive agent hears a vague answer and asks the question a good human interviewer would ask next. That is covered in more depth in our explainer on adaptive interviewing.
How an AI interviewer evaluates candidates
An AI interviewer evaluates candidates by scoring each answer against pre-defined rubric criteria, not by forming an overall impression. This is the part most explanations skip, and it is the part that determines whether the output is defensible.
A rubric criterion, written out
Take a customer-facing role and the competency "de-escalation under pressure." A usable rubric does not say "rate 1 to 5." It defines what each level looks like:
Scroll to see all columns
| Level | What the answer contains |
|---|---|
| 1 | No specific incident. Describes intention rather than action. |
| 2 | A real incident, but the candidate's own actions are vague or passive. |
| 3 | A specific incident with named actions taken by the candidate. |
| 4 | Specific actions plus the reasoning behind them, and what the candidate weighed at the time. |
| 5 | All of the above, plus the outcome and what the candidate changed afterwards. |
Scroll to see all columns
Scores against anchors like these are reviewable. A recruiter can open the transcript, see the sentence that produced a 4, and disagree with it. A single composite number with no anchors and no trace cannot be reviewed, only accepted.
What to demand from any scoring engine
- Per-criterion rationale. Every score links to the specific response text behind it.
- Separation of concerns. The model that conducts the conversation should not be the model that grades it, and scoring should return a fixed schema rather than free prose.
- Explicit non-assessment. The report should state what was not evaluated, so nobody infers a judgment the system never made.
- A human review trigger. Defined rules for when a case must go to a person, such as a technical failure, a very short interview, or a candidate-flagged issue.
- No auto-rejection. The system should not be able to end a candidacy on its own.
"Communication is multi-modal: presence, energy, articulation, and confidence are legitimate signals that hiring managers weigh in every final-round interview," writes Adam Jackson, Braintrust founder and CEO, in a comparison of AI voice and video interviews.
How it compares with the alternatives
Scroll to see all columns
| Format | Consistency | Follow-up depth | Cost per candidate | Candidate reaction |
|---|---|---|---|---|
| Recruiter phone screen | Low, varies by interviewer and time of day | High when the recruiter is strong | Highest | Generally positive |
| One-way recorded video | High | None, the question set is fixed | Low | Consistently poor |
| Structured human interview | High | High | High, and hard to sustain at volume | Positive |
| AI interviewer | High, identical rubric for every candidate | Moderate to high, through adaptive follow-up | Low, and roughly flat with volume | Mixed, and improves when a human review step is disclosed |
Scroll to see all columns
The structural advantage is consistency at volume. Structured interviews carry an operational validity of .42, against .19 for unstructured interviews. Human recruiters are capable of running structured interviews. They are rarely resourced to run them for every applicant.
What an AI interviewer does not evaluate
- Truthfulness. No system detects lying reliably. Claims of deception detection should be treated as a reason to disqualify the vendor.
- Personality from a face. Inferring traits from facial expression has weak scientific support and carries direct legal exposure.
- Credentials. A license, degree, or certification is verified with the issuing body, not in an interview.
- Culture fit. An unstructured impression, which is what a rubric exists to remove.
- Hands-on execution. Talking about a skill is not the same as demonstrating it, which is why technical roles pair the interview with a work sample.
Does a human ever see it
They should, and in most jurisdictions the design assumption is that they do. Public expectation on this is unambiguous. Pew Research found 71 percent of US adults oppose AI making a final hiring decision, against 7 percent in favor, and 66 percent said they would not want to apply to a job that used AI in hiring decisions.
Candidate trust is similarly low. Gartner found 26 percent of applicants trust AI to evaluate them fairly, and 62 percent said they were more likely to apply when an employer requires in-person interviews at some stage.
There is a measured brand cost too. A meta-analysis of technology-mediated interviews found an small negative effect on how attractive the employer looked, at a correlation of -.17, while the interview ratings themselves showed no significant difference at -.01. Read carefully, that is a useful result. The format does not appear to change who scores well. It changes how the employer is perceived. The mitigation is disclosure, a fast human follow-up, and a clearly communicated review step.
Fairness, accessibility, and the law
An AI interviewer is a selection procedure, and selection procedures are regulated whoever or whatever administers them.
Scroll to see all columns
| Requirement | What it means in practice |
|---|---|
| NYC Local Law 144 | An independent bias audit within the past year, published results, and 10 business days' notice to candidates |
| Illinois AI Video Interview Act | Notice, explanation, and consent before AI evaluation, plus deletion within 30 days on request |
| EU AI Act, Annex III 4(a) | Candidate evaluation is high risk, bringing documentation, oversight, and conformity duties |
| 29 CFR 1607.4(D) | Adverse impact where a group's selection rate falls below 80 percent of the highest group's rate |
| 29 CFR 1630.11 | A test must measure the skill, not an impaired sensory, manual, or speaking ability, which makes an accommodation path mandatory rather than optional |
Scroll to see all columns
Accessibility deserves its own line because almost nobody writing about AI interviews addresses it. Any voice-based interview inherits the error profile of its speech recognition layer, and that profile is uneven. Across five commercial systems, researchers in 2020 measured an average word error rate of 0.35 for Black speakers against 0.19 for white speakers. Candidates with speech differences, strong regional accents, or non-native fluency face the same mechanism. An employer running these interviews needs a documented alternative route, offered before the interview and not only on request.
Compliance in the market is thin. New York State's Comptroller reported that enforcement of Local Law 144 produced two complaints and 17 instances of potential non-compliance among 32 firms reviewed over a two-year audit period. Braintrust publishes an independent third-party bias audit and maps its obligations by jurisdiction. Ask any vendor, Braintrust included, for the impact ratios by group behind the headline audit result.
Does it actually work
Until recently the only evidence was vendor evidence. That changed with a randomized field experiment covering roughly 70,000 job applicants, in which candidates routed to an AI voice interview were 12 percent more likely to receive a job offer, with higher job starts and higher retention than the human-recruiter control group. Coverage of the same study reports that, when given the choice, 78 percent of applicants chose the AI interviewer over a human recruiter, and that about 7 percent of AI-led interviews hit a technical difficulty.
Two caveats belong with that result. It is one study in one high-volume context, and the 7 percent technical failure rate is a real operational cost that argues for a human fallback path.
Where the return shows up
For high-volume employers, the economics change at the top of the funnel rather than at the offer stage. One enterprise retail client running seasonal surges previously hired temporary contractors to absorb call volume. After automating the first interview, cost per hire fell by more than 60 percent and time to hire moved from 24 days to 6, according to Braintrust client data.
The full model, including which line items survive scrutiny from a finance partner and which do not, is set out in our quantitative ROI analysis.
Braintrust AIR runs live, adaptive interviews, scores against a configured rubric, and returns a ranked scorecard with the recording attached. Candidates preparing for one will find our candidate guide useful, and the product specifics are covered in the AIR FAQ. To experience one from the candidate's side, run a live interview, or book a demo.
Frequently Asked Questions
How does an AI interviewer score an answer
It maps the answer to each rubric criterion defined for the role, assigns a level against written anchors, and attaches the specific response text that produced the score. The recruiter reads the score, the rationale, and the recording together. For the category definition itself, see what AI interview software is.
How does an AI interview agent evaluate candidates
It maps each answer to pre-defined rubric criteria and assigns a score with the supporting response text attached. Good systems publish per-criterion rationale, keep scoring separate from the conversational model, and state what was not assessed.
How is an AI interviewer different from a one-way video interview
A one-way video tool asks fixed questions and records answers with no interaction. An AI interviewer listens and generates follow-up questions in real time, which is what produces depth on vague answers and makes rehearsed scripts harder to sustain.
Does a human review the AI interview
With a properly configured system, yes. Braintrust AIR is built so that it cannot auto-accept or auto-reject, and a recruiter reviews the scorecard and recording before any decision. Candidates in several jurisdictions also have notice and explanation rights.
Are AI interviews fair
Fairness depends on the rubric, the audit, and the accommodation path rather than on the model. Speech recognition error rates differ measurably across speaker groups, so an employer needs impact ratios by group and an alternative route for candidates who need one.
Can an AI interviewer detect lying or cheating
It cannot reliably detect deception, and any vendor claiming otherwise should be removed from your shortlist. Adaptive follow-up questions do make rehearsed or model-generated answers harder to sustain, because they probe specifics a script does not contain.
Do AI interviews use facial recognition
Some systems have historically analyzed facial expression, a practice with weak scientific support and significant legal exposure. Ask any vendor directly whether facial analysis contributes to the score, and get the answer in writing.
What happens to my interview recording
That depends on the employer's configuration and the vendor's policy. Ask about retention period, who inside the employer can view it, whether it is used to train models, and where it is stored. Illinois grants a deletion right within 30 days on request.
What if I have an accent or a speech difference
Ask the employer for an alternative assessment route before the interview. Automatic speech recognition error rates vary by speaker group, and under 29 CFR 1630.11 a test must measure the skill in question rather than an impaired speaking ability.
Do AI interviewers replace recruiters
They replace the first-round screen, not the recruiter. Sourcing strategy, candidate relationships, offer negotiation, and every advance or reject decision stay with people.


