How the assessment works

An explainability statement for candidates, employers and their lawyers: what the AI interviewer looks at, how it arrives at its numbers, what it cannot do, and where the person comes in.

Last updated 9 September 2026.

What the interview is

A 30-minute conversation with an AI interviewer, conducted from a plan the employer configured: the role, the level, the interview type (technical, behavioral or similar), a focus area, the job description, topics to make sure to cover, and up to five questions the employer wants asked word for word. By default every candidate in a screening gets the same plan; an employer can instead have a fresh plan generated for each candidate from the same configuration, to discourage question sharing. There is an optional 5-minute practice run first that is never scored.

What is scored, and what is not

  • Scored: the transcript of your answers, typed or spoken and transcribed, read against the plan. An answer counts as spoken only when our own speech-to-text produced most of its text for this interview, or when the employer required spoken answers; nothing your browser sends can mark an answer as spoken. The scorer is told, answer by answer, which were spoken and machine-transcribed and which were typed: transcription artifacts are not held against a spoken answer, and a typed answer is read as written. If you introduce yourself by name, that is in the transcript the scorer reads; nothing removes it.
  • Not scored: your face, expressions, tone of voice, accent, or anything from the camera or screen recording. Those recordings are shown to the employer as evidence; the scorer never receives them. The scorer is given no timing rule, and no per-turn times of any kind are in the transcript it reads.
  • Not scored: the session events (tab switches, time away, camera state, connection changes). They are shown to the employer as context with timestamps. They are not score inputs and never cause a rejection. One thing derived from the timeline does reach the scorer: whether you stopped taking part before the end (no answer and no microphone, editor or whiteboard activity for the final third or more). It is told so, grades what you answered on its merits, and counts the unused time.
  • Language: for interviews conducted in a language other than English, the scorer is instructed that the language and your fluency in it are not assessment dimensions and may not be named in the reasoning. For English interviews there is no separate fluency rule: spoken answers are graded as speech-to-text with transcription artifacts excluded, and the communication dimension grades clarity of thought and structure.

How the numbers are produced

The scorer is a large language model from OpenAI, prompted with the plan, the transcript (including any code or drawings), and a list of which parts of the plan the interviewer reached. It returns four dimension scores with written reasoning for each, and an overall number from 0 to 100:

  • Objectives alignment (up to 30): how well the answers address the objectives the plan set for the interview.
  • Level expectations (up to 30): whether the answers show the depth expected at the role's level.
  • Technical proficiency (up to 20): depth and correctness in the subject, or the equivalent experience dimension for non-technical interviews.
  • Communication and problem solving (up to 20): how clearly the reasoning is laid out.

Anything in the plan the interviewer never reached is listed as "not assessed" rather than scored. A question the interviewer did not ask is the interviewer's miss, not the candidate's. If a session ends before its time, the answers that were given are graded on their merits, and the unused share of the session is stated, not hidden.

The employer sees the overall number as an "AI signal" with the caveat "a signal, not a decision", beside the transcript, the recording and the four reasoned dimensions. The scorer's own free-text recommendation is removed before anything reaches the employer: no hire or no-hire label, and no ranking computed by us; the employer's own table can be sorted by any column, including the signal.

Where the person comes in

Nothing is decided by software. The employer's recruiter reads the evidence and records Advance or Reject themselves. They can record their own score with a written reason beside the AI signal, and hiring managers can vote. If you are turned down, ask the employer for the reasoning: they can see it for every dimension, and they can reconsider.

Known limitations

  • It is a language model. The same transcript can score differently on two runs: in our test, six unchanged runs of one transcript spread from 87 to 89.
  • Less said, less evidence. Short or vague answers score low for lack of evidence, not because they are wrong. The interviewer asks follow-up questions for that reason.
  • It cannot verify facts about you. It assesses what you said, not whether it is true. The employer's later stages do that.
  • It cannot see you. That is deliberate (see how interviews are monitored), and it means the scorer cannot make allowances a human interviewer might make from context.

What we test

On 9 September 2026 we ran an internal test on synthetic transcripts: 17 transcripts (technical and behavioral, strong to weak), each scored under 8 variants that changed only the candidate's self-introduced name, drawn from name sets associated in US research with white, Black, Hispanic and East Asian candidates, each in a male and a female form, and on 6 of the transcripts a gendered pronoun in one sentence. 136 scoring calls through the production scorer, unchanged.

  • The largest gap between group averages was 1.1 points on the 0-100 scale, inside the scorer's own run-to-run noise: the same transcript scored six times unchanged spread from 87 to 89.
  • The four-fifths (adverse impact) ratio was 1.00 at a threshold of 60 and 0.88 at 70 (95% interval 0.71 to 1.00). A ratio under 0.80 would indicate adverse impact.
  • 3.6% of paired scores moved by more than 5 points between variants of the same transcript, and none by more than 9.

What this cannot show. With 17 transcripts the interval at the 70 threshold reaches below 0.80, so a small effect is not ruled out. Most synthetic transcripts scored above both thresholds, which limits what a selection-rate test can detect; the next run uses transcripts built to straddle them. The six-transcript pronoun subset was too small to interpret (its ratio was 0.67 with an interval from 0.33 to 1.00; recorded, not dropped). Only names and pronouns were varied, not dialect, phrasing or background details. It is the start of an anti-bias log, not an independent audit under NYC Local Law 144 or similar. The full write-up, with the name sets and every number, is available on request.

Also tested on 9 September 2026, and not shipped: an instruction telling the scorer that fluency in English is not an assessment dimension. On nine borderline transcripts that were already fluent English it raised scores by about four points on average and moved three or four of the nine across the 70 line, all upward; a version that told the scorer the instruction changed nothing for fluent transcripts did the same. A rule that inflates borderline scores regardless of fluency is not a fairness rule, so the English fluency gap stays open and is recorded in the anti-bias log with the numbers.

We will support an independent audit when a customer's jurisdiction requires one, including supplying synthetic test data, which NYC Local Law 144 permits; none has been performed yet.

Changes

  • 9 September 2026: spoken answers are marked as transcriptions answer by answer from our own transcription record, also when speaking is optional. Per-turn times removed from the transcript the scorer reads.
  • 9 September 2026: first version.

For employers with statutory duties

This page is the plain-language description of the system you can give a candidate. The Colorado documentation lists the developer documentation and the deployer checklist. The candidate privacy notice is shown to every candidate before they consent.

How the assessment works | InterviewStack for Recruiters