The evidence on structured interviews has been stable for four decades and is not seriously disputed. Ask every candidate the same questions, rate against defined anchors, and you roughly double the predictive validity of the interview relative to an unstructured conversation. It is one of the few genuinely settled findings in the field.
Almost nobody does it. Not because hiring managers disagree with the research. Most of them can cite it. The reason is that the format asks a human being to run a rubric and hold a conversation at the same time, and human beings are bad at that.
What the research actually says
Two things get conflated. Structure is not one dial, it is two.
- Question structure. Every candidate gets the same questions in the same order, drawn from the requirements of the role rather than from whatever the interviewer thought of on the way in.
- Evaluation structure. Every answer is rated against defined levels, written down before the interview, with an example of what a two looks like versus a four.
The second matters more than the first, and it is the one that gets dropped. Plenty of teams have a shared question list and a scorecard with five competencies rated one to five, with nothing anywhere defining what a three means. That is question structure with the evaluation left as an exercise for the reader, and it recovers very little of the predictive gain.
Why it collapses in practice
Watch someone run a structured interview properly and the problem is obvious within ten minutes. They are doing four jobs at once:
- Asking the question as written, without leading
- Listening well enough to ask a good follow-up
- Taking notes detailed enough to justify a rating later
- Holding the anchors in mind so the rating is against the rubric rather than against the last candidate
Three and four lose. Notes degrade into fragments, ratings get filled in afterwards from memory, and the memory is dominated by the most recent five minutes and the candidate's warmth. The scorecard gets completed, the process is described as structured, and the actual mechanism that produces the validity gain never ran.
We have a name for this internally. Scorecard theatre: all the artefacts of structure, none of the measurement.
Bookkeeping is the enemy of listening
The instinct is to fix this with discipline: better training, stricter templates, a reminder to fill the scorecard in within an hour. It helps at the margin and it does not survive a busy week.
The realistic fix is to remove the bookkeeping from the interviewer entirely. If the interview is recorded and transcribed, then the mapping from what was said to what the rubric asks about is a retrieval problem, not a memory problem. The interviewer's only job becomes the one they are good at: asking a real question and listening to the answer.
That produces a scorecard where each competency arrives with the passages that bear on it, timestamped, and the interviewer's job is to agree, disagree, or push back on evidence rather than to reconstruct an hour from four lines of handwriting.
Making the structure invisible
The design goal we ended up with is that a well-structured interview should feel less structured to the candidate, not more.
This sounds contradictory and is not. What makes a structured interview feel like a deposition is the interviewer's visible bookkeeping: the eyes going to the notes, the pause while something is typed, the mechanical transition to the next item. Remove those and what is left is a conversation that happens to cover the same ground every time.
The things we deliberately did not automate:
- Follow-ups. The second question is where the information is, and it depends on what was just said. Scripting it defeats the purpose.
- The rating itself. Evidence is assembled automatically. The judgement is made by the interviewer, on the record, and it is theirs.
- The decision. A scorecard is an input to a debrief, not a replacement for one.
Calibration is the whole game
The last piece is the one teams skip and then wonder why their scores do not mean anything across interviewers.
Take three recorded interviews. Have everyone who will run the loop rate them independently against the anchors. Then compare. The first time a team does this, the spread is usually two full points on at least one competency, and the conversation that follows, about what a four actually looks like on this competency for this role, is worth more than any amount of interviewer training.
Do it again after twenty interviews. If the spread has not closed, the anchors are the problem, not the people.
Structured interviewing is not a form to fill in. It is a shared definition of the thing being measured, maintained by argument, with the paperwork moved somewhere it cannot interrupt the listening.




