01The validity problem hiding in plain sight
Most hiring decisions feel rigorous. A manager meets a candidate, a conversation flows, an impression forms. That impression feels meaningful — rich, holistic, human. What the research has shown, repeatedly, since at least the 1950s, is that it largely isn't. The unstructured interview — the free-ranging chat that dominates hiring — is one of the weakest predictors of job performance available. Validity coefficients for unstructured interviews typically hover around 0.20 or below on a scale where 1.0 would be perfect prediction. Work sample tests and structured interviews consistently do better, often substantially so.
The problem is not the interview format itself. It's the variability. When different candidates get different questions, answers can't meaningfully be compared. When interviewers follow their instincts about what to probe, they're often chasing signals that feel diagnostic — confident manner, shared background, a familiar sense of humour — but that the evidence doesn't support. Frank Schmidt and John Hunter's landmark 1998 meta-analysis, aggregating decades of selection research across hundreds of studies, found that structured interviews predicted job performance with clearly higher validity than their unstructured equivalents. That gap has held up in subsequent reviews.
02What structure actually means
Structure in interviewing is not a single switch but a spectrum of practices, and the research distinguishes two key dimensions: question standardisation and evaluation standardisation.
Question standardisation means every candidate faces the same questions, in the same sequence. This sounds obvious but it removes the single largest source of noise in interviewing: the interviewer's improvisation. Two broad question types have accumulated the most evidence. Situational questions ask candidates what they would do in a hypothetical scenario directly relevant to the role ("A client escalates a complaint at the end of a long day — walk me through how you'd handle it"). Behavioural questions ask what they actually did in a real past situation ("Tell me about a time you had to deliver difficult feedback"). Behavioural questions draw on the idea that past behaviour predicts future behaviour; situational questions tap intentions and reasoning. Used together, they cover more ground than either alone.
Evaluation standardisation means scoring answers against pre-defined criteria, typically using a behaviourally anchored rating scale (BARS) — a rubric specifying what a weak, adequate and strong answer looks like for each question. Without this, two interviewers assessing the same answer can reach wildly different conclusions, and the same interviewer can apply different standards across a long day of candidates.
A further evidence-backed addition is the panel interview: having two or more independent raters reduces the idiosyncratic bias any single interviewer carries, and averaging across raters improves reliability. This isn't always practical, but where it is, it pays off.
| Structured | Unstructured | |
|---|---|---|
| Questions | Same set, same order, role-derived | Improvised per candidate |
| Scoring | Anchored rubric (BARS), scored before discussion | Overall impression, recalled later |
| Comparability | Answers line up across candidates | Every interview is its own event |
| Bias exposure | Constrained by shared criteria | Open door to affinity and anchoring effects |
| Predictive validity | Roughly twice that of the free-form chat | ~0.20 — barely better than chance |
03Why gut feel keeps winning anyway
If the research is this consistent, why do unstructured interviews remain the norm? Several forces work against structure. One is the illusion of insight: unstructured conversations feel more revealing because they're more varied. An interviewer who probes freely believes they've found something a standard question wouldn't have uncovered. This confidence is usually unwarranted, but it's hard to argue against a feeling.
A second force is the discomfort with looking rigid. Structured interviews can seem cold, bureaucratic, unfair to candidates who perform differently under formulaic conditions. In practice, the evidence runs the other way: structured interviews are more equitable, precisely because every candidate is evaluated on the same evidence. Unstructured interviews give more room for affinity bias — favouring people who look, sound or think like the interviewer — and for the anchoring effects of early impressions that then colour everything that follows.
A third factor is that most interviewers receive little or no training. Building a good structured interview — identifying the competencies the role requires, writing questions that genuinely surface those competencies, constructing BARS, calibrating raters — takes real work upfront. Organisations that invest in that work see the returns; those that don't tend to keep mistaking confidence for competence.
Gut feel without structure is mostly measuring itself.
04Putting it into practice
The practical floor is low. Even partial structure — standardising five or six core questions while scoring on a simple 1–5 rubric with written anchors — beats the improvised conversation. The ceiling, achieved by well-designed competency-based systems with multiple trained raters and BARS, approaches the predictive validity of the best selection tools available.
A few principles hold across the research. Keep the interview focused on the role's critical requirements, not a kitchen-sink list of qualities. Write questions that require specific, concrete answers, not generalities. Train interviewers to record and score before discussing — panel discussion before scoring is one of the quickest ways to lose independent judgment. And resist adding unstructured time at the end to "just get a feel" — that part tends to undo a lot of the work that came before.
The finding is not that gut feel is worthless. It's that gut feel without structure is mostly measuring itself. Give interviewers the same window onto each candidate, and the view gets clearer for everyone.
05Who did the work
Frank Schmidt
industrial-organisational psychologist
co-author of major meta-analyses on selection validity
John Hunter
psychologist
co-author with Schmidt on personnel selection research
EVIDENCE RATING — STRONG
strong. The superiority of structured over unstructured interviewing is one of the most replicated findings in occupational psychology, with multiple large meta-analyses across five decades.
