01The validity problem hiding in plain sight

Most hiring decisions feel rigorous. A manager meets a candidate, a conversation flows, an impression forms. That impression feels meaningful — rich, holistic, human. What the research has shown, repeatedly, since at least the 1950s, is that it largely isn't. The unstructured interview — the free-ranging chat that dominates hiring — is one of the weakest predictors of job performance available. Validity coefficients for unstructured interviews typically hover around 0.20 or below on a scale where 1.0 would be perfect prediction. Work sample tests and structured interviews consistently do better, often substantially so.

The problem is not the interview format itself. It's the variability. When different candidates get different questions, answers can't meaningfully be compared. When interviewers follow their instincts about what to probe, they're often chasing signals that feel diagnostic — confident manner, shared background, a familiar sense of humour — but that the evidence doesn't support. Frank Schmidt and John Hunter's landmark 1998 meta-analysis, aggregating decades of selection research across hundreds of studies, found that structured interviews predicted job performance with clearly higher validity than their unstructured equivalents. That gap has held up in subsequent reviews.

Work sample tests.54General mental ability.51Structured interviews.51Job knowledge tests.48Unstructured interviews.38Reference checks.26Years of experience.18Graphology.02
Predictive validity of common selection methods (correlation with job performance), after Schmidt & Hunter's 1998 meta-analysis. Structured interviews sit near the top of the table; the unstructured conversation does not.

02What structure actually means

Structure in interviewing is not a single switch but a spectrum of practices, and the research distinguishes two key dimensions: question standardisation and evaluation standardisation.

Question standardisation means every candidate faces the same questions, in the same sequence. This sounds obvious but it removes the single largest source of noise in interviewing: the interviewer's improvisation. Two broad question types have accumulated the most evidence. Situational questions ask candidates what they would do in a hypothetical scenario directly relevant to the role ("A client escalates a complaint at the end of a long day — walk me through how you'd handle it"). Behavioural questions ask what they actually did in a real past situation ("Tell me about a time you had to deliver difficult feedback"). Behavioural questions draw on the idea that past behaviour predicts future behaviour; situational questions tap intentions and reasoning. Used together, they cover more ground than either alone.

Evaluation standardisation means scoring answers against pre-defined criteria, typically using a behaviourally anchored rating scale (BARS) — a rubric specifying what a weak, adequate and strong answer looks like for each question. Without this, two interviewers assessing the same answer can reach wildly different conclusions, and the same interviewer can apply different standards across a long day of candidates.

A further evidence-backed addition is the panel interview: having two or more independent raters reduces the idiosyncratic bias any single interviewer carries, and averaging across raters improves reliability. This isn't always practical, but where it is, it pays off.

STRUCTURED VS UNSTRUCTURED — WHERE THE DIFFERENCE LIVES
StructuredUnstructured
QuestionsSame set, same order, role-derivedImprovised per candidate
ScoringAnchored rubric (BARS), scored before discussionOverall impression, recalled later
ComparabilityAnswers line up across candidatesEvery interview is its own event
Bias exposureConstrained by shared criteriaOpen door to affinity and anchoring effects
Predictive validityRoughly twice that of the free-form chat~0.20 — barely better than chance

03Why gut feel keeps winning anyway

If the research is this consistent, why do unstructured interviews remain the norm? Several forces work against structure. One is the illusion of insight: unstructured conversations feel more revealing because they're more varied. An interviewer who probes freely believes they've found something a standard question wouldn't have uncovered. This confidence is usually unwarranted, but it's hard to argue against a feeling.

A second force is the discomfort with looking rigid. Structured interviews can seem cold, bureaucratic, unfair to candidates who perform differently under formulaic conditions. In practice, the evidence runs the other way: structured interviews are more equitable, precisely because every candidate is evaluated on the same evidence. Unstructured interviews give more room for affinity bias — favouring people who look, sound or think like the interviewer — and for the anchoring effects of early impressions that then colour everything that follows.

A third factor is that most interviewers receive little or no training. Building a good structured interview — identifying the competencies the role requires, writing questions that genuinely surface those competencies, constructing BARS, calibrating raters — takes real work upfront. Organisations that invest in that work see the returns; those that don't tend to keep mistaking confidence for competence.

Gut feel without structure is mostly measuring itself.

04Putting it into practice

The practical floor is low. Even partial structure — standardising five or six core questions while scoring on a simple 1–5 rubric with written anchors — beats the improvised conversation. The ceiling, achieved by well-designed competency-based systems with multiple trained raters and BARS, approaches the predictive validity of the best selection tools available.

A few principles hold across the research. Keep the interview focused on the role's critical requirements, not a kitchen-sink list of qualities. Write questions that require specific, concrete answers, not generalities. Train interviewers to record and score before discussing — panel discussion before scoring is one of the quickest ways to lose independent judgment. And resist adding unstructured time at the end to "just get a feel" — that part tends to undo a lot of the work that came before.

The finding is not that gut feel is worthless. It's that gut feel without structure is mostly measuring itself. Give interviewers the same window onto each candidate, and the view gets clearer for everyone.

05Who did the work

Frank Schmidt

industrial-organisational psychologist

co-author of major meta-analyses on selection validity

John Hunter

psychologist

co-author with Schmidt on personnel selection research

EVIDENCE RATING — STRONG

strong. The superiority of structured over unstructured interviewing is one of the most replicated findings in occupational psychology, with multiple large meta-analyses across five decades.