Structured interviews: guide + rating scales for 2026
In 2022, researchers re-checked a century of hiring studies and the leaderboard flipped: structured interviews now show the highest validity of any selection method examined, ahead of cognitive ability tests.
A structured interview asks every candidate the same job-related questions, in the same order, and scores the answers on the same anchored scale. Same criteria, same scale, written evidence per rating.
This guide walks through the evidence, six steps to run your first structured interview, the rating-scale mechanics, and a bias checklist. One clarification up front: you standardize the core questions, their sequence, the scale, and the approved follow-ups.
Small talk and clarifying questions stay human.
Key takeaways
The hiring playbook, in your inbox
One email a week - benchmarks, AI screening tactics, and short interview templates from the 100Hires team. No product pitches.
- Definition: same core questions, same order, same anchored rating scale for every candidate on a role
- Evidence: a 2022 meta-analytic revision puts structured interviews at .42 validity, the highest among the selection methods it examined
- You build four artifacts: a competency list, a question bank, an anchored scale, and a scorecard
- Structure the scoring, not the small talk: Google found rejected candidates were 35% more satisfied after structured interviews
- Independence beats headcount: individual structured interviews predicted performance better (.46) than panel interviews (.38) in McDaniel's meta-analysis
What is a structured interview
The U.S. Office of Personnel Management defines it as an interview where every candidate gets the same job-related questions and every answer is rated against predefined criteria.
In plainer terms: same questions, same order, graded on a scoring system you wrote before anyone walked in.
Teams rarely call it that. Greenhouse calls the artifact an interview kit. Recruiters on X say scorecard or rubric. 100Hires calls it an evaluation form. The labels overlap; the method underneath is the same one.
Three quick disambiguations. Researchers also use "structured interview" for a survey technique in academic studies, a different use of the term. Employee evaluations are post-hire performance reviews, a different concept.
And structured does not mean behavioral-questions-only: U.S. Customs and Border Protection runs structured interviews built on hypothetical scenarios, scored against defined competencies by a trained panel.
Here is how the two interview styles compare on the dimensions that matter to a hiring team.
| Dimension | Structured interview | Unstructured interview |
|---|---|---|
| Questions | Same core set, same sequence, written in advance | Improvised per candidate and per interviewer |
| Scoring | Anchored rating scale, evidence per rating | Overall impression, often after the fact |
| Interviewer role | Asks, probes with approved follow-ups, records evidence | Steers freely, relies on memory and feel |
| Candidate experience | Predictable and fair; rejected candidates rate it higher | Varies by interviewer mood and skill |
| Legal exposure | Consistent, documented record | Loose process that is harder to defend |
| Validity | .42 (Sackett) to .44 (McDaniel), sourced below | .33 in McDaniel's data |
That validity row deserves its own section, because the research consistently favors this method even though hiring teams use it far less than they should.
Do structured interviews actually work
Yes, and the strength of the evidence is the surprising part.
For decades the textbook answer was that cognitive ability tests predict job performance best. In 2022, Sackett, Zhang, Berry, and Lievens re-examined the underlying meta-analyses in the Journal of Applied Psychology and corrected a long-standing statistical overcorrection.
In their revised estimates, structured interviews came out on top of the methods examined at .42, ahead of cognitive ability tests at .31.
Honest caveat: the .42 is an average. The same paper reports an 80% credibility interval from .24 to .66, so the payoff depends on how well you build and run the process. The six steps below are what "run it well" means.
The older evidence points the same direction. McDaniel and colleagues analyzed 245 validity coefficients covering 86,311 people in 1994: structured interviews predicted job performance at .44 versus .33 for unstructured ones.
The operational payoff is measurable too. Google's internal research, published in its re:Work guide, found structured interviewing saved interviewers about 40 minutes per interview, since questions and rubrics exist before the calendar invite.
Using Google's 40-minute estimate, conducting 20 interviews a month saves roughly 13 hours.
Candidates prefer it too. In the same Google data, rejected candidates were 35% more satisfied than rejected candidates from unstructured interviews. A fair process reads as fair even when the answer is no.
There is a legal thread as well. The OPM guide cites research linking loose, inconsistent interviews to discrimination challenges, including Terpstra, Mohamed, and Kethley (1999) and the U.S. Merit Systems Protection Board (2003).
A documented, consistent process reduces that risk. It does not eliminate it, and nothing here is legal advice.
One limit to the claim: structure improves the information you gather during interviews, not everything that happens afterward. Recruiters on practitioner threads point out that many "bad hires" trace to onboarding gaps or manager misalignment, which no interview format prevents.
How to conduct a structured interview in 6 steps
Every major methodology source recommends the same basic sequence: define competencies, write questions, anchor a scale, then calibrate the people using it.
We split that into six steps and build one running example along the way: hiring a support specialist, starting from a single competency called customer communication.
Step 1: define 4 to 6 competencies from the job
Start from the job description, not from your favorite questions. List the qualifications, then group them into three buckets: must-have technical skills, must-have behavioral dimensions, and one or two role-specific outcomes.
Cap the list at 4 to 6 competencies. Overloaded scorecards die fast: interviewers skim, boxes go blank, and the process quietly reverts to gut feel.
Recruiters who have rolled scorecards out at startups report that adding more questions just produces more one-character comments; the real gap is competency definition, not form length.
Write a two-sentence definition of what great looks like for each competency. For our support hire: "Customer communication: explains a technical fix in plain words without blaming the customer. Stays calm and specific when the customer is angry."
Methodologies like Topgrading build a full job scorecard before any candidate is seen, which is the right instinct. The same methodology runs chronological interviews that stretch across most of a workday, which is the extreme end. You do not need it to get the validity gain.
Step 2: build the question bank
Two question types carry a structured interview. Behavioral questions ask about the past: "Tell me about a time an angry customer escalated an issue to you. What did you do?"
Situational questions pose a hypothetical: "A customer says the fix you shipped broke something else. Walk me through your next ten minutes."
Write both types for each competency. Workable's rule of thumb is about two questions per competency, so a six-competency role needs roughly twelve. That is a vendor guideline, not research, but it matches what a complete interview process can cover.
Skip brain teasers. Google said publicly years ago that they predict nothing, and its own research backs the ban.
Values questions belong in the bank too, decomposed into observable behaviors.
Recruiters debating the long-car-ride compatibility test landed on the fix: candidates are auditioning to be coworkers, not friends, so ask about the behavior behind the value instead of checking vibes.
Plenty of teams now draft question banks with ChatGPT or Claude, and that is a fine starting point as long as humans own the criteria. One hiring lead we spoke with ran roughly 600 applicants through a rubric that lived inside a chat window, hand-feeding batches of ten.
The questions were decent; the process around them is what broke. For inspiration beyond AI drafts, our sample interview questions with scoring notes cover seven proven ones.
Step 3: anchor the rating scale
This is the step top-ranking guides skip, and it is where most scorecards fail. An unanchored 1-10 scale invites safe middle scores from everyone. The OPM guide recommends five to seven rating levels, with at least three of them carrying concrete behavioral anchors.
Here is our customer communication competency, anchored at levels 1, 3, and 5.
| Level | Anchor for customer communication |
|---|---|
| 5 | Reframed the escalation in plain words, owned the miss without blaming anyone, and gave the customer a specific next step with a time |
| 3 | Explained the fix accurately but leaned on jargon; stayed professional under pushback with some prompting |
| 1 | Blamed the customer or another team; answer stayed vague after two follow-up probes |
Anchors turn "good communicator" from a comment into a rating two interviewers can agree on. If a five-level scale feels heavy for a first rollout, practitioners note that even a three-tier scale beats an unanchored number.
One approach that scales well: keep a standard core (the same 1-5 anchors and 3 to 4 values questions for every role) and let hiring managers customize only the role-specific questions. One team describes this as an 80/20 split that cut prep time without losing comparability.
If you want the ready-made version of this artifact, our structured interview scorecard page covers the form itself in depth.
Step 4: score on a shared scorecard, evidence per rating
Every interviewer involved in that stage completes the same form: a rating for each competency plus the evidence behind it. A rating of 3 with no evidence cannot be meaningfully challenged at debrief.
A 3 with "used the word 'deprecated' three times to an angry customer" is something the panel can argue about.
Our worked example ends like this. Evidence note: "Walked me through the broken-fix scenario calmly, named the apology, the workaround, and the follow-up time. Jargon-free."
Rating: 5 on customer communication. Debrief decision: advance, with the panel probing time management next round.
Question order is part of the structure. A CodeSignal founder describes easy-to-hard ordering as a fairness decision: changing the order for one candidate undermines the comparison.
The same discipline applies to the debrief: review one rubric area at a time instead of collecting overall verdicts.
Decide attribute by attribute rather than reducing each candidate to an overall thumbs-up or thumbs-down. The author of Cracking the Coding Interview scores each dimension separately and accepts some false negatives to avoid false positives.
Two more conventions worth adopting: the common 70/30 rule, where the candidate does most of the talking, and the practitioner guideline of about three core questions per 30-minute conversation.
Step 5: calibrate the panel and keep ratings independent
Here is the counterintuitive finding: in McDaniel's meta-analysis, individual structured interviews predicted performance at .46 versus .38 for panels. More interviewers in the room did not mean better judgment. What seems to matter is that each judgment is formed independently.
Earlier ratings anchor the ones that follow. One recruiter tracking their own debriefs realized the first candidate to score a 7 became the bar, and the next five interviews were comparisons to her. Six interviews, one real evaluation.
Three fixes. Run a kickoff where the hiring manager assigns each panelist a focus area, so four people stop asking the same favorite question. Require written scores before any discussion.
And keep ratings blind until submitted: in 100Hires, the Blind Evaluations setting hides peers' scores until an interviewer files their own, so earlier ratings cannot influence later ones.
Collaborative hiring works better when the collaboration starts after the independent judgment.
Train the people too. Occupational psychologist Memory Nguwi's informal survey put untrained panelists above 80%, and recruiter threads echo it: most interviewers were never taught to interview.
Even one calibration session, where the panel scores a recorded or mock interview together, gets everyone applying the same standard.
One cautionary anecdote on consistency: a recruiter who tracked interview times across identical roles and rubrics found 71% of before-noon candidates advanced versus 26% after 3pm.
That is one person's data, not a study, but a reason to watch when you schedule and how tired your panel is.
Step 6: audit and improve the process
A structured process drifts unless someone checks it. Quarterly, look at four numbers: scorecard completion rate, interviewer disagreement patterns, candidate feedback on the process, and pass-through rates by stage for signs of adverse impact.
Persistent outliers are a calibration signal, not noise. And set an override protocol now: exceptions to the rubric are allowed with documented job-related evidence, and every override triggers a rubric review.
If judgment keeps beating the rubric, the rubric is measuring the wrong things: fix it, do not abandon it.
How do structured interviews reduce bias
By replacing the moments where bias enters: undefined criteria, improvised questions, and memory-based scoring. The OPM guide names the rating errors a structured process is built to catch. Use them as a literal checklist during calibration:
- Similar-to-me: rating people higher for resembling the interviewer
- Halo effect: one strong answer lifting every other rating
- Central tendency: everything scored a safe middle number
- Leniency and strictness: one interviewer's 4 is another's 2
- First-impression rater bias: the opening minute deciding the final score
The fairness evidence goes beyond error names. A 2023 Cambridge-published analysis by Huffcutt and colleagues found that structured interviews have strong validity with lower adverse impact on racial groups than most alternatives.
Two cautions. A rubric can encode bias if built carelessly: questions framed around hobbies such as chess or fantasy football quietly favor whoever already dominates the team's demographics, so frame questions around the skills the role actually requires.
And take-home assignments carry a fairness cost of their own: caregivers and people with second jobs pay for them in scarce hours. A live, structured exercise measures the same skill more evenly.
Structure means consistent criteria, not identical logistics. Reasonable accommodations, such as extra time or a format change, fit inside a structured process and are required where the law applies.
On the legal side, treat structure as risk reduction: documented, consistent evaluations are easier to stand behind months later, and interview notes are discoverable either way.
Common structured interview mistakes
- Reading a verbatim script. Candidates describe fully scripted interviews as mind-numbing. Hold the core questions, sequence, anchors, and approved probes constant; deliver them like a person.
- Using the rubric as a hard gate. An auto-reject rubric selects for people who are good at interviews. The rubric structures evidence; humans still make the call.
- Stage creep. On a recruiting podcast, one People leader recalls a delivery company finding 10-12 stage loops in an internal audit, then capping stages per level - an insider-told story, not a study. The practitioner sweet spot is 4 to 6 stages.
- Blank evidence boxes. Most scorecard failures are adoption failures. Reduce friction, require evidence only on decisive ratings, and chase completion instead of adding questions.
- Sharing scorecards with candidates. Working recruiters treat this as a legal liability, not a transparency win. Share the competencies you evaluate; keep the internal ratings internal.
- Hiding the criteria internally too. Rubric criteria the panel never aligned on mostly test who had insider access. The kickoff meeting exists to fix this.
When does gut feel beat the rubric
The case against structured interviews deserves a fair hearing, because it gets far more attention. Viral clips of famous founders urging leaders to rely on instinct outperform any research summary posted next to them.
Four objections are legitimate. Some leaders' best hires would have failed the rubric; exceptional profiles exist. Stale rubrics select for last year's job when nobody revisits the criteria. A two-decimal average of gut guesses is still a gut guess dressed as precision.
And rubric upkeep is real work a five-person team may not prioritize.
Notice what those objections attack: bad structure. Verbatim scripts, auto-reject gates, criteria written once and never audited. None of them argues that improvising different questions for each candidate produces better evidence.
The maintenance point is fair, though. If you hire twice a year, a full rubric with quarterly audits is overkill, and a lightweight one-page scorecard is the rational tradeoff. The gain scales with volume.
Wharton professor Ethan Mollick has repeatedly summarized the research: unstructured interviews with no preset criteria can make hiring decisions worse than not interviewing at all.
Unstructured interviewing persists for a simpler reason: it takes no work up front. No design work, no training, no scoring. The costs appear later, in poor hires and indefensible decisions.
So give gut feel a vote, not a veto. An exception to the rubric needs documented, job-related evidence, and it triggers the Step 6 rubric review. That keeps human judgment in the process without letting the loudest instinct run it.
Are structured interviews still worth it in the AI era
More than before. Resumes tell you less now that AI writes them, and take-home tests stopped being reliable once AI could complete them. Finance recruiters report firms dropping take-homes for exactly this reason.
The live, structured conversation is becoming the part of the process AI cannot fake for the candidate.
AI still earns a place in the process. It drafts question banks and anchor language well, and it can score responses inside a process humans designed: in 100Hires, the AI score sits next to the human scorecards as a second opinion, never a replacement.
Candidates, meanwhile, describe interviews run entirely by AI agents as an ordeal. The structure should come from software; the conversation should stay human.
How to run structured interviews in 100Hires
Full disclosure: 100Hires is our ATS. Here is exactly how the six steps map to it, verified against our own product documentation.
You build Evaluation Forms in Settings. You build a form once with custom questions and several question types, including rating scales, then assign it to a job or attach it when scheduling an interview.
Interviewers get daily email reminders until they submit, and completed evaluations sit side by side on the candidate profile for the debrief.
For informal feedback, interviewers can also use star ratings on the Discussion tab. For structured feedback, the evaluation form is the tool, and the distinction matters for record-keeping.
Scheduling connects to the same flow, so the right form is attached when the interview gets booked. Evaluation data is retrievable over the API for custom reporting.
Here is the before-and-after for a team running interviews out of spreadsheets.
| Task | With 100Hires | Spreadsheet or docs |
|---|---|---|
| Store the rubric | Evaluation form built once, reused per job | Reusable template file, copied per role |
| Distribute to the panel | Form attached at scheduling | Manual link distribution per interview |
| Chase completion | Automatic daily reminders | Manual reminder follow-ups |
| Compare candidates | Evaluations side by side on the profile | Cross-referencing separate sheets |
| Keep independent ratings | Blind Evaluations toggle, peers' scores hidden until submitted | Early entries visible to later raters |
Spreadsheets deserve respect: they are how most teams start, and they build the habit before the software does. The upgrade is about removing the friction that kills the habit at scale.
Limitations worth knowing. 100Hires does not offer candidate anonymization: Blind Evaluations hides peer scores until submission (admins keep visibility), and it does not hide names or identities from screeners.
If you need deep scorecard analytics of the kind Greenhouse is known for, that is Greenhouse's strength; ours is helping a small team run structured interviews inside the same tool that already posts the jobs and books the interviews.
Start a free 14-day trial and build your first evaluation form in an afternoon, or book a demo to see the workflow live.
Frequently asked questions
What is the difference between structured and unstructured interviews?
A structured interview uses the same job-related questions, order, and anchored rating scale for every candidate; an unstructured one improvises per candidate. The Sackett 2022 revision put structured validity at .42, and McDaniel's 1994 meta-analysis put it at .44 versus .33 for unstructured. In 100Hires, the structured version is one evaluation form every interviewer completes for a job.
What are examples of structured interview questions?
Behavioral: "Tell me about a time you missed a deadline. What happened next?" Situational: "A customer reports your fix broke something else. What are your first ten minutes?" Write about two per competency, and store them in a reusable form; 100Hires evaluation forms keep the bank attached to the job so every panel uses the same questions.
How do you score candidates in a structured interview?
Rate each competency on a short anchored scale (OPM recommends five to seven levels with at least three labeled anchors) and record the evidence behind every rating. 100Hires evaluation forms pair rating-scale questions with text fields for evidence, and completed evaluations line up side by side on the candidate profile.
Do structured interviews feel robotic to candidates?
Only when interviewers read a script. Standardize the questions, sequence, and scale; keep rapport and follow-ups conversational. Google's research found rejected candidates were 35% more satisfied after structured interviews. 100Hires structures the scoring in the background, so the conversation itself stays human.
Are panel interviews better than one-on-one interviews?
Not automatically. McDaniel's meta-analysis found individual structured interviews predicted performance better (.46) than panels (.38). Independence is what matters: written scores before discussion, and blind ratings. The 100Hires Blind Evaluations setting hides peers' scores until each interviewer submits their own.
Are structured interviews more legally defensible?
Research cited by OPM links loose, inconsistent interviews to more discrimination challenges, so a consistent, documented process reduces risk without guaranteeing anything. Submitted evaluations in 100Hires create a retrievable record of who rated what and why, months after the decision.
You now have the four artifacts: competencies, a question bank, an anchored scale, and a scorecard. Start your free 100Hires trial and make structured interviews your team's default.
Try 100Hires for free
No credit card. 14-day trial. Forbes Advisor #1 ATS for SMBs.