The interview scorecard, done properly: three role-specific examples

Generic scorecards get filled in generically. These three are written for specific roles, with the criteria defined tightly enough that two interviewers would score the same candidate the same way.

Interview scorecards fail for a predictable reason: they are written to be reusable, and reusable means vague. "Communication skills: 1-5" produces a number that two interviewers will disagree about while believing they agree.

A scorecard works when the criteria are specific enough that two people watching the same interview arrive at the same score. That means writing one per role family, with defined anchors.

The four rules

1. Four to six criteria, no more

Beyond six, interviewers stop scoring and start pattern-matching to an overall impression, then distributing scores to justify it. Fewer criteria, scored carefully, beat more criteria scored casually.

2. Define what each score means

A 3 has to be described, not assumed. Without anchors, one interviewer 3 is another interviewer 4, and the aggregate is noise.

3. Require evidence, not impressions

Each criterion needs a free-text field, and the instruction is to write what the candidate said or did — not how they seemed. This single change does more for interview quality than any other. It also matters if the notes are ever disclosed: evidence of behaviour is defensible, impressions of a person are not.

4. Score before the debrief

Scores submitted after discussion are contaminated by the loudest voice in the room. Independent scores first, then discuss the disagreements — that is where the signal is.

Example 1: Software engineer, mid-level

Criterion: Problem decomposition

  • 1 — Jumped to implementation without clarifying the problem; missed a stated constraint.
  • 2 — Asked some clarifying questions; approach needed correction to proceed.
  • 3 — Clarified requirements, broke the problem into parts, chose a workable approach unprompted.
  • 4 — Identified an edge case or constraint the interviewer had not raised, and adjusted the approach for it.

Criterion: Code quality under discussion

  • 1 — Code did not work as described; could not explain what it does.
  • 2 — Worked for the main case; structure made changes difficult.
  • 3 — Worked, readable, sensible naming, handled the obvious failure cases.
  • 4 — Made a deliberate trade-off and articulated what it costs.

Criterion: Technical communication — can a colleague follow the reasoning in real time, and does the candidate say when they do not know something.

Criterion: Response to feedback — when challenged, do they defend reflexively, capitulate immediately, or evaluate the point.

Example 2: Account executive

Criterion: Discovery

  • 1 — Pitched before understanding the situation.
  • 2 — Asked surface questions, accepted the first answer.
  • 3 — Asked follow-ups that reached the underlying problem, not the stated one.
  • 4 — Uncovered a constraint the buyer had not articulated and reframed around it.

Criterion: Handling objections — does the objection get acknowledged and addressed, deflected, or argued with.

Criterion: Pipeline discipline — can they describe how they qualify, forecast and disqualify, with a real example of a deal they walked away from.

Criterion: Coachability — ask for a specific piece of feedback they received and what changed afterwards. "I improved my time management" scores 1; a concrete before and after scores 3 or above.

Example 3: Operations manager

Criterion: Process design — given a broken process, do they diagnose the cause or add a control on top of it.

Criterion: Prioritisation under constraint — presented with more work than capacity, do they sequence explicitly and say what they are dropping.

Criterion: Working through others — evidence of getting outcomes from people who do not report to them.

Criterion: Measurement — do they describe outcomes in numbers without being prompted to.

What breaks scorecards in practice

  • Culture fit as a criterion. Undefined, unevidenced, and the single most reliable route to biased hiring. If there is a behaviour you actually need, name and define the behaviour.
  • Scores with no evidence field. Numbers alone cannot be reviewed, challenged, or defended later.
  • A five-point scale with an unused middle. Four points forces a decision; five invites 3 as an escape.
  • Everyone scores everything. Assign criteria across the loop. Four interviewers each scoring the same six criteria produces redundancy, not reliability.

Scorecards only work inside a loop that uses them consistently — see how to run a structured interview loop and debriefs that produce decisions. To build these into the process itself rather than a document, see scorecards in Hireall.

Get hiring insights in your inbox

One email a week — the best playbooks, benchmarks and product updates. No spam.

More from Resources

Everything you need to hire better — in one place.