Ask a vendor whether their screening AI is biased and you will get a reassurance. Ask a researcher and you will get a qualified yes. Both are unhelpful, because the question is not answerable in general — only about a specific tool, used on a specific population, for a specific role.
The useful version is: how would I find out about mine? That is measurable, and the measurement is not difficult.
Where bias enters
Training data
A model learning from historical hiring decisions learns the pattern of those decisions, including any systematic preferences in them. If a company historically hired mostly one profile for a role, a model trained on that outcome will rate that profile highly — and will be accurate about the past while being wrong about capability.
Proxy variables
The subtler and more common route. Removing name and gender does not remove their influence, because other fields correlate with them: university, postcode, career gaps, sports and society membership, language patterns, employment continuity. A model with no access to gender can still reproduce a gendered outcome through fields that track it.
Requirement definition
Frequently the largest effect, and it has nothing to do with the model. If the job requires ten years of continuous experience, the tool will faithfully disadvantage anyone who took time out for caring responsibilities. The tool did what it was told; the requirement was the discriminatory element.
Language and fluency
Systems that reward polished writing penalise non-native speakers and candidates without access to CV coaching, for roles where writing fluency may be irrelevant.
Why removing names is not enough
Name removal is the most-cited intervention and the least sufficient. It addresses direct discrimination by an evaluator, which matters for human review — see does blind hiring work — but it does almost nothing about proxy effects in an automated system.
The only reliable test is outcome-based: measure what the system actually does to different groups, rather than reasoning about what it should do given its inputs.
How to measure it
Step 1: Get demographic data, separately
Voluntary self-identification collected at application, stored separately from the evaluation record, never visible to evaluators or to the scoring system. If you do not have this, it is the first thing to build, and it takes a quarter to accumulate enough data to be meaningful.
Step 2: Pick a resolved cohort
One role family, one period, candidates who reached a terminal state. Mixing role families produces a number that describes nothing.
Step 3: Compute selection rates
For each group, the proportion advancing past the stage where the tool is used. Compute at each stage separately — bias frequently appears at one stage and is invisible in an end-to-end number.
Step 4: Compute impact ratios
Each group selection rate divided by the highest group rate. A commonly used rule of thumb treats ratios below four-fifths as warranting investigation. It is a screening heuristic, not a legal standard, and it is unreliable on small samples.
Step 5: Check the score distribution, not just pass rates
Average scores by group, and the shape of the distribution. Pass rates can look equal while the underlying scores differ systematically — which will surface as soon as you change the threshold.
Step 6: Compare against the pre-tool baseline
The most informative comparison and the most often skipped. If your human process had a disparity and the tool reproduced it, the tool did not create the problem — but it did automate and scale it, which is a different and larger exposure.
What to do with a disparity
- Check the requirements first. The most common cause is a requirement that is not genuinely necessary. Removing it is cheaper and more effective than any model adjustment.
- Look at which factors drove the scores. If a single factor explains most of the gap, you have a specific and fixable problem.
- Check the sample size before acting. Below a few hundred per group, impact ratios move on noise. Do not restructure a process on twelve candidates.
- Record what you found and what you did. A dated review that found a disparity and a documented response is a far stronger position than no review.
How often
Quarterly for most teams. Per intake for high-volume hiring, where a small systematic effect reaches a large number of people fast. Whether you need an external audit on top of this is a separate question — the decision tree is here.
For what our own scoring includes and excludes, see how the fit score works.