Most vendors describe AI candidate scoring the same way: advanced machine learning matches the best candidates to your role. That sentence is compatible with almost any implementation, including bad ones, which is why it should not reassure anyone.
Here is what ours actually does.
What the score is
The fit score estimates how closely a candidate stated experience matches the requirements of a specific role. It is a similarity estimate against a job definition, not a prediction of job performance.
That distinction is the most important thing on this page. No system that reads a CV can predict how someone will perform, because the CV does not contain the information that would determine it. What a scoring system can do is order a large pile so the most plausible matches are read first.
Used as a reading order, it saves substantial time. Used as a decision, it will make errors you cannot see.
What goes in
- Role requirements as defined in the job — required skills, experience level, and any hard prerequisites you have marked as such.
- Candidate experience extracted from the application and CV: roles held, duration, described responsibilities, skills and tools.
- Structured answers to screening questions, where the answer is a defined value rather than free text.
- Recency, because experience five years ago carries different weight from current experience.
What is deliberately excluded
The exclusions matter more than the inclusions.
- Name, gender, age, nationality, photograph. Excluded from scoring. Where blind hiring is enabled, they are hidden from reviewers too.
- Institution prestige. Which university someone attended is a strong proxy for socioeconomic background and a weak proxy for capability.
- Employer prestige as an independent signal, for the same reason.
- Career gaps, which correlate with caring responsibilities, health and immigration status.
- Writing style and fluency beyond what is needed to extract the content, which would penalise non-native speakers.
Excluding a field is not the same as excluding its influence — proxies exist, and that is why outcome testing is not optional. But excluding the direct signals removes the largest and most avoidable effects.
What comes out
A score, and — the part that matters — the reasons for it. Every scored candidate carries a breakdown showing which requirements were matched, which were not, and how much each contributed.
This exists because a number with no explanation cannot be reviewed. A recruiter shown 82 and nothing else has no basis to disagree, so the only rational response is to accept it. Showing that the score is high because four of five required skills matched, and low on the fifth, gives a reviewer something to exercise judgement about — which is the difference between oversight and rubber-stamping, as covered in what meaningful review requires.
Where the human sits
The score orders a list. It does not reject anyone. There is no configuration in which a candidate is rejected because of a fit score alone.
Automatic rejection is available only on binary hard requirements you define explicitly — no right to work, missing a required licence, unavailable for the only shift pattern. Those are facts, not estimates, and they are the only things we consider safe to automate. The reasoning is in screening questions.
How to audit it
You should not take any of the above on trust. Four checks, all of which you can run yourself.
1. Compare scores against outcomes
Export scores and final outcomes for a closed requisition. Did the people you hired score well? Did high scorers who were rejected get rejected for reasons the score could not see? If there is no relationship at all, the role definition is probably too vague for scoring to work.
2. Test outcomes by group
Selection rates and average scores by demographic group, where you collect that data separately. This is the check that catches proxy effects, and it is the one that matters for every regime we operate under. Method in is CV screening AI biased.
3. Review below the cut line
Take twenty candidates the system ranked low and review them properly. If you find people who should have advanced, the role definition or the extraction is missing something — this is a calibration exercise, not an accusation.
4. Check the override rate
How often do reviewers disagree with the ranking and act on it? A rate near zero means either the score is excellent or nobody is really reviewing. The second is far more common, and review time will tell you which.
What we will not claim
We will not claim the score is unbiased. No system trained on human hiring data can be assumed neutral, and any vendor claiming otherwise is describing a marketing position rather than a technical one. What we claim is narrower and testable: the score is explainable at the individual level, the direct demographic signals are excluded, the outputs and reviewer actions are logged, and you can audit all of it against your own outcomes.
For the record-keeping side, see security and compliance. For the regulatory context, the EU AI Act and hiring.