An interview score is only useful if it means the same thing across candidates, across weeks, and across the model versions running underneath it. A score of 3.4 in March that would have been a 3.9 in October is not a score. It is noise with a decimal point.
Nova Recruiter generates structured interviews from a job description and scores the responses. This post describes the three mechanisms that keep those scores stable: how we write rubrics, how we freeze a calibration set, and how we detect drift before a client does.
Rubrics that constrain the judgment
Most scoring inconsistency is a rubric problem, not a model problem. A criterion like "communicates clearly" gives an evaluator, human or otherwise, nothing to anchor against. Ask five people to score the same answer on that criterion and you get five distributions.
Every Nova Recruiter rubric criterion has to satisfy five constraints before it ships. These came out of eighteen months of watching which criteria produced tight agreement and which produced arguments.
- Behaviorally anchored. Each score level describes an observable thing the candidate said or did, not a trait they possess.
- Four levels, never five. Five-point scales collapse to the middle. Four forces a directional call.
- Independently scorable. If criterion B cannot be scored without knowing criterion A's score, they are one criterion.
- Job-linked. Every criterion traces to a requirement in the job description, which is also what makes it defensible under EEOC scrutiny.
- Written with a negative anchor. The level-1 description has to be as concrete as the level-4 description.
The negative anchor rule matters more than it sounds. Rubrics written only from the top down describe excellence in detail and failure in a sentence, and the resulting scores cluster in the upper half. Writing the floor as carefully as the ceiling spreads the distribution back out.
The calibration set is frozen
We maintain a held-out set of 1,240 recorded interview responses spanning 31 role families, each scored independently by three trained human evaluators against the shipped rubric. Where the three disagreed by more than one level, a fourth senior evaluator adjudicated and the item was flagged as ambiguous. Ambiguous items stay in the set. They are the ones that catch regressions.
The set is frozen. It does not get refreshed when scores look bad, and no calibration item is ever used to tune the scoring model. That separation is the entire point. The moment your evaluation set becomes training data, you lose your only independent read on whether the system still works.
A calibration set you are allowed to edit when results disappoint is not a measurement. It is a mirror.
Every candidate scoring model version runs the full calibration set before release. We gate on three things: quadratic weighted kappa against the human consensus at 0.78 or above, mean absolute error under 0.42 on the four-point scale, and no subgroup where the mean score shifts more than 0.15 from the previous shipped version. That last gate has blocked more releases than the first two combined.
Drift is continuous, so monitoring is too
Nothing about a production scoring system stays still. Candidate populations shift with the labor market. Clients add new role families. A model provider updates a version underneath you. Rubrics get edited by client admins. Any of these moves the score distribution without anyone changing the scoring logic.
We run four monitors on a weekly cadence, each with an alerting threshold that pages the evaluation team rather than filing a ticket someone reads on Friday.
- Score distribution shift, measured as population stability index per rubric criterion per role family. Alert above 0.2.
- Calibration replay. The frozen set reruns weekly against production config, not just at release. Alert on kappa dropping below 0.75.
- Human override rate. Recruiters can override any AI score. A rising override rate on one criterion is the earliest signal we get.
- Subgroup mean deltas, computed on the aggregate level for adverse impact monitoring under NYC Local Law 144 and EEOC guidance.
The override monitor has been the most valuable of the four. It is the only one grounded in a human who looked at the same interview and disagreed. When overrides on a criterion cross 12 percent for two consecutive weeks, we pull the criterion for review regardless of what the statistical monitors say.
What we do when drift is real
Detection is the easy half. The response has to be decided before the alert, or you end up making a scoring policy call under time pressure with a client on the phone.
Our runbook has three tiers. A distribution shift with stable kappa means the population changed, and we rebaseline the reporting without touching scoring. A kappa drop with a stable population means the model or config regressed, and we roll back to the last gated version. A rubric-specific override spike means the criterion is ambiguous in a context we did not anticipate, and it goes back to rubric authoring.
The human stays in the loop
None of this makes an AI score a hiring decision. Nova Recruiter produces a structured, rubric-anchored evaluation with the supporting evidence attached to each criterion. A human recruiter makes the advance-or-reject call, and that call is logged alongside the score.
That design is partly a compliance posture under the EU AI Act and Local Law 144. It is also just better engineering. The override signal is our best drift detector, and you only get that signal if a person is genuinely empowered to disagree with the machine.
Sofia MarchettiStaff ML Engineer, Evaluation at Novexhire