Skip to content
Playbook
AdvancedHRFounderCEO

Criterion Contamination: When Performance Reviews Measure Things Employees Can't Control

A performance rating should measure performance. In practice it absorbs luck, territory quality, the manager's quirks, and the team someone happened to join.

7 min read
On this page▾
60-Second Summary
  • Criterion contamination: a performance measure includes things that are not part of the job performance you meant to measure.
  • Its twin, criterion deficiency, is when the measure misses important parts of the job.
  • Scullen, Mount and Goff (2000) found that the rater's own idiosyncrasies explained more variance in ratings than the ratee's actual performance.
  • Deming argued most variation in outcomes comes from the system, not the individual.
  • Fixes: adjust for context, separate outcomes from behaviours, use multiple raters, and ask 'could this person have controlled this?' for every metric.

Two account managers. Maya gets the Kathmandu enterprise territory with three renewing clients. Dev inherits a region where the largest customer just went bankrupt. Maya finishes at 130% of quota and is rated 'exceptional'. Dev finishes at 70% after saving two accounts nobody thought were savable, and is rated 'needs improvement'. The review system worked exactly as designed. It just measured the territory, not the people.

The technical idea

Industrial-organisational psychologists distinguish between the 'ultimate criterion' — the true, complete concept of job performance — and the 'actual criterion' — the measure you can collect. The gap produces two errors. Criterion deficiency is when your measure misses important parts of the job (for example, rating engineers only on tickets closed and ignoring code reviews). Criterion contamination is when your measure includes things that are not job performance — luck, resources, the rater's mood, or factors outside the employee's control. Austin and Villanova's history of 'the criterion problem' (1992) shows this has troubled the field for a century.

Relevance, deficiency and contamination
Relevance
What the measure captures that is truly performance
Deficiency
Real performance the measure misses
Contamination
What the measure captures that isn't performance
Goal
Maximise relevance, minimise the other two

The five biggest contaminants

What sneaks into a performance rating
  1. 1
    The rater
    Scullen, Mount and Goff (2000) analysed multi-source ratings of thousands of managers and found that idiosyncratic rater effects accounted for more variance in ratings than the actual performance of the person rated. A rating often tells you as much about the manager as the employee.
  2. 2
    Luck and market conditions
    Bertrand and Mullainathan (2001) showed CEOs are rewarded for 'luck' — for example oil-price movements beyond their control — about as much as for general performance. The same thing happens lower down.
  3. 3
    Resources and territory
    Budgets, tools, headcount, and the quality of the patch someone inherits all show up as 'performance'.
  4. 4
    Team and system effects
    W. Edwards Deming argued that most variation in results comes from the system, not the worker. Ranking individuals on system-driven outcomes punishes people for the system's design.
  5. 5
    Visibility and proximity
    Remote or quieter employees can be rated lower for work that was simply less seen.
Why this is a pay and legal issue

When contaminated ratings drive bonuses, promotions, or redundancy selection, the contamination becomes pay inequity. If a contaminant — such as territory allocation or flexible-working visibility — correlates with gender, ethnicity, or disability, it can become a discrimination risk.

The controllability test

The simplest practical tool is to ask, for every metric in a review: 'To what extent could this person have changed this result through their own choices?' Rate each metric high, medium, or low on controllability. Low-controllability metrics can still be tracked — they matter to the business — but they should carry less weight in individual evaluation, or be adjusted for context.

Controllability audit example: account manager
MetricControllabilityTreatment
Revenue vs quotaMedium — depends on territoryAdjust quota by territory potential
Pipeline quality and activityHighWeight heavily
Client retentionMediumReview alongside client-risk context
Customer satisfactionMedium to highUse, with product issues flagged separately
Market-wide price changesLowExclude from individual rating

How to decontaminate reviews

  1. Separate 'what happened' (outcomes) from 'what the person did' (behaviours and decisions).
  2. Normalise outcomes for context: territory potential, team size, inherited problems.
  3. Use multiple raters and calibration so one manager's idiosyncrasies don't dominate.
  4. Train raters on specific behavioural anchors rather than general impressions.
  5. Evaluate decision quality: given what they knew at the time, was the call sound?
  6. Audit rating outcomes for patterns by team, location, work arrangement, and demographic group.

Limits and caveats

  • Some roles are legitimately paid for outcomes, luck included — for example, commission sales. The question is whether that is a conscious choice.
  • Over-adjusting can create endless excuses. Context should inform ratings, not replace accountability.
  • Exact rater-effect sizes vary by study and instrument; the consistent finding is that they are large.

Contamination and deficiency: the two ways a measure fails

In industrial-organisational psychology, the 'criterion' is the measure you use to judge performance. A criterion can fail in two ways. It is contaminated when it includes things that are not really performance — luck, territory quality, the rater's mood. It is deficient when it leaves out things that are part of performance, such as helping colleagues or preventing problems. Most review systems suffer from both at once.

What your rating actually captures
Relevant
True performance the rating measures correctly
Contamination
Measured, but not performance: luck, bias, context
Deficiency
Performance the rating misses: helping, prevention
Noise
Random error from one-off events and memory

How big is the problem?

A widely cited study by Scullen, Mount and Goff (2000) analysed multi-rater data for thousands of managers. They found that the largest share of variance in ratings — around 62% — was explained by the rater's own tendencies, while actual performance explained far less. In other words, a rating often tells you as much about the person giving it as the person receiving it. This 'idiosyncratic rater effect' is one of the clearest forms of contamination.

~62%
Share of rating variance linked to rater tendencies
Scullen, Mount & Goff (2000)
2
Types of performance often confused: task and contextual
Borman & Motowidlo (1993)
8
Performance dimensions in Campbell's model
Campbell (1990)

Common contaminants in real reviews

Things reviews measure that employees cannot control
ContaminantExampleFix
Territory or portfolioA rep given the strongest regionAdjust targets to opportunity
Team and toolsAn engineer blocked by broken infrastructureRecord constraints alongside results
TimingResults depend on a market cycleCompare with peers in the same conditions
VisibilityRemote staff are less seen by leadersUse evidence logs, not memory
Rater styleSome managers rate everyone higherCalibration and rater training
RecencyThe last month dominates the reviewRegular check-in notes through the year

A cleaner review design

Four steps to reduce contamination
  1. 1
    Separate outcomes from behaviours
    Rate what the person did (behaviours, decisions, quality of work) separately from results that depend on context.
  2. 2
    Record constraints
    Ask employees and managers to note what helped or blocked performance during the period.
  3. 3
    Calibrate raters
    Compare rating patterns across managers and discuss differences before ratings are final.
  4. 4
    Include contextual performance
    Borman and Motowidlo's work shows helping, cooperating, and volunteering matter; measure them explicitly to reduce deficiency.

Questions to ask in calibration

  • What would this person's result have been with an average territory, team, or tool set?
  • Is this rating based on evidence from the whole period or the last few weeks?
  • Would another manager with the same evidence give the same rating?
  • What did this person do that does not appear in the numbers?
  • Review forms separate behaviours from results.
  • Constraints are documented during the year.
  • Rater distributions are compared across managers.
  • Pay decisions are not based on a single rating number alone.

Worked case: the rating that measured visibility

Illustrative composite

A constructed example of a common pattern.

Two engineers deliver similar work. One works in the office, presents at every demo and replies to messages within minutes. The other works remotely, ships quietly and writes the documentation everyone relies on. At calibration, the first is rated 'exceeds', the second 'meets'. When the manager is asked for evidence, most of what he lists for the first engineer is visible activity: demos, meetings, quick responses.

That is criterion contamination: the rating has absorbed things that are not the job's real outcomes. Visibility is not worthless — communication matters — but if it is not in the written expectations, it should not quietly decide the rating.

A contamination check for every review

  • Every rating comment points to an outcome or behaviour listed in the role's expectations.
  • Evidence comes from the whole period, not only the last six weeks.
  • The reviewer has asked: would I rate this the same if this person worked in another location?
  • Results outside the person's control (a lost client, a frozen budget) are noted and set aside.
  • Peer input is from people who saw the work, not only people who sit nearby.

Contamination vs deficiency

Contamination (measuring too much)
  • Office presence
  • Personality similar to the manager
  • Luck with accounts or projects
  • Recent events only
Deficiency (measuring too little)
  • Mentoring and documentation ignored
  • Risk avoided, not counted
  • Team-level outcomes missing
  • Customer impact not tracked

Most review systems have both problems at once. Fixing contamination without fixing deficiency just leaves a narrower, still unfair measure.

What each role can do

  • HR: rewrite the rating form so each rating requires a linked piece of evidence.
  • Calibration leaders: ask 'what is the evidence?' every time a rating is defended with a general impression.
  • Managers: keep a short monthly note per person so the review is not built from memory.
  • Employees: keep your own record of outcomes and share it before the review.

Frequently asked questions

What is criterion contamination in simple terms?

When a performance measure includes things that aren't really the person's performance, like luck, resources, or rater bias.

What is criterion deficiency?

When a performance measure misses important parts of the job.

Is it the same as rater bias?

Rater bias is one source of contamination. Luck, resources, and system effects are others.

What is the quickest improvement?

Run a controllability audit on your review metrics and separate outcomes from behaviours.

The takeaway

A review that measures the territory, the manager, and the weather is not a performance review. Measure what people control, and treat everything else as context.

Written by Pawan Joshi.Sources cited inline.
First published 29 Sept 2026See site changelog →