Criterion Contamination: When Performance Reviews Measure Things Employees Can't Control
A performance rating should measure performance. In practice it absorbs luck, territory quality, the manager's quirks, and the team someone happened to join.
On this page▾
- The technical idea
- The five biggest contaminants
- The controllability test
- How to decontaminate reviews
- Limits and caveats
- Contamination and deficiency: the two ways a measure fails
- How big is the problem?
- Common contaminants in real reviews
- A cleaner review design
- Questions to ask in calibration
- Worked case: the rating that measured visibility
- A contamination check for every review
- Contamination vs deficiency
- What each role can do
- Frequently asked questions
- Criterion contamination: a performance measure includes things that are not part of the job performance you meant to measure.
- Its twin, criterion deficiency, is when the measure misses important parts of the job.
- Scullen, Mount and Goff (2000) found that the rater's own idiosyncrasies explained more variance in ratings than the ratee's actual performance.
- Deming argued most variation in outcomes comes from the system, not the individual.
- Fixes: adjust for context, separate outcomes from behaviours, use multiple raters, and ask 'could this person have controlled this?' for every metric.
Two account managers. Maya gets the Kathmandu enterprise territory with three renewing clients. Dev inherits a region where the largest customer just went bankrupt. Maya finishes at 130% of quota and is rated 'exceptional'. Dev finishes at 70% after saving two accounts nobody thought were savable, and is rated 'needs improvement'. The review system worked exactly as designed. It just measured the territory, not the people.
The technical idea
Industrial-organisational psychologists distinguish between the 'ultimate criterion' — the true, complete concept of job performance — and the 'actual criterion' — the measure you can collect. The gap produces two errors. Criterion deficiency is when your measure misses important parts of the job (for example, rating engineers only on tickets closed and ignoring code reviews). Criterion contamination is when your measure includes things that are not job performance — luck, resources, the rater's mood, or factors outside the employee's control. Austin and Villanova's history of 'the criterion problem' (1992) shows this has troubled the field for a century.
The five biggest contaminants
- 1The raterScullen, Mount and Goff (2000) analysed multi-source ratings of thousands of managers and found that idiosyncratic rater effects accounted for more variance in ratings than the actual performance of the person rated. A rating often tells you as much about the manager as the employee.
- 2Luck and market conditionsBertrand and Mullainathan (2001) showed CEOs are rewarded for 'luck' — for example oil-price movements beyond their control — about as much as for general performance. The same thing happens lower down.
- 3Resources and territoryBudgets, tools, headcount, and the quality of the patch someone inherits all show up as 'performance'.
- 4Team and system effectsW. Edwards Deming argued that most variation in results comes from the system, not the worker. Ranking individuals on system-driven outcomes punishes people for the system's design.
- 5Visibility and proximityRemote or quieter employees can be rated lower for work that was simply less seen.
When contaminated ratings drive bonuses, promotions, or redundancy selection, the contamination becomes pay inequity. If a contaminant — such as territory allocation or flexible-working visibility — correlates with gender, ethnicity, or disability, it can become a discrimination risk.
The controllability test
The simplest practical tool is to ask, for every metric in a review: 'To what extent could this person have changed this result through their own choices?' Rate each metric high, medium, or low on controllability. Low-controllability metrics can still be tracked — they matter to the business — but they should carry less weight in individual evaluation, or be adjusted for context.
| Metric | Controllability | Treatment |
|---|---|---|
| Revenue vs quota | Medium — depends on territory | Adjust quota by territory potential |
| Pipeline quality and activity | High | Weight heavily |
| Client retention | Medium | Review alongside client-risk context |
| Customer satisfaction | Medium to high | Use, with product issues flagged separately |
| Market-wide price changes | Low | Exclude from individual rating |
How to decontaminate reviews
- Separate 'what happened' (outcomes) from 'what the person did' (behaviours and decisions).
- Normalise outcomes for context: territory potential, team size, inherited problems.
- Use multiple raters and calibration so one manager's idiosyncrasies don't dominate.
- Train raters on specific behavioural anchors rather than general impressions.
- Evaluate decision quality: given what they knew at the time, was the call sound?
- Audit rating outcomes for patterns by team, location, work arrangement, and demographic group.
Limits and caveats
- Some roles are legitimately paid for outcomes, luck included — for example, commission sales. The question is whether that is a conscious choice.
- Over-adjusting can create endless excuses. Context should inform ratings, not replace accountability.
- Exact rater-effect sizes vary by study and instrument; the consistent finding is that they are large.
Contamination and deficiency: the two ways a measure fails
In industrial-organisational psychology, the 'criterion' is the measure you use to judge performance. A criterion can fail in two ways. It is contaminated when it includes things that are not really performance — luck, territory quality, the rater's mood. It is deficient when it leaves out things that are part of performance, such as helping colleagues or preventing problems. Most review systems suffer from both at once.
How big is the problem?
A widely cited study by Scullen, Mount and Goff (2000) analysed multi-rater data for thousands of managers. They found that the largest share of variance in ratings — around 62% — was explained by the rater's own tendencies, while actual performance explained far less. In other words, a rating often tells you as much about the person giving it as the person receiving it. This 'idiosyncratic rater effect' is one of the clearest forms of contamination.
Common contaminants in real reviews
| Contaminant | Example | Fix |
|---|---|---|
| Territory or portfolio | A rep given the strongest region | Adjust targets to opportunity |
| Team and tools | An engineer blocked by broken infrastructure | Record constraints alongside results |
| Timing | Results depend on a market cycle | Compare with peers in the same conditions |
| Visibility | Remote staff are less seen by leaders | Use evidence logs, not memory |
| Rater style | Some managers rate everyone higher | Calibration and rater training |
| Recency | The last month dominates the review | Regular check-in notes through the year |
A cleaner review design
- 1Separate outcomes from behavioursRate what the person did (behaviours, decisions, quality of work) separately from results that depend on context.
- 2Record constraintsAsk employees and managers to note what helped or blocked performance during the period.
- 3Calibrate ratersCompare rating patterns across managers and discuss differences before ratings are final.
- 4Include contextual performanceBorman and Motowidlo's work shows helping, cooperating, and volunteering matter; measure them explicitly to reduce deficiency.
Questions to ask in calibration
- What would this person's result have been with an average territory, team, or tool set?
- Is this rating based on evidence from the whole period or the last few weeks?
- Would another manager with the same evidence give the same rating?
- What did this person do that does not appear in the numbers?
- Review forms separate behaviours from results.
- Constraints are documented during the year.
- Rater distributions are compared across managers.
- Pay decisions are not based on a single rating number alone.
Worked case: the rating that measured visibility
A constructed example of a common pattern.
Two engineers deliver similar work. One works in the office, presents at every demo and replies to messages within minutes. The other works remotely, ships quietly and writes the documentation everyone relies on. At calibration, the first is rated 'exceeds', the second 'meets'. When the manager is asked for evidence, most of what he lists for the first engineer is visible activity: demos, meetings, quick responses.
That is criterion contamination: the rating has absorbed things that are not the job's real outcomes. Visibility is not worthless — communication matters — but if it is not in the written expectations, it should not quietly decide the rating.
A contamination check for every review
- Every rating comment points to an outcome or behaviour listed in the role's expectations.
- Evidence comes from the whole period, not only the last six weeks.
- The reviewer has asked: would I rate this the same if this person worked in another location?
- Results outside the person's control (a lost client, a frozen budget) are noted and set aside.
- Peer input is from people who saw the work, not only people who sit nearby.
Contamination vs deficiency
- Office presence
- Personality similar to the manager
- Luck with accounts or projects
- Recent events only
- Mentoring and documentation ignored
- Risk avoided, not counted
- Team-level outcomes missing
- Customer impact not tracked
Most review systems have both problems at once. Fixing contamination without fixing deficiency just leaves a narrower, still unfair measure.
What each role can do
- HR: rewrite the rating form so each rating requires a linked piece of evidence.
- Calibration leaders: ask 'what is the evidence?' every time a rating is defended with a general impression.
- Managers: keep a short monthly note per person so the review is not built from memory.
- Employees: keep your own record of outcomes and share it before the review.
Frequently asked questions
What is criterion contamination in simple terms?
When a performance measure includes things that aren't really the person's performance, like luck, resources, or rater bias.
What is criterion deficiency?
When a performance measure misses important parts of the job.
Is it the same as rater bias?
Rater bias is one source of contamination. Luck, resources, and system effects are others.
What is the quickest improvement?
Run a controllability audit on your review metrics and separate outcomes from behaviours.
A review that measures the territory, the manager, and the weather is not a performance review. Measure what people control, and treat everything else as context.
- Scullen, S., Mount, M. & Goff, M. (2000). Understanding the latent structure of job performance ratings — Journal of Applied Psychology
- Austin, J. & Villanova, P. (1992). The criterion problem: 1917–1992 — Journal of Applied Psychology
- Bertrand, M. & Mullainathan, S. (2001). Are CEOs Rewarded for Luck? — Quarterly Journal of Economics
- Deming, W. E. (1986). Out of the Crisis — MIT Press
- Scullen, S., Mount, M. & Goff, M. (2000). Understanding the latent structure of job performance ratings — Journal of Applied Psychology
- Borman, W. & Motowidlo, S. (1993). Expanding the criterion domain to include elements of contextual performance — Personnel Selection in Organizations
- Opportunity-to-Perform Bias: Are Your 'Top Performers' Actually Getting the Best Opportunities?
- The Ratchet Effect: Why Your Best Employees Learn to Hide Their True Capacity
- The McNamara Fallacy: How People Analytics Loses the War It's Winning on the Dashboard
- Boundary Spanners: The Employees Holding Your Company Together Without Anyone Noticing
Read next
All playbooksPerformance depends on ability, motivation — and opportunity. When some people get the best projects, clients, and tools, their results look like talent.
When this year's great result becomes next year's baseline, smart people stop showing you what they can really do.
Robert McNamara ran Vietnam by the numbers he could measure and lost the war he couldn't. His fallacy has four stages: measure what you can, disregard what…