What the data actually shows

The most striking finding is the 'idiosyncratic rater effect.' When researchers decompose performance ratings, a large portion of the variance turns out to be specific to the individual rater — their personal tendencies and standards — rather than to the performance being measured. The classic analysis by Scullen, Mount and Goff (2000) found rater idiosyncrasy was the single largest source of variation in ratings, larger than actual performance. In plain terms, your rating tells you a great deal about who your manager is and comparatively little about how you did.

Ratings are also distorted in predictable directions. Recency bias means recent events count far more than the months that preceded them. Leniency and central-tendency biases compress scores toward 'meets expectations.' Halo effects let one strong or weak trait color everything else. None of these are signs of bad managers in particular — they are well-documented features of human judgment under vague criteria.

On effectiveness, the picture is mixed but unflattering for the traditional model. Reviews of the appraisal literature find that formal annual reviews frequently fail to raise subsequent performance, and that feedback delivered alongside a numerical rating or ranking can backfire, because the score triggers defensiveness that crowds out the learning. This is part of why many organizations have shifted toward more frequent check-ins and forward-looking feedback instead of a single retrospective scoring event.

Why this feels different from how it actually is

Performance reviews feel objective because they come with numbers, forms, and a calendar slot. A 3.5-out-of-5 looks like a measurement. But a number is not the same as accurate measurement, and the research shows much of that number is rater noise. The format lends an air of precision the underlying judgment does not have.

They also feel high-stakes because pay, promotion, and standing are often attached to them, so a single annual conversation carries the weight of a year. That weight makes the recency and defensiveness problems worse: people brace for a verdict rather than absorb feedback, and managers, knowing the stakes, often soften or compress their scores.

And because almost every organization runs some version of this ritual, it feels like the natural, even inevitable, way to evaluate work. Its ubiquity is easy to mistake for evidence that it works. The critique is old — quality thinkers like W. Edwards Deming argued decades ago that annual ratings mostly measure variation in the system and the rater, not the individual — but the ritual has proven remarkably durable regardless.

Your rating tells you a great deal about who your manager is and comparatively little about how you did.
On the idiosyncratic rater effect

What the research says to do about it

The most consistent signal is that feedback should be frequent, specific, and forward-looking rather than annual and retrospective. Many organizations that dropped or de-emphasized formal annual ratings replaced them with regular lightweight check-ins focused on what to do next, and report that the conversations became more useful even where the paperwork shrank.

Separating the developmental conversation from the pay-and-ranking decision tends to help, because the score is what triggers defensiveness. When feedback is not fused to a number that determines your raise, people are measurably more able to actually hear and act on it.

On the measurement side, the research favors structure: clear, behavior-based criteria defined in advance, input from multiple sources rather than a single rater, and asking raters to describe specific behaviors rather than assign a global impression. None of this makes appraisal perfectly accurate, but it reduces the share that is pure rater idiosyncrasy.

What the research says does not help

Adding more precision to the rating scale does not fix the core problem. Moving from a 5-point to a 10-point scale, or adding decimal scores, gives the rater idiosyncrasy more room to express itself rather than less. The noise is in the judgment, not the granularity of the form.

Forced ranking — sorting people into a fixed distribution and culling the bottom — has largely fallen out of favor for good reason: the evidence suggests it damages collaboration and morale without reliably improving performance, and it amplifies the very rater biases that make ratings unreliable in the first place.

Saving up feedback for the annual review is one of the least effective approaches. Feedback decays in value the longer it waits, and an end-of-year surprise is both harder to act on and more likely to feel like a verdict than a course correction. Frequency is doing more of the work than the formal event ever does.

A number is not the same as accurate measurement. Much of that number is rater noise — treat any single score as one noisy reading from one instrument, not a verdict.

What this looks like in real life

The mechanism

What a 3.5-out-of-5 actually measures

A score arrives with numbers, forms, and a calendar slot, so it feels like a measurement. But research decomposing ratings finds that much of that number reflects the rater's personal standards rather than the work. In plain terms, the rating often says more about who your manager is than about how you did — the format lends an air of precision the underlying judgment does not have.

Illustrative

Feedback saved up for December

A manager holds back a useful observation from March until the year-end review, where it lands as an end-of-year surprise. By then it is harder to act on and reads as a verdict rather than a course correction. Feedback decays in value the longer it waits — which is why frequency, not the formal event, tends to do most of the work.

Real numbers in context

The headline figure from the appraisal literature is about variance, not a single statistic: in decompositions of multi-rater performance data, rater-specific idiosyncrasy is repeatedly found to be the largest single component of a person's ratings — in the Scullen, Mount and Goff (2000) analysis, larger than the contribution of actual performance. Treat any single review score as one noisy reading from one instrument, not a verdict.

It is worth holding these numbers loosely. Exact percentages vary by study, job type, and how performance is measured, and good feedback genuinely does help people improve. The honest summary is that the annual rating is a weak and biased measurement, the ongoing feedback is the part with value, and the two have been bundled together in a way that serves neither well.

Largest source
Rater idiosyncrasy as a share of variation in performance ratings
Scullen, Mount & Goff, 2000
Mixed / weak
Effect of formal annual reviews on subsequent performance
Performance appraisal research reviews
Recency + leniency
Common biases that distort annual ratings
Appraisal bias literature