Twenty minutes into a demo, the prospect stopped smiling. “Can you flip this? Make 1 the best score and 5 the worst.” The rep said “let me check” and logged a ticket that afternoon.
By the time it reached the product manager, someone had attached a screenshot to explain why: the customer's own tracker, 1 to 5, best to worst, unchanged for years. One look was enough. A leaderboard, dressed up as a performance scale.
And a leaderboard asks a different question than a performance review should: are you comparing this person to who they used to be, or to the person next to them.
Take two managers filling out the same form. One scores Maya, who jumped from "meets expectations" to "exceeds expectations" this year, worth a real conversation. The other scores Arjun, who came in 4th on a five-person team, same as last year. Maya's score tells you something happened. Arjun's tells you almost nothing.
A rating measures how much of a quality someone has, against a standard. A ranking measures where someone stands next to everyone else, right now. A 4 out of 5 this year and a 4.5 next year still track the same person, improving. A rank of 4th out of 5 just tells you the room he was standing in, and moving him to a stronger team can drop that rank to 5th without him changing at all.
Researchers who study how grant panels score applications ran into this exact tension and gave it a name: absolute versus relative judgment. Same applicant, two different scores, depending only on which question the reviewer was answering.
Flip through enough of these requests and a pattern shows up. “1 is best, 5 is worst” is race-result logic. First place, second place, last place. Nobody finishes a marathon with a 4.2. They finish 4th.
That's exactly the logic Jack Welch built into GE's performance system in the 1980s: the vitality curve, later shorthanded as the 20-70-10 rule. The top 20% of a team were the stars, the middle 70% were adequate, and the bottom 10% were managed out, regardless of what the bottom 10% had actually accomplished that year.
GE cared about standing against peers, not performance against a standard, on purpose. The system existed to force a decision: who gets the bonus pool, who gets cut. Ranking is well suited to that job. It's a terrible fit for the question most performance reviews are actually supposed to answer: is the person improving?
Microsoft ran a version of the GE model for years. Managers scored employees on a scale of 1 to 5, and the scores had to fall on a bell curve regardless of how the team actually performed. In November 2013, HR head Lisa Brummel ended it in a company-wide memo with two lines worth reading twice: “No more curve.” “No more rankings.”
Both phrases were needed, because both problems were live. The curve was the forced distribution: a fixed percentage had to land in each bucket. The ranking was the logic underneath it: employees graded against each other rather than against a standard. Former employees described the day-to-day effect as colleagues quietly working against each other instead of the market, because a neighbor's good quarter was suddenly a threat to your own score.
That isn't an isolated complaint. A 2021 study in the Journal of Accounting Research put forced, relative rating systems in a lab and measured what changed. Forced ratings didn't raise output. They raised stress, measured through both surveys and stress biomarkers, and that stress reduced the very creativity the review was meant to reward. Under the forced system, how well someone wrote up their work mattered more to their score than what they'd actually done.
Ranking has its place. Forced calibration, mostly. It's just the wrong tool for what most performance reviews are trying to do.
Next time a scale request shows up looking like a cosmetic tweak, it's worth sitting with a few questions first:
Most of the time, what the customer actually needs is a rating: a consistent, absolute scale that lets a manager track the same person's growth from one cycle to the next, and lets HR compare across teams honestly. Ranking has a real job, forced calibration for a fixed bonus pool, for instance, but it belongs as a separate, deliberate step downstream of the rating, not baked into the base scale.
As for the prospect from the demo, the one whose screenshot started this whole conversation: the scale stayed absolute, cycle over cycle. The flip they'd asked for showed up somewhere else instead, as a separate calibration step downstream, built for the one decision that actually needed a leaderboard: sizing the bonus pool. The rep who logged the ticket had a five-minute answer ready the next time a prospect asked for the same flip.
greytHR's Performance Management is built on a rating scale designed to stay comparable across cycles and teams, so the number a manager gives an employee this year still means something next year, without turning into a ranking.
References
Ranking versus rating in peer review of research grant applications, PMC / BMC journal article on absolute vs. relative judgment in reviewer scoring.
Vitality curve (“20-70-10” forced ranking), popularized by Jack Welch at General Electric in the 1980s.
Microsoft ends stack ranking, November 2013: Lisa Brummel company memo, reported by the Wall Street Journal, CNNMoney, and Bloomberg.
Cardinaels et al. (2021), Forced Rating Systems from Employee and Supervisor Perspectives, Journal of Accounting Research.
