BUSFACTOR.TECH
Team Health

Performance Reviews for Engineers, Minus Stack Ranking

Signals, not ranks

REVIEWS WITHOUT A CURVE

How to run engineering performance reviews without a forced curve: evidence-linked scorecards, system context first, and calibration on receipts.

2 receipts in this article ↓

TL;DR: Most engineering performance reviews fail the same three ways: they measure activity instead of impact, they judge people for what the system did to them, and they collapse a human into one sortable number. The fix isn't abolishing reviews; engineers deserve to know how they're doing. The fix is rebuilding them as evidence-linked scorecards: strengths, growth areas, unique contributions, receipts. Review people against expectations, never against each other. The curve stays dead.

Microsoft ended forced ranking in 2013 with three words - "No more curve" - after years of commentary tied it to a collaboration-killing culture. Why stack ranking backfires is settled. What's less settled is what a good review looks like instead, because "we don't rank" too often decays into "we don't say anything useful," and engineers end up with annual vibes, a surprise rating, and no idea what would change it.

A performance review is worth running. It just has to answer the right question. The wrong question is "where does this person rank?" The right one is "is this person growing, supported, and recognized for what they actually carry?"

Rule one: system context before personal conclusions

Before any sentence about a person, audit the system they worked in. The engineer who "shipped nothing" for six weeks may have spent them wedged behind a review queue nobody watches, an unowned dependency, and two production fires. Write the review without that context and you're grading someone on the weather.

This is the review-season version of a daily truth: most individual-performance stories are system stories on inspection. Check flow, queues, and interruptions first. What's left after the system explains its share - that's the part that belongs in a conversation about a person.

Rule two: receipts, not vibes

The failure mode of curve-free reviews is adjective soup: "solid quarter," "could be more visible," "great team player." Unfalsifiable praise is as useless as unfalsifiable criticism. Every claim in a review should carry a receipt, a linkable and specific example:

  • The design that held. The migration plan that survived contact with production.
  • The review that mattered. Comments that caught a real defect, not nitpicks.
  • The person they multiplied. The hire who ramped in half the expected time because someone built them a runway.
  • The fire that didn't spread. Incident handling, and what changed afterward.

What does not count as a receipt: commit counts, lines of code, tickets closed, story points burned. The SPACE research is blunt about this: productivity can't be captured by a single metric, and raw activity is the weakest, most gameable dimension available. Activity metrics don't just miss glue work, mentoring, and review load. They punish them.

The code-review matrix: a heatmap of who reviews whom, with the reviewers carrying the heaviest load standing out.The code-review matrix: a heatmap of who reviews whom, with the reviewers carrying the heaviest load standing out.
The review matrix - who carries the loadLive product · fictional demo org

The scorecard structure

Replace the rating-plus-paragraph format with a multi-dimensional scorecard. Four sections, each with evidence:

  1. Strengths. What this person is genuinely excellent at, with the receipts above. Praise is half the review: specific praise, the kind that tells someone what to do more of.
  2. Growth areas. Two, at most three, each with a concrete example and a concrete next step. "Improve communication" is a horoscope; "design docs land without an explicit rollback section - here are two examples, let's fix the template together" is coaching.
  3. Unique contributions. What this person carries that no metric shows: the domain only they understand, the coordination they quietly absorb, the on-call calm. This section is also your key-person-risk early-warning system. If it's long, you have a resilience problem to fix, not a hero to burn out.
  4. Impact, in outcomes. What changed for users, for the team, for the system, not what activity occurred.

Two hard rules. No single composite score. The moment the dimensions collapse into one sortable number, you've rebuilt a leaderboard with better branding. No surprises. If anything in the review is news, the review is late; the document should be a summary of conversations already had.

Calibration without a curve

Cross-team fairness reviews are healthy. Quotas of failure are not. Run calibration as a receipts exercise: managers bring their evidence-linked scorecards, peers pressure-test the examples ("is that senior-level impact, or solid mid-level work?"), and expectations get aligned across teams. What nobody brings is a target distribution. If two teams both performed excellently, both reviews say so - the thing Microsoft's curve literally forbade.

And handle underperformance the boring, fair way: written expectations, specific examples of the gap, an honest system check, a support plan, time to improve. That process doesn't need the other nine people ranked against each other. It never did.

The organization overview: a health index dial with the six sub-scores behind it and the top findings underneath.The organization overview: a health index dial with the six sub-scores behind it and the top findings underneath.
The overview - the whole org in one dialLive product · fictional demo org

Make it continuous, not annual archaeology

An annual review written from memory is a recency-bias generator with a signature line. Instead: keep a running evidence file per person (managers and engineers should both keep one), make regular survey signal part of the picture so you hear about system problems before they wear a person's face, and treat the formal review as the quarterly summary of an ongoing conversation.

If you're a new manager inheriting a team, this is the ritual to rebuild first, because it's the one where trust is won or torched. Run it as ranking, and your best people start reading the leaderboard and updating their CVs. Run it as evidence-linked support, and the review becomes the rarest thing in engineering management: a meeting people actually want on the calendar.

Frequently asked

How should software engineer performance reviews work without stack ranking?

Review each person against expectations for their role and level - never against each other. Build the review from evidence-linked examples collected all cycle, check system context before personal conclusions, and structure it as a scorecard: strengths, growth areas, unique contributions, and impact with receipts. Calibrate across teams on concrete examples, not distribution quotas.

What evidence belongs in an engineering performance review?

Specific, linkable examples: the design that survived contact with production, the review comments that caught real issues, the incident that got handled, the teammate who ramped faster because of them. What doesn't belong: raw activity counts like commits, lines of code, or tickets closed - those measure motion, are trivially gameable, and systematically miss review work, mentoring, and glue work.

Why shouldn't reviews use a single overall score?

Because a single sortable number turns a support tool into a leaderboard. The moment every engineer collapses to one figure, people compare, rank, and compete - and the collaborative work engineering depends on becomes a donation to a rival. Keep the dimensions separate: a person can be strong on delivery, growing on communication, and irreplaceable on domain knowledge all at once. That's information; an average of it is noise.

How do you handle genuine underperformance without a ranking system?

The boring, fair way: clear written expectations, specific evidence-linked examples of the gap, an honest check that the system isn't the real cause (blocked dependencies, review queues, unowned areas), and a concrete support plan with time to improve. None of that requires ranking teammates against each other - and all of it works better without the fear tax.

Receipts

Keep reading