Busy Isn't the Same as Good: How Engineering Reviews Got Broken
Here's a scenario that'll feel familiar to anyone who's spent time inside a mid-size engineering org.
It's Q4. Review season is approaching. Two engineers are up for promotion consideration. The first has spent the last six months quietly refactoring a core module — cleaning up debt, improving test coverage, reducing the surface area for bugs. The work is invisible to everyone except people who understand the codebase well. The second has shipped five features, attended every cross-functional meeting, and made a lot of noise in Slack. Their PRs are frequent, their names appear in standup notes constantly, and they have a tidy list of things they can point to.
Guess who gets promoted.
If you've watched this play out before, you already know the answer. And if you've been the first engineer in that story, you probably also know the feeling that follows: a slow, quiet decision to stop doing the invisible work.
The Metrics We Reach For (and Why They Fail)
Performance measurement in engineering is genuinely hard. Unlike sales, where you can count closed deals, or support, where you can track ticket resolution, engineering output is tangled, interdependent, and often deliberately delayed in its value. The things that matter most are frequently the hardest to quantify.
So teams reach for proxies. Commit frequency. PR count. Lines of code changed. Story points closed. Tickets moved across a board. These things are measurable, which makes them feel rigorous. They're not. They're just easy.
The problem with optimizing for proxies is that smart people figure out how to optimize for the proxy instead of the underlying thing. When commit frequency is visible and valued, engineers start making smaller, more frequent commits — not because it's the right approach for the work, but because it signals activity. When PR count matters, reviews get rushed. When story points drive performance conversations, estimation becomes a game.
This is Goodhart's Law doing what it always does: the moment a measure becomes a target, it stops being a useful measure.
The Refactoring Theater Problem
One of the more insidious manifestations of this dynamic is what I'd call refactoring theater — the performance of technical improvement without the substance of it.
It looks like this: a developer notices that a piece of code could be cleaner. Instead of making a targeted, purposeful improvement tied to an actual business need, they embark on a sprawling refactor that touches dozens of files, generates a massive PR, consumes weeks of review time, and ultimately produces... code that's marginally different. Maybe slightly cleaner. Maybe not. Hard to tell.
But the PR is visible. The commits are there. The activity is documented. And in a review culture that rewards the appearance of craft over its actual practice, that refactor looks like serious engineering work.
Meanwhile, the engineer who spent the same two weeks carefully scoping a performance improvement that cut infrastructure costs by 20% — but did it in a small, focused change — has less to show on paper. Less noise. Less visible industry.
Which one gets called out in the all-hands? You already know.
What This Does to Morale (and Retention)
The damage here isn't just strategic — it's deeply personal. Engineers who care about their craft are acutely sensitive to whether the environment they're in actually values craft. When the reward systems consistently favor visibility over substance, the message received is clear: the work that matters to you doesn't matter here.
That realization doesn't usually trigger an immediate resignation. It triggers something slower and more corrosive — a gradual withdrawal of discretionary effort. The engineer stops proposing the systemic fix and just patches the symptom. They stop pushing back on the architectural shortcut because they've learned that pushback doesn't get rewarded. They start optimizing for their own metrics rather than the team's outcomes.
You don't lose them all at once. You lose them in increments, until one day they accept a recruiter's LinkedIn message they would have ignored six months ago.
What Actually Worth Measuring Looks Like
None of this means engineering performance is unmeasurable. It means we need to be more honest about what we're actually trying to measure.
A few principles that hold up better in practice:
Measure outcomes, not outputs. Did the feature ship? Did it work? Did users actually use it? Did the system become more reliable after the refactor? Outputs — PRs, commits, tickets — are inputs to outcomes. Don't confuse the two.
Measure lead time over activity. How long does it take for an idea to get from conception to production? That number captures team health, process friction, and individual contribution in a way that commit counts never will. A team with a short lead time is probably doing a lot of things right. A team with a long lead time has problems that no amount of Jira velocity will diagnose.
Make the invisible visible explicitly. If mentorship, documentation, code review quality, and architectural thinking matter to your organization — and they should — then create explicit space for them in how you evaluate people. If you don't name them, you're telling your team they don't count.
Talk about impact in the past tense. In review conversations, ask engineers to describe the actual effect of their work on the product, the team, or the codebase. Not what they did — what changed because of it. That reframe alone surfaces a different kind of thinking.
The Harder Conversation
Fixing this isn't primarily a metrics problem. It's a management problem. And it requires managers to do something genuinely uncomfortable: develop the technical judgment to evaluate engineering work on its actual merits, not just its surface area.
That's a high bar. It requires staying close enough to the work to know the difference between a clever refactor and a meaningful one. It requires resisting the organizational pressure to reward what's easy to see. It requires advocating for the engineer doing the quiet, important work even when they don't have a highlight reel to present.
But here's the thing: the teams that figure this out don't just retain better engineers. They build better software. Because when your incentives are aligned with actual quality — not the performance of quality — people stop optimizing for the review and start optimizing for the work.
That's not a small thing. That's the whole game.