How to Build a Performance Management System Managers Will Actually Use
Most performance systems are designed for the People team. Design for the manager holding a difficult conversation on a Thursday afternoon and almost everything changes.
There is a reliable way to find out whether a performance system works. Ask a manager to explain, without notes, why one of their people received the rating they did. If the answer starts with a description of the process rather than of the person's work, the system is producing paperwork instead of judgement.
Performance systems fail for two structural reasons, and neither is the framework document itself.
Failure one: nobody defined what is being assessed
A rating scale is a measuring instrument. It cannot create a standard that does not exist. If the organisation has never written down what is expected of someone at this level, in this role, then "exceeds expectations" means whatever each manager privately believes it means.
This is why performance work so often has to start upstream, in role definitions and level expectations. The three inputs a manager needs before any assessment is possible:
- Outcomes — what this person is responsible for delivering, stated so that delivery is verifiable.
- Level expectations — what scope, autonomy and complexity look like at their level, distinct from the level below and above.
- Behaviours — the small number of observable behaviours the company will actually act on, not a list of virtues.
A useful editing rule: if you cannot describe what evidence would demonstrate a competency, remove it. Every unfalsifiable line in a framework becomes a place where managers substitute their own impression, and impressions differ systematically by manager.
Failure two: it was designed for the wrong user
The People team is not the user of a performance system. They are its administrator. The user is a manager with eight people, a delivery commitment, and ninety minutes in the week to do this properly.
Designing for that person produces different decisions than designing for a complete process:
| Designed for the process | Designed for the manager |
|---|---|
| Twelve competencies, each with five proficiency levels | Four to six expectations that distinguish this level from the next |
| A form that collects data for reporting | A structure that helps prepare the conversation |
| Guidance published on the intranet | Guidance inside the template, at the point of use |
| A cycle that assumes uninterrupted preparation time | A cycle where every step has a defined time cost |
The evidence standard is the whole game
Consistency across managers does not come from the scale. It comes from agreeing what counts as evidence before the cycle starts.
Define it explicitly: specific instances rather than general impressions, work from the whole period rather than the last six weeks, outcomes the person could influence, and at least one source beyond the manager's own observation for senior levels. Then hold the standard in calibration, which is the only place it can actually be enforced.
Calibration, sized to your company
Calibration is not a distribution exercise and it should not be run as one. Forcing a curve on a team of nine produces a decision that cannot be explained to the person affected, which is the definition of a bad system.
What calibration is for: making sure "exceeds" means the same thing in two different teams. A workable format for a company under 300 people:
- Managers submit ratings with evidence, before the session.
- Peers within a function calibrate first, in groups of four to six managers.
- Each manager presents only their proposed top and bottom, with evidence.
- The group challenges the evidence, not the manager.
- A facilitator records every change and the reason for it.
- Cross-function calibration happens at the leadership level, on senior levels only.
The recorded reasons matter more than the ratings. They become the precedent that makes next year's cycle faster, and the material a new manager reads to understand the standard.
Decide the link to pay before the first cycle
Whatever you decide, decide it in advance and say it out loud. The failure mode is not "ratings affect pay" or "ratings do not affect pay" — both work. The failure mode is that nobody knows which is true, so everybody assumes the version that suits them.
If performance informs pay, the mapping should be explicit: which rating produces which range of outcomes, how position in the salary band modifies it, and who approves exceptions. If it does not, be clear about what does drive pay, or the rating will be read as a pay signal anyway.
A rating people cannot connect to any consequence is not neutral. It is read as a consequence nobody will explain.
Prepare managers for the conversation, not the form
Most training covers how to complete the assessment. Almost none covers the part managers actually find hard: delivering a message someone does not want to hear, without either softening it into ambiguity or turning it into a formal process.
Practise on real cases with the managers who will run them. Ninety minutes of rehearsal changes more than a well-produced guide, because the difficulty is not comprehension.
A minimum viable first cycle
If you are starting from nothing, do not build the complete system. Build this, run it once, and correct it with what you learn:
- Level expectations for the three or four levels you actually have.
- A simple assessment against outcomes and expectations, with an evidence field that cannot be left empty.
- A three or four point scale with written definitions.
- One calibration round per function.
- A defined conversation, with a template and one rehearsal session.
- A stated position on how this connects to pay.
That is a working performance management system. Everything else is refinement, and refinement is much easier once you have seen where your organisation actually struggles.
Co-founder — People & Management
Ana Reis
Builds the People foundations of growing companies — and stays with them until they work day to day.
More from Ana