Performance calibration: how to run a session that is fair and quick
Calibration is where ratings stop being one manager's opinion and become comparable across teams. Ninety minutes, a fixed agenda, and every rating explainable to the person who receives it.
Performance calibration is a short, structured meeting where managers compare the ratings they propose for their people before anything is communicated. It exists because a rating written by one manager, alone, only describes that manager's standard. Calibration turns a set of private judgements into a set of comparable ones, so that the same words mean the same thing in engineering and in customer support. It is not a meeting to distribute grades. It is a meeting to test evidence.
Run well, it takes ninety minutes for a function. Run badly, it takes a full day and produces decisions nobody can explain afterwards. The difference is almost entirely in the preparation and in the agenda.
What performance calibration is for
A calibration session has three jobs. First, to check that the evidence behind each proposed rating meets the standard the company agreed before the cycle started. Second, to compare people at the same level across teams, which no individual manager can do alone. Third, to record the reason for every rating that changes, so the decision survives the meeting and can be repeated next year.
Anything that is not one of those three belongs somewhere else. Career conversations, salary decisions and succession planning all matter, and all of them will stretch the session past the point where anyone is still thinking clearly.
A ninety-minute agenda
Publish the agenda with the invitation and hold to it. The timings assume a group of four to six managers from the same function, with ratings and evidence submitted at least two days before.
| Minutes | Step | Who speaks | Output |
|---|---|---|---|
| 0–10 | Purpose, scale definitions and the evidence standard, read aloud | Facilitator | Everyone starts from the same definitions |
| 10–25 | The distribution as submitted, by level and by team | Facilitator | A visible pattern: where two teams disagree |
| 25–55 | Outliers first: the highest and lowest proposed ratings | Owning manager, then the group | Each outlier confirmed or changed, with a reason |
| 55–75 | Same-level comparison across teams, one level at a time | The group | Ratings that mean the same thing in every team |
| 75–85 | Changes read back, with the reason for each | Facilitator | A written record the managers agree with |
| 85–90 | Who tells whom, and by when | Facilitator | A named owner and a date for every conversation |
Rules of the room
Say these out loud at the start of the first session of the cycle, and once more at the start of the second. They are what keeps ninety minutes from becoming three hours.
- Evidence over impressions. A manager who cannot point to specific work from across the whole period has not made an assessment yet. "She is very strong" is not a contribution; "she took over the migration in March and closed it without escalating once" is.
- Discuss outliers first. The top and the bottom of each list is where the disagreement lives and where the consequences are largest. The middle of the range rarely needs the group and can be handled in writing.
- No forced distribution unless it was decided in advance. If the company has chosen to constrain the shape of the ratings, everyone must know before writing a single assessment. Introducing it inside the room turns the session into a negotiation about scarcity.
- Record the reason for every change. Not the vote and not who argued hardest: the evidence that moved the rating. That sentence is what the manager will use in the conversation, and it is the precedent for next year.
- Challenge the evidence, not the manager. The question is always "what makes you say that", never "are you being too generous".
Justifying a rating in one sentence
A rating a manager cannot defend in one sentence is a rating the person will not accept. The sentence has a reliable shape: what the person was responsible for, what actually happened, and how that compares with what the level expects. The examples below are written generically, for a four-level scale — the wording changes with each company's framework, the structure should not.
How to justify a performance rating
| Level | Justification from outcomes | Justification from scope | Justification from behaviour |
|---|---|---|---|
| Exceeds expectations | Delivered the two hardest commitments of the year and one that was never planned, without moving the dates on either. | Took on work a level above their own for most of the period and did it without supervision. | Changed how the team works, and the change held after they stopped pushing it. |
| Meets expectations | Delivered what was agreed, at the quality agreed, across the whole period rather than in the final quarter. | Owned the scope described at their level and handled the complications inside it without escalating. | Acted consistently with what the company asks for, including when it cost them something. |
| Partially meets | Delivered most of what was agreed, with one commitment that slipped and was not flagged in time. | Needed help with parts of the role that the level expects to be handled alone. | The behaviour is not in question; the consistency is, and it has been named more than once. |
| Does not meet | The main commitments of the period were not delivered, and the reasons were within their control. | Worked at the scope of the level below for most of the year, with the support already in place. | A specific expectation was raised, in writing, and the pattern did not change. |
Where calibration goes wrong
Calibrating to the budget instead of to performance
This is the failure that does the most damage, because it is invisible to everyone outside the room. The increase pool is fixed, ratings map to increases, and so the ratings are quietly adjusted until the arithmetic closes. What leaves the room is not an assessment; it is a budget wearing an assessment's clothes. Decide the pool separately, and if the money will not stretch, say that to people as a fact about the money rather than encoding it as a judgement about their work.
The loudest manager wins
In an unfacilitated room, ratings drift towards whoever argues hardest, and their teams work it out within two cycles. The correction is structural rather than cultural: a facilitator who is not one of the managers, a fixed speaking order, equal time per outlier, and an explicit instruction that silence is not agreement. Ask the quietest person in the room directly, every time.
Changing a rating without telling the manager why
A rating altered over the manager's head, with no reason attached, destroys two things at once: the manager's ability to hold the conversation, and their willingness to give an honest assessment next time. If a rating changes, the manager who proposed it leaves the room with the sentence they will use. If they cannot say that sentence and mean it, the change has not been made properly.
Common questions
Does calibration require a forced distribution?
No. Forcing a curve on a team of nine produces a decision that cannot be explained to the person affected. Calibration and forced distribution are separate choices: the first makes ratings comparable, the second constrains their shape, usually to control cost. If the company decides to constrain the shape, decide it before the cycle opens, say so publicly, and apply it to a population large enough to be defensible — never inside a single small team.
Who takes part in calibration?
Managers from the same function who assess people at comparable levels, in groups of four to six, plus a facilitator with nobody under assessment in that session. Six is the point at which the ninety minutes stops working. Everyone in the room must have read the evidence beforehand. Senior levels are calibrated once more at leadership level, across functions, because that is the only place those comparisons can be made. The People team facilitates and records; it does not set the ratings.
What do you tell the person afterwards?
The rating, the evidence, and what would change it — in the manager's own words, in a conversation, before it appears in any system. Never present calibration as the author of the decision. "It was changed in calibration" tells a person that their manager did not decide and cannot explain, which is worse than any rating. The manager owns the message, including the parts they argued against.
Calibration is a small meeting carrying a lot of weight: it is the point where an assessment stops being one person's opinion and becomes something the company can stand behind. It only works on top of a cycle with clear level expectations and an agreed evidence standard, which is what our performance management work builds. If you are earlier than that, start with a performance system managers will actually use and with career frameworks for scale-ups, which is where the level definitions come from.

Co-founder — People & Management
Ana Reis
Builds the People foundations of growing companies — and stays with them until they work day to day.
More from Ana