Skip to content

Performance management framework managers will actually use (with a template)

A performance management framework is only as good as the conversation it produces. Here is the whole process, with a review template, a rating scale and twelve example comments.

Ana ReisUpdated 8 September 20268 min read

A performance management framework is the small set of written rules a company uses to judge how well someone is doing their job: what each level is expected to deliver, what counts as evidence, and how the two are turned into a rating. It works when a manager can explain a rating without opening the form. It fails when the answer describes the process rather than the person's work.

This is the whole performance management process, from level expectations to the conversation itself, with a review template, a rating scale and example comments you can adapt. Start with the two structural reasons these systems fail, because neither of them is the framework document.

Failure one: nobody defined what is being assessed

A rating scale is a measuring instrument. It cannot create a standard that does not exist. If the organisation has never written down what is expected of someone at this level, in this role, then "exceeds expectations" means whatever each manager privately believes it means.

This is why performance work so often has to start upstream, in role definitions and level expectations. The three inputs a manager needs before any assessment is possible:

  • Outcomes — what this person is responsible for delivering, stated so that delivery is verifiable.
  • Level expectations — what scope, autonomy and complexity look like at their level, distinct from the level below and above.
  • Behaviours — the small number of observable behaviours the company will actually act on, not a list of virtues.

A useful editing rule: if you cannot describe what evidence would demonstrate a competency, remove it. Every unfalsifiable line in a framework becomes a place where managers substitute their own impression, and impressions differ systematically by manager.

Failure two: it was designed for the wrong user

The People team is not the user of a performance system. They are its administrator. The user is a manager with eight people, a delivery commitment, and ninety minutes in the week to do this properly.

Designing for that person produces different decisions than designing for a complete process:

Designed for the processDesigned for the manager
Twelve competencies, each with five proficiency levelsFour to six expectations that distinguish this level from the next
A form that collects data for reportingA structure that helps prepare the conversation
Guidance published on the intranetGuidance inside the template, at the point of use
A cycle that assumes uninterrupted preparation timeA cycle where every step has a defined time cost

A performance review template

The template below is the entire assessment: six sections, each with a rule about what goes in it. The rule is what makes two different managers produce comparable documents.

SectionWhat goes in itExample line
Outcomes for the periodThe three to five things the person owned, each with what actually happenedOwned the billing service migration; delivered inside the quarter, with two weeks of slippage the team absorbed
EvidenceSpecific, dated instances from the whole period — not impressionsMarch: rewrote the incident runbook after the February outage; two other teams have used it since
Level expectationsHow the work compares with what the level asks for, expectation by expectationWorks independently on scoped problems: met. Frames the problem before someone else does: not yet consistent
BehavioursThe two or three behaviours the company acts on, each with one observed exampleRaised the risk to the delivery date early enough for the plan to change, in front of the client
Rating and justificationThe rating, then the sentence that would survive being read aloud to the personAt the level: delivered the level in full, and the scope of the work matched the level rather than exceeding it
What changes next periodOne development priority, and one thing the manager will do differentlyTakes the design review on the next project; the manager hands over the client relationship for it

One field should be impossible to leave empty: evidence. Everything else can be corrected in calibration. A blank evidence field cannot.

A four-level rating scale

Four levels, each with a written definition and an observable signal. Three works too. Five is where the middle two stop meaning anything different.

LevelDefinitionObservable signal
Below the levelThe work has not met what the level asks for, and the gap is specific rather than generalWork at this level is regularly reviewed or redone by someone else
Approaching the levelMost of the level is met; one or two expectations are not yet consistentStrong once the work is scoped, still needs help framing ambiguous work
At the levelThe level is met in full, across the whole period, without a caveatThe manager hands over work at this level and stops thinking about it
Above the levelThe work matched the level above, in scope and in autonomy, more than onceOther teams route problems of this type to this person before they reach the manager

Note what the top level is not: working hard, being well liked, or having had a good quarter. It is doing the next level's job.

Example performance review comments

Writing the comment is where most managers stall. Twelve to adapt, three per level. Each one names the work, the effect it had, and what happens next.

  1. Below the level. The last two projects needed a colleague's second pass before shipping. The pattern is in the testing, not the design.
  2. Below the level. Planning commitments moved three times without being flagged. The cost landed on the teams downstream, who found out late.
  3. Below the level. The technical work is sound; decisions get written down after the fact. We agreed a format and a check-in in six weeks.
  4. Approaching the level. Delivers reliably once the work is scoped. Not yet met: framing the problem before someone else does.
  5. Approaching the level. The second half of the period met the level and the first half did not. The rating covers both.
  6. Approaching the level. Strong on execution, hesitant on escalation. Two of this period's delays were visible weeks before being raised.
  7. At the level. Owned the busiest area of the roadmap and delivered it without the manager holding the plan.
  8. At the level. Handled a difficult client conversation without escalating it, and wrote up the outcome for the next person.
  9. At the level. Consistent across the whole period, including the quarter the team spent short-handed. Nothing rests on the last six weeks.
  10. Above the level. Took the ambiguous half of the migration, defined the scope, and ran it with two people alongside.
  11. Above the level. Other teams now bring problems of this type here before they reach the manager. That changed during the period.
  12. Above the level. Rewrote how the team estimates work, which changed decisions outside their own project.

How to justify a top rating

The top rating is the one most likely to be challenged, and the one most often written badly. Three sentences, in this order, survive the room:

  1. Evidence. The specific, dated work the rating rests on. Two instances rather than one, and not both from the same month.
  2. Impact. What changed because of it, beyond the person's own delivery — for another team, for a client, or for a decision the company took.
  3. Comparison with the expected level. The sentence explaining why this is the level above rather than a good year at this level. If that sentence cannot be written, the rating is one level too high.

The same method is what makes calibration short. A rating with those three sentences attached is either agreed quickly or corrected quickly.

The evidence standard is the whole game

Consistency across managers comes from agreeing what counts as evidence before the cycle starts, not from the scale. Define it explicitly: specific instances rather than impressions, work from the whole period rather than the last six weeks, outcomes the person could influence, and at senior levels at least one source beyond the manager's own observation.

Calibration, sized to your company

Calibration is not a distribution exercise: forcing a curve on a team of nine produces a decision nobody can explain to the person affected. Its job is narrower and more useful — making sure the top rating means the same thing in two different teams. A workable format below 300 people:

  1. Managers submit ratings with evidence, before the session.
  2. Peers within a function calibrate first, in groups of four to six managers.
  3. Each manager presents only their proposed top and bottom, with evidence.
  4. The group challenges the evidence, not the manager.
  5. A facilitator records every change and the reason for it.
  6. Cross-function calibration happens at the leadership level, on senior levels only.

The recorded reasons matter more than the ratings: they are the precedent a new manager reads to understand the standard.

Whatever you decide, decide it in advance and say it out loud. The failure mode is not "ratings affect pay" or "ratings do not affect pay" — both work. The failure mode is that nobody knows which is true, so everybody assumes the version that suits them.

If performance informs pay, make the mapping explicit: which rating produces which range of outcomes, and who approves exceptions. If it does not, say what does drive pay — otherwise the rating gets read as a pay signal anyway.

A rating people cannot connect to any consequence is not neutral. It is read as a consequence nobody will explain.

Prepare managers for the conversation, not the form

Most training covers how to complete the assessment. Almost none covers the hard part: delivering a message someone does not want to hear without softening it into ambiguity. Rehearse it on real cases; the difficulty is not comprehension.

A minimum viable first cycle

If you are starting from nothing, do not build the complete system. Build this, run it once, and correct it with what you learn:

  • Level expectations for the three or four levels you actually have.
  • The review template above, with an evidence field that cannot be left empty.
  • A three or four point scale with written definitions.
  • One calibration round per function.
  • A defined conversation, with a template and one rehearsal session.
  • A stated position on how this connects to pay.

That is a working performance management framework. Everything else is refinement, and refinement is much easier once you have seen where your organisation actually struggles. If you would rather not start from a blank page, that is the work we do in the Performance Cycle.

Ana Reis

Co-founder — People & Management

Ana Reis

Builds the People foundations of growing companies — and stays with them until they work day to day.

More from Ana

Related reading

Organisation Design & Scaling5 min read

The People Operating System: what sits beyond HR

HR delivers processes. An operating system connects them. The difference shows up the first time a performance rating has to justify a pay decision.

Ana Reis

Your company has changed. Has your organisation caught up?

The People Diagnostic establishes where the organisation is constrained, and what to fix first.