Running a calibration meeting in an AI-first team
Performance calibration breaks when half the output is AI-assisted. Here is the rubric, the evidence standard, the bias controls and the agenda we use to keep ratings fair when leverage — not hours — is the unit of work.
A calibration meeting exists for one reason: to make sure a rating means the same thing across managers. Manager A's "exceeds" and Manager B's "exceeds" should describe comparable work. Without calibration, ratings measure the manager, not the person.
AI-assisted work makes this harder, because the traditional proxies quietly stop working.
What breaks when the team goes AI-first
- Volume stops being a signal. An assistant can multiply anyone's output. Shipping more no longer distinguishes anyone.
- Polish stops being a signal. Documents, decks and code all read well now. Quality of prose says nothing about quality of thought.
- Speed compresses. Tasks that separated strong from average performers now take everyone an afternoon.
- The real differentiator moves upstream — to problem selection, system design, and whether the work made other people faster.
If your rubric still rewards throughput, you will calibrate a room full of people who all look excellent.
The rubric: leverage over output
Rate on four dimensions, each with observable evidence:
1. Problem selection. Did they work on what mattered? Evidence: which problems they chose when unassigned, and what they declined.
2. System building. Did the work leave behind something reusable — a template, a pipeline, a skill, a documented process? Evidence: artifacts other people now use.
3. Multiplier effect. Did colleagues get faster because of them? Evidence: named people, named workflows, before-and-after.
4. Judgment under AI. Did they catch what the model got wrong? Evidence: a specific instance where they overruled or verified output, and what it prevented.
A rating of "exceeds" requires evidence on at least three, including the multiplier effect. That single rule kills most rating inflation.
The evidence standard
Enforce one sentence in the room: a claim without an artifact is an opinion.
Acceptable evidence:
- A shipped artifact and who reused it
- A named colleague whose workflow changed, with the before and after
- A decision with the discarded option and the cost of the chosen one
- A metric with a baseline
Not acceptable: "very reliable", "great attitude", "everyone loves working with them", "always available". These describe a feeling about a person, not their work.
Bias controls that actually change outcomes
- Read the ratings before the discussion. Distribute the written cases in advance; no verbal pitches first. Charismatic managers otherwise anchor the room.
- Discuss the evidence before revealing the rating. State what happened, then the proposed rating — not the reverse.
- Rotate who speaks first. The first case sets the calibration bar for the whole session.
- Flag recency out loud. If every example comes from the last six weeks, the case is incomplete.
- Name the visibility gap. Remote, part-time and back-office contributors generate less ambient evidence. Ask explicitly what you would be missing.
- Have someone own the counter-argument. One person per case argues the lower rating. Not to be adversarial — to force the evidence out.
A 90-minute agenda
| Time | Segment | Output |
|---|---|---|
| 0-10 | Restate the rubric and what each rating means | Shared bar |
| 10-20 | Two anchor cases read aloud, one strong, one solid | Calibrated reference points |
| 20-70 | Case-by-case: evidence, counter-argument, rating | Provisional ratings |
| 70-80 | Distribution review — look for manager and visibility skews | Adjustments |
| 80-90 | Development actions and who delivers each message | Owners and dates |
Cap it at eight cases per session. Ratings degrade badly in hour three.
After the meeting
Calibration is worthless if the outcome never reaches the person. Within a week, each manager delivers: the rating, the two pieces of evidence that drove it, and one concrete leverage behaviour to develop next quarter.
That last part is what closes the loop back to hiring. The behaviours you calibrate on should be the same ones you screened for — the six pillars of AI-first talent — and the same ones your onboarding plan set expectations around on day one.
If the rubric in your calibration room doesn't match the scorecard in your interview loop, one of the two is lying. Our AI-first hiring framework exists to keep them the same document.
Want this installed in your hiring process?
We run a 30-minute intro call to map your roles against the AI-first framework.
Book a 30-min intro callRelated reading
AI Engineering: How to Land Your Dream Job (Using the Skills Map)
A practical playbook for landing an AI engineering job — built on Andrew Ng's four-skill map, with the portfolio, interview answers and signals hiring teams actually reward.
AI-First vs. AI-Native vs. Digital-First: What's the Difference?
AI-first, AI-native and digital-first are often used interchangeably — they are not. Clear definitions of each term, and why the difference matters for hiring.
What Is AI-First Talent? The Complete Definition
AI-first talent is a subject matter expert who multiplies their domain expertise with AI. The complete definition, what it is not, and how to hire for it.