Article3 min readBy SmartScale

    Running a calibration meeting in an AI-first team

    Performance calibration breaks when half the output is AI-assisted. Here is the rubric, the evidence standard, the bias controls and the agenda we use to keep ratings fair when leverage — not hours — is the unit of work.

    A calibration meeting exists for one reason: to make sure a rating means the same thing across managers. Manager A's "exceeds" and Manager B's "exceeds" should describe comparable work. Without calibration, ratings measure the manager, not the person.

    AI-assisted work makes this harder, because the traditional proxies quietly stop working.

    What breaks when the team goes AI-first

    • Volume stops being a signal. An assistant can multiply anyone's output. Shipping more no longer distinguishes anyone.
    • Polish stops being a signal. Documents, decks and code all read well now. Quality of prose says nothing about quality of thought.
    • Speed compresses. Tasks that separated strong from average performers now take everyone an afternoon.
    • The real differentiator moves upstream — to problem selection, system design, and whether the work made other people faster.

    If your rubric still rewards throughput, you will calibrate a room full of people who all look excellent.

    The rubric: leverage over output

    Rate on four dimensions, each with observable evidence:

    1. Problem selection. Did they work on what mattered? Evidence: which problems they chose when unassigned, and what they declined.

    2. System building. Did the work leave behind something reusable — a template, a pipeline, a skill, a documented process? Evidence: artifacts other people now use.

    3. Multiplier effect. Did colleagues get faster because of them? Evidence: named people, named workflows, before-and-after.

    4. Judgment under AI. Did they catch what the model got wrong? Evidence: a specific instance where they overruled or verified output, and what it prevented.

    A rating of "exceeds" requires evidence on at least three, including the multiplier effect. That single rule kills most rating inflation.

    The evidence standard

    Enforce one sentence in the room: a claim without an artifact is an opinion.

    Acceptable evidence:

    • A shipped artifact and who reused it
    • A named colleague whose workflow changed, with the before and after
    • A decision with the discarded option and the cost of the chosen one
    • A metric with a baseline

    Not acceptable: "very reliable", "great attitude", "everyone loves working with them", "always available". These describe a feeling about a person, not their work.

    Bias controls that actually change outcomes

    • Read the ratings before the discussion. Distribute the written cases in advance; no verbal pitches first. Charismatic managers otherwise anchor the room.
    • Discuss the evidence before revealing the rating. State what happened, then the proposed rating — not the reverse.
    • Rotate who speaks first. The first case sets the calibration bar for the whole session.
    • Flag recency out loud. If every example comes from the last six weeks, the case is incomplete.
    • Name the visibility gap. Remote, part-time and back-office contributors generate less ambient evidence. Ask explicitly what you would be missing.
    • Have someone own the counter-argument. One person per case argues the lower rating. Not to be adversarial — to force the evidence out.

    A 90-minute agenda

    TimeSegmentOutput
    0-10Restate the rubric and what each rating meansShared bar
    10-20Two anchor cases read aloud, one strong, one solidCalibrated reference points
    20-70Case-by-case: evidence, counter-argument, ratingProvisional ratings
    70-80Distribution review — look for manager and visibility skewsAdjustments
    80-90Development actions and who delivers each messageOwners and dates

    Cap it at eight cases per session. Ratings degrade badly in hour three.

    After the meeting

    Calibration is worthless if the outcome never reaches the person. Within a week, each manager delivers: the rating, the two pieces of evidence that drove it, and one concrete leverage behaviour to develop next quarter.

    That last part is what closes the loop back to hiring. The behaviours you calibrate on should be the same ones you screened for — the six pillars of AI-first talent — and the same ones your onboarding plan set expectations around on day one.

    If the rubric in your calibration room doesn't match the scorecard in your interview loop, one of the two is lying. Our AI-first hiring framework exists to keep them the same document.

    #performance#calibration#ai-first-hiring#leadership

    Want this installed in your hiring process?

    We run a 30-minute intro call to map your roles against the AI-first framework.

    Book a 30-min intro call

    Related reading