
Why Do Two Managers Rate 'Meets Expectations' Differently? Calibration in Reviews
Two managers rate the same performance differently because each carries a different internal ruler, not because either is acting in bad faith. Performance appraisal calibration is a structured session that lines managers' ratings up side by side to align that ruler, so it separates the signal (real performance) from the noise (the rater's own tendency). The scale of the problem is uncomfortable. In the landmark study by Scullen, Mount and Goff, published in the Journal of Applied Psychology in 2000, a rater's idiosyncratic tendencies explained about 62% and 53% of the variance in ratings across two datasets, while the employee's actual performance explained only 21% and 25%. Put plainly: the rating tells you more about the manager than about the employee. This piece is for CEOs, CHROs and group HR leaders. It explains why this happens, how calibration turns fairness from a feeling into a distribution you can see, and how to make every score change defensible if a Saudi labor dispute ever reopens it.
Key takeaways
Performance appraisal calibration lines managers' ratings up together to align judgment standards, not to impose quotas on teams.
More than half of the variance in a rating reflects the rater, not the ratee (Scullen, Mount & Goff, 2000).
Fairness is not an impression. It is a visible distribution that separates a genuinely strong team from a lenient manager.
A strict forced curve imposes ratios; sound calibration guides the distribution and leaves the decision to people.
In Saudi Arabia, any later performance action (a PIP or a termination) only survives if it rests on a documented reason, an approval chain, and an audit trail.
Why do two managers rate the same employee differently?
Picture two employees who perform at the same level in the same role. One gets "exceeds expectations" and the other gets "meets expectations." The gap is usually not in the work. It's in who held the pen. Each manager reads the same rating scale their own way. A lenient one lifts everybody. A strict one lowers everybody. A third bunches everyone in the middle to avoid awkward conversations. A fourth lets one strong impression color every criterion.
These are not moral accusations. They are well documented patterns in industrial psychology, and they have names. Here they are, so you can name what you see in the next review cycle:
Rater tendency | What it does to the score | Effect on fairness |
Leniency or strictness | Lifts or lowers everyone by a fixed amount | The same performance earns two different scores by manager |
Central tendency | Crowds everyone into "meets expectations" | The gap between a star and a straggler disappears |
Halo effect | One impression colors every criterion | Rating accuracy across distinct competencies is lost |
How much does this actually weigh? The number that should worry every CEO comes from Scullen, Mount and Goff (2000), one of the most thorough studies ever run on what a performance rating really measures. The researchers analyzed ratings of roughly 4,492 managers across two datasets, each rated from seven sources spanning bosses, peers and subordinates.

The finding: idiosyncratic rater effects explained 62% and 53% of the variance, against 21% and 25% for the employee's actual performance. In fairness, these were developmental 360 ratings rather than pay decisions, so the number doesn't mean 62% of your annual review is bias. But the mechanism, a rater's ruler overwhelming the ratee's performance, is exactly what calibration exists to correct.
This is where Solvait Wise comes in. Its calibration session lines up managers' ratings for a single cycle side by side, then applies a MAX_DEVIATION rule that caps how far one manager's ratings can drift from the rest. The personal ruler shrinks, and the score moves closer to the performance and further from the mood of whoever wrote it.
Is fairness a feeling, or a distribution you can see?
Fairness is not a feeling a manager settles into. It is a distribution you look at to know where each team stands. Without seeing ratings distributed across teams, you can't tell a genuinely strong team from a lenient manager who lifted everyone. When everything inflates, the whole company "exceeds." When central tendency takes over, the whole company "meets." Either way, the grade stops meaning anything.
The numbers back this up. Gartner's 2021 research found that only 18% of employees work in a high-fairness environment, and that those who do perform 26% higher and are 27% less likely to leave. Fairness here is procedural: how the score is made, not just what it turns out to be.
The table below shows how the two common tendencies distort the distribution, and how calibration restores its meaning:
Distribution state | What the report shows | What it actually means |
Rating inflation | Most people "exceed expectations" | You can't tell a star from an average performer |
Central tendency | Most people "meet expectations" | The signal disappears entirely |
Calibrated distribution | Clear, justified gaps between teams | The score carries information you can discuss |
A leadership point for every CHRO or group HR head: there is a real difference between guiding grades and imposing a strict forced curve. A forced curve sets quotas per grade in advance, penalizes small teams, and turns the review into a numbers game. That trap is toxic, and McKinsey's 2018 research noted that removing forced ranking was one of the most common changes companies made to their performance systems. The goal is consistency, not quotas.

In Solvait Wise, that philosophy becomes configurable rules. A BAND_DISTRIBUTION rule shows how ratings spread across bands, and a DEPT_MEAN_GAP rule surfaces the gap between each department's mean and the company mean, so a lenient or strict team stands out. Both run in one of two modes: GUIDELINE only warns, while ENFORCED halts until review. And because quotas are meaningless on very small teams, a minParticipants rule shields them from mandatory distributions. In short, a tool that guides rather than corners.
What makes a score change defensible?
The most dangerous kind of calibration is the kind an employee feels as "politics." When the message lands as "HR changed my score" with no written reason, trust in the whole cycle collapses. What turns calibration from a top down verdict into a fair procedure is three things for every adjustment: a documented justification, a clear approval chain, and a complete audit trail.
In the Saudi context this is not administrative nicety. It is governance. Any later performance action that leans on the score, from a performance improvement plan to a termination, can be reopened before the labor courts. Saudi Labor Law (Royal Decree M/51), administered by the Ministry of Human Resources and Social Development, sets a clear ceiling. Article 80 limits dismissal without award to specific grounds, and each ground must be proven with documentation and a documented disciplinary process that gives the employee a chance to respond. Where a dismissal lacks valid cause, compensation follows Article 77. The working rule anyone who has been through a labor dispute knows: the employer carries the burden of proof, and weak documentation is the leading reason employers lose these cases.

This is where Solvait Wise settles the difference. Every calibrated rating carries a written justification: the AI suggests the draft, and a human reviews and approves it. The score then passes through an approval chain and a re-confirmation step, and every change is captured in a full activity log. The result is a record that holds up if the discussion reopens later. Which is why we describe it precisely: calibration in Wise is AI assisted, not run by an autonomous agent. The AI drafts; the decision stays with your team.
How Solvait Wise brings the three ideas together
Calibration is not one feature. It's a chain that starts with a structured appraisal cycle and ends with a defensible record. Solvait Wise runs appraisal cycles with defined criteria, clear approver identity and an endorsement flow, on top of Microsoft Dynamics 365. The three rules work together: MAX_DEVIATION limits how far one manager drifts, BAND_DISTRIBUTION with DEPT_MEAN_GAP makes the distribution visible and discussable, and justification plus approval plus audit trail makes every adjustment defensible. Where the bigger decision needs broader compliance across the whole HR platform, Solvait AI HR provides the regulatory backbone around it.
If you're writing appraisal goals from scratch before the cycle, try the free performance appraisal goal generator in Arabic and English.
See how calibration and approvals work in Solvait Wise.
Want to see calibration run on your own team's data? Book a demo.
FAQ
What is performance appraisal calibration?
Performance appraisal calibration is a structured session where managers and HR review a single cycle's ratings together, so the same standard applies regardless of who the manager is. The goal is to align what "exceeds," "meets" and "needs improvement" mean across teams, moving the score closer to real performance and away from rater tendency.
Why do two managers rate the same employee differently?
Because each manager has a different internal ruler: leniency or strictness, central tendency, or a halo effect from one impression. Scullen, Mount and Goff (2000) found rater tendencies explained more than half the variance in ratings, against roughly a fifth for actual performance. Calibration is what narrows that gap.
Is calibration just a forced curve?
No. A forced curve imposes mandatory ratios per grade and penalizes small teams. Sound calibration guides the distribution, makes it visible and discussable, shields very small teams, and leaves the final decision to people. The aim is consistent standards, not imposed quotas.
How do I make a score change defensible in Saudi Arabia?
Tie every change to a written justification, an approval chain and an audit trail. Any later action such as a PIP or termination may be reviewed under the Labor Law (Articles 80 and 77), and the employer carries the burden of proof. A complete record is what protects the decision if it reopens.
Is Solvait Wise agentic?
No. Solvait Wise is AI-assisted. In calibration, the AI suggests a draft justification for an adjustment, while approval and the decision stay with your team. The only agentic product in the Solvait suite is Solvait AI HR.
References
Scullen, S. E., Mount, M. K., & Goff, M. Understanding the latent structure of job performance ratings, Journal of Applied Psychology, 2000 (the 62%/53% vs 21%/25% figures).
Gartner. HR Research Reveals 82% of Employees Report Working Environment Lacks Fairness, 2021 (18% high-fairness, +26% performance, 27% lower attrition).
Gartner. Only 30% of Managers Who Participate in Talent Reviews Believe Their Leadership Bench is Strong, 2024 (calibration time per employee).
McKinsey & Company. Harnessing the power of performance management, 2018 (procedural fairness and removing forced ranking).
Ministry of Human Resources and Social Development: Saudi Labor Law (Royal Decree M/51), Articles 77 and 80, Saudi Arabia.
Ready to see Solvait in action?
Book a personalized demo and see how Solvait's AI-powered HR platform can transform the way your team works.
Tags
Related Posts

AI CV Screening: How It Ranks Arabic and English CVs
How AI CV screening reads Arabic and English CVs, scores each applicant with strengths and gaps, and what Saudi recruiters should still decide themselves.
Sep 17, 2026

The Real Cost of Manual HR in a Saudi Enterprise
Manual HR quietly costs a Saudi enterprise six figures a year. See the EY per-task numbers, the payroll-error math, and what automation actually recovers.
Sep 6, 2026

Free Performance Appraisal Goal Generator | Solvait
Generate professional performance appraisal goals for any role in under a minute. Free tool, SMART framed, editable, built for Saudi HR teams.
Aug 27, 2026

