Evans Learning Labs
LinkedIn Email
Cornerstone Research Monograph

What Gets Measured Gets Gamed

Measurement systems do not neutrally capture performance. They shape it. The behavioral consequences of how organizations measure work, and what they choose not to measure, determine more about what employees actually do than any stated priority or leader exhortation.
Abstract
The relationship between organizational measurement and employee motivation has been characterized by a fundamental tension identified in management research since Kerr's (1975) foundational analysis: the measurement practices that produce the most tractable performance data are frequently those that most powerfully distort the behavior they are designed to capture. This tension is structural: the outcomes organizations most value resist quantification, while the outcomes most easily measured are frequently proxies systematically optimized at the expense of the underlying capabilities they represent. Deci, Koestner, and Ryan's (1999) meta-analysis found that extrinsic rewards undermine intrinsic motivation, a finding this monograph engages alongside Cameron and Pierce's (1994) competing meta-analysis, which found the effect considerably smaller and more conditional than commonly portrayed. This monograph reviews the evidence on how measurement influences motivation, examines the measurement-behavior connection in high-stakes contexts, considers the self-determination theory account of how extrinsic measurement can undermine intrinsic motivation and the genuine academic debate over how large and generalizable that effect actually is, and addresses the design principles most associated with measurement systems that produce the behaviors organizations genuinely want rather than sophisticated proxies for those behaviors.

The Kerr Problem and Its Structural Persistence

Kerr (1975) documented with specificity a phenomenon that any experienced organizational practitioner recognizes but that most organizational design processes fail to prevent: organizations routinely articulate priorities and then design measurement and reward systems that ignore or actively punish the behaviors those priorities require. His examples remain clear in their illustration: organizations that value teamwork but measure and reward only individual performance; those that espouse long-term thinking but reward only quarterly results; those that claim to value innovation but measure and reward only successful execution of existing products. Each measurement-aspiration misalignment produces predictable behavioral outcomes: employees rationally optimize for the measured and rewarded behaviors, producing the outcomes the measurement system incentivizes rather than the outcomes the organization genuinely wants but did not design its measurement system to produce.

The persistence of the Kerr problem across decades of organizational research and practitioner awareness reflects the structural difficulty of measuring what organizations most value. The organizational outcomes most associated with long-term competitive advantage, including the development of genuine expertise, the quality of client relationships, the accumulation of institutional knowledge, and the cultivation of organizational culture, are precisely those that resist reliable quantification. The outcomes most easily measured, including quarterly revenue, project completion rates, call volume, and activity counts, are frequently proxies for performance rather than performance itself, and the proxies are systematically optimized at the expense of the underlying capabilities they were designed to represent. The closer the measurement is to the surface of performance, the more easily it can be produced without the underlying performance it was designed to measure, which is why measurement systems concentrated on observable activity metrics consistently produce gaming rather than the performance improvement they were intended to enable.

Goodhart's Law, described by Strathern (1997) as applying the principle that when a measure becomes a target, it ceases to be a good measure, captures the behavioral dynamic that organizational measurement consistently produces when measurement is poorly designed. Once employees understand that their organizational standing, compensation, or advancement depends on a specific measured indicator, they direct effort toward producing that indicator rather than toward the broader performance it was intended to represent. Employees who optimize measured proxies rather than underlying performance are not primarily acting from bad intentions; they are acting rationally within the incentive architecture the organization has created. The organizational design problem is not individual employee character but the measurement system that makes proxy optimization the individually rational strategy, and the solution is therefore measurement system redesign rather than management of individual employee behavior.

Campbell and Fiske (1959) established validity requirements for measurement systems that apply directly to organizational performance measurement: a valid measurement system should capture multiple dimensions of performance through multiple measurement methods, should demonstrate convergent validity by producing consistent findings about the same underlying performance through different methods, and should demonstrate discriminant validity by distinguishing meaningfully between high and low performers on the dimensions that actually affect organizational outcomes. Most organizational measurement systems fail all three criteria: they measure narrow dimensions through a single method with insufficient discrimination between the performance levels that most affect organizational outcomes. The result is a measurement architecture that is internally consistent and administratively tractable but that produces systematically distorted performance information and correspondingly distorted behavioral optimization.

How Extrinsic Measurement Undermines Intrinsic Motivation

Reward type x autonomy: motivation quality matrix
Low Task Autonomy High Task Autonomy
High Development Orientation Controlled motivation: compliance without discretionary effort or creative investment Highest performance: intrinsic motivation amplified by supportive developmental context
High Extrinsic Reward Salience Lowest engagement: neither autonomy nor reward salience; minimum viable contribution Transactional performance: extrinsic reward sufficient; measurable outputs deliverable
Figure 1. Reward type and task autonomy interact to determine motivation quality and behavioral output. Extrinsic rewards dominate in low-autonomy, high-routinization contexts. Intrinsic motivation predicts performance quality and discretionary effort most powerfully in high-autonomy, judgment-intensive roles.
Deci, Koestner and Ryan, 1999; Ryan and Deci, 2000

Ryan and Deci (2000) developed self-determination theory as an account of the conditions under which human motivation is most and least productive, distinguishing intrinsic motivation, engaging in an activity for its inherent interest and satisfaction, from extrinsic motivation, engaging in an activity because of its instrumental consequences including rewards and punishments. Their research established that these are not simply different points on a single motivation continuum but qualitatively different motivational orientations with distinct behavioral, cognitive, and performance consequences. Intrinsically motivated activity produces more creative performance, greater depth of engagement, more persistent effort in the face of difficulty, and higher-quality outcomes in judgment-intensive tasks than extrinsically motivated activity producing equivalent levels of observable effort.

Deci, Koestner, and Ryan (1999) conducted a meta-analysis of 128 studies examining the effects of extrinsic rewards on intrinsic motivation and found that tangible, contingent, expected extrinsic rewards significantly undermined intrinsic motivation for activities that were intrinsically interesting before the reward was introduced. The undermining effect operates through the psychological process of internalization: introducing external evaluation and control on an activity that the person was engaging in for its inherent value converts the activity from a source of genuine engagement into a transactional exchange, reducing the discretionary investment and creative problem-solving that intrinsic motivation produces and replacing them with the minimum performance necessary to earn the extrinsic reward.

The organizational application of the intrinsic motivation research requires distinguishing between the types of performance where intrinsic motivation most significantly determines outcomes and those where extrinsic motivation is adequate. For high-routinization, low-judgment tasks where excellent performance means consistent execution rather than creative problem-solving, extrinsic motivation and contingent measurement produce adequate performance outcomes. For high-autonomy, judgment-intensive roles where excellent performance requires genuine intellectual engagement, creative problem-solving, and sustained investment beyond the minimum, intrinsic motivation substantially determines the quality difference between adequate and excellent performance. Organizations applying extrinsic measurement and contingent reward structures to the latter category of work are systematically undermining the motivational conditions that produce the performance quality those roles require.

The practical measurement implication is that measurement system design should distinguish by role type: for high-routinization roles, measure outputs with sufficient specificity to distinguish good from poor execution; for high-judgment roles, measure the conditions and behaviors that are theoretically associated with excellent judgment-intensive performance rather than the specific outcomes that the individual may produce as much through circumstance as through judgment. Measurement approaches that provide informational feedback, accurate performance information without controlling evaluative pressure, support intrinsic motivation in high-judgment roles while providing the performance information needed to identify and address genuine performance problems. Measurement approaches experienced as controlling, as external pressure to produce specific measured numbers, undermine the intrinsic motivation that most directly determines performance quality in the judgment-intensive roles that organizations most want to improve.

Measurement Design Principles

The measurement design principles most consistently associated with sustaining intrinsic motivation alongside appropriate accountability distinguish between measurement for development and measurement for evaluation, recognizing that the two functions require conditions that are not simultaneously achievable in the same measurement context. Developmental measurement provides accurate information about current performance relative to development goals, framed in ways that reduce ego threat and activate learning orientation rather than defensive self-protection. Evaluative measurement provides an accurate comparative assessment of performance against organizational standards, framed to produce accurate and consistent evaluation without the relational softening that compromises accuracy. Organizations that attempt to serve both functions with a single measurement instrument or a single periodic evaluation conversation consistently fail at both.

Behavioral specificity in measurement, describing specific observable behaviors and their consequences rather than aggregating across behavioral patterns into global performance verdicts, serves both developmental and motivational functions simultaneously. Behavioral specificity increases the information value of feedback for behavioral adjustment, because it identifies what specifically to change rather than how the overall performance is judged. It also reduces the ego-threat component of measurement by separating specific behavioral evaluation from global person evaluation, making it possible to engage analytically with the behavioral finding rather than defending against it as a characterization of one's overall performance or professional identity, a distinction Kluger and DeNisi's (1996) meta-analysis of feedback interventions found to be practically consequential: feedback directed at the task or behavior level improved subsequent performance, while feedback that shifted attention to the self, global judgments about the person rather than the specific behavior, frequently produced no improvement or measurable performance decline, even when the substantive content of the feedback was accurate.

The measurement of behaviors that organizations cannot directly observe, including the quality of organizational relationships, the degree to which a leader develops others, the extent to which a professional's advice reflects genuine client interest rather than client preference, requires multisource assessment that aggregates the observations of those in the best position to observe those behaviors. Multisource measurement, collecting behavioral observations from those who experience the behaviors being assessed, is more costly and administratively complex than supervisor-only or outcome-only measurement, but it is the only measurement approach that can reliably capture the behavioral dimensions that most directly determine organizational outcomes in high-autonomy, high-judgment, and high-relational-demand roles. Organizations willing to make the measurement investment required by these roles produce more accurate performance information and correspondingly better performance management, development, and talent decisions.

Building Measurement Systems That Align with Organizational Values

Measurement design: intrinsic motivation implications
Measurement design featureSupports intrinsic motivationUndermines it
Information vs. controlFeedback framed as accuracy informationContingent reward tied to specific metric
Behavioral specificityDescribes specific behaviors and consequencesGlobal ratings or aggregate competency scores
Multisource observationCollects behavioral observation from multiple vantage pointsSingle-rater judgment subject to idiosyncratic bias
Separate dev from evalDifferent instruments for development vs. administrativeSame conversation serves both evaluation and development
Figure 2. The measurement design principles that best sustain intrinsic motivation alongside accountability separate informational from controlling measurement and specify behavioral targets precisely enough to distinguish performance from proxy optimization.
Ryan and Deci, 2000; Campbell and Fiske, 1959

The most consequential test of an organizational measurement system is not whether it accurately captures what it measures, a test most well-designed systems pass, but whether what it measures accurately represents what the organization most needs from the people being measured, a test that most organizational measurement systems fail. Building measurement systems that pass both tests requires a design process that begins with the identification of the specific behaviors and outcomes that most determine the performance the organization needs, and that works backward from that identification to the measurement approach that would most accurately and most motivationally capture those behaviors and outcomes.

This design process is rare in organizational practice because it requires articulating with precision what excellent performance looks like at the behavioral level, which is organizationally difficult both intellectually and politically. Intellectually, identifying the specific behavioral patterns that distinguish excellent from adequate performance in complex, judgment-intensive roles requires careful observation and analysis of what high performers actually do differently, not just what outcomes they produce. Politically, the behavioral specificity that makes a measurement system motivationally valid and behaviorally accurate also makes explicit what was previously implicit, creating organizational accountability for specific behaviors that some stakeholders may prefer to leave ambiguous.

Organizations that complete this design process, beginning with behavioral specificity rather than measurement convenience, produce measurement systems that are simultaneously more accurate as performance assessments, more motivationally appropriate for the roles being measured, and more resistant to the gaming behavior that proxy-heavy measurement systems reliably generate. The investment in the design process is substantial but is recovered through better performance management decisions, better talent decisions, and the productivity improvement that comes from employees directing their effort toward the behaviors that genuinely matter rather than toward the measured proxies that the convenience of measurement has substituted for them.

The Validity Problem Beneath the Motivation Problem

Campbell and Fiske's (1959) validity criteria, introduced earlier in this monograph's discussion of measurement architecture, connect to the motivation research in a way worth making explicit rather than leaving as two separate arguments. A measurement system that fails discriminant validity, that cannot reliably distinguish high performers from low performers on the dimensions that actually matter, does not merely produce inaccurate performance information; it produces exactly the kind of arbitrary, poorly calibrated evaluative pressure that self-determination theory research identifies as most corrosive to intrinsic motivation. Employees experience low-validity measurement not as objective performance information but as unpredictable, poorly justified evaluation, and unpredictability in evaluative consequences is itself a documented source of the controlling, ego-threatening measurement experience that undermines the motivational conditions this monograph has reviewed throughout.

This connection matters practically because it means the validity problem and the motivation problem are not independent design challenges requiring separate solutions; they are the same underlying design failure viewed from two different angles. A measurement system built with genuine behavioral specificity and multisource convergent validation, the design this monograph has advocated on Campbell and Fiske's validity grounds, simultaneously produces the informational, non-controlling feedback experience that self-determination theory research identifies as motivationally protective. Organizations do not face a genuine trade-off between measurement accuracy and employee motivation, despite how the two concerns are frequently discussed as competing organizational priorities requiring balance; well-designed measurement serves both simultaneously, and poorly designed measurement, regardless of intent, undermines both simultaneously as well.

This unification also clarifies why the multisource measurement this monograph has recommended for high-judgment, high-relational-demand roles carries a motivational benefit beyond its accuracy benefit. A single-source measurement, one supervisor's judgment of a complex, multidimensional performance, is inherently more vulnerable to the reliability problems Campbell and Fiske's framework warns against, and that reduced reliability translates directly into the unpredictable, poorly justified evaluative experience that undermines intrinsic motivation. Multisource assessment improves validity by triangulating across observers and simultaneously improves the employee's motivational experience of being measured, because convergent evaluation from multiple credible sources is experienced as more legitimate and less arbitrary than a single evaluator's judgment, regardless of how capable that single evaluator is.

Limitations: The Magnitude Debate

The case this monograph has built on Deci, Koestner, and Ryan's (1999) meta-analytic finding that extrinsic rewards undermine intrinsic motivation requires acknowledging a genuine and substantive academic controversy the finding generated rather than treating it as uncontested. Cameron and Pierce's (1994) competing meta-analysis of the same general research literature reached a materially different conclusion: they found that the undermining effect, while real under specific conditions, was considerably smaller and more conditional than Deci and colleagues' analysis suggested, occurring reliably only for tangible, expected rewards tied to task completion regardless of quality, and largely absent or even reversed for rewards tied to performance quality or offered as unexpected recognition rather than pre-announced incentive.

The methodological disagreement between these research programs turned substantially on which studies were included and how reward conditions were categorized, a genuine scientific dispute this monograph cannot resolve by asserting one meta-analysis is simply correct and the other simply wrong. What the controversy does clarify, and what both research programs' more careful readings support, is that the undermining effect is not a blanket argument against extrinsic measurement of any kind; it is specifically strongest for tangible rewards tied mechanically to task completion in ways divorced from performance quality, precisely the kind of proxy-driven, activity-based measurement this monograph's opening discussion of the Kerr problem already identified as poorly designed for independent reasons. Measurement tied to genuine performance quality, and recognition that is not experienced as controlling pre-announced incentive, appears considerably less vulnerable to the undermining effect in both research programs' findings, which narrows rather than eliminates this monograph's practical recommendation: the risk is concentrated in exactly the proxy-heavy, activity-based measurement approaches that are already poorly designed on independent, non-motivational grounds.

Synthesis

The evidence this monograph has reviewed, spanning Kerr's foundational documentation of measurement-aspiration misalignment, Strathern's articulation of Goodhart's Law, and the self-determination theory research on intrinsic and extrinsic motivation, converges on a specific practical distinction that most organizational measurement practice does not make: the difference between measuring genuine performance and measuring a tractable proxy for performance. The proxies are easier to measure precisely because they are further from the underlying performance organizations actually care about, and that same distance is what makes them systematically gameable and, per the magnitude debate this monograph has engaged directly, most vulnerable to the motivational undermining effect when tied mechanically to reward.

Building measurement systems that avoid this trap requires the harder, more behaviorally specific design work this monograph has described: identifying what excellent performance actually looks like at the level of observable behavior, tolerating the administrative cost of multisource assessment where single-source measurement cannot reach the relevant behaviors, and distinguishing developmental from evaluative measurement rather than asking one instrument to serve both functions. Organizations willing to make this investment produce measurement systems that motivate rather than merely track the performance they were designed to capture, avoiding both the gaming behavior Kerr's original analysis documented and the motivational undermining that concentrated, proxy-driven extrinsic measurement most reliably produces.

References
Explore the Evans Learning Labs diagnostic toolkit
Structured diagnostic tools covering leadership, team, and organizational performance.
Browse the toolkitBack to articles
Welcome back
Sign in to access your assessments
No account?
Terms of Use and Disclaimer

Informational and Educational Use Only

The diagnostic tools, assessments, profiles, and indexes offered by Evans Learning Labs are designed for informational and educational purposes only. Results do not constitute professional consulting advice, legal advice, psychological assessment, clinical evaluation, or any form of certified professional guidance.

Self-Reported Results

All results are based entirely on the responses provided by the individual completing the assessment. Evans Learning Labs makes no representation that scores or profiles accurately reflect objective organizational conditions or any other measurable external reality.

No Guarantee of Outcomes

Evans Learning Labs does not guarantee that use of these tools will produce any specific organizational, leadership, or performance outcome. Recommendations are general in nature and may not be appropriate for every individual, team, or organizational context.

Limitation of Liability

To the fullest extent permitted by applicable law, Evans Learning Labs, its principals, employees, and affiliates shall not be liable for any direct, indirect, incidental, consequential, or punitive damages arising from the use of or reliance on these tools or their results.

Governing Law

These terms are governed by the laws of the United States and Commonwealth of Kentucky.