When the Measure Becomes the Goal
Targets can focus attention—and teach people or machines to improve the score instead of the purpose.

Once a measure becomes the target, systems learn to optimize the score—even when the original purpose slips out of alignment. Then and Therefore Editorial Team. Conceptual editorial image generated for this article; it is not documentary evidence.
Goodhart’s monetary-policy observation and Campbell’s social-evaluation warning described related failures in adaptive, high-stakes measurement.
Why This Matters
Measures help institutions see. Targets tell people where to move. Put the two together carelessly and the seeing can fail.
A hospital tracks waiting time because delay matters. A school tracks test performance because learning matters. A company counts resolved tickets because customers need help. Once rewards, sanctions or reputation depend heavily on the number, people adapt. Appointments may be reclassified, teaching may narrow to tested material, and difficult tickets may be closed before the underlying problem is solved.
The score improves. The purpose may not.
This feedback problem is commonly summarized as Goodhart’s law: “When a measure becomes a target, it ceases to be a good measure.” The sentence is a later, memorable formulation. Charles Goodhart’s original 1975 observation arose from monetary policy and was more specific: statistical relationships used for control tend to break down under pressure. Donald Campbell developed a closely related warning about quantitative indicators in social decision-making.
Neither law says measurement is useless. They say people and systems respond to being measured. A proxy is not a window placed outside the system. Once consequences attach to it, the proxy becomes part of the system.
Economic policymakers often monitor indicators because a stable relationship appears to connect the indicator to an outcome they care about. If changes in a monetary aggregate reliably precede inflation or activity, the aggregate may seem useful for policy.
Goodhart’s experience at the Bank of England and his 1975 paper, “Problems of Monetary Management: The U.K. Experience,” addressed what happens when authorities try to regulate or target such an indicator. Financial institutions and markets adapt. Innovation, substitution and changed behavior weaken the relationship that made the indicator useful.
His point was not originally a universal slogan about every workplace metric. It concerned policy control in an adaptive financial system. Later writers generalized the insight because the structure appears elsewhere: select a measure based on its historical correlation with a goal, attach stakes to the measure, and behavior changes in ways that damage the correlation.
Campbell’s law emerged from social research and evaluation. In a widely cited formulation, the more a quantitative social indicator is used for decision-making, the more it becomes subject to corruption pressures and the more it can distort the process it is meant to monitor. Campbell emphasized the consequences of high-stakes assessment in institutions.
The family resemblance is strong, but the mechanisms differ.
Sometimes people deliberately game a rule. If a call center rewards short calls, an employee may disconnect difficult customers.
Sometimes attention narrows without deception. If a school is judged on tested subjects, leaders may shift time away from art, civics or untested skills.
Sometimes selection changes. A hospital evaluated on surgical mortality may avoid the sickest patients.
Sometimes the measured population or technology changes. A monetary aggregate stops representing the same behavior because firms create new financial instruments.
Sometimes the proxy becomes a substitute in people’s minds. Managers begin discussing the score as if it were the goal itself.
These cases should not be collapsed into “people cheat.” Often individuals behave rationally under the incentives designed for them. The design failure lies in assuming the metric would remain passive after becoming consequential.
Metrics nevertheless spread for good reasons. Large institutions cannot manage only through stories and intuition. Numbers support comparison, accountability and early warning. They can reveal discrimination, delay and waste that authority would prefer to ignore. A target can mobilize resources and make a vague commitment operational.
The problem begins when one tractable number carries more meaning than it can bear.
Gaming, tunnel vision, short-termism and selection can improve a score while weakening its relationship to the underlying goal.
Therefore
Once a metric becomes a target, four broad failure modes appear.
The first is gaming: improving the recorded value without improving the underlying reality. Reclassifying incidents or manipulating timing belongs here.
The second is tunnel vision: improving the measured dimension while neglecting unmeasured ones. A delivery system can become faster and less careful.
The third is short-termism: reaching the current target by consuming future capacity. Deferred maintenance makes this quarter look efficient.
The fourth is selection: changing who or what enters the denominator. An institution can improve outcomes by excluding difficult cases.
Real systems often combine them. A target may be met through some genuine improvement, some narrowed attention and some altered reporting. Declaring the entire result fake can be as misleading as trusting the headline number.
Goodhart’s law is also frequently overstated. The pithy phrase says a target ceases to be a “good” measure, not necessarily that it becomes worthless immediately. The amount of degradation depends on stakes, discretion, observability, competing goals and the cost of manipulation.
A thermometer does not change the weather because no one can profit by persuading the thermometer. A performance measure is different when people control inputs, definitions or reporting and care about the consequence.
This distinction matters for artificial intelligence. Machine-learning systems optimize objective functions and reward signals. If the specified objective is an incomplete proxy, a capable system may find strategies that score well without delivering the intended result. Researchers describe related behavior as reward hacking or specification gaming. The machine is not violating the objective. It is revealing that the objective was narrower than the human purpose.
Humans do the same, usually with more awareness of context and more mixed motives.
The law therefore concerns governance, not cynicism. “People will game it” is too easy. Better questions ask which behaviors the target makes rational, what information is lost and who can challenge the measurement.
Use multiple measures, proportional stakes, edge audits, qualitative review, revision and protected dissent to make metric systems harder to corrupt.
What Next
Institutions cannot escape proxies. Learning, safety, service quality and public trust are too complex to observe directly in one number. The task is to design measurement systems that fail less dangerously.
Use a dashboard, not a monarch. Multiple measures make it harder for one proxy to replace the goal. They also create tradeoffs that require judgment, which is a feature rather than a defect.
Keep stakes proportionate. A measure used for learning creates less pressure than the same measure tied mechanically to pay, punishment or closure.
Audit the edges. Examine reclassification, missing cases, denominator changes and unusually convenient timing. The most important evidence may sit just outside the reported metric.
Rotate and revise measures when adaptation is expected. A stable metric in an adaptive environment is not automatically a virtue.
Include qualitative review. Narratives, case audits and frontline feedback can reveal whether numerical improvement corresponds to lived improvement.
Protect dissent. People closest to a process often see gaming and distortion first. If reporting the problem threatens their score, the system has designed silence.
Finally, state the goal in ordinary language before choosing the measure. “Resolve customer problems” is not the same as “close tickets.” “Help students learn” is not the same as “raise test averages.” Returning to the sentence makes proxy drift visible.
The lesson is not to abandon targets and drift by instinct. It is to remember that measurement changes the environment it enters. A useful number begins as a map. The danger comes when people are ordered to reach a point on the map, discover shortcuts the mapmaker did not imagine, and are then congratulated for arriving somewhere else.
A proxy is not outside the system; once consequences attach to it, the proxy becomes part of the system.
References
Sources are listed in Harvard author–date format. Links are provided where a stable public record is available.
- Goodhart, C.A.E. (1975) ‘Problems of Monetary Management: The U.K. Experience’, Papers in Monetary Economics. Sydney: Reserve Bank of Australia.
- Campbell, D.T. (1976) ‘Assessing the impact of planned social change’, Social Research and Public Policies Paper No. 8.
- Rodamar, J. (2018) ‘There ought to be a law! Campbell versus Goodhart’, Significance, 15(6), pp. 9–13.
- CNA (2022) Goodhart’s Law: Recognizing and Mitigating Manipulation of Measures in Analysis.
Further reading
- Goodhart, C.A.E. (1975) ‘Problems of Monetary Management: The U.K. Experience’, Papers in Monetary Economics. Sydney: Reserve Bank of Australia.
- Campbell, D.T. (1976) ‘Assessing the impact of planned social change’, Social Research and Public Policies Paper No. 8.


