A dashboard does not report on reality. It reports on the handful of things somebody decided were worth watching, measured in a way somebody decided was acceptable, against thresholds somebody decided were safe. Every one of those decisions was reasonable when it was made. None of them is revisited. That is how a company runs six green quarters and loses its two largest accounts in the seventh.

A dashboard reports on choices, not on the business

Green means that the things you chose to watch are inside the range you chose to accept. It does not mean that the business is healthy. The gap between those two statements is where most expensive surprises live, and the gap is invisible by construction: a dashboard cannot display the absence of a metric.

This matters most for executives who cannot audit the underlying system themselves. A chief executive with a commercial or financial background has no independent way to test whether an availability figure, a ticket count or a delivery percentage still stands for the thing it was created to stand for. They read the number. The number is correct. The conclusion they draw from it may be wrong by a wide margin, and nothing on the screen will tell them.

The people who built the dashboard are not hiding anything. They built a set of indicators that could be read in fifteen minutes by a board that does not speak the language of the system. Legibility was the design goal. Discomfort was not.

The most famous health target in the world started as a product name

The 10,000 step goal was never a clinical finding. It comes from the name of a Japanese pedometer launched in 1965 around the Tokyo Olympics, the Manpo-kei, literally the 10,000 step meter. The number was chosen because it sounded right and the character for 10,000 looks like a walking figure. Research since then, including work led by I-Min Lee at Harvard, found that mortality benefits in older women level off well below that, around 7,500 steps a day.

Sixty years later, hundreds of millions of wrists carry a device that turns green at a figure invented by a marketing department.

The interesting part is not the trivia. It is what the figure did to behaviour. People walk in circles in their kitchen at eleven at night to close a ring. They do not do this because they are stupid. They do it because a target was set and they are meeting it. The target was a proxy for something broad and hard to measure: move, get outside, do not spend the day in a chair. Once the proxy became the objective, the underlying thing stopped being the subject of the evening question. The question became whether the ring was closed.

Nobody in that chain lied. A product name became a target, a target became a proof, and the proof replaced the thing it was standing in for.

Goodhart’s law describes an incentive, not a character flaw

Charles Goodhart, writing about UK monetary policy in 1975, observed that a statistical regularity tends to collapse once pressure is placed on it for control purposes. Marilyn Strathern gave the formulation that travelled, in a 1997 paper on audit in British universities: when a measure becomes a target, it ceases to be a good measure. Donald Campbell had described the same dynamic for social indicators in 1979, with a sharper edge: the more an indicator is used for decision making, the more it will be corrupted, and the more it will distort the process it was meant to monitor.

Three disciplines, three decades, one mechanism. That is usually a sign that the mechanism is structural rather than cultural.

The word corruption is misleading to a modern ear. It suggests fraud. In practice almost nothing that degrades an indicator is fraud. It is optimisation, performed by competent people who were asked for a number and delivered it.

A support team measured on ticket resolution time will split large tickets into smaller ones. That is a defensible operational decision, taken for good reasons, and it moves the number without moving a single customer’s experience. A delivery team measured on velocity will stop estimating the genuinely hard work, because hard work poisons a velocity chart. A sales organisation measured on products per customer will open products customers did not ask for, which is how a very large American bank ended up with several million unauthorised accounts and a cross-selling metric that looked excellent right up to the moment it did not.

At each step, someone did what the organisation asked. The organisation asked for the number.

Surrogation is the moment the proxy replaces the objective in people’s heads

Management researchers have a name for the specific failure that follows. Willie Choi, Gary Hecht and Bill Tayler called it surrogation: the tendency of people to mentally substitute the measure for the strategy it represents. Harris and Tayler laid it out for executives in Harvard Business Review in 2019, and their diagnosis is uncomfortable because it is not about bad actors. It is about cognition. Strategy is abstract, hard to hold in mind, and impossible to check on a Tuesday afternoon. The metric is concrete, visible, and settled by lunchtime.

So the metric wins. Not in the reporting pack, in the heads of the people doing the work. They are not tracking a proxy for customer trust any more. They are tracking the number, and they believe the number is the thing.

Incentive compensation accelerates this, which is why the effect shows up most violently where bonuses are attached. It does not require compensation. A weekly review in which one figure is always the first slide is enough. Attention is a currency, and the organisation spends it on whatever is displayed.

Six green quarters and two lost anchor accounts

What follows is a composite. The shape is common enough that the details do not need to belong to any single company, and none of them do.

A B2B services business, roughly three hundred people, sells a platform into mid-market customers with long contracts and slow renewal cycles. The chief executive came from the commercial side. The board pack carries four technology indicators: platform availability, support ticket volume and closure rate, engineering velocity, and a customer satisfaction score. Every one of those indicators was chosen sensibly, three years earlier, by a technical director who has since left.

For six quarters, all four are green.

Availability sits above 99.9 percent. The figure is accurate. It is also computed on a definition of downtime that excludes planned maintenance windows, and the windows have been getting longer. Two years earlier maintenance meant a quiet hour on a Sunday. It now means a four hour window most Wednesday evenings, because a core integration has become fragile enough that changes are batched and staged carefully. Customers in the Americas experience this as the product being unreliable on Wednesday afternoons. The availability metric experiences it as nothing at all, because planned maintenance was excluded from the definition on the day the definition was written, when maintenance was genuinely rare.

Ticket closure rate is excellent and improving. The improvement started shortly after closure rate became a quarterly objective for the support organisation. Tickets are now scoped more tightly, which is good practice, and multi part problems are logged as separate tickets, which is also good practice. A customer whose integration breaks now generates four tickets, each closed inside the target, instead of one that stayed open for nine days. The board sees a rising closure rate. The customer sees the same nine days.

Velocity is up eighteen percent year on year. Part of that is real. Part of it is that the team has quietly routed around the oldest and most dangerous part of the codebase, the batch job that reconciles customer entitlements every night. Nobody has decided to avoid it. It is simply that work touching it is unpredictable, unpredictable work wrecks a sprint commitment, and sprint commitments are visible. The reconciliation job now runs on a version of a runtime that has been out of support for over a year, and exactly one engineer understands it. That fact appears on no indicator, because there is no indicator for it.

The customer satisfaction score holds at a strong number. The survey is sent thirty days after a support interaction is closed, to accounts with no open critical incident. The exclusion was added for a reasonable purpose: not to ask angry customers to rate you in the middle of an outage. Its practical effect is that the most damaged relationships are systematically outside the sample.

None of this is fraud. Every single choice above was made by a competent person for a defensible reason, most of them years before the consequence appeared. There is no one to blame, which is precisely what makes it survive.

There is one more layer, and it is the one that makes the situation durable. Each of these four indicators is owned by a different person, and each owner is reporting accurately on their own scope. The head of infrastructure reports availability as defined. The support manager reports closure rate as defined. The engineering lead reports velocity as measured by the tool. The customer success lead reports the satisfaction score the survey returns. No individual has visibility across all four, and no individual is wrong.

The technical director who would have held the whole picture left two years earlier and was replaced by a delivery manager with a narrower remit. That single organisational change, made for sensible cost reasons, removed the only role in the company whose job was to look across the indicators rather than at one of them.

What the chief executive experiences is six quarters of a technology function performing well. Investment decisions follow: the technology budget is held flat for two years, because the indicators show a system in good shape and the growth argument is being made in sales. Two commercial hires are made instead of the platform engineering role the departing technical director had asked for.

In the seventh quarter, the two largest accounts, together around a quarter of recurring revenue, decline to renew. Both cite reliability and responsiveness. One of them has been running a parallel evaluation for five months. The chief executive learns the reasons in the exit conversation, not in the board pack, and the reasons are things that were happening continuously for two years, in plain sight, on the other side of four green indicators.

The cost is not the outage that never happened. The cost is two years of capital allocated on the strength of a picture that measured activity and called it health, and a renewal conversation that was lost before anyone knew it had started.

The recovery is worse than the loss. Fixing the reconciliation job, rebuilding the availability definition, and re-establishing trust with a mid-market reference base takes longer and costs more than the platform engineering hire would have. The company survives. The two years do not come back.

Vanity metrics survive because they are easy to defend, not because they are useful

There is an economy to this, and it is worth stating plainly because it explains why the problem regenerates after every clean-up.

Every indicator in a board pack has to survive a specific test: it must be explicable in under a minute to an audience that does not share the vocabulary of the system being described. That test favours a particular kind of number. Counts over conditions. Averages over distributions. Percentages of things completed over assessments of whether the completed thing works. Anything that requires a caveat loses to anything that does not.

The people who assemble the pack know this. They are not cynical about it. They are doing the reasonable thing, which is choosing figures that will not derail a meeting with a definitional argument. The result is a systematic bias toward metrics that are legible rather than metrics that are informative, and the bias operates every quarter, in every organisation, without anyone intending it.

There is a second force pushing the same way. An indicator that can make you look bad is an indicator you have an interest in not proposing. Nobody in the chain that produces the reporting has a professional incentive to introduce a number that says the system is more fragile than it looks. The absence is not a conspiracy. It is the sum of many individually sensible omissions, made by people who each assumed someone else was covering it.

Put those two forces together and you get a dashboard that drifts steadily toward reassurance. Not through any single decision. Through the accumulated gravity of a hundred small, defensible ones.

This is also why asking the same chain to review its own indicators does not work. The review will be honest and it will change nothing important, because the people conducting it share the constraints that produced the problem. Every metric will be confirmed as accurate, because every metric is accurate. Accuracy was never the issue.

Four questions that separate a health indicator from a decorative one

The audit that matters is not arithmetic. Take each indicator in the pack and answer four things about it.

Who built it, and to persuade whom. Indicators built to reassure a board behave differently from indicators built to run an operation, and after a few years nobody remembers which kind they were dealing with.

What it was originally standing in for. Availability stands in for customers being able to work. Ticket closure stands in for problems being resolved. Write down the underlying thing, then ask whether the indicator would still move if the underlying thing changed.

What it excludes today. Every metric has a definition, every definition has exclusions, and exclusions written years ago under different conditions are where the drift accumulates. Planned maintenance. Accounts with open incidents. Work that was not estimated.

Who gains if it goes up. Not to catch anyone. To know which figures are under pressure, because those are the ones that will degrade first.

Four questions, applied across a dozen indicators, takes a couple of days. It has to be done by somebody outside the chain that produces the numbers, and that is the entire difficulty, because the organisation has no natural role for that person.

A better dashboard is not the fix

The instinct after discovering a decorative metric is to replace it with a better one. That instinct is right in principle and it fails in practice, because the replacement enters the same system that degraded the original.

A new indicator arrives clean. It describes something real, it is not yet attached to an objective, and for two or three quarters it tells the truth. Then it gets a target, because an indicator that nobody is accountable for tends to disappear from the pack. Once it has a target it acquires the same gravity as everything else, and the same slow drift begins.

What survives is not a metric. It is a habit: someone from outside the chain that produces the numbers looks at the definitions on a fixed cadence and asks what changed. Not the values, the definitions. What was excluded that was not excluded last year. Which threshold was adjusted, and by whom, and what was happening that month. Which number started improving unusually fast, and what became an objective shortly before.

That review is unglamorous and it takes two days a year. It is also the only mechanism that catches drift, because drift does not announce itself in the values. It hides in the definitions, and definitions are the one part of a reporting pack that nobody reads twice.

Organisations that do this well tend to keep a short written record of why each indicator exists and what it was meant to stand for. Not a governance artefact. A paragraph per metric, written when it was introduced, readable by the person who inherits it three years later and has no idea what the original question was.

Dashboard analytics on a laptop screen

The bill for a green dashboard always arrives late and in full

The particular cruelty of this failure mode is its timing. Nothing goes wrong while it is happening. There is no incident to investigate, no threshold breached, no moment where a reasonable executive should have known. The degradation is continuous and the reporting is stable, so the only signal available is the absence of surprise, which is not a signal anyone can act on.

By the time it surfaces, it surfaces as revenue.

A dashboard is not a view of your business. It is a model of your business, built by people with constraints, in a context that has since changed, and it decays the way any unmaintained model decays. The decay is silent and the colour stays the same.

Green is not an answer. It is a claim, made by an instrument, about a question somebody else chose.

Sources

FAQ

Why can a dashboard show all green indicators while a business is actually deteriorating?

Because a dashboard reports on the handful of things someone decided were worth watching, measured in a way someone decided was acceptable, against thresholds someone decided were safe — not on the business itself. Definitions drift silently over time (a maintenance window definition written when outages were rare, a satisfaction survey that excludes accounts with open incidents), and each individual owner reports their own metric accurately. No one has visibility across all the indicators at once, so nobody notices the aggregate picture no longer matches reality.

What is Goodhart's law and why does it explain vanity metrics?

Formulated by Charles Goodhart in 1975 and sharpened by Marilyn Strathern in 1997 (‘when a measure becomes a target, it ceases to be a good measure’), it describes how any statistical regularity collapses once people are pressured to hit it. This isn’t fraud — a support team that splits large tickets into smaller ones to hit a closure-rate target is making a defensible operational choice that happens to move the number without improving the customer’s actual experience. The fix isn’t a better metric; a clean replacement metric enters the same system and starts drifting the moment it gets its own target.