Sampling, Extrapolation, and the Statistics Cricket Fans Already Understand

Sampling, Extrapolation, and the Statistics Cricket Fans Already Understand

Cricket supporters are, without particularly noticing it, unusually comfortable with statistical inference. We accept that a strike rate over twelve balls means less than one over twelve hundred. We argue about Duckworth-Lewis-Stern precisely because we understand it is estimating a full innings from a partial one. We instinctively distinguish a sample from a population, which is more than most people manage.

That fluency turns out to be useful for understanding one of the biggest financial stories in American healthcare, because the mechanism at its centre is exactly the kind of statistical reasoning cricket has been training us on for years.

The setup

American private insurers cover more than thirty million older adults, receiving monthly government payments that rise with how ill each member’s medical records show them to be. Federal auditors check whether those records actually support the diagnoses claimed.

Checking every member is impossible; there are millions. So auditors do what any sensible statistician does, and what every cricket analyst does instinctively. They take a sample.

The bit that changes everything

Here is where cricket fans get an advantage in understanding this story. For years, the audit programme worked like a batting average calculated only from the innings you happened to watch. Auditors examined a sample, found errors, and recovered only those specific errors. The sample told you about the sample and nothing more.

The current methodology does the thing every cricket statistician does automatically: it treats the sample as an estimate of the whole. A sample of between 35 and 200 members is drawn, the error rate is calculated, and that rate is then applied across the insurer’s entire contract population.

Think about what that does to the arithmetic. Under the old approach, a 30 percent error rate found in 150 charts cost you 150 charts’ worth of repayment. Under extrapolation, it estimates a 30 percent error rate across hundreds of thousands of members and recovers accordingly. Same sample, same fieldwork, vastly different consequence.

Cricket fans will recognise this as the difference between reporting what happened in the overs you saw and estimating what would have happened across the full fifty.

Why the estimate is defensible

The obvious objection is the one every DLS argument produces: how can you charge someone for members you never examined?

The answer is the same as it is in cricket. If the sample is drawn properly, it is representative by construction, and the estimate it produces is the best available description of the population. An insurer arguing that its unexamined records are cleaner than its examined ones is making a claim it cannot support, exactly like a side arguing that its unbowled overs would have been unusually productive.

There is a second layer that cricket analysts will appreciate. Auditors do not draw uniformly at random. They deliberately sample high-risk categories: conditions that are frequently miscoded, acute events coded outside acute settings, single-occurrence diagnoses. It is targeted sampling, and it produces a higher error rate than a uniform sample would, which is precisely the point. You test where failure is likely, not where it is comfortable.

What the sampling found

Federal reviews of three insurance plans published this spring found that between 81 and 91 percent of the sampled high-risk diagnosis codes were not adequately supported by the medical records. In individual audits, certain acute condition categories reached 100 percent error rates. In March, a major insurer agreed to pay 117.7 million dollars to settle federal claims about its records.

The audit workforce has grown from roughly forty reviewers to around two thousand certified coders, working on a rolling quarterly cycle through successive payment years.

The behavioural effect, which is the real point

Extrapolation does something no amount of enforcement rhetoric achieves. It makes every record matter, because any record might be in the sample.

When only found errors are recovered, the rational strategy is to fix problems when caught. When sampled errors are extrapolated, the only rational strategy is to keep the underlying error rate genuinely low everywhere, all the time. Insurers have responded by building RADV audit readiness for payers around the government’s own sampling logic, running internal audits quarterly and correcting what fails before anyone officially arrives.

That is a well-designed incentive, and cricket fans should appreciate the elegance. You do not need to watch every ball. You need the balls you watch to be genuinely representative, and everyone to know they are.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *