Uncertain Terms

If a vendor promises us that their tool is 99% accurate, what should we be asking of them?

A 99% accurate model can still hurt people

Accuracy without a disease rate is essentially meaningless; at actual TB rates, a 99% accurate screening tool would flag hundreds of healthy people for every sick person.

Sep 28, 20264 min readPost 1 of 1

According to provisional data released by the CDC in March of this year, there were an estimated 10,260 cases of tuberculosis in the United States in 2025. That works out to about 3 cases per 100,000 people. Save this number. It's the one that turns "99% accurate" into a problem.

Imagine a company in Miami has built an AI tool that screens patients for tuberculosis from a chest X-ray. The company says the tool is 99% accurate. Take them at their word.

Now run the tool on 100,000 patients. Three of them actually have tuberculosis, and the tool correctly identifies all three. Of the remaining 99,997 patients, it gets 99% right. That sounds fine until you do the math. One percent of 99,997 is about 1,000 patients who get flagged as possibly having tuberculosis when they don't.

So the tool flags about 1,003 patients, and only 3 of them actually have tuberculosis. The chance that a flagged patient really has the disease is about 3 in 1,003, which is less than half a percent. The tool is 99% accurate, yet 99.7% of the patients it flags do not have tuberculosis.

That isn't a flaw in the tool. It's a mathematical result. The fraction of flagged patients who actually have the disease is called the positive predictive value, or PPV.

PPV depends on two things: how well the tool performs and how common the disease is in the population being tested. The vendor controls the first. The second is out of their hands, and it is rarely mentioned in the brochure.

A 2022 clinician's guide in the Journal of Clinical Psychiatry walks through the classic textbook example. Say a screening test correctly identifies 90% of patients who have the disease (that's its sensitivity) and correctly clears 90% of patients who don't (that's its specificity). If half the patients being tested have the disease, the PPV is 90%. If only 5% have it, the PPV drops to 32%.

Take a group of 200 patients where 10 have the disease. The test correctly identifies 9 of those 10, but it also flags 19 of the 190 healthy patients. Of the 28 flagged patients, only 9 actually have the disease. That's a PPV of 32%, and most flagged patients are healthy. Same test, same 90 and same 90, but what a positive result means has changed completely.

Now go to the other end. According to the World Health Organization, about 10.7 million people became ill with tuberculosis in 2024. That's 131 cases per 100,000 worldwide. Suppose the screening tool has 90% sensitivity and 99% specificity. In a population of 100,000, it would find about 118 real cases and also flag about 1,000 healthy people. The PPV comes out to about 10.6%. That isn't great, but it's one in ten, and a clinic could reasonably use it to decide who should get a sputum test. In Miami, the same tool would be right on one flag in three hundred.

The tool hasn't changed. The disease rate has. That's one reason I think "how accurate is it?" is a nearly meaningless question. Accurate on whom, and at what disease rate?

I know the pull of the big number myself. In my research at Florida Atlantic University, I'm training a stacked ensemble on chest X-ray images to detect tuberculosis. On the images it was trained on, the models score over 99%.

As a teenager writing my first paper, I watched a Colab run finish with 99.4% on the screen, and I wanted to put it in the abstract. I know that's a true number. I also know it's the least informative true number in the project.

When the same model is tested on chest X-rays from hospitals it has never seen, it scores between 67% and 86%, depending on the source. That spread is the actual result. The confidence interval is part of the actual result too.

A Clopper-Pearson interval is a common way to show that, given the number of test images, the model's true performance probably falls within a certain range. A model tested on a few hundred images can score 99% and still carry a confidence interval several points wide. A number without a confidence interval is a guess dressed up as a fact.

I think the honest way to report results is to show the errors rather than the accuracy. How many real TB patients were missed, and how many healthy patients were flagged, using data from sources other than the training data, with a confidence interval. A model can score 99% on a dataset where 99% of patients are healthy just by predicting "healthy" every time. That model has never found a case of TB in its life.

There's also a human cost on both sides of the error. A missed case of TB is a patient who keeps coughing on the bus. A flagged healthy patient is someone waiting two weeks for a sputum culture and losing sleep over it, multiplied by a thousand. None of that is captured in the "99%."

So here's the question I think a hospital should ask instead of "how accurate is it?" At the disease rate in our hospital, out of the patients your tool flags, how many actually have TB, and how many patients with TB do you miss? And a follow-up: were those numbers determined, with a confidence interval, on patients from hospitals that weren't used to train the model?

If the vendor can't answer, they may still have a good model. They just haven't yet shown that it works on patients from hospitals other than the ones it learned from.

Sources

  1. CDC, provisional 2025 tuberculosis data, 2026, https://cdc.gov/tb-data/aboutprovisionaldata/2025-provisional-data.html
  2. CIDRAP, summary of the CDC provisional 2025 TB data, 2026, https://www.cidrap.umn.edu/tuberculosis/cdc-data-suggest-small-decline-us-tb-cases
  3. World Health Organization, Global Tuberculosis Report 2025, TB incidence, 2025, https://www.who.int/teams/global-programme-on-tuberculosis-and-lung-health/tb-reports/global-tuberculosis-report-2025/tb-disease-burden/1-1-tb-incidence
  4. Zimmerman, clinician's guide to positive predictive value and screening tests, Journal of Clinical Psychiatry 83(5), 2022, https://www.psychiatrist.com/jcp/positive-predictive-value-clinician-guide-avoid-misinterpreting-results-screening-tests/