Uncertain Terms

Why does it matter whether an AI's confidence score means anything and how would a regulator go about checking it?

What it means for a model to say how sure it is

Modern neural networks tend to be overconfident. A hospital should ask for a calibration curve that shows whether the confidence reported by the model actually means anything.

Oct 8, 20264 min readPost 2 of 2

A 2017 paper titled "On Calibration of Modern Neural Networks" looked at something very few people think to check: whether the confidence score a neural network attaches to an image actually means anything. On a standard test set, one deep learning model (ResNet-110) was on average 16.5 percentage points more confident than it was correct. When the model said "I'm sure," it was wrong far more often than "sure" should allow.

The best way I know to explain this is a weather forecaster. If a good forecaster says there's a 70% chance of rain, look at every day they said that. It should have rained on about 70% of those days. If it rained on 95% of them, the forecaster is underconfident. If it rained on only 40%, the forecaster is overconfident.

When stated confidence matches how often the prediction is actually right, we call that calibration. A calibrated model is one where 90% means 90%.

A chest X-ray screening AI makes the same kind of claim. Instead of saying "TB," it says something like "0.91," and the clinician is meant to read that as a 91% chance of TB. For that to make sense, the model has to be calibrated. If it isn't, the confidence score is meaningless.

Guo and colleagues found that older, smaller models were fairly well calibrated. A small model from the 1990s (LeNet) was roughly as accurate as it was confident. The modern deep networks that replaced them are far more accurate, but also far more overconfident. The authors concluded that "modern neural networks are no longer well-calibrated." The design choices that made these networks better at classification, more depth and more width, also seem to have made them worse at knowing how sure they should be.

They measured this with a number called expected calibration error (ECE). In short, ECE takes a model's outputs, groups them by confidence, and then averages how far the stated confidence is from the true accuracy in each group. An ECE of 0% means a perfect forecaster. On the CIFAR-100 test set, ResNet-110 had an ECE of 16.5%.

The solution that worked

What I found interesting about this paper is how easy the problem was to fix. A method called temperature scaling (a one-number version of Platt scaling) reduced ResNet-110's ECE from 16.5% to 1.3%. On a much larger network tested on the ImageNet dataset, ECE went from 6.3% to 2.0%. Accuracy didn't change. The ordering of the predictions didn't change. Only the confidence changed.

The fix itself is small. After a model is trained, you hold back a set of data it hasn't seen. Then you learn a small adjustment on top of the model's outputs. Platt scaling learns two values, a stretch and a shift. Temperature scaling learns one, a divisor that softens or sharpens every confidence value uniformly. What matters is that the adjustment is learned on data that wasn't used for training. That way it measures how overconfident the model really is on new data, rather than how confident the model is about its own training performance.

In my lab we use Platt scaling on top of a stacked ensemble. Several pretrained models (DenseNet, EfficientNet, ConvNeXt and Xception) all look at the same X-ray. A small logistic regression model combines their outputs into one. Then a Platt scaling adjustment is applied to that single output using held-out data. Stacking makes the model more accurate. Platt scaling makes it more honest. Without it, an ensemble of confident models just becomes one very confident model.

The plot I'd want a hospital to ask for is called a calibration curve, or reliability diagram. The x-axis shows what the model predicted, from 0% to 100%. The y-axis shows how often it was actually correct at each value.

A perfectly calibrated model draws a diagonal across the plot. An overconfident model draws a line that sags below the diagonal. For example, a model claiming 90% confidence might turn out to be right only 70% of the time. The shape of this plot tells you more about a model than a single accuracy number, because it shows where the model's confidence can be trusted and where it can't.

I haven't seen one in a vendor's brochure yet.

This matters for regulation because of what a clinician can do with the model's output. A model that says "88% TB" and is wrong is a bad outcome. A model that says "55% TB" and is wrong is a different kind of outcome. A radiologist looking at a 55 might take a closer look at the image or order a sputum test. A 55 that really means 55 is useful. An 88 that really means 60 is worse than no confidence score at all, because it borrows the authority of precision without earning it.

Uncertainty is something a human can act on. False certainty takes that option away.

I've noticed this in my own work with model outputs. A 0.97 gets questioned less often than a 0.61, even when I know the model isn't calibrated and the 0.97 isn't really a probability. If that's true of a student glancing at a notebook between classes, I assume it's true of a tired radiologist at the end of a long day, and the model should be held accountable for it.

So if a regulator asked me how to check whether an AI's confidence score means anything, I wouldn't start with accuracy. I'd ask for a calibration curve and the ECE, computed on patients from hospitals that weren't used in training. I'd also ask what the vendor does after training to fix the confidence values, and on what data. If the answer is "nothing, the raw output is the probability," then we already know what to expect from the 2017 paper.

I'd also want the calibration curve recomputed after the model has been running for a year. Patient populations change, and a model calibrated on one hospital's patients isn't necessarily calibrated on another's. That's the part of this topic I'm least certain about, and the part I expect to get wrong first.

Sources

  1. Guo, Pleiss, Sun, and Weinberger, "On Calibration of Modern Neural Networks," Proceedings of Machine Learning Research, ICML, 2017, https://proceedings.mlr.press/v70/guo17a/guo17a.pdf
  2. Guo, Pleiss, Sun, and Weinberger, "On Calibration of Modern Neural Networks," arXiv preprint 1706.04599, 2017, https://arxiv.org/abs/1706.04599