Uncertain Terms

About

Why this exists

I ran one of my models on chest X-rays from a hospital that wasn't in my training data, and the accuracy went from above 99% to somewhere in the sixties. There were two columns in a spreadsheet. The first was the kind of number a vendor might feature in a brochure. The second looked like a coin that had been tilted slightly.

Then came a long debate over which column to use in the paper. Both were true. The 99% was the model's performance on data similar to what it had been trained on. The sixties were its performance on strangers. I concluded that a hospital only ever meets strangers.

That debate is what this site is about. I train deep learning models on chest X-rays and CT scans, and then I spend most of my time trying to break them. I run them on images from sources they weren't trained on. I check whether a 90% is really 90%. I keep each patient out of either the training data or the test data, never both, and I put a confidence interval next to every number.

Each of those is a choice, and each choice comes with a question someone outside the lab could ask. A hospital might want to know who a model was tested on. A regulator might want to know what "tested" means. A patient might want to know whether the model that read their image had ever seen a patient like them. Uncertain Terms turns the choices I make at my desk into those questions, in plain language. Expect a post about every two weeks and a podcast episode about every three.

Who I am

I am currently a senior at Coral Glades High School in Coral Springs, Florida, and dual-enrolled at Broward College. Since November 2025 I have been a research assistant at Florida Atlantic University in the lab of Dr. Harshal Sanghvi, working on deep learning applied to medical image classification with an emphasis on calibration and on performance on data that the model has never seen before. I work on everything from data preprocessing and patient-level splits through calibration and evaluation.

I have three projects accepted at peer reviewed conferences. CXR-StackNet is a calibrated ensemble of four convolutional networks for tuberculosis detection from chest X-rays and was accepted at ICITA 2026, where I presented it in Sydney. A hybrid model for early lung cancer detection from CT was awarded the Best Paper in its category at IEEE AIC 2026. PneuStack is a calibrated ensemble for pneumonia detection and was accepted for an oral presentation at ICECCME 2026 in October. A paper on differentiating tuberculosis from the diseases that mimic it is in preparation.

Since March 2026 I have also shadowed Dr. Gupta at a private orthopedic practice, where I watch how a clinician handles an ambiguous radiograph. I log the cases and compare them to my models' behavior on difficult images. One thing the clinic has that my models do not is "let me look again."

At school I am Director of Operations of Mu Alpha Theta. I was previously president for a year and during that time I helped revive a dormant chapter and created the first peer-led math tutoring program in the school, with twelve volunteer tutors serving over a hundred students. I am also vice president of activities for DECA and a voting student member of the School Advisory Council. Outside school I coordinate youth programs at BAPS in Miami. I plan to major in computer science.

How I write these

Each post has its sources listed at the bottom. Every number that appears in a post can be traced back to a source. In the margin there is a note called "What I'm unsure about" which describes which claim I would drop first if the evidence were to change. If I make a mistake, the error is corrected in the location of the error and I put the date next to it. I do not present unpublished numbers as final and I do not show any patient images.

Mission

Uncertain Terms takes the decisions inside medical AI and translates them into questions a hospital administrator, a regulator, a legislator or a patient can ask. It argues that the rules should require a model to say how sure it is and to admit the limits of what it has seen.

CadenceA post about every two weeks, an episode about every three.
SourcesListed at the bottom of every post and episode.
UncertaintyEvery piece carries a note on what I'm least sure about.
CorrectionsNoted in place with a date. Email me.