Machine learning · Methodology · 2026
30-day readmission in diabetic patients
An end-to-end clinical ML pipeline, plus an experiment asking whether the validation protocol matters as much as the model. It doesn't, and the negative result is the most useful thing the project produced.
- AUC-PR on test (prevalence 0.114, lift 2.02×)
- 0.230
- readmissions caught in the top-10% risk group
- 24.3%
- the only resolved protocol effect
- −8.2%
Problem
Predict which hospital stays of diabetic patients will end in a readmission within 30 days, to decide who gets enrolled in a capacity-limited transitional-care programme. The data is Diabetes 130-US Hospitals (UCI #296, CC BY 4.0): 101,766 stays, 71,518 patients, up to 40 stays per patient, about 11% positives.
I chose it over two other medical datasets because it is full of real problems: repeated patients that force a group-aware split, missing values that aren’t random, ICD-9 categoricals with 700+ levels, and a target definition that has a flaw you have to find yourself.
What I built
- Two data problems, found and fixed. 2,752 stays end in death or hospice, so those patients cannot be readmitted. Keeping them teaches the model to recognise mortality and passes it off as predictive power. Separately, in two lab columns the string
'None'means “test not ordered”, which is a clinical decision and therefore a signal. pandas turns it intoNaNby default. Only'?'is truly missing. - Group-aware split (
StratifiedGroupKFoldon patient, 64/16/20), with anassert_disjointguardrail that raises if a patient lands in two partitions. - Every fitted step inside the Pipeline. No
dropna(): missingness gets an explicit category plus a*_recordedflag. - Metrics chosen for the decision. AUC-PR rather than accuracy, a decision threshold set by expected cost (assumed 10:1, with a sensitivity analysis from 1:1 to 50:1), and calibrated probabilities, because the ward needs to know how many patients to enrol.
- The protocol experiment. Comparing two protocols directly was the wrong design: they produce different test sets, and sampling noise (about 5 points of standard deviation) drowns a 2% effect. The final design is paired, and it adds a data-volume control arm for patient leakage.
- FastAPI service, a test suite (data, leakage, evaluation, model contract, API) and a Makefile that encodes the phases as a small DAG.
Results
Final model: HistGradientBoosting, evaluated once on the test set (19,867 stays from 13,951 patients never seen in training).
| Metric | Value |
|---|---|
| AUC-PR | 0.230 (95% CI 0.217–0.245), prevalence 0.114 → lift 2.02× |
| AUC-ROC | 0.671 (95% CI 0.659–0.683); the literature on this dataset reports 0.65–0.70 |
| Brier | 0.096 |

Following the 10% of patients with the highest risk catches 24.3% of readmissions, against 10% at random. Four out of five still go unflagged. Model families barely differ: about 10 thousandths of AUC-PR separate logistic regression from boosting in cross-validation.
The hypothesis was rejected. Patient leakage (net of data volume): −0.5% [−1.6, +0.6]. Preprocessing fitted before the split: +0.0% [−1.4, +1.4]. Evaluating on patients whose outcome was impossible: −8.2% [−10.1, −6.3], the only resolved effect, and in the opposite direction from what I expected. None of them comes close to the +4.7% gained by switching from logistic regression to boosting.
Limits
The signal is weak by nature: no variable separates the classes (|Cohen’s d| < 0.4). The model mostly identifies heavy users of the health system rather than the sickest patients. The cohort is 75% Caucasian, so confidence intervals on small groups are too wide to conclude anything about fairness. The data is US-only, 1999–2008, ICD-9. It is a prioritisation tool, not a medical device.