Asthma risk datasets are small and lopsided — far more "low risk" records than "high risk" ones.
That imbalance quietly sabotages machine learning models. A model can score well on paper while barely ever catching the high-risk cases that actually matter, because it learns to just predict the majority class most of the time. The usual fix is generating synthetic samples of the minority class (SMOTE and its variants), but how well those techniques actually work hadn't been carefully compared against each other on real patient data.
We evaluated five existing SMOTE-style oversampling techniques and proposed three new ones, testing all of them across 24 real asthma patient datasets from Soonchunhyang University Bucheon Hospital, South Korea.
We tested every method across four classifiers (Decision Tree, Logistic Regression, k-Nearest Neighbors, Naïve Bayes) plus a transfer learning framework, measuring weighted accuracy, sensitivity, precision, F1, and specificity.
SMOTEBoostCC, one of our proposed methods, came out on top overall — highest weighted accuracy (0.645) and sensitivity (0.545) among every method tested, meaningfully ahead of the untouched original data and standard SMOTE. The proposed methods held their own against established techniques, which was the core question we set out to answer.
This was my first real experience with the gap between "a model that scores well" and "a model that's actually useful" — in healthcare specifically, missing a high-risk case is a much bigger problem than a false alarm, so optimizing for the wrong metric has real consequences. It's also where I first got comfortable presenting technical findings to people who weren't in the room for the math, which turned out to matter far more in my career since than the SMOTE variants themselves.