← Back to portfolio

Why I optimized for recall, not F1, on an imbalanced education dataset

Nafe Abubaker · Capstone retrospective

For my capstone I built a model to predict whether a US high school student would drop out, using the HSLS:09 longitudinal study — about 15,900 students tracked from 9th grade through 2016. The class split was roughly 5.5 to 1: most students don't drop out, which is good for the kids and bad for the model.

The obvious trap with imbalanced data is that your baseline looks fine on accuracy and quietly fails at the one thing that matters. My first pass — plain Logistic Regression, Random Forest, XGBoost, LightGBM, all at the default 0.50 threshold — hit ROC-AUC around 0.78, which sounds respectable, and recall around 0.14, which is a disaster. The model was catching one out of every seven at-risk students and confidently telling everyone else they were fine.

So I had to pick a metric that matched the actual use case. An early-warning system for school counselors isn't a precision problem. If the model flags a kid who turns out to be fine, a counselor has a slightly awkward check-in conversation. If the model misses a kid who drops out, nobody intervenes and a real person's life gets worse. False negatives are the expensive error. F1 treats precision and recall as equally important, and in this setting they aren't — not even close.

Two changes moved the needle. First, class weighting. Just adding class_weight='balanced' to the baseline Logistic Regression pushed recall from 0.14 to 0.73 on the same features. That one flag was worth more than every model I tried afterward. The default threshold on imbalanced data was the actual bottleneck, not the model family or the features.

Second, threshold tuning. On the tuned XGBoost, lowering the decision threshold from 0.50 to 0.30 pushed recall from 0.71 to 0.87 — catching 423 out of 489 at-risk students in the test set. Precision dropped to 0.24, which means roughly three out of four flagged students aren't actually at risk. In a normal classification problem that would be embarrassing. For a counselor with a caseload and a list of "kids to check in on," it's fine. You'd rather over-flag and triage than under-flag and lose people.

The lesson I'd give my past self: pick the metric before you pick the model. If I had opened the project by writing down "we care about recall because false negatives are kids we don't help," I would have skipped two weeks of tuning F1 on baselines that were never going to work. The model is downstream of what you're actually trying to do, and "what you're actually trying to do" is almost never "maximize F1."

Next iteration I'd spend real time on the 3,000+ raw columns in HSLS:09 before modeling. Most of the lift in this project came from weighting and thresholding, which suggests the feature set was leaving signal on the table.