← Projects

healthcare risk · payer analytics

MEPS Healthcare Risk Stratification

Prospective models that rank people by next-year medical cost and prescription burden, built for limited outreach capacity and audited by subgroup.

RtidymodelsXGBoostRandom ForestSurvey Weights
0.883ROC-AUC, next-year prescription burden
39.3%of high-cost members captured at 10% outreach
74.6%precision at that same 10% capacity

The question

Can prior-year survey, utilization, expenditure, insurance, access, chronic-condition and prescription information identify people at elevated risk of high medical spending or high prescription out-of-pocket burden in the following year?

The data

Medical Expenditure Panel Survey (MEPS) Panel 24 longitudinal records and Prescribed Medicines event files. Features from 2021 predict outcomes in 2022 for the same people, which avoids the temporal leakage in a same-year baseline. The weighted targets are next-year medical expenditure at or above the 90th percentile (~$15,256) and next-year prescription self/family payment at or above the 90th percentile (~$517).

How it was done

  1. Preserve MEPS survey weights in descriptive estimates and model evaluation.
  2. Compare logistic regression, random forest, and gradient boosting.
  3. Evaluate discrimination, calibration, precision, recall, and lift at fixed outreach capacity.
  4. Audit performance by age, income, insurance, race or ethnicity, and chronic-condition burden.
  5. Treat thresholds as resource-allocation choices rather than defaulting to 0.50.

What the evidence showed

Model performance on prospective targets
TargetBest modelROC-AUCPR-AUCPrecisionRecall
Next-year high medical expenditureRandom forest0.8350.6110.4640.667
Next-year high prescription burdenRandom forest0.8830.6160.4900.740
Value of ranking when outreach capacity is constrained
TargetOutreach shareCapture ratePrecision
High medical expenditure10%39.3%74.6%
High medical expenditure20%57.6%55.2%
High prescription burden10%47.5%66.4%
High prescription burden20%71.9%52.2%

What I recommend

What this does not prove