Archived
Fraud Detection
MSc thesis on card fraud detection over 590,540 IEEE-CIS transactions: leakage-safe chronological evaluation, Optuna tuning, MLflow tracking, SHAP-driven feature reduction, and a FastAPI service that returns an explanation with every score. Graded 10/10.
Python LightGBM XGBoost CatBoost scikit-learn Optuna MLflow SHAP FastAPI
Fraud Detection is the ML pipeline built for my MSc thesis: a leakage-safe experimental study of
gradient boosting models on the IEEE-CIS dataset, where roughly 3.5% of transactions are fraudulent
and a model that approves everything scores 96.5% accuracy while being worthless.
Most of the work went into the evaluation rather than the modelling, because most of the convenient
ways to make a fraud model look good are also ways to fool yourself. Transactions are split
chronologically rather than at random, so the model never sees the future; the holdout is never
downsampled and never informs tuning, feature selection or threshold choice. Optuna runs inside the
training window with expanding time-series cross-validation, every run is tracked in MLflow, and
every metric carries a bootstrap confidence interval, which is what showed that the gap between the
two best models was not distinguishable from noise. SHAP agreement across the three models cut the
feature space from 748 to 215 at a cost of 0.0002 ROC AUC.
A FastAPI service loads the model logged in MLflow and returns a fraud probability together with the
SHAP contributions behind it, because a score nobody can interrogate is not much use to a fraud team.
The threshold, not the model, turned out to dominate the system’s behaviour.
This project sits at the intersection of the academic side of the MSc in Business Administration,
Data Analytics & Computer Systems and the applied ML work referenced throughout the rest of this
site. It carries the same instinct toward building something that runs, not just something
that scores well on a slide.