About 3.5% of the transactions in the IEEE-CIS dataset are fraudulent. A model that approves everything is right 96.5% of the time and is completely worthless. That is the problem in one line: accuracy is not a target here, and most of the convenient ways to make a fraud model look good are also ways to fool yourself.
This was my MSc thesis at the University of Athens, over 590,540 labelled card transactions. I spent much more of it on the evaluation than on the models, and that turned out to be the right split.
Getting the evaluation right first
Transactions are sorted by TransactionDT, the first 80% used for training and the final 20% held out. A random split scores better and lies, because fraud patterns shift over time and a shuffled split lets the model see the future. The holdout never informed tuning, feature selection or threshold choice, and it was never downsampled, so it keeps the original imbalanced distribution. Class downsampling at 1:5 was applied to training data only.
Hyperparameter search ran through Optuna with expanding time-series cross-validation, entirely inside the training window. Every run went to MLflow with its parameters, metrics and artifacts, which is the thing that makes “which configuration produced this number” answerable three months later, when you have forgotten.
The confidence intervals changed the conclusion
Every metric carries a bootstrap confidence interval over 200 resamples. On the held-out set, with the original imbalanced distribution:
| Run | ROC AUC | 95% CI | Precision | Recall |
|---|---|---|---|---|
| LightGBM, full features | 0.9193 | [0.9147, 0.9242] | 0.3919 | 0.6471 |
| LightGBM, reduced | 0.9191 | [0.9141, 0.9236] | 0.4027 | 0.6368 |
| CatBoost, reduced | 0.9168 | [0.9118, 0.9212] | 0.2749 | 0.7311 |
| XGBoost, full features | 0.9069 | [0.9019, 0.9123] | 0.2960 | 0.6811 |
LightGBM ranked highest, and if I had reported point estimates alone I would have written a clean three-way ranking. The intervals do not support one. LightGBM at [0.9141, 0.9236] and CatBoost at [0.9118, 0.9212] overlap across most of their range, so on this data that gap is not distinguishable from noise. XGBoost’s interval does not overlap LightGBM’s, so that difference is real. Same numbers, a materially weaker claim, and the weaker claim is the true one.
This is the part I would keep if I could only keep one thing from the whole project. A leaderboard of point estimates is very easy to produce and quietly implies a precision the experiment does not have.
Feature reduction was close to free
SHAP values were computed per model, and a feature kept only if at least two of the three models ranked it in their top 30%. That cut the input space from 748 engineered features to 215, and moved LightGBM’s ROC AUC by 0.0002. Roughly seventy percent of the feature space was redundant as far as three different gradient boosting models were concerned, which says more about how easy it is to generate features than about how useful they are.
The threshold dominates everything
Same LightGBM model, same predictions, three operating points:
| Threshold | Precision | Recall | F1 |
|---|---|---|---|
| 0.5 (default) | 0.4027 | 0.6368 | 0.4934 |
| 0.1 | 0.1412 | 0.8583 | 0.2426 |
| 0.02 | 0.0662 | 0.9624 | 0.1238 |
Dropping the threshold from 0.5 to 0.02 takes recall from 64% to 96%, and takes precision to the point where the review team sees about fourteen false alarms for every fraud they catch. Nothing about the model changed between those rows. One number, chosen by a person, moves the behaviour of the system much further than picking a different gradient boosting library does.
So the model ranks risk and somebody else decides where to cut, based on review capacity and the relative cost of a missed fraud against a blocked customer. The evaluation code has a cost-optimal threshold search in it for that reason, but a real answer needs a real cost model, and I did not have one.
Serving a score you can interrogate
The inference service is FastAPI. It loads the model logged in MLflow and returns a fraud probability together with the SHAP contributions behind it, because a score nobody can question is not much use to a fraud analyst who has to justify blocking a customer. Prediction endpoints sit behind an API key, and there is an open endpoint serving example payloads so the demo can be driven without hunting for valid input.
What this is not
It is a controlled experimental study on a public benchmark, not a production fraud system. The gaps that matter most, in order:
The dataset is anonymised, so there is no real customer identifier. Behavioural aggregate features are anchored on a simulated user proxy built from card1, addr1 and an account-age proxy, which is a reasonable stand-in and definitely not the real thing. There is no probability calibration, so the scores rank well but should not be read as true probabilities. There is no cost model behind the threshold, no drift monitoring, and no evaluation over a long enough horizon to say anything about degradation. SHAP explains what the model attributed, not what caused anything. And the serving layer is a demo: no containerisation, no CI, no test suite, no data versioning.
Those are the next things worth building rather than caveats to wave away, and most of them are engineering problems rather than modelling ones.
Graded 10/10. Building the model was the smallest part of it.
Code, results and the full write up: github.com/koutsompinask/MSC-thesis