AI
Catching Health Insurance Fraud with Machine Learning
How machine learning models can flag suspicious insurance claims automatically, and what it takes to wrap them in a system an insurer can actually use.
Rahi Bulbul6 min read
Every year, insurers pay out billions on claims that should never have been approved. Some of it is outright invention: bills for treatments that never happened. Some of it is subtler: inflated amounts, upcoded procedures, the same claim submitted twice. Whatever the flavour, the cost doesn't stay with the insurer. It comes back to policyholders as higher premiums and thinner coverage, and it quietly erodes trust in the whole system.
The obvious fix is to check the claims. The problem is that there are millions of them. No team of human investigators can review that volume, and the ones they do review take time and still get things wrong.
This post walks through a project that tried a different approach: training machine learning models to flag suspicious claims automatically, and wrapping them in a working system an insurer could actually use.
Why rule-based systems run out of road
Most fraud detection still works on hand-written rules. Flag anything over a certain amount. Flag this combination of procedure codes. Flag providers who bill more than their peers.
Rules are easy to explain and easy to audit, which is why they've lasted. But they have two failure modes that get worse over time. First, they're brittle: a fraudster who learns the threshold simply bills just under it. Second, they're noisy. Rules tend to catch a lot of legitimate claims that happen to look unusual, and every false positive costs an investigator's afternoon.
Rules also only know about fraud that someone has already seen and written down. They can't find a pattern nobody has described yet.
The data
The project used three linked datasets: beneficiary records, inpatient claims, and outpatient claims.
The beneficiary file held demographics and chronic condition flags: age, gender, whether the patient has diabetes, heart disease, osteoarthritis, and so on. The claims files held the transactional side: claim IDs, admission dates, diagnosis codes, procedure codes, and reimbursement amounts.
Merging inpatient and outpatient claims and joining them to beneficiary records on patient ID matters more than it sounds. A claim in isolation looks fine or it doesn't. A claim in the context of a patient's full treatment history is where the anomalies show up: the diagnosis that doesn't fit the patient's condition profile, the reimbursement that doesn't fit the diagnosis.
Pre-processing was the usual work: fill missing values, label-encode the categorical codes so the models could read them, and select a working feature set.
Which model won
Three candidates were tested: logistic regression, decision trees, and random forests.
Random forest came out ahead. That's not a surprising result, and it's a useful one. A random forest is an ensemble of decision trees, each trained on a slightly different slice of the data, voting on the outcome. That structure handles a lot of what makes claims data awkward: mixed numeric and categorical features, non-linear relationships, and interactions between variables that you'd never think to specify by hand.
It also gives you feature importances for free, which turned out to matter.
Three problems worth talking about
The interesting part of any fraud project isn't the model choice. It's what breaks.
Imbalance. Fraudulent claims are a tiny minority of all claims. A model that labels everything legitimate can hit a very impressive accuracy number while catching exactly zero fraud. The fix here was SMOTE (Synthetic Minority Over-sampling Technique), which generates synthetic fraudulent examples to balance the training set so the model has enough signal to learn from. Worth noting: SMOTE is applied to the training data only, after the split, or you've leaked information into your test set.
Too many features. The raw data had far more fields than were useful. Irrelevant features slow training down and give the model more opportunity to memorise noise. The approach was two-stage: drop features that domain knowledge said were unlikely to matter (patient birthdate, for instance), then use the random forest's own feature importance scores to see which of the survivors were actually carrying the prediction.
Overfitting. Early versions scored beautifully on training data and poorly on test data: the classic symptom. Cross-validation gave a more honest read on generalisation, and GridSearchCV tuned the hyperparameters that control model complexity: number of trees, maximum tree depth, minimum samples per split.
From model to system
A model in a notebook doesn't catch anyone. The project built the surrounding system:
- Front end in React.js, where insurers upload claim data, view flagged results, and generate reports.
- API layer in Flask, accepting claim data, running the same pre-processing the model was trained on, and returning a prediction as JSON.
- Storage in Firebase Realtime Database, holding claims, results, and audit history, with indexing so retrieval stayed fast as the dataset grew.
On security (and healthcare data leaves no room for shortcuts here), communication between client, API, and database was encrypted with SSL, and Firebase's authentication and access control rules restricted who could read or write claims data.
The detail that's easy to get wrong: the API has to apply exactly the same encoding transformations used during training. Mismatched pre-processing between training and inference is one of the most common ways a model that tested well fails in production.
What this doesn't do
Worth being clear about the boundaries. This is retrospective analysis on historical claims, not real-time scoring at the point of submission. It uses structured data only: the free-text clinical notes and provider reports, where a lot of fraud signal actually lives, are outside scope. And the models are trained on known fraud, which means they are better at recognising established patterns than inventing categories for new ones.
Where it goes next
Three directions stand out.
- Unstructured data. Natural language processing on clinical notes and claim descriptions would add a dimension that numbers alone cannot capture.
- Explainability. In fraud detection, “the model says so” is not an answer an investigator can act on, and it certainly is not one that survives a dispute. Explainable AI methods that show why a claim was flagged turn a score into something usable.
- Adaptation. Fraud changes. Detection has to change with it. Models need regular retraining on recent data, and ideally a feedback loop where confirmed and dismissed flags feed back into the next training cycle.
The takeaway
None of this replaces human investigators. What it does is change what they spend their day on: reviewing a ranked shortlist of genuinely suspicious claims instead of sampling blindly from millions. The random forest approach proved accurate enough to be useful and cheap enough to run at scale, and the surrounding system showed the whole thing can be packaged into something an insurance team could sit down in front of.
The infrastructure is the easy part. The hard part, as always, is the data, and the fact that the people you're trying to catch are learning too.
Work with NextByte
Ready to build something that works?
From AI and automation to full platforms, tell us what you're planning and we'll help you build it.
Start a project