Tanner Ingalls · Oct 2, 2026

Churn Prediction: Finding Actionable Drivers

churn customer retention uplift modeling feature importance machine learning

A churn score tells you who might leave. It does not tell you what to do about it.

Most churn projects stop at the score. The list goes to customer success, and a CSM learns that one of her accounts has a 72% risk. Her next question is the right one: why, and what should I do? If the answer is "because they are on a monthly contract," the model was accurate and useless at the same time.

The value sits in three places: drivers your team can act on, a specific play for each, and proof that the play works.

Predictive is not the same as actionable

Some inputs predict churn well and give you nothing to do. New accounts churn more than year-three accounts, so tenure helps the model rank. Nobody can make a customer older. Monthly contracts churn more than annual ones, and you are not rewriting a contract mid-term because a model flagged it.

Other inputs are levers: onboarding completion, support friction (repeat tickets, slow resolution), usage decline against the account's own baseline, and price changes like a renewal increase or an expiring discount.

Keep both kinds in the model. The fixed inputs sharpen the ranking, so a new logo is compared with other new logos. But label every input fixed or actionable, and let only the actionable ones decide the play. (Building these inputs from CRM and product data: Feature Engineering for Tabular Business Data.)

Feature importance shows association, not cause

Churn models are often explained with SHAP, a method Scott Lundberg and Su-In Lee introduced at NeurIPS 2017. It assigns each feature an importance value for a particular prediction.[1] It is good at explaining why the model scored an account the way it did.

It does not tell you what happens if you change that input. An article in the SHAP documentation, written by Lundberg with colleagues at Microsoft, puts it plainly: SHAP makes the correlations picked up by predictive models transparent, but making correlations transparent does not make them causal.[2]

Their subscription example makes the point. In simulated data, customers who reported more bugs were more likely to renew, because the customers who need the product most use it more and report more bugs. Customers with bigger discounts were less likely to renew, because sales gave bigger discounts to accounts it thought were less interested. Taken at face value, the model says ship more bugs and cut discounts. In the simulation, reporting a bug had no causal effect on renewal, and discounts had a small positive one.[2]

Classic feature importance has the same limit. The scikit-learn documentation says permutation importance measures how much a model relies on a feature, and reflects the feature's importance to that particular model, not its intrinsic predictive value.[3] Reliance is not leverage. When important drivers go unmeasured, the SHAP article's authors conclude, randomized experiments remain the gold standard for finding causal effects.[2]

The model tells you where to look. Only a test tells you what works.

Tie every driver to a play and an owner

A driver is actionable only if someone owns the response. Illustrative examples:

  • Onboarding stalled (setup incomplete after 30 days): guided setup session, owned by the onboarding lead.
  • Support friction (three tickets on one issue in 60 days): escalation and root cause fix, owned by the support manager.
  • Usage decline (active seats down 30% from the 90-day average): usage review with the admin, owned by the CSM.
  • Price change (renewal within 90 days of an increase): value review before the quote, owned by the account manager, with finance setting discount limits.

Set thresholds from your own history. A driver with no play goes to the product or pricing backlog. A play with no owner will not happen.

Be careful with price plays. A Journal of Marketing Research field experiment with 65,000 wireless customers tested proactively recommending plans that would lower their costs. Over the next three months, 10% of the treated group churned, against 6% of the control group. The authors found support for two explanations: the outreach lowered customers' inertia and made past usage patterns more salient.[4] A sensible save play made churn worse, and only the control group showed it.

Measure the save play, not just the model

A model can be accurate while the retention program loses money. (Judging the model itself: Evaluating Models for Real Decisions.)

Nicholas Radcliffe and Patrick Surry, in a 2011 paper on uplift modeling, call it well-established best practice to measure a campaign's incremental impact against a randomly chosen control group. They warn that negative effects are not uncommon in retention, especially where interventions are intrusive and customers are already unhappy.[5]

The minimum version for a CS team: randomly hold out a slice of flagged accounts, give them the standard motion, and compare churn after 90 days. The difference is what the play saved. Without it, the play gets credit for customers who were never leaving. It is the holdout from Data Collection That Improves Models, pointed at the intervention.

Then you can go further. Uplift modeling is the set of techniques used to model the incremental impact of an action on a customer outcome.[6] It asks whose outcome changes if you act, not who will churn. Plan for volume: Radcliffe and Surry's rule of thumb is that the control group needs to be at least ten times larger for modeling than for simple measurement.[5]

The payoff can be large. Eva Ascarza ran two field experiments, with a wireless provider and a membership organization. The overlap between the highest-risk customers and the most responsive ones was about 50%, no better than chance. Targeting by sensitivity to the intervention instead of by risk would have cut churn by an additional 4.1 and 8.7 percentage points in the two studies. Her summary: across the two applications, half of the retention money is wasted.[7]

Prioritize by value at risk, not probability

Probability also ignores what an account is worth. An illustrative pair:

  • Account A: $8,000 a year at 60% risk. Value at risk: $4,800.
  • Account B: $90,000 a year at 20% risk. Value at risk: $18,000.

Sorted by probability, A comes first. Sorted by value at risk, B does, with almost four times the exposure. With fixed CS capacity, that ordering decides the week.

Aurélie Lemmens and Sunil Gupta, in Marketing Science, note that targeting on churn probability or responsiveness alone ignores that some customers contribute more to campaign profit than others. Their approach ranks customers by the intervention's incremental impact on churn and later cash flows, after its cost, and sets the campaign size. In two field experiments it produced significantly more profitable campaigns than competing models.[8]

The operator version is one line: annual value, times the churn reduction the play delivers for accounts like this one, minus the cost of running it. Rank by that. Stop where it turns negative.

What to ask for

  1. Every input labeled fixed or actionable.
  2. One play and one owner per actionable driver.
  3. Reason codes on each flagged account, from actionable drivers only.
  4. A random holdout on every save play, reviewed at 90 days.
  5. Ranking by value at risk, then by measured impact once holdout results exist.
  6. A stop rule for any play that does not beat its holdout.

If you want churn models that point to drivers your team can act on, with save plays measured against a holdout, talk to Gamify Data. We start with the history you already have and build around the decisions your CS team makes every week.

Sources

  1. Scott M. Lundberg and Su-In Lee, "A Unified Approach to Interpreting Model Predictions," Advances in Neural Information Processing Systems 30 (NeurIPS 2017), December 2017. https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html
  2. Eleanor Dillon, Jacob LaRiviere, Scott Lundberg, Jonathan Roth, and Vasilis Syrgkanis, "Be careful when interpreting predictive models in search of causal insights," SHAP documentation (originally published May 2021), accessed September 29, 2026. https://shap.readthedocs.io/en/latest/example_notebooks/overviews/Be%20careful%20when%20interpreting%20predictive%20models%20in%20search%20of%20causal%20insights.html
  3. scikit-learn developers, "Permutation feature importance," scikit-learn 1.9.1 documentation, accessed September 29, 2026. https://scikit-learn.org/stable/modules/permutation_importance.html
  4. Eva Ascarza, Raghuram Iyengar, and Martin Schleicher, "The Perils of Proactive Churn Prevention Using Plan Recommendations: Evidence from a Field Experiment," Journal of Marketing Research, vol. 53, no. 1, February 2016. https://doi.org/10.1509/jmr.13.0483
  5. Nicholas J. Radcliffe and Patrick D. Surry, "Real-World Uplift Modelling with Significance-Based Uplift Trees," Stochastic Solutions White Paper (Portrait Technical Report TR-2011-1), 2011. https://stochasticsolutions.com/pdf/sig-based-up-trees.pdf
  6. Pierre Gutierrez and Jean-Yves Gérardy, "Causal Inference and Uplift Modelling: A Review of the Literature," Proceedings of The 3rd International Conference on Predictive Applications and APIs, PMLR vol. 67, 2017. https://proceedings.mlr.press/v67/gutierrez17a.html
  7. Eva Ascarza, "Retention Futility: Targeting High-Risk Customers Might Be Ineffective," Journal of Marketing Research, vol. 55, no. 1, February 2018. https://doi.org/10.1509/jmr.16.0163
  8. Aurélie Lemmens and Sunil Gupta, "Managing Churn to Maximize Profits," Marketing Science, vol. 39, no. 5, 2020 (published online August 27, 2020). https://doi.org/10.1287/mksc.2020.1229

Ready to fit models to your data? Get in touch with Gamify Data.