Production ML Without Massive New Infrastructure
Most companies do not need a new platform to get value from machine learning. They need one good model, scored on a schedule, showing up on a screen someone already opens every morning.
The common failure is not the math. Gartner surveyed 644 respondents from organizations in the U.S., Germany, and the U.K. in late 2023 and found that, on average, only 48% of AI projects make it into production, and it takes 8 months to go from AI prototype to production. The top barrier, cited by 49% of participants, was the difficulty of estimating and demonstrating the value of AI projects.[1]
Our view at Gamify Data: much of that time goes into machinery the first model did not need, while the value question waits.
The model is the small part
A 2015 NeurIPS paper by Google engineers put it in one diagram: only a small fraction of real-world ML systems is composed of the ML code, and the required surrounding infrastructure is vast and complex. A mature system might end up being (at most) 5% machine learning code and (at least) 95% glue code.[2] Their opening line is the one to keep: developing and deploying ML systems is relatively fast and cheap, but maintaining them over time is difficult and expensive.[2]
Every new component is something you pay to maintain. So the first question is not "which platform do we need?" It is "what is the least new machinery that puts a score in front of the decision maker?"
Start where the data already lives
Churn and lead models draw on the CRM (accounts, opportunities, activity), billing (plans, invoices, renewal dates), and product usage. In many companies those already land in a warehouse for reporting.
The major warehouses can train models in place. Google's documentation says BigQuery ML lets you create and run machine learning models by using GoogleSQL queries, and that it speeds up model development by removing the need to move data from the data warehouse. Its model types include logistic regression and boosted trees.[3] Snowflake's classification function lists customer churn prediction as a common use case.[4]
Neither product is required; a scheduled Python job that reads from the warehouse works too. What matters is that training data, input definitions, and scores sit in the same governed place as the reporting finance already trusts. Fewer copies means fewer places for "active customer" to mean two different things. (Building those inputs: Feature Engineering for Tabular Business Data.)
Score on a schedule, not in real time
Real-time serving means the model runs behind an API and answers each request in milliseconds. Batch scoring means it runs over every account or lead on a schedule and saves the results in a table.
Martin Zinkevich's Rules of Machine Learning, published by Google, treats this as a basic choice: apply the model live, or pre-compute the model on examples offline and store the results in a table.[5] Google Cloud's documentation shows the cost difference. Online inference requires deploying the model to an endpoint, which associates compute resources with the model. Batch inference is for when you don't require an immediate response and want to process accumulated data by using a single request.[6]
Match the cadence to the decision. A CSM reviews her book weekly, so a nightly churn score is plenty. Reps work their queue each morning, so a nightly lead score covers most of the value. Real-time earns its cost when the decision happens in the moment, like routing an inbound demo request while the prospect is still on the page. Even the scheduler can live in the warehouse: BigQuery lets you schedule queries to run on a recurring basis.[7]
If the decision is made weekly, a model that answers in milliseconds is paying for speed nobody uses.
Put the score where the work happens
A score in a dashboard nobody opens is not a score. Write it back to the account or lead record in the CRM as a few plain fields: score, top reason, recommended play, date scored. Then use the CRM's own list views and alerts. Nobody learns a new tool.
The plumbing is mature. Salesforce's Bulk API 2.0 is built to insert, update, upsert, or delete many records asynchronously, and Salesforce says any data operation that includes more than 2,000 records is a good candidate for it.[8] A nightly update of a few thousand scores is a routine job, not a platform project.
In our experience, write-back is where adoption is won or lost. For churn, show reason codes from actionable drivers only (Churn Prediction: Finding Actionable Drivers). For leads, show the top two reasons so a rep can see why a lead jumped the queue.
Small, well-understood models first
The Rules of Machine Learning are blunt. Rule #4: keep the first model simple and get the infrastructure right, because the first model provides the biggest boost and doesn't need to be fancy. Rule #14: starting with an interpretable model makes debugging easier.[5]
On CRM and billing data, simple does not mean weak. Leo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux compared deep learning with tree-based models across 45 datasets at NeurIPS 2022. Tree-based models remain state-of-the-art on medium-sized data (around 10,000 samples) even without accounting for their superior speed.[9]
Our usual order: a logistic regression baseline to set the bar, then gradient boosted trees if they beat it on recent periods the model never saw. If nothing beats a simple rule, keep the rule. (Running that comparison: Evaluating Models for Real Decisions.)
Keep monitoring simple, but do it
The NeurIPS paper suggests a check it calls a surprisingly useful diagnostic: in a system working as intended, the distribution of predicted labels should usually equal the distribution of observed labels.[2] If the model says 8% of accounts will churn this quarter, roughly 8% should. The Rules of Machine Learning add a warning about silent failures, like a feature populated in 90% of examples that suddenly drops to 60%.[5]
For a scheduled churn or lead model, we start with four checks:
- Did it run? Records scored today against yesterday.
- Are the inputs filled? Missing rate per input against its usual level.
- Does it match reality? Predicted against actual churn or conversion each month, by segment.
- Did the scores shift? Share of accounts in each risk band, week over week.
Each is a query on tables you already have, with an alert someone owns. (Why models drift after launch: Why Forecasting Models Degrade.)
Know when it is time to invest
This is not an argument against infrastructure. It is an argument for buying it when you need it. Google Cloud's MLOps guidance calls a manual, data-scientist-driven process "level 0," says it is common in many businesses that are beginning to apply ML, and says it might be sufficient when models are rarely changed or trained. It also cautions that, in practice, models often break when they are deployed in the real world.[10]
Our rule of thumb is to invest when one of these is true:
- The decision happens in seconds, like inbound routing.
- Manual retraining has become a job the team cannot keep up with.
- Accuracy decays between retrains, because the market moves faster than your refresh cadence.
- A holdout has proven the value, so the spend scales something that works.
What this looks like for churn
An illustrative setup: CRM, billing, and usage sync to the warehouse nightly. A boosted trees model trains there monthly. A 2 a.m. job scores every active account and writes churn risk, value at risk, top driver, and scored date to the CRM. CSMs work a list sorted by value at risk. Lead scoring is the same pattern with a rep queue. Total new infrastructure: one scheduled job and one write-back integration.
What to ask for
- An input inventory showing every source is already in the warehouse or CRM.
- A decision cadence for each model, and a scoring schedule that matches it.
- Write-back fields designed with the reps and CSMs who will use them.
- A simple baseline the production model has to beat.
- Four monitoring checks with a named owner.
- A written trigger for when more infrastructure is justified.
If you want churn and lead models that run on the systems you already have and show up where your team already works, talk to Gamify Data. We start with your warehouse and CRM and add only what the decision needs.
Sources
- Gartner, "Gartner Survey Finds Generative AI Is Now the Most Frequently Deployed AI Solution in Organizations," press release, May 7, 2024. https://www.gartner.com/en/newsroom/press-releases/2024-05-07-gartner-survey-finds-generative-ai-is-now-the-most-frequently-deployed-ai-solution-in-organizations
- D. Sculley et al., "Hidden Technical Debt in Machine Learning Systems," Advances in Neural Information Processing Systems 28 (NeurIPS 2015), December 2015. https://proceedings.neurips.cc/paper_files/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf
- Google Cloud, "Introduction to ML in BigQuery," BigQuery documentation, last updated October 5, 2026, accessed October 6, 2026. https://docs.cloud.google.com/bigquery/docs/bqml-introduction
- Snowflake, "Classification (Snowflake ML Functions)," Snowflake documentation, accessed October 6, 2026. https://docs.snowflake.com/en/user-guide/ml-functions/classification
- Martin Zinkevich, "Rules of Machine Learning: Best Practices for ML Engineering" (Rules #4, #10, and #14), Google for Developers, last updated August 25, 2025. https://developers.google.com/machine-learning/guides/rules-of-ml
- Google Cloud, "Overview of getting inferences on Agent Platform," Gemini Enterprise Agent Platform documentation, last updated October 6, 2026, accessed October 6, 2026. https://docs.cloud.google.com/gemini-enterprise-agent-platform/machine-learning/predictions
- Google Cloud, "Scheduling queries," BigQuery documentation, last updated October 6, 2026, accessed October 6, 2026. https://docs.cloud.google.com/bigquery/docs/scheduling-queries
- Salesforce, "Introduction to Bulk API 2.0 and Bulk API," Bulk API 2.0 and Bulk API Developer Guide, accessed October 6, 2026. https://developer.salesforce.com/docs/atlas.en-us.api_asynch.meta/api_asynch/asynch_api_intro.htm
- Leo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux, "Why do tree-based models still outperform deep learning on typical tabular data?," Advances in Neural Information Processing Systems 35 (NeurIPS 2022), Datasets and Benchmarks Track, December 2022. https://proceedings.neurips.cc/paper_files/paper/2022/hash/0378c7692da36807bdec87ab043cdadc-Abstract-Datasets_and_Benchmarks.html
- Google Cloud, "MLOps: Continuous delivery and automation pipelines in machine learning," Cloud Architecture Center, last updated August 28, 2024. https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning
Ready to fit models to your data? Get in touch with Gamify Data.