Home/FAQ/Analytics & AI/Machine Learning

Analytics & AI Machine Learning - FAQs

16 questions answered by the Hopton Analytics team.

Yes — machine learning can predict sales pipeline conversion, with enough historical data. Models trained on past won and lost opportunities can predict win probability based on deal characteristics: account size, deal value, stage progression speed, activity volume, the rep working the deal. The output is a per-opportunity score that supplements the human judgement of the sales rep. The score is a flag for review, not a decision; sales is too relationship-dependent for the model to act autonomously. We typically deploy these as monitoring agents on top of certified pipeline models.

Yes — you can use Python and standard data science tools throughout. Fabric Data Science notebooks support Python and R natively, with the standard libraries (pandas, scikit-learn, TensorFlow, PyTorch, statsmodels, Prophet) available. The development experience is familiar to data science teams who have worked in Jupyter or Databricks. The integration advantage is that the notebooks have direct access to OneLake data without separate connectors, and trained models deploy into the same Fabric environment without separate infrastructure.

Machine learning is one of the capabilities we deliver alongside Power BI, Microsoft Fabric, Azure Data Factory, and Business Central work. ML engagements span demand forecasting, customer segmentation, churn prediction, anomaly detection, and project margin forecasting (for construction clients particularly). The team includes data scientists and analytics engineers with practical experience deploying ML in mid-market production environments. ML is not the entirety of what we do, but it is a core capability and we deliver it regularly.

CLV predictions are variable in accuracy, by individual customer. Probabilistic CLV models have substantial uncertainty for individual customers because future behaviour is genuinely uncertain. The aggregate-level predictions (total CLV across a segment or cohort) are usually more accurate than any individual prediction because the individual errors cancel out. The honest framing is: do not over-rely on the individual customer estimate, but the aggregate is reliable enough to inform strategy. The Trust Storytelling Delivery Checklist applies here particularly; the model output should travel with limits and confidence framing.

We know if a machine learning model is good enough to deploy through specific evaluation against the use case it serves. Generic accuracy metrics (precision, recall, AUC) are necessary but not sufficient. The deeper question is whether the model performs in the segments that matter most: the high-value customers, the new product lines, the regions with different patterns. A model that is 92 per cent accurate overall but wrong on the top ten customers is not a deployable model. The Trust Storytelling Delivery Checklist (covered in our Governance FAQ) is the operational standard we apply to every ML output going into a decision.

ML model maintenance works by managing drift: models drift as the underlying data and business conditions change. The maintenance pattern includes regular performance monitoring against ground truth, scheduled retraining (typically monthly or quarterly depending on data volatility), and explicit version control so model changes are visible to stakeholders. Silent changes to models destroy trust faster than initial deployment failures. Every ML output going into a decision should travel with the Model Trust One-Pager (the seven-block framework in our AI/ML whitepaper) showing freshness, limits, and change history.

Customer churn prediction works through a classification model trained on historical customer data: who churned, who did not, and what was different about them. Common features include time since last purchase, change in purchase frequency, support ticket history, contract value, and engagement with marketing. The model produces a churn probability per customer, which feeds the retention team's prioritisation. The trust framework matters here particularly: false positives waste retention spend, false negatives lose revenue. The prediction is most useful as a flag for review, not an automated trigger.

Machine learning is used in finance and FP&A in several specific patterns. Cash flow forecasting using historical patterns and pipeline data. Anomaly detection in expense and supplier data for finance integrity work. Driver-based forecast modelling for FP&A. Working capital optimisation. The use cases pay back through faster, better forecasts and through finding issues that manual review misses. ML in finance benefits from the FP&A foundations being clean first; ML on top of inconsistent finance data produces confident-sounding rubbish, which is worse than no ML at all.

Machine learning is genuinely useful for mid-market businesses, where the data foundations support it. ML works particularly well for pattern recognition tasks (segmentation, churn prediction, anomaly detection), forecasting (demand, sales, cash flow), and classification (lead scoring, content tagging, risk categorisation). The use cases that pay back are the ones with sufficient historical data to learn from and a specific decision the prediction informs. ML works less well as a vague aspiration without a defined problem. The pattern is to identify a concrete, repeated decision that better information would change, then build the ML on top.

For a business new to machine learning, price and inventory optimisation is usually not the right starting point: we would generally recommend starting with core reporting and demand forecasting foundations first, and treating price and inventory optimisation as a next step once those are proven and trusted, rather than the first machine learning initiative attempted.

For most mid-market CLV implementations, use Fabric Data Science rather than Azure Machine Learning. The notebook environment is sufficient, the integration with the lakehouse is direct, and the operational overhead is low. Azure ML is the better choice when the CLV is part of a broader ML pipeline with sophisticated MLOps requirements, when the model retraining cadence is high, or when the model serving needs to be exposed as an API for real-time scoring. Most mid-market CLV use cases do not need Azure ML; the Fabric implementation is cleaner and cheaper.

Use Fabric Data Science for most mid-market ML work. The advantage is integration: the data lives in OneLake, the notebooks run against it directly, and the trained models can be deployed back into Fabric for inference. The setup is simpler and the operating cost is lower than running a separate Azure ML environment. Azure ML is the better choice for very high-volume training, complex MLOps requirements, or model serving patterns that need dedicated compute. For the typical mid-market use case (segmentation, forecasting, churn), Fabric Data Science is sufficient and cleaner.

Messy or fragmented historical data is a common starting point. The cleansing and consolidation work in the Silver layer of the lakehouse is part of every ML engagement. It is more work than the modelling itself; the modelling techniques are well-known, the data preparation is where the engineering effort concentrates. Treat the data foundation work as the precondition for ML rather than as something to skip. Clients who try to skip it discover the hard way that bad data plus ML produces worse outcomes than bad data alone, because the ML adds confidence to the wrong answers.

A previous machine learning investment that did not deliver is a common situation. Most failed ML investments are recoverable with the right framework. The honest assessment usually shows that the technology worked but the trust framework was missing, the data foundations were inadequate, or the use case was poorly defined. We have rebuilt failed forecasting models and abandoned classification work using the same underlying technology with the foundations corrected. The relaunch is usually faster and lighter than the original because the lessons are visible.

The difference between ML, AI, and predictive analytics is that ML and predictive analytics are largely the same thing under different names. AI is the broader category that includes ML, generative AI (large language models), and agent systems. Most of what people in mid-market call AI today is actually ML, and most of what they call predictive analytics is also ML. The distinctions matter for marketing more than for delivery. We use 'machine learning' to mean the statistical and pattern-recognition techniques that have been delivering value in analytics for decades, regardless of whether the marketing wraps them as AI.

Simple predictive CLV is good enough when you need a working number quickly, when the customer base behaves homogeneously, or when the audience for the CLV is more interested in the magnitude than the precision. A simple formula (average revenue per customer per year times average customer lifespan times margin percentage minus acquisition cost) produces a single CLV figure that informs acquisition spend decisions. The number is approximate but defensible. For most mid-market businesses early in their analytics journey, simple predictive CLV is the right starting point.

Still have questions?

Can’t find what you’re looking for?

The first conversation is exploratory and carries no obligation. We’ll give you an honest answer to any question you have.

Book a free audit