Use Fabric Data Factory for new builds in most cases, rather than Azure Data Factory (ADF). The capabilities are similar (and growing closer) but the Fabric version integrates natively with OneLake and the rest of the Fabric platform without separate connection management. For existing ADF investments, the migration path to Fabric Data Factory is straightforward but not always urgent; existing ADF pipelines continue to work and can stay on standalone ADF for the foreseeable future. For new mid-market analytics builds, the default is Fabric Data Factory.
Yes — Microsoft Fabric and Databricks are increasingly competing for the same workloads. Both platforms can do most things: lakehouse, data warehouse, BI, ML, real-time. The competition is real. The differences are in optimisation and culture rather than capability. Fabric is built around BI-led mixed workloads. Databricks is built around code-first lakehouse and ML. Both can do the other thing, but each is designed for a different centre of gravity, and that is what shapes the decision.
Yes, for the workloads most mid-market businesses actually run. Both are credible cloud data platforms, both use a lakehouse-style architecture, and both bill on consumption rather than fixed licence tiers. The differences are in ecosystem fit and specialisation rather than raw capability. Fabric's advantage is native integration with Power BI, Azure, and M365, plus OneLake as a single storage layer. Snowflake's advantage is genuine multi-cloud portability across AWS, Azure, and GCP, and a data-sharing model some organisations rely on for external data collaboration. Neither has a decisive edge in AI engineering specialisation over the other today. The decision is rarely about which is technically better. It is usually about which fits your team, your existing estate, and whether multi-cloud is a real requirement or a hypothetical one.
Yes — both platforms have non-obvious costs worth knowing. Fabric capacity sizing is its own learning curve, and over-provisioning is common in early Fabric deployments. Databricks workspaces can sprawl, with notebook clusters left running and unmanaged. Both vendors have monitoring tools. Use them from day one rather than discovering the bill at month-end.
If you scored nine or above, you are likely ready to consider Fabric. The Building stage is where structured platform investment usually starts paying back. Below nine, fix the foundations first. A new platform amplifies the problems, not solves them. The Fabric Adoption Playbook covers the platform decision in detail; the assessment tells you whether to start that conversation now or later.
Azure Data Factory (and Synapse) suits teams with an established Azure estate and heavy, cost-sensitive orchestration that want granular control. Microsoft Fabric suits teams who want ingestion, storage, transformation and Power BI unified in one governed SaaS platform and value simplicity. For most mid-market organisations, Fabric’s consolidation is the better fit; the mistake to avoid is buying a large multi-cloud platform you do not need.
Use Azure SQL Database for most use cases, and Azure SQL Managed Instance for specific compatibility requirements. Azure SQL Database is the modern, fully-managed PaaS database with the lowest operational overhead. Managed Instance offers near-100 per cent compatibility with on-premises SQL Server, which matters for migrations from existing SQL Server environments where breaking changes would be costly. For new mid-market data platforms, Azure SQL Database is usually the default. For lift-and-shift migrations of existing SQL Server estates, Managed Instance may be the easier path.
CDC (change data capture) and event streaming are different patterns serving different purposes. CDC captures changes from a transactional database and streams them to the analytical platform; the source is a database. Event streaming captures discrete events from applications, devices, or APIs; the source is an event producer. CDC is appropriate for analytical platforms that need to track changes in transactional data. Event streaming is appropriate for systems that produce events natively. Many real-time implementations use both, with CDC bringing in operational system changes alongside event streams from applications and devices.
Yes — Fabric can coexist with existing Power BI and Azure investments. Fabric does not require ripping out existing Power BI investments; it extends them. Existing Power BI Pro licences continue to work, existing reports continue to run, and the migration path from Power BI Premium to Fabric capacity is well-documented. Existing Azure data investments (ADF, Synapse, Azure SQL, Storage) can be progressively migrated to Fabric or kept where they are with Fabric reading from them. The transition is incremental rather than a forklift replacement, which suits mid-market budgets and risk appetite.
Yes — Fabric can do ML. Fabric's Data Science workload covers the end-to-end ML workflow in one environment: Spark-based notebooks for data ingestion and experimentation, AutoML for automated model selection and tuning, built-in MLflow for experiment tracking and model registry, and direct deployment of model output into Power BI reports. Real-time intelligence and Copilot-based conversational analytics agents sit alongside this, so a business user can query a model's output in natural language rather than writing code. The tools are credible. The ecosystem maturity is younger than Databricks. For most mid-market ML use cases (forecasting, classification, basic recommendations), Fabric is enough. For deep ML workflows with serious MLOps requirements, Databricks remains stronger.
Fabric usually cannot replace your planning system. Planning systems handle the workflow (scenario building, approval cycles, exception management, S&OP processes) that Fabric does not. Fabric handles the analytical layer (statistical forecasting, accuracy tracking, multi-source data integration) that the planning systems often handle weakly. The two work together. For very small operations a Fabric-only approach might work; for serious supply chain operations the planning system stays in place and Fabric augments it.
Yes — Hopton can build the architecture without delivering the reports. Some clients engage us for the data engineering and architecture workstream specifically, with the reporting and visualisation handled by their internal team or another partner. The architectural deliverables (lakehouse structure, ingestion pipelines, semantic model design, deployment pipelines) are independent of the report build. The engagement scope is flexible. We are equally willing to deliver the full stack or a specific layer.
Yes — Hopton can design your Azure data architecture. Architecture design is part of every Establish phase. The output is a written architecture document covering the Azure services to be used, how they fit together, the cost model, the security posture, the operational model, and the migration path from existing systems. The architecture is the foundation of the build phase that follows, but the architecture document is also useful as a standalone deliverable. Some clients use us for architecture only and deliver the build with internal teams or other partners.
Yes — Hopton can help you decide whether Fabric is right for your business. The Reporting Modernisation Assessment (a four-week fixed-price engagement) is the structured way to get an independent view. Output is a written assessment with a clear recommendation: Fabric, alternatives, or a phased approach. We have written assessments recommending against Fabric where the business case did not support it, and we have written several recommending it. The framework decides, not the commercial bias toward selling more work.
Yes — Hopton can review your existing data architecture. The Power BI Health Check (a two-week fixed-price engagement) covers the data architecture alongside the Power BI estate. For deeper architectural reviews on existing Fabric or Azure data platforms, we run extended assessments (typically four to six weeks) covering medallion structure, ingestion patterns, SCD handling, deployment maturity, and observability. Output is a written report with prioritised recommendations and a roadmap. The assessment is independent of any subsequent build engagement.
Yes — market basket analysis can inform promotional design, in two patterns. First, promote anchor products (those with strong outgoing associations) to drive baskets, accepting the margin sacrifice on the anchor for the basket-building uplift on associated products. Second, avoid promoting strongly co-purchased products simultaneously: if customers already buy A and B together, promoting both is unnecessary spending. The MBA-informed promotional plan is usually meaningfully different from the intuition-based plan, and the difference is measurable in basket-level uplift. We have run promotional optimisation work for retailers using MBA as the foundation.
Yes — anomaly detection can support customer behaviour monitoring. Pattern detection on customer behaviour can surface accounts at risk of churn (sudden behaviour change), accounts at risk of fraud or compromise (unusual access patterns), and accounts that may be growing through organisational change (revenue patterns shifting). The challenge is distinguishing meaningful customer-level anomalies from natural variation, which requires careful baseline construction. Per-customer baselines work better than population baselines for established customer relationships.
Real-time can sometimes replace daily reporting, for the right use cases. The high-frequency operational dashboards benefit from real-time. The strategic management reporting (monthly P&L, quarterly board pack) does not. The right pattern is a layered approach: real-time dashboards for operational decisions, refreshed-hourly dashboards for tactical management, and traditional batch reporting for strategic and statutory needs. Trying to put everything on real-time produces an expensive and unnecessary platform; trying to use only batch produces lost operational opportunities. Both have a place.
Whether the forecast can run autonomously or needs human oversight depends on the use case. For high-volume, low-stakes decisions (routine replenishment of stable products), autonomous operation with monitoring is appropriate. For high-stakes decisions (new product launches, promotional commitments, capacity decisions), human oversight is essential. The pattern that works is the forecast as input to a human-in-the-loop process for important decisions, and as direct driver for routine decisions, with monitoring agents flagging anomalies for review. The design of the operational workflow matters as much as the forecasting itself.
Yes — you can move from Fabric to Snowflake later, or vice versa, with effort. Open table formats make data portability easier than it was. The harder parts are semantic models, BI tooling, security and the operational tooling around the platform. A migration is a real project, not a button-push, but it is no longer a one-way door. Choose well now and you will rarely need to migrate. Choose poorly and migration is a recoverable mistake.
Yes — you can reduce Fabric cost materially through reservations. Annual reservations save around 20 per cent off PAYG list prices. Three-year commitments through an Enterprise Agreement save up to 41 per cent. For workloads that only run during business hours, pausing capacity outside hours saves another 50 per cent on top. Combined, these can reduce Fabric capacity costs by 60 per cent or more against PAYG-without-pausing for typical mid-market patterns. The specific savings depend on your workload pattern; we cover sizing methodology in the True Cost FAQ.
Yes — you can run a quick MBA proof of concept. We extract a sample of transaction data, run MBA in Fabric Data Science, produce the rule output, and walk you through the results. The proof of concept produces real rules on your real transaction data, not a generic demonstration. Often the rules surfaced in the proof of concept directly inform commercial decisions, even before the full implementation. Two to three weeks is the typical duration.
Yes — you can run an anomaly detection proof of concept. A three to four week proof of concept on a representative data slice produces a working detector, validates the technique choice, and demonstrates the operational signal-to-noise ratio. The proof of concept output is often usable directly for some monitoring use cases, with the full implementation extending the coverage. Several of our anomaly detection engagements started as proofs of concept and grew into permanent monitoring capabilities.
Yes — you can use both platforms together, and some organisations do. Fabric for BI and analyst-led work, Databricks for engineering and ML. The integration is workable through OneLake shortcuts and shared open table formats. The cost is operational complexity. Most mid-market businesses find that one platform serves them better than two. Larger organisations with distinct teams sometimes choose both deliberately.
On open table formats, both Fabric and Databricks support Delta. Databricks supports Iceberg natively. Fabric currently leans on Delta. Open formats reduce vendor lock-in and make migration easier. They do not decide the platform choice on their own. The interoperability story is improving fast on both sides.
For most serious implementations, yes. The capability that Purview provides (sensitivity labels, classification, lineage tracking, access governance) is genuinely valuable and increasingly expected by stakeholders, auditors, and regulators. Smaller implementations can get by with less; once the data estate spans multiple sources and the user base is broader than a small team, governance discipline supported by tooling becomes important. The cost of Purview is reasonable; the cost of governance failure is much higher.
Both KQL and SQL work, with different strengths. KQL is the native language for eventhouses and Real-Time Intelligence; the queries run faster and the language better fits the time-series patterns. SQL is supported through cross-Fabric queries and is more familiar for most analysts. For exploratory work and ad-hoc queries, KQL fluency makes a meaningful difference in productivity. For dashboards and reports built on top, the language choice matters less because the queries are written once and consumed many times. Most teams use a mix.
No — we do not work on Azure outside of data and analytics. Our scope is the Microsoft data and analytics stack. Azure as a broader cloud platform (compute, networking, identity, security, application services) is outside our specialism. We work with Azure data services because they are part of the data and analytics stack, but we do not deliver general Azure infrastructure work, application modernisation, or cloud migration outside the data context. For broader Azure work, we coordinate with specialist Azure partners.
For now Databricks does have better ML and AI tooling, on the ML side. MLflow, Feature Store, Unity Catalog and the wider MLOps ecosystem are more mature in Databricks than in Fabric. Fabric is closing the gap on the data science workload, but mature MLOps is a multi-year build and Databricks has been at it longer. For organisations with serious ML ambitions, this is a meaningful advantage.
Yes — Hopton does implement real-time analytics. Real-time and streaming analytics is part of our delivery scope when client use cases justify it. We have built real-time operational dashboards, supply chain visibility, and event-driven monitoring on Fabric Real-Time Intelligence. The team holds the relevant Microsoft certifications and has practical experience with the workload. Real-time is not the bulk of our work (most mid-market analytics is batch-led) but it is a capability we deliver when the use case is right.
Azure data services are at the centre of our work alongside Power BI, Microsoft Fabric, and Microsoft Dynamics 365 Business Central. Most of our engagements involve Azure Data Factory, Azure SQL, Azure Storage, and Microsoft Purview at minimum. Many include Azure Machine Learning, Azure OpenAI, or Azure-hosted infrastructure. The team holds individual Microsoft certifications across the Azure data stack. Azure is not a separate specialism for us; it is part of the Microsoft data and analytics stack we work in daily.
Yes — Hopton does specialise in data engineering and architecture. Data engineering and architecture is at the centre of our work. Most of our engagements include the architectural design and implementation work alongside the analytical and reporting layer. The team includes data architects and engineers with practical experience building Fabric and Azure data platforms across mid-market sectors. Several team members hold individual Microsoft certifications including the DP-700 (Fabric Data Engineer) and the broader Azure data certifications. Architecture is not a sideline for us.
Hopton works with the broader Microsoft stack, not Fabric only. Fabric is one of our core technologies alongside Power BI, Azure Data Factory, Azure SQL, Azure Machine Learning, Microsoft Purview, and Microsoft Dynamics 365 Business Central. The strength is in knowing how these components fit together. Many of our engagements span Fabric and other Azure or Microsoft 365 services. Specialising in Fabric without specialising in the surrounding Microsoft estate would limit the value we deliver.
Yes — market basket analysis works for non-retail businesses, with adjustments. The technique applies wherever transactions contain multiple items: financial products bought together, professional services purchased in combination, content consumed in sequence, courses taken in combination, pharmaceutical prescriptions in patient histories. The retail roots are visible in the terminology (basket, items) but the underlying maths is general. The data shape is the same: a transaction identifier and a set of items per transaction. Where the data fits the shape, MBA produces useful rules.
Not every organisation needs all three layers as separate physical layers on day one, but the three responsibilities of keeping raw data untouched, validating and conforming it, and modelling it for reporting exist in every platform whether they are named or not. Smaller estates sometimes combine Silver and Gold early on, but the moment reporting logic starts living in more than one place, splitting them out usually pays for itself.
Demand forecast accuracy is variable, by item and by horizon. For stable mature SKUs at weekly horizon, forecast errors of 10 to 20 per cent (mean absolute percentage error) are typical and good. For volatile or low-volume SKUs, errors of 30 to 50 per cent are common. For new products, forecast errors of 50 to 100 per cent or more are realistic. Aggregate accuracy is always better than SKU-level accuracy because errors cancel out. Setting realistic accuracy expectations is part of the early engagement; over-promising on accuracy produces disappointed stakeholders.
Azure ML, Azure OpenAI, and Copilot serve different layers of AI capability. Azure ML is for traditional machine learning (predictions, classifications, forecasting) on your own data. Azure OpenAI is for generative AI through direct API access for custom applications. Copilot products are the user-facing AI assistants embedded in Microsoft applications, built on top of OpenAI models with Microsoft's product-specific tuning. Most mid-market businesses end up using Copilot for the user-facing AI, Fabric Data Science for traditional ML, and Azure OpenAI selectively for custom generative AI applications. The choice depends on the specific use case.
Snowflake has had a longer head start on cross-organisation data sharing through Snowflake Marketplace and direct sharing. Fabric is catching up through OneLake shortcuts and external sharing. If multi-organisation data sharing is core to your business model, this currently favours Snowflake. For most mid-market businesses, it is a non-issue.
On open table formats, both Fabric and Snowflake support Delta Lake. Snowflake also supports Iceberg natively. Fabric uses Delta as the storage format underneath the Lakehouse component. Open formats reduce vendor lock-in and make some migration scenarios easier, but they do not decide the platform choice on their own. The decision is mostly about platform fit, not file format.
Email hello@hoptonanalytics.com with a brief description of your current Azure estate, what you are trying to achieve, and any existing investments. The first conversation is exploratory and free. If there is a fit, we propose a four-week Establish phase that produces an architecture, a priority list, and a written delivery plan. From there you can decide whether to proceed with the build.
Email hello@hoptonanalytics.com with a brief description of your current data and analytics estate, your main pain points, and what you are trying to improve. The first conversation is exploratory and free. If there is a fit, we propose a four-week Establish phase that produces an architecture, a priority list, and a written delivery plan. From there you can decide whether to proceed with the build.
Email hello@hoptonanalytics.com with a brief description of the use case (what kind of unusual behaviour are you trying to catch), the data sources, and the operational team that would respond to the alerts. The first conversation is exploratory and free. If there is a fit, we propose either a focused anomaly detection proof of concept or a broader analytics engagement that includes anomaly detection as one component.
Email hello@hoptonanalytics.com with a brief description of your current data estate (sources, current platform, scale), the architectural challenges you are facing, and what you are trying to achieve. The first conversation is exploratory and free. If there is a fit, we propose either a Health Check (lighter-touch review) or an architectural assessment (deeper review) as the structured next step.
Email hello@hoptonanalytics.com with a brief description of your supply chain operation, the products and SKUs you forecast, the planning systems already in place, and the specific accuracy or operational issues you are trying to solve. The first conversation is exploratory and free. If there is a fit, we propose either a four-week forecasting proof of concept or a broader supply chain analytics engagement that includes forecasting as one component.
Email hello@hoptonanalytics.com with a brief description of your transaction data, your basket sizes, and what commercial decisions you are looking to inform. The first conversation is exploratory and free. If there is a fit, we propose either a focused MBA implementation or a broader customer and product analytics engagement that includes MBA as one component.
To start a conversation about real-time analytics, email us at hello@hoptonanalytics.com with a brief description of the use case (what decision needs to be made faster, what events need to trigger action, what data sources are streaming-capable), and the current state of your data platform. The first conversation is exploratory and free. If there is a fit, we propose either a focused real-time scoping engagement or a broader analytics engagement that includes real-time alongside batch.
MBA rules inform store layout by revealing which products customers naturally buy together, so layouts can either encourage the association (placing the products near each other for convenience and uplift) or deliberately separate them (placing products at opposite ends of the store to extend the shopping journey). Both strategies have been used effectively. The decision depends on the business: convenience-led retailers tend to group; destination retailers tend to separate. The rules surface the choice; the merchandising team makes the call. Most mid-market retailers do not formally use MBA for layout decisions and many would benefit from doing so.
Basket size and product range affect market basket analysis through two practical considerations. Basket size: MBA needs baskets with multiple items to find associations. Single-item baskets contribute nothing. Retailers with average basket size of 1.0 to 1.5 items get less from MBA than retailers with baskets of 3+ items. Product range: very large ranges (tens of thousands of SKUs) produce sparse data where most pairs occur infrequently. Aggregating to category or sub-category level often produces more useful rules than analysis at SKU level. The right granularity is a business decision, usually middle-out.
Training and operability costs factor in heavily, and most pricing comparisons miss this. A platform your team can run with existing skills costs less in practice than one that requires hiring. Fabric tends to win on this for organisations with Power BI developers. Snowflake tends to win for organisations with strong SQL engineering teams. The platform you can run is almost always cheaper than the platform you have to staff up to run.
Training and team costs are heavily underweighted in most BI platform pricing comparisons. The platform your team can run with existing skills costs less in practice than one requiring hiring. Fabric tends to win for organisations with Power BI developers and analyst skillsets. Databricks tends to win for organisations with Python and Spark engineering capability. Hire the platform you have a team for, or have a plan to staff the one you need.
Avoiding alert fatigue in anomaly detection takes five disciplines. Calibrate thresholds against historical false positive rates so the alert volume is manageable. Tier alerts by severity so the most critical anomalies stand out from routine ones. Suppress duplicate alerts within a defined window. Route different alert types to the right teams rather than everything to one inbox. Track alert response and outcome so the calibration can improve over time. Without these disciplines, anomaly detection generates noise that gets ignored, defeating the point of the system. With them, the alerts that fire are the ones genuinely worth acting on.
To compare BI platform pricing fairly, build a representative workload and price both. Sample queries, refresh patterns, user counts. Most pricing comparisons are unfair because they price one platform at a typical workload and the other at peak. Both vendors have pricing calculators. Use them with realistic numbers, including non-production environments and growth assumptions.
Forecasting when products change frequently (range churn, NPI) is a common challenge in retail and consumer goods. New products without history cannot be forecast directly; the forecast has to come from comparable products (analogue forecasting), category-level patterns, or planning team input. The forecasting platform should handle the new-product case explicitly rather than treating it as a normal forecast. Mature products with stable histories are the easy part; the work is in handling the long tail of new and changing products well.
We handle data quality issues that look like anomalies carefully, because the line between 'data quality issue' and 'genuine anomaly' is sometimes thin. The pattern that works is two-stage: a data quality layer that filters out known data issues (missing values, duplicate records, schema violations) before the anomaly detection runs, and an anomaly detection layer that flags genuinely unusual events for review. Without the separation, anomaly detection produces a flood of false positives driven by data quality rather than business unusualness. The Trust Storytelling Delivery Checklist applies here particularly.
We handle promotional baskets in market basket analysis carefully. Promotions distort buying patterns: customers buy products together because they are on promotion, not because they have a natural affinity. Including promotional baskets without flagging them produces rules that reflect promotional planning rather than customer behaviour. The right approach is usually to tag promotional transactions, run MBA both with and without them, and compare the rules. The 'organic' rules (without promotions) are usually more useful for ranging and store layout; the 'with promotion' rules are useful for promotional design.
We handle promotional volatility in demand forecasting by treating promotions as features in the forecasting model. The historical promotional flags become predictors, and the forecast for future periods incorporates the planned promotional calendar. This produces promotion-aware forecasts that distinguish baseline demand from promotional uplift. The technique requires the historical promotional data to be captured cleanly, which is often the bottleneck. Without clean promotional history, the models attribute promotional uplift to other factors and produce poor forecasts in promotional periods.
We measure forecast quality through standard accuracy metrics: mean absolute error (MAE), mean absolute percentage error (MAPE), bias (whether the forecast systematically over- or under-shoots), and forecast value-add (whether the model is better than a naive baseline). The metrics are tracked over time and across SKU groups. Models that are getting worse trigger investigation. Models that are systematically biased need correction. The accuracy tracking is operational discipline, not just initial validation.
We test data pipelines in Microsoft Fabric through three test types. Unit tests on the transformation logic itself (does this Python function produce the expected output for known inputs). Integration tests on pipeline behaviour (does this pipeline correctly read from the source and write the expected schema to the destination). Data quality tests on the output (does the produced data meet the quality expectations: row counts in expected ranges, no null values in critical columns, foreign keys resolve). The Microsoft Fabric ecosystem supports each of these through standard tools (pytest, Great Expectations, SQL assertions). Mature data teams run all three; ad-hoc teams run none.
There are two main patterns for using market basket analysis rules in cross-sell. Real-time recommendations: when a customer adds A to a basket online or at the till, the system suggests B based on the rules. The implementation requires the rules table to be accessible at transaction time, usually through an API. Batch recommendations: customers who have bought A but not B receive a targeted email or marketing communication suggesting B. The batch pattern is simpler to implement and works well for email-driven businesses. Both patterns produce measurable lift in attach rate when the rules are good.
To visualise MBA rules in Power BI, three views work well. A rule table sorted by lift, showing the top rules with their measures. A network graph showing the strongest associations as edges between items, useful for understanding the structure of the relationships. A category heatmap showing aggregate association strength between categories rather than individual items. The right view depends on the audience: buying teams use the rule table, range planners use the network graph, leadership uses the heatmap. Power BI handles all three.
A single view of pipeline health comes from a central monitoring view rather than checking each pipeline individually. Fabric's monitoring hub gives one place to see every pipeline run across a workspace: its status, how long it took compared with previous runs, and where in the pipeline it failed if it failed. That execution history matters as much as the current status, because a pipeline that succeeds but takes three times as long as usual is an early warning of a problem that has not caused a visible failure yet. On top of that, most pipelines we build include retry logic for transient failures, such as a source system timing out briefly, so a single blip does not require a person to intervene. That is the closest real equivalent of self-healing: automatic recovery from the failure types that are expected and understood, combined with an alert to a person for the failure types that are not.
Each layer is a set of Delta tables in dedicated lakehouse workspaces or schemas. Bronze tables are append-only with transaction timestamps. Silver tables are mutable with merge logic for deduplication. Gold tables are typically rebuilt or merged depending on the modelling pattern. Direct Lake mode in Power BI reads Gold tables directly without copying. The Delta Lake format provides ACID transactions, schema evolution, and time travel across all three layers. The technology supports the pattern cleanly; the discipline is in the design.
We handle the training-versus-detecting tension in Fabric anomaly detection by treating anomaly detection as a continuously learning system. The model trained on the most recent stable period is used for current detection. As new observations accumulate, they are evaluated against the current model. Confirmed anomalies are excluded from future training data; confirmed normal observations are included. The model retrains periodically (weekly or monthly typically) on the cleaned recent data. This pattern keeps the model adapted to evolving normal behaviour while preserving sensitivity to genuine anomalies.
Training data with embedded anomalies is a real challenge in anomaly detection. Anomaly detection that trains on data containing past anomalies treats them as part of normal, which means similar future events will not be flagged. Two approaches. First, manual labelling: identify and remove known anomalies from the training data before fitting the model. Second, robust techniques (median-based statistics, robust ML methods) that resist contamination by outliers in training data. Both approaches require some manual investment in understanding the historical data. The investment pays back through reduced false negatives in production.
Every pipeline is built with monitoring and alerting from day one, not added afterwards. That means logging at each stage of the pipeline, automated failure alerts rather than someone noticing a report looks wrong, retry logic for transient source-system issues, and data quality checks that run before a load is considered complete. Across the Azure and Fabric pipelines we run for clients, this approach holds an average 96% pipeline reliability rate. We treat a pipeline the same way we treat a report: it needs an owner, a defined expected behaviour, and a way to know quickly when it has drifted from that behaviour.
ADF handles on-premises systems through the Self-hosted Integration Runtime, an agent installed on a server with access to the on-premises systems. ADF in the cloud orchestrates the pipeline; the integration runtime executes the source-side data movement. The model handles secure, performant integration without exposing the on-premises systems to the public internet. For mid-market businesses with significant on-premises data (older ERPs, internal databases, file shares), the integration runtime is the standard pattern. The same pattern applies to Fabric Data Factory.
Fabric is built on Azure infrastructure but presented as an integrated platform with its own pricing model and user experience. Several Azure services have Fabric equivalents: Azure Data Factory has Fabric Data Factory; Azure Synapse has Fabric Data Engineering and Data Warehouse; Power BI Premium is now part of Fabric capacity. The Azure services continue to exist for use cases that need them; the Fabric versions are usually preferred for new analytical workloads. The Fabric for Mid-Market FAQ in our library covers when Fabric is the right choice over assembling Azure components separately.
Data Activator triggers actions when configured anomaly conditions are met. The pattern: the anomaly detection scoring runs continuously or in micro-batches, the scores are written to a stream that Data Activator monitors, and rules evaluate the stream and trigger actions. For high-stakes anomalies, the action is human notification (Teams message, email, ticket creation) for immediate review. For lower-stakes anomalies, the action is logging for periodic review. Data Activator is the operational layer that connects detection to response.
Databricks handles multi-cloud natively. Databricks runs on AWS, Azure and GCP with the same product. If your organisation has workloads across more than one cloud, this is a genuine advantage. Fabric is Azure only and pulls Azure-adjacent workloads with it. Multi-cloud is rarely a small consideration when it applies to you.
Databricks is consumption-based, priced in DBUs (Databricks Units) per workload type. Compute scales when needed and stops when idle. Costs are visible per workload but less predictable per month. For organisations with intermittent workloads, this can be cheaper than Fabric's fixed capacity. For continuously running workloads, Fabric capacity often comes out ahead.
Delta Lake tables in OneLake keep a transaction log and version history for every change, so a pipeline can be pointed back at an earlier point in time and reprocessed deterministically. Combined with version-controlled notebooks or Dataflows Gen2, that history is what makes a genuine end-to-end replay possible rather than just a restore from backup.
Direct Lake mode changes the architecture significantly. Direct Lake reads Power BI data directly from Delta tables in OneLake without import or pass-through. The result combines Import-mode performance with near-real-time freshness because the lake is the storage layer rather than a separate copy. For new Fabric implementations, Direct Lake is the default for most semantic models. The exceptions are models that need calculated columns or complex DAX patterns that Direct Lake does not yet support, in which case Import mode remains the right choice.
Fabric workspaces can be connected to Azure DevOps Git or GitHub repositories. Notebooks, semantic models, pipelines, and reports are stored as code in the repository. Changes are made through pull requests with code review. Deployments to other workspaces (development to test to production) happen through the deployment pipelines feature or through CI/CD automation. The integration is the foundation for serious data engineering practice on Fabric.
Fabric is the productised, integrated path. Raw Azure (ADF plus Synapse plus Power BI plus Storage plus Purview, configured separately) is more flexible but more work. For most mid-market businesses, Fabric's integration advantages outweigh the flexibility gains of raw Azure. For very specific high-volume or specialised workloads, raw Azure still has a place, but those situations are uncommon at mid-market scale. The default for mid-market new builds is now Fabric. Existing investments on Synapse or raw Azure can usually be migrated cleanly into Fabric.
Fabric integrates natively with the rest of M365. Entra ID for identity, Purview for governance, Sensitivity Labels for data classification, Teams for collaboration, Copilot Studio for AI. Power BI is part of the platform, not bolted on. This is the strongest single argument for Fabric in M365-heavy organisations: the integration tax of running anything else compounds over time.
Fabric pricing affects the per-user Power BI cost decision through an interaction that is the most commonly missed cost pattern. On Fabric capacity below F64, every author and viewer of Power BI content still needs a Pro licence (£11 per user per month). At F64 and above, viewers are free; authors still need Pro. The crossover point where F64 becomes cheaper than F8 plus per-user Pro is around 500 viewers at PAYG, around 370 with annual reservations, around 250 with three-year EA discounts. Most mid-market businesses sit below the crossover and should stay there.
Hopton helps you choose between the platforms in two ways. A decision review, two to three weeks at fixed price, where we walk a structured framework with your team and produce a written recommendation. Or, if you have already chosen Fabric, full delivery. We do not deliver Databricks projects ourselves. If Databricks is the right answer, we will tell you and can introduce you to people who do.
Hopton helps you choose between the platforms in two ways. A decision review, two to three weeks at fixed price, where we walk a structured framework with your team and produce a written recommendation. Or, if you have already chosen Fabric, full delivery: architecture, build, governance and training. We do not deliver Snowflake projects ourselves. If Snowflake is the right answer, we will tell you and can introduce you to people who do.
Hopton typically builds MBA as part of a broader customer or product analytics engagement, or as a focused four to six week implementation when the data foundations are in place. The implementation includes the data preparation, the rule generation in Fabric Data Science, the Power BI visualisation, and the activation pattern with the chosen channel (online recommendation engine, email marketing, store planning, range review). The first useful rules are typically available within three to four weeks; the activation takes another two to four weeks.
Hopton typically builds anomaly detection as part of a broader analytics engagement covering finance integrity, operations monitoring, or risk management; or as a focused six to ten week implementation when the data foundations are in place. The implementation includes the data review, the technique selection (statistical versus ML based on the use case), the build in Fabric Data Science, the operational pattern (batch versus streaming), the activation through Data Activator and the existing alerting tools, and the calibration discipline. Working detection is typically running within four to six weeks; full operational integration takes another four to six weeks.
Hopton typically builds demand forecasting as part of a broader supply chain or commercial analytics engagement, or as a focused 8 to 12 week implementation when the foundations are in place. The implementation includes the data review, the technique selection (Prophet versus gradient-boosted versus classical based on the business shape), the build in Fabric Data Science, the validation and accuracy testing, the Power BI visualisation, and the integration with the planning system. Working forecasts are typically available within six to eight weeks; full operational integration takes another four to six weeks.
KQL is more concise for time-series operations, more powerful for pattern matching, and faster on the time-series workloads it was designed for. SQL is more universal, more familiar, and stronger for complex joins across many tables. For an analyst whose work is mostly time-series and event data, KQL pays back the learning investment quickly. For an analyst whose work is mostly traditional dimensional modelling, SQL remains the better tool. Microsoft Fabric supports both well; the choice can be made per use case rather than imposed across the team.
Market basket analysis, RFM, and CLV are complementary techniques. RFM segments customers by behaviour. MBA finds product associations. CLV models customer value over time. Together they form the core of customer and product analytics. Most of our customer analytics engagements deliver RFM first (because the data foundations are usually in place faster), MBA second (because it benefits from the customer-level view), and CLV third (because it needs the most data and is the most sophisticated of the three).
MBA supports range and listing decisions by identifying products that anchor baskets versus products that come along. Anchor products (high lift to many other products) are essential to the range. Pure 'comes along' products (high confidence given other purchases but low independent demand) survive on the back of the anchors. Stand-alone products (high independent demand, low association) survive on their own. The range review meeting becomes data-informed: which products are essential, which are dispensable, which need promoting harder. The conversation moves from anecdote to evidence.
Both are capable data platforms, and both use a lakehouse approach with consumption-based pricing rather than fixed licence tiers. The practical differences are in specialisation and ecosystem fit. Fabric is the better choice for organisations already invested in the Microsoft ecosystem (Power BI, Azure, M365) because the integration is native, the governance model is consistent, and Fabric capacity billing sits alongside your existing Azure consumption. Databricks is often the stronger choice for organisations with heavy Python/Spark workloads, dedicated data or AI engineering teams who want deeper control over the ML lifecycle, or genuinely multi-cloud architectures spanning AWS, Azure, and GCP. We assess which platform fits your specific situation rather than defaulting to a position.
The Fabric forecast is the analytical backbone of the S&OP (sales and operations planning) or IBP (integrated business planning) process. Each cycle, the team starts with the statistical forecast, adjusts for known factors not in the model, and arrives at a consensus plan. Fabric supports the cycle by providing the latest forecast, the historical accuracy data, and the analytical context (recent trends, anomalies, comparisons). The S&OP discipline runs on top; Fabric provides the data foundation.
Real-Time Intelligence is well-suited to IoT workloads. The eventstream ingests device telemetry, the eventhouse stores it efficiently, KQL queries surface patterns, and Data Activator triggers responses. Common patterns: equipment monitoring (predictive maintenance), fleet tracking, asset tracking, environmental monitoring. The architecture supports millions of devices and billions of events. For mid-market businesses with IoT data, the Fabric path is usually cleaner and cheaper than building a custom IoT analytics stack.
Power BI surfaces the forecast for business users through dashboards combining actuals, forecasts, and accuracy. The standard views include: forecast versus actual over time, accuracy metrics by SKU group and horizon, exception views showing SKUs with the largest forecast errors, and the planning view showing the forecast as input to the operational plan. Different audiences use different views: planners use the SKU-level detail, leadership uses the aggregate trends, finance uses the accuracy metrics for forecast credibility. Power BI handles all the views from the same underlying data.
Power BI works well on Databricks. Power BI connects to Databricks SQL through the standard connector and works fine. The integration is solid. The difference compared to Fabric is the same as Snowflake: it is two systems working together rather than one unified system. For BI-led organisations, the integration tax is real but manageable. For organisations where BI is one of many concerns, it is rarely the deciding factor.
Purview integrates natively with Fabric and Power BI. Sensitivity labels applied in Purview propagate to Fabric workspaces, semantic models, and Power BI reports. Lineage tracks data movement across the analytical estate. Classification rules apply consistently. The integration is one of the genuine advantages of staying within the Microsoft ecosystem; equivalent integration with non-Microsoft governance tools requires more work. For Microsoft-stack mid-market analytics, Purview is usually the right governance tool.
Snowflake handles multi-cloud natively. Snowflake runs on AWS, Azure and GCP and supports cross-cloud data sharing. If your organisation has workloads across more than one cloud, this is a genuine advantage. Fabric is Azure only and pulls Azure-adjacent workloads with it. Multi-cloud strategy is rarely a small consideration when it applies.
A real-time engagement differs from a batch one in three ways. The use case definition matters more, because real-time has higher costs and tighter benefits. The architectural design is more complex, with eventstream routing, eventhouse design, and Data Activator rules to specify. The operational model is different, with streaming-specific monitoring and recovery patterns rather than batch retry logic. The engagement timeline can be similar (eight to twelve weeks for a focused real-time implementation) but the discovery and design phases are usually more involved.
Anomaly detection is complementary to the other ML techniques — RFM, MBA, CLV, and forecasting. RFM segments customers; anomaly detection flags unusual customers. MBA finds product associations; anomaly detection flags transactions outside expected patterns. Forecasting predicts expected demand; anomaly detection flags actuals deviating from forecast. The five techniques together cover most of the mid-market ML opportunity. Many of our customer analytics and operations engagements deliver several of them in sequence as the foundations stabilise.
Anomaly detection integrates with your existing alerting tools through Data Activator's flexible action configuration, alerts can route to existing operational tools: ServiceNow for incident management, Teams for collaborative response, email for distribution, Power Automate flows for custom workflow integration, REST endpoints for proprietary tools. The detection runs in Fabric; the action lands wherever the operational team works. The integration model means anomaly detection extends existing workflows rather than requiring teams to adopt new tools.
Anomaly detection supports finance integrity through monitoring of transactional data for unusual patterns. Examples: journal entries outside normal patterns for the user or period, expense submissions outside the user's historical range, supplier payments to unusual destinations, account reconciliations with unusual variances, period-end accruals outside historical ranges. The patterns surface to the finance team for review, supplementing manual review. The technique is particularly valuable for mid-sized finance teams where manual review of every transaction is not feasible.
Fabric is native to Entra ID, Purview and M365 Sensitivity Labels. The governance story is genuinely tighter for organisations already on M365. Snowflake supports the same integrations but they are plumbed in rather than built in. The difference is operational rather than functional. Day-to-day, Fabric requires fewer touchpoints. For organisations with mature governance functions, the difference is smaller. For organisations relying on out-of-the-box governance, it is larger.
Both Fabric and Databricks are lakehouse platforms and handle the lakehouse architecture similarly. Databricks invented the term and the architecture. Fabric implements lakehouse via OneLake using Delta as the storage format. The architectures are functionally similar. The differences are in tooling, governance and the surrounding ecosystem. Both let you store once and query many ways. The platform decision is rarely about the lakehouse architecture itself.
Governance works in Fabric through Microsoft Purview, integrated natively with Fabric, plus three governance layers we configure explicitly: access models, regional deployment, and security alignment. Sensitivity labels, data classification, lineage tracking, and access governance all run through Purview, and Fabric workloads inherit that configuration automatically. Access models are workspace-based: read, write, and admin roles are assigned per workspace, with row-level security layered on top for report-level restriction. Regional deployment matters because a Fabric capacity is pinned to an Azure region at creation and cannot be moved later, so data residency and region are agreed before provisioning, not after. Security alignment covers tenant-level settings (external sharing, guest access, capacity admin roles) plus Entra ID conditional access, agreed with your IT and security team before go-live. The integration is meaningful: previous Microsoft data platforms required separate governance configuration for each component, and the discipline often slipped. Fabric centralises it. Purview setup, access model design, and regional and security sign-off are all part of the Establish phase in our standard implementation, not an afterthought.
Real-time fits alongside existing batch analytics through a unified architecture where batch and streaming feed the same lakehouse, with each workload using the data through whichever path suits the latency requirement. Batch analytics reads Gold-layer Delta tables for traditional reporting. Real-time analytics queries eventhouses for streaming dashboards and Data Activator for event-driven actions. The two coexist on the same Fabric capacity, sharing storage in OneLake, sharing governance through Purview, and sharing identity through Entra ID. Real-time is an extension of the platform, not a separate platform.
Real-time fraud detection works through Data Activator rules monitoring transaction streams for patterns that indicate fraud. The patterns are usually combinations: unusually high transaction amount, unusual location, unusual time, or unusual frequency relative to the customer's history. When the rule triggers, Data Activator routes the event to the fraud team for review or, in tightly-scoped automated cases, takes immediate action (hold the transaction for review). The trust framework matters here particularly: false positives cost customer relationships, false negatives cost actual money. Conservative threshold setting and human review on the marginal cases is the right pattern.
The Power BI/Fabric forecast integrates with your planning system as an input rather than a replacement: most mid-market businesses with serious supply chain operations use a dedicated planning system (RELEX, Anaplan, o9, Slimstock, similar) for the actual planning workflow. The Fabric forecast feeds the planning system as input. The pattern: Fabric generates statistical forecasts, the planning system adds management adjustments, the result becomes the operational plan. The integration is bidirectional: actuals flow back from the planning system into Fabric for accuracy tracking and model refinement. The integration architecture is straightforward but requires deliberate design.
The trust framework applies strongly to demand forecasts. Demand forecasts are model outputs going into operational decisions, often automated ones (replenishment orders, production schedules). The Model Trust One-Pager from our AI/ML whitepaper applies directly: every forecast should travel with sources, freshness, limits, confidence, recommended action, and change history. The 'limits' block matters particularly because forecasts are weakest in exactly the situations where the operational risk is highest (new products, promotional events, regime changes). Hiding the weakness produces operational disasters; surfacing it produces sensible decisions.
How important notebook-native development is depends on your team. For engineers who live in notebooks, it is core. For analysts who live in BI tools, it is rarely used. If notebooks are the daily working environment for more than a few people on your team, Databricks fits better. If notebooks are an occasional tool for a small group, Fabric's notebooks are sufficient.
CDC is implemented in Microsoft Fabric through Fabric Data Factory pipelines that consume change streams from source systems and write incremental updates to Bronze. The destination Bronze tables append the change records with metadata (timestamp, operation type, source identifier). Silver layer transformations apply the changes to produce the current-state view. The pattern preserves the change history in Bronze (useful for audit and recovery) while presenting the consumed Silver view as the working state. Mature CDC implementations also support point-in-time queries through Delta Lake's time travel.
Fabric capacity for mid-market is priced as F-SKU capacity, from F2 (around £200 per month) to F2048 for enterprise workloads. Mid-market implementations typically start on F8 (around £1,050 per month) or F16 (around £2,100 per month) for production. F64 (around £6,400 per month) becomes economical when viewer counts exceed roughly 500 because free Power BI viewing kicks in at that tier. The True Cost FAQ in our library covers the licensing economics in detail, including the five overspending patterns we see most often.
MBA is implemented in Fabric Data Science through a Fabric Data Science notebook reading transaction data from the lakehouse. The standard implementation uses Python with the mlxtend library, which provides Apriori and FP-Growth as ready-to-use functions. The notebook prepares the data (typically as a binary item-presence matrix or list-of-items format), runs the algorithm, filters the rules by support, confidence, and lift thresholds, and writes the rules back to the lakehouse as a Delta table. Power BI semantic models read the rules table for reporting and activation.
OneLake is the unified storage layer underneath every Fabric workload. One copy of the data, in open Delta format, accessible by all Fabric tools without copying or moving. This is genuinely different from previous Microsoft data platforms where Power BI had its own storage, Synapse had its own, and integrations meant duplicating data. The 'one copy' principle reduces storage cost, simplifies governance, and removes the synchronisation work that previously consumed engineering time. For most mid-market implementations, OneLake is the architectural feature with the highest practical impact.
Type 2 SCD is implemented in Microsoft Fabric through MERGE operations in Delta Lake (using SQL or Python notebooks) that compare incoming records to current dimension state, insert new versions when attributes change, and update effective-date and current-flag columns. The pattern is well-established and there are reusable code templates. The implementation is more involved than Type 1 (which is a simple overwrite) but the effort is bounded; once built, the same pattern handles all Type 2 dimensions in the platform.
Anomaly detection is built in Microsoft Fabric through Fabric Data Science notebooks for the modelling, with deployment patterns ranging from scheduled batch evaluation to streaming evaluation through Real-Time Intelligence. The standard implementation: a notebook trains the model on historical data, persists it, and runs scoring as a scheduled pipeline against new data. Anomaly scores are written to a Delta table in the lakehouse.
Threshold alerts fire when a value crosses a fixed boundary (revenue below £100k, response time above 2 seconds). Anomaly detection learns what 'normal' looks like and flags deviations from that pattern. The difference matters because thresholds work for one-dimensional, stable metrics; anomaly detection works for complex, multi-dimensional, or seasonally varying patterns. Threshold alerts on retail revenue would fire false alarms every quiet Tuesday and miss the anomaly of a Christmas Eve dropping 30 per cent below expected. Anomaly detection adapts to the pattern.
Demand forecasting is built in Microsoft Fabric through Fabric Data Science notebooks reading historical demand data from the lakehouse. The notebooks fit forecasting models (Prophet, gradient-boosted regression, ARIMA depending on the use case), generate forecasts for the relevant horizon, and write the forecasts back to the lakehouse as a Delta table. Power BI semantic models read the forecasts for reporting alongside actuals. The pipeline runs on schedule (typically weekly) so the forecasts stay current. The architecture is consistent with the broader Bronze/Silver/Gold pattern.
Demand forecasting predicts what customers want; sales forecasting predicts what will be sold. The two diverge when supply constraints, stock-outs, or pricing decisions affect what actually gets sold. A pure-play retailer with available stock has demand and sales aligned. A retailer with frequent stock-outs has demand exceeding sales for the missing periods. The distinction matters because demand forecasts should drive supply decisions independently of past sales constraints. Sales-only forecasting can produce a self-reinforcing cycle of under-stocking. Demand forecasting models, properly built, account for stock-out periods and predict the underlying demand.
Real-time differs from batch in three structural ways, from a data engineering point of view. Batch processes accumulated data on schedule; streaming processes individual events as they arrive. Batch tolerates failure with retry; streaming requires careful state management to avoid losing or duplicating events. Batch performance is measured in throughput; streaming performance is measured in latency. The skills, the tools, and the operational model differ. Most data engineering teams need to learn streaming separately from batch; they are not interchangeable.
A modern lakehouse differs from a traditional data warehouse in three structural shifts. Traditional warehouses combined storage and compute in a single appliance; modern lakehouses separate them, storing data once in cloud object storage and applying compute on demand. Traditional warehouses kept only structured business-ready data; lakehouses keep raw data alongside transformed data, supporting machine learning and exploration alongside reporting. Traditional warehouses were schema-on-write (transformations happened on the way in); lakehouses are schema-on-read or hybrid, allowing schema evolution without rebuilding the warehouse. Each shift reduces operational pain and increases flexibility.
A Fabric implementation in mid-market takes eight to twelve weeks for a production-ready first release covering the priority reporting domains. Three to six months for a full estate including all priority workloads, governance, and adoption work. The eight-to-twelve week figure assumes a focused engagement with prioritised reporting needs and a working Microsoft estate underneath. Stretched implementations (multiple competing priorities, scope creep, parallel workstreams) take longer. The Establish phase (typically four weeks) sets the architecture and the priorities before the build phase starts.
A Fabric lakehouse or warehouse implementation, including pipelines, governance, and Power BI integration, typically takes ten to sixteen weeks in the Build phase. Migrations from existing platforms depend on the complexity of the current estate. The Establish phase (four weeks) always precedes the Build to define the architecture and scope precisely.
How long a Microsoft Fabric implementation takes depends on scope, and it should be incremental rather than a single migration. A focused pilot typically takes around 4 to 8 weeks, a mid-sized deployment across several subject areas around 3 to 6 months, and a full enterprise migration around 9 to 18 months. The biggest drivers of timeline are governance complexity, migration scope and data quality, testing and reconciliation, and change management — not the technology itself.
The decision phase should take two to four weeks for most mid-market organisations. Longer for larger or more complex situations. The trap is letting the decision phase drag for months while the team debates internally. Set a fixed timeline, run a structured framework, and commit to a decision. Indecision costs more than a slightly imperfect choice.
The number of rules MBA produces is variable, controlled by the thresholds. With permissive thresholds (low minimum support and confidence), MBA can produce thousands or tens of thousands of rules, most of them not commercially useful. With strict thresholds, the rule count drops to hundreds or tens. The right thresholds depend on the business: a 50,000-SKU retailer needs different thresholds than a 500-SKU specialist. The pattern that works is to start permissive, observe the rule distribution, and tighten the thresholds until the output is the right size to act on.
A data pipeline has no fixed number of transformation steps, but most well-run Fabric estates settle on four: landing raw data untouched, conforming it into a consistent shape, applying business logic once, and serving it through a semantic model. Fewer steps usually means logic is being duplicated somewhere it shouldn't be, and many more often means the pipeline is solving problems that belong in governance rather than transformation.
Demand forecasting needs three years of history as the recommended minimum, five years as the comfortable target. Three years captures typical seasonal patterns and enough variation for the models to learn. Two years works for stable businesses but produces less reliable forecasts for the seasonal periods near the boundaries of the history. Less than two years makes proper seasonal forecasting difficult; the techniques work but the accuracy is materially lower.
Market basket analysis needs at least six months of transaction history for stable rules, ideally a year or more. The technique needs enough transaction volume to produce statistically meaningful associations. A retailer with one million transactions per year produces good rules on six months of data; a B2B business with ten thousand transactions per year needs years to find similar patterns. The threshold is transaction volume rather than time. Below roughly 50,000 transactions, MBA produces unstable rules that change between runs; above that, the rules stabilise.
Across our engagements, automated Azure and Fabric pipelines run at a 96% average reliability rate, meaning scheduled refreshes complete successfully without manual intervention in the large majority of runs. We get there through standard engineering discipline rather than luck: retry logic on transient failures, alerting the moment a pipeline fails rather than letting it fail silently, and reconciliation checks that catch data quality problems before they reach a report. Pipeline reliability is monitored from the first week of go-live, not assumed.
Snowflake works well with Power BI, but as two systems. Power BI connects to Snowflake via the standard connector and works fine. The integration is solid. What you do not get is the unified experience Fabric provides: shared semantic models, shared identity, integrated lineage. For most mid-market workloads, that integration tax is real but manageable. For BI-led organisations, the tax adds up over time.
Databricks is better for engineering teams if engineering means code-first, version-controlled, notebook-led work. Databricks gives engineers an environment that feels familiar. Fabric is improving on this front but the visual designers are still the centre of the experience. Engineering-led organisations tend to find Databricks lower friction. The tipping point is roughly when more than half your team writes code daily.
Databricks is not just for big data anymore. Databricks Serverless SQL and the wider product investments have made it credible for mid-market workloads. The historic perception that Databricks was only for large enterprises with Spark engineering teams is out of date. Cost is no longer a knock-out for mid-market. The remaining question is whether the platform's centre of gravity (code-first, notebook-led) suits how your team actually works.
Fabric is generally better for analyst-led teams. Fabric's visual designers, low-code paths and Copilot integration are genuinely strong for analyst-led work. Databricks is improving here but the centre of gravity remains code-first. If your team includes more analysts than engineers, Fabric reduces friction. If your team is engineer-heavy, Fabric can feel constraining.
Fabric is not just a rebadged Synapse, though the marketing makes it tempting to think so. Synapse components are inside Fabric, but the experience, governance, BI integration and capacity model are genuinely new. Treat Fabric on its own merits rather than judging it by Synapse's history. Several of the friction points that affected Synapse have been addressed in Fabric.
Fabric is sometimes over-specified for specific use cases, but rarely overall, at mid-market data volumes. Fabric is sized to scale up to enterprise workloads, but the entry tier (F2 capacity at around £200 per month) is genuinely affordable for mid-market and gives you the full platform capability at low volume. The scale headroom is useful as the business grows; the low entry point makes the start manageable. The cost concern in mid-market Fabric is rarely 'is this too much capacity'; it is 'is the per-user Power BI Pro licensing being properly managed alongside the capacity'. The True Cost FAQ in our library covers this in detail.
Fabric is ready for production workloads for most mid-market workloads. Fabric has been generally available for over a year and the components inside it are mature. The integrated experience is newer than the parts. There are still edge cases where the integration is incomplete, but for the workloads most mid-market organisations run, Fabric is production-ready. We have multiple clients on Fabric in production today.
Yes — Microsoft Fabric is ready for mid-market production use. Fabric reached general availability in late 2023 and has been in production across our client base since 2024. The platform is mature enough for mission-critical workloads when implemented properly. The early adopter risks (feature gaps, version churn, performance unpredictability) have largely settled. The remaining decision is not whether Fabric is ready but whether it is the right fit for your specific situation. For most mid-market businesses on the Microsoft stack, the answer is increasingly yes.
On warehouse-shaped workloads, Snowflake is often faster than Fabric. On the round trip from raw data to a dashboard a business user opens, less often. The benchmark wars are real but mostly relevant at scale most mid-market businesses do not reach. For a typical mid-market workload, both platforms are fast enough. The performance difference matters most when you are running enterprise-scale analytical queries continuously.
Fraud detection is one application of anomaly detection. Not all anomalies are fraud (a genuine extraordinary transaction is also anomalous), and not all fraud is anomalous (sophisticated fraud may look normal). The right framing is that anomaly detection surfaces unusual events, and human judgement determines whether the unusualness indicates fraud, error, exception, or legitimate variation. The technique supports fraud detection workflows but is not a complete fraud system on its own.
Demand forecasting is an ongoing capability, not a one-off project. Demand patterns evolve, new products launch, ranges change, market conditions shift. A one-off forecast is useful for the moment it is built and then degrades. The right framing is to build demand forecasting as a capability with regular refresh, monitoring, and accuracy tracking. The implementation is a project; the ongoing operation is a discipline. Most of our demand forecasting engagements include the ongoing operational pattern alongside the initial build.
Modern data architecture is less of an overkill for mid-market businesses than buyers usually expect. The architectural patterns scale down as well as up. A mid-market business does not need every component of an enterprise data platform, but the principles (separation of concerns, version-controlled pipelines, deliberate medallion layering) apply at every scale. The cost of doing it properly is moderate; the cost of doing it badly compounds quickly. Mid-market data platforms built without architectural discipline tend to need rebuilding within three to five years; platforms built with discipline keep growing.
Yes — real-time analytics is often overkill for mid-market businesses. The honest assessment for many mid-market businesses is that 'real-time' requirements turn out to mean 'within an hour' or 'within a day', which standard batch refresh patterns handle. The cost of building real-time architecture (development effort, infrastructure cost, operational complexity) is substantial; the cost should be justified by genuine business need rather than by the technical novelty. We have built real-time platforms for clients where the use case justified it; we have also recommended against real-time investment for clients where batch was sufficient.
For most mid-market workloads, the Fabric, Snowflake and Databricks comparison is a fair fight rather than one platform being obviously better. There are scenarios where one is obviously better and we cover those in the deck. Pure data warehouse performance at extreme scale tends to favour Snowflake. BI-led organisations on M365 tend to favour Fabric. In between, the decision is genuine and depends on your situation more than on the platforms.
Use a Lakehouse for most cases. The Lakehouse stores Delta tables in OneLake, which all the other Fabric workloads can read directly. It supports notebooks (Python, Spark) for transformation and runs Power BI semantic models with Direct Lake mode for high-performance reporting without copying data. The Warehouse is better when your team works in SQL exclusively, when you have specific T-SQL requirements, or when performance characteristics demand it. Most mid-market implementations use the Lakehouse as the default with the Warehouse for specific workloads.
Snowflake excels at elastic, multi-cloud SQL warehousing and data sharing. Databricks leads for large-scale data engineering and machine learning on the lakehouse. Microsoft Fabric excels when you are Microsoft-centric and want data integration, warehousing, governed self-service BI in Power BI, and AI in one platform rather than integrated afterwards. Choose by your dominant workload — warehouse-first, ML-first, governance-heavy, or multi-cloud — rather than by brand.
For MBA in Fabric, use Python with mlxtend for most mid-market implementations. The data volumes are usually manageable on a single notebook instance. Spark with MLlib is the right choice for very large transaction volumes (hundreds of millions of transactions or more) where the parallelism matters. The output format and the rule interpretation are the same regardless of the implementation. Most retailers and wholesalers we work with run successful MBA on Python without needing Spark.
For forecasting in Fabric, use Python for most mid-market implementations. The data volumes per SKU are usually moderate, and Python on a Fabric Data Science notebook handles thousands of independent time-series comfortably. Spark with MLlib becomes worthwhile at very large scale (tens of thousands or hundreds of thousands of SKUs forecast simultaneously). The output format is the same; the implementation choice is mostly about scale. Most mid-market retailers and wholesalers run successful forecasting on Python without needing Spark.
Most mid-market implementations use batch CDC at intervals of minutes to hours. Real-time CDC (where changes flow continuously into the platform) is achievable through Fabric's eventstreams and Real-Time Intelligence workload but is overkill for most analytical use cases. The right interval depends on the business need: customer-facing operational dashboards may justify minute-level CDC, monthly management reporting works fine on overnight CDC. Choose deliberately rather than defaulting to real-time.
Anomaly detection should run almost always with human review for the actions that matter, rather than fully autonomously. The pattern that works: the system flags anomalies and routes them to a human queue for review, the human confirms or dismisses, and the action follows the human decision. Fully autonomous anomaly response is appropriate only for narrow well-defined cases where false positives are cheap and the response is easily reversible. The Trust Storytelling framework applies particularly here: anomaly detection outputs going into operational decisions need provenance, freshness, and explicit limit framing.
Whether anomaly detection should run in batch or real-time depends on the use case. Finance integrity work usually runs in batch (overnight or weekly review of recent transactions). Operational monitoring usually runs in real-time or near-real-time (the value is in catching the issue while it is still recoverable). Customer behaviour monitoring varies. Most mid-market implementations include both: batch detection for periodic review with deeper analysis, and real-time detection for high-stakes immediate response. The Fabric architecture supports both within the same platform.
Data should be transformed after it lands in Fabric, as a general rule. We extract and load first, raw and unaltered, into the bronze layer, rather than transforming data in flight before it arrives. This matters practically: if a transformation rule turns out to be wrong, the original data is still there to reprocess from, rather than already overwritten by a flawed transformation on the way in. Transformation happens afterwards, moving data from bronze to silver to gold inside Fabric itself, where it is easier to test, version and fix. For on-premises source systems, such as an on-site SQL Server or an ERP database that is not internet-facing, a self-hosted integration runtime acts as the bridge: a lightweight agent installed inside your network that lets Fabric pipelines reach that data securely without opening the source system directly to the internet. This is the standard pattern for hybrid estates where some systems are cloud-based and others are not.
For new analytical platform investments, Fabric is usually the right choice. The integration advantages outweigh the flexibility losses for typical mid-market workloads. For specific use cases (high-volume custom data engineering, specialist AI workloads, integration with non-Microsoft systems where Azure connectors are stronger), the underlying Azure services may be the better fit. Most mid-market businesses end up with a mix: Fabric for the analytical estate, with specific Azure services where they add value. The architecture decision should follow the use case rather than picking a platform first.
Real-time data should sometimes be stored separately from batch data, depending on the use case. High-volume time-series data (IoT, application logs, audit events) lives best in eventhouses with Kusto-style optimisation. Streaming data that needs to be combined with batch data for unified reporting lives best in lakehouse Delta tables. Many real-time implementations route the same source events to both, with the eventhouse serving real-time queries and the lakehouse serving combined batch and streaming analytics. The decision is per source; the architecture supports both patterns.
You should forecast at both SKU and aggregate level, with reconciliation. SKU-level forecasts inform stock decisions; aggregate forecasts (category, total) inform broader planning and reconcile against expected revenue. Independent SKU and aggregate forecasts often disagree (the SKU forecasts add up to a different total than the aggregate forecast). Hierarchical forecasting techniques (using libraries like hts) reconcile the levels mathematically. For most mid-market businesses, simpler approaches (forecast at the most useful level, use management judgement to reconcile) work well enough.
For most mid-market organisations, migrating from Azure Synapse to Microsoft Fabric makes sense - but the timing and approach depend on your current investment. Fabric supersedes Synapse Analytics and Microsoft has signalled it is the strategic direction for the platform. The migration path is well-defined for most workloads. We assess the specific components you are using, the dependencies involved, and the right sequencing before recommending a migration approach.
For most mid-market Synapse customers, yes, eventually. Fabric is Microsoft's strategic direction for the data and analytics platform; Synapse is in maintenance rather than active development. New investments should go into Fabric, and existing Synapse workloads should be migrated as the right opportunities arise (renewals, refresh cycles, major project moments). The migration from Synapse to Fabric is well-documented and usually less work than the original Synapse build.
Starting with a forecasting proof of concept is often the right move. A four-week proof of concept on a representative SKU subset produces real forecasts on real data, validates the technique choice, and demonstrates the achievable accuracy. The proof of concept output is usable directly for some decisions and informs the scoping of the full implementation. We have run forecasting proof of concepts that became permanent operational forecasts because the accuracy was strong enough to act on without additional implementation effort.
Copilot in Fabric is genuinely ahead of Databricks Genie for analyst-led use. Natural language to DAX, Copilot in Power BI, semantic-model-aware AI assistants: this is where Microsoft's broader AI investment shows up first. For organisations whose AI ambition is mostly about analyst productivity and BI augmentation, Fabric has the lead today.
Copilot and Databricks Genie are both AI assistants for data work. Copilot in Fabric is more mature for analyst-led, BI-shaped use. Databricks Genie is improving and benefits from the broader AI investment Databricks has made. For organisations whose AI ambition is mostly analyst productivity, Copilot is ahead. For organisations doing serious ML and looking for AI-assisted engineering, Databricks is closing fast.
Unity Catalog is Databricks' answer to data governance and is genuinely strong. Fabric uses Purview for the same role. Both work. Unity Catalog is more tightly integrated into Databricks itself. Purview integrates with the wider M365 estate. Which is better depends on whether your governance perimeter is the data platform or the wider Microsoft estate.
Deep learning for forecasting is useful in specific scenarios. Deep learning models (DeepAR, Temporal Fusion Transformer, N-BEATS) handle very large numbers of related time-series simultaneously with shared learning across them. The technique pays back when you have many SKUs (thousands or tens of thousands) with related patterns. For mid-market businesses with smaller SKU counts, deep learning is usually overkill; gradient-boosted regression or Prophet produces comparable accuracy at a fraction of the complexity.
Real-time intelligence in Fabric is available, but used selectively in mid-market. The Real-Time Intelligence workload supports streaming data from event sources (IoT devices, application telemetry, transactional systems) with KQL-based analytics on top. The use cases that pay back are narrow: live operational dashboards where the value is in seconds rather than minutes, anomaly detection on fast-moving data, and event-driven workflows. Many mid-market 'real-time' requirements turn out to mean 'within an hour', which is usually fine on standard Fabric refresh patterns rather than streaming.
Anomaly detection techniques fall into three main categories. Statistical methods (control charts, Z-scores, MAD-based scoring) which are simple, interpretable, and effective for univariate or weakly-multivariate data. Machine learning methods (isolation forest, one-class SVM, autoencoder-based) which handle multivariate and complex patterns. Hybrid methods combining the two for specific use cases. Each has its place. Most mid-market implementations start with statistical methods because they are interpretable and easy to deploy, then layer on ML methods for use cases where the statistical approach is insufficient.
Slowly changing dimensions (SCDs) are dimensions whose attributes change over time, requiring deliberate handling to preserve historical accuracy. A customer changes address, a product changes category, an employee changes department. Without SCD handling, historical reports look correct but use current attributes for all periods, producing subtly wrong answers (the customer's December purchase shown against their February address). SCDs are one of the most important and most often skipped data architecture concerns. Getting them right is part of a serious modern data architecture.
Support, confidence, and lift are three measures of how meaningful a rule is. Support is the proportion of all transactions containing both items: high support means the combination occurs often. Confidence is the proportion of transactions containing item A that also contain item B: high confidence means B reliably follows A. Lift is the ratio of observed confidence to the baseline frequency of B: lift greater than 1 means A increases the likelihood of B; lift below 1 means A decreases it. Useful rules typically have lift well above 1, confidence above some practical threshold, and support high enough to be worth acting on.
There are six Fabric workloads. Data Factory for integration and pipelines (used in every implementation). Lakehouse for analytical storage with notebook-based transformation (used in most). Data Warehouse for SQL-based analytics (used where the team prefers SQL to Spark). Real-Time Intelligence for streaming and event data (used where genuinely needed, less common in mid-market). Data Science for ML (used in maybe a third of mid-market implementations). Power BI for visualisation (used in every implementation). Most mid-market Fabric estates run heavily on Data Factory, Lakehouse, and Power BI, with the others added as the use cases emerge.
There are six SCD types. Type 1 overwrites the attribute with the latest value (no history kept). Type 2 inserts a new row for each change with effective dates (full history). Type 3 keeps current and previous values in separate columns (limited history). Type 4 stores history in a separate table. Type 6 combines Type 1, 2, and 3 in one table. Most mid-market implementations use Type 1 for non-historical attributes and Type 2 for attributes that need full history. Type 3 and beyond are specialist patterns for specific cases.
Five migration risks come up repeatedly when moving to Microsoft Fabric, and we plan for each explicitly before a migration starts. Unsupported features: some Synapse and Power BI Premium capabilities do not have a direct Fabric equivalent yet, so we check your existing estate against current feature parity before committing to a timeline. Redesign requirements: pipelines and semantic models built for the old platform usually need genuine rework, not a lift-and-shift, particularly around Direct Lake and medallion layering. Dependency inventorying: every report, dataset, pipeline, and downstream consumer is catalogued before anything is touched, because the most common migration failure is breaking something nobody realised depended on the source system. Rollback planning: every migration phase has a defined rollback point, so a failed cutover does not take live reporting down with it. Reconciliation testing: before any workload goes live on Fabric, its output is checked against the legacy system at the row level, not just visually, until the numbers match exactly. None of this is unique to Fabric, but skipping it is the single biggest cause of migrations that overrun on time and budget.
Demand forecasting needs historical demand or sales by item per period at minimum. The granularity depends on the use case: daily for short-cycle retail, weekly for most retail and consumer goods, monthly for industrial. Additional data improves accuracy: pricing history, promotional history, channel breakdown, stock-out flags, marketing spend, weather data for weather-sensitive categories. The minimum dataset produces a working forecast; the richer dataset produces a more accurate one. Most ERPs and ecommerce systems carry the minimum; the additional data often requires integration work.
Anomaly detection needs historical observations of the metric or pattern being monitored. The volume needs to be high enough for the technique to learn what normal looks like: typically thousands of observations for univariate statistical methods, tens of thousands for multivariate ML methods. The history needs to span at least one full cycle of any seasonality (annual cycles need at least one year of data). Additional features (context, dimensions, related metrics) improve the precision of multivariate methods. Most operational and financial systems carry sufficient data for anomaly detection; the bottleneck is usually the modelling decisions rather than the data.
Market basket analysis needs two columns at minimum: transaction identifier and item identifier, with multiple rows per transaction (one per item in the basket). For richer analysis, additional columns help: customer identifier (for personalised rules), date (for temporal analysis), store or channel (for context-specific rules), price and quantity (for value-aware rules). The minimum dataset is small but the richness of the analysis grows with additional columns. Most ERPs and ecommerce systems carry the necessary data already.
Azure cost for mid-market data work is variable, because the services are consumption-priced. Typical mid-market Azure data spend ranges from around £1,000 per month for a modest implementation (small ADF environment, small Azure SQL database) up to £10,000 or more per month for a substantial estate. Reservations and Enterprise Agreements reduce list prices materially (typically 20 to 41 per cent for committed usage). The True Cost FAQ in our library covers the licensing economics; this FAQ covers the architectural choices that determine which services you actually need.
CI/CD means continuous integration and continuous deployment: automated testing and release of changes to the data platform. CI runs validation tests on every pull request (does the pipeline still work, does the model still load, do the measures still produce expected outputs). CD automates the promotion of approved changes through development, test, and production environments. The pattern reduces deployment risk, increases release frequency, and makes the platform more responsive to business needs. CI/CD is the engineering operating model that distinguishes mature data teams from ad-hoc ones.
Databricks is a cloud lakehouse platform built around Apache Spark. It runs notebooks, ML workflows, SQL warehouses and data engineering pipelines. It runs on AWS, Azure and GCP. Databricks champions open table formats (Delta, Iceberg) and an open architecture. The historic strength has been ML and data engineering. The recent investment has been in SQL and BI to compete more directly with Fabric and Snowflake.
Fabric is usually cheaper at mid-market scale once integration costs are factored in. Buying ADF, Synapse, Power BI Premium, and Storage separately produces a list-price total similar to or higher than equivalent Fabric capacity, with significantly more configuration work. The integration tax is real: separate components require integration engineering, separate governance, and separate operational overhead. Fabric removes that. For mid-market businesses, the all-in cost of Fabric is usually lower than the all-in cost of equivalent capability assembled from separate Azure components.
Fabric cost is capacity-based. You buy a capacity tier (F2, F4, F8, and so on) and run workloads against it. Costs are predictable per month, less so per query. Mid-market organisations typically land on F4 or F8 for production, costing several hundred to a few thousand pounds a month. Power BI Premium licences flow into Fabric capacity, so existing investment is not lost.
Fabric is an integrated platform combining data warehousing, lakehouse, data engineering, real-time intelligence, data science and Power BI under one capacity model. It is Azure-only and tightly integrated with M365. The components were previously separate (Synapse, Power BI Premium, Data Factory) and Microsoft has unified them under one billing and governance model.
Snowflake is a cloud data warehouse that has expanded into a broader data platform. It separates storage and compute, runs on AWS, Azure or GCP, and is known for elastic scaling. It also handles data engineering, data sharing, and increasingly data science workloads. Snowflake is platform-neutral about your BI tool: Power BI, Tableau, Looker, and others all work on top.
Snowflake cost is consumption-based. You pay for compute when queries run and storage for data at rest. Costs are visible per query, less predictable per month. Mid-market organisations often land on a few hundred to a few thousand pounds a month, depending on workload patterns. Spiky workloads can be cheaper than fixed Fabric capacity. Steady workloads are often more expensive.
A Fabric implementation in mid-market delivers a unified analytics platform with one storage layer (OneLake), one compute model (Fabric capacity), and one governance layer (Purview). The workloads sit on top: data integration through Data Factory, lakehouse and warehouse for analytical storage, real-time intelligence for streaming, data science for ML, Power BI for visualisation, and Data Activator for event-driven actions. The integration is what differentiates Fabric from buying these components separately. Organisations using Fabric get a coherent platform; organisations using the components separately get an integration tax.
A data engineering consultancy delivers six things in our engagements: warehouse and lakehouse design using proper schemas rather than a flat data dump, cloud platform build on Azure or Microsoft Fabric sized to what the business actually needs, automated integration pipelines that pull data from ERP, CRM and other source systems on a schedule rather than by hand, governance built into the platform from day one including cataloguing, lineage and access control, semantic modelling that translates raw tables into business terms analysts can use directly, and migration work for businesses moving off legacy warehouses or another cloud. The output is a governed analytics layer your own team can run and extend, not a black box that depends on the consultancy to touch it again.
A data engineering engagement with Hopton covers more than pipeline code. A typical engagement starts with an assessment of the current estate, moves through a migration plan for the workloads being moved onto Azure or Fabric, adds a governance layer built on Microsoft Purview for classification and lineage, and ends with a modernisation path that retires the parts of the old stack that no longer earn their cost. The architecture document produced during the Establish phase ties these together in writing: which Azure services are used and why, the cost model, the security posture, and the operational model for running it day to day. That operational model is where cost optimisation and workload tuning live in practice - right-sizing capacity, monitoring spend against actual usage, and re-tuning pipelines and Eventstream jobs as data volumes grow, rather than treating cost as a one-off sizing exercise at kickoff. Security governance is not bolted on afterwards either: sensitivity labels, access boundaries, and audit logging sit inside the same Purview-based governance layer used across Data Factory, Data Lake Storage, and the Fabric lakehouse, so the whole estate is governed consistently instead of pipeline by pipeline.
A typical Microsoft Fabric implementation runs phase by phase. Establish (weeks 1 to 4): discovery of source systems, priority reporting domains, architecture design, capacity sizing, governance approach, written delivery plan. Build (weeks 5 to 12 typically): Bronze layer extractions, Silver layer transformations, Gold layer business models, certified semantic models, priority reports, governance configuration. Adoption (weeks 12 to 16): training, change management, the Trust Storytelling Delivery Checklist applied to executive outputs, Decision Adoption Rate baseline. Continuity is the optional ongoing engagement after launch.
An end-to-end machine learning workflow inside Microsoft Fabric stays in one environment rather than moving data between separate tools for each stage. Data already sits in the lakehouse's silver or gold layer, so a notebook can read it directly using Spark, without a separate export step. Experimentation happens in that same notebook: Fabric's built-in MLflow tracking logs every run automatically, so different feature sets or model versions can be compared without a data scientist building their own tracking spreadsheet. For standard problems, AutoML can test a range of algorithms against the training data and return the best-performing candidate rather than requiring every model to be hand-coded from scratch. Once a model is selected, it is registered in the same workspace and can be called from a notebook, a pipeline, or scored directly against the same lakehouse tables it was trained on. The output, a prediction or a score, lands in a gold-layer table like any other governed dataset, so it inherits the same row-level security and appears in Power BI through the same certified semantic model as everything else, rather than as a separate AI report bolted on afterwards. This matters more than the individual tools: the reason to run machine learning inside Fabric rather than a separate platform is that the model's output stays inside the governance boundary already built, instead of becoming a new, ungoverned data source of its own.
An operational real-time dashboard is live views of operational KPIs refreshed continuously rather than on schedule. Examples include call centre dashboards showing live queue depth and agent status, warehouse dashboards showing live picking rates and order status, retail dashboards showing live till activity by store, and supply chain dashboards showing live shipment positions. The value is the ability to act on developing situations: a queue building, a warehouse falling behind, an unusual sales pattern emerging. Power BI dashboards on Direct Lake against eventhouse data deliver the live view.
Data observability covers pipeline health monitoring (are jobs running on schedule and succeeding), data quality monitoring (is the data arriving with expected characteristics), lineage tracking (where did this data come from and what transformations did it pass through), and freshness monitoring (is the data current). Microsoft Purview provides several of these capabilities natively for Fabric. Third-party tools (Monte Carlo, Soda, Great Expectations) extend the coverage. Observability is the operational discipline that distinguishes data platforms that work reliably from platforms that suffer recurring outages.
Data orchestration means coordinating the execution of data pipelines: when each pipeline runs, what it depends on, how failures are handled, how the orchestration recovers when something breaks. Fabric Data Factory provides the orchestration layer with pipeline definitions, scheduling, dependency management, and monitoring. The orchestration is what makes the data platform reliable: pipelines that run individually in isolation are fragile; orchestrated pipelines with proper dependencies and recovery are robust.
Market basket analysis (MBA) output is a table of rules, each row showing antecedent (the items in the 'if' part of the rule), consequent (the items in the 'then' part), support, confidence, and lift. A typical rule might read: 'antecedent: bread, butter; consequent: jam; support: 0.04; confidence: 0.65; lift: 8.2'. This means 4 per cent of transactions contain bread, butter, and jam together; among transactions with bread and butter, 65 per cent also contain jam; the lift of 8.2 means the bread-and-butter combination is 8.2 times more likely to contain jam than a random transaction. Rules with lift greater than 2 are usually the actionable ones.
Real-time means three different things in practice, and the distinction matters. True real-time means data flowing continuously with latencies measured in milliseconds to seconds. Near real-time means data flowing in micro-batches with latencies measured in seconds to minutes. Operational reporting often described as real-time means refresh cycles of minutes to hours. The right answer depends on what the business needs. True real-time is appropriate for operational systems, fraud detection, and IoT. Near real-time covers most operational dashboards. Anything an hour or longer is batch, regardless of what the requirements document calls it.
Replayability is the ability to run a pipeline again from the original raw data and land at the same, or a deliberately improved, result. It depends on Bronze staying untouched and on transformation logic living in version-controlled notebooks or Dataflows rather than being applied once by hand. When a quality rule is added or a bug is fixed, a replayable pipeline can reprocess the full history through the new logic in one run, rather than patching only the records someone happened to notice were wrong.
The Hopton Fabric architecture is Bronze, Silver, Gold layers in OneLake. Bronze captures raw data from source systems unchanged. Silver cleans, structures, and standardises. Gold contains certified business-ready facts and dimensions. Power BI semantic models point at Gold using Direct Lake. Reports are built on the semantic models and inherit certified definitions. The architecture is the same we use across implementations regardless of source system or sector. Specifics vary in the source extraction patterns and the Gold-layer model design; the structure is consistent.
The decision review for a Fabric vs Databricks choice covers workload analysis, team and skill assessment, estate review, scoring against the six dimensions in the comparison deck, and a written recommendation. The review is independent of any subsequent build. A meaningful share of reviews end with us recommending Databricks, and that is fine. Helping you choose well matters more than winning the build.
The decision review for a Fabric vs Snowflake choice covers workload analysis, team and skill assessment, estate review, weighted scoring against the six dimensions in the comparison deck, and a written recommendation. Output is a document you can take to anyone, including in-house or another partner. The review is independent of any subsequent build engagement. About half of decision reviews end with us not doing the build, and that is fine.
In a medallion architecture, Bronze contains raw data exactly as extracted from source systems. JSON files from APIs, CSV exports, database snapshots, change logs. The format preserves source structure even when awkward. Silver contains cleaned data with deduplication, standardisation, and quality rules applied. Customer records are unified, product codes are consistent, dates are normalised. Gold contains business-ready facts and dimensions with the modelling decisions made. Sales fact tables, customer dimensions, time-intelligence dates. Power BI semantic models consume from Gold; ML notebooks usually consume from Silver or Gold depending on the use case.
Fabric implementations go wrong in three patterns we see repeatedly. Inadequate Silver layer cleansing, where Bronze data flows too quickly into Gold without proper standardisation, producing reports that almost reconcile but not quite. Capacity over-provisioning at launch, where teams buy F32 'for headroom' and run at 18 per cent utilisation. And inadequate governance configuration, where Purview is set up but not operationally maintained, so the labels and access governance drift. Each is preventable. None is unique to Fabric, but Fabric implementations exhibit the patterns more visibly because the platform is more integrated than what came before.
A previous forecasting investment that did not deliver is a common situation. The honest assessment usually shows that the technical work was sound but the operational integration was missing, the accuracy expectations were unrealistic, or the trust framework was absent. Forecasts that nobody acts on are wasted regardless of how accurate they are. Forecasts that everyone acts on without understanding the limits produce operational mistakes. The relaunch usually focuses on the operational and trust workstreams rather than rebuilding the technical core. Several of our forecasting engagements have been recoveries of previously failed projects.
Azure Data Factory is used for moving and transforming data between systems. ADF orchestrates pipelines that extract data from source systems (ERPs, CRMs, databases, APIs, files), transform it as needed, and load it into target systems (data lakes, data warehouses, downstream applications). It is the data integration backbone for most mid-market Azure data implementations. The capability is comparable to enterprise data integration tools (Informatica, Fivetran, Airbyte) with the advantage of being native to Azure and integrated with the wider Microsoft ecosystem.
Azure Data Lake Storage Gen2 is object storage optimised for analytical workloads. It is the storage layer underneath traditional Azure data architectures (ADF plus Synapse plus Power BI). For new builds, OneLake within Fabric is usually the right choice instead, because it provides similar capability with cleaner integration to the rest of the analytical stack. ADLS Gen2 remains the right choice when you have specific tooling that integrates with it directly, when you need to share storage across non-Microsoft analytical tools, or when existing investments make migration impractical.
Azure Machine Learning is Microsoft's dedicated ML platform with full lifecycle support: data preparation, model training, deployment, monitoring, and MLOps. It supports Python, R, automated ML, custom code, and prebuilt cognitive services. For mid-market ML workloads, Fabric Data Science is usually sufficient and more cleanly integrated with the wider analytical estate. Azure ML is the better choice for very high-volume training, complex MLOps requirements, model serving patterns that need dedicated compute, or ML workloads that are part of a broader Azure-only architecture.
Azure OpenAI Service is direct API access to OpenAI's large language models (GPT-4, GPT-3.5, embeddings) hosted in Azure with enterprise-grade security and governance. Azure OpenAI is the path for custom AI applications outside the Copilot interaction pattern: bespoke generative AI features in your own applications, custom retrieval-augmented generation systems, advanced agent patterns. For typical mid-market use cases, Microsoft Copilot products usually cover the need without requiring direct Azure OpenAI integration. Azure OpenAI becomes the right choice for specific applications that need direct API access.
Synapse Analytics was Microsoft's previous-generation integrated analytics platform, combining data warehousing, big data, and BI. It is now in maintenance rather than active development; Microsoft Fabric is the strategic successor. Existing Synapse customers can continue to use it, but new investments should generally go into Fabric. The migration from Synapse to Fabric is well-documented and usually less work than the original Synapse build. For mid-market businesses on Synapse, the migration question is when, not if.
Azure for data and analytics is Microsoft's cloud platform of services for data integration, storage, processing, machine learning, and AI. The relevant services for mid-market data and analytics work include Azure Data Factory (data integration and orchestration), Azure SQL Database and Managed Instance (relational databases), Azure Storage and Data Lake Storage (object storage), Azure Synapse Analytics (legacy analytics platform, now in maintenance), Azure Machine Learning (the ML platform), Azure OpenAI Service (generative AI), and Microsoft Purview (governance). Microsoft Fabric is the newer integrated platform that brings several of these together; the underlying Azure services remain available and useful.
Data Activator is the Fabric capability that triggers actions when conditions are met in streaming data. The pattern: define a rule (when temperature exceeds 30 degrees, when stock level drops below 100 units, when an account shows churn signals), and Data Activator evaluates the rule continuously against incoming events and triggers the configured action (Teams notification, email, Power Automate flow, custom webhook). The capability moves Fabric from analytics to action: not just observing what is happening but doing something about it.
DataOps is the application of DevOps principles to data engineering: automation, version control, monitoring, fast feedback cycles. Yes, it is relevant for mid-market, scaled appropriately. A mid-market business does not need an enterprise DataOps platform with dozens of tools, but the underlying disciplines (version control, automated testing, deployment pipelines, observability) apply at any scale. The pattern that works is selective adoption: pick the practices that produce the most value for your team size and skip the ones that add overhead without proportional benefit.
Direct Lake is a Power BI query mode unique to Microsoft Fabric. It allows Power BI to query data directly from the OneLake delta parquet files without importing or caching the data, and without the latency of a live DirectQuery connection. The result is near-import-speed performance on data that is always current - removing the trade-off between performance and freshness.
Fabric Copilot Capacity (FCC) is a separate capability that lets you point all Copilot usage across the tenant to one centralised capacity for billing and management. Still requires F64 or higher because it is designed to be shared across many workspaces. Most mid-market clients do not need this until they are operating at scale. Useful when you have many distinct Power BI workspaces and want to centralise the Copilot cost rather than carry it on each capacity.
KQL is Kusto Query Language. The query language for eventhouses and the broader Kusto ecosystem (Azure Data Explorer, Azure Monitor, Azure Sentinel). KQL is designed for time-series and log analytics workloads. The syntax is pipeline-style (data flows through a series of operators) rather than nested-query SQL. The language is genuinely well-designed for its domain: queries that would be awkward in SQL are concise in KQL, and the performance on time-series data is excellent. Worth learning if you are doing serious work with eventhouses.
Microsoft Fabric is a unified, SaaS analytics platform that brings data integration, a lakehouse and warehouse, real-time intelligence, data science, and Power BI together on a single storage layer called OneLake, with Copilot and AI woven in. It replaces the assembled stack many teams run today — a separate integration tool, warehouse, lake, data-science environment and BI tool — and the integration glue between them, consolidating everything into one governed environment with one security model and one bill.
Microsoft Fabric is Microsoft's unified, SaaS data platform, launched in 2023. It consolidates what previously required five or more separate Azure services (Synapse Analytics, Data Factory, Power BI Premium, Purview, and more) into a single, integrated platform. Fabric covers data engineering, data warehousing, real-time intelligence, data science, and Power BI reporting - all governed through a single OneLake storage layer. Compared with a fragmented traditional stack built from separately licensed, separately managed tools, this consolidation matters in three practical ways: one copy of data in OneLake rather than duplicate copies scattered across a warehouse, a lake, and a BI tool; far less integration engineering to connect those tools together; and one consumption-based capacity and governance model instead of several overlapping licences to manage.
Microsoft Purview is Microsoft's data governance platform, covering sensitivity labels, data classification, lineage, glossary, and access governance. Purview integrates with Fabric, Azure data services, and Microsoft 365 to provide consistent governance across the data estate. For mid-market businesses building serious analytical platforms, Purview is the standard governance layer. The Data Governance FAQ in our library covers the governance methodology; Purview is the toolset that supports it.
Prophet (open-source from Facebook/Meta) is a forecasting library designed for business time-series with strong seasonality, holiday effects, and trend changes. It is more accessible than ARIMA (less parameter tuning) and handles common business patterns out of the box. Prophet is particularly well-suited to retail and consumer goods forecasting where weekly and yearly seasonality dominate. The technique is a strong default for mid-market demand forecasting because it produces reasonable results without expert tuning.
Real-Time Intelligence in Microsoft Fabric is the Fabric workload for streaming and event-driven analytics. It includes eventstreams (for ingesting streaming data), eventhouses (for storing time-series and event data), KQL queryset (for querying the data), and Data Activator (for triggering actions based on events). The workload is built on Microsoft's Azure Data Explorer (Kusto) technology, mature and battle-tested at hyperscale within Microsoft's own services. Real-Time Intelligence is one of the six core Fabric workloads alongside Data Factory, Data Engineering, Data Warehouse, Data Science, and Power BI.
A data pipeline SLA is a documented commitment about the data a pipeline delivers, not just whether the servers were running. A good one covers freshness (how recent the data must be), availability (how often it is expected to be delivered on time), completeness (which sources and records must be present), the alerting that fires when a target is missed, and the incident-response expectation — who responds and how quickly. We tier SLAs by reporting criticality, so board and finance data carries a far tighter commitment than exploratory datasets.
A good false positive rate depends on the use case and the cost of investigation. For high-stakes anomalies (potential fraud, equipment failure), a false positive rate of 10 to 20 per cent is acceptable because the cost of missing a true anomaly is high. For lower-stakes anomalies (routine outliers in finance review), false positive rates need to be much lower (5 per cent or below) or the operational team stops engaging. The right threshold is tuned to balance the false positive rate against the missed-anomaly rate, with the balance set by the business impact of each.
A medallion architecture organises data into three progressively refined layers inside OneLake. Bronze holds an unaltered copy of what arrived from the source system. Silver holds that data once it has been validated, deduplicated and modelled at a usable grain. Gold holds business-ready, curated datasets built specifically for reporting and analytics. Each layer has a distinct job, so a problem at one stage rarely means rebuilding the whole pipeline.
A modern data architecture is a data platform built around three principles: separation of storage from compute (lakehouse architecture), separation of raw data from business-ready data (medallion architecture), and separation of code from configuration (deployment pipelines and version control). The result is a platform that scales with the business, supports multiple analytical workloads from the same data, and remains maintainable as the team and the data grow. The architecture is the difference between a collection of pipelines that work today and a platform that keeps working in five years.
An eventhouse is a streaming-optimised database in Fabric, built on the Kusto engine. It stores time-series and event data with sub-second query latency on billions of rows. The schema-on-read model handles semi-structured data well (JSON events, IoT telemetry, application logs) without requiring pre-defined schemas. Eventhouses are the right destination for high-volume streaming data; lakehouse Delta tables are the right destination for streaming data that needs to be combined with batch data for unified analytics. Most real-time implementations use both, with eventstreams routing events to each based on the use case.
An eventstream is the ingestion mechanism for streaming data into Fabric. Sources include Azure Event Hubs, IoT Hub, Kafka, custom REST endpoints, and Microsoft 365 audit logs. The eventstream routes events to destinations: into an eventhouse for time-series queries, into a lakehouse Delta table for combined batch and streaming analytics, or into a custom endpoint for application integration. The eventstream handles the streaming pipeline mechanics so the engineering team can focus on the analytical logic rather than the plumbing.
Anomaly detection is a pattern recognition technique that identifies observations differing significantly from expected behaviour. The output is typically a flag (anomaly or not) or a score (how unusual the observation is) per record or per time period. Anomaly detection is one of the most useful and underused ML techniques in mid-market analytics.
CDC is the discipline of capturing changes (inserts, updates, deletes) from source systems incrementally rather than re-reading the full source on every refresh. It matters for two reasons: efficiency (full reloads of large tables are expensive and slow) and fidelity (CDC captures the history of changes, enabling slowly changing dimensions and audit trails). Modern data platforms run on CDC for any source where it is available. Sources without CDC support are accommodated through other patterns, but CDC is the default for transactional systems.
Demand forecasting is a predictive technique that estimates future demand for products, SKUs, or services over a defined horizon. The output is typically a time-series of expected demand by item per period (week, month, quarter), with confidence intervals quantifying the uncertainty. Demand forecasting drives stock planning, production scheduling, capacity decisions, and supplier ordering. It is one of the highest-value ML use cases for product-led businesses because the cost of being wrong (stock-outs or over-stock) is usually material.
Isolation Forest is an ML technique that identifies anomalies based on how easily a record can be 'isolated' from the rest through random splits. Anomalous records are easier to isolate (fewer splits needed) than normal records. The technique handles multivariate data well, runs efficiently on large datasets, and produces interpretable scores. It is one of the standard go-to ML techniques for general-purpose anomaly detection.
Market basket analysis is a pattern-finding technique that identifies which products tend to be bought together. The output is a set of association rules of the form 'customers who buy A also buy B', each with three statistical measures: support (how often the combination appears), confidence (how often B is bought when A is), and lift (how much more likely B is bought when A is present, versus the baseline rate). MBA has been used in retail since the 1990s and remains one of the most reliable techniques for cross-sell, range planning, and store layout decisions.
Medallion architecture is a data organisation pattern with three layers: Bronze (raw data ingested from source systems, unchanged), Silver (cleaned, deduplicated, conformed data ready for analytical use), and Gold (business-ready data with the modelling, aggregations, and semantic structure for reporting and analytics). Each layer has a defined purpose, defined ownership, and defined consumption patterns. The pattern was popularised by Databricks but applies equally to Microsoft Fabric implementations.
Lambda architecture is the historical pattern of running parallel batch and streaming pipelines, with a serving layer combining the two. It was developed when batch and streaming systems were genuinely separate. Modern unified platforms (Fabric, Databricks, Snowflake) reduce the need for explicit Lambda architecture because the same storage layer can be queried by both batch and streaming workloads. Most new mid-market real-time implementations use a unified pattern rather than explicit Lambda. The principle of supporting both batch and streaming remains; the implementation has simplified.
The cost of switching BI platforms later is lower than it used to be, because both platforms support open table formats (Delta, Iceberg) and standard SQL, which makes switching easier than it once was. Switching is still meaningful work, mostly in semantic models, BI tooling and security. Plan as if you are choosing for five years. Switching is possible but not free, and choosing well now is cheaper than fixing it later.
Bronze is the ingestion layer: an immutable, as-is copy of source data, kept exactly as it arrived so nothing is lost or reinterpreted too early. Silver is the validation layer: records are deduplicated, conformed to agreed types and keys, checked against quality rules, and modelled at a clear grain. Gold is the consumption layer: business logic, aggregations and denormalisation are applied so the data is shaped for a specific report, semantic model or downstream application. Responsibility moves from preserving truth, to cleaning it, to presenting it.
The famous beer and nappies example is an often-cited story (whose details vary) about a US retailer discovering through MBA that beer and nappies were frequently bought together, allegedly because young fathers picking up nappies also picked up beer for the weekend. The story has acquired apocryphal status; the original analysis is debated. The example illustrates the principle: MBA surfaces non-obvious associations that the buying team would not have found through intuition alone. Whether or not the original story is fully accurate, the technique genuinely produces unexpected, actionable rules across most categories.
Power BI Copilot and Fabric Copilot both run on Fabric capacity. Since April 2025 the minimum capacity is F2 (around £200 per month). Before that change, the minimum was F64. The reduction is significant for mid-market organisations that were previously priced out. The capacity covers all Fabric workloads, not just Copilot, so the cost should be assessed against the whole platform value, not Copilot in isolation.
The medallion (bronze, silver, gold) architecture is a way of organising a data platform into three layers, each with one job. Bronze is raw, immutable landing that preserves data exactly as it arrived along with its ingestion metadata, which makes reprocessing possible. Silver is cleaned and validated — deduplicated, schema-standardised, quality-checked. Gold is business-modelled and analytics-ready, feeding semantic models and reports. Because each layer has a clear responsibility, you get end-to-end lineage, auditability and the ability to replay history from source.
ADF can handle most common mid-market sources. SaaS applications (Salesforce, Dynamics, ServiceNow, Workday) through native connectors. Databases (SQL Server, Oracle, MySQL, PostgreSQL) through database connectors. Files (CSV, JSON, Parquet) on cloud storage or on-premises. APIs through REST connectors. SAP through dedicated connectors. The connector library is broad and growing. For the rare source where no native connector exists, custom integration is achievable through Azure Functions or Logic Apps. Source system access is rarely the constraint.
Microsoft Fabric anomaly detection can monitor operations and supply-chain problems including: equipment performance monitoring (deviations from historical patterns indicating maintenance need), supplier delivery performance (lead time anomalies indicating disruption), warehouse productivity (unusual processing times), order pattern monitoring (demand anomalies in advance of stock-out risk), and customer service quality (response time or satisfaction anomalies). Each surfaces an issue while it is recoverable. The supply chain use cases tend to be the highest-value for mid-market because of the cost of stockouts and supplier failures.
Three orchestration patterns keep Azure and Fabric pipelines reliable as they scale: beyond the basic bronze, silver and gold layering, three patterns do most of the work. Parameterised pipelines: one pipeline definition with source, destination and load type passed in as parameters, rather than a hand-built pipeline per table or source system, which is what lets a platform scale past a handful of sources without becoming unmaintainable. Incremental loading: pipelines track a watermark, typically a last-modified timestamp or change-tracking column, so a scheduled run pulls only what changed rather than reprocessing an entire source every time, which is both faster and less likely to strain the source system it is reading from. Control flow with defined recovery: activities are chained with explicit success, failure and completion paths, so a failure partway through a load triggers a specific response, retry, alert, or skip-and-continue, instead of leaving bronze, silver and gold out of sync with each other silently. We also deploy pipeline changes through a proper release process, so a change tested in a development workspace is promoted to production deliberately rather than edited live, and every run is logged centrally so a failure is visible within minutes rather than discovered when someone notices a stale dashboard.
Three principles work best for pipeline design. Idempotency: a pipeline should produce the same result whether it runs once or many times, so reruns after failures are safe. Modularity: pipelines should be small and focused, composable into larger workflows rather than monolithic. Observability: pipelines should emit structured logs, metrics, and lineage information so failures can be diagnosed quickly. These principles produce data platforms that operate reliably with minimal intervention; absence of any one of them produces platforms that consume engineering time disproportionate to the value they deliver.
Three patterns of rules should be ignored, or at least treated with scepticism. Trivially obvious rules (peanut butter and jam, salt and pepper) confirm common knowledge without adding insight. Rules driven by single events or promotional periods rather than recurring patterns. Rules involving very low-frequency items where the support is too low to be statistically reliable. Filtering these out concentrates attention on the rules that genuinely surface non-obvious patterns. Most useful MBA outputs require some manual review by someone who understands the business context.
Demand forecasting uses techniques from three main categories. Classical time-series models (ARIMA, exponential smoothing, ETS) which extrapolate from historical patterns. Machine learning models (gradient-boosted regression, random forests, neural networks) which can incorporate features beyond the time-series. Hybrid and modern approaches (Prophet, DeepAR, Temporal Fusion Transformer) which combine elements of both. Each has its place. The right choice depends on data volume, the complexity of the demand patterns, and the team's analytical capability.
Transformation typically happens through Dataflows Gen2, for lower-code, business-owned logic, or notebooks running Spark for more complex engineering work. Data Factory pipelines orchestrate the sequence, triggering Bronze ingestion, then Silver validation and cleansing, then Gold aggregation, on a schedule or an event trigger. Everything lands as Delta tables in OneLake, and it is the Gold layer that a certified semantic model reads from directly using Direct Lake mode in Power BI, without duplicating the data again.
ML methods are worth the additional complexity in three patterns. Multivariate anomalies where unusualness depends on combinations of features (a customer's transaction is anomalous in this geographic location at this time even if neither factor alone is unusual). Highly seasonal data with complex patterns that simple statistical methods cannot capture. Use cases where the cost of false positives is high enough to justify the better precision of ML methods.
Gradient-boosted models are the right choice when the forecast benefits from features beyond the time-series itself: pricing, promotions, weather, competitor activity, marketing spend, listing changes. Gradient-boosted regression (LightGBM, XGBoost) or its variants can incorporate dozens of features per SKU per period and learn the relationships. The technique outperforms classical time-series when the data and features support it. The complexity is real: feature engineering and model tuning are substantial work, and the output is less interpretable than classical models.
Statistical methods are sufficient for univariate or low-dimensional data with clear distributions. Daily revenue, weekly transaction counts, response times. The statistical methods (X-bar charts, EWMA, Z-score thresholds, median absolute deviation) handle these cases well, are interpretable to non-technical stakeholders, and run with minimal computational cost. For most operational monitoring use cases, statistical methods are sufficient and the additional cost of ML methods is not justified.
Real-time analytics pays back in five recurring patterns. Operational dashboards where the value of acting in seconds exceeds the cost of the streaming infrastructure. Fraud detection where the window to prevent loss is short. Supply chain visibility where shortages or delays need to surface before they affect customers. IoT and connected products where the device data is inherently streaming. Customer-facing operations (call centres, dispatch, warehouse) where decisions are made in real time anyway. Outside these patterns, batch analytics is usually sufficient and significantly cheaper.
ARIMA is still the right choice for relatively stable products with regular demand patterns and a single primary driver (time and seasonality). ARIMA is well-understood, statistically defensible, and fast to fit. The output is interpretable. For mature products with several years of stable history and no major external influences, ARIMA produces accurate forecasts at low computational cost. The technique is venerable but not obsolete; it remains the right answer for a meaningful subset of forecasting use cases.
Databricks is clearly the right answer when ML is the centre of gravity, your team is code-first, and you have a multi-cloud reality. Notebook-native development, mature MLOps tooling, and platform-neutrality across clouds are real advantages. If your engineering team would rather write Python than open a designer, Databricks fits how they actually work. Fabric will feel constraining.
Fabric is clearly the right answer when BI is the centre of gravity, your organisation is on M365, and your team is small or mixed-skill. The integrated experience reduces the surface area to learn and run. Fabric is also the lower-effort answer when you are already on Power BI Premium, because the licence path flows directly into Fabric capacity. For most UK mid-market businesses we work with, this describes them, and Fabric is the lower-effort answer.
Snowflake is clearly the right answer in three scenarios. Multi-cloud or AWS-first estates, because Fabric is Azure only. Heavy data engineering teams who prefer SQL-first, code-led tooling. And spiky, unpredictable query workloads, where Snowflake's per-query consumption model fits better than Fabric's fixed capacity. If any of these describe you, Snowflake is the better starting point.
Type 1 for attributes that should always reflect current state (descriptive labels, internal categorisations that get refined). Type 2 for attributes where historical context matters (geographic territory, sales rep assignment, customer tier, product hierarchy). The decision should be deliberate per attribute, not a blanket policy. Most dimensions are mixed: some attributes Type 1, some attributes Type 2. The Silver layer is where the SCD logic lives in our standard architecture.
Microsoft Learn has a free KQL learning path that covers the language well in a few hours. The Microsoft documentation for Kusto is comprehensive. Several books (most notably 'Mastering KQL' and 'Kusto Query Language Pocket Reference') cover the language in more depth. For mid-market teams adopting Real-Time Intelligence, the standard pattern is one or two team members learning KQL deeply and the rest of the team picking it up by example. The learning curve is moderate; the productivity gain is meaningful for streaming workloads.
The full comparison deck is on hoptonanalytics.com under Resources. The deck covers the six dimensions that decide it, the common myths and pitfalls, six clear scenarios for picking each platform, and two real worked examples. To discuss a specific situation, email hello@hoptonanalytics.com.
Cosmos DB is Azure's distributed NoSQL database, suited to specific use cases: massive scale, global distribution, flexible schema, document or graph workloads. For typical mid-market data and analytics work, Cosmos DB is rarely the right choice; the use cases that justify it are uncommon at mid-market scale and the cost can rise quickly without careful design. Mid-market businesses occasionally use Cosmos DB for specific application backends (mobile apps, IoT scenarios) but rarely as the analytical data store.
Market basket analysis pays back in five common applications. Cross-sell recommendations (when a customer puts A in the basket, suggest B). Email campaign segmentation (target customers who have bought A but not B). Store layout decisions (place strongly associated products near each other or deliberately apart). Range planning (identify products that anchor baskets versus products that come along). Promotional design (price the anchor product to drive baskets). Each pays back when the marginal margin from the rule-driven action exceeds the cost of implementing the rule. For most retailers and wholesalers, several of the rules surfaced are commercially material.
Anomaly detection pays back fastest in five recurring patterns. Finance integrity (unusual journals, expense outliers, supplier payment anomalies). Operations monitoring (equipment, process, service quality). Supply chain (delivery exceptions, demand spikes, stock anomalies). Customer behaviour (unusual purchase patterns, account compromise). Cyber security (unusual access patterns, data exfiltration).
Demand forecasting pays back fastest in five business shapes. Retailers with significant stock investment where over-stock and stock-out costs are material. FMCG and consumer goods with multi-channel distribution. Wholesalers and distributors with thousands of SKUs and multi-warehouse stock. Manufacturers with long production lead times. Subscription businesses needing capacity planning. The pay-back is typically through stock optimisation savings (lower stock levels at the same service level, or higher service levels at the same stock level) and through fewer stock-outs reaching customers.
For anomaly detection, use Scikit-learn for isolation forest, one-class SVM, and basic statistical methods. PyOD (Python Outlier Detection) for a wider range of algorithms including local outlier factor, autoencoders, and ensemble methods. Statsmodels for control charts and time-series anomaly detection. Prophet's anomaly detection capability for seasonal time-series. Microsoft Azure Anomaly Detector (a cognitive service) for use cases where the prebuilt service is the right answer. The choice of library depends on the specific technique; the standard data science stack covers the breadth.
For demand forecasting, use Prophet for the default implementation. Statsmodels for ARIMA and classical time-series. LightGBM or XGBoost for gradient-boosted regression. Sktime for unified time-series API across techniques. The lifetimes ecosystem (lifelines for survival analysis, lifetimes for CLV) is adjacent rather than directly used in forecasting. The standard data science stack covers demand forecasting end to end. We have notebook templates for each technique that we adapt to specific client engagements.
The two standard algorithms are Apriori and FP-Growth, both available through Python libraries. Apriori is the classic algorithm: simple, well-understood, slower on large datasets. FP-Growth is the more modern algorithm: faster, especially on large transaction volumes, with similar output. For mid-market data volumes either works; for very large datasets FP-Growth is meaningfully faster. The choice is mostly a performance one. The output is the same: a set of association rules with support, confidence, and lift.
Which platform is cheaper overall is highly situation-specific. Both can be cheaper. Fabric tends to be cheaper for steady BI-led workloads on M365 because Power BI licences flow in and capacity is predictable. Snowflake tends to be cheaper for spiky workloads where compute can scale to zero. The total cost picture also includes integration cost, training and team operability, and those often outweigh raw platform cost.
Which platform is cheaper is highly situation-specific. Fabric tends to be cheaper for steady, BI-led workloads on M365 because Power BI licences flow into Fabric capacity. Databricks tends to be cheaper for spiky compute workloads where serverless SQL can scale to zero between queries. The total cost picture also includes integration and team operability, which often outweigh raw platform cost.
Most modern transactional databases support CDC. SQL Server (with Change Tracking or CDC features), PostgreSQL (logical replication), MySQL (binlog-based replication), Oracle (LogMiner or GoldenGate), Microsoft Dynamics 365 (Dataverse change tracking), and most cloud data sources expose change streams. SaaS applications increasingly expose webhooks or change feeds. The patterns vary; the principle is universal. Sources without native CDC can be approximated through delta loads (extract only records changed since last run) or full snapshots compared period to period.
In a well-run Fabric estate, only the ingestion pipelines themselves write to Bronze, never analysts or ad hoc scripts editing files directly. Restricting write access to the automated pipeline is what keeps the layer trustworthy as a permanent record, and it's usually enforced through workspace roles and OneLake access policies rather than convention alone.
Microsoft Fabric matters for modern data architecture because Fabric implements the modern architecture as a productised platform rather than as components to assemble. OneLake provides the unified storage. The Lakehouse and Warehouse workloads provide the analytical compute. Fabric Data Factory provides the orchestration. Microsoft Purview provides the governance. The integration is what differentiates Fabric from assembling these capabilities from raw Azure components. For most mid-market businesses, Fabric is the right modern data architecture platform; for enterprises with very specific needs, raw Azure or hybrid approaches still have a place.
Version control matters for data engineering because data pipelines are code, and code without version control is a liability. Version control provides three capabilities every serious data platform needs: history of changes (who changed what and when), branching for parallel work and review, and reproducibility (the ability to roll back to a previous working state). Without version control, data engineering teams accumulate technical debt invisibly and lose the ability to recover from mistakes. The Microsoft Fabric Git integration brings version control to the platform natively.
You can trust our Fabric vs Databricks comparison, even though we are a Microsoft consultancy, because the work that fails costs us more than the work we do not win. A misplaced Fabric recommendation that fails in production damages our reputation and our renewal pipeline far more than a clean Databricks recommendation costs us in this one project. We can show specific situations where we recommended Databricks and explain why.
You can trust our Fabric vs Snowflake comparison, even though we are a Microsoft consultancy, for two reasons. We do not benefit from selling you the wrong platform. A failed Fabric project costs us reputation and renewal far more than a clean Snowflake recommendation costs us in lost revenue. And we have actively recommended Snowflake to clients where it was the right answer, including to clients who came to us assuming Fabric was the inevitable choice. The recommendations and the rationale are documented and we are happy to walk through them.
Hopton is a strong fit for Fabric implementations for three reasons. We are a Microsoft data and analytics specialist; Fabric is in the centre of our work, not a side capability. We have delivered Fabric implementations across mid-market sectors (retail, construction, FMCG, healthcare, professional services, recruitment) so the architectural patterns are well-established. We hold the Microsoft accreditations relevant to Fabric work and the team holds individual certifications including the DP-600 (Fabric Analytics Engineer) and DP-700 (Fabric Data Engineer). The work is in the centre of what we do, not at the edge.
A medallion architecture uses three layers rather than two or four because the three layers correspond to three distinct concerns. Bronze is about capture (getting data in reliably without losing fidelity). Silver is about quality (making the data trustworthy without yet imposing business semantics). Gold is about meaning (encoding business rules and structure for the consumption layer). Two layers conflate quality and meaning into one stage, which produces brittle pipelines that need rebuilding when business rules change. Four layers add an intermediate stage that rarely earns its place at mid-market scale. Three is the right balance.
Yes — Hopton will recommend against real-time when batch is sufficient. Real-time is one of the technologies most often over-specified in scoping conversations because it sounds modern. Our discovery work usually surfaces whether the requirements genuinely need real-time or whether near-real-time batch would deliver the value at much lower cost. We have written scoping reports recommending against real-time investment; the recommendation costs us a more complex engagement and earns trust for the simpler one we deliver instead.
Still have questions?
Can’t find what you’re looking for?
The first conversation is exploratory and carries no obligation. We’ll give you an honest answer to any question you have.
Book a free audit