No — Power BI and Fabric cannot fix data quality problems on their own. Both platforms can surface data quality issues clearly and can apply cleansing logic during transformation, but neither can decide what your definition of an active customer should be, or reconcile two genuinely conflicting golden records. That is a business decision, not a technical one, and it has to be made by people who own the data, not by the reporting layer sitting on top of it.
Purview includes data quality scoring and profiling capability that can flag issues such as missing values, out-of-range figures, and format inconsistencies against rules you define. It surfaces and scores issues; resolving them - deciding which record wins, fixing the source system, or building a reconciliation rule - is still a deliberate design and governance decision.
Yes — you can get the one-page framework as a template. We share our working version with clients on request. No commitment required. The framework belongs to whoever is using it: tailor it to your environment, change the naming patterns to suit your conventions, extend it where the basics do not cover your situation. The point is for it to fit your team, not to be copied verbatim.
Most mid-market businesses have an MDM problem well before they have MDM tooling. Duplicate customer records across a CRM and an ERP, inconsistent product hierarchies between a webshop and a warehouse system, and conflicting cost centre structures between finance and operations are common at businesses far smaller than the enterprises MDM software is traditionally marketed to. The problem exists regardless of whether you have named tooling for it.
Purview has value even in a Power BI-only estate, particularly for lineage (tracing where a number in a report actually originated) and for applying consistent sensitivity labels across reports. Its value increases significantly once Fabric is in play, because Purview then has visibility across the wider lakehouse, warehouse, and pipeline estate, not just the semantic model layer.
Master data management is folded into wider Power BI, Fabric, and data governance engagements rather than sold as a separate, generic MDM project. In our experience, master data problems are best fixed in the context of the specific reports and decisions they affect, rather than as an abstract cleanup exercise disconnected from a concrete business use.
Hopton handles both Irish data protection requirements and UK GDPR, as first-class requirements rather than an afterthought. Irish clients sit under the EU GDPR framework enforced by the Data Protection Commission rather than the UK's ICO, and where an engagement spans both jurisdictions - a UK head office with an Irish subsidiary, or vice versa - we build the governance model, including row-level security, data classification, and retention rules, so it holds up against either regulator rather than defaulting to whichever one we know best. This is the same governance discipline we already apply to clients running multi-currency consolidation and cross-border reporting between UK and Irish entities.
Hopton uses Microsoft Purview where it adds genuine value, particularly for lineage and sensitivity labelling on Fabric-based estates. We do not deploy Purview by default on every engagement; smaller Power BI-only estates sometimes get equivalent value from simpler governance documentation without the additional platform overhead.
No — data governance does not need a Chief Data Officer, particularly not in mid-market organisations. Governance does not need senior structure to begin. It needs a small set of rules, a few people who agree to follow them, and the discipline to keep doing so. Waiting for a CDO is one of the most common reasons governance does not start. The environment continues to sprawl while the role is being scoped, and the longer the wait, the harder the eventual cleanup.
Fixing data quality means both cleaning up source systems and handling issues in the reporting layer, and the right mix depends on severity. Minor formatting inconsistencies can reasonably be normalised during data transformation in Fabric or Power Query. Genuine duplicate or conflicting master records need to be fixed at the source, or reconciliation logic needs to be built and maintained deliberately, because patching the same problem invisibly in every downstream report is fragile and eventually breaks.
Yes — the Checklist applies particularly to AI outputs. AI outputs land on a desk with confidence and no provenance, and the Checklist forces the same trust signals human-produced analysis used to carry. The four questions (Sources, Freshness, Limits, Action) are the same four blocks that appear in the Model Trust One-Pager from our AI guide. Both come from Nick Kelly's work on trust storytelling.
No — the Project Gate does not apply to AI projects only. We use it for every project, not just AI. The risk of admiration projects is universal. The Project Gate is particularly important for AI projects because they tend to attract enthusiasm and budget faster than they accumulate operational readiness, but the framework applies anywhere project work is being committed to.
Duplicate records happen in the first place usually through independent data entry across systems that were never designed to be the master for the same entity - a sales rep creating a new CRM contact instead of searching for the existing one, a new supplier being set up in the ERP without checking whether they already exist under a slightly different name. Left unmanaged, this compounds over years.
To find out how bad your own data quality problem is, ask us for a discovery conversation. Most clients underestimate the scale of the issue until we walk through a specific reconciliation exercise on their actual customer or product data, which is usually more revealing - and quicker to do - than it sounds.
Email hello@hoptonanalytics.com describing your current systems and where you suspect (or have already found) inconsistencies. We will scope an initial discovery conversation and, where useful, a specific reconciliation exercise on your most business-critical entity before recommending a wider programme of work.
To make governance routine rather than ceremonial, embed it — do not announce it. Avoid the all-hands meeting, the kick-off, the policy email. None of those change behaviour. Instead, embed the rules into work the team is already doing. The next workspace someone creates uses the naming convention. The next dataset deployed has a named owner. The next report promoted goes through the checklist. After a few weeks, the rules are how things get done. After a few months, the team would not know how to work any other way.
We decide which entities need formal master data management first by starting with whichever entity appears most often across the reports that matter most, and where a mismatch would be most damaging if it went unnoticed. For most mid-market businesses, that is customer and product; for some it is cost centre or location. We prioritise based on where a data quality issue would actually change a decision, not by trying to fix everything at once.
We handle data security and compliance by building everything within your own Azure tenant - your data never leaves your environment. We adhere to ISO 27001-aligned practices and can work within your existing governance and compliance frameworks. Row-level security, workspace management, and endorsed datasets are standard deliverables in every Power BI engagement, not optional extras.
Access control answers who can see or change what; KPI governance answers whether the number itself can be trusted, and both need the same discipline. Each KPI is documented once: its formula, the table and column it draws from, a named owner accountable for it, and where it is meant to be used, so two departments cannot end up with two different definitions of active customer or gross margin without anyone noticing. Changing a KPI's definition goes through the same promotion process as any other change to a certified model, not a quiet edit inside one report. Ownership sits with a named person rather than a team, for the same reason access control uses row-level security rather than ad-hoc filters: a rule with nobody accountable for it degrades over time. In practice this pairs directly with access control. Workspace roles and row-level security decide who can open a report and what rows they see inside it; the KPI's documented definition and owner decide whether what they are looking at means what they think it means. Get one right without the other and the result is either a secure report showing the wrong number, or a correct number that anyone can quietly redefine.
Hopton uses the BVQ at four points in our process. In discovery: we open with the BVQ and ask the client to complete it with us. In proposals: every proposed workstream has a BVQ attached. In sprints: we review the BVQ at the start of every sprint to check whether the decision, the owner, or the deadline have shifted. In delivery: we present findings against the BVQ. Not what we built, but the decision it supports, the evidence behind it, and what the client should do next.
Hopton handles a client whose master data is a mess directly, without pretending the reporting layer alone can fix it. We identify the highest-impact reconciliation work, agree who owns the decision on each entity, and sequence the analytics build around getting the most business-critical entities right first, rather than promising a dashboard can paper over unresolved data quality problems.
For a mid-market business, data governance should be lightweight, business-led and incremental — not the enterprise committee model. Pick the metrics that cause the most disputes, give each a business owner, agree and certify the definition behind it, add quality controls, then expand. Executive sponsorship keeps it honest, but the day-to-day stays with the people who use the data. You are building a habit, not a bureaucracy.
Data lineage in Purview helps with data quality by making it possible to trace a suspicious number backwards through every transformation step to its source, rather than debugging blind. This turns "the numbers don't match" from a multi-day investigation into a lineage trace that usually takes minutes, which matters enormously when a finance director is asking why a report has changed.
Governance fits into the Analytics Acceleration Programme naturally, as an ongoing discipline rather than a one-off. Governance is not a one-off project. The framework needs maintenance, the quarterly reviews need running, the rules need extending as the environment grows. The AAP gives you a dedicated allocation of consulting days each month over twelve or twenty-four months, which is the right shape for ongoing governance work. Most AAP engagements have governance as a standing component rather than a separate workstream.
Poor data quality shows up to a business user as a dashboard number that "doesn't look right" and gets quietly ignored, or worse, a decision made on a number that looks right but is not. The most damaging data quality problems are the ones that are subtle enough to go unchallenged rather than obviously broken.
Transactional data is the record of events: an order, an invoice, a stock movement. Master data is the relatively stable reference information those events point back to: which customer placed the order, which product was invoiced, which warehouse the stock moved through. Transactional data changes constantly; master data should change rarely and deliberately.
The BVQ is broader and applies to every piece of work, including small ones. The Project Gate is a sharper version applied to projects specifically, with an emphasis on operational readiness. A project with a good BVQ might still fail the Project Gate if the data owner cannot unblock access in time, or if the decision window closes before the project finishes. Both checks are useful. The Project Gate is what we use when the question is 'is this project ready to start' rather than 'is this work worth doing'.
The Trust Storytelling Checklist differs from the BVQ in its position in the workflow. The BVQ sits at the start of work, before anything begins, ensuring the work is worth doing. The Trust Storytelling Checklist sits at the end of work, before anything is delivered, ensuring the work is safe to act on. Both are essential. Both are short. Both make the difference between consultancy that produces deliverables and consultancy that produces decisions.
Governance documentation only needs to be one page to start. The framework we use covers the Business Value Question, the four practices, and the quarterly review on a single page. The argument for the page is longer, naturally. But the page itself, the thing your team lives by, has to be short. Nobody refers to a fifty-page policy document at 9am on a Tuesday. They refer to a one-page reference pinned to the team's shared workspace.
Once set up, MDM and data quality require meaningful but bounded ongoing effort: new records still get created imperfectly, and someone needs to own periodic review and reconciliation. This is one of the reasons we build data governance and data quality maintenance into ongoing support arrangements rather than treating it as a one-off project that is "done" at go-live.
Access controls come down to two questions: who can see what data, and who can change what content. For data access, row-level security in the semantic model is the right pattern. For content access, workspace permissions follow a small number of standard roles (viewer, contributor, member, admin). Avoid bespoke role structures: they become unmanageable quickly. Standard roles are easier to maintain, easier to audit, and easier to understand for everyone in the team.
No — Purview is not the same thing as master data management. Purview is primarily a cataloguing, lineage, and governance tool; it helps you see and control your data estate. MDM is the broader discipline of actively defining and maintaining golden records for your core entities. Purview can support an MDM programme by surfacing where inconsistencies exist, but it does not itself decide or enforce which record is the true one.
Data quality assessment is part of the Establish phase discovery work in our AAP. The scope and depth of remediation work needed varies significantly by client, and where it is substantial, we scope it explicitly rather than absorbing an open-ended cleanup exercise inside a fixed-price Build phase.
No — governance should not be a separate workstream from delivery. Governance must be embedded in delivery, applied to the dataset being deployed this week, the report being promoted, the model being shared. Anything else is theatre. A governance team that writes rules separately from the delivery team that builds reports rarely produces governance that anyone follows. The rules have to live where the work happens.
The four data-governance practices that follow the Business Value Question are naming conventions, named ownership, a promotion process, and access controls. Together with the BVQ and a quarterly review, they form a six-section framework that fits on one page. Each practice is simple. Together they prevent most of the problems that ungoverned environments create. The four practices are intentionally minimal. They cover the basics every environment needs without the heavy structure most governance attempts get caught up in.
The most common data quality issues in mid-market analytics estates are duplicate customer or supplier records created by inconsistent entry across systems, inconsistent product or account hierarchies between operational systems, missing or stale reference data (old cost centres, discontinued products still appearing in reports), and definitional drift, where the same metric name means different things in different systems because no one agreed a single definition.
A single source of truth means every team is looking at the same numbers, built from the same definitions. In practice it comes from three things working together: a certified semantic model where measures like "revenue" or "active customer" are defined once and reused everywhere, documented ownership so everyone knows which report is authoritative, and governance that stops parallel, unofficial versions taking root. It is an outcome of governed architecture, not a tool you switch on. Clients who reach this point consistently describe the same shift: leadership stops debating whose numbers are right and starts spending that time on the decision itself.
Purview is Microsoft's data governance and cataloguing service. It scans data estates (across Azure, Fabric, and beyond) to build a searchable catalogue of what data exists and where, tracks data lineage so you can see where a number in a report actually came from, applies sensitivity labels and access policies consistently, and gives a central place to see data quality metrics across your estate.
A good naming convention is short enough that everyone remembers it. We typically use a domain-purpose-environment pattern: domain (Finance, Sales, Operations), purpose (Pipeline, Cash Flow, Stock Performance), and environment (Dev, Test, Prod). 'Finance - Sales - Production' is meaningful. 'Final Report v3' is not. The convention should cover workspaces, datasets, reports, and measures consistently. Longer or more complex conventions get followed inconsistently and quickly become noise.
A good promotion process uses three environments: development, test, and production. New work is built in Dev. It is validated in Test by people other than the builder. It is promoted to Prod only after that validation passes. For a small team, the promotion process can be a checklist: model documented, security tested, refresh schedule confirmed, two stakeholders signed off. The point is to prevent accidents, not to slow delivery.
A practical first step towards MDM, for a business with no formal programme today, is a reconciliation exercise: identify your core entities, compare how each is represented across your two or three most important systems, and agree which system (or which combination) is authoritative for each field. This does not require buying dedicated MDM software; it requires a decision-making exercise that most mid-market businesses have simply never done formally.
For a mid-market organisation, data governance means: agreed definitions for every key metric (what does 'revenue' include?), clear ownership of datasets and reports (who is responsible for accuracy?), documented data lineage (where does each number come from?), access controls (who can see what?), and a process for managing changes. Governance is not a one-time project - it is the operational discipline that keeps analytics trustworthy over time.
The first 90 days of a governance rollout run week by week: in weeks 1-2, agree the one-page framework, identify owners for every existing dataset and shared report, and sort the most obvious naming inconsistencies. Weeks 3-6: apply the framework to all new content, set up the promotion process, establish workspace roles, document the rules. Weeks 7-10: bring existing content into compliance, rename what needs renaming, reassign orphans, retire content no longer used. Weeks 11-12: run the first quarterly review, confirm what is working and what is not, adjust where the rules turned out to be wrong or missing.
The quarterly data-governance review checks four things. Ownership refresh: are the named owners still the right people, has anyone left or changed role. Naming compliance: has new content followed the convention, are there outliers to clean up. Sprawl audit: what new workspaces, datasets and reports have appeared, are they all needed. Framework fitness: are the rules still matching the environment, anything missing or unnecessary. Thirty minutes is usually enough. The point is to keep the framework alive, not to redesign it every quarter.
Each block of the Business Value Question (BVQ) captures one thing, starting with the decision: what specific decision will this work inform, in plain language. The decision owner: a named person, not a team or committee, with authority to act. The action: what will the decision owner do differently, specific and observable. The measurement: how will we know if the decision was successful, with concrete metrics. The deadline: when does this decision need to be made. The cost of inaction: what happens if this decision is not made or is made without data.
If you cannot name a decision owner for a piece of analytics work, then the work is not ready to start. A specific named person, not a team, not a committee. They must have the authority to approve, defer, or reject the recommendation. If you cannot name them, the project is still conceptual. This is one of the most valuable findings from running the BVQ: it surfaces work that everybody assumes will be useful but nobody can defend at the level of who will actually act on it.
A "golden record" is the single, agreed, trusted version of a master data entity - the one definition of a given customer, product, or supplier that all systems should ultimately treat as authoritative, even if the raw data about that entity still lives in several source systems.
A certified dataset is one with an explicit owner, documented lineage, quality checks, monitoring, and a recertification cadence, so anyone building on it knows it is trustworthy and current. Certification turns governance from a policy document into something operational, and its lineage lets you answer why did this number change with a fact rather than a guess.
A single source of truth in analytics is one agreed definition for each key metric, more than it is a single physical database. When departments calculate active customer or net revenue differently, no report can be trusted. A single source of truth centralises the business logic — standardised definitions held in a governed semantic model — so every report inherits the same calculation and the numbers finally agree.
Data lineage is the documented record of where each piece of data comes from, how it has been transformed, and where it is used. It matters because when a number looks wrong, lineage tells you exactly where to look. Without it, debugging a data quality issue means tracing the problem manually through every system and transformation. Fabric and Power BI both have built-in lineage views that we configure and maintain as part of every engagement.
Incremental data quality is the practice of treating quality rules as something that grows over time rather than a one-off cleansing project. A new rule gets added whenever a new failure mode is discovered, and because the raw layer is immutable and the pipeline replayable, that rule can be run back across the full history immediately rather than only applying from the day it was written.
Master data management (MDM) is the discipline of agreeing, in one place, what the core reference entities in your business actually are - your customer list, your product list, your chart of accounts, your cost centres - so that every system and every report refers to the same version of the truth rather than each holding its own slightly different copy.
The Business Value Question is the single discipline that anchors every piece of work we do. Six blocks: the decision, the decision owner, the action, the measurement, the deadline, and the cost of inaction. We complete the BVQ with the client at the start of every engagement, every proposal, every sprint. If we cannot complete it, the work stays in the framing stage. It does not start until it can. The BVQ prevents technically impressive work that changes nothing in practice.
The Hopton Governance Workshop is a one-day structured workshop with your team and ours. Morning: assess your current environment together. What exists, who owns it, where the gaps are, what has been tried before. Afternoon: work through the BVQ and the four practices for your environment. Output: a working one-page framework, a named-owner spreadsheet for your existing estate, and a 90-day plan. Fixed-scope, no obligation to continue afterwards. Many clients do, but the framework belongs to them either way.
The Project Gate is a four-point check that prevents what we call 'admiration projects': work that looks impressive in demos but changes nothing in practice. All four must be answered before work is committed to. Decision owner: a specific named person with authority and a decision window open during the project. Concrete action: the decision owner can describe what they will do differently. Measurement window: impact can be measured within the project timeline. Data owner: a named person who can unblock data access. If a request cannot pass this gate, it is not ready to consume the team's capacity.
The Trust Storytelling Delivery Checklist is four questions every deliverable must pass before it goes to a stakeholder: Sources (can we show exactly which data sources this analysis is based on), Freshness (can we tell the client how current this data is), Limits (have we stated where this analysis is weak or based on assumptions), and Action (does the deliverable say what the client should do next). If any answer is no, the deliverable goes back to the team. The Checklist ensures the work is safe to act on, where the BVQ ensures the work is worth doing.
You should extend the data-governance framework beyond the basics when the cost of not extending becomes obvious. Sensitive data classification becomes worth adding once you have more than a couple of domains. Lineage tracking helps when datasets feed each other in non-trivial chains. A lightweight data dictionary helps when more than a handful of people are building reports. Premature governance is as bad as no governance: both produce overhead without value. The discipline is to add things only when the absence becomes visible as a real problem.
The full data governance guide is on hoptonanalytics.com under Resources. The guide covers why governance fails, the BVQ in detail, the four practices, the Trust Storytelling Delivery Checklist, the Project Gate, the 90-day rollout, the quarterly review, and how to extend the framework as the environment grows. To discuss your specific situation or request the one-page template, email hello@hoptonanalytics.com.
MDM work fits in an analytics project before building dashboards, ideally, at least for the entities that matter most (customer, product, and whichever master entity your specific business runs on). Building dashboards on top of unreconciled master data means rebuilding the semantic model later once the inconsistencies surface, which almost always costs more than addressing them up front.
Hidden complexity in analytics projects usually comes from four places, most commonly: ETL logic that has to handle far more source exceptions than anyone estimated, governance requirements that surface once real users and real permissions are involved, validation rules that reveal disagreements between departments about what a number means, and cross-department alignment that turns out to need more meetings than the plan allowed for. None of these are unusual. They are just rarely priced in up front.
Most data governance initiatives fail in five recurring patterns. First, the 50-page policy document that becomes evidence governance exists rather than something that shapes behaviour. Second, waiting for a Chief Data Officer or formal committee before starting. Third, treating governance as a separate workstream from delivery. Fourth, confusing governance with control, which kills delivery and gets routed around. Fifth, no review cadence, so the framework drifts out of date and becomes irrelevant.
Organisations struggle with data governance because governance tends to be treated as a compliance exercise rather than a delivery discipline. Most analytics projects focus on building reports and move on - the governance layer gets skipped because it adds time and does not produce a visible output. The result is an analytics estate that works at the start and degrades as the business changes. We embed governance into every engagement from day one because retrofitting it later is significantly harder.
Master data quality matters so much for Power BI and Fabric because a semantic model is only as trustworthy as the master data underneath it. If "customer" means something slightly different in your CRM than it does in your finance system, a Power BI report joining the two will silently produce numbers that do not reconcile, and no amount of good dashboard design fixes a broken join underneath it.
Every dataset, every workspace, every shared report needs a named owner: a specific person, not a team or 'IT' or 'the data team'. The owner is responsible for refresh schedules, documentation, security, and decisions about change. Without named ownership, content becomes orphaned, refresh schedules drift, and nobody can authorise change. When that person leaves the organisation or moves to another role, the ownership transfers explicitly.
The Business Value Question (BVQ) is so central to governance because before any naming convention or ownership rule, governance starts with a question: what decision is this work supposed to inform? If you cannot answer that, the work is not ready, regardless of how well it is named, owned, or promoted. Naming and ownership protect work that is worth doing. The BVQ decides whether the work is worth doing in the first place. Without it, you can run a perfectly governed environment full of technically excellent reports that change nothing.
The Trust Storytelling Checklist is so important at the moment of delivery because the moment of delivery is when trust is built or broken. Not policy documents. Not steering groups. The handover of a dashboard, a report, or an analytical finding to a stakeholder. A confident, well-sourced delivery makes the decision feel safe to take. A vague, hedged delivery does the opposite, regardless of how good the underlying analysis is. Stakeholders do not resist data because they do not understand it. They resist it because acting on it carries personal risk. The Checklist reduces that risk.
You should not report directly off raw source data because raw operational data is not written for reporting, it is written for the system that captures it. It contains duplicates from retries, records mid-transaction, inconsistent keys across source systems, and no agreed business logic for things like an active customer or a completed order. Reporting directly against it means every analyst rebuilds those rules themselves, differently, which is exactly how two dashboards end up disagreeing. Validating and modelling the data once, in Silver and Gold, means everyone downstream works from the same trusted answer.
The raw data layer should be immutable because it is your only honest record of what the source system actually sent you, and your only way back if something downstream goes wrong. If a transformation bug corrupts the Silver layer, or a new business rule needs applying to two years of history, an immutable Bronze layer means you can simply reprocess from the original data. If Bronze has already been cleaned, reshaped or overwritten, that option is gone, and you are stuck trying to reconstruct history from a Gold table that was never meant to hold it.
Yes — Hopton will tell you if your existing governance is fine and you don't need help, and we sometimes do. If your governance is working, we will say so and recommend you stick with it. The audit of your current state is part of the workshop and we report honestly on what we find. The point of the engagement is to help you make governance work, not to sell our help if you do not need it.
Still have questions?
Can’t find what you’re looking for?
The first conversation is exploratory and carries no obligation. We’ll give you an honest answer to any question you have.
Book a free audit