A company can use POS software.
Marketplaces.
Digital payments.
CRM.
Accounting.
Inventory tools.
Web analytics.
Spreadsheets.
And still be unable to answer a basic question:
How many active customers do we actually have?
The problem is not data scarcity.
It is data hygiene.
One customer appears under three identities.
The same product has five SKU names.
Refunded orders remain in revenue.
Transactions are duplicated.
Marketing uses order date.
Finance uses settlement date.
The dashboard looks sophisticated.
The underlying data may not be.
More data does not mean better data
Each system can define reality differently.
A POS identifies customers by phone number.
A CRM uses email.
A marketplace uses platform-specific IDs.
Accounting sees invoices.
One person can therefore appear to be four customers.
Products suffer the same problem.
Human employees may recognise that several names refer to the same product.
Software does not unless the mapping exists.
When inconsistent data enters analytics, technology visualises inconsistency.
When it enters AI, the problem can become more consequential because the system may use the information to recommend or perform actions.
NIST treats data quality as an AI-risk issue
The NIST AI Risk Management Framework explicitly calls for attention to data collection and selection, including availability, representativeness and suitability.[1]
NIST’s measurement guidance also notes that AI systems depend on data and methods in ways that directly connect performance with data quality and representativeness.[2]
Data hygiene is therefore not merely an IT housekeeping exercise.
It is part of risk management.
Dashboards face the same problem before AI does
Consider a revenue dashboard.
The same purchase is recorded twice.
Refunds are missing.
Cancelled orders remain counted.
Marketplace transactions arrive from both an API and a spreadsheet.
A more advanced chart does not fix the numbers.
Google Analytics uses a unique transaction ID to deduplicate purchases and correctly process refunds.[3]
That is not a universal architecture requirement.
It illustrates an important principle:
every transaction should have a stable identity.
Start with identifiers
A company needs a minimum set of unique IDs.
Transaction ID.
Customer ID where appropriate.
Product or SKU ID.
Supplier ID.
Invoice ID.
Labels may change.
IDs ideally should not.
Without stable identifiers, reconciliation becomes a matching exercise based on names and assumptions.
Do not use unnecessary personal data as IDs
Google specifically advises that transaction IDs should not contain information capable of identifying individual customers.[3]
The broader principle is useful.
A transaction does not need to use a customer’s phone number, email address or national ID as its identifier.
Use technical identifiers and collect only what is necessary.
Define transaction status
Many dashboards become unreliable because “sale” is poorly defined.
Is an order:
created?
paid?
processing?
shipped?
completed?
cancelled?
refunded?
A business needs consistent logic around when an order becomes revenue for each operational use case.
Without that, two teams can report different sales numbers while both believe they are correct.
Refunds are not a minor detail
Refunds affect revenue, product performance, customer analysis and marketing attribution.
They should link back to the original transaction.
Otherwise, businesses lose the ability to answer basic questions about what was sold, reversed and retained.
Build product master data
Products frequently have different names across marketplaces, stores and inventory systems.
A product master should provide a canonical structure:
product ID;
official name;
category;
variant;
unit;
active status.
Display names can vary by channel.
The underlying product identity should not.
This is how reliable product-mix analysis begins.
Customer data requires more restraint
Not every company needs a complete customer master.
A high-volume physical retailer may legitimately serve anonymous customers.
Subscription, loyalty, CRM and B2B businesses have different requirements.
Where customer identity is needed, companies should collect data proportionate to the purpose.
More information is not automatically better.
Indonesia’s PDP law makes accuracy relevant too
Indonesia’s Personal Data Protection Law does not deal only with cybersecurity.
It requires personal data processing to be accurate, complete, non-misleading, current and accountable.[4]
Article 29 specifically requires controllers to ensure the accuracy, completeness and consistency of personal data and undertake verification.[4]
Customer-data quality can therefore have legal and governance dimensions, not only analytics implications.
Assign ownership
Ask a simple question:
Who owns customer data?
Marketing?
Sales?
IT?
Finance?
Operations?
The most useful answer often separates business ownership from technical stewardship.
Business owners define meaning and permitted use.
Data stewards monitor quality.
IT manages systems.
Security manages access.
Privacy and legal teams manage lawful processing.
Without accountability, everyone can modify data while nobody owns its reliability.
One metric needs one definition
Consider “active customer”.
Marketing may mean someone who opened the app.
Sales may mean someone who bought recently.
Finance may mean someone with an open invoice.
Customer success may mean someone under contract.
All can be valid.
The problem is showing one dashboard label called “active customers” without defining it.
Important metrics need:
a name;
definition;
formula;
source;
owner;
refresh cadence.
That becomes the metric dictionary.
Lineage explains where a number came from
Data lineage tracks information from origin to output.
POS → transaction database → warehouse → transformation → dashboard.
When revenue suddenly changes, the team can investigate whether the source changed, a mapping failed, refunds disappeared or time-zone logic shifted.
NIST also encourages documentation of data provenance including sources, origins, transformations, dependencies and metadata.[5]
Time zones can move revenue into another day
A marketplace may store UTC.
A POS may use Jakarta time.
A cloud system may follow a server default.
An order made shortly after midnight can therefore fall into different reporting dates across systems.
That changes daily sales, campaign attribution and service metrics.
Set a canonical timestamp and conversion rule.
Missing does not mean zero
A blank field can mean many things.
Unknown.
Not applicable.
Not collected.
Pending.
Treating every missing value as zero creates artificial certainty.
Businesses should distinguish absence of information from a real value of zero.
AI does not magically resolve business definitions
AI can help identify duplicates.
Match records.
Classify text.
Detect anomalies.
But it does not automatically know whether two similar product names refer to the same SKU.
That requires a business rule.
AI can assist data cleaning.
It cannot replace organisational definitions.
Do not start with the data lake
When businesses discover data problems, they often jump to large technology projects.
Warehouses.
Lakehouses.
AI platforms.
Master-data tools.
These can be useful.
But many SMEs and mid-sized businesses need something simpler first:
one SKU master;
one transaction identifier;
one channel taxonomy;
one refund rule;
one customer definition;
one chart of accounts.
Technology should support the definitions—not substitute for them.
Audit five common error types
Duplicate
Are transactions, customers or products recorded twice?
Missing
Are required fields absent?
Inconsistent
Do categories, channels or statuses use multiple formats?
Invalid
Are dates, amounts or codes outside expected rules?
Stale
Is data too old for the decision being made?
A simple score across those areas can make data quality visible.
Dashboards should show their own reliability
A good dashboard can also display:
last refresh time;
data completeness;
unmatched transactions;
failed imports;
unknown product mappings;
reconciliation differences.
Managers then see not only the number but how trustworthy it is.
AI needs an authoritative-source hierarchy
Internal AI systems should know which source is authoritative.
Product price → ERP.
Order status → order-management system.
Invoice → accounting platform.
Consent → CRM or privacy system.
Policies → approved document repository.
When two systems disagree, a hierarchy prevents AI from arbitrarily choosing whichever value it finds first.
Data hygiene never ends
Products change.
Channels are added.
Integrations break.
Employees create spreadsheets.
Customers update details.
Data quality is therefore a recurring operating discipline.
NIST frames AI risk management as continuous throughout the system lifecycle.[1]
Data governance needs similar continuity.
Eight foundations before scaling analytics
Before investing heavily in dashboards or AI, make sure the business has:
- a unique transaction ID;
- product/SKU master data;
- a clear customer definition;
- standard transaction statuses;
- refund and cancellation logic;
- channel taxonomy;
- a metric dictionary;
- clear ownership and access rules.
Those eight elements are not glamorous.
They are what make sophisticated systems trustworthy.
Dashboards cannot repair upstream chaos
A dashboard is an output layer.
If upstream systems are inconsistent, it produces a cleaner-looking version of the inconsistency.
AI can go one step further by turning poor data into recommendations that sound persuasive.
The strategic question should therefore not begin with:
“Which AI should we buy?”
or:
“Which dashboard should we build?”
It should begin with:
“Is our source data consistent, traceable and trustworthy enough to support a decision?”
Businesses do not lack data.
What remains scarce is data clean enough to trust.
And in an AI-enabled organisation, that scarcity becomes increasingly important.
- [1] NIST. Artificial Intelligence Risk Management Framework 1.0. (airc.nist.gov)
- [2] NIST AI RMF Playbook — Measure guidance on data quality and representativeness. (airc.nist.gov)
- [3] Google Analytics Help. “Minimize duplicate key events with transaction IDs.” (support.google.com)
- [4] Republic of Indonesia. Law No. 27/2022 on Personal Data Protection, including data accuracy, completeness, consistency and verification requirements. (jdih.komdigi.go.id)
- [5] NIST AI RMF Playbook — data provenance documentation. (airc.nist.gov)
- [6] Google Analytics Help. Ecommerce event measurement guidance. (support.google.com)
- Editorial Notes:
- Google Analytics transaction-ID practices are used as an illustrative data-engineering principle, not a mandatory architecture for every business system.
- Data quality and data privacy overlap but are not identical disciplines.
- Clean data does not guarantee correct AI output; model behaviour, governance, context, human oversight and system design still matter.
Published: September 15, 2026




