Skip to content
Lorendix

Data engineering · Master data and entity resolution

One customer, however many systems they appear in

Matching and deduplicating customers, products and suppliers across systems, establishing one trusted record for each, and putting in place the process that stops the duplicates returning next quarter.

Where this starts

Deduplication that has to be done again

Duplicate records are rarely a one-off mess. They are the steady output of a process: a customer created again because the search did not find them, a supplier added under a slightly different name, a product coded differently in each system. Cleaning them once without changing the process that creates them buys about a quarter before the count is back where it began.

What we usually find

  • The same customer held several times under different spellings
  • Customer counts that differ by system, with no way to say which is right
  • Suppliers paid twice through two records for one company
  • Products coded differently in sales, stock and finance
  • An annual cleansing exercise that never stays clean
  • The same person contacted three times by the same campaign

Our position

Cleaning duplicates without fixing the point where they are created is a recurring cost dressed up as a project.

What the work covers

Match, merge and prevent

Matching is the visible part. Stewardship of the trusted record, and prevention at the point of entry, are what make it last.

Matching

Exact rules where shared identifiers exist and probabilistic matching where they do not, tuned against your own records so the balance between missed and false matches is a deliberate choice.

Survivorship

Rules for which source wins each attribute when records merge, so the trusted record takes the most reliable address, the latest contact and the correct legal name.

The trusted record

One authoritative record per customer, product or supplier, with every source record linked to it, so each system keeps its own copy while agreeing on identity.

Reference data

Shared lists such as countries, units and categories managed once and distributed, rather than maintained separately and differently in every system.

Prevention at entry

Matching applied at the moment a record is created, so a user is shown the likely existing match before they add a new one.

Stewardship

A review queue for uncertain matches, owned by people who know the customers, with their decisions feeding back into the matching rules.

How it stays clean

Matching as a standing process

A trusted record is only as good as its most recent update. These keep it current.

Continuous
New and changed records matched as they arrive rather than in an annual cleanse, so a duplicate is caught while there is still only one of it.
Match tiers
Automatic merges only above a strict certainty, a review tier beneath it, and every reviewer decision recorded to sharpen the rules.
Unmerge
Every merge reversible with its history intact, because a wrong merge does more harm than a missed one and has to be recoverable.
Duplicate rate
New duplicates tracked by source and entry point, which shows precisely where the process still needs to change.
Shared keys
The trusted record’s identifier sent back to every source system, so reporting and operations join on the same key.

What you are left holding

  • One trusted record per customer, product and supplier
  • Survivorship rules agreed per attribute
  • Duplicate checks at the point of entry
  • A stewardship queue owned by the business
  • A duplicate rate tracked by source

Tell us how many customers you have

If the answer depends on which system we ask, that difference is exactly where we start.