Услуга · Data and analytics

Managing data and reference lists

Figures in reports disagree not because the software is bad but because the same counterparty is entered as three records and a product's unit is sometimes a piece and sometimes a box. The cure is not buying software but rules and named people responsible.

Duplicates
we find and merge them
Rules
who creates them and how
The owner
for every reference list
Validation
at entry rather than afterwards

What the work includes

The tools are secondary here. Everything rests on organisation: every list has a specific employee answerable for its condition.

Discuss the scope

A review of the reference data

Duplicates, empty mandatory fields, inconsistency in names and units.

Rules for filling in

How a name is written, which fields are mandatory, who is allowed to create new records.

Merging duplicates

We merge them so the documents do not lose their link to the right record.

The owners

A specific person is responsible for each list, not an abstract department.

Validation at entry

Validation that stops a duplicate being created or a mandatory field being skipped.

A regular report

A cleanliness summary: how many new duplicates appeared, where the empty fields are.

How it goes

The review takes a week or two. The clean-up itself depends on the volume and usually stretches over several months in stages.

01

We assess the scale

We count the duplicates and the records filled in against the rules.

02

We write the rules

We agree the wording with the people who actually fill the lists in, rather than issuing an order from above.

03

We remove

We merge the duplicates in turn, starting with those that distort the reporting.

04

We keep

Validation at entry and a regular report on data cleanliness.

Removing duplicates without introducing rules is pointless. Within six months they will be back: a manager will fail to find an existing record and create a new one. Rules and validation at entry first, the clean-up afterwards. Otherwise the whole job has to be done twice.

Questions and answers

Because finding an existing record is harder than creating a new one. A person searches for a familiar name, fails to find it because of a stray quotation mark or space, and creates another record. The cure is decent search and a ban on creating a record without checking first.

Finding suspicious pairs, yes, a machine can do that. Merging them is better done by eye: similar names do not always mean the same organisation, and wrongly merged records will muddle the document links for good.

The assessment takes a few days, while the time to tidy up is set by the volume. A partner list of several thousand entries takes more than a week to work through, but it can be done without stopping current work, moving from active counterparties to archived ones.

We will sort out your reference data

Tell us which reports disagree and which lists look neglected. As a first step we will assess the scale of the mess.