Here is a scenario that comes up often in conversations with HR Managers.
Someone asks for a headcount report. It should be straightforward. But then there are the contractors. Are they in? What about the two people on long-term sick? The three on secondment? The person whose leaver date is not yet in the system? The apprentice who is technically employed by the training provider?
By the time the report is pulled together, it has taken twice as long as it should, there are caveats on every number, and the confidence in the final figure is, let us say, partial.
If this feels familiar, you are not alone. It is not a reflection of how well HR is run. It is a reflection of how data tends to accumulate in real organisations, and the headcount example is only the most visible version of it. Ask the same organisation for a turnover rate, an absence figure or a training spend and you will usually find the same pattern underneath: a number that is roughly right, produced by someone who knows exactly which bits of it they would not want to defend in a meeting.
Where the mess usually comes from
HR data does not arrive messy. It becomes messy over time, and it does so quietly, through a series of decisions that each made perfect sense at the time.
Most organisations are running more than one source of employee data. There is the HR Information System (HRIS), there is payroll, there is at least one spreadsheet that someone built during a crisis three years ago and never retired, and there is a shared drive with a folder called “Final” and another called “Final v2”. These sources are rarely reconciled in full. Each one is trusted for something, and nobody is entirely sure which one wins when they disagree.
Then processes change and the data does not catch up. A new onboarding flow is introduced, but the old fields are still there and still half-populated. A restructure creates new departments, and half the workforce is recoded while the other half keeps their old cost centre because nobody got round to it.
Definitions are the other big one, and in my experience the most underrated. What counts as a leaver? Does a fixed-term contract ending count the same as a resignation? When does someone’s start date begin, the day they signed or the day they walked in? Who is “active”? Is a person on maternity leave active? Is a person on a career break? These sound like pedantic questions until two people produce two different headcount figures for the same board pack, and both of them are right by their own definition.
Add to that a system migration that did not transfer everything cleanly, a handful of manual workarounds that became permanent, and several different people recording the same information in slightly different ways over several years, and you have the situation most HR teams are actually working in.
None of these is catastrophic on its own. Together, they create a situation where you are not entirely confident in your own numbers, which makes it very difficult to report upwards with any authority. That is the bit that bothers me. The number is usually the easy bit. It is the confidence behind the number that has gone missing.
Why it matters more than it used to
For a long time, messy HR data was an inconvenience rather than a risk. You produced the report, you added the caveats, and everyone moved on. Two things have changed that.
The first is regulatory. The Employment Rights Act and the Fair Work Agency mean that employers may be asked to produce records to demonstrate compliance, and to produce them at short notice rather than after a fortnight of tidying up. If your working time records, contract types, holiday accruals or leaver dates are inconsistent across systems, that is now a practical risk as well as a reporting problem. “We think it is about right” is not a position you want to be in when someone with enforcement powers asks the question.
The second is strategic. Boards and senior leadership teams are increasingly expecting HR to contribute meaningful insight, not just activity reports. They want to know whether turnover in a particular team is a problem or a blip, whether the absence trend is seasonal or structural, and what the workforce will cost next year under two or three different scenarios. That is very hard to do when you are not confident in your underlying data, because every insight has to be built on a number you would rather not be asked about.
There is a quieter third reason as well. Messy data is expensive in time. Every report that takes twice as long as it should, every reconciliation done by hand, every “can you just check that figure” email is capacity that HR does not have to spare, and it is one of the reasons turnover, absence and training spend so rarely get tracked consistently even when everyone agrees they matter. Tidying the foundations is not glamorous, but it is one of the few things that gives time back rather than taking it.
Where to start
The temptation is to try to fix everything at once. That rarely works. It produces a very long list, a lot of good intentions and, six months later, the same spreadsheets. A more useful approach is to start with one or two areas and get those right before moving on.
Three questions tend to be a useful starting point.
What is your primary system of record for employee data? And is it actually being used as such, or is the spreadsheet on someone’s desktop quietly doing the real job? Most organisations can name a system of record. Fewer can say with a straight face that it is the place people actually go when they need the answer.
How confident are you in your headcount figure right now? Could you produce it in five minutes with no caveats? If the honest answer involves the word “depends”, that is worth noticing. Headcount is the number everything else is divided by. If it is soft, every rate you calculate from it is soft too.
When did you last check whether your HRIS data matches your payroll data? Not roughly. Actually matches, person by person. Payroll is usually the most accurate data in the building, because people tend to notice when it is wrong. If the HRIS disagrees with it, the HRIS is usually the one that needs attention.
You do not have to answer these perfectly. But asking them honestly will tell you where to focus first, and it will usually tell you more quickly than you expect.
What a data audit actually involves
A structured HR data audit is the most efficient way I know to move from “we think our data is mostly okay” to “we know what our data looks like and we have a plan”. The word “audit” puts people off, so it is worth being clear about what it does and does not mean.
It does not mean someone arriving with a clipboard to find fault. It means looking, systematically, at four things: what you collect, how and where it is stored, how consistent it is across sources, and where the gaps are. In practice that is a mix of talking to the people who actually use the data, comparing a sample of records across systems, and writing down the definitions everyone has been carrying around in their heads.
The output should be short and usable: a clear picture of where things stand, a handful of definitions agreed and written down, and a practical set of next steps in order of priority. It does not need to be lengthy or complicated. If it produces a forty-page report that nobody reads, it has missed the point.
From there, building better habits around data quality is much easier, because you know what you are building on. Reporting starts to have a rhythm to it rather than being reassembled from scratch each time. And the next time someone asks for a headcount figure, you can give them one.
Ready to get a clear picture of where your HR data actually stands?
If your reporting is not telling you what you need to know, book an HR Data Clarity Call and we will work out where it is getting stuck.


