
When people imagine a large international research project, they usually picture researchers interviewing participants, collecting health data, or publishing exciting findings.
Almost nobody thinks about data harmonisation. Yet without it, none of those findings could be trusted.
The R-NEET project follows 1,600 young people across South Africa and Nigeria to understand resilience and mental health among young adults who are not currently in employment, education or training. Over several years we are collecting hundreds of different pieces of information about each participant, from demographic information and housing conditions to physical health and detailed measures of psychological wellbeing.
That sounds straightforward, but it really isn’t
Two countries. Multiple teams. Hundreds of variables.
Our data are collected by different organisations in two countries, using different survey software, different database structures and, inevitably, slightly different ways of recording information.
Even something as apparently simple as age or gender can be stored differently between databases. More complex psychological questionnaires introduce even greater challenges, particularly when they have been translated into different languages or adapted for different contexts.
By the end of the project we will have:
- background demographic information
- housing assessments
- physical health measurements
- around 200 psychosocial questions collected repeatedly over four waves.
Multiply that across 1,600 participants and suddenly we’re dealing with millions of individual data points.
Before we can answer a single research question, all of those pieces need to fit together.
What is harmonisation?
People often imagine harmonisation as copying data from one spreadsheet into another, whereas in practice it is much closer to translating between languages. Every variable has to be checked. Every coding decision has to be documented.
If one site records “Yes” and “No” while another uses 1 and 0, that’s relatively easy.
More challenging are situations where questions have changed over time, instruments work differently across languages, or data have been entered using different conventions. Even small inconsistencies can have major consequences when analysing depression, wellbeing or other psychological measures.
What we are trying to do is to ensure that when researchers compare participants across countries, they can be confident they are genuinely comparing like with like.
Good data harmonisation is largely invisible. Readers of a journal article rarely see the hundreds of hours that have gone into checking variables, documenting decisions, testing reliability, or independently verifying the work.
However, this hidden effort is what makes the science credible. It allows other researchers to understand exactly how the data were prepared, reproduce analyses, and have confidence that published findings genuinely reflect participants’ experiences rather than differences in data processing.
In large international projects, this level of transparency is essential.
Building a resource for the future
The R-NEET project is designed to produce far more than a single set of publications. Our ambition is to create a research resource that will support:
- collaborative papers led by the international research team;
- doctoral research using carefully curated subsets of the data;
- future data sharing in line with FAIR principles and the Nagoya Protocol.
To make that possible, we are building detailed documentation alongside the datasets themselves. Every variable is mapped, every transformation recorded, and every scoring decision explained. Original raw data are preserved alongside harmonised versions so that every analytical decision remains transparent and reversible.
Looking ahead
Perhaps the biggest lesson from this work is that harmonisation should never be an afterthought. The easiest data to harmonise are data that have been designed to work together from the very beginning. As R-NEET moves into its next phase, one of our priorities will be establishing shared data standards before data collection even begins. Investing time upfront will make future analyses faster, more transparent and ultimately more reliable. Because in international, interdisciplinary research, good science doesn’t begin when the statistical analysis starts. It begins long before that, with the careful, often invisible work of making sure every piece of data can be trusted.















