A unified dataset of UK census variables for small areas: Harmonised data tables from the 2021 England, Wales, and Northern Ireland censuses and the 2022 Scotland census
Owen Goodwin; Alex Singleton (2026). Environment and Planning B: Urban Analytics and City Science. DOI: 10.1177/23998083261429563
Abstract
This paper describes a new dataset release containing harmonised census tables from the 2021 and 2022 UK Censuses. The release is the first unified dataset covering all four UK nations at the smallest available geographic level: Output Areas in England, Wales, and Scotland, and Data Zones in Northern Ireland. The UK’s three census agencies: ONS (England and Wales), NRS (Scotland), and NISRA (Northern Ireland) release their data separately, each with distinct variables, formats, and disclosure controls. Through a process of matching, standardisation, and aggregation, 190 comparable variables are produced. The dataset is made available as a series of topic tables indexed across all 239,023 of the UK’s small-area geographies. By providing a standardised dataset, this work enables seamless UK-wide analyses, facilitating cross-national comparisons and supporting research and public policy development.
Extended Summary
How can researchers and policymakers analyse UK-wide census data when England, Wales, Scotland, and Northern Ireland each release their statistics separately, using different variables, formats, and disclosure rules? This paper addresses that problem by creating the first unified small-area census dataset covering all four UK nations from the 2021/22 census round.
The research draws on open-access census data published by the three UK statistical agencies: the Office for National Statistics (ONS) for England and Wales, National Records of Scotland (NRS), and the Northern Ireland Statistics and Research Agency (NISRA). Data were collected via bulk downloads and, for Northern Ireland, automated web scraping of NISRA’s Flexible Table Builder, since no application programming interface (API) was available. The team then undertook a detailed harmonisation process: standardising variable naming conventions, matching comparable census tables across the three releases, and aggregating differently categorised variables (such as ethnicity groupings) into consistent UK-wide measures. This was a largely manual task, as census questions, response categories, and disclosure control methods varied considerably between nations. For example, measures of household rooms and bedrooms could not be fully harmonised due to differing data collection approaches, while ethnicity categories differed because of local demographic priorities and disclosure limits.
The study also grappled with statistical disclosure control differences—methods used by agencies to protect individual privacy through techniques like cell key perturbation (small random adjustments to counts) and targeted record swapping. Because these methods were applied separately by each country before harmonisation, combining data introduced further complexity, and table totals had to be recalculated consistently across all nations to avoid discrepancies.
The result is a dataset of 190 harmonised variables organised into 25 thematic tables, covering topics such as ethnicity, health, employment, housing tenure, and language proficiency, indexed across 239,023 small-area geographies (Output Areas and Data Zones). Of 52 comparable topic tables initially identified for England and Wales, only 28 could be matched across all three census releases, and just four of these had directly equivalent categories without requiring aggregation. Validation checks, including distribution comparisons across countries, confirmed that harmonised variables were broadly compatible, though some distortion from aggregating perturbed data is unavoidable.
This work carries significant implications for UK-wide demographic research and policy development, offering the first accessible, reproducible framework for cross-national 2021/22 census analysis since no official equivalent exists. It also acknowledges that pandemic-related disruptions and the one-year delay in Scotland’s census affect comparability of some measures, such as employment and commuting patterns. Beyond the UK, the fully open code pipeline and documentation could serve as a template for harmonising census or population data in other multi-jurisdictional contexts, such as the European Union, Australia, or Canada, where similar cross-border data linkage challenges exist.
Key Findings
- First unified UK-wide census dataset combining 2021 England/Wales/Northern Ireland and 2022 Scotland data at smallest geographic level
- 190 harmonised variables created across 25 topic tables, covering 239,023 small-area Output Areas and Data Zones
- Only 28 of 52 England and Wales topic tables could be matched across all three census releases, with just 4 directly comparable
- Significant variable aggregation was required for 21 of 25 matched tables due to differing categorisations, such as ethnicity and language proficiency measures
- Differences in disclosure control methods (cell key perturbation and record swapping) between agencies introduce distortion when combining data across nations
Citation
@article{goodwin2026unified,
author = {Owen Goodwin; Alex Singleton},
title = {A unified dataset of UK census variables for small areas: Harmonised data tables from the 2021 England, Wales, and Northern Ireland censuses and the 2022 Scotland census},
journal = {Environment and Planning B: Urban Analytics and City Science},
year = {2026},
doi = {10.1177/23998083261429563}
}