Re-identification of anonymous data
In plain terms
Data presented as anonymous, cross-referenced with other sources until someone works out who it's about. The name wasn't there; it gets reconstructed.
Definition
Re-identification means recovering the identity of people within a dataset whose direct identifiers have been removed, by cross-referencing it with other sources. It shows that "anonymous" describes an operation performed on a file, not a stable property of that file once it circulates.
How it works
Removing a name doesn't remove what distinguishes someone. Attributes remain — a postal code, an age range, an occupation, schedules, routes — whose combination may match very few people, sometimes only one. All it then takes is a second source where those same attributes appear alongside a name: a breach, a directory, a public profile, an open registry. The cross-reference restores the name. Location data illustrates the mechanism well: with no identifier at all, a route repeated between two points at the same times identifies a person, because few people sleep in one place and work in another on exactly that rhythm. The practical takeaway for a reader is not to treat "anonymized" as a permanent guarantee — it's a risk reduction, whose value depends on what exists elsewhere, today and later.
Warning signs
- A service that claims to collect "anonymous" data without saying which data or at what level of detail
- A study or dataset published with precise attributes: municipality, exact age, occupation
- A health, fitness, or route app you don't know what it publishes or shares
- A result that visibly concerns you within a set presented as anonymous
- Location history kept and shareable, often turned on by default
How to verify
In a privacy policy, look not for the word "anonymous" but for the list of what's collected and how detailed it is: a town rather than a region, a date of birth rather than an age range, a precise location rather than a zone. It's that level of detail that determines whether cross-referencing is possible, not the adjective used.
What to do
Limit the precision you provide when the service allows it: an approximate location rather than an exact one, an age range rather than a date, history turned off. Check what sport, health, and route apps publish by default, where sharing is often on without an explicit prompt.
If it already happened
You can ask the organization for access to data about you, its erasure, and object to its processing; the CNIL is the authority to turn to if the request goes unanswered. On a dataset already published, erasing existing copies isn't achievable — the useful step is to get it withdrawn at the source and reduce what you feed into it afterward.
Frequently asked questions
- So "anonymized" isn't a guarantee?
- It's a relative guarantee, one that depends on what exists elsewhere. A dataset genuinely impossible to re-identify today can stop being so tomorrow, if another source appears and enables cross-referencing. The term describes a process, not a permanent state.
- What can I actually do?
- Act on precision rather than on the principle: an approximate location, an age range, history turned off considerably reduce how unique your combination of attributes is — and it's that uniqueness, not the presence of your name, that makes re-identification possible.
Related attacks
Official sources
This article is part of the Data and digital identity family. Last updated: 2026-09-03.