[Version anglaise de l’article « Quelle place pour les données humaines dans la science ouverte ? Un équilibre fragile entre protection de la vie privée, souveraineté et partage des données »]
In late April 2026, the UK Biobank, the largest research database of UK health data, suffered a major breach: the health data of 500,000 volunteers was found for sale on the website Alibaba. The UK Biobank’s directors immediately had the listing taken down, in collaboration with Alibaba and the Chinese government. Access to the UK Biobank was then temporarily suspended while new restrictions were implemented for the 20,000 researchers (my current lab included) who use the UK Biobank in their daily work.
The leak made headlines, but it isn’t the first. Luc Rocher, an associate professor and principal investigator at the Synthetic Society Lab at the University of Oxford, has observed an increase in UK Biobank data leaks on GitHub since August 2025. Up to now, these leaks appeared to be errors stemming from inexperience rather than malicious acts (a master’s student at Yale University famously uploaded data for 96,000 participants on their GitHub by accident). To prevent such situations recurring, two British economists, Elizabeth Green and Felix Ritchie, advocate for obligatory training for scientists in sensitive data handling, while others, such as the Chinese researcher Shenglong Li, suggest real-time monitoring of data use with severe penalties for misconduct. Some people even question the inclusion of sensitive health data -particularly genetic data- in open science in general.
At the same moment as the UK Biobank suspension, a News Feature in Nature Medicine reported on the increasing promotion of “data sovereignty” at the expense of open science. In this context, data sovereignty refers to nations’ control over their citizens’ data as a resource, with very limited access for foreign researchers. While this idea appears to contradict international initiatives promoting open science, Jingyuan Fu, professor at the University of Groningen in the Netherlands, suggests that its growing popularity is due to heavily lopsided data sharing in the current political climate. She gives the example from her own lab: over 30 years, they have built a database of microbiome profiles of 167,000 people (transcriptomics, proteomics, cell types, etc.). This work, requiring to millions of euros from public funding, is now accessible and used by countries which collect but do not share similar data of their own. In the era of artificial intelligence, this unequal sharing disadvantages countries which support open science policies, as they generate data which is integrated into closed databases elsewhere. This practice then permits researchers using the closed databases to create more powerful and better-performing models than their open-science counterparts, thanks to training on larger and more diverse datasets.
Indeed, unequal data sharing seems to have a domino effect leading to increasingly closed science policies: in 2019, China designated human genetic data as a “strategic resource”, which resulted in tight restrictions on its sharing. In April 2025, the National Institutes of Health (NIH) in the United States blocked access to 21 databases for researchers from China and other “countries of concern”. In December 2025, the European Union released new criteria which explicitly exclude Chinese collaborators from certain Horizon Europe grants, notably in the field of artificial intelligence. In the Nature Medicine article and a Nature World View, both European and Chinese researchers expressed their concern that policies preventing data sharing risk diminishing the statistical power and thus the efficacy of health research. In the long term, these policies will also be more expensive as they will lead to the multiplication of national-level studies to reach the same conclusions. Equally, such policies risk exacerbating the bias in databases and AI models due to the lack of diversity in the data.
On the other hand, citizens can and should expect their governments and research institutions to protect their privacy and health data. This protection is even more important if the potential for data sharing with foreign entities, or the risks of data breach, are not explicitly discussed with database participants when seeking their informed consent. Indeed, a recent article criticizes the new Canadian genetic database, the Canadian Precision Health Initiative, for asking participants to sign data sharing agreements without defining envisioned research projects. Akin to the UK Biobank, this database should become a major resource for health research in Canada and the world. Yet how can we reconcile open science with essential participant consent, when it is impossible to imagine all possible future research studies?
In post-colonial contexts, such as in Canada, the questions of data sovereignty are even more complex. In order to best serve the entire population, these databases require the maximum genetic diversity possible. However, in sharing their DNA, indigenous populations fear to be newly exploited and discriminated, without the opportunity to play a role in directing the research itself. As a result, there is a growing movement to enable indigenous groups to manage, study, and control their genetic data at the level of their communities, and not at the level of the ‘nation’.
Personally, I am biased by my research subject (human genetics of neurodevelopmental conditions): I remain convinced that the study and the sharing of genetic data is essential to better understand our individual risk and resilience in the face of all kinds of diseases, be they infectious or genetic. This knowledge will then allow us to best accompany every person through healthcare which is increasingly personalized. That said, it is also our responsibility to ensure that data anonymization and protection measures are respected and enforced. I sincerely hope that we will continue to improve the management and protection of health research databases, to enable the diverse data sharing essential for the research breakthroughs to come.
Caitlin Martin, postdoctoral researcher at the Institut Pasteur
References:
- Webster P. Who owns my health data? Nat Med. 2026 Jun 1;32(6):1942–5. doi:10.1038/s41591-026-04378-7
- Li S. UK Biobank breach prompts the field of genomics to rethink open science. Nature. 2026 May 12;653(8114):642–642. doi:10.1038/d41586-026-01520-w
- Xu S. Open data is key to genomics research — if the information can be kept safe. Nature. 2026 May 12;653(8114):332–332. doi:10.1038/d41586-026-01475-y
- Su R, Quinn P. Comparing and contrasting the definitions of genetic data in Chinese and EU law. Int Data Priv Law. 2026 Jun 1;16(2):ipaf036. doi:10.1093/idpl/ipaf036
- Kolopenuk J, Smith RWA. Indigenous sovereignty and the limits of the Canadian Precision Health Initiative. Nat Commun. 2026 Mar 25;17(1):2956. doi:10.1038/s41467-026-71192-7


