Skip to content
Wednesday, October 7, 2026
Dark BiotechnologyBIOTECH · GENETICS · DEVICES
Tech News

All of Us Release Makes the Genomics Platform the Largest Integrated Health Database

NIH's All of Us Research Program became the world's largest integrated genomic and electronic health record database with its June 30, 2026 release, which per the agency's announcement makes data from more than 747,000 participants available to scientists and includes more than 535,000 whole…

Oliver Strnad · September 21, 2026 · 5 min read
ShareXFacebookLinkedInTelegramEmail
A data scientist scrolling a genomic variant dashboard on dual monitors above a steel lab bench washed in cool teal light.
A data scientist scrolling a genomic variant dashboard on dual monitors above a steel lab bench washed in cool teal light.

NIH's All of Us Research Program became the world's largest integrated genomic and electronic health record database with its June 30, 2026 release, which per the agency's announcement makes data from more than 747,000 participants available to scientists and includes more than 535,000 whole genome sequences linked to nearly 482,000 electronic health records. The follow-on Curated Data Repository version 9 landed in the Researcher Workbench in August 2026.

What Was Actually Released?

The June release is the most expansive in the program's history, per NIH's own announcement, and its composition matters more than its headline size. The NIH release enumerates more than 1.3 billion genetic variants, 553,000 genotyping arrays, 96,000 structural variant records, and roughly 600,000 physical measurements, alongside 747,000 survey responses covering social circumstances, behaviors, and environments. Enrolled-participant count passed 883,000, growth of more than 114,000 since the previous data version, and EHR data grew 22% in this release. NIH also states that All of Us data has fueled more than 1,400 peer-reviewed publications by nearly 23,000 researchers. Figures beyond the disclosed set are not yet disclosed.

What Makes This a Platform Rather Than a Dataset?

The distinction is integration and access mechanics. The program's CDRv9 support article describes the repository as the world's largest integrated dataset combining genomic data with real-world clinical and wearable data, delivered through an updated Researcher Workbench with tiered access. Tiering separates aggregate from individual-level data, and the versioned releases mean analyses can be reproduced against a fixed data cut. The platform framing also implies a maintenance obligation, since EHR linkages, reconsent rules, and re-identification protections must be engineered into the repository rather than appended. This is infrastructure that behaves like software, with releases, versions, and deprecation cycles.

Why Does Cohort Composition Change the Science?

Because variant interpretation is population-dependent, and this cohort is deliberately not a convenience sample. Per NIH, more than 645,000 participants, 86% of the total, come from communities historically underrepresented in biomedical research, and participants span all 50 states and territories, reflecting more than 98% of U.S. three-digit ZIP codes. NHGRI's glossary definition of population genomics frames the field as the large-scale application of genomic technologies to study populations, and the statistical power of that application depends on who is inside the population. For target discovery and risk modeling, diversity is a data-quality parameter. Findings still require replication before they support clinical claims, and the program's own publications record is the honest measure of output so far.

What Does Entry Into the Multiomics Era Mean?

The release adds molecular layers beyond DNA for the first time, per NIH: proteomics data from nearly 10,000 participants, RNA sequencing from nearly 9,000, and long-read whole genome sequences from more than 14,500. The table below shows the layers as disclosed.

Data layerParticipants (June 2026 release, per NIH)
Whole genome sequencesMore than 535,000
Linked electronic health recordsNearly 482,000
ProteomicsNearly 10,000
RNA sequencingNearly 9,000
Long-read whole genome sequencingMore than 14,500

NIH states that additional multiomic data releases are planned later in 2026. The gap between the half-million-scale DNA layers and the ten-thousand-scale omics layers is the honest picture of where the platform stands.

What Should Industry Readers Take From It?

The platform is now the reference cohort for U.S. precision medicine research, and its scale makes it a practical substrate for rare variant discovery, biomarker identification, and drug target validation, uses the CDRv9 article explicitly names alongside AI and machine learning innovation. The disclosure basis matters: every figure above is an NIH or program statement, not an independent audit, and the largest-integrated-database claim is the agency's own. Access runs through the Researcher Workbench under controlled terms, which shapes who can build on the data and how quickly. For competitive analysis, the calendar item is the planned multiomic releases, which will indicate whether the platform's omics layers scale toward its DNA layer or plateau.

How Does Access Work in Practice?

The repository is not an open download, and its versioning is part of the science. Per the program's support documentation, CDRv9 ships in two tiers, a Controlled Tier designated C2025Q4R6 and a Registered Tier designated R2025Q4R6, with individual-level data confined to the controlled environment. Researchers work inside the Researcher Workbench rather than exporting raw records, and analyses run against a versioned data cut that can be cited. The tier names encode the quarter of the underlying data refresh, which is what makes cross-study comparisons reproducible. For platform watchers, the versioning discipline is the signal that the resource is being run as software infrastructure. Access terms, not just data volume, determine how much of the platform's value escapes into the wider literature.

What Are the Open Questions for the Platform?

Three questions follow directly from the disclosed record. First, whether the multiomics layers scale from their current four- and five-figure participant counts toward the genomic layer's half-million scale, which NIH says will be answered by releases planned later in 2026. Second, how the linkage quality between genomes and health records behaves as EHR sources diversify, since the release attributes its 22% EHR growth partly to participant-mediated submissions and health information exchange data. Third, how independent researchers assess data quality, since every figure in the release is the agency's own statement. None of these questions diminishes the scale achievement. They define the difference between a large database and a durable research platform.

Dark Biotechnology is an independent industry publication. This article is explanatory journalism, not medical advice, and does not recommend or evaluate any treatment, test, or device for individual patients. Readers should consult qualified clinicians and the primary regulatory documents linked above before making decisions that affect patient care.

Sources

  1. NIH's All of Us Research Program is now the largest integrated genomics and health database in the world — National Institutes of Health
  2. Our Largest Genomic Dataset: Curated Data Repository version 9 — All of Us Research Program Support (NIH)
  3. Population Genomics - Talking Glossary of Genomic and Genetic Terms — National Human Genome Research Institute (NIH)

More from our brands

Part of the VUGA Network