The problem AlphaGenome was built to solve has a number attached to it, and the number is larger than most people expect.
The reference annotation of the human genome, GENCODE release 50, lists 19,442 protein-coding genes — and 35,885 long non-coding RNA genes alongside them, plus 7,608 small non-coding RNA genes. Of 78,733 annotated genes, roughly a quarter code for protein. The rest are transcribed into RNA that never becomes one.
That is the territory this system is designed to read. AlphaFold solved a problem about proteins, which are the well-mapped quarter. AlphaGenome is aimed at the three quarters we can transcribe but cannot yet interpret — the regulatory machinery where most disease-associated variants actually sit.
For most of the twentieth century, geneticists studied the less than 2% of the human genome that codes for proteins and set the rest aside. Much of the other 98% was written off as “junk DNA”, a phrase the geneticist Susumu Ohno popularised in 1972 — inert evolutionary residue, clutter accumulating in the genome like boxes in an attic.
That picture has changed, though not as completely as headlines sometimes claim. The non-coding genome, sometimes called “dark DNA” in recognition of how little we understand it, contains a vast regulatory landscape of switches and dials that decide when, where, and how strongly each gene is expressed. How much of the rest matters is still argued over, and the argument is worth knowing about.
In 2012 the ENCODE consortium announced that it could assign a “biochemical function” to 80 per cent of the genome, and newspapers declared junk DNA dead. Evolutionary biologists objected sharply: a stretch of DNA that is copied into RNA or bound by a protein is not necessarily doing anything useful, and a 2013 critique led by Dan Graur pointed out that less than a tenth of the genome shows the signs of being preserved by natural selection. The largest comparison yet, the Zoonomia project’s 2023 alignment of 240 mammal genomes, found that at least 10.7 per cent of the human genome is unusually conserved. Strikingly, 80 per cent of the most conserved single letters lie outside protein-coding genes, and half have no annotation in ENCODE at all. The honest summary: far more of the non-coding genome matters than the junk label implied, but probably not most of it.
These switches decide whether a stem cell becomes a neuron or a liver cell, whether an immune cell attacks or stands down, whether a cancer-suppressing gene stays active or falls silent. The same DNA, read differently in different cells, gives rise to every tissue in the body.
This is where most disease hides. When researchers at the US National Human Genome Research Institute catalogued the results of genome-wide association studies in 2009, 88 per cent of the variants linked to diseases and traits sat in introns or between genes, not in the protein-coding sequence itself.
Cancer, diabetes, heart disease and schizophrenia are all shaped this way. For decades we could see the clues — a variant statistically tied to a condition — without being able to read what they actually did.
In June 2025, Google DeepMind released AlphaGenome, an artificial-intelligence model built specifically to decode that landscape, and the full peer-reviewed description appeared in Nature on 28 January 2026. From the lab that produced AlphaFold, it represents a qualitative leap in our ability to read the regulatory genome.
Its reach extends across medicine, drug discovery, evolutionary biology and our most basic understanding of how life is controlled at the molecular level.
What AlphaGenome Does

AlphaGenome is a deep-learning model that takes a DNA sequence as input and predicts its regulatory output — how that stretch of sequence shapes gene activity across different cell types and tissues. Its reach is what sets it apart.
The model reads sequences up to one million DNA letters long while resolving their effects down to a single base pair. That combination matters: earlier tools had to trade breadth for precision, either seeing long stretches at low resolution or short stretches sharply, but never both at once.
It was trained on public experimental data from the ENCODE, GTEx, 4D Nucleome and FANTOM5 consortia, spanning human and mouse genomes. From these it learned to predict thousands of molecular readouts: gene expression, where transcription starts, chromatin accessibility, histone marks, transcription-factor binding, RNA splicing, and chromatin contact maps, which show how the DNA strand folds back to touch itself.
One model, many modalities
What sets AlphaGenome apart is breadth of output. Earlier tools were specialists, each trained for a single task. AlphaGenome predicts all of the assessed regulatory readouts at once, from a single input sequence.
Those readouts include where transcription starts and ends, how RNA is spliced, how much is produced, how open the DNA is, and where regulatory proteins bind. In the published benchmarks it was the only model able to predict every assessed modality jointly.
This matters because real biology is not modular. A single variant can shift splicing, expression and chromatin state at once, and a tool that sees all of these together can reveal a mechanism a single-task model would miss.
Reading each layer separately can hide the very interaction that causes the disease. Seeing them jointly is what turns a scattered set of signals into a coherent biological story.
The practical payoff is what geneticists call variant-effect prediction. Given a sequence with a single letter changed, AlphaGenome can forecast how that change alters gene regulation in specific cell types.
That is the central problem in making sense of the thousands of disease variants GWAS keeps finding but cannot yet explain. In the peer-reviewed evaluation, the model matched or beat the strongest specialised tools on 25 of 26 variant-prediction benchmarks; the June 2025 preprint had reported 24 of 26. It can score one variant’s effect on all of its readouts in about a second.
Building on AlphaFold: DeepMind’s Genomics Programme
To grasp why AlphaGenome matters, it helps to recall what AlphaFold achieved — and why the same lab turned from protein structure to gene regulation next.
AlphaFold, released in 2020 and expanded in 2022, cracked the protein-folding problem: predicting a protein’s three-dimensional shape from its amino-acid sequence with accuracy rivalling experiment. The work earned DeepMind’s Demis Hassabis and John Jumper a share of the 2024 Nobel Prize in Chemistry.
But structure is only half the story. Knowing a protein’s shape tells you what it can do; it says nothing about when it is made, in which cells, or in what quantity.
That is the domain of gene regulation, and it is exactly what AlphaGenome sets out to read. The AlphaFold database now holds predicted structures for nearly every human protein, but structure alone leaves the question of control unanswered.
The two systems are complementary. AlphaFold tells you what a protein looks like. AlphaGenome tells you when and where it is switched on. AlphaGenome is more directly a successor to Enformer, DeepMind’s earlier regulatory model, and a companion to AlphaMissense, which handles variants inside protein-coding genes.
It is also strikingly efficient. According to DeepMind, a single AlphaGenome model was trained in about four hours using roughly half the computing power its predecessor Enformer required — a reminder that progress in this field comes from better architecture, not only bigger machines.
Non-Coding DNA and Disease
The medical stakes are hard to overstate. Genome-wide association studies compare the genomes of people with and without a disease to find variants statistically tied to it. They have flagged many thousands, but with a catch: the overwhelming majority fall outside protein-coding genes.
Catalogues of association results put the figure close to 90 per cent: the overwhelming majority of disease-associated variants sit in non-coding, regulatory DNA. That has created a deep interpretive gap between knowing a variant is linked to disease and knowing what it actually does.
Without the mechanism — which gene, which cell type, which pathway — a statistical association is a clue without an explanation, and you cannot build a targeted therapy on a clue alone. This is the bottleneck AlphaGenome is designed to break.
Consider how a single regulatory change can act. A one-letter difference in an enhancer — a distant control element — might weaken the grip of a transcription factor, quietly lowering a gene’s output in one cell type while leaving every other tissue untouched.
Effects this specific are why the non-coding genome resisted decoding for so long. The signal is real but subtle, buried in context, and invisible unless a model can hold both the long-range sequence and the fine-grained detail in view at the same time.
Laboratory methods can now test thousands of variants at once, but each experiment covers a chosen cell type and a chosen set of sequences, and a 2024 review of these techniques describes how slowly the evidence accumulates. The gap between what is flagged and what is understood is exactly the space these AI tools are built to close.
By predicting a variant’s regulatory effect in a specific cell type, it converts raw associations into testable biological hypotheses. For a cancer variant it can suggest which cell types are affected; for a psychiatric one, which brain cells show altered regulation.
In one striking demonstration, the model recapitulated the mechanism of clinically relevant variants near the TAL1 oncogene, a known driver in certain leukaemias. This kind of reasoning connects directly to the wider story in our guide to the genetics of cancer.
TAL1 makes a good test because the answer was already known from the laboratory. In 2014 a team at the Dana-Farber Cancer Institute found that in some cases of T-cell acute lymphoblastic leukaemia, small acquired mutations in a stretch of DNA upstream of TAL1 create a new landing site for a protein called MYB. MYB then gathers the machinery for a powerful control switch, a super-enhancer, that drives the cancer gene. Given only the DNA sequence, AlphaGenome recovered that chain of effects across several of its readouts at once.
The Size of the Problem It Is Trying to Solve
To understand why a system like AlphaGenome is being built at all, it helps to see how lopsided our knowledge of the genome actually is.
GENCODE release 50 annotates 78,733 genes in the human genome. Of these, 19,442 code for proteins. 35,885 are long non-coding RNA genes and 7,608 are small non-coding RNA genes, with 14,702 pseudogenes, broken copies of once-working genes, making up most of the remainder. The protein-coding portion — the part biology understood first and understands best — is under a quarter of the total.
That imbalance is the whole reason this problem is hard. AlphaFold’s achievement was to predict the folded shape of a protein from its amino acid sequence — a difficult problem with a well-defined answer, in a domain with decades of structural data to learn from. The non-coding genome offers neither. There is no single output to predict, no equivalent of a folded structure, and far less labelled ground truth about what any given regulatory sequence does.
This matters clinically because of where disease-associated variants tend to fall. Genome-wide association studies repeatedly find that the majority of variants linked to common diseases lie outside protein-coding regions — in promoters, enhancers and other regulatory elements whose function we can rarely read directly from sequence. A variant is identified as statistically associated with a condition, and then the trail goes cold, because nobody can say what the sequence it sits in was doing.
One honest qualification belongs with the annotation figures themselves. Counting 35,885 long non-coding RNA genes means transcripts have been detected from those loci; it does not establish that each has a biological function. Some fraction almost certainly represents transcriptional noise. Distinguishing meaningful regulation from background is itself unfinished work — which means a model trained on this annotation is learning from a map that is still being drawn.
Set against that, the ambition is reasonable in scale. Reading the regulatory genome is the largest unfinished task in human genetics, and it is a pattern-recognition problem over enormous sequence datasets — which is precisely the shape of problem where these methods have earned their reputation.
AlphaGenome and Drug Discovery
One of the most immediate uses is in finding and validating drug targets. Most drugs work by tuning the activity of a protein, and the hardest part of the process is knowing which protein to aim at, and whether nudging it will help without causing harm.
The economics are unforgiving. A drug candidate can absorb a decade and vast sums before failing in trials. Many fail because the biological target was never truly causal.
Better evidence at the very start of that pipeline is worth more than speed anywhere later, which is why causal insight into disease regulation carries such weight for the companies that develop medicines.
AlphaGenome helps by pointing to which regulatory variants actually drive a disease, and therefore which genes are causally involved rather than merely correlated.
That distinction is worth a great deal. A gene whose regulation is causally disturbed by a disease variant is a far more credible drug target than one flagged only by weaker, indirect evidence.
The approach is most valuable for complex, multifactorial illnesses — cancer, cardiovascular disease, neurodegeneration, metabolic disorders — where no single obvious target exists and the disease emerges from regulatory shifts across many genes and cell types at once.
Pharmaceutical teams have begun folding this style of regulatory prediction into their target-identification pipelines. Because so many failed drug candidates fail for lack of a genuine causal link to disease, any tool that strengthens that link early is commercially as well as scientifically valuable.
Personalised Medicine and the Regulatory Genome
The temptation is to leap from here to a fully personalised readout of your own genome. That leap needs a firm caveat, and DeepMind has been explicit about it.
With that boundary in mind, the longer-term direction is still real. As whole-genome sequencing becomes routine, tools of this kind could help interpret which disease mechanisms are active in a patient’s regulatory variants, extending the precision-medicine logic already reshaping oncology to a far wider range of common diseases.
Two people with the same diagnosis may have reached it by different regulatory routes — different variants, different genes, different cell types — and may respond differently to the same drug. Reading those routes is what could eventually make treatment genuinely individual.
For a broader look at how gene-editing tools like CRISPR change medicine by rewriting DNA directly, see our report on gene editing in 2026. And for how the environment shapes gene expression above the sequence itself, see our explainer on epigenetics.
Together, AlphaGenome, gene editing and epigenetics form three converging routes to reading and steering gene regulation — one computational, one molecular, one environmental.
That AI now parses genomes at all is part of a wider shift explored in our piece on how large language models work, whose sequence-modelling ideas underpin tools like this one.
A Growing Toolkit for Reading Life
AlphaGenome does not stand alone. It joins a widening family of DeepMind biology models, each reading a different layer of the molecular story, from DNA sequence up to protein shape and interaction.
AlphaFold and its successor AlphaFold 3 handle protein and molecular structure. AlphaMissense scores the effects of variants inside protein-coding genes. AlphaProteo designs new binding proteins. AlphaGenome fills the gap none of them addressed: the vast regulatory country between the genes.
Nor is DeepMind alone. Borzoi, from Calico Life Sciences, published in Nature Genetics in 2025, predicts how much RNA each stretch of DNA produces in specific cells and tissues, and scores variants across transcription, splicing and the tail-end processing of RNA. Evo 2, from the Arc Institute, published in Nature in 2026, took a different route: trained on 9 trillion letters of DNA from every domain of life with a one-million-letter window, it predicts the effects of variants such as those in the breast cancer gene BRCA1 without being trained for that task, and its makers released everything, including the training data.
Independent geneticists have described it as a kind of Swiss-army knife for non-coding DNA, able to systematically predict the molecular consequences of every possible variant across a disease-linked region. Used that way, it can help prioritise which variants deserve costly laboratory follow-up.
Its predictions can also complement established scoring tools such as CADD, sharpening the assessment of whether a given variant is likely to be harmful. The value is not that AlphaGenome replaces experiments, but that it tells scientists which experiments are worth running first.
There is an evolutionary dimension too. Much of what makes species distinct lies in regulation, not in the protein-coding genes themselves, which are often highly conserved. The Zoonomia comparison found 4,552 ultraconserved elements, stretches almost identical across mammals, and linked changes in genes and regulatory elements to unusual traits such as hibernation. A tool that reads regulatory sequence opens a new window onto how forms diverge over deep time.
Limitations and What Comes Next

AlphaGenome is powerful, but its limits deserve clear statement. Its predictions are probabilistic and can be wrong, particularly for novel variant combinations or for cell types thinly represented in its training data.
It does predict chromatin contact maps, a readout of how chromosomes fold in the nucleus, but only within its one-million-letter window, and DeepMind says the influence of control elements more than about 100,000 letters away remains hard to capture. Nor does it model how regulation shifts dynamically through development, ageing or the course of a disease.
DeepMind first offered AlphaGenome through an online interface for non-commercial research in June 2025, then released the model’s code and trained weights alongside the Nature paper in January 2026, again under non-commercial terms. Running it locally takes a high-end data-centre graphics processor. Feedback from users will steer the next generation; DeepMind has named sharper cell- and tissue-specific predictions as a priority.
Openness is part of the strategy. By putting the model, its weights and its scoring tools into researchers’ hands, DeepMind turns the wider community into a distributed test bed, surfacing failures and edge cases far faster than any single lab could.
The first large product of that openness arrived in September 2026, posted as a preprint that has not yet been peer reviewed. DeepMind researchers ran AlphaGenome on every possible single-letter change in the human genome and condensed the predictions into one AlphaGenome Variant Impact score. They report that the resulting atlas helped solve a rare-disease case of epileptic encephalopathy, a severe childhood epilepsy, and increased the statistical power to detect rare non-coding variants that shape traits across whole populations. If those results hold up under review, the atlas could become a standard first stop for anyone trying to make sense of a variant outside the genes.
It is worth pausing on how recent all of this is. For most of the genomic era, the non-coding majority of our DNA was a source of embarrassment more than insight. The label “junk” was a confession of ignorance dressed up as a conclusion.
What has changed is not the DNA but our ability to interrogate it. Consortia spent two decades painstakingly measuring how the genome behaves across hundreds of cell types, and models like AlphaGenome now distil those measurements into predictions that arrive in hours rather than years.
If AlphaFold made the protein world legible, AlphaGenome is an early attempt to do the same for the regulatory world — a far larger and messier territory. It will be wrong often, revised repeatedly, and eventually superseded by something better.
That is how the illumination of the dark genome is likely to proceed. Not in a single flash, but steadily, one readable stretch at a time, until the part of our DNA we once called junk becomes one of the best-understood layers of biology.
The broader trajectory is clear. The dark DNA is being illuminated. The regulatory genome, for decades the least legible part of biology, is becoming readable — and the consequences for medicine, evolution and our understanding of life are difficult to overstate.
What Scientists Say
The reception among geneticists has been enthusiastic but measured. In Trends in Genetics, Judit García-González and Krzysztof Gogolewski of the Icahn School of Medicine at Mount Sinai described AlphaGenome as a “Swiss army knife” for non-coding DNA, notable for keeping single-letter resolution while holding long-range context. They also set out its limits in prioritising and interpreting the variants that underlie human traits and diseases.
Writing in Nature Structural & Molecular Biology, Dennis Gankin and Pedro Beltrao called it the largest multimodal DNA sequence model for non-coding regions so far, advancing the state of the art in almost all prediction tasks while leaving clear room to improve. The field’s verdict: a significant step, not a finished solution.
That balance — genuine excitement paired with insistence on experimental validation — is the healthiest sign for a tool of this kind. The specialists closest to the work are the ones most careful to say that a prediction is a hypothesis, not a verdict.
DeepMind’s own framing has been similarly disciplined. The team has repeatedly emphasised what the model is not for — personal genome prediction, clinical diagnosis, complex-trait forecasting — even while showcasing what it can do, an unusually candid posture for a headline product launch.
Frequently Asked Questions
Further Reading on Web News For Us
Sources
Primary peer-reviewed research:
- Avsec, Ž., et al. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature, 649, 1206–1218. doi.org/10.1038/s41586-025-10014-0
- Avsec, Ž., et al. (2025). AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model. bioRxiv preprint. doi.org/10.1101/2025.06.25.661532
- Avsec, Ž., et al. (2021). Effective gene expression prediction from sequence by integrating long-range interactions (Enformer). Nature Methods, 18, 1196–1203. doi.org/10.1038/s41592-021-01252-x
- Christmas, M.J., et al. (2023). Evolutionary constraint and innovation across hundreds of placental mammals. Science, 380, eabn3943. doi.org/10.1126/science.abn3943
- ENCODE Project Consortium (2012). An integrated encyclopedia of DNA elements in the human genome. Nature, 489, 57–74. doi.org/10.1038/nature11247
- Graur, D., et al. (2013). On the immortality of television sets: “function” in the human genome according to the evolution-free gospel of ENCODE. Genome Biology and Evolution, 5, 578–590. doi.org/10.1093/gbe/evt028
- Hindorff, L.A., et al. (2009). Potential etiologic and functional implications of genome-wide association loci for human diseases and traits. PNAS, 106, 9362–9367. doi.org/10.1073/pnas.0903103106
- Mansour, M.R., et al. (2014). An oncogenic super-enhancer formed through somatic mutation of a noncoding intergenic element. Science, 346, 1373–1377. doi.org/10.1126/science.1259037
- Linder, J., et al. (2025). Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation (Borzoi). Nature Genetics, 57, 949–961. doi.org/10.1038/s41588-024-02053-6
- Brixi, G., et al. (2026). Genome modelling and design across all domains of life with Evo 2. Nature, 652, 1349–1361. doi.org/10.1038/s41586-026-10176-5
- Cheng, J., et al. (2026). AlphaGenome Atlas: in silico mutagenesis of the entire human genome improves prioritization and interpretation of non-coding variants. medRxiv preprint. doi.org/10.64898/2026.09.16.26363192
Institutional / science journalism:
- Google DeepMind (2025). AlphaGenome: AI for better understanding the genome. deepmind.google
- García-González, J. & Gogolewski, K. (2026). AlphaGenome, a Swiss-army knife for exploring non-coding DNA. Trends in Genetics, 42(1), 4–6. doi.org/10.1016/j.tig.2025.11.007
- Peña-Martínez, E.G. & Rodríguez-Martínez, J.A. (2024). Decoding non-coding variants: recent approaches to studying their role in gene regulation and human diseases. Frontiers in Bioscience (Scholar Edition), 16(1), 4. ncbi.nlm.nih.gov
- Gankin, D. & Beltrao, P. (2026). The AlphaGenome deep learning model predicts effects of non-coding variants. Nature Structural & Molecular Biology, 33, 373–374. doi.org/10.1038/s41594-026-01763-1
- Google DeepMind (2026). AlphaGenome research code and model weights. github.com/google-deepmind/alphagenome_research
- GENCODE release 50 — human gene annotation statistics, EMBL-EBI (retrieved 8 October 2026)
Baryon. (2025, November 9). Decoding the Dark DNA: How DeepMind’s AlphaGenome is Revolutionizing Genetic Research. Web News For Us. https://webnewsforus.com/decoding-the-dark-dna-alphagenome/
Baryon. “Decoding the Dark DNA: How DeepMind’s AlphaGenome is Revolutionizing Genetic Research.” Web News For Us, 9 November 2025, https://webnewsforus.com/decoding-the-dark-dna-alphagenome/. Accessed 11 October 2026.
