Iceland’s Pangenome Reveals Hidden Disease Variants
A remarkable scientific expedition unfolded across the human genome much as an archaeological survey unfolds across a buried landscape. For years, many studies had relied on a single standard human genome as though it were one master map of a vast ancient world. But one map can miss whole settlements, smooth away local tracks, and hide real human variation.
The solution was to build an Icelandic pangenome: a grand atlas made from many individual genetic lineages. This effort brought together 788 haplotypes, including 698 from Icelandic people and others from wider global datasets. Iceland's exceptional genealogical records gave the project special depth, with parent-child trios helping researchers separate maternal and paternal inheritances like tightly layered archaeological deposits.
More than 57,000 Icelanders had their genome data mapped to the new pangenome, yielding nearly 99 million variants — over 6% more than older single-reference methods. The pangenome included over 51 million small phased variants: single-letter changes, insertions, and deletions, ranging from widely shared to singletons seen only once. It illuminated previously dark genomic regions, bringing medically important genes into clearer view. The human genome, it turns out, resembles an ancient road system with many branches rather than one straight Roman road.
Behind the pangenome lay a powerful new assembly method called Emblask. Every person carries chromosomes from both parents, and untangling those two sets resembles sorting goods from two families thrown together in the same long-abandoned storeroom. Emblask used parent-offspring trios to determine which DNA stretches belonged to the paternal line and which to the maternal.
The method combined long reads — broad but error-prone — with accurate but fragmented short reads. Long reads were first corrected using short reads, dramatically reducing errors while preserving the information needed to keep maternal and paternal sequences distinct. The sequences were then progressively assigned to each parent based on similarity with parental DNA.
Across 329 Icelandic trios, the resulting assemblies achieved roughly one error per 100,000 bases and nearly 99% completeness. Raw long reads with mean error rates around 7% were transformed into polished assemblies of remarkable quality — like taking shattered inscriptions and half-buried blocks and recreating a temple plan with confidence. Each resolved haplotype represented a genuine inherited lineage rather than a blurred average, ensuring that rare and medically relevant Icelandic variation was directly represented in the genomic framework.
Some genomic regions are repetitive or so similar to neighbouring stretches that short DNA reads struggle to find their correct home. These dark regions are where ordinary tools lose confidence and important details disappear. A new method called Weaver was designed to work precisely in this difficult terrain, mapping short reads not to a single reference but to the richer pangenome graph.
Instead of forcing every read onto one path, Weaver allows different routes through the pangenome, using small matching sequence anchors linked into plausible paths. Crucially, it preserves rare variants rather than discarding them for tidiness — accepting complexity because reality is complex. In performance tests, Weaver mapped more reads correctly than established methods, especially in low-mappability regions, while also running faster despite working with far more haplotypes.
The dark territory shrank dramatically. Millions of previously ambiguous bases became accessible, including many protein-coding and medically relevant genes. Better placement of reads also improved variant calling accuracy, particularly in regions where duplicated sequences cause one location to masquerade as another — allowing hidden signals to finally emerge from camouflage.
The most dramatic discoveries came when the new methods exposed disease-linked variants that older approaches had missed entirely. The gene GBA1 sits beside a near-identical pseudogene, GBAP1, causing reads to be systematically misassigned. Within this difficult terrain, the pangenome uncovered a missense variant, p.Leu483Pro, already documented as pathogenic for Parkinson's disease yet effectively invisible to standard short-read workflows.
In Iceland this variant appeared at roughly 0.22% frequency, with a particularly strong association with early-onset Parkinson's disease before age 60. The finding was then replicated across more than 429,000 British and Irish participants in the UK Biobank after targeted remapping — a hidden signal in one population becoming a reproducible result in another. The variant had been present all along; the methods had simply been blind to it, like a corroded inscription concealing a name that was always there.
A second discovery involved CBS, associated with homocystinuria — a serious disorder affecting the eyes, skeleton, blood vessels, and nervous system. The variant p.Gly307Ser, known to be pathogenic and linked to Celtic ancestry, lay in a low-mappability region and had not been called in older datasets. The pangenome recovered it at roughly 0.31% frequency in Iceland, and the only homozygous carrier identified had already received a diagnosis consistent with homocystinuria. A nearby linked marker also showed higher frequency in Ireland and Scotland than elsewhere in the British Isles — a quiet reminder that genes, like languages, can carry traces of population history across centuries.
These findings illustrate that the problem was never simply insufficient sample size. Enormous datasets already existed. The problem was that medically important variants sat in genomic regions where standard methods could not see clearly. The pangenome did not invent these variants — it restored them to visibility.
Applying the new framework across more than 57,000 Icelanders produced a callset of nearly 99 million variants — a full population-scale operation, not a polished demonstration. Compared with an earlier Icelandic dataset using the old linear reference, the pangenome-based approach yielded over 5.7 million additional reliable small variants despite having fewer sequenced samples. Method and reference choice can matter as much as simply adding more participants; a better map reveals more than a larger expedition with poorer charts.
These extra variants were concentrated in previously difficult regions, including within coding sequences of medically connected genes. The benefits extended further through imputation, yielding nearly 12% more well-imputing variants genome-wide and about 10% more in coding regions — spreading rare and useful genetic information more effectively across broader population samples for association studies and medical interpretation.
A quiet but important philosophical shift underlies all of this. The single human reference genome had stood as a monumental centrepiece: useful but inevitably limited. The pangenome approach replaces it with something more collective and historically faithful — a structure built from many inherited lines that acknowledges variation as the rule, not the exception. A genomic reference that ignores local variation is like writing history from imperial capitals alone, preserving the broad outline while losing the real texture.
The genome, often presented as a fixed and settled text, turns out to resemble an archaeological landscape continually reinterpreted as methods improve. New features once dismissed as noise become central. Iceland provided the setting for a major rethinking of how that landscape should be mapped — and the result is not just more data, but a fuller, livelier, and more truthful account of human inheritance.

Comments