Rubinacci Lab · FIMM · University of Helsinki
Rubinacci - Lab
We develop efficient statistical and computational methods to decode human genetic variation at scale, applying them to biobanks to understand the genetic basis of disease.
EHSG 2026 · Gothenburg
We study deletions, duplications and other large rearrangements, and how they affect genes and disease risk. Population-scale sequencing and functional data help us separate causal changes from background variation.
Long shared haplotypes carry information about ancestry and recent relatedness. We use identity-by-descent methods to trace that structure and improve rare-variant discovery and disease mapping.
We build methods that remain practical as cohorts grow from thousands to millions of genomes, with attention to memory, runtime and the awkward edge cases in real datasets.
We ask how much can be recovered from a low-coverage genome. Our methods use a reference panel to infer missing genotypes, making 0.1× sequencing useful for large cohorts and rare variants.
Phasing tells us which variants sit on the same chromosome. We develop methods that do this accurately for rare and singleton variants, even when no relatives are available.
Genetic associations are only a starting point. We link variants to expression, protein levels and other molecular measurements to follow the path from DNA change to disease.
ESHG 2025 · Milan, Italy
Modern genomic datasets contain an enormous amount of information about human health, but much of it remains difficult to access and interpret. We develop computational methods to turn increasingly complex genomic data into useful biological insight.
Our software is open, efficient and designed for real-world studies, enabling researchers to analyse genetic variation at population scale and discover new links between genomes and disease.
We introduced a lcWGS imputation method that sublinearly to millions of haplotypes. Applied to 150,119 UK Biobank genomes.
Documentation →Statistical phasing method that allows <5% switch error for ultra-rare variants.
Documentation →Method allowing SNP array imputation scale to millions reference individuals.
Documentation →Talks and conversations from a week in Gothenburg.
Certain repeat sequences in our DNA actually expand and contract as we age, driving serious diseases. A recent study uses massive biobanks and advanced computational methods to map where and why these genomic changes occur.
Population-based phasing methods have long struggled with singleton variants, where limited sharing of haplotype information makes accurate phase inference very difficult. SHAPEIT5 addresses this by leveraging the local haplotype context rather than relying solely on allele sharing.
The UK Biobank's WGS release created an unusually large reference panel — large enough to strain conventional imputation pipelines. Here's how GLIMPSE2 was redesigned to handle it efficiently.