Population-Specific Reference Genomes for Genome Interpretation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for interpreting whole genome sequencing data face challenges such as reliance on an outdated human reference genome, difficulty in assigning phase to genetic variants, and accurately predicting genetic risks due to computational limitations and biased reference sequences.
Innovation Solution
Development of ethnicity-specific human reference genome sequences and integration of these into interpretation pipelines, along with methods for resolving long-range haplotype phase and prioritizing variants relevant to Mendelian diseases, using Hidden Markov Models and population linkage disequilibrium data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the current NCBI human reference genome is used for sequence alignment and variant calling, then the process is computationally feasible and straightforward, but the interpretation is biased due to the reference sequence containing both common and rare disease risk variants and representing only a small sampling of human genetic variation
Solution Approach 1:
The patent segments the single reference genome into multiple population-specific reference genomes (e.g., EUR, AFR, EAS). Each reference genome is constructed by selecting the major allele at each position from sequencing data of individuals from that specific population, thereby dividing the universal reference into targeted, population-adapted references that reduce bias for each group.
Solution Approach 2:
The patent applies local quality by creating reference genomes tailored to specific population groups rather than using a single universal reference. Each population-specific reference genome has locally optimized quality for that population, with major alleles selected from that population's sequencing data, ensuring better alignment and variant calling accuracy for individuals from that specific population.
2Adaptability or versatility
If de novo assembly of genome sequences from raw sequence reads is performed, then an alternative to reference-based alignment is achieved, but computational limitations and the large amount of mapping information encoded in relatively invariant genomic regions make this unattractive
Solution Approach 1:
The patent introduces population-specific reference genomes as an intermediary between raw sequence reads and variant interpretation. This intermediary provides a population-adapted framework for alignment and variant calling, avoiding the computational burden of de novo assembly while eliminating the bias of the ancestral reference genome. The reference genomes serve as a mediator that enables efficient, accurate, and population-appropriate genome analysis.
3Productivity
If high-throughput whole genome sequencing is performed to generate massive amounts of genetic data, then population-wide genome sequencing becomes feasible, but technologies for interpretation of the data must advance in step to handle the volume and complexity
Solution Approach 1:
The patent performs preliminary action by pre-processing sequencing data from multiple individuals from each population to construct population-specific reference genomes before actual variant interpretation. This preliminary step captures population-specific major alleles and creates optimized references that simplify subsequent interpretation of individual genomes, reducing the complexity of the interpretation pipeline while enabling high-throughput analysis.
Data Source
AI summary
In an embodiment of the present invention, three novel human reference genome sequences were developed based on the most common population-specific DNA sequence (“major allele”). Methods were developed for their integration into interpretation pipelines for highthroughput whole genome sequencing.


