Haplotype-Block Imputation for Low-Depth Genomic Markers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomic prediction methods face challenges in balancing cost, precision, and throughput, particularly with low-depth genomic sequencing data, leading to inaccurate genomic marker imputation and high costs for high-throughput applications such as genome-wide association studies and breeding projects.
Innovation Solution
A method using haplotype-block-guided marker imputation to supplement low-quality genomic data, increasing read coverage and accuracy by aligning haplotypes with haplotype-blocks and imputing missing or ambiguous markers based on shared haplotype information, thereby improving genomic feature prediction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If low-depth genomic sequencing is used, then costs are reduced, but measurement precision of genomic markers deteriorates
Solution Approach 1:
The patent introduces haplotype-blocks as an intermediary structure that mediates between low-depth sequencing data and accurate genomic marker detection. By organizing reference data into haplotype-blocks and using them to guide imputation, the system achieves high-precision marker detection without requiring high-depth sequencing, thus resolving the contradiction between cost reduction and measurement precision.
Solution Approach 2:
The patent creates virtual copies of genomic information through imputation. Instead of directly observing all markers through expensive high-depth sequencing, the system copies marker information from reference haplotype-blocks to target individuals, achieving accurate marker detection at low cost. This copying mechanism allows low-depth sequencing data to be supplemented with imputed marker data.
2Quantity of substance
If low-depth genomic sequencing data is used, then costs are reduced, but reliability of genomic predictions deteriorates
Solution Approach 1:
The patent performs preliminary organization of reference data into haplotype-blocks before the actual genomic prediction task. By pre-structuring reference data into meaningful haplotype blocks with known marker patterns, the system enables reliable imputation during prediction, ensuring that even low-depth sequencing data can be accurately supplemented, thus maintaining prediction reliability while reducing costs.
Solution Approach 2:
Haplotype-blocks serve as an intermediary that bridges low-depth sequencing data and reliable genomic predictions. The structured haplotype-blocks provide a framework for accurate imputation, allowing the system to achieve reliable predictions without requiring expensive high-depth sequencing, thereby resolving the contradiction between cost and reliability.
3Device complexity
If traditional imputation methods are used with low-depth data, then computational simplicity is maintained, but measurement precision of imputed markers deteriorates
Solution Approach 1:
The patent segments the genome into haplotype-blocks, which are discrete units with specific marker patterns. This segmentation allows the imputation process to work with structured, manageable blocks rather than attempting to impute individual markers across the entire genome, thereby improving imputed marker accuracy while maintaining computational feasibility through localized, block-based processing.
Solution Approach 2:
The patent changes the fundamental parameter of imputation from individual marker-based approaches to haplotype-block-guided approaches. By shifting the unit of imputation from single markers to structured haplotype blocks with known patterns, the system achieves superior imputation accuracy while the block-based framework actually simplifies the computational process through pattern matching and reduced search space.
Data Source
AI summary
The invention relates to a computer-implemented method for predicting a genome-related feature (458) from genomic data of multiple individuals (402), the method comprising: —receiving (102) genomic marker data (434, 442) of each of the individuals, the genomic marker data being indicative of a plurality of first marker positions assigned to identified marker variants (1140-1142) and multiple second (1144) marker positions have a missing or ambiguous marker variant assignment; —computing (104) a haplotype-block library (448) comprising a plurality of haplotype-blocks (1126-1136), each haplotype-block comprising start and stop coordinates and a series of marker positions referred to as ‘comparison marker positions’ lying within the start and stop coordinates; —performing (106) a haplotype-block-guided marker imputation; —supplementing (108) the genomic marker data with the imputed marker variants; and—using the supplemented genomic marker data (454) for computationally predicting the feature (458) of the individuals.


