Haplotype Phasing Models Using Markov Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing phasing algorithms become intractable when dealing with hundreds of thousands of genomic samples, requiring rebuilding of models with each new sample addition, making them impractical for continuous or periodic new data.
Innovation Solution
The method involves training and updating haplotype cluster Markov models using a reference set of phased genomic samples, allowing new samples to be phased quickly and accurately without rebuilding the entire model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional phasing algorithms are used to phase genomic samples, then phasing accuracy can be maintained, but the computational complexity becomes intractable when dealing with hundreds of thousands of samples
Solution Approach 1:
The patent segments the genome into smaller windows or regions, and constructs separate hidden Markov models for each window rather than building one comprehensive model for the entire genome. This segmentation reduces the computational complexity of model construction while maintaining phasing accuracy through localized analysis.
Solution Approach 2:
The patent performs preliminary action by pre-training hidden Markov models on reference haplotypes before actual phasing. The models are trained offline on reference data, and then reused for phasing new samples without retraining, which significantly reduces the computational burden when processing large numbers of samples.
2Measurement precision
If traditional phasing algorithms are used, then existing samples can be phased accurately, but new samples require complete model rebuilding which is impractical for continuous data addition
Solution Approach 1:
The patent performs preliminary training of hidden Markov models on reference haplotypes in advance. Once trained, these models can be reused for phasing new samples without requiring complete model rebuilding, enabling efficient adaptation to continuous new data while maintaining accuracy.
Solution Approach 2:
The patent creates reusable hidden Markov models that can be copied and applied to multiple different sample batches. Instead of rebuilding models for each new sample batch, the trained models are copied and reused, making the system practical for continuous data addition scenarios.
3Measurement precision
If models are rebuilt with each new sample addition, then phasing accuracy can be maintained, but processing time increases significantly
Solution Approach 1:
The patent performs model training in advance as a preliminary action, so that when new samples arrive, the models are already trained and ready for immediate use. This eliminates the time-consuming model rebuilding step for each new sample batch while maintaining phasing accuracy.
Solution Approach 2:
The hidden Markov models, once trained, serve themselves by being reused across multiple sample batches without requiring external retraining. The models automatically adapt to new data through the phasing process itself, eliminating the need for repeated model construction and significantly reducing processing time.
Data Source
AI summary
Novel haplotype cluster Markov models are used to phase genomic samples. After the models are built, they rapidly and accurately phase new samples without requiring that the new samples be used to re-build the models. The models set transition probabilities such that the probability for an appearance of any allele within any haplotype is a non-zero number. Furthermore, the most unlikely pairs of haplotypes are discarded from each model at each level until ε of the likelihood mass at each level is discarded. The models are also constructed such that contributing windows of SNPs partially overlap so that phasing decisions near one of the extreme ends of any model is are not significantly determinative of the phase. Additionally, the models are configured such that two or more nodes can be merged during the building/updating procedure to consolidate haplotype clusters having similar distributions.


