IBD Segment Estimation Using HMM and Digital Templates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating identical-by-descent (IBD) segments between individuals are hindered by genotyping errors and phase switch errors, particularly in large genetic data sets and among closely related individuals.
Innovation Solution
The implementation of computer-implemented methods and systems that process haplotype data using digital templates to identify IBD segments, reducing the impact of genotyping errors, and employing a hidden Markov model (HMM) to correct phase switch errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If IBD estimation is performed using phased haplotypes to improve relationship inference and trait inheritance tracing, then measurement precision is improved, but device complexity increases due to the need for phase-aware processing and error correction mechanisms
Solution Approach 1:
The patent segments the haplotype data into focal segments (from the query individual) and reference segments (from population references). By dividing the complex task of phase-aware IBD estimation into separate processing steps - identifying focal segments, finding matching reference segments, and merging results - the system achieves high measurement precision while managing computational complexity through modular processing.
Solution Approach 2:
The patent introduces reference segments as intermediary elements that mediate between the query individual's phased haplotypes and the final IBD estimation. These reference segments serve as a bridge to handle phase switch errors and genotyping errors, allowing the system to achieve accurate IBD estimation without requiring completely new complex algorithms for error correction.
2Reliability
If multiple digital templates with different masked and unmasked site arrangements are applied to reduce genotyping errors, then reliability is improved, but device complexity increases due to multiple template processing requirements
Solution Approach 1:
The patent divides the haplotype sites into masked and unmasked segments across multiple digital templates. Each template processes different portions of the data (different masked/unmasked arrangements), and the results are merged to produce reliable IBD segment identification. This segmentation allows the system to handle genotyping errors by cross-validating across multiple templates without requiring a single overly complex processing system.
Solution Approach 2:
The patent merges the results from multiple digital templates with different masked and unmasked site arrangements to produce the final set of IBD segments. By combining the strengths of each template (each handling different error patterns), the system achieves high reliability in IBD segment identification while distributing the computational complexity across multiple simpler processing steps rather than one complex step.
3Measurement precision
If hidden Markov models are employed to correct phase switch errors in haplotype data, then measurement precision is improved, but device complexity increases due to additional computational models required
Solution Approach 1:
The patent uses hidden Markov models as intermediary computational tools that operate on the phased haplotype data to correct phase switch errors. The HMM serves as a probabilistic mediator between the raw phased data and the final IBD estimation, allowing the system to achieve high measurement precision through statistical modeling while keeping the complexity manageable through standard HMM algorithms rather than custom complex models.
Solution Approach 2:
The patent replaces direct mechanical comparison of haplotypes with a probabilistic HMM-based approach to handle phase switch errors. Instead of using complex deterministic algorithms to correct errors, the system substitutes a statistical model that infers the most likely correct haplotype configuration, achieving high precision error correction through probabilistic reasoning rather than complex computational mechanics.
Data Source
AI summary
Example embodiments relate to identity-by-descent (IBD) relatedness based on focal and reference segments. An example method includes determining, by a services platform based on personal information of a focal individual, a focal string. The method also includes retrieving, by the services platform from a reference database, a reference string of a reference individual. Additionally, the method includes computationally identifying, by the services platform, IBD segments between the focal string and the reference string. Further, the method includes determining, by the services platform and based on the merged set of IBD segments, a degree of relatedness between the focal individual and the reference individual. In addition, the method includes providing, by the services platform, access to the degree of relatedness via a user interface.


