Haplotype Analysis Using Molecular Markers and Hidden Markov Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining haplotypes, such as PGT-M and PGT-SR-Balanced assays, face challenges with high false-positive and false-negative rates due to allele dropout and uneven whole-genome amplification, and are unable to accurately detect normal CNVs using low-depth whole-genome sequencing or gene chips.
Innovation Solution
A method and device that analyze genome data sets from a descendant object, their parents, and a genetically related reference object using molecular marker genotyping, Hidden Markov models, and Viterbi dynamic programming to determine haplotype genetic flow and composition, incorporating multiple family members' genetic information for accurate haplotype analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If direct pathogenic site detection is performed in PGT-M assay, then detection speed is improved, but false-positive rate and false-negative rate increase due to allele dropout and uneven whole-genome amplification
Solution Approach 1:
The patent introduces molecular markers as intermediary elements that flank the target site. Instead of directly detecting the pathogenic site which suffers from amplification biases, the method detects molecular markers upstream and downstream of the target site. These markers serve as proxies that can be reliably detected and used to infer the haplotype status of the intervening target region, thereby avoiding the allele dropout problem while maintaining detection speed.
Solution Approach 2:
The patent replaces direct mechanical detection of the pathogenic site with a computational inference system. By detecting molecular markers and using Hidden Markov models to calculate likelihood ratios, the system substitutes direct observation with probabilistic reasoning. This allows the system to overcome the limitations of direct detection while maintaining efficiency through automated computational analysis.
2Reliability
If linkage and crossover theory is used to infer disease state by comparing haplotypes, then reliability of disease state inference is improved, but analysis complexity increases
Solution Approach 1:
The patent transforms the complex haplotype comparison problem into a simplified likelihood ratio calculation by changing the parameter representation. Instead of comparing entire haplotype sequences, the method calculates a single likelihood ratio L based on molecular marker genotypes at flanking positions. This parameter transformation maintains the reliability of linkage analysis while dramatically reducing computational complexity and making the analysis more tractable.
Solution Approach 2:
The patent extracts only the essential information needed for haplotype inference by focusing on molecular markers at the flanking positions upstream and downstream of the target site. Rather than analyzing the entire genome or all possible haplotype combinations, the method isolates and analyzes only the critical marker positions that provide sufficient information for reliable disease state inference, thereby simplifying the overall analysis.
3Quantity of substance
If low-depth whole-genome sequencing or gene chips are used to detect CNVs, then detection cost is reduced, but measurement precision of normal CNVs deteriorates
Solution Approach 1:
The patent introduces molecular markers as intermediary detection targets that provide more precise measurements than direct CNV detection. By genotyping specific molecular markers at flanking positions rather than attempting to directly measure CNV regions with low-depth sequencing, the method achieves superior measurement precision for detecting copy number abnormalities while maintaining the cost-effectiveness of targeted genotyping approaches.
Data Source
AI summary
The invention provides an analysis method and a device for determining a haplotype of a descendant object. Particularly, the invention provides a data analysis method for determining a haplotype genetic flow, comprising the following steps: (a) providing data sets for the analysis, the data sets being data sets related to genome information; (b) performing molecular marker genotyping in the upstream and downstream regions of Y1 target sites in each of the data sets, thereby obtaining molecular marker genotyping data, wherein Y1 is a positive integer greater than or equal to 1; (c) constructing a binary genetic vector of (0, 1) for each molecular marker site upstream and downstream of each target site in each of the data sets; (d) determining a maximum likelihood estimation value L using a Hidden Markov model for each target site; (e) determining a haplotype genetic flow direction of the descendant object and the family members through a Viterbi dynamic programming algorithm.


