HLA Haplotype Determination via Segmented PCR and Computational Phasing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA-based technologies for HLA typing face challenges such as limited read lengths in sequencing, the need for cloning haplotypes separately, and high error rates, which hinder accurate determination of HLA haplotypes.
Innovation Solution
The method involves selectively amplifying nucleic acid molecules, particularly exons and adjacent introns of HLA genes, and using paired-end sequencing to obtain non-overlapping sequencing reads. These reads are then partitioned using algorithms like k-means clustering to determine haplotypes accurately.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Length of moving object
If traditional sequencing technologies are used for HLA typing, then sequencing can be performed, but read lengths are insufficient to sequence both exons together and cloning is required separately
Solution Approach 1:
The invention divides the HLA gene region into two segments: exon 2 and exon 3. By designing separate PCR amplification reactions for each exon, the method enables sequencing of individual exons with standard read lengths, avoiding the need for long-read sequencing or cloning while still allowing haplotype determination through computational phasing of the segmented data.
Solution Approach 2:
The invention transitions from a single-dimension approach (attempting to sequence the entire exon 2-exon 3 region in one read) to a multi-dimensional approach by performing separate amplifications and sequencing for each exon, then combining the information computationally. This dimensional shift allows standard sequencing technologies to achieve what would otherwise require advanced sequencing capabilities.
2Measurement precision
If sequencing based typing is performed to provide higher information content, then typing ambiguities are avoided, but high error rates are associated with sequencing technologies
Solution Approach 1:
The invention implements a feedback mechanism through iterative computational analysis. Sequencing reads are initially assigned to haplotypes based on sequence similarity, then consensus sequences are generated from assigned reads, and this process repeats iteratively. Each iteration refines the haplotype assignments by comparing new reads against updated consensus sequences, allowing error correction through cumulative evidence and statistical phasing algorithms that leverage linkage disequilibrium patterns.
Solution Approach 2:
The invention creates multiple copies of sequencing reads through PCR amplification and generates redundant sequencing data by sequencing both exons separately. This redundancy allows computational algorithms to identify and correct sequencing errors by comparing multiple reads against each other and against reference haplotype patterns, thereby improving reliability despite inherent sequencing error rates.
3Loss of information
If complete sequencing of both exons together is attempted, then direct haplotype information is obtained, but read lengths required are not available with current technologies
Solution Approach 1:
The invention performs preliminary action by separately amplifying and sequencing exon 2 and exon 3 before computational combination. By obtaining sequence data for each exon independently with standard read lengths, the method preserves phase information through statistical phasing algorithms that analyze linkage disequilibrium patterns, avoiding the need for physical long-read sequencing while still recovering complete haplotype information.
Data Source
Figure 1
Figure 2A~2B
Figure 2C
AI summary
Presented herein are methods and compositions for determining haplotypes in a sample. The methods are useful for obtaining sequence information regarding, for example, HLA type and haplotype. Also presented herein are methods of determining haplotypes in a sample based on a plurality sequence reads.