Polymorphic Gene Typing Using Posterior Probability Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for high-resolution HLA gene typing from whole exome sequencing data face challenges due to sequence ambiguities, polymorphism complexity, and capturing biases, leading to laborious and expensive processes with limited accuracy.
Innovation Solution
A method involving the alignment of sequencing reads to a reference set of allele variants, calculation of posterior probabilities, and application of weighting factors to determine allele variants, enabling precise polymorphic gene typing and mutation detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If whole exome sequencing is used for HLA gene typing, then high-throughput and cost-effectiveness are improved, but sequencing accuracy and allele differentiation capability deteriorate due to capturing biases and short read length
Solution Approach 1:
The patent segments the HLA gene sequencing task into multiple informative exon regions (exons 2-4 for Class I, exons 2-3 for Class II) that are separately captured and sequenced. This segmentation allows targeted enrichment of polymorphic regions while avoiding non-informative regions, improving allele differentiation accuracy without requiring complete gene sequencing.
Solution Approach 2:
The patent applies local quality by designing capture probes specifically for polymorphic exon regions with high allelic discrimination value. The capture efficiency and sequencing depth are optimized locally for these specific regions rather than uniformly across the entire HLA gene, maximizing the information content from short reads.
2Measurement precision
If directed experimental protocols with PCR are used for HLA typing, then allele typing accuracy is improved, but labor intensity and cost increase
Solution Approach 1:
The patent merges multiple experimental steps (DNA extraction, PCR amplification, sequencing library preparation) into a single targeted capture sequencing workflow. The captured amplicons are directly prepared for sequencing without requiring separate PCR enrichment steps, simplifying the overall process while maintaining high typing accuracy.
Solution Approach 2:
The patent replaces manual, labor-intensive PCR-based typing methods with automated next-generation sequencing technology. The sequencing platform automatically performs parallel amplification and sequencing of multiple samples, eliminating the need for manual PCR setup, gel electrophoresis, and Sanger sequencing steps.
3Loss of information
If complete HLA gene sequencing is performed, then comprehensive allele information is obtained, but sequencing cost and time increase significantly
Solution Approach 1:
The patent extracts and sequences only the informative exon regions (exons 2-4 for HLA Class I, exons 2-3 for HLA Class II) that contain the majority of polymorphic sites useful for allele differentiation. Non-informative regions are excluded from the capture design, reducing sequencing requirements while maintaining sufficient typing resolution.
Solution Approach 2:
The patent applies partial action by sequencing only the necessary portions of the HLA genes that provide sufficient discriminatory power for allele typing. This partial sequencing approach achieves adequate typing resolution without the excessive time and cost of complete gene sequencing, balancing information completeness with resource efficiency.
Data Source
AI summary
A system and method for determining the exact pair of alleles corresponding to polymorphic genes from sequencing data and for using the polymorphic gene information in formulating an immunogenic composition. Reads from a sequencing data set mapping to the target polymorphic genes in a canonical reference genome sequence, and reads mapping within a defined threshold of the target gene sequence locations are extracted from the sequencing data set. Additionally, all reads from the set data set are matched against a probe reference set, and those reads that match with a high degree of similarity are extracted. Either one, or a union of both these sets of extracted reads are included in a final extracted set for further analysis. Ethnicity of the individual may be inferred based on the available sequencing data which may then serve as a basis for assigning prior probabilities to the allele variants. The extracted reads are aligned to a gene reference set of all known allele variants. The allele variant that maximizes a first posterior probability or posterior probability derived score is selected as the first allele variant. A second posterior probability or posterior probability derived score is calculated for reads that map to one or more other allele variants and the first allele variant using a weighting factor. The allele that maximizes the second posterior probability or posterior probability score is selected as the second allele variant.A system and method for identifying somatic changes in polymorphic loci using WES data. The exact pair of alleles corresponding to the polymorphic gene are determined as described using a normal or germline sample from an individual. A tumor or otherwise diseased sample is also retrieved from the individual and the corresponding WES data is generated. Reads corresponding to the polymorphic gene are extracted as described in the paragraph above. These reads are then aligned to the inferred pair of allele sequences. The alignment of the germline or normal reads to the inferred pair of alleles, along with the alignment of the tumor or diseased reads to the inferred pair of alleles are simultaneously used as inputs to somatic change detection algorithms to identify somatic changes with greater precision and sensitivity.


