Tumor Mutation Phasing With Statistical Haplotype Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computational methods for phasing mutations in tumors are inadequate due to assumptions that do not hold in tumor sequencing data, particularly for polyploid genomes and RNA sequencing, and are hindered by short-read limitations and complex tumor cell subpopulations, making it difficult to accurately predict neoantigens for personalized immunotherapies.
Innovation Solution
A statistical model is used to analyze tumor DNA and RNA sequence reads, estimating haplotype-existence probabilities, prevalences, and transcript prevalences by enumerating unique mutation patterns and determining their probabilities, allowing for accurate phasing of somatic and germline variants in tumor samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If existing computational methods for phasing mutations are used, then the process is simpler, but the accuracy of neoantigen prediction deteriorates due to assumptions that do not hold in tumor sequencing data
Solution Approach 1:
The patent changes the fundamental parameters of the phasing approach by developing a statistical model specifically tailored for tumor sequencing data characteristics, including polyploid genomes and RNA sequencing. This model accounts for tumor-specific features such as subclonal populations and varying allele frequencies, thereby improving prediction accuracy without requiring overly complex procedures
Solution Approach 2:
The patent replaces traditional read-backed phasing approaches with a statistical modeling framework that uses probability distributions to estimate haplotype configurations. This substitution allows the system to handle the complexities of tumor data (polyploidy, subclonality) that traditional methods cannot accommodate, improving accuracy while maintaining computational feasibility
2Measurement precision
If statistical modeling is used to account for polyploid genomes and RNA sequencing, then the accuracy of haplotype phasing is improved, but the computational complexity increases
Solution Approach 1:
The patent segments the complex phasing problem into manageable components by separately modeling DNA and RNA sequencing data, then integrating the results. The statistical model divides the genome into regions with different ploidy characteristics and applies appropriate phasing strategies to each segment, reducing overall computational complexity while maintaining accuracy
Solution Approach 2:
The patent implements partial phasing by focusing computational resources on regions of the genome that contain somatic mutations and are relevant to neoantigen prediction, rather than attempting to phase the entire genome. This selective approach maintains high accuracy for clinically relevant regions while reducing unnecessary computational burden
3Loss of time
If short-read sequencing data is used, then the sequencing cost and time are reduced, but the ability to accurately phase mutations deteriorates due to read length limitations
Solution Approach 1:
The patent introduces statistical modeling as an intermediary layer that connects short-read sequencing data to haplotype phase information. The model uses probability distributions and population genetics principles to infer phase relationships from indirect evidence in short reads, effectively bridging the gap between limited read lengths and accurate phasing requirements
Solution Approach 2:
The patent replaces direct physical linkage information (which would require long reads) with statistical inference mechanisms. By using probabilistic models that incorporate linkage disequilibrium, recombination rates, and mutation patterns, the system achieves accurate phasing without relying on long physical reads, thus maintaining the advantages of short-read sequencing
Data Source
AI summary
This application relates generally to analyzing mutations in tumors, and more particularly, to systems and methods for phasing mutations in tumors of subjects (e.g., cancer patients). An exemplary method for phasing mutations in a tumor of a subject comprises enumerating, based on tumor DNA and/or RNA sequence reads, a set of unique mutation patterns observed in the plurality of sequence reads; counting the set of unique patterns observed in the sequence reads to calculate a quantity of each of the unique mutation patterns and/or a quantity of each combination of unique mutation pattern and a transcript group; determining mutation pattern probabilities; and inputting the mutation pattern quantities and the mutation pattern probabilities into a statistical model to estimate at least one of a set of haplotype-existence probabilities that each of the haplotypes exists, a set of haplotype prevalences, and a set of haplotype-transcript prevalences.


