Tumor Purity Estimation From B-Allele Frequency Without Matched Controls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for estimating tumor purity in biological samples are subjective and inaccurate, often requiring histopathologic evaluation or a matched normal control sample, which may not be available, leading to reduced accuracy in detecting somatic mutations and copy number changes.
Innovation Solution
A method using a trained machine-learning model to process B-allele frequency distribution from nucleic acid sequence data, allowing for accurate estimation of tumor purity without a matched normal control, employing fully connected neural networks, one-dimensional convolutional neural networks, and two-dimensional convolutional neural networks to generate an estimated metric.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional histopathologic evaluation methods are used to estimate tumor purity, then the estimation can be performed with existing techniques, but the results are subjective and inaccurate
Solution Approach 1:
The patent replaces manual histopathologic evaluation (mechanical/visual inspection) with computational analysis of B-allele frequency distributions using machine learning models. This substitution eliminates subjectivity while maintaining accessibility, as the computational method processes sequencing data through automated algorithms rather than human expert judgment.
Solution Approach 2:
The patent introduces B-allele frequency distribution as an intermediary metric that bridges raw sequencing data and tumor purity estimation. This intermediary allows the system to derive purity information from allele frequency patterns without requiring direct visual inspection or matched normal samples, thereby improving accuracy while keeping the method accessible.
2Measurement precision
If matched normal control samples are required for tumor purity estimation, then the accuracy of somatic mutation detection can be improved, but the method becomes inapplicable when normal controls are unavailable
Solution Approach 1:
The patent extracts the essential information needed for tumor purity estimation (B-allele frequency distribution) from tumor-only sequencing data, separating this requirement from the need for matched normal control samples. By focusing on allele frequency patterns inherent in the tumor sample itself, the method achieves accurate purity estimation without requiring external control samples.
Solution Approach 2:
The patent creates a universal method that functions with tumor-only samples, making the technique applicable to all clinical scenarios regardless of normal sample availability. The B-allele frequency-based approach serves multiple purposes: estimating tumor purity, detecting somatic mutations, and assessing copy number changes, all without requiring matched normal controls.
3Measurement precision
If machine learning models are used to process B-allele frequency distribution, then tumor purity estimation accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent performs preliminary processing of sequencing data to generate B-allele frequency distributions before applying machine learning models. This preprocessing step organizes the data into a standardized format that simplifies the subsequent machine learning analysis, reducing the computational complexity of the model while maintaining high estimation accuracy.
Data Source
AI summary
The disclosure provides methods for estimating tumor purity from tumor samples without use of matched-normal controls. A set of genomic regions are identified based on a nucleic acid sequence data that is aligned to a reference genome. Each genomic region of the set of genomic regions includes one or more nucleotide-sequence variants relative to a corresponding genomic region of the reference genome. A B-allele frequency distribution for the biological sample is determined based on a B-allele frequency determined for each genomic region of the set of genomic regions. The B-allele frequency distribution is processed using a trained machine-learning model to estimate a metric identifying tumor purity in the biological sample.


