cfDNA Variant Significance Modeling for False Positive Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep sequencing of circulating cell-free nucleotides for cancer detection faces challenges such as errors during sample preparation and sequencing, leading to difficulties in accurately identifying rare variants, high compute-time and memory usage, and low allele frequencies, which can result in false positives.
Innovation Solution
A significance model is trained to predict noise levels in sequencing read information, using two distributions to assess the likelihood of false positives, and is applied to filter out false positives based on stratification-specific parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep sequencing is performed to detect rare variants in cell-free nucleotides, then sensitivity for variant detection is improved, but the number of false positives increases due to noise in the data
Solution Approach 1:
A significance model acts as an intermediary between the raw sequencing data and the final variant calls. The model processes the read frequency information and stratification data through statistical distributions to generate significance scores, which then determine whether a variant call is retained or filtered out, thereby reducing false positives while maintaining sensitivity
Solution Approach 2:
The patent replaces conventional mechanical filtering methods with a statistical significance modeling approach. Instead of using fixed thresholds or simple count-based filtering, the system uses probability distributions and significance scores to dynamically assess the reliability of each variant call, enabling more accurate distinction between true variants and noise
2Adaptability or versatility
If conventional variant calling methods are used on deep sequencing data, then computational resources are consumed, but the methods are not suitable for cell-free nucleotide samples
Solution Approach 1:
The significance model is trained specifically for cell-free nucleotide data characteristics, with different statistical parameters and distributions optimized for this particular sample type. The model stratifies variant calls based on local characteristics such as read frequency and genomic context, applying appropriate significance thresholds to each category rather than using a uniform approach
Solution Approach 2:
The system changes the parameters of variant calling by introducing significance scores derived from statistical distributions specific to cell-free data. The model uses parameters such as mean and standard deviation calculated from blank samples to determine significance thresholds, adapting the calling methodology to the unique characteristics of cfDNA/cfRNA sequencing
3Measurement precision
If sequencing depth is increased to detect low allele frequency variants, then detection capability is improved, but errors from sample preparation and sequencing become more prominent
Solution Approach 1:
The patent converts the harmful effect of noise and errors into a beneficial filtering mechanism. By modeling the distribution of read frequencies in blank samples (which contain only noise), the system creates a significance model that can distinguish between true variants and error-induced signals. The noise characteristics are used to define the null distribution, enabling statistical differentiation between signal and noise
Data Source
AI summary
A system and a method are described for applying a noise model for predicting the occurrence and a level of noise that is present in cfDNA read information. The significance model is trained for a plurality of stratifications of called variants using training data in the stratification. Stratifications may include a partition and a mutation type. The significance model predicts the likelihood of observing a read frequency for a called variant in view of two distributions of the significance model. The first distribution predicts a likelihood of noise occurrence in the sample. The second distribution predicts a likelihood of observing a magnitude of the read frequency for the called variant. The two distributions may further depend on a baseline noise level of blank samples. With these two distributions, the significance model, for a particular stratification, more accurately predicts the likelihood of a false positive for a called variant.


