Read-Tier Noise Models for ctDNA Variant Calling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting circulating tumor DNA in blood samples are hindered by low levels of tumor DNA relative to other molecules, leading to difficulties in distinguishing true positives from false positives, resulting in unreliable variant calling and analysis.
Innovation Solution
The development of site-specific noise models classified into multiple read tiers, using Bayesian inference to determine noise levels and account for covariates like trinucleotide context and sequencing depth, allowing for improved identification of true positives and filtering out false positives by stratifying sequencing data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single noise model is used for all sequence reads, then the device complexity is low, but the measurement precision and reliability of variant calling deteriorate due to inability to distinguish true positives from false positives
Solution Approach 1:
The patent divides the sequence reads into multiple read tiers based on quality metrics such as mapping quality, base quality, and alignment characteristics. Each tier is assigned a specific noise model trained on data from that tier, allowing differentiated analysis of high-quality versus low-quality reads. This segmentation resolves the contradiction by improving measurement precision through tier-specific modeling while managing complexity through systematic categorization.
Solution Approach 2:
The patent implements local quality assessment by training separate noise models for different read tiers, each model tailored to the specific characteristics and error profiles of reads in that tier. High-quality reads receive more sensitive detection thresholds while low-quality reads use more conservative thresholds, optimizing variant calling accuracy for each local context rather than applying a uniform approach.
2Reliability
If read-tier specific noise models are implemented, then the reliability of distinguishing true positives from false positives improves, but the device complexity and computational requirements increase
Solution Approach 1:
The patent performs preliminary stratification of sequence reads into read tiers before variant calling, using pre-defined quality metrics and thresholds. This preliminary action organizes the data into manageable groups that can be processed by appropriate noise models, reducing the complexity of real-time decision-making during variant calling while improving reliability through targeted analysis.
Solution Approach 2:
The patent changes key parameters such as noise thresholds, sequencing depth requirements, and quality score cutoffs based on the read tier being analyzed. Each noise model uses parameter values optimized for its specific tier, allowing the system to adapt its sensitivity and specificity to the quality characteristics of different read populations, thereby improving true positive identification without requiring a completely separate system for each tier.
3Measurement precision
If noise models account for multiple covariates and parameters, then the measurement precision of noise level determination improves, but the computational complexity and data processing requirements increase
Solution Approach 1:
The patent segments the analysis into discrete read tiers, each with its own noise model and parameter set. This segmentation allows the system to focus computational resources on modeling the specific covariates relevant to each tier rather than attempting to model all possible interactions across all read types simultaneously, reducing overall computational complexity while maintaining precision within each segment.
Solution Approach 2:
The patent adjusts model parameters such as sequencing depth thresholds and quality score weights based on the specific read tier being analyzed. By changing parameters to match the characteristics of each tier (e.g., using higher depth thresholds for low-quality tiers), the system achieves precise noise level determination for each category without requiring a single overly complex model to handle all scenarios.
Data Source
AI summary
Noise models for processing nucleic acid datasets can stratify processed sequence reads into different read tiers. Each read tier can be defined based on whether a potential variant location is at an overlapping region and/or a complementary region of the sequence reads. A processing system can determine, for each read tier, a stratified sequencing depth at the variant location. The processing system can determine, for reach read tier, one or more noise parameters conditioned on the stratified sequencing depth of the read tier. The noise parameters can be associated with a noise distribution. The processing system can generate an output for each noise model based on the noise parameters conditioned on the stratified sequencing depth. The processing system can combine the output for each stratified noise model to generate a combined result, which can represent a likelihood that an event would be as or more extreme than the observed data.


