Read-Tier Noise Models for ctDNA Variant Calling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting circulating tumor DNA in blood samples are hindered by low levels of tumor DNA relative to other molecules, leading to difficulties in distinguishing true positives from false positives, resulting in unreliable variant calling and analysis.

Innovation Solution

The development of site-specific noise models classified into multiple read tiers, using Bayesian inference to determine noise levels and account for covariates like trinucleotide context and sequencing depth, allowing for improved identification of true positives and filtering out false positives by stratifying sequencing data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single noise model is used for all sequence reads, then the device complexity is low, but the measurement precision and reliability of variant calling deteriorate due to inability to distinguish true positives from false positives

Engineering Contradiction:
Improvevariant calling accuracyVSAvoidnoise model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the sequence reads into multiple read tiers based on quality metrics such as mapping quality, base quality, and alignment characteristics. Each tier is assigned a specific noise model trained on data from that tier, allowing differentiated analysis of high-quality versus low-quality reads. This segmentation resolves the contradiction by improving measurement precision through tier-specific modeling while managing complexity through systematic categorization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality assessment by training separate noise models for different read tiers, each model tailored to the specific characteristics and error profiles of reads in that tier. High-quality reads receive more sensitive detection thresholds while low-quality reads use more conservative thresholds, optimizing variant calling accuracy for each local context rather than applying a uniform approach.

Inventive Principle:
Principle #3Local quality

2Reliability

If read-tier specific noise models are implemented, then the reliability of distinguishing true positives from false positives improves, but the device complexity and computational requirements increase

Engineering Contradiction:
Improvetrue positive identificationVSAvoidmulti-model system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary stratification of sequence reads into read tiers before variant calling, using pre-defined quality metrics and thresholds. This preliminary action organizes the data into manageable groups that can be processed by appropriate noise models, reducing the complexity of real-time decision-making during variant calling while improving reliability through targeted analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes key parameters such as noise thresholds, sequencing depth requirements, and quality score cutoffs based on the read tier being analyzed. Each noise model uses parameter values optimized for its specific tier, allowing the system to adapt its sensitivity and specificity to the quality characteristics of different read populations, thereby improving true positive identification without requiring a completely separate system for each tier.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If noise models account for multiple covariates and parameters, then the measurement precision of noise level determination improves, but the computational complexity and data processing requirements increase

Engineering Contradiction:
Improvenoise level determinationVSAvoidmodel parameter complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the analysis into discrete read tiers, each with its own noise model and parameter set. This segmentation allows the system to focus computational resources on modeling the specific covariates relevant to each tier rather than attempting to model all possible interactions across all read types simultaneously, reducing overall computational complexity while maintaining precision within each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adjusts model parameters such as sequencing depth thresholds and quality score weights based on the specific read tier being analyzed. By changing parameters to match the characteristics of each tier (e.g., using higher depth thresholds for low-quality tiers), the system achieves precise noise level determination for each category without requiring a single overly complex model to handle all scenarios.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220336044A1Read-Tier Specific Noise Models for Analyzing DNA Data
Publication Date: 2022.10.20 GRAIL INC
  • US20220336044A1 patent drawing
  • US20220336044A1 patent drawing
  • US20220336044A1 patent drawing

AI summary

Noise models for processing nucleic acid datasets can stratify processed sequence reads into different read tiers. Each read tier can be defined based on whether a potential variant location is at an overlapping region and/or a complementary region of the sequence reads. A processing system can determine, for each read tier, a stratified sequencing depth at the variant location. The processing system can determine, for reach read tier, one or more noise parameters conditioned on the stratified sequencing depth of the read tier. The noise parameters can be associated with a noise distribution. The processing system can generate an output for each noise model based on the noise parameters conditioned on the stratified sequencing depth. The processing system can combine the output for each stratified noise model to generate a combined result, which can represent a likelihood that an event would be as or more extreme than the observed data.