Sequencing Read-Pair Filtering for Cross-Contamination Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting cross-sample contamination in sequencing data from cancer detection samples, which require high-depth sequencing studies, are inadequate for identifying tumor-derived mutations and suffer from false positive calls due to contaminating DNA.

Innovation Solution

A system and method that filters sequence read pairs using dual-strand rulesets and applies a contamination model based on a negative binomial distribution to determine contamination probability, identifying and removing sequence read pairs indicative of contamination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing contamination detection methods are applied to high-depth sequencing data from cancer detection samples, then the detection process can be performed, but the methods fail to accurately identify tumor-derived mutations and produce false positive calls due to contaminating DNA

Engineering Contradiction:
Improvecontamination detection accuracyVSAvoidcancer detection specificity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the contamination detection process into multiple independent components: (1) identifying candidate SNPs from sequencing data, (2) filtering SNPs based on population frequency thresholds, (3) calculating contamination probability using a negative binomial distribution model, and (4) comparing against a significance threshold. This segmentation allows each component to be optimized independently, improving overall detection accuracy while maintaining cancer detection specificity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary statistical model (negative binomial distribution) that acts as a mediator between the raw sequencing data and the final contamination determination. This intermediary model accounts for overdispersion in sequencing data and provides a more accurate probability calculation, reducing false positives while maintaining sensitivity to true contamination events.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If high-depth sequencing is performed to detect low tumor burden, then cancer detection sensitivity is improved, but the risk of false positive calls from contaminating DNA increases

Engineering Contradiction:
Improvetumor burden detection sensitivityVSAvoidfalse positive calls from contamination
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent performs preliminary filtering of SNPs based on population frequency before conducting the main contamination analysis. By pre-filtering out common SNPs that are likely to be polymorphic rather than contaminating, the method reduces the pool of candidate false positives before the main detection process, thereby maintaining sensitivity to true low-level tumor signals while reducing false positives from contamination.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces simple threshold-based contamination detection with a probabilistic statistical model (negative binomial distribution). This substitution allows the system to distinguish between true contamination signals and random variations in high-depth sequencing data, maintaining sensitivity to low tumor burden while reducing false positives from contaminating DNA.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If contamination probability is calculated for each SNP using population minor allele frequency, then contamination detection can be performed, but the computational complexity increases with the number of SNPs analyzed

Engineering Contradiction:
Improvecontamination probability estimation accuracyVSAvoidcomputational processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the SNP population into frequency-based groups and applies different filtering criteria to each segment. By dividing SNPs into common and rare categories and applying appropriate thresholds to each, the method reduces the total number of SNPs requiring full probabilistic analysis, thereby reducing computational complexity while maintaining detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies a two-stage filtering approach where not all SNPs undergo the full computational contamination probability calculation. Instead, SNPs are first filtered by population frequency, and only those passing the initial filter undergo the more computationally intensive negative binomial analysis. This partial application of the full analysis to only necessary SNPs reduces overall computational complexity while maintaining accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4193362B1Detecting cross-contamination in sequencing data
Publication Date: 2025.10.29 GRAIL INC
  • EP4193362B1 patent drawingFigure 1
  • EP4193362B1 patent drawingFigure 2
  • EP4193362B1 patent drawingFigure 3

AI summary

Detecting cross-contamination between test samples used for determining cancer in a subject is beneficial. To detect cross-contamination, test sequences including at least one single nucleotide polymorphism are prepared using genome sequencing techniques. Some of the test sequences can be filtered to improve accuracy and precision. A prior contamination probability for each test sequence is determined based on a minor allele frequency. A contamination model including a likelihood test is applied to a test sequence. The likelihood test obtains a current contamination probability representing the likelihood that the test sample is contaminated. The contamination model can also determine a likelihood that the sample includes loss of heterozygosity representing the likelihood that the test sequence is contaminated. Test samples that are contaminated are removed. A source for the contaminated test sample can be found by comparing contaminated test sequences to other test sequences.