Machine Learning Models for Noisy Nucleotide Read Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing MRD detection systems are inflexible and inaccurate due to their inability to accommodate noisy nucleotide reads, leading to false positives and inadequate distinction between erroneous and actual ctDNA supporting reads.

Innovation Solution

The system employs a combination of whole genome sequencing (WGS) and machine learning models to create a unique fingerprint for cancer presence, allowing for flexible accommodation of noisy reads and accurate MRD detection by comparing post-treatment samples to a panel of normals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If existing systems directly analyze nucleotide reads to detect ctDNA, then the detection process is simple, but sequencing imperfections cause noisy reads to be misidentified as ctDNA leading to false positives

Engineering Contradiction:
Improvedetection process simplicityVSAvoidMRD detection accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces machine learning models as intermediaries between the raw nucleotide reads and the MRD detection decision. These models process the noisy reads and distinguish true ctDNA signals from sequencing errors, thereby maintaining operational simplicity while significantly improving detection reliability and reducing false positives

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the direct mechanical analysis approach (simple read counting) with a computational intelligence system (machine learning models). This substitution enables the system to handle noisy data and complex patterns that simple mechanical methods cannot distinguish, improving reliability without substantially increasing operational complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If existing systems fail to distinguish noisy reads from actual ctDNA supporting reads, then the system operates inflexibly, but accurate distinction requires sophisticated methods that existing systems lack

Engineering Contradiction:
Improvesystem flexibilityVSAvoidread distinction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic machine learning models that can adapt to different sequencing data characteristics and noise patterns. The models are trained to flexibly distinguish between noisy reads and true ctDNA signals, providing both adaptability to various conditions and high precision in read distinction through learned patterns rather than rigid rules

Inventive Principle:
Principle #15Dynamics

3Productivity

If existing systems provide inaccurate detection results due to noisy reads, then the detection speed is fast, but the results require validation that slows down the overall process

Engineering Contradiction:
Improvedetection speedVSAvoiddetection result accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by training machine learning models in advance on labeled data before actual MRD detection. This pre-training enables the models to quickly and accurately process new sequencing data without requiring time-consuming validation steps, thereby maintaining fast detection speed while ensuring high result accuracy through pre-learned discrimination capabilities

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250201346A1Using machine learning models for detecting minimum residual disease (MRD) in a subject
Publication Date: 2025.06.19 ILLUMINA INC
  • US20250201346A1 patent drawing
  • US20250201346A1 patent drawing
  • US20250201346A1 patent drawing

AI summary

This disclosure describes methods, non-transitory computer-readable media, and systems that detect minimal residual disease (MRD) within a sample of interest. For example, in some cases, the disclosed systems identify, for an initial genomic sample of a subject infected with cancer, a tumor fingerprint comprising variants at a target genomic region. The disclosed systems further determine, for a sample of interest of the subject, a set of sample of interest nucleotide reads associated with the target genomic region. The disclosed system process the set of sample of interest nucleotide reads using a first machine learning model and process panel of normals nucleotide reads using one or more additional machine learning models. The disclosed systems compare scores determined from the outputs of the machine learning models to predict whether the sample of interest has minimal residual disease related to the cancer.