Machine Learning for Somatic Mutation Detection in ctDNA
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Next-generation sequencing (NGS) technologies face challenges in reliably detecting sequence mutations or structural alterations in circulating tumor DNA (ctDNA) due to the presence of abundant normal DNA and the generation of very short sequence reads, which interferes with the detection of medically significant alterations.
Innovation Solution
A computer-trained machine learning algorithm, such as a random forest or neural network, is used to analyze sequence reads from NGS technologies, comparing them to a reference to detect and validate sequence mutations or structural alterations, even in samples with mixed tumor and normal DNA, allowing for the accurate reporting of tumor-specific mutations and alterations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If NGS technologies are used to sequence ctDNA, then large amounts of genetic information can be sequenced rapidly, but the detection of sequence mutations or structural alterations becomes unreliable due to abundant normal DNA and very short sequence reads
Solution Approach 1:
The patent introduces machine learning algorithms as an intermediary between NGS data generation and mutation detection. The algorithm processes short NGS reads, distinguishes ctDNA mutations from normal DNA background, and filters sequencing artifacts, thereby enabling reliable mutation detection despite the limitations of short reads and abundant normal DNA
Solution Approach 2:
The patent changes the analytical parameters by using machine learning models that can detect mutations at very low allele frequencies (down to 0.1% or lower) and by processing extremely short sequence reads (30-150 bases) that traditional methods cannot reliably analyze, thus adapting the detection parameters to match the constraints of ctDNA sequencing
2Productivity
If NGS technologies generate very short sequence reads, then sequencing throughput increases, but assembly of reads to reconstruct the subject's genome becomes challenging
Solution Approach 1:
Instead of attempting to assemble short reads into complete genome reconstructions, the patent extracts and detects specific mutation signals directly from the short reads by comparing them to a reference genome. This approach bypasses the need for complex de novo assembly while still achieving the goal of identifying medically significant alterations
Solution Approach 2:
Rather than assembling reads to reconstruct the genome and then finding mutations, the patent inverts the approach by directly comparing short reads to a reference genome and detecting mutations without requiring full genome assembly, thus simplifying the workflow for ctDNA analysis
3Ease of operation
If ctDNA is present among excess non-target DNA, then the sample represents a minimally invasive approach, but the presence of abundant normal DNA interferes with finding sequence mutations
Solution Approach 1:
The machine learning algorithm serves as an intermediary that processes the mixed cfDNA data, distinguishing true ctDNA mutations from normal DNA variants and sequencing artifacts. The algorithm uses training data including known somatic mutations to learn patterns that identify tumor-derived mutations even in the presence of abundant normal DNA
Solution Approach 2:
The patent optimizes detection parameters for low-allele-fraction mutations by using machine learning models trained to detect mutations at frequencies as low as 0.1% or lower, enabling precise mutation detection in the challenging context of cfDNA where tumor DNA may constitute a very small fraction of total DNA
Data Source
AI summary
A machine learning system and method for somatic mutation discovery are provided that provides improved identification of tumor-specific mutations. The improved identification of tumor-specific mutations may affect discovery of alterations and therapeutic management of cancer patients.


