cfDNA Screening Using UMI Error Profiling for MSI Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting somatic mutations and microsatellite instability in cell-free DNA (cfDNA) are limited by low sensitivity and high noise levels, making it difficult to distinguish true mutations from artifacts and complicating early cancer detection and analysis.
Innovation Solution
A computer-implemented method that normalizes and analyzes nucleic acid and white blood cell-derived sequence reads to identify microsatellite loci with distinct distance metrics, using unique molecular identifiers (UMIs) to enhance detection of microsatellite instability and somatic mutations, and employs machine-learning classifiers to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional cfDNA sequencing methods are used, then detection can be performed, but sensitivity is low and noise levels are high making it difficult to distinguish true mutations from artifacts
Solution Approach 1:
The method segments the cfDNA analysis into multiple independent steps: UMI assignment to individual molecules, molecular cloning reconstruction, error profile generation from normal samples, and mutation calling. This segmentation allows each step to be optimized independently and reduces the propagation of errors through the workflow.
Solution Approach 2:
The method performs preliminary actions by generating error profiles from normal cfDNA samples before analyzing tumor samples. UMIs are assigned to molecules during library preparation, and molecular clones are reconstructed in advance, allowing the system to establish baseline error rates and improve sensitivity before actual mutation detection.
2Measurement precision
If ctDNA blood levels are extremely low as in early-stage solid tumors, then cancer detection becomes more challenging, but early detection is critical for treatment outcomes
Solution Approach 1:
The method uses UMIs as molecular copies to track individual cfDNA molecules through the sequencing process. By assigning unique identifiers to each molecule and tracking its descendants through PCR amplification and sequencing, the system can distinguish true low-abundance mutant molecules from sequencing errors even when ctDNA concentration is extremely low.
Solution Approach 2:
The method replaces traditional mutation detection mechanics with a molecular tracking system based on UMIs. Instead of relying on signal intensity or frequency alone, the system uses molecular identity tracking through UMI sequences to detect and validate mutations, enabling detection at much lower concentrations.
3Measurement precision
If mutation fractions in cfDNA approach noise levels of next-generation sequencing workflows, then true somatic mutations cannot be distinguished from artifacts
Solution Approach 1:
The method introduces UMIs as an intermediary between the original cfDNA molecules and the sequencing reads. These intermediaries serve as molecular passports that allow the system to trace each sequencing read back to its parent molecule, enabling distinction between true mutations and artifacts even when mutation fractions approach noise levels.
Solution Approach 2:
The method implements feedback by using error profiles generated from normal samples to inform the analysis of tumor samples. The system continuously refines its understanding of sequencing artifacts and molecular clone structures, using this feedback to improve mutation calling accuracy in subsequent analyses.
4Productivity
If high-throughput detection is implemented, then more samples can be analyzed, but sensitivity and accuracy may be compromised
Solution Approach 1:
The method achieves universality by creating a standardized pipeline that handles both high-throughput requirements and sensitivity needs through modular components. The UMI assignment, molecular cloning reconstruction, and error profile generation steps work consistently across different sample types and throughput levels, maintaining sensitivity while enabling high throughput.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
A genomic data processing system can be configured to process next-generation sequencing information. The genomic data processing system described herein can accurately detect mutations in nucleic acid (e.g., cell free DNA (cfDNA) sequence reads associated with plasma nucleic acid samples. The genomic data processing system of the present disclosure also detects microsatellite instability in nucleic acid sequence reads with a higher degree of sensitivity compared to existing genomic data processing systems.