Deep-Learning Whole-Genome Sequencing for Low-Abundance ctDNA Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting circulating tumor DNA (ctDNA) in low tumor burden settings face significant challenges due to limited input material, leading to reduced detection sensitivity and false negatives, especially in non-invasive liquid biopsies, which are crucial for early-stage cancer diagnosis and monitoring.
Innovation Solution
A deep learning-based approach utilizing a convolutional neural network to analyze sequence fragments from plasma samples, enhancing the detection of ctDNA by integrating regional and local probabilities through a tensor-based classification system, improving sensitivity and specificity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If ultra-deep sequencing is applied to detect ctDNA, then sequencing depth is improved, but detection sensitivity deteriorates due to limited input material
Solution Approach 1:
The patent changes the approach from increasing sequencing depth to increasing the number of genomic loci analyzed. By transitioning from targeted sequencing of a few genes to whole-genome sequencing, the system analyzes thousands of potential mutation sites simultaneously, effectively compensating for the limited number of ctDNA molecules through statistical aggregation across the entire genome.
Solution Approach 2:
The patent shifts the detection dimension from depth (sequencing coverage at a single locus) to breadth (number of loci analyzed). Instead of sequencing the same region multiple times, the system sequences the entire genome at lower depth, using the aggregate signal across all genomic locations to achieve detection sensitivity that overcomes the limited input material constraint.
2Measurement precision
If statistical methods require multiple independent observations, then measurement precision is improved, but detection capability deteriorates in low tumor fraction settings
Solution Approach 1:
The patent merges the detection signals from thousands of genomic loci into a unified analysis framework. By combining the observations across the entire genome rather than requiring multiple observations at each individual locus, the system achieves sufficient statistical power to detect low-frequency mutations even when each individual site has limited coverage.
Solution Approach 2:
The patent creates a universal detection framework that applies the same analytical approach across all genomic loci simultaneously. This multi-functional system can detect mutations throughout the entire genome using the same methodology, eliminating the need for locus-specific deep sequencing while maintaining detection accuracy through aggregate statistical analysis.
Data Source
AI summary
Systems, methods, and computer program products are provided for classifying sequence fragments and labelling sequence fragments that represent tumor markers. A plurality of reference sequences are read. A plurality of sequence fragments obtained from a biological sample of a patient are read. A first read and a second read are selected from the plurality of sequence fragments. A regional probability based on a plurality of regional features from the patient is received from a first trained classifier. A tensor is generated comprising a corresponding reference sequence, the first read, the second read, a first position, a second position, and an alt position. A local probability based on the tensor is received from a second trained classifier comprising a convolutional neural network. A label associated with a tumor marker is determined when the regional probability is above a first predetermined threshold and the local probability is above a second predetermined threshold.


