Deep-Learning Whole-Genome Sequencing for Low-Abundance ctDNA Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting circulating tumor DNA (ctDNA) in low tumor burden settings face significant challenges due to limited input material, leading to reduced detection sensitivity and false negatives, especially in non-invasive liquid biopsies, which are crucial for early-stage cancer diagnosis and monitoring.

Innovation Solution

A deep learning-based approach utilizing a convolutional neural network to analyze sequence fragments from plasma samples, enhancing the detection of ctDNA by integrating regional and local probabilities through a tensor-based classification system, improving sensitivity and specificity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ultra-deep sequencing is applied to detect ctDNA, then sequencing depth is improved, but detection sensitivity deteriorates due to limited input material

Engineering Contradiction:
Improvesequencing depthVSAvoiddetection sensitivity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the approach from increasing sequencing depth to increasing the number of genomic loci analyzed. By transitioning from targeted sequencing of a few genes to whole-genome sequencing, the system analyzes thousands of potential mutation sites simultaneously, effectively compensating for the limited number of ctDNA molecules through statistical aggregation across the entire genome.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent shifts the detection dimension from depth (sequencing coverage at a single locus) to breadth (number of loci analyzed). Instead of sequencing the same region multiple times, the system sequences the entire genome at lower depth, using the aggregate signal across all genomic locations to achieve detection sensitivity that overcomes the limited input material constraint.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If statistical methods require multiple independent observations, then measurement precision is improved, but detection capability deteriorates in low tumor fraction settings

Engineering Contradiction:
Improvemutation detection accuracyVSAvoidctDNA detection capability
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent merges the detection signals from thousands of genomic loci into a unified analysis framework. By combining the observations across the entire genome rather than requiring multiple observations at each individual locus, the system achieves sufficient statistical power to detect low-frequency mutations even when each individual site has limited coverage.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal detection framework that applies the same analytical approach across all genomic loci simultaneously. This multi-functional system can detect mutations throughout the entire genome using the same methodology, eliminating the need for locus-specific deep sequencing while maintaining detection accuracy through aggregate statistical analysis.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250250636A1Ultra-sensitive liquid biopsy through deep learning empowered whole genome sequencing of plasma
Publication Date: 2025.08.07 NEW YORK GENOME CENT
  • US20250250636A1 patent drawing
  • US20250250636A1 patent drawing
  • US20250250636A1 patent drawing

AI summary

Systems, methods, and computer program products are provided for classifying sequence fragments and labelling sequence fragments that represent tumor markers. A plurality of reference sequences are read. A plurality of sequence fragments obtained from a biological sample of a patient are read. A first read and a second read are selected from the plurality of sequence fragments. A regional probability based on a plurality of regional features from the patient is received from a first trained classifier. A tensor is generated comprising a corresponding reference sequence, the first read, the second read, a first position, a second position, and an alt position. A local probability based on the tensor is received from a second trained classifier comprising a convolutional neural network. A label associated with a tumor marker is determined when the regional probability is above a first predetermined threshold and the local probability is above a second predetermined threshold.