Parallel Cancer Source of Origin Classification via Methylation Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cancer detection methods are limited by low detection rates and high false positive rates, particularly in early-stage cancers, and lack cross-applicability across different cancer types, making them ineffective for general population screening.

Innovation Solution

The development of parallel cancer signal origin (CSO) classifiers that separately predict the organ type affected by cancer and the tumor biology type, using methylation sequencing data and avoiding the need for multiple sequencing assays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a single binary cancer detection test is used, then the test is simple and cost-effective, but it provides insufficient granularity for diagnostic workup and treatment decisions

Engineering Contradiction:
Improvecancer source of origin informationVSAvoidclassification system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The single cancer detection test is segmented into two independent parallel classifiers: one for organ type prediction and one for tumor biology type prediction. Each classifier receives the same methylation sequencing data but processes it independently to provide specialized predictions, thereby increasing information granularity without requiring separate sequencing assays for each prediction type.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If multiple sequencing assays are performed to obtain both organ type and tumor biology type predictions, then comprehensive cancer information is obtained, but the assaying process becomes more complex and costly

Engineering Contradiction:
Improvecancer classification informationVSAvoidsequencing assay complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges the data processing workflows by using a single methylation sequencing assay to generate feature vectors that are then fed to both the organ type classifier and the tumor biology type classifier. This combining approach eliminates the need for separate sequencing assays while maintaining comprehensive cancer information through parallel independent predictions from the unified data source.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If parallel classifiers are trained to separately predict organ type and tumor biology type, then prediction granularity increases, but training time and computational resources increase

Engineering Contradiction:
Improvecancer source of origin prediction precisionVSAvoidclassifier training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary feature extraction from methylation sequencing data to create feature vectors that capture methylation patterns. These pre-processed feature vectors are then used by both classifiers during training, eliminating the need for both classifiers to process raw sequencing data independently. This preliminary action reduces redundant computational work and accelerates the training process while maintaining high prediction precision.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250132055A1Parallel cancer source of origin classification for organ type and tumor biology type
Publication Date: 2025.04.24 GRAIL INC
  • US20250132055A1 patent drawing
  • US20250132055A1 patent drawing
  • US20250132055A1 patent drawing

AI summary

Methods for cancer source of origin (CSO) prediction are disclosed to predict CSO characteristics. The CSO prediction may include the affected organ or organ group and tumor biology. The method for training parallel CSO classifiers includes obtaining training samples derived from subjects with known cancer diagnosis, each training sample comprising methylation sequence reads corresponding to nucleic acid fragments in a biological sample collected from each subject and each known cancer signal origin including a known affected organ or organ group a plurality of organs or organ groups and a known tumor biology from a plurality of tumor biology classes. The method includes generating, for each training sample, a feature vector based on the methylation sequence reads. The method includes generating a first training data set comprising the feature vectors for the training samples and the known organs or organ groups, and training an organ or organ group classifier with the first training data set to predict organ or organ group from the plurality of organs or organ groups based on an input feature vector. The method includes generating a second training data set comprising the feature vectors for the training samples and the known tumor biology classes, and training a tumor biology classifier with the second training data set to predict tumor biology from the plurality of tumor biology classes based on input feature vector.