Multi-Label Tumor Origin Classification Using RNA and Pathology Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining the origin of tumors are computationally demanding and lack accuracy, particularly in classifying tumors of unknown origin, leading to suboptimal clinical treatment recommendations.
Innovation Solution
A multi-label classification system that integrates RNA sequence features and pathology report data, utilizing machine learning to model tumor patterns and improve classification accuracy by incorporating adaptable classifier ensembles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing methods for determining tumor origin are used, then the process can be performed, but the computational demand is excessive and accuracy is insufficient
Solution Approach 1:
The patent segments the classification task into multiple specialized classifier ensembles, each trained on specific subsets of tumor types or data features. This divides the computationally intensive monolithic problem into smaller, more manageable components that can be executed more efficiently while maintaining or improving overall classification accuracy through the coordinated output of multiple specialized models.
Solution Approach 2:
The patent transforms the input data by extracting and emphasizing specific gene expression parameters and pathological features that are most discriminative for tumor origin classification. By changing the parameter representation from raw comprehensive data to curated feature sets with higher signal-to-noise ratio, the system achieves improved accuracy with reduced computational burden.
2Measurement precision
If existing classification methods are applied, then tumors can be classified, but the precision rate remains below 93-96% for tumors of unknown origin
Solution Approach 1:
The patent merges multiple classification approaches by integrating RNA sequence features with pathology report data through an ensemble framework. This combination of diverse data sources and classification methods produces complementary results that collectively achieve precision rates of 93-96% for tumors of unknown origin, while the diversity of the ensemble members enhances overall classification reliability through error cancellation and consensus building.
Data Source
AI summary
Systems and methods are provided for determining a cancer type of a somatic tissue in a subject. A first plurality of sequence reads is obtained from a plurality of RNA molecules in a biopsy of the subject. A first set of sequence features comprising relative mRNA abundance values of genes is determined from the first plurality of sequence reads. Sequence features are applied to a classification model trained to distinguish between each cancer type in a set of at least 50 cancer types, thus determining the cancer type of the somatic tissue in the subject. The classification model provides an indication that the somatic tissue is or is not a respective cancer type, and the set of cancer types includes at least two cancer types from one or more classes of cancer selected from the group consisting of hematological cancers, squamous cancers, endometrial cancers, sarcoma cancers, and neuroendocrine cancers.


