Hierarchical ML Classifiers for Molecular Category Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying effective therapies for cancer, particularly cancer of unknown primary (CUP), are limited by the inability to accurately determine the primary site of origin, leading to ineffective treatment selection.

Innovation Solution

The use of machine learning techniques to analyze RNA and DNA expression data through a hierarchy of classifiers corresponding to molecular categories, allowing for the identification of candidate molecular categories for a biological sample.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods are used to identify cancer primary site of origin, then the classification is based on organ or tissue location, but the accuracy of determining primary site is insufficient leading to ineffective treatment selection

Engineering Contradiction:
Improveaccuracy of primary site identificationVSAvoideffectiveness of treatment selection
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transforms the classification approach by changing the parameters used for cancer identification. Instead of relying on anatomical location parameters, the invention uses molecular expression parameters (RNA and DNA expression profiles) to characterize tumors. This parameter transformation enables more precise identification of cancer subtypes and molecular categories, directly addressing the inaccuracy of traditional primary site identification methods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces molecular expression data as an intermediary between the tumor sample and the classification outcome. By using RNA and DNA expression profiles as intermediate indicators, the system can infer cancer characteristics and molecular categories without directly observing the primary site. This intermediary approach resolves the contradiction by providing accurate molecular classification even when primary site identification fails.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a single classifier is used for cancer classification, then the system is simple, but it cannot accurately identify multiple molecular categories and their hierarchical relationships

Engineering Contradiction:
Improveaccuracy of molecular category identificationVSAvoidcomplexity of classifier system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the classification task into multiple specialized classifiers, each trained to identify specific molecular categories or subcategories. This segmentation allows the system to handle different molecular categories with dedicated expertise, improving overall classification accuracy. The hierarchical structure divides the complex classification problem into manageable segments that can be processed independently and then integrated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a hierarchical dimension to the classification system by organizing classifiers in a tree structure with parent-child relationships. This dimensional transformation from a flat single-classifier approach to a hierarchical multi-classifier system enables the model to capture nested molecular categories and their relationships, significantly improving the precision of molecular category identification while systematically managing the complexity through structured organization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If comprehensive RNA expression data from multiple gene sets is analyzed, then the molecular category identification becomes more accurate, but the processing complexity and computational requirements increase

Engineering Contradiction:
Improveaccuracy of molecular category characterizationVSAvoidcomplexity of data processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive RNA expression data into multiple distinct gene sets, with each set associated with specific molecular categories or biological pathways. This segmentation allows the system to process large volumes of expression data in manageable portions, reducing computational complexity while maintaining the ability to accurately characterize molecular categories through the aggregated information from multiple gene sets.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal hierarchical classifier system that can process multiple types of input data (different gene sets, RNA and DNA expression data) through a unified framework. This multi-functional system handles diverse data inputs using the same hierarchical classification architecture, reducing the need for separate processing pipelines for each data type and thereby managing complexity while comprehensively analyzing all expression data.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250125013A1Hierarchical machine learning techniques for identifying molecular categories from expression data
Publication Date: 2025.04.17 BOSTONGENE CORP
  • US20250125013A1 patent drawing
  • US20250125013A1 patent drawing
  • US20250125013A1 patent drawing

AI summary

Described herein in some embodiments is a method comprising: obtaining expression data previously obtained by processing a biological sample obtained from a subject; processing the expression data using a hierarchy of machine learning classifiers corresponding to a hierarchy of molecular categories to obtain machine learning classifier outputs including a first output and a second output, the hierarchy of molecular categories including a parent molecular category and first and second molecular categories that are children of the parent molecular category in the hierarchy of molecular categories, the hierarchy of machine learning classifiers comprising first and second machine learning classifiers corresponding to the first and second molecular categories; and identifying, using at least some of the machine learning classifier outputs including the first output and the second output, at least one candidate molecular category for the biological sample.