Multi-Modal Disease Progression Prediction via Feature Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in accurately predicting medical condition progression and treatment response due to high heterogeneity across subjects, tumors, and regions within a tumor, as well as the complexity of large biological data sets, which often leads to subjective and inaccurate predictions.

Innovation Solution

The integration of digital pathology data, gene-expression data, and radiology data using machine-learning models to select informative features and generate predictive results, employing techniques like variable focusing and bootstrapping to reduce data dimensionality and prevent overfitting, thereby improving prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple data types (digital pathology, gene-expression, radiology) are integrated to improve prediction accuracy, then the comprehensiveness and accuracy of medical condition progression prediction is improved, but the complexity of data processing and model integration increases

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex multi-modal data integration problem into distinct processing modules: digital pathology data processing, gene-expression data processing, and radiology data processing. Each modality is processed independently through its own pipeline before being integrated, reducing the overall complexity by dividing the system into manageable segments that can be developed and validated separately.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary representations for each data modality that translate complex multi-omic and imaging data into standardized feature formats. These intermediaries serve as mediators between the diverse input data types and the final prediction model, enabling integration without requiring direct processing of all raw data types simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If variable focusing techniques are applied to reduce data dimensionality, then the computational efficiency and model training speed are improved, but the risk of information loss from reducing tens of thousands of genes to smaller subsets increases

Engineering Contradiction:
Improvemodel training speedVSAvoidgene expression information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies preliminary filtering and selection techniques to gene-expression data before main model training, pre-identifying and prioritizing genes with higher relevance to the medical condition. This preliminary action reduces the dimensionality early in the pipeline, allowing faster subsequent processing while preserving the most informative genetic features.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical feature selection methods with machine learning-based approaches that can identify non-linear relationships and interactions among genes. This substitution enables more intelligent dimensionality reduction that preserves information about complex biological relationships while reducing computational burden.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If bootstrapping techniques are used to prevent overfitting, then the model generalization capability is improved, but the computational time and processing resources required for training increase

Engineering Contradiction:
Improvemodel generalizationVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a partial bootstrapping approach where resampling is applied selectively to certain data subsets or at specific training stages rather than uniformly across all training data. This partial application maintains the overfitting prevention benefits while reducing the computational overhead of full bootstrapping across the entire dataset.

Inventive Principle:
Principle #16Partial or excessive action

4Ease of operation

If subjective human assessments and static rules are used to select data portions, then the simplicity and ease of implementation are maintained, but the accuracy and objectivity of disease progression predictions deteriorate

Engineering Contradiction:
Improveimplementation simplicityVSAvoidprediction accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent systematically replaces subjective human assessment and static rule-based selection with machine learning models that objectively analyze multi-omic and imaging data. These models automatically identify patterns and relationships in the data that would be difficult for human experts to detect, providing more accurate and consistent predictions while eliminating observer bias.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240038393A1Predicting disease progression based on digital-pathology and gene-expression data
Publication Date: 2024.02.01 GENENTECH INC
  • US20240038393A1 patent drawing
  • US20240038393A1 patent drawing
  • US20240038393A1 patent drawing

AI summary

In some embodiments, a current state of a medical condition or a progression of the medical condition is predicted by processing one or more digital pathology images and expression levels of genes using a machine-learning model. In some embodiments, one or more predicted gene-expression levels are generated by processing a data set corresponding to one or more digital pathology images using a machine-learning model. In some embodiments, one or more predicted digital pathology metrics are generated by processing a data set that corresponds to expression levels of a set of genes using a machine-learning model.