Knowledge Graph Gene Expression Perturbation Prediction for Drug Response

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing gene expression prediction models lack the simultaneous integration of multi-omics data and prior domain knowledge of molecular interactions, leading to inaccurate predictions of drug responses due to the omission of patient-specific contextual factors and complex molecular interactions.

Innovation Solution

A method involving the generation of a knowledge graph (KG) that integrates domain knowledge about gene and perturbation agent interactions, followed by training a machine-learning model to predict perturbed gene expression using learned embeddings, thereby incorporating prior knowledge of molecular interactions and improving prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing prediction models are used, then the model complexity is low and ease of operation is maintained, but the prediction accuracy is insufficient due to lack of multi-omics data integration and domain knowledge

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple data sources (multi-omics data including gene expression, protein abundance, and metabolite levels) with domain knowledge (molecular interaction networks, pathway databases) into a unified prediction framework. This integration combines heterogeneous information types to improve prediction accuracy while managing complexity through systematic data fusion strategies

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces intermediate representation layers that translate raw multi-omics data and domain knowledge into unified feature vectors. These intermediary representations serve as bridges between diverse data modalities and the prediction model, enabling accurate integration without overwhelming model complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If existing models are used, then the model structure is simple, but the ability to capture patient-specific contextual factors and molecular interactions is insufficient

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by incorporating patient-specific contextual factors into the prediction model. Instead of using a uniform approach for all patients, the model adapts to individual patient characteristics, molecular interaction networks, and omics data profiles, allowing personalized predictions that reflect unique biological contexts

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces dynamic elements that allow the model to adapt to different patient contexts and molecular interaction patterns. The framework dynamically weights and integrates various data sources based on their relevance to each specific prediction task, enabling versatility without requiring completely separate models for each application

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250226058A1Method and system for predicting gene expression perturbations
Publication Date: 2025.07.10 NEC LAB EURO GMBH
  • US20250226058A1 patent drawing
  • US20250226058A1 patent drawing
  • US20250226058A1 patent drawing

AI summary

A computer-implemented method for predicting gene expression perturbations includes generating a knowledge graph (KG) from domain knowledge, where the KG describes relations including associations, similarities and/or interactions between a plurality of entities, the plurality of entities including at least a number of genes and perturbation agents. A machine-learning (ML) model is trained to predict perturbed gene expression from pre-perturbed gene expression data and learned embeddings of the plurality of entities of the KG. Gene expression data obtained from a subject-derived gene sample is provided and the trained ML model is used to predict a response of the gene sample in terms of gene expression changes effected by applying one or more perturbation agents to the gene sample.