Compound-Protein ML Representation for Bioactivity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for predicting bioactivity using machine learning models face challenges in accuracy, efficiency, and operational flexibility, particularly in modeling complex biological interactions and identifying contributing proteins.

Innovation Solution

The system utilizes a compound-protein interaction machine learning model to generate a compound-protein machine learning representation, which serves as a unique proteome fingerprint. This representation is used to train target machine learning models for predicting bioactivity and to employ explainability models to identify contributing proteins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning models are used to predict bioactivity, then predictions can be generated, but accuracy is insufficient due to inability to effectively model complex compound-protein interactions

Engineering Contradiction:
Improveprediction accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces compound-protein interaction representations as an intermediary layer between input compound data and output bioactivity predictions. These representations capture complex interaction patterns through proteome fingerprinting, serving as a mediator that translates molecular features into biologically meaningful predictions without requiring excessively complex model architectures

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The prediction system is segmented into distinct functional components: compound feature extraction, protein interaction modeling, and bioactivity prediction. This segmentation allows each component to be optimized independently, improving overall accuracy while managing complexity through modular design

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If large volumes of training data are used to improve prediction accuracy, then model performance increases, but computational resources and training time increase significantly

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent extracts essential interaction patterns from training data through compound-protein interaction representations. By capturing the most relevant features through proteome fingerprinting, the system achieves high prediction accuracy without requiring exhaustive use of all available training data, thereby reducing computational resource consumption

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary action by pre-computing compound-protein interaction representations and storing them as reusable features. This pre-processing step transforms raw training data into condensed interaction patterns that can be efficiently utilized during prediction, reducing the computational burden of processing large datasets during model training and inference

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If conventional machine learning models are used, then predictions can be generated, but operational flexibility is limited in identifying contributing proteins and underlying biological mechanisms

Engineering Contradiction:
Improveoperational flexibilityVSAvoidbiological mechanism information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms that trace predictions back to contributing proteins and interaction patterns. By maintaining compound-protein interaction representations throughout the prediction process, the system can provide feedback about which proteins and biological mechanisms drive specific predictions, enhancing operational flexibility and interpretability without losing critical biological information

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250157569A1Utilizing compound-protein machine learning representations to generate bioactivity predictions
Publication Date: 2025.05.15 RECURSION PHARMACEUTICALS INC
  • US20250157569A1 patent drawing
  • US20250157569A1 patent drawing
  • US20250157569A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods that utilizing compound-protein machine learning representations to generate target results. For example, the disclosed systems can utilize a compound-protein interaction machine learning model to generate a compound-protein machine learning representation for compound protein pairs. The disclosed systems can utilize the compound-protein machine learning representation to train and utilize other target machine learning models in generating predicted bioactivity results. For example, the disclosed systems train a target machine learning model from compound-protein machine learning representations to generate ADMET predictions and/or biological perturbation program predictions. Furthermore, the disclosed systems can utilize one or more explainability models in conjunction with target machine learning models trained based on compound-protein machine learning representations to identify proteins that contribute to predicted bioactivity results.