Multi-Domain Tractability Scoring for Protein–Compound Predictions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning systems suffer from inaccuracies, inefficiencies, and operational inflexibility in generating predictions, particularly in multi-domain models for compound-protein interactions, leading to wasteful computational resources and unreliable bioactivity programs.

Innovation Solution

A multi-domain tractability system utilizing a tractability machine learning model generates protein-model tractability scores to assess the accuracy of compound-protein interaction models, providing context and flexibility in downstream applications through protein-model tractability scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional machine learning systems are used to generate predictions, then operational simplicity is maintained, but prediction accuracy deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system is divided into multiple independent machine learning models, each specialized for a specific domain (e.g., protein binding affinity prediction, compound property prediction). This segmentation allows each model to be optimized for its specific task, improving overall prediction accuracy while maintaining manageable complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A tractability machine learning model serves as an intermediary that generates tractability scores to assess the reliability of predictions from other models. This intermediary layer filters and weights predictions before final decision-making, improving accuracy by accounting for model uncertainty without requiring complete redesign of the prediction system.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If large volumes of training data are used to improve prediction accuracy, then measurement precision improves, but loss of time and computational resources increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of training all models exhaustively on complete datasets, the system uses partial training approaches where models are trained on representative subsets of data or using incremental learning. This reduces training time while maintaining sufficient accuracy for practical applications, balancing computational investment with prediction quality.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary data processing and feature extraction before model training, and uses pre-trained models where applicable. This preliminary action prepares data in advance and leverages existing knowledge, reducing the computational burden and time required during actual model training and prediction phases.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If conventional machine learning systems are used, then device complexity is maintained, but reliability of predictions deteriorates

Engineering Contradiction:
Improveprediction reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The tractability machine learning model provides feedback on the reliability of predictions generated by other models. By continuously assessing prediction quality and adjusting tractability scores based on model performance, the system improves overall reliability through iterative refinement without requiring complete system redesign.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system combines multiple machine learning models with different strengths and characteristics into a composite prediction system. Each model contributes its specialized capabilities, and their predictions are integrated through the tractability scoring mechanism, creating a more reliable overall system than any single model could achieve alone.

Inventive Principle:
Principle #40Composite materials

4Measurement precision

If multi-domain machine learning models are implemented to improve prediction accuracy, then measurement precision improves, but operational flexibility deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidoperational flexibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The multi-domain prediction system is segmented into independent, specialized models for different domains (protein interactions, compound properties, etc.). This segmentation maintains operational flexibility by allowing individual models to be selected and executed based on specific prediction needs, while still achieving high accuracy through domain-specific expertise of each model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The tractability machine learning model serves as a universal intermediary that can assess the reliability of predictions from any of the specialized domain models. This multi-functional component maintains operational flexibility by providing a consistent evaluation framework across different domains without requiring domain-specific customization of the assessment mechanism.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250272607A1Utilizing a tractability machine learning model to generate tractability scores for a multi-domain machine learning model for improved machine learning predictions
Publication Date: 2025.08.28 RECURSION PHARMACEUTICALS INC
  • US20250272607A1 patent drawing
  • US20250272607A1 patent drawing
  • US20250272607A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods that utilize a multi-domain tractability machine learning model to generate tractability scores for a multi-domain machine learning model and further generate improved bioactivity predictions. Indeed, in one or more implementations, the disclosed systems generate a predicted match score between a target protein and a target compound using a compound-protein interaction machine learning model. For instance, the disclosed systems generate a protein-model tractability score that indicates a measure of accuracy of the compound-protein interaction machine learning model relative to the target protein. Moreover, in some instances, the disclosed systems utilize the protein-model tractability score by providing the protein-model tractability score in conjunction with the predicted match score or the target protein or the disclosed systems generate a bioactivity prediction from the predicted match score and the protein-model tractability score.