Structure-Based Ligand Activity Prediction With Binding Mode Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing structure-based activity prediction methods for ligands suffer from dataset bias and misrepresentation of ligand-protein interactions due to randomly distributed 'incorrect' poses, leading to inaccurate activity predictions.

Innovation Solution

A system and method that incorporates a binding mode prediction model with transfer learning techniques to improve activity prediction by selecting reliable binding modes using a binding mode selector, enhancing the performance of activity prediction models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If docking programs are used to sample binding modes, then binding mode predictions can be obtained, but incorrect poses are randomly distributed leading to misrepresentation of ligand-protein contacts

Engineering Contradiction:
Improvebinding mode prediction accuracyVSAvoidrepresentation of ligand-protein contacts
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces an intermediate step between docking and activity prediction: a binding mode prediction model that evaluates and selects the most accurate binding modes from docking outputs. This intermediary filters out incorrect poses before they can corrupt the activity prediction training data, thereby resolving the contradiction between obtaining binding mode predictions and ensuring reliable ligand-protein contact representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The binding mode prediction model performs preliminary evaluation and selection of binding modes before the activity prediction model processes the data. By pre-filtering and pre-ranking binding modes based on their accuracy, the system ensures that only high-quality binding mode data is used for training activity predictions, thus preventing the propagation of errors from randomly distributed incorrect poses.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If deep learning neural networks are combined with structural data, then activity prediction performance can be improved, but dataset bias occurs making protein-related features irrelevant

Engineering Contradiction:
Improveactivity prediction accuracyVSAvoidrelevance of protein-related features
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by focusing the neural network's attention on specific, relevant local features within the binding mode data. Rather than using all structural data equally, the system identifies and weights important local interactions (such as key ligand-protein contacts) more heavily, making the model adapt to the specific task of activity prediction while filtering out irrelevant protein features through the binding mode selection mechanism.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If all generated docked structures are used for training, then more training data is available, but incorrect poses lead to misrepresented contacts and inaccurate predictions

Engineering Contradiction:
Improvetraining data volumeVSAvoidactivity prediction accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies partial action by selectively using only a subset of the generated docked structures for training - specifically, only those binding modes that are predicted to be accurate by the binding mode prediction model. Rather than using all available docked structures (excessive action), the system filters to use only the necessary high-quality portion, thereby maintaining adequate training data volume while ensuring accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12412638B2Structure-based, ligand activity prediction using binding mode prediction information
Publication Date: 2025.09.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12412638B2 patent drawing
  • US12412638B2 patent drawing
  • US12412638B2 patent drawing

AI summary

A system and method for structure-based, small molecule activity prediction using binding mode prediction information. Binding scores between ligands and target molecules, (e.g. proteins, RNA, DNA, lipids, sugars) are first generated using molecular docking. A first machine learned deep neural network (DNN) model is developed using data representing the molecular ligand-target pair 3D structures and docking features to predict binding modes. Using transfer learning, weights of layers learned in the first machine learned model are used as weights in layers of a second machine learned DNN model used to more accurately improve the performance of activity prediction of the second machine learned model. For a target newly paired ligand-target complex, the method further implements a binding mode selector for selecting one or more particular binding poses for input to the activity prediction model for use in activity mode prediction of an activity of the target paired ligand-protein complex.