Drug-Target Binding Prediction Using Synthetic Ghost Ligands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computational methods for predicting drug-target interactions face challenges due to limited data coverage of the human proteome, high-dimensional feature spaces, and vulnerability to overfitting, leading to inaccurate predictions, especially when applied to new protein systems or drug scaffolds.

Innovation Solution

A method that combines protein-based and ligand-based predictions by generating synthetic data using ghost ligands projected onto 3D protein structures, creating localized features, and training a machine learning model with augmented DTI features to predict drug-target interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If protein-based predictions use 3D molecular structures of ligands co-crystallized with proteins, then biophysical compatibility can be learned, but the method is highly data-bound and computationally demanding

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates synthetic protein-ligand complex structures by copying known ligand binding modes from experimentally determined structures and applying them to template proteins. This generates large amounts of training data without requiring actual experimental structures for every protein-ligand pair, thus reducing computational complexity while maintaining prediction accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The method pre-generates synthetic training data by projecting ligands onto template proteins before actual prediction tasks. This preliminary data generation creates a ready-to-use training set that reduces computational burden during actual prediction while preserving biophysical compatibility information.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If neural networks are trained on 100's to 1000's of different proteins, then the model can learn from diverse data, but the high feature space to data ratio produces false negatives and positives

Engineering Contradiction:
Improvemodel generalizabilityVSAvoidprediction precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a new dimension by creating synthetic 3D structural data from 2D ligand representations and template protein structures. This additional dimensional information enriches the training data without proportionally increasing feature space complexity, improving the feature space to data ratio and reducing false predictions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The method transforms ligand representations by projecting them onto 3D protein templates, changing the parameter space from flat 2D descriptors to spatially-aware 3D coordinates. This parameter transformation increases data information density without linearly increasing model complexity, improving prediction precision.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If ligand-based predictions use DTI databases with millions of records, then large data volume is available, but only ~2,000 of 20,000 human proteins are represented

Engineering Contradiction:
Improvedata volumeVSAvoidprotein coverage
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training framework where synthetic data generated from template proteins can be applied to predict interactions across all 20,000 human proteins. The synthetic projection method is protein-agnostic, allowing the same approach to generate training data for any target protein, thus expanding coverage from 2,000 to 20,000 proteins while utilizing the same data volume.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If individual models are derived for each of the 2,000 proteins, then successful predictions can be made for represented proteins, but computational resources increase and models do not learn physical attributes of drug-protein compatibility

Engineering Contradiction:
Improveprediction reliabilityVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges individual protein-specific models into a single unified model that learns general drug-protein compatibility principles. By training one model on synthetic data from multiple template proteins rather than 2,000 separate models, the approach reduces model complexity while maintaining reliability through learned biophysical patterns that generalize across proteins.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3906556B1Method and system for predicting drug binding using synthetic data
Publication Date: 2026.04.22 CYCLICA INC
  • EP3906556B1 patent drawingFigure 1A
  • EP3906556B1 patent drawingFigure 1B
  • EP3906556B1 patent drawingFigure 1C

AI summary

A method for predicting drug-target binding using synthetically-augmented data involves generating a multitude of ghost ligands for a multitude of proteins in a protein structure database, generating a multitude of drug-target interaction (DTI) features for proteins and ligands in a DTI database, using the multitude of ghost ligands, generating a machine learning model using the multitude of DTI features, and predicting a likelihood of interaction for a combination of a query protein and a query ligand using the machine learning model.