Binding Affinity Prediction Using Protein Sequence Tokens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting binding affinity between chemical or biological molecules and their protein targets are inefficient due to the complexity of protein-protein and protein-small molecule interactions, reliance on experimentally verified 3D structural data, and the time-consuming nature of conventional approaches, which fail to accurately predict binding affinities especially for novel proteins and disordered states.

Innovation Solution

A system and method using machine learning models to preprocess and convert protein and molecule data into tokens, generating representation models and attention maps to predict binding affinity, allowing for accurate prediction without relying on experimental 3D structure information, and incorporating data preprocessing techniques like outlier correction and data augmentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional binding affinity prediction tools are used, then structural information can be processed, but they fail to accurately predict binding affinities for novel proteins and disordered states

Engineering Contradiction:
Improvebinding affinity prediction accuracyVSAvoidapplicability to novel proteins and disordered states
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent uses language model pre-trained weights and knowledge from large corpora of protein sequences to create a representation model that captures general protein characteristics. This pre-trained model is then fine-tuned on binding affinity data, allowing the system to generalize to novel proteins and disordered states without requiring experimental 3D structure information.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the input representation from traditional 3D structural parameters to sequence-based token representations. By changing the fundamental input parameters from structural coordinates to amino acid sequences processed through language model embeddings, the system can handle novel proteins and disordered regions that lack stable 3D structures.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If experimentally observed data is collected, then accurate binding affinity information is obtained, but it requires a lot of effort and time

Engineering Contradiction:
Improvebinding affinity data accuracyVSAvoidtime and effort for experimental observation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the language model on large corpora of protein sequences and binding data before actual prediction. This pre-training phase captures general patterns and relationships, so that when new protein-molecule pairs need to be evaluated, the system can make accurate predictions without requiring time-consuming experimental measurements for each specific case.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical experimental system (wet lab binding assays) with an computational language model-based system. Instead of physically measuring binding affinities through experimental procedures, the system uses pre-trained language model representations to predict binding affinities, dramatically reducing time and effort while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If 3D structural information is obtained, then binding affinity can be predicted using minimized energy models, but it is hard to predict 3D structure from protein sequence and proteins may change shape

Engineering Contradiction:
Improvebinding affinity predictionVSAvoidcomplexity of 3D structure prediction
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential information needed for binding affinity prediction directly from protein sequences without requiring the intermediate step of 3D structure determination. By taking out the dependency on 3D structural information and working directly with sequence-based language model representations, the system simplifies the overall process while maintaining prediction accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of following the traditional approach of predicting 3D structure from sequence and then using that structure for binding affinity prediction, the patent inverts the approach by directly predicting binding affinity from sequence representations. This reversal eliminates the complex intermediate step of 3D structure prediction while achieving the same or better results.

Inventive Principle:
Principle #13The other way round (Inversion)

4Extent of automation

If virtual screening methods are used to shortlist compounds, then the process can be automated, but they are time consuming and lack generalization and accuracy

Engineering Contradiction:
Improveautomation of compound screeningVSAvoidscreening speed and accuracy
Core Design Contradiction:
Extent of automationVSProductivity

Solution Approach 1:

The patent replaces traditional virtual screening methods (molecular docking, energy minimization) with a language model-based prediction system. This substitution maintains automation while dramatically improving both speed and accuracy, as the language model can evaluate binding affinities without the computationally intensive steps of structural alignment and energy calculation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the evaluation parameters from traditional structural parameters (binding pose, interaction energy) to language model-based semantic representations. By transforming the screening criteria into token-based predictions, the system achieves faster evaluation while improving generalization across different protein-molecule pairs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230326545A1System and method for predicting biological activity of chemical or biological molecules and evidence thereof
Publication Date: 2023.10.12 PEPTRIS TECH PTE LTD
  • US20230326545A1 patent drawing
  • US20230326545A1 patent drawing
  • US20230326545A1 patent drawing

AI summary

A system 100 for predicting binding affinity of chemical or biological molecules and their protein targets and generating pair-wise attention map as an evidence of binding between the chemical or biological molecules and their protein targets is provided. The system 100 includes a binding activity predicting system 104 receives the knowledge data of the chemical or biological molecules and their protein targets from the global knowledge database 102 and processes the knowledge data to convert into tokens of proteins and tokens of molecules. The tokens of protein and tokens of molecules are used to train a protein and molecule representation model to predict biological activity. The protein and molecule representation model is used to train a binding activity prediction model to predict binding affinities and to generate pair-wise attention maps as likelihoods of biological activity between amino acid residues and fragments involved in binding.