Drug target binding affinity prediction method and system based on multi-level connection diagram

By adopting a multi-level connection map framework and attention mechanism feature fusion method in drug target binding affinity prediction, the problem of low prediction accuracy in cold start scenarios is solved, and more efficient and robust drug target binding affinity prediction is achieved.

CN120108488APending Publication Date: 2025-06-06SHENZHEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510090269.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology has low accuracy in predicting drug target binding affinity in cold start scenarios and cannot effectively reflect the true generalization performance in the face of brand new and unseen data.

Method used

A drug target binding affinity prediction method based on multi-level connection map is used to obtain drug small molecules, protein targets and affinity maps, and pretreat the multi-level connection maps, extract node features and perform feature fusion, and integrate drug and target features using attention mechanisms, and finally input into a classifier for prediction.

Benefits of technology

It significantly improves the accuracy and robustness of drug target binding affinity prediction in cold-start scenarios, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108488A_ABST
    Figure CN120108488A_ABST
Patent Text Reader

Abstract

The invention provides a drug target binding affinity prediction method and system based on a multi-level connection diagram, and relates to the technical field of computational biology, the method comprises the following steps: preprocessing drug small molecules, protein targets and affinity diagrams to obtain the multi-level connection diagram, and performing node feature extraction; obtaining drug isomorphic connection diagram embedding, drug internal connection diagram embedding, protein isomorphic connection diagram embedding, protein internal connection diagram embedding and heterogeneous connection diagram embedding; fusing the heterogeneous connection diagram embedding with the drug internal connection diagram embedding and the protein internal connection diagram embedding respectively to obtain drug internal connection diagram node features and protein internal connection diagram node features; and performing feature fusion based on an attention mechanism on the drug internal connection diagram node features, the protein internal connection diagram node features, drug isomorphic connection diagram embedding and protein isomorphic connection diagram embedding to obtain a drug-protein interaction feature input classifier, and outputting a prediction result of the drug target binding affinity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational biology, and in particular to a method and system for predicting drug target binding affinity based on a multi-level connection graph. Background Art

[0002] Drug discovery refers to the process of identifying drug molecules with therapeutic potential and developing safe and effective drug candidates, which has a profound impact on society, the economy and personal health. Drug development is a long and complex process, and each stage from target identification to clinical trials requires a lot of time and resources. As the R&D process progresses, the cost of each step is increasing. Therefore, it is particularly important to accurately screen the most promising drug candidates before entering the next stage [1].

[0003] At present, drug-target affinity prediction is a computational method for evaluating the binding strength between drug candidate molecules and biological targets (such as proteins, nucleic acids, etc.). If the binding strength between drug molecules and targets is high, it usually means that the drug has significant affinity for the target, thereby increasing its possibility as a candidate drug for treating diseases related to the target. The prior art discloses deep learning strategies for drug-target binding affinity prediction methods, but although deep learning strategies have shown excellent prediction effects on drug-target binding affinity prediction problems, the evaluation of these methods is mostly based on randomly divided data sets. In other words, the drug molecules and target proteins involved in the test data set may have been included in the training data set, and the model trained by such data segmentation cannot fully reflect its true generalization performance when facing new and unseen data. [In practical applications, the challenge we face is often to predict targets that are not associated with any known drugs, or when developing new drugs, we do not know which targets they may act on. This type of situation is usually defined as a "cold start problem." When the data set is segmented according to this more realistic scenario, the drug-target binding affinity prediction performance of many models tends to drop sharply. For this reason, some researchers have turned their attention to the cold start problem, making some progress in drug-target affinity prediction in cold start scenarios. However, their focus is mainly on optimizing the mechanism for extracting interactive information between drugs and targets, but they ignore the introduction of additional auxiliary information to enhance the generalization ability of the model, resulting in low accuracy in drug-target binding affinity prediction in cold start scenarios. Summary of the invention

[0004] In order to solve the problem of low accuracy in predicting drug target binding affinity in the above-mentioned prior art in the cold start scenario, the present invention proposes a drug target binding affinity prediction method and system based on a multi-level connection graph, which can effectively improve the accuracy of drug target binding affinity prediction in the cold start scenario.

[0005] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0006] 1. A method for predicting drug target binding affinity based on a multi-level connection graph, comprising the following steps:

[0007] S1. Obtaining a drug small molecule, a protein target, and an affinity map for representing the interaction relationship between the drug small molecule and the protein target;

[0008] S2. Preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map;

[0009] S3. Extract node features from the multi-level connection graph to obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding;

[0010] S4. fusing the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding respectively to obtain drug internal connection graph node features and protein internal connection graph node features;

[0011] S5. Perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph, and the embedding of the protein isomorphic connection graph to obtain drug-protein interaction features;

[0012] S6. Input the drug-protein interaction features into a classifier and output a prediction result of drug-target binding affinity.

[0013] Preferably, the nodes of the affinity graph include drug small molecule nodes and protein target nodes, and the edges of the affinity graph are the binding affinities of the drug small molecule nodes and the protein target nodes.

[0014] Preferably, the multi-level connection diagram includes a drug internal connection diagram, a drug isomorphic connection diagram, a protein internal connection diagram, a protein isomorphic connection diagram and a heterogeneous connection diagram.

[0015] Preferably, the pre-processing of the drug small molecule, the protein target and the affinity map comprises:

[0016] Inputting the drug small molecule into a pre-trained molecular representation learning model, extracting the characteristic representation of the drug small molecule through the molecular representation learning model, and outputting the drug internal connection graph and the drug isomorphic connection graph respectively;

[0017] Inputting the protein target into a pre-trained protein representation deep learning model, extracting feature representation of the protein target through the protein representation deep learning model, and outputting the protein internal connection graph and the protein isomorphic connection graph respectively;

[0018] The binding affinity matrix in the affinity graph is used as the feature of the node in the heterogeneous connection graph to obtain a matrix feature vector, and one-hot encoding is used to distinguish the drug small molecule nodes and the protein target nodes in the affinity graph to obtain a one-hot encoding vector, and the matrix feature vector and the one-hot encoding vector are concatenated to obtain the initial features of the heterogeneous connection graph.

[0019] Preferably, the extracting node features from the multi-level connection graph comprises:

[0020] The drug isomorphic connection graph G D and protein isomorphic connection graph G T Sparse, the sparse drug isomorphic connection graph G D and protein isomorphic connection graph G T Input a preset graph attention network, which outputs an initial drug isomorphic connection graph embedding and a protein isomorphic connection graph embedding, and inputs the initial drug isomorphic connection graph embedding into a multi-layer perceptron, which outputs a final drug isomorphic connection graph embedding. The initial protein isomorphic connection graph is embedded into the multi-layer perceptron, and the multi-layer perceptron outputs the final protein isomorphic connection graph embedding

[0021] The internal connection diagram of the drug G d and protein internal connectivity diagram G t The graph attention network is used to output the initial drug internal connection graph embedding. and the initial protein internal connectivity graph embedding

[0022] The heterogeneous connection graph is input into a preset graph convolutional neural network, and the preset graph convolutional neural network outputs the heterogeneous connection graph and embeds it into M A .

[0023] Preferably, said fusing said heterogeneous connection graph embedding with said drug internal connection graph embedding and said protein internal connection graph embedding respectively comprises:

[0024] S41. Embedding the heterogeneous connection graph into the initial drug internal connection graph The fusion is performed as follows:

[0025]

[0026] in, It represents the atomic level representation of a drug molecule after feature fusion. represents the drug features of the corresponding nodes embedded in the heterogeneous connection graph, || represents the splicing operation, and f c Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the drug internal connection graph, represents element-wise addition, represents element-wise subtraction;

[0027] The heterogeneous connectivity graph embedding is fused with the protein internal connectivity graph embedding as follows:

[0028]

[0029] in, represents the atomic level representation of a protein after feature fusion, f d Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the protein internal connection graph, Represents the protein features of the nodes corresponding to the embedding of the heterogeneous connection graph;

[0030] S42. Atomic-level representation of the same drug molecule and atomic-level representation of proteins Perform average pooling operation and further feature mapping through multi-layer perceptron to extract higher-level features, and finally obtain the node features of the internal connection graph of the drug and protein internal connectivity graph node features

[0031] Preferably, the node features of the internal connection graph of the drug The node characteristics of the protein internal connection graph The drug isomorphic connectivity graph is embedded in and the protein isomorphic connectivity graph embedded in Perform feature fusion based on attention mechanism, including:

[0032] S51. The node features of the internal connection graph of the drug and the drug isomorphic connection graph embedding Perform splicing to obtain the drug splicing feature M D as follows:

[0033]

[0034] S52. The node features of the protein internal connection graph are and the protein isomorphic connectivity graph embedded in Splice and obtain protein splicing feature M T as follows:

[0035]

[0036] S53. Introduce a strategy based on the attention mechanism to combine the drug splicing features M D and the protein splicing feature M T Integrate and get the attention weight W 1 The calculation expression is as follows:

[0037]

[0038] Among them, f a represents the a-th multilayer perceptron;

[0039] S54. Using the attention weight W 1 For the drug splicing feature M D and the protein splicing feature M T Perform the initial weighted summation to obtain the initial weighted summation result F out The calculation expression is as follows:

[0040] F out =W 1 ×M D +(1-W 1 )×M T

[0041] S55. According to the initial weighted summation result F out Update the attention weight and get the new attention weight W 2 The calculation expression is as follows:

[0042] W 2 =f b (F out )

[0043] Among them, f b represents the b-th multilayer perceptron;

[0044] S56. Using the new attention weight W 2 For the drug splicing feature M D and the protein splicing feature M T Perform the final weighted summation to obtain the drug-protein interaction feature F interact The calculation expression is as follows:

[0045] F interact =W 2 ×M D +(1-W 2 )×M T .

[0046] Preferably, the prediction result The calculation expression is as follows:

[0047]

[0048] Among them, f predict Represents a classifier.

[0049] Preferably, the calculation expression of the loss function MSE for training the classifier is as follows:

[0050]

[0051] in, represents the predicted value for the i-th drug-target pair, y represents its true value, and n represents the number of samples in the test set.

[0052] The present invention also proposes a drug target binding affinity prediction system based on a multi-level connection graph, comprising:

[0053] An acquisition module, used for acquiring a drug small molecule, a protein target and an affinity map for representing the interaction relationship between the drug small molecule and the protein target;

[0054] A preprocessing module, used for preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map;

[0055] A multi-level connection graph and feature extraction module is used to extract node features from the multi-level connection graph, and obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding respectively;

[0056] A first feature fusion module is used to fuse the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding to obtain drug internal connection graph node features and protein internal connection graph node features;

[0057] The second feature fusion module is used to perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph and the embedding of the protein isomorphic connection graph to obtain the drug-protein interaction features;

[0058] The prediction module is used to input the drug-protein interaction characteristics into a classifier and output the prediction results of drug-target binding affinity.

[0059] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0060] The present invention proposes a drug-target binding affinity prediction method and system based on a multi-level connection graph. First, drug small molecules, protein targets and affinity graphs are preprocessed to obtain a multi-level connection graph. Then, node features are extracted through a multi-level connection graph framework to obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding, respectively. The purpose is to introduce similarity information between drugs and targets and heterogeneous connection graph information of affinity. Then, the drug-protein interaction features between drugs and targets are learned through a feature fusion mechanism. Finally, the prediction results of drug-target binding affinity are output through a classifier, which effectively improves the performance of drug-target binding affinity prediction in a cold start scenario and obtains more robust prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A flowchart showing a method for predicting drug target binding affinity based on a multi-level connection graph proposed in an embodiment of the present invention;

[0062] Figure 2 Another flowchart of a drug target binding affinity prediction method based on a multi-level connection graph proposed in an embodiment of the present invention is shown;

[0063] Figure 3 2 is a comparison diagram of the ablation experiment proposed in this embodiment;

[0064] Figure 4 The figure shows a structural block diagram of a drug target binding affinity prediction system based on a multi-level connection graph proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0066] It is understandable to those skilled in the art that some well-known contents may be omitted in the drawings;

[0067] In order to facilitate understanding of this embodiment, first, the prior art information of this embodiment is introduced as follows:

[0068] Traditional molecular docking technology that relies on virtual screening usually requires high time and financial costs. However, with the rapid advancement of computer technology in the past few decades, the application of machine learning in predicting drug-target affinity (DTA) has begun to show its potential, providing a more efficient and lower-cost alternative for molecular docking.

[0069] Drug-target affinity prediction is a computational method to evaluate the binding strength between a drug candidate molecule and a biological target (such as protein, nucleic acid, etc.). If the binding strength between a drug molecule and a target is high, this usually means that the drug has a significant affinity for the target, thereby increasing its likelihood of being a candidate drug for treating diseases associated with the target. Many computational methods for this task have been proposed in the past three decades. Pahikkala et al. proposed the Kronecker regularized least squares method (KronRLS), which defines the similarity score of a drug-target pair by the Kronecker product of a similarity matrix. He et al. proposed Simboost, a crossover method that uses a gradient booster to predict drug-target affinity. et al. proposed a deep learning model DeepDTA with two independent convolutional blocks to learn representations from drug SMILES sequences and protein sequences. This was the first time that deep learning was applied to drug-target affinity prediction. After that, the application of deep learning models in this field has become increasingly widespread. Abbasi et al. proposed a deep learning-based method DeepCDA, which combines convolutional layers and long short-term memory (LSTM) layers to effectively encode local and global temporal patterns for deep cross-domain complex protein affinity prediction. GraphDTA proposed by Nguyen et al. was the first to introduce graph neural networks (GNNs) into the drug-target affinity prediction task, in which drug molecular graphs were used as drug representations. However, they still regarded proteins as one-dimensional sequences. On this basis, DGraphDTA proposed by Jiang et al. [9] used the residue contact graph of proteins as the protein molecular graph and utilized a pair of GNNs to process the drug graph and protein graph.

[0070] Although some deep learning strategies have shown excellent predictive effects in drug-target binding affinity prediction, the evaluation of these methods is mostly based on randomly divided data sets. In other words, the drug molecules and target proteins involved in the test data set may have been included in the training data set. The model trained by this data segmentation method cannot fully reflect its true generalization performance when facing new and unseen data. In practical applications, the challenge we face is often to predict targets that are not associated with any known drugs, or when developing new drugs, we do not know which targets they may act on. This kind of situation is usually defined as a "cold-start problem". When the data set is segmented according to this more realistic scenario, the predictive performance of many models tends to drop sharply. In the past five years, some researchers have turned their attention to the cold start problem. FusionDTA proposed by Yuan et al. uses a new multi-head linear attention mechanism to replace the coarse pooling method, which improves the model's ability to capture the interactive information between drugs and targets. At the same time, it adds a knowledge distillation module to alleviate the overfitting problem to a certain extent. HGRL-DTA proposed by Chu et al. establishes a hierarchical graph learning architecture, which effectively integrates the coarse-grained and fine-grained information in the affinity graph, thereby improving the generalization performance of the model. The NHGNN-DTA model proposed by He et al. adaptively captures the feature representation of drugs and proteins, realizes information interaction at the graph level, and cleverly integrates the sequence features of drug molecules and target molecules and the graph-based structural information, improving its performance in the cold start scenario. Although the above methods have made some progress in drug-target affinity prediction in the cold start scenario, their focus is mainly on optimizing the interactive information extraction mechanism between drugs and targets, but ignores the introduction of additional auxiliary information to enhance the generalization ability of the model. By integrating additional auxiliary information, especially chemical sequence similarity information and the features of large-scale pre-trained models, the present invention can significantly improve the performance and prediction accuracy in the cold start scenario without significantly increasing the computational cost.

[0071] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0072] Example 1

[0073] like Figure 1 As shown, this embodiment proposes a drug target binding affinity prediction method based on a multi-level connection graph, comprising the following steps:

[0074] S1. Obtaining a drug small molecule, a protein target, and an affinity map for representing the interaction relationship between the drug small molecule and the protein target;

[0075] The nodes of the affinity graph include drug small molecule nodes and protein target nodes, and the edges of the affinity graph are the binding affinities of the drug small molecule nodes and the protein target nodes.

[0076] S2. Preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map;

[0077] The multi-level connection graph includes a drug internal connection graph G d , drug isomorphism connection graph G D 、Protein internal connection diagram G t 、Protein isomorphic connection graph G T and heterogeneous connection graph G A ; Define drug pre-training feature X d , protein pre-training feature X t .

[0078] The pre-processing of the drug small molecule, the protein target and the affinity map comprises:

[0079] Inputting the drug small molecule into a pre-trained molecular representation learning model, extracting the characteristic representation of the drug small molecule through the molecular representation learning model, and outputting the drug internal connection graph and the drug isomorphic connection graph respectively;

[0080] Inputting the protein target into a pre-trained protein representation deep learning model, extracting feature representation of the protein target through the protein representation deep learning model, and outputting the protein internal connection graph and the protein isomorphic connection graph respectively;

[0081] The binding affinity matrix in the affinity graph is used as the feature of the node in the heterogeneous connection graph to obtain a matrix feature vector, and one-hot encoding is used to distinguish the drug small molecule nodes and the protein target nodes in the affinity graph to obtain a one-hot encoding vector, and the matrix feature vector and the one-hot encoding vector are concatenated to obtain the initial features of the heterogeneous connection graph.

[0082] S3. Extract node features from the multi-level connection graph to obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding;

[0083] The step of extracting node features from the multi-level connection graph includes:

[0084] The drug isomorphic connection graph G D and protein isomorphic connection graph G T Sparse, drug isomorphic connection graph G DIt is derived from the similarity matrix of atoms in drug molecules, and the weight of the edge is derived from the similarity between nodes, ranging from 0 to 1. We use 0.5 as the threshold and only retain the edges with weight greater than the threshold;

[0085] We also do the same for protein isomorphic graphs. After that, we use graph attention network and multi-layer perceptron to get the embedding of isomorphic graphs. and More specifically, the sparsely connected drug isomorphic graph G D and protein isomorphic connection graph G T Input a preset graph attention network, which outputs an initial drug isomorphic connection graph embedding and a protein isomorphic connection graph embedding, and inputs the initial drug isomorphic connection graph embedding into a multi-layer perceptron, which outputs a final drug isomorphic connection graph embedding. The initial protein isomorphic connection graph is embedded into the multi-layer perceptron, and the multi-layer perceptron outputs the final protein isomorphic connection graph embedding

[0086] The internal connection diagram of the drug G d and protein internal connectivity diagram G t The graph attention network is used to output the initial drug internal connection graph embedding. and the initial protein internal connectivity graph embedding

[0087] The heterogeneous connection graph is input into a preset graph convolutional neural network, and the preset graph convolutional neural network outputs the heterogeneous connection graph and embeds it into M A After that, we use a feature fusion mechanism to leverage the learned heterogeneous connectivity graph node representations to further guide and enhance the representations of individual atom or residue nodes;

[0088] S4. fusing the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding respectively to obtain drug internal connection graph node features and protein internal connection graph node features;

[0089] The embedding of the heterogeneous connection graph is respectively fused with the embedding of the internal connection graph of the drug and the embedding of the internal connection graph of the protein, comprising:

[0090] S41. embedding the heterogeneous connection graph into the initial drug internal connection graph d1 The fusion is performed as follows:

[0091]

[0092] in, It represents the atomic level representation of a drug molecule after feature fusion. represents the drug features of the corresponding nodes embedded in the heterogeneous connection graph, || represents the splicing operation, and f c Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the internal connection graph of the drug, represents element-wise addition, represents element-wise subtraction;

[0093] The heterogeneous connectivity graph embedding is fused with the protein internal connectivity graph embedding as follows:

[0094]

[0095] in, represents the atomic level representation of a protein after feature fusion, f d Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the protein internal connection graph, The protein features of the corresponding nodes of the heterogeneous connection graph are embedded; for the fusion of protein residue features, we use the same method. Through this design, each atomic-level node representation can absorb its corresponding coarse-grained level information, and the vector addition and subtraction operations also ensure the adaptive balance between coarse-grained and fine-grained representations.

[0096] S42. Atomic-level representation of the same drug molecule and atomic-level representation of proteins Perform average pooling operation and further feature mapping through multi-layer perceptron to extract higher-level features, and finally obtain the node features of the internal connection graph of the drug and protein internal connectivity graph node features

[0097] It is worth mentioning that in the prediction stage of the model, we cannot directly obtain the affinity information of the drug-protein pair to be predicted. Therefore, we select the K drugs or proteins that are closest to it as references based on the chemical sequence similarity (the specific K value is determined by the total number of reference drugs and proteins). Then, we process the representations of these reference drugs or proteins with a multi-layer perceptron and perform an average pooling operation on their output features to generate heterogeneous graph node features.

[0098] S5. Perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph, and the embedding of the protein isomorphic connection graph to obtain drug-protein interaction features;

[0099] The node features of the internal connection graph of the drug The node characteristics of the protein internal connection graph The drug isomorphic connectivity graph is embedded in and the protein isomorphic connectivity graph embedded in Perform feature fusion based on attention mechanism, including:

[0100] S51. The node features of the internal connection graph of the drug and the drug isomorphic connection graph embedding Perform splicing to obtain the drug splicing feature M D as follows:

[0101]

[0102] S52. The node features of the protein internal connection graph are and the protein isomorphic connectivity graph embedded in Splice and obtain protein splicing feature M T as follows:

[0103]

[0104] In order to deeply capture the interaction information between drugs and target proteins, we did not adopt the method of directly splicing drug and protein features commonly used in most existing algorithms, but instead introduced a strategy based on an attention mechanism in S53 to integrate the feature representations of drugs and proteins.

[0105] S53. Introduce a strategy based on the attention mechanism to combine the drug splicing features M D and the protein splicing feature M T Integrate and get the attention weight W 1 The calculation expression is as follows:

[0106]

[0107] Among them, f a represents the a-th multilayer perceptron;

[0108] S54. Using the attention weight W 1 For the drug splicing feature M D and the protein splicing feature M T Perform the initial weighted summation to obtain the initial weighted summation result F out The calculation expression is as follows:

[0109] F out =W 1 ×M D +(1-W 1 )×MT (6)

[0110] S55. In order to learn deeper interactive information, we repeat the above operation and calculate the initial weighted summation result F out Update the attention weight and get the new attention weight W 2 The calculation expression is as follows:

[0111] W 2 =f b (F out ) (7)

[0112] Among them, f b represents the b-th multilayer perceptron;

[0113] S56. Using the new attention weight W 2 For the drug splicing feature M D and the protein splicing feature M T Perform the final weighted summation to obtain the drug-protein interaction feature F interact The calculation expression is as follows:

[0114] F interact =W 2 ×M D +(1-W 2 )×M T (8)

[0115] S6. Input the drug-protein interaction features into a classifier and output a prediction result of drug-target binding affinity.

[0116] The prediction results The calculation expression is as follows:

[0117]

[0118] Among them, f predict Represents a classifier.

[0119] The calculation expression of the loss function MSE for training the classifier is as follows:

[0120]

[0121] in, represents the predicted value for the i-th drug-target pair, y represents its true value, and n represents the number of samples in the test set.

[0122] We evaluate the performance of our model in predicting drug-target binding affinity on the Davis dataset and the Kiba dataset. The Davis dataset focuses on the interaction between kinases and inhibitors, contains 442 proteins and 68 ligands, and provides 30,056 binding affinity values; the Kiba dataset combines multiple kinase inhibitor bioactivity data, contains 229 proteins and 2,111 ligands, and provides 118,254 binding affinity values. Our experiment splits the training set, validation set, and test set in a ratio of 8:1:1. The cold start scenario can be divided into a cold drug scenario (the drug in the test set has not appeared in the training set), a cold target scenario (the target in the test set has not appeared in the training set), and a fully cold scenario (neither the drug nor the target in the test set has appeared in the training set). For the cold drug scenario, the drugs in the Davis dataset are divided into three non-overlapping groups, with 54, 7, and 7 drugs, respectively, which are used for the training set, validation set, and test set; the drugs in the Kiba dataset are divided into 1654, 207, and 207 non-overlapping drugs. In the cold target scenario, the Davis dataset is divided into 354, 44, and 44 different protein targets, while the Kiba dataset is divided into 182, 23, and 23 different protein targets. For the full cold scenario, we will combine the cold drug and cold target scenarios to allocate the dataset. At the same time, in order to ensure the reliability and consistency of the experimental results, we used different random seeds to cut the dataset five times, and calculated and reported the average results of these five experiments. With such a dataset and experimental design, we can effectively evaluate the model's ability to predict drug target binding affinity in the cold start scenario.

[0123] The principles of the above steps are analyzed and summarized here. A drug target binding affinity prediction method based on a multi-level connection graph proposed in this embodiment mainly includes:

[0124] Pre-trained models and feature initialization: In order to enhance the ability to extract feature representations of unknown drug molecules, we used a geometrically enhanced molecular representation learning model GEM. The molecular representation learning model GEM was pre-trained on a large number of organic small molecule compound libraries. For protein targets, we used the Transformer-based protein representation deep learning model ESM-1b. The protein representation deep learning model ESM-1b was pre-trained on a large-scale sequence database. We used these two pre-trained models to generate the initial features of drugs and protein targets in the internal connection graph and the isomorphic connection graph. For the heterogeneous connection graph constructed from the affinity graph, we used the affinity matrix as the feature of the node in the heterogeneous graph, and used one-hot encoding to distinguish between drug and protein nodes, and finally spliced ​​these two features together as the initial feature.

[0125] Multi-level connection graph and feature extraction: We introduce internal connection graph, isomorphic connection graph and heterogeneous connection graph to learn the feature information of drugs and targets at different levels respectively. At the same time, we use graph attention network to extract node feature information in internal connection graph and isomorphic connection graph, and use graph convolutional neural network to extract node feature information in heterogeneous connection graph.

[0126] Feature fusion: The node feature information extracted from the heterogeneous connection graph will be fused with the node feature information learned in the internal connection graph to guide the learning of node features in the internal connection graph; at the same time, we will use an attention-based fusion mechanism to fuse the final learned drug features and protein target features.

[0127] Result prediction: The fused feature vector is input into the classifier through a multi-layer perceptron for prediction.

[0128] Currently, most drug-target binding affinity prediction algorithms are not designed to fully consider the introduction of additional auxiliary information, which has potential value for improving the performance of the algorithm in cold start scenarios. In this patent, we use pre-trained models to generate initial feature representations. At the same time, through a multi-level connection graph framework, we introduce similarity information between drugs and targets and affinity heterogeneous graph information, and through a reasonable feature fusion mechanism, learn the key interaction information between drugs and targets. Compared with existing technologies, this patented technology can improve the performance of drug-target protein binding affinity prediction in cold start scenarios and obtain more robust prediction results.

[0129] Example 2

[0130] This example further explains the principles and steps of a method for predicting drug target binding affinity based on a multi-level connection graph proposed in the above example.

[0131] A. Explanation of symbols in the algorithm:

[0132] Drug internal connection diagram using G d =(V d , E d ) indicates that V d represents the node set in the internal connection graph of the drug, E d represents the edge set in the drug internal connection graph; the protein internal connection graph constructed by k nearest neighbors is represented by G t =(V t , E t ) indicates that V t represents the node set in the internal connection graph of the drug, E t represents the edge set in the internal connection graph of the drug; the drug isomorphic connection graph is represented by G D =(V D ,E D) indicates that V D represents the node set in the drug isomorphic connection graph, E D represents the edge set in the drug isomorphic connection graph; the protein molecule graph is represented by G T =(V T , E T ) indicates that V T represents the node set in the drug isomorphic connection graph, E T represents the edge set in the drug isomorphic connection graph;

[0133] Define drug pre-training features X d , protein pre-training feature X t , where X d ∈R m×p , X t ∈R n×q , m and n are the number of atoms or residues in a drug or protein molecule, and p and q are the embedding dimensions of its pre-trained features;

[0134] Drug internal connection diagram embedding The initial protein internal connection graph is embedded with Indicates that the node features of the internal connection graph of the drug are represented by The node features of the protein internal connection graph are represented by Denotes, where |D| and |T| represent the number of drugs and proteins in the training set, respectively, and h represents the dimension of drug and protein node embedding;

[0135] Drug isomorphic connection graph embedding Represented by, the protein isomorphic connection graph is embedded with Representation; heterogeneous connection graph embedding is M A ∈R (|D|+|T|)×(|D|+|T|+2) express.

[0136] B. Multi-level connection graph: The above-mentioned embodiment method includes three levels of connection graphs: internal connection graph (including drug internal connection graph and protein internal connection graph), isomorphic connection graph (drug isomorphic connection graph and protein isomorphic connection graph) and heterogeneous connection graph.

[0137] The internal connection map refers to the connection pattern between atoms or residues inside the drug molecule or protein target, which reflects the intrinsic structural characteristics of the drug or target protein;

[0138] Isomorphic connectivity graph refers to a connectivity graph constructed between drug molecules or protein targets based on additional chemical sequence similarity information. This graph can reflect the degree of similarity in chemical structure between multiple drugs or target proteins.

[0139] A heterogeneous connection graph refers to a graph structure between drugs and target proteins based on the known affinity relationships in the training set. By constructing a multi-level connection graph and iteratively learning node features, our model can more accurately capture and understand the complex interaction information between drugs and targets.

[0140] C. Feature Extraction: In order to better capture the complex relationships and dependencies between nodes in the internal connection graph and the isomorphic connection graph, we use a graph attention network to extract node embeddings in the internal connection graph and the isomorphic connection graph;

[0141] For heterogeneous connection graphs, we mainly focus on the connection relationship between the two nodes, so we use a more computationally efficient graph convolutional neural network to extract node embeddings.

[0142] D. Pooling: Figure 2 As shown in Figure 3, we use a pooling layer after the internal connection graph to perform aggregation operations on node features, while no pooling operation is required in the homogeneous connection graph and the heterogeneous connection graph.

[0143] This is because each node in the internal connection graph represents an atom or residue in a drug or protein, and only after pooling can the embedding of the entire drug or protein be represented. For homogeneous or heterogeneous connection graphs, each node represents a complete drug or protein, so no pooling operation is required.

[0144] E. Feature fusion mechanism: This patent uses two feature fusion mechanisms. The first feature fusion mechanism aims to use the graph-level node representation of drugs or proteins learned from heterogeneous connection graphs to guide the learning of atomic-level node representations in internal connection graphs. We use formulas 1 and 2 to fuse these two node features; while the second feature fusion mechanism aims to better capture the interaction information between drugs and proteins. We use an attention-based fusion mechanism to fuse the learned final drug representation and protein final representation, as shown in formulas 5 to 8.

[0145] F. Loss function setting: The drug-target binding affinity prediction task is usually regarded as a regression problem. Therefore, the mean square error, which is widely used in regression analysis, is adopted as the loss function to guide the optimization process of the classifier parameters.

[0146] Example 3

[0147] This example further evaluates the drug target binding affinity prediction method based on a multi-level connection graph proposed in the above example.

[0148] Data collection: This example conducts experiments on Davis and Kiba datasets. The data set division follows the method described above. At the same time, we use the GEM and ESM-1b models to obtain the pre-trained features X of drug atoms, respectively. d ∈R m×p and the pre-trained features X of protein residues t ∈R n×p , where m represents the number of atoms contained in a drug, n represents the number of residues contained in a protein, and p and q represent the embedding dimensions of their pre-trained features.

[0149] Build the model: Place X d and X t As the input features of the drug internal connection map and protein internal connection map, X d and X t Perform average pooling to get X D ∈R |D|×p and X T ∈R |T|×q As the input features of the drug isomorphic connection graph and protein isomorphic connection graph respectively. At the same time, the input feature X of the heterogeneous connection graph is obtained by splicing affinity information and the unique hot encoding that distinguishes drug or protein nodes A ∈R (|D|+|T|)×(|D|+|T|+2) . Then build the model according to the steps described above.

[0150] Select verification methods and evaluation indicators: According to the data set cutting method mentioned above, we used different random seeds to cut the data set five times, calculated and reported the average results of these five experiments. We mainly use mean square error (MSE) to evaluate the performance of the model, and also use consistency index (CI), Pearson correlation coefficient (Pearson) and modified R square (Rm2) as evaluation indicators.

[0151] Parameter setting: The node embedding dimension h of drugs and proteins is set to 128. The number of neighbors K of drugs is set to 2, and the number of neighbors K of proteins is set to 5. During model training, we use the Adam optimizer with a learning rate of 0.0005 and 500 iterations.

[0152] In order to verify the effectiveness of the present invention, the present invention set up an ablation experiment in the cold drug scenario of the Davis dataset, and removed the isomorphic connection graph and the heterogeneous connection graph in the present invention and then compared them. At the same time, in order to verify the effectiveness of the pre-trained features, we replaced the initial pre-trained features with more commonly used features based on chemical sequence information and structural information, and conducted comparative experiments. The experimental results are shown in Figure 2. Figure 3This indicates that the use of the homogeneous connection graph, heterogeneous connection graph and pre-training features in the present invention improves the performance of the drug target binding affinity prediction model to a certain extent.

[0153] Example 4

[0154] See also Figure 4 , this embodiment proposes a drug target binding affinity prediction system based on a multi-level connection graph, including:

[0155] An acquisition module, used for acquiring a drug small molecule, a protein target and an affinity map for representing the interaction relationship between the drug small molecule and the protein target;

[0156] A preprocessing module, used for preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map;

[0157] A multi-level connection graph and feature extraction module is used to extract node features from the multi-level connection graph, and obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding respectively;

[0158] A first feature fusion module is used to fuse the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding to obtain drug internal connection graph node features and protein internal connection graph node features;

[0159] The second feature fusion module is used to perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph and the embedding of the protein isomorphic connection graph to obtain the drug-protein interaction features;

[0160] The prediction module is used to input the drug-protein interaction characteristics into a classifier and output the prediction results of drug-target binding affinity.

[0161] In this embodiment, drug small molecules, protein targets and affinity graphs are first preprocessed to obtain a multi-level connection graph, and then node features are extracted through the multi-level connection graph framework to obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding, respectively. The purpose is to introduce the similarity information between drugs and targets and the heterogeneous connection graph information of affinity, and then learn the drug-protein interaction characteristics between drugs and targets through a feature fusion mechanism. Finally, the prediction results of drug-target binding affinity are output through a classifier, which effectively improves the performance of drug-target binding affinity prediction in cold start scenarios and obtains more robust prediction results.

[0162] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not intended to limit the implementation methods of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation methods here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A drug target binding affinity prediction method based on a multi-level connection graph, characterized in that: The following steps are involved: S1. Obtaining a drug small molecule, a protein target, and an affinity map for representing the interaction relationship between the drug small molecule and the protein target; S2. Preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map; S3. Extract node features from the multi-level connection graph to obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding; S4. fusing the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding respectively to obtain drug internal connection graph node features and protein internal connection graph node features; S5. Perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph, and the embedding of the protein isomorphic connection graph to obtain drug-protein interaction features; S6. Input the drug-protein interaction features into a classifier and output a prediction result of drug-target binding affinity.

2. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 1, characterized in that: The nodes of the affinity graph include drug small molecule nodes and protein target nodes, and the edges of the affinity graph are the binding affinities of the drug small molecule nodes and the protein target nodes.

3. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 2, characterized in that: The multi-level connection diagram includes a drug internal connection diagram, a drug isomorphic connection diagram, a protein internal connection diagram, a protein isomorphic connection diagram and a heterogeneous connection diagram.

4. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 3, characterized in that: The pre-processing of the drug small molecule, the protein target and the affinity map comprises: Inputting the drug small molecule into a pre-trained molecular representation learning model, extracting the characteristic representation of the drug small molecule through the molecular representation learning model, and outputting the drug internal connection graph and the drug isomorphic connection graph respectively; Inputting the protein target into a pre-trained protein representation deep learning model, extracting feature representation of the protein target through the protein representation deep learning model, and outputting the protein internal connection graph and the protein isomorphic connection graph respectively; The binding affinity matrix in the affinity graph is used as the feature of the node in the heterogeneous connection graph to obtain a matrix feature vector, and one-hot encoding is used to distinguish the drug small molecule nodes and the protein target nodes in the affinity graph to obtain a one-hot encoding vector, and the matrix feature vector and the one-hot encoding vector are concatenated to obtain the initial features of the heterogeneous connection graph.

5. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 4, characterized in that: The step of extracting node features from the multi-level connection graph includes: The drug isomorphic connection graph G D and protein isomorphic connection graph G T Sparse, the sparse drug isomorphic connection graph G D and protein isomorphic connection graph G T Input a preset graph attention network, which outputs an initial drug isomorphic connection graph embedding and a protein isomorphic connection graph embedding, and inputs the initial drug isomorphic connection graph embedding into a multi-layer perceptron, which outputs a final drug isomorphic connection graph embedding. The initial protein isomorphic connection graph is embedded into the multi-layer perceptron, and the multi-layer perceptron outputs the final protein isomorphic connection graph embedding The internal connection diagram of the drug G d and protein internal connectivity diagram G t The graph attention network is used to output the initial drug internal connection graph embedding. and the initial protein internal connectivity graph embedding The heterogeneous connection graph is input into a preset graph convolutional neural network, and the preset graph convolutional neural network outputs the heterogeneous connection graph and embeds it into M A .

6. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 5, characterized in that: The embedding of the heterogeneous connection graph is respectively fused with the embedding of the internal connection graph of the drug and the embedding of the internal connection graph of the protein, comprising: S41. Embedding the heterogeneous connection graph into the initial drug internal connection graph The fusion is performed as follows: in, It represents the atomic level representation of a drug molecule after feature fusion. represents the drug features of the corresponding nodes embedded in the heterogeneous connection graph, || represents the splicing operation, and f c Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the drug internal connection graph, represents element-wise addition, represents element-wise subtraction; The heterogeneous connectivity graph embedding is fused with the protein internal connectivity graph embedding as follows: in, represents the atomic level representation of a protein after feature fusion, f d Represent a two-layer multilayer perceptron to coordinate the representation of nodes in the heterogeneous connection graph and the representation of nodes in the protein internal connection graph, Represents the protein features of the nodes corresponding to the embedding of the heterogeneous connection graph; S42. Atomic-level representation of the same drug molecule and atomic-level representation of proteins Perform average pooling operation and further feature mapping through multi-layer perceptron to extract higher-level features, and finally obtain the node features of the internal connection graph of the drug and protein internal connectivity graph node features 7. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 6, characterized in that: The node features of the internal connection graph of the drug The node characteristics of the protein internal connection graph The drug isomorphic connectivity graph is embedded in and the protein isomorphic connectivity graph embedded in Perform feature fusion based on attention mechanism, including: S51. The node features of the internal connection graph of the drug and the drug isomorphic connection graph embedding Perform splicing to obtain the drug splicing feature M D as follows: S52. The node features of the protein internal connection graph are and the protein isomorphic connectivity graph embedded in Splice and obtain protein splicing feature M T as follows: S53. Introduce a strategy based on the attention mechanism to combine the drug splicing features M D and the protein splicing feature M T After integration, the calculation expression of attention weight W1 is as follows: Among them, f a represents the a-th multilayer perceptron; S54. Using the attention weight W1 to the drug splicing feature M D and the protein splicing feature M T Perform the initial weighted summation to obtain the initial weighted summation result F out The calculation expression is as follows: F out =W1×M D +(1-W1)×M T S55. According to the initial weighted summation result F out Update the attention weight and get the calculation expression of the new attention weight W2 as follows: W2=f b (F out ) Among them, f b represents the b-th multilayer perceptron; S56. Using the new attention weight W2 to the drug splicing feature M D and the protein splicing feature M T Perform the final weighted summation to obtain the drug-protein interaction feature F interact The calculation expression is as follows: F interact =W2×M D +(1-W2)×M T 。。 8. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 7, characterized in that: The prediction results The calculation expression is as follows: Among them, f predict Represents a classifier.

9. The method for predicting drug target binding affinity based on a multi-level connection graph according to claim 7, characterized in that: The calculation expression of the loss function MSE for training the classifier is as follows: in, represents the predicted value for the i-th drug-target pair, y represents its true value, and n represents the number of samples in the test set.

10. A drug target binding affinity prediction system based on a multi-level connection graph, characterized in that: include: An acquisition module, used for acquiring a drug small molecule, a protein target and an affinity map for representing the interaction relationship between the drug small molecule and the protein target; A preprocessing module, used for preprocessing the drug small molecule, the protein target and the affinity map to obtain a multi-level connection map; A multi-level connection graph and feature extraction module is used to extract node features from the multi-level connection graph, and obtain drug isomorphic connection graph embedding, drug internal connection graph embedding, protein isomorphic connection graph embedding, protein internal connection graph embedding and heterogeneous connection graph embedding respectively; A first feature fusion module is used to fuse the heterogeneous connection graph embedding with the drug internal connection graph embedding and the protein internal connection graph embedding to obtain drug internal connection graph node features and protein internal connection graph node features; The second feature fusion module is used to perform feature fusion based on the attention mechanism on the node features of the drug internal connection graph, the node features of the protein internal connection graph, the embedding of the drug isomorphic connection graph and the embedding of the protein isomorphic connection graph to obtain the drug-protein interaction features; The prediction module is used to input the drug-protein interaction characteristics into a classifier and output the prediction results of drug-target binding affinity.

Citation Information

Cited By

  • Method and system for predicting protein-ligand binding affinity

    CN121306235A