Deep learning prediction method for drug-protein interaction

Through deep learning prediction methods, the high-dimensional structural similarity of drugs and proteins is calculated and high-dimensional information is introduced using contrast learning methods, which solves the problem of limited information in high-dimensional sparse data and low-dimensional data in drug-protein interaction prediction, and improves the prediction accuracy and applicability.

CN120148604APending Publication Date: 2025-06-13CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510147086.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has problems in drug-protein interaction prediction that high-dimensional sparse data is difficult to model and limited low-dimensional data information lead to insufficient model performance.

Method used

Deep learning prediction methods are used to calculate the similarity of drugs and proteins in high-dimensional structures and to use contrast learning methods in the similarity space to implicitly introduce high-dimensional structural information into features.

Benefits of technology

It improves the accuracy of drug-protein interaction prediction, overcomes the problem of difficult modeling of high-dimensional sparse data and limited information of low-dimensional data, and shows stronger applicability and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148604A_ABST
    Figure CN120148604A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning prediction method for drug-protein interaction, and the method comprises the steps: S1, collecting drug and protein data, and calculating the similarity of drugs and the similarity of proteins; s2, extracting feature vectors of proteins and drugs in each batch; s3, calculating the contrast loss in each batch based on the similarity of the drugs, the similarity of the proteins and the feature vectors of the proteins and the drugs in each batch; s4, splicing the feature vectors of the proteins and the drugs in each batch, and calculating the cross entropy loss in each batch according to the splicing result to obtain the final loss; and S5, performing model training according to the final loss, and inputting protein codes and drug codes into a trained model to obtain a prediction result. The high-dimensional information is introduced into the features while direct modeling of sparse high-dimensional data is avoided, and the accuracy of drug-protein interaction prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence-assisted drug discovery, and particularly relates to a deep learning prediction method for drug-protein interaction. Background Art

[0002] Drug-Target Interactions (DTI) refer to the process in which a drug binds to specific protein target molecules such as enzymes, nuclear receptors, G protein-coupled receptors, and ion channels in the body, changes the function of the protein target, and thus achieves the treatment of diseases and the change of the body state. Therefore, DTI prediction is a very important step in the drug discovery process. Traditional chemical experimental screening methods have high determination costs and long time consumption. On the contrary, computer-aided drug screening is more efficient and can help drug R & D personnel narrow the experimental scope and accelerate drug R & D.

[0003] Traditional DTI prediction methods include docking-based methods and ligand-based methods. These methods have achieved some achievements in drug discovery, but have great limitations. Docking-based methods cannot perform virtual screening on proteins lacking three-dimensional structure data. Ligand-based methods compare candidate ligands with known ligands of the target protein and may perform poorly for target proteins with a small number of known ligands. Due to the lack of three-dimensional structure data and ligand data, a strategy called chemogenomics has been used for DTI prediction, which realizes the prediction of DTI by integrating chemical space and genomics information. Subsequently, a large number of machine learning algorithms based on this strategy have been proposed. However, machine learning-based DTI prediction methods require manual feature setting, which requires expert experience and may also lose important information.

[0004] In recent years, due to its end-to-end characteristics, deep learning has become the mainstream method in this field. DTI prediction methods based on deep learning can be divided into two categories: DTI prediction based on network data mining and DTI prediction based on structural data mining. The DTI prediction method based on network data mining is based on an empirical rule: "Drugs with similar structures are likely to have the same targets, and proteins with similar structures or functions are likely to have the same ligands". This method first needs to analyze and construct complex networks of various biological entities such as drugs, proteins, and diseases, and then uses graph neural networks such as GCN and GAT to mine the topological information between nodes and model the drug-target interaction relationship. For example, Peng et al. used a three-layer GCN to extract the features of proteins and drugs in a drug-protein-disease-side effect heterogeneous graph network, and used the inner product to calculate the probability of the binding of proteins and drug molecules. The method they proposed has good performance on each dataset. However, the network-based method cannot reasonably model isolated nodes lacking topological structure information and can only perform DTI prediction on proteins and drug molecules in the network. Moreover, the similarity principle relied on by the network-based method is an empirical rule, and its prediction results are likely to have false positive problems.

[0005] The DTI prediction method based on structural data mining is based on the theory of "structure determines function" and models DTI by mining the structural information of drugs and proteins. Obviously, the theoretical basis relied on by this method is more rigorous, and this method is free from the limitations of the network and can generalize to unknown data. This method usually can be divided into three stages: data representation, feature extraction, and DTI prediction. Initially, protein descriptors and drug molecule fingerprints were used to represent proteins and drug molecules respectively. However, these manually designed features often lose some information and cannot perfectly represent proteins and drugs.

[0006] Although previous work has greatly improved the DTI prediction results, most of the work is the innovation of the model architecture, especially for the interpretability of the model. They ignored such a problem: low-dimensional data such as amino acid sequences, SMILES strings of drug molecules, and topological graphs provide limited information, and using such data will lose many key high-dimensional structural information, which is likely to affect the performance of the model. Taking drug molecules as an example, such as Figure 1As shown, both molecule A and molecule B are active against the protein vascular endothelial growth factor receptor 2 (VEGFR2), which indicates that although molecule A and molecule B have significant differences in topological structure, they have certain similarities in three-dimensional space. On the contrary, molecule C is not active against VEGFR2, which means that even though molecule C and molecule B are very similar in topological structure, their conformations in three-dimensional space are quite different. Therefore, it is considered that in the process of drug-protein interaction modeling, introducing some high-dimensional structure information using appropriate methods will help improve the prediction effect of the model. Summary of the Invention

[0007] Aiming at the above deficiencies in the prior art, a deep learning prediction method for drug-protein interaction provided by the present invention solves the problems of difficult modeling of high-dimensional sparse data and insufficient modeling performance of low-dimensional data with limited information.

[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a deep learning prediction method for drug-protein interaction, comprising the following steps:

[0009] S1. Collect drug and protein data, and calculate the similarity of drugs and the similarity of proteins;

[0010] S2. Encode the drug and protein data, and extract the feature vectors of proteins and drugs in each batch;

[0011] S3. Based on the similarity of drugs, the similarity of proteins, and the feature vectors of proteins and drugs in each batch, calculate the contrastive loss in each batch;

[0012] S4. Concatenate the feature vectors of proteins and drugs in each batch, calculate the cross-entropy loss in each batch according to the concatenated result, and obtain the final loss according to the cross-entropy loss and the contrastive loss;

[0013] S5. Train the model according to the final loss, input the protein encoding and drug encoding into the trained model, and obtain the prediction result.

[0014] Further: in the above S1, the similarity of drugs is specifically: the Tanimoto similarity of drug molecules on the set fingerprint, where the set fingerprint includes the substructure fingerprint, topological structure fingerprint, and pharmacophore fingerprint of drug molecules;

[0015] The similarity of the proteins includes the sequence similarity of the proteins , domain similarity and molecular function similarity , and its expression is specifically the following formula:

[0016]

[0017]

[0018]

[0019] Wherein, are the sequence information of proteins i and j respectively, are the domain information of proteins i and j respectively, and are the molecular function annotation sets of proteins i and j respectively, is the similarity between the molecular function annotation t and the molecular function annotation set MF(j), is the similarity between the molecular function annotation o and the molecular function annotation set MF(i), is the molecular function annotation set is the number of elements of, is the molecular function annotation set is the number of elements of.

[0020] Furthermore: The S2 includes the following sub-steps:

[0021] S21. Encode the amino acid sequence of the protein and the SMLES data of the drug molecule to obtain a protein encoding and a drug encoding;

[0022] S22. Input the protein encoding into the protein CNN module to obtain the initial features of the protein, and input the drug encoding into the drug CNN module to obtain the initial features of the drug;

[0023] S23. Input the initial features of the protein and the initial features of the drug into the cross-attention mechanism module to obtain a protein representation and a drug molecule representation;

[0024] S24. Input the protein representation into the first max pooling layer to obtain the feature vector of the protein within each batch, and input the drug molecule representation into the second max pooling layer to obtain the feature vector of the drug within each batch.

[0025] Furthermore: In the S22, the structures of the protein CNN module and the drug CNN module are the same, and both include the first to third layers of one-dimensional convolution connected in sequence;

[0026] The method for obtaining the initial features of the protein is specifically:

[0027] Search for the protein encoding to obtain the protein embedding tensors within each batch, input them into the first to third layers of one-dimensional convolution in sequence, and use the protein hidden layer representation extracted by the third layer of one-dimensional convolution as the initial features of the protein;

[0028] Among them, the expression of the protein hidden layer representation extracted by each layer of one-dimensional convolution is as follows:

[0029]

[0030] In the formula, the ordinal number l ∈ {0, 1, 2}, is the protein embedding tensor within a batch, is the protein hidden layer representation extracted by the first layer of one-dimensional convolution, is the protein hidden layer representation extracted by the second layer of one-dimensional convolution, is the protein hidden layer representation extracted by the third layer of one-dimensional convolution, is the weight matrix in the CNN convolution kernel of the (l + 1)-th layer of one-dimensional convolution in the protein CNN module, is the bias vector in the CNN convolution kernel of the (l + 1)-th layer of one-dimensional convolution in the protein CNN module, is the one-dimensional convolutional neural network operation, is the non-linear activation function;

[0031] The method for obtaining the initial features of the drug is specifically as follows:

[0032] Search for the drug code to obtain the drug embedding tensors within each batch, input them into the first to third layers of one-dimensional convolution in sequence, and use the drug hidden layer representation extracted by the third layer of one-dimensional convolution as the initial features of the drug;

[0033] Among them, the expression of the drug hidden layer representation extracted by each layer of one-dimensional convolution is as follows:

[0034]

[0035] In the formula, is the drug embedding tensor within a batch, is the drug hidden layer representation extracted by the first layer of one-dimensional convolution, is the drug hidden layer representation extracted by the second layer of one-dimensional convolution, is the drug hidden layer representation extracted by the third layer of one-dimensional convolution, is the weight matrix in the CNN convolution kernel of the (l + 1)-th layer of one-dimensional convolution in the drug CNN module, is the bias vector in the CNN convolution kernel of the (l + 1)-th layer of one-dimensional convolution in the drug CNN module.

[0036] Furthermore: in S23, the expression for obtaining the protein representation is specifically as follows:

[0037]

[0038] In the formula, is, and are both similar space embedding functions, which are composed of two different linear functions and Relu activation functions. ( ) is the normalized exponential function. is the query matrix of the current batch of proteins in the protein-drug similarity space. is the key matrix of the drugs in the batch in the protein-drug similarity space. is the value matrix of the current batch of drugs in the similarity space. is the dimension of, and T represents the matrix transpose operation;

[0039] Obtain the drug molecule representation The specific expression of is as follows:

[0040] .

[0041] Furthermore: In S3, calculating the contrast loss within any batch includes the following sub-steps:

[0042] S31. Calculate the contrast loss of proteins within a batch according to the feature vectors and similarities of proteins in the batch by the contrast learning method;

[0043] S32. Calculate the contrast loss of drugs within a batch according to the feature vectors and similarities of drugs in the batch by the contrast learning method;

[0044] S33. Obtain the contrast loss within any batch according to the contrast loss of proteins and the contrast loss of drugs within a batch.

[0045] Furthermore: In S31, the specific contrast learning method is as follows:

[0046] S311. Map the feature vectors of proteins within a batch to the molecular function contrast space, domain annotation contrast space, and sequence annotation contrast space through 3 independent modules with two fully connected layers;

[0047] S312. Construct positive and negative pairs in the contrast space according to the similarities of proteins, and calculate the contrast loss of proteins in the three contrast spaces;

[0048] S313. Obtain the contrast loss of proteins within a batch according to the contrast losses of proteins in the three contrast spaces.

[0049] Furthermore: In S311, the specific expressions for mapping the feature vectors of proteins within a batch to the molecular function contrast space, domain annotation contrast space, and sequence annotation contrast space are as follows:

[0050]

[0051] In the formula, k is the comparison space, is the molecular function comparison space, is the domain annotation comparison space, and seq is the sequence annotation comparison space. are the weights and biases of the first linear layer mapped by the comparison space k, is the bias of the first linear layer mapped by the comparison space k, is the weight of the second linear layer mapped by the comparison space k, is the bias of the second linear layer mapped by the comparison space k, is the hidden feature of the protein in the current batch in the comparison space k, , where N is the number of drug-protein pairs in a batch, and D represents the feature dimension;

[0052] In the above S312, calculate the comparison loss of the protein in the comparison space k The specific expression of is:

[0053]

[0054] In the formula, represents the set of proteins in a batch, is the positive pair of protein i in the current batch in the comparison space k, is the hidden feature of protein i in the comparison space k, and sim is the function for calculating similarity. is the hidden feature of protein j in the comparison space k, is the temperature coefficient in the contrast learning, is the hidden feature of protein z in the comparison space k.

[0055] Furthermore: In the above S4, the cross-entropy loss The specific expression of is:

[0056]

[0057] In the formula, y represents the true label of the drug-protein pair interaction, and p represents the probability of the drug-protein pair interaction predicted by the model.

[0058] Furthermore: In the above S5, the method for model training is specifically:

[0059] With the goal of minimizing the final loss, AdamW is used as the optimizer for gradient descent, the model parameters are updated with a learning rate of 0.0001, the maximum number of training epochs is set to 200, and early stopping technology is used to determine whether the model training is completed, with the early stopping step set to 50 steps.

[0060] The beneficial effects of the present invention are as follows:

[0061] (1) The present invention provides a deep learning prediction method for drug-protein interactions. By calculating the similarities between drugs and between proteins in the high-dimensional structure, and then using the method of contrastive learning in these similarity spaces to implicitly introduce this high-dimensional structure information into the features of proteins and drug molecules. Thus, while avoiding direct modeling of sparse high-dimensional data, this high-dimensional information is introduced into the features, improving the accuracy of drug-protein interaction prediction and solving the problem that it is difficult to model high-dimensional sparse data and the modeling performance is insufficient due to limited information in low-dimensional data.

[0062] (2) In order to verify the effectiveness of the model CLSF-DTI proposed by the present invention, the following five experiments were conducted. The results of these experiments show that the present invention has the following advantages:

[0063] CLSF-DTI has stronger applicability in drug-protein interaction prediction. Different from some methods that can only be used on large-scale datasets, CLSF-DTI can not only be used on large-scale datasets, but also has good performance on small-scale datasets. This is because this method extracts features with CNN, and the parameters of the model are relatively few, and a small amount of data can also complete the optimization of the model.

[0064] Compared with other benchmark methods, CLSF-DTI has better performance in drug-protein interaction prediction. This is because after introducing high-dimensional structure data, the model can more accurately identify the differences between drugs and proteins.

[0065] CLSF-DTI has stronger prediction ability for unknown data than previous methods. This is because after using contrastive learning to introduce high-dimensional structure data, the model can more easily capture the patterns between data, thus obtaining stronger generalization ability.

[0066] CLSF-DTI has the potential for virtual screening of drugs. During the experiment of virtual screening of active drugs for protein PKA-Cα, CLSF-DTI successfully found 22 effective ligands. This benefits from the strong generalization ability brought by the contrastive learning module.

[0067] CLSF-DTI has certain insight capabilities into drug-protein interactions. The cross-attention module can simulate the interaction between proteins and drug molecules. In visualization experiments, it can be found that the model can correctly predict the binding sites of some drugs on proteins. Description of the Drawings

[0068] Figure 1 It is a diagram showing the relationship between the two-dimensional topological structure of the molecule and its VEGFR2 activity.

[0069] Figure 2 It is a flowchart of a deep learning prediction method for drug-protein interactions of the present invention.

[0070] Figure 3 It is a structural diagram of the model CLSF-DTI of the present invention.

[0071] Figure 4 It is a diagram of the ablation experiment results.

[0072] Figure 5 It is a diagram of the visualization result of PDB: 2GU8.

[0073] Figure 6 It is a diagram of the visualization result of PDB: 5N23. Detailed Embodiments

[0074] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0075] As Figure 2 shown, in an embodiment of the present invention, a deep learning prediction method for drug-protein interactions includes the following steps:

[0076] S1. Collect drug and protein data, and calculate the similarity of drugs and the similarity of proteins;

[0077] In the above S1, the drug data includes the SMLES data of drug molecules, and the protein data includes the amino acid sequence, domain data, and molecular function data of proteins;

[0078] S2. Encode the drug and protein data, and extract the feature vectors of proteins and drugs in each batch;

[0079] S3. Calculate the contrastive loss within each batch based on the similarity of drugs, the similarity of proteins, and the feature vectors of proteins and drugs within each batch.

[0080] S4. Concatenate the feature vectors of proteins and drugs within each batch, calculate the cross-entropy loss within each batch based on the concatenated result, and obtain the final loss based on the cross-entropy loss and the contrastive loss.

[0081] S5. Perform model training based on the final loss, input the protein encoding and drug encoding into the trained model, and obtain the prediction result.

[0082] The present invention proposes a method for predicting drug-protein interactions by fusing high-dimensional structural information using contrastive learning. The overall CLSF-DTI model is as Figure 3 shown. For the collected drug and protein data, the present invention uses a CNN module to extract the feature vectors of proteins and drugs within each batch, then uses a cross-attention mechanism module to enhance the representation of proteins and drug molecules, then uses max pooling to obtain the features of proteins and drug molecules, and then uses a Contrastive Module to add various structural similarity information to the features of proteins and drug molecules respectively. Finally, through a fully connected layer, the features of proteins and drug molecules are fused and the prediction probability of DTI is output.

[0083] As Figure 3 shown, in this embodiment, the present invention encodes the protein sequence and drug SMILES string through Encoding and Embeding, extracts features using a CNN module, and enhances the feature representation using a cross-attention mechanism. Then, the extracted features are fed into a contrastive learning module based on structural and functional similarity to learn structural information, and finally, the prediction result is output through a fully connected layer, where ⊕ represents concatenation.

[0084] In the above S1, the similarity of drugs specifically refers to the Tanimoto similarity of drug molecules on the set fingerprints, where the set fingerprints include the substructure fingerprint, topological fingerprint, and pharmacophore fingerprint of drug molecules.

[0085] In this embodiment, the RDkit tool is used to obtain the substructure fingerprint, topological fingerprint, and pharmacophore fingerprint of drug molecules, and then the Tanimoto similarity of each drug on these three fingerprints is calculated as the Tanimoto similarity of the drug in these three cases.

[0086] The similarity of the proteins includes the sequence similarity of proteins , domain similarity and molecular function similarity. , and its expression is specifically as follows:

[0087]

[0088]

[0089]

[0090] In the formula, are the sequence information of proteins i and j respectively, are the domain information of proteins i and j respectively, and are the molecular function annotation sets of proteins i and j respectively, is the similarity between the molecular function annotation t and the molecular function annotation set MF(j), is the similarity between the molecular function annotation o and the molecular function annotation set MF(i), is the molecular function annotation set is the number of elements, is the molecular function annotation set is the number of elements.

[0091] The said S2 includes the following sub-steps:

[0092] S21. Encode the amino acid sequence of the protein and the SMLES data of the drug molecule to obtain the protein encoding and the drug encoding;

[0093] S22. Input the protein encoding into the protein CNN module to obtain the initial features of the protein, and input the drug encoding into the drug CNN module to obtain the initial features of the drug;

[0094] S23. Input the initial features of the protein and the initial features of the drug into the cross-attention mechanism module to obtain the protein representation and the drug molecule representation;

[0095] S24. Input the protein representation into the first max-pooling layer to obtain the feature vector of the protein in each batch, and input the drug molecule representation into the second max-pooling layer to obtain the feature vector of the drug in each batch.

[0096] The said S21 is specifically as follows:

[0097] S211. Set the first length threshold of the protein, intercept the amino acid sequence of the protein exceeding the first length threshold, and perform digital 0 padding on the amino acid sequence of the protein less than the first length threshold to obtain a learnable first embedding matrix, and use it as the protein encoding;

[0098] S212. Set the length threshold of the drug, intercept the SMLES data of drug molecules exceeding the second length threshold, and perform digital 0 padding on the SMLES data of insufficient drug molecules to obtain a learnable second embedding matrix, which is used as the drug code.

[0099] In this embodiment, numbers are used to encode these characters. For the characters in the amino acid sequence string, its digital encoding range is from 1 to 24, and the digital encoding range of the SMILES string is from 1 to 64. For the convenience of calculation, the length of the protein is fixed at 1000, and the drug is 100. Molecules exceeding this length are intercepted, and molecules shorter than this length are padded with the number 0.

[0100] According to the idea of word embeding, a learnable embedding tensor is constructed for the protein , and a learnable embedding tensor is constructed for the drug molecule , where and respectively represent the embedding dimensions of amino acid characters and SMILES characters. R represents a matrix. By looking up the corresponding embedding tensor, the proteins within a batch can be represented as the embedding tensor , and the drug is represented as the embedding tensor , where N represents the number of drug-protein pairs within a batch, and respectively represent the fixed lengths of the amino acid sequence and the SMILES string.

[0101] In S22, the structures of the protein CNN module and the drug CNN module are the same, both including the first to third layers of one-dimensional convolution connected in sequence;

[0102] The method for obtaining the initial features of the protein is specifically as follows:

[0103] Look up the protein code to obtain the protein embedding tensors within each batch, and input them into the first to third layers of one-dimensional convolution in sequence. The protein hidden layer representation extracted by the third layer of one-dimensional convolution is used as the initial feature of the protein;

[0104] Among them, the expression of the protein hidden layer representation extracted by each layer of one-dimensional convolution is as follows:

[0105]

[0106] In the formula, the ordinal number l ∈ {0, 1, 2}, is the protein embedding tensor within a batch, is the protein hidden layer representation extracted by the first layer of one-dimensional convolution, The protein hidden layer representation extracted by the second - layer one - dimensional convolution The protein hidden layer representation extracted by the third - layer one - dimensional convolution The weight matrix in the CNN convolution kernel of the (l + 1)-th layer one - dimensional convolution of the protein CNN module The bias vector in the CNN convolution kernel of the (l + 1)-th layer one - dimensional convolution of the protein CNN module The one - dimensional convolutional neural network operation The non - linear activation function

[0107] The method for obtaining the initial features of the drug is specifically as follows:

[0108] Search for the drug code to obtain the drug embedding tensors within each batch, and input them into the first to third layers of one - dimensional convolution in sequence. Use the drug hidden layer representation extracted by the third - layer one - dimensional convolution as the initial features of the drug;

[0109] Among them, the expression of the drug hidden layer representation extracted by each layer of one - dimensional convolution is as follows:

[0110]

[0111] In the formula, is the drug embedding tensor within a batch, is the drug hidden layer representation extracted by the first - layer one - dimensional convolution, is the drug hidden layer representation extracted by the second - layer one - dimensional convolution, is the drug hidden layer representation extracted by the third - layer one - dimensional convolution, is the weight matrix in the CNN convolution kernel of the (l + 1)-th layer one - dimensional convolution of the drug CNN module, is the bias vector in the CNN convolution kernel of the (l + 1)-th layer one - dimensional convolution of the drug CNN module.

[0112] In S23, the expression for obtaining the protein representation is specifically as follows:

[0113]

[0114] In the formula, is, and are all similarity space embedding functions, which are composed of two layers of different linear functions and the Relu activation function, ( ) is the softmax function, is the query matrix of the current batch of proteins in the protein - drug similarity space, is the key matrix of the drugs in the batch in the protein - drug similarity space, is the value matrix of the drug in the current batch in the similarity space, is the dimension of, and T represents the matrix transpose operation;

[0115] obtain the drug molecule representation The specific expression of is as follows:

[0116] .

[0117] In this embodiment, a cross-attention mechanism is adopted to simulate the interaction between proteins and drug molecules. The cross-attention mechanism is a variant of the Self-Attention mechanism, which can fuse two different types of data for cross-representation to improve the representation ability of the model. To ensure the symmetry of the model, only one cross-attention module is adopted here, and the representation of the other type of data can be obtained by simply changing the input positions of the protein and the drug molecule in the module.

[0118] In S24, the protein representation and the drug molecule representation obtained in the above process are respectively fed into the max-pooling layer to calculate the feature vector of the protein and the feature vector of the drug , where N represents the number of drug-protein pairs in a batch, D represents the feature dimension, and its specific expression is:

[0119]

[0120]

[0121] In the formula, is the max-pooling function.

[0122] In S3, calculating the contrastive loss in any batch includes the following sub-steps:

[0123] S31. Calculate the contrastive loss of proteins in a batch according to the feature vectors of proteins and the similarity of proteins by the contrastive learning method;

[0124] S32. Calculate the contrastive loss of drugs in a batch according to the feature vectors of drugs and the similarity of drugs by the contrastive learning method;

[0125] S33. Obtain the contrastive loss in any batch according to the contrastive loss of proteins and the contrastive loss of drugs in a batch.

[0126] In S31, the contrastive learning method is specifically:

[0127] S311. Map the feature vectors of proteins in a batch to the molecular function comparison space, the domain annotation comparison space, and the sequence annotation comparison space through three independent modules with two fully connected layers;

[0128] S312. Construct positive and negative pairs in the comparison space according to the similarity of proteins, and calculate the comparison loss of proteins in the three comparison spaces;

[0129] S313. Obtain the comparison loss of proteins in a batch according to the comparison losses of proteins in the three comparison spaces.

[0130] In S311, the expressions for mapping the feature vectors of proteins in a batch to the molecular function comparison space, the domain annotation comparison space, and the sequence annotation comparison space are specifically as follows:

[0131]

[0132] In the formula, k is the comparison space, is the molecular function comparison space, is the domain annotation comparison space, seq is the sequence annotation comparison space, are the weights and biases of the first linear layer mapped by the comparison space k, is the bias of the first linear layer mapped by the comparison space k, is the weight of the second linear layer mapped by the comparison space k, is the bias of the second linear layer mapped by the comparison space k, is the hidden feature of the protein in the current batch in the comparison space k, , where N is the number of drug-protein pairs in a batch, and D represents the feature dimension;

[0133] In S312, the expression for calculating the comparison loss of proteins in the comparison space k is specifically as follows:

[0134]

[0135] In the formula, batch represents the set of proteins in a batch. is the positive pair of protein i in the current batch in the comparison space k, is the hidden feature of protein i in the comparison space k, sim is the function for calculating similarity, and the cosine similarity is taken here, is the hidden feature of protein j in the comparison space k, is the temperature coefficient in contrastive learning, is the hidden feature of protein z in the comparison space k.

[0136] In this embodiment, constructing positive pairs and negative pairs in the corresponding space according to the similarity of proteins is specifically as follows:

[0137] For each similarity of proteins and drugs, different similarity thresholds are set. When the similarity is higher than the similarity threshold, this drug pair or protein pair is regarded as a positive pair; otherwise, it is regarded as a negative pair. The specific similarity thresholds are measured through experiments by means of discrete values. During the experiment of the similarity matrix threshold, it is considered that different similarity thresholds are independent. Therefore, the present invention constructs 6 contrastive learning modules each containing only one similarity matrix, then uses a series of discrete values to conduct experiments on the DrugBank dataset, and finally selects the value with the optimal accuracy as the similarity threshold for the corresponding similarity matrix.

[0138] In S313, obtaining the contrastive loss of proteins within a batch The specific expression is as follows:

[0139]

[0140] In the formula, is the contrastive loss of proteins in the sequence annotation contrast space, is the contrastive loss of proteins in the domain annotation contrast space, is the contrastive loss of proteins in the molecular function contrast space.

[0141] In S32, the principle of calculating the contrastive loss of drugs within a batch by the contrastive learning method is the same as that of calculating the contrastive loss of proteins within a batch in S31. Therefore, the method of calculating the contrastive loss of drugs within a batch will not be elaborated here.

[0142] In S33, obtaining the contrastive loss within any batch The specific expression is as follows:

[0143]

[0144] In the formula, is the contrastive loss of drugs within a batch.

[0145] The specific content of S4 is as follows:

[0146] Concatenate the feature vectors of proteins and drugs within each batch, input the concatenated result into the fully connected layer for prediction, calculate the interaction probability between proteins and drugs, obtain the cross-entropy loss within each batch, and obtain the final loss according to the cross-entropy loss and the contrastive loss.

[0147] Among them, the cross-entropy loss The specific expression is as follows:

[0148]

[0149] In the formula, y represents the true label of the drug-protein pair interaction, and p represents the probability of the drug-protein pair interaction predicted by the model.

[0150] Final loss The specific expression of is as follows:

[0151]

[0152] In S5, the method for model training is specifically as follows:

[0153] Taking the minimization of the final loss as the optimization goal, AdamW is used as the optimizer for gradient descent, the parameters are updated with a learning rate of 0.0001, the maximum number of training rounds is set to 200 rounds, and the early stopping technique is used to determine whether the model training is completed, and the early stopping step is set to 50 steps.

[0154] When the trained model is used for prediction, the protein and drug data with unknown interactions are encoded, and after encoding, they are input into the trained model to obtain the prediction result of the model. Similarity data is no longer required during prediction.

[0155] To verify the beneficial effects of the present invention, the following experiments are given in the present invention.

[0156] Experiment 1: Performance comparison and analysis of CLSF-DTI and benchmark methods

[0157] First, the benchmark method and CLSF-DTI are trained on 5 datasets: DrugBank, KIBA, Enzymes, Ion Channels, and GPCRs. For the two larger datasets, DrugBank and KIBA, the present invention uses 5 metrics for comprehensive comparison: Accuracy, AUC, AUPR, Precision, Recall. The experimental results are shown in Tables 1-2. For the three small datasets, Enzymes, Ion Channels, and GPCRs, the performance of each model is not good enough. Therefore, the present invention uses three comprehensive metrics, Accuracy, AUC, and AUPR, to compare the performance of each model. The experimental results are shown in Table 3. The following conclusions can be drawn from the experimental results.

[0158] On four datasets, the method of the present invention significantly outperforms other methods. This indicates that introducing structural and functional information through contrastive learning can better help the model to model DTI prediction. Compared with the MCANet model, the accuracy, AUC, and AUPR of the method of the present invention on the Enzymes dataset are increased by 4.3%, 4.9%, and 4.4% respectively, and on DrugBank, the accuracy, AUC, and AUPR are increased by 2.5%, 2.4%, and 2.5% respectively, and on the GPCRs dataset, the accuracy, AUC, and AUPR are increased by 3.1%, 3.5%, and 2.9% respectively.

[0159] On the KIBA dataset, the performance of all methods is not very different. This shows that there are still certain difficulties in modeling these extremely imbalanced datasets such as KIBA with the current methods.

[0160] The two models, TransformerCPI and HyperAttention, have good performance on the Drugbank and KIBA datasets, but perform poorly on the three small datasets, Enzymes, Ion Channels, and GPCRs, while DeepConv-DTI, MCANet, and CLSF-DTI have good performance on these 5 datasets. This indicates that for small datasets, CNN-based models are more competitive than Transformer-based models.

[0161] Table 1 Performance of CLSF-DTI and baselines on DrugBank

[0162]

[0163] Table 2 Performance of CLSF-DTI and baselines on KIBA

[0164]

[0165] Table 3 Performance of CLSF-DTI and baselines on Enzymes, Ion Channels, and GPCRs

[0166]

[0167] Experiment 2: Analysis of the generalization ability of CLSF-DTI for unknown data

[0168] To evaluate the generalization ability of CLSF-DTI, the present invention re-partitions the DrugBank dataset in three ways to form three sub-datasets: the unseen drug dataset, the unseen protein dataset, and the unseen both dataset. First, the present invention splits 10% of the DrugBank dataset as the test sets for the above three datasets. Then the present invention removes the DTIs containing the drugs in the test sets from the remaining data to form the unseen drug dataset. In a similar way, the present invention can construct the unseen protein and unseen both datasets. In other words, the training set of unseen drug does not contain the drug data of the test set, the training set of unseen protein does not contain the protein data of the test set, and the training set of unseen both does not contain either the drug data or the protein data in the test set. Then, the above training sets are evenly divided into 5 parts for five-fold cross-validation. Finally, the average value of the model on the test set is used as the experimental result, as shown in Table 4.

[0169] Compared with the complete DrugBank dataset, the performance of all models has significantly decreased on these unseen datasets, and the performance degradation is the most severe on the unseen both dataset. However, CLSF-DTI still has the optimal performance, which indicates that CLSF-DTI has stronger generalization ability.

[0170] The performance of all models on the unseen drug dataset is better than that on the unseen protein dataset, which shows that removing protein data has a greater impact on the model. This is because compared with drug molecules, the chemical space of proteins is more extensive and it is more difficult to model proteins. Therefore, the generalization ability of the model for proteins is worse than that for drugs.

[0171] Table 4 Performance of CLSF-DTI and baselines on unseen datasets

[0172]

[0173] Experiment 3: Conduct ablation experiments to analyze each component of CLSF-DTI

[0174] To evaluate the utility of each module in the model of the present invention, the present invention decomposes the CLSF-DTI model into the following several variants:

[0175] NoAttention: The attention mechanism in the CLSF-DTI model is removed

[0176] NoDrugContrast: The contrastive learning module for drugs in CLSF-DTI is removed.

[0177] NoProteinContrast: The contrastive learning module for proteins in CLSF-DTI is removed.

[0178] NoContrast: All contrastive learning modules in CLSF-DTI are removed.

[0179] The present invention evaluates the above variant models on the DrugBank dataset, and the experimental results are as Figure 4 shown. The present invention can draw the following conclusions:

[0180] Both the contrastive learning module and the attention mechanism contribute to improving the performance of the model in DTI prediction. However, the attention mechanism has a relatively small impact on the model performance. After removing the attention mechanism, the accuracy of the model only drops by 0.6%, while after removing all contrastive learning modules, the accuracy of the model drops by 2.7%. This shows that using contrastive learning to introduce the functional and structural information of proteins and drug molecules does help to improve the performance of DTI prediction.

[0181] The performance of NoDrugContrast is better than that of NoProteinContrast, which indicates that after removing the protein contrastive learning module, the model performance drops more. This may be because the extraction of protein features is more difficult, and adding the protein contrastive learning module can greatly improve this problem, thus bringing a greater gain to the model performance.

[0182] As Figure 4 shown, where NoContrast means removing all contrastive learning modules in CLSF-DTI, NoProteinContrast means removing the contrastive learning module for proteins in CLSF-DTI, NoDrugContrast means removing the contrastive learning module for drugs in CLSF-DTI, and NoAttention means removing the CrossAttention module in CLSF-DTI. ALL means CLSF-DTI.

[0183] Experiment 4: Conduct a virtual screening experiment of active drugs on the target protein PKA-Cα using CLSF-DTI

[0184] To test the usability of CLSF-DTI in the real world, the present invention performed virtual screening on cAMP-dependent protein kinase catalytic subunit alpha (PKA-Cα, uniprot id: P17612) in KIBA. PKA-Cα is encoded by the PRKACA gene and can phosphorylate a large number of substrates in the cytoplasm and nucleus. Previous studies have shown that PKA-Cα is related to the formation of cardiovascular diseases and breast cancer. First, the present invention collected all the small molecule data in the Chembl database and removed the molecules that could not be represented by the drug embedding matrix E d After that, there were 1,913,075 molecules left. Then, these drug molecules and proteins were numerically encoded, and finally, they were sent into the fourth-fold model trained by KIBA for prediction. The present invention deleted the effective binders of PKA-Cα in the KIBA dataset from the screening results, and finally obtained 34,749 drug molecules predicted to be positive. Then, the top 10,000 drug molecules ranked by prediction were selected and compared with the Chembl database, and it was found that 22 drugs among them were verified to interact with PKA-Cα. Their specific information is shown in Table 5. The results show that the model of the present invention has a certain predictive ability for potential DTI, and those drugs that have not been verified are good candidates for wet experiments.

[0185] Table 5 Drugs that can interact with PKA-Cα among the top ten thousand results of virtual screening

[0186]

[0187] Experiment Five: Visualize the interaction between proteins and drugs on CLSF-DTI

[0188] To further demonstrate the effectiveness of the virtual screening results of CLSF-DTI, the present invention visualized the prediction results of the model on the protein PKA-Cα. The present invention first searched the PDB database for the 3D structure resolution results of PKA-Cα and found that two working PDBs: 2GU8 and PDB: 5N23 were involved in the screened molecules. Then, the present invention calculated the attention weights of the protein to the drug molecules in the cross-attention module. Finally, the present invention used Pymol for visualization. In this process, the present invention regarded the amino acids within 5 Å around the drug molecules as the binding sites of the drugs. The visualization results are as Figures 5-6As shown, the parts marked in red are the binding sites with high weight scores given by the model, which are the binding regions correctly predicted by the model. The blue regions are the true binding sites that are not given high weight scores. And the yellow fragments are non-binding site regions but are given relatively high weight scores. It can be observed that the model of the present invention can correctly predict some of the binding sites (red regions). Moreover, some of the yellow fragments are very close to the binding sites, and these fragments are likely to be potential binding sites. And although the two PDB complexes have the same protein, the model of the present invention gives very different attention weight distributions. The above results show that CLSF-DTI indeed has a certain insight ability in DTI prediction.

[0189] The beneficial effects of the present invention are as follows: The present invention provides a deep learning prediction method for drug-protein interactions. By calculating the similarities between drugs and between proteins in the high-dimensional structure, and then using the method of contrast learning in these similarity spaces to implicitly introduce this high-dimensional structure information into the features of protein and drug molecules. Thus, while avoiding direct modeling of sparse high-dimensional data, this high-dimensional information is introduced into the features, improving the accuracy of drug-protein interaction prediction and solving the problems of difficult modeling of high-dimensional sparse data and insufficient modeling performance of low-dimensional data with limited information.

[0190] To verify the effectiveness of the model CLSF-DTI proposed by the present invention, the following five experiments were conducted. The results of these experiments show that the present invention has the following several advantages:

[0191] CLSF-DTI has stronger applicability in drug-protein interaction prediction. Different from some methods that can only be used on large-scale datasets, CLSF-DTI can not only be used on large-scale datasets but also has good performance on small-scale datasets. This is because this method extracts features with CNN, and the parameters of the model are relatively few, and a small amount of data can also complete the optimization of the model.

[0192] Compared with other benchmark methods, CLSF-DTI has better performance in drug-protein interaction prediction. This is because after introducing high-dimensional structure data, the model can more accurately identify the differences between drugs and proteins.

[0193] CLSF-DTI has stronger prediction ability for unknown data than previous methods. This is because after using contrast learning to introduce high-dimensional structure data, the model can more easily capture the patterns between data, thus obtaining stronger generalization ability.

[0194] CLSF-DTI has the potential for virtual screening of drugs. During the experiment of virtual screening of active drugs against the protein PKA-Cα, CLSF-DTI successfully found 22 effective ligands. This benefits from the strong generalization ability brought by the contrastive learning module.

[0195] CLSF-DTI has certain insight ability into drug-protein interactions. The cross-attention module can simulate the interaction between proteins and drug molecules. It can be found in the visualization experiment that the model can correctly predict the binding sites of some drugs on the protein.

[0196] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "thickness", "upper", "lower", "horizontal", "top", "bottom", "inner", "outer", "radial", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of technical features. Therefore, the features defined by "first", "second", "third" may explicitly or implicitly include one or more of such features.

Claims

1. A deep learning prediction method for drug-protein interaction, characterized in that: The following steps are involved: S1. Collect drug and protein data and calculate drug similarity and protein similarity; S2, encode the drug and protein data and extract the feature vectors of proteins and drugs in each batch; S3, calculating the contrast loss within each batch based on the similarity of drugs, the similarity of proteins and the feature vectors of proteins and drugs within each batch; S4, concatenating the feature vectors of proteins and drugs in each batch, calculating the cross entropy loss in each batch based on the concatenated results, and obtaining the final loss based on the cross entropy loss and the contrast loss; S5. Train the model based on the final loss, input the protein code and drug code into the trained model to obtain the prediction results.

2. The deep learning prediction method for drug-protein interaction according to claim 1, characterized in that: In S1, the similarity of the drug is specifically: the Tanimoto similarity of the drug molecule on the set fingerprint, wherein the set fingerprint includes the substructure fingerprint, topological structure fingerprint and pharmacophore fingerprint of the drug molecule; The protein similarity includes the sequence similarity of the protein , domain similarity Similarity in molecular function , its specific expression is as follows: In the formula, are the sequence information of proteins i and j respectively, are the structural domain information of protein i and j respectively, and are the molecular function annotation sets of proteins i and j, is the similarity between the molecular function annotation t and the molecular function annotation set MF(j), is the similarity between the molecular function annotation o and the molecular function annotation set MF(i), Molecular function annotation collection The number of elements of Molecular function annotation collection The number of elements of .

3. The deep learning prediction method for drug-protein interaction according to claim 2, characterized in that: The S2 comprises the following sub-steps: S21, encoding the amino acid sequence of the protein and the SMLES data of the drug molecule to obtain a protein code and a drug code; S22, inputting the protein code into the protein CNN module to obtain the initial features of the protein, and inputting the drug code into the drug CNN module to obtain the initial features of the drug; S23, inputting the initial features of the protein and the initial features of the drug into the cross-attention mechanism module to obtain protein representation and drug molecule representation; S24. Input the protein representation into the first maximum pooling layer to obtain the feature vector of the protein in each batch, and input the drug molecule representation into the second maximum pooling layer to obtain the feature vector of the drug in each batch.

4. The deep learning prediction method for drug-protein interaction according to claim 3, characterized in that: In S22, the protein CNN module and the drug CNN module have the same structure, both including the first to third layers of one-dimensional convolutions connected in sequence; The method for obtaining the initial characteristics of the protein is as follows: Look up the protein code to obtain the protein embedding tensor in each batch, input it into the first to third layers of one-dimensional convolution in sequence, and use the protein hidden layer representation extracted by the third layer of one-dimensional convolution as the initial feature of the protein; Among them, the expression of the protein hidden layer extracted by each one-dimensional convolution is as follows: Where, the ordinal number l∈{0, 1, 2}, is the protein embedding tensor within a batch, is the protein hidden layer representation extracted by the first layer of one-dimensional convolution, The protein hidden layer representation extracted by the second layer of one-dimensional convolution, The protein hidden layer representation extracted by the third layer of one-dimensional convolution, is the weight matrix in the CNN convolution kernel of the l+1th layer of the protein CNN module, is the bias vector in the CNN convolution kernel of the l+1th layer of the protein CNN module, is a one-dimensional convolutional neural network operation, is a nonlinear activation function; The method for obtaining the initial characteristics of the drug is specifically: Look up the drug code to obtain the drug embedding tensor in each batch, input it into the first to third layers of one-dimensional convolution in sequence, and use the hidden layer representation of the drug extracted by the third layer of one-dimensional convolution as the initial feature of the drug; Among them, the expression of the hidden layer representation of the drug extracted by each one-dimensional convolution is as follows: In the formula, is the drug embedding tensor within a batch, is the hidden layer representation of the drug extracted by the first layer of one-dimensional convolution, The hidden layer representation of drugs extracted by the second layer of one-dimensional convolution, The hidden layer representation of drugs extracted by the third layer of one-dimensional convolution, is the weight matrix in the CNN convolution kernel of the l+1th layer of the drug CNN module, It is the bias vector in the CNN convolution kernel of the l+1th layer of the one-dimensional convolution of the drug CNN module.

5. The deep learning prediction method for drug-protein interaction according to claim 4, characterized in that: In S23, the protein expression The specific expression is: In the formula, for, and They are all similar space embedding functions, which consist of two layers of different linear functions and Relu activation functions. ( ) is the normalized exponential function, is the query matrix of the current batch of proteins in the protein-drug similarity space, is the bond matrix of drugs in the batch in the protein-drug similarity space, is the value matrix of the current batch of drugs in the similarity space, for The dimension of , T represents the matrix transpose operation; Get drug molecule representation The specific expression is: 。 6. The deep learning prediction method for drug-protein interaction according to claim 5, characterized in that: In S3, calculating the contrast loss in any batch includes the following steps: S31, calculating the contrast loss of proteins in a batch according to the feature vectors of proteins in a batch and the similarity of proteins by contrastive learning method; S32, calculating the contrast loss of the drugs in a batch according to the feature vectors of the drugs in a batch and the similarity of the drugs by a contrastive learning method; S33. According to the comparative loss of protein and the comparative loss of drug in a batch, the comparative loss in any batch is obtained.

7. The deep learning prediction method for drug-protein interaction according to claim 6, characterized in that: In S31, the contrastive learning method is specifically: S311, mapping the feature vectors of proteins in a batch to the molecular function comparison space, domain annotation comparison space and sequence annotation comparison space through three independent modules with two fully connected layers; S312, constructing positive pairs and negative pairs in the comparison space according to the similarity of the proteins, and calculating the comparison loss of the proteins in the three comparison spaces; S313. Obtain the contrast loss of proteins in a batch according to the contrast loss of proteins in three contrast spaces.

8. The deep learning prediction method for drug-protein interaction according to claim 7, characterized in that: In S311, the expression for mapping the feature vector of proteins in a batch to the molecular function comparison space, the domain annotation comparison space and the sequence annotation comparison space is specifically: Where k is the contrast space, For molecular function comparison space, is the domain annotation comparison space, seq is the sequence annotation comparison space, To compare the weights and biases of the first linear layer of the spatial k-mapping, To compare the deviation of the first linear layer of the spatial k-mapping, To compare the weights of the second linear layer of the spatial k-mapping, To compare the deviation of the second linear layer of the spatial k-mapping, is the hidden feature of the proteins in the current batch in the comparison space k, , where N is the number of drug-protein pairs in a batch, and D represents the feature dimension; In S312, the contrast loss of the protein in the contrast space k is calculated The specific expression is: In the formula, represents the protein set within a batch, is the positive pair of protein i in the current batch in comparison space k, is the hidden feature of protein i in comparison space k, sim is the function of similarity calculation, is the hidden feature of protein j in the contrast space k, To compare the temperature coefficient in the study, is the hidden feature of protein z in the comparison space k.

9. The deep learning prediction method for drug-protein interaction according to claim 8, characterized in that: In S4, the cross entropy loss The specific expression is: In the formula, y represents the true label of the drug-protein interaction, and p represents the probability of the drug-protein interaction predicted by the model.

10. The deep learning prediction method for drug-protein interaction according to claim 1, characterized in that: In S5, the method for performing model training is specifically as follows: Taking minimizing the final loss as the optimization goal, AdamW is used as the gradient descent optimizer, the model parameters are updated at a learning rate of 0.0001, the maximum number of training rounds is set to 200, and the early stopping technique is used to determine whether the model training is complete, and the number of early stopping steps is set to 50.

Citation Information

Cited By

  • Drug protein binding rate prediction method

    CN120741402A

  • Method and equipment for establishing target protein acidity coefficient prediction model, medium and program product

    CN120853685A

  • Antibody drug conjugate property prediction method based on multi-modal fusion

    CN121281625A

  • A method for predicting the properties of antibody-drug conjugates based on multimodal fusion

    CN121281625B