Drug-target interaction prediction method assisted by graph neural network
Through the combination of graph neural network and dynamic hypergraph, the problem of insufficient local integration in the existing methods is solved, and more comprehensive and accurate drug-target interaction prediction is achieved, improving the efficiency and accuracy of drug design.
Patent Information
- Application Number
- CN202510419815.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
Existing drug-target interaction prediction methods mostly focus on physical combination at the local level, and fail to effectively capture global dependencies, making it difficult for the model to accurately predict the type of drug-target interaction relationship.
The graph neural network is used to polymerize the local neighborhood information of drug molecules and target proteins, and their global embedding representations are obtained through dynamic hypergraphs. The local and global embedding representations are combined for comparison learning and fusion. The graph prompt learning technology is used to capture the atomic distribution and functional group structural characteristics, and a two-part graph of the drug-target relationship is constructed.
It improves the comprehensiveness and accuracy of drug-target interaction prediction, and enhances the prediction ability of the model by capturing the complex higher-order interaction between drug molecules and target proteins.
Smart Images

Figure CN120279983A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technology of predicting the relationship between drugs and targets, and particularly relates to a method for predicting drug-target interactions assisted by a graph neural network. Background Art
[0002] Predicting drug-target interactions (DTIs) is of great significance in drug research and development because it can help discover new targets, improve research and development efficiency, and reduce risks. However, traditional high-throughput screening technologies are time-consuming and expensive, making it an efficient and economical choice to assist DTI prediction with artificial intelligence technologies. In recent years, graph neural networks (GNNs) have demonstrated excellent performance in analyzing and understanding graph-structured data. They can effectively capture the complex relationships between nodes and are widely used in multiple fields, including social networks, biomedicine, etc. In the field of biomedicine, GNNs are particularly suitable for modeling the three-dimensional structures of drug molecules and protein targets and their interaction patterns, providing a new perspective and technical means for drug design.
[0003] In recent years, researchers have proposed various methods for the DTI prediction problem. For example, the fragment-based DTI prediction method uses a convolutional neural network to learn the synergistic effects between various fragments of drugs and targets, thereby reducing the training cost and improving the generalization ability; for example, the multi-feature fusion method based on deep learning uses the BPE algorithm and a multi-attention framework for multi-granularity encoding to achieve the encoding learning of multi-character atoms and chemical functional groups, thereby improving the representation ability of drug molecules; for example, the cross-attention mechanism combined with the Transformer network and domain adaptation technology (CAT-DTI) enhances the performance of the DTI prediction model through global feature extraction and cross-domain adaptation.
[0004] Although existing methods have achieved certain results in DTI prediction, there are still the following main limitations. On the one hand, the extraction of multiple structural information in drug molecules is insufficient. Existing methods cannot fully utilize the characteristics of molecules in terms of atomic distribution and functional group distribution, resulting in bias in molecular representation and affecting the accuracy of model prediction. For example, the aromatic ring of gefitinib (an anti-cancer drug) enhances its binding stability to the EGFR active site through key π-π stacking interactions, while the hydroxyl group forms hydrogen bonds to further consolidate the binding. If only the atomic composition of the drug is considered here and the role of the aromatic ring is ignored, it will lead to a prediction bias of the inhibitory activity of gefitinib. On the other hand, the interaction relationship between drugs and targets in DTI prediction has dual attributes of local and global. Existing methods mostly focus on the physical binding at the local level and fail to effectively capture the global dependence relationships, such as the regulation of target expression levels by drugs and functional regulation through allosteric sites, resulting in limitations in the comprehensive characterization of interaction relationships by the model. For example, gefitinib binds to the ATP site of the EGFR protein and inhibits it by inducing its conformational change to an inactive state. If only the local physical interaction between the drug and the ATP binding site is considered and this global dependence relationship is ignored, it will be difficult for the model to comprehensively characterize the inhibitory mechanism of the drug. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the graph neural network-assisted drug-target interaction prediction method provided by the present invention solves the problem that existing methods mostly focus on the physical binding at the local level and fail to effectively capture the global dependence relationships, resulting in the difficulty for the model to accurately predict the type of drug-target interaction relationship.
[0006] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:
[0007] Provide a graph neural network-assisted drug-target interaction prediction method, which includes the steps of:
[0008] S1. Obtain the drug molecule and target protein information in the dataset and construct a bipartite graph of the drug-target relationship;
[0009] S2. Use a graph neural network to aggregate the local neighborhood information of the bipartite graph and obtain the local embedding representation matrices of drug molecules and target proteins respectively;
[0010] S3. According to the local embedding representation matrices of drug molecules and target proteins, use a dynamic hypergraph to obtain the global embedding representation matrices of drug molecules and target proteins;
[0011] S4. Perform contrastive learning and fusion on the local embedding representation matrices and global embedding representation matrices of drug molecules and target proteins to obtain the embedding representations of drug molecules and target proteins;
[0012] S5. Use a fully connected layer to concatenate the embedded representation of the drug molecule and the embedded representation of the target protein, and then input it into the prediction model to obtain the probability of the drug-target interaction relationship type.
[0013] Further, step S1 further includes:
[0014] S11. Obtain the drug molecule and target protein information in the dataset, use one-hot encoding to obtain the vector representation of the target protein, and use a pre-trained 3DGNN model to obtain the initial vector representation of the drug molecule graph;
[0015] S12. Update the initial vector of the drug molecule according to the learnable vector related to the atomic type and the attention mechanism to obtain the updated representation of each atomic node in the drug molecule;
[0016] S13. Identify the functional groups in the drug molecule, and combine its learnable prompt vector with the updated representations of all atomic nodes in the drug molecule to form the final drug molecule representation;
[0017] S14. Use multiple drug molecule representations and vector representations of multiple target proteins to construct a bipartite graph of the drug-target relationship.
[0018] The beneficial effects of the above technical solutions are as follows: By using the prompt learning technology to model the drug molecule structure, it is possible to capture the atomic distribution and functional group structure features, so as to better distinguish various types of atoms and functional groups, thereby improving the utilization rate of the drug-target structure information.
[0019] Further, step S12 further includes:
[0020] S121. Calculate the atomic node representation of the atoms in the drug molecule according to the initial vector representation of the drug molecule and the learnable vector related to the atomic type:
[0021]
[0022] where, is the atomic node representation of node v in the drug molecule graph, d is the dimension of E v ; is the initial vector representation of node v; T(v) is the atomic type of node v, is the learnable vector with atomic type T(v);
[0023] S122. Fuse the atomic node representation E v and the atomic node representation E u of the neighbor node u of node v in the drug molecule graph to obtain the fusion information of node v and u;
[0024] S123. Aggregate the fusion information of node v and different neighbor nodes using the attention mechanism to obtain the updated representation of atomic node v
[0025]
[0026] wherein, and are both learnable parameters, and N(v) is the set of neighbor nodes of node v in the drug molecular graph; E vm is the fusion information between node v and its neighbor node m; exp(·) represents the exponential function with base e.
[0027] The beneficial effects of the above technical solution are as follows: By designing a set of learnable vectors related to atomic types, the heterogeneity of nodes (i.e., atoms in the molecule) in the drug molecular graph can be captured.
[0028] Furthermore, step S13 is further included:
[0029] S131. Use the RDKit tool to identify the functional groups in the SMILES sequence of the drug molecule;
[0030] S132. Use average pooling to perform a readout operation on the updated representations of all atomic nodes of the drug molecule to generate a preliminary global representation of the drug molecule;
[0031] S133. Add the cue vector of the identified functional group to the preliminary global representation to generate the final drug molecule representation.
[0032] The beneficial effects of the above technical solution are as follows: By integrating the functional group characteristics of the drug molecule, the higher-order topological properties of the drug molecule can be reflected, and the structural characteristics of the drug molecule can be captured more comprehensively.
[0033] Furthermore, the expression of the final drug molecule representation is:
[0034]
[0035] wherein, is the final drug molecule representation; READOUT(·) represents the readout operation of average pooling on the representations of all atomic nodes of the drug molecule; is the updated representation of atomic node v; V is the set of all atomic nodes in the drug molecule; S(G) is the set of functional groups in the drug molecular graph G; p j is the learnable cue vector of the j-th functional group in S(G).
[0036] Furthermore, use a graph neural network to aggregate the local neighborhood information of the bipartite graph, and the propagation process of the l-th layer is expressed as:
[0037]
[0038] Among them, and are respectively the local embedding representation matrices of the drug molecules and target proteins aggregated by the l-th propagation layer; and are respectively the local embedding representation matrices of the drug molecules and target proteins aggregated by the (l + 1)-th propagation layer; is the adjacency matrix after normalizing the interaction matrix A of the drug molecules and target proteins; is the drug embedding representation matrix containing M drug molecules; is the target embedding representation matrix containing N target proteins;
[0039] The graph neural network iterates to the last L-th propagation layer to obtain the local embedding representation matrix of the drug molecules and the local embedding representation matrix
[0040] The beneficial effect of the above technical solution is that by using the graph neural network for message propagation in the constructed bipartite graph, the local interaction between the drug molecules and target proteins is fully captured, and the local embedding representation matrices of the drug molecules and target proteins with richer semantics are learned.
[0041] Furthermore, the expression for obtaining the global embedding representation matrices of the drug molecules and target proteins by using the dynamic hypergraph is:
[0042]
[0043] Among them, is the global embedding representation matrix learned from the dynamic hypergraph at the (l + 1)-th layer; is the global embedding representation matrix of the target proteins generated from the dynamic hypergraph at the (l + 1)-th layer; are respectively the drug-hyperedge matrix and the target-hyperedge matrix, and K is the number of hyperedges; and are respectively the hyperedge embedding representations of the drug molecules and the hyperedge embedding representations of the target proteins; T is the transpose symbol; W (d) and respectively represent the trainable parameter matrices of the drug-hyperedge matrix and the target-hyperedge matrix.
[0044] The beneficial effect of the above technical solution is that by using the dynamic hypergraph learning to explore the global interaction between the drug molecules and target proteins, the global embedding representation matrices of the drug molecules and target proteins with high-order structural information are learned.
[0045] Furthermore, the expressions for the embedded representation of the drug molecule and the embedded representation of the target protein are as follows:
[0046]
[0047] Among them, is the embedded representation of the i-th drug molecule; is the embedded representation of the j-th target protein; L is the total number of propagation layers of the graph neural network; is and are used to construct the fused embedded representation of drug molecule i by element-wise addition; is and are used to construct the fused embedded representation of target protein j by element-wise addition; is the local embedded representation matrix of drug molecule i aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedded representation matrix of drug molecule i at the l-th layer obtained through hypergraph learning; is the local embedded representation matrix of target protein j aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedded representation matrix of target protein j at the l-th layer obtained through hypergraph learning.
[0048] The beneficial effects of the above technical solution are as follows: By combining the local and global embedded representations of drug molecules and target proteins respectively, the semantic representations of drug molecules and target proteins can be captured more comprehensively and accurately.
[0049] Furthermore, the loss function of the dynamic hypergraph is:
[0050]
[0051] Among them, and are the loss function values of drug molecules and target proteins in the dynamic hypergraph respectively; sim(·) is the cosine similarity; τ is the temperature hyperparameter; is the global embedded representation matrix of drug molecule i' at the l-th layer obtained through hypergraph learning; is the global embedded representation matrix of target protein j ′ at the l-th layer.
[0052] The beneficial effects of the above technical solution are as follows: By optimizing the above loss function, the similarity of the local and global embedded representations of drug molecules and target proteins can be ensured, and the learned semantic representations of drug molecules and target proteins are more discriminative.
[0053] Furthermore, the expression for the probability of the drug-target interaction relationship type is as follows:
[0054]
[0055] Wherein, is the probability of the predicted drug-target interaction relationship type, C is the total number of relationship types; both W1 and W2 are learnable weight matrices, and both b1 and b2 are bias vectors; is the embedding representation of the i-th drug molecule; is the embedding representation of the j-th target protein; RELU(·) is the activation function; || is the concatenation operation of the fully connected layer;
[0056] The loss function of the prediction model is:
[0057]
[0058] Wherein, is the loss function of the prediction model; y c ∈{0,1} is whether the class label c is the true label of the drug-target pair; is the probability that the predicted drug-target interaction relationship type is type c.
[0059] The beneficial effects of the present invention are as follows: Firstly, this solution obtains the local embedding representation matrices of drug molecules and protein targets through a graph neural network, then obtains the global embedding representation matrices of drug molecules and target proteins through a dynamic hypergraph, and then conducts contrastive learning and fusion on the local embedding representation matrix and the global embedding of the drug-target interaction, and better models the complex high-order interaction relationship between drugs and targets according to the local and global relationships, thereby improving the comprehensiveness and reliability of drug-target interaction prediction.
[0060] This solution models the molecular structure through a graph hint learning method based on a graph neural network, and captures the atomic distribution and functional group structure features to better distinguish various types of atoms and functional groups, improve the utilization rate of the structural information of drug-targets, and improve the accuracy of drug-target interaction prediction. Brief Description of the Drawings
[0061] Figure 1 is the flowchart of the drug-target interaction prediction method assisted by a graph neural network.
[0062] Figure 2 is the schematic diagram of the initial vector representation of a drug molecule.
[0063] Figure 3 is the schematic diagram of a drug-target relationship bipartite graph. Detailed Embodiments
[0064] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0065] Referring to Figure 1 , Figure 1 which shows a flowchart of a method for predicting drug-target interactions assisted by a graph neural network; as Figure 1 shown, the method S includes steps S1 to S5.
[0066] In step S1, information on drug molecules and target proteins in the dataset is obtained, and a bipartite graph of the drug-target relationship is constructed; preferably, step S1 of this solution may further include the following sub-steps:
[0067] S11. Obtain information on drug molecules and target proteins in the dataset, use one-hot encoding to obtain a vector representation of the target protein, and use a pre-trained 3DGNN model to obtain an initial vector representation of the drug molecule graph; One-hot encoding is a method of converting categorical variables into a format that can be understood by machine learning models. For a protein sequence, each amino acid can be represented by a vector, and the length of the vector is equal to the number of all possible amino acid types.
[0068] For example, if the 20 amino acids are arranged in the following order: A, R, N, D, C, E, Q, G, H, I, L, K, M, F, P, S, T, W, Y, V, and each amino acid is assigned an index position (starting from 0 to 19), then: the one-hot encoding of Alanine (A) is [1, 0, 0,..., 0] (only the first position is 1), the one-hot encoding of Arginine (R) is [0, 1, 0,..., 0] (only the second position is 1), the one-hot encoding of Asparagine (N) is [0, 0, 1, 0,..., 0] (only the third position is 1), and so on. For any amino acid, there will be a 1 only at its corresponding position, and the rest of the positions are 0.
[0069] As Figure 2 shown, for the drug molecule, the drug molecule graph can be output as an initial vector representation through an existing pre-trained 3DGNN model where d is the dimension of the vector.
[0070] S12. Update the initial vector of the drug molecule according to the learnable vector related to the atomic type and the attention mechanism to obtain the updated representation of each atomic node in the drug molecule.
[0071] To capture the heterogeneity of the nodes in the drug molecule graph (in the solution, the drug molecule is represented by a "graph", and the nodes in the graph represent the atoms of the drug molecule), a set of learnable prompts p related to the atomic type are designed in this solution T(v) to alleviate the heterogeneity problem. The detailed implementation method for this step S12 is as follows:
[0072] S121. Calculate the atomic node representation of the atoms in the drug molecule according to the initial vector representation of the drug molecule and the learnable vector related to the atomic type:
[0073]
[0074] Among them, is the atomic node representation of the node v in the drug molecule graph, d is the dimension of E v ; is the initial vector representation of node v; T(v) is the atomic type of node v, is the learnable vector with atomic type T(v);
[0075] S122. Fuse the atomic node representation E v and the atomic node representation E u of the neighbor node u of node v in the drug molecule graph to obtain the fusion information of node v and u:
[0076] E vu = E v ||E u
[0077] Among them, E vu is the fusion information of node v and u;
[0078] S123. Use the attention mechanism to aggregate the fusion information of node v and different neighbor nodes to obtain the updated representation of atomic node v
[0079]
[0080] Among them, and are both learned parameters, N(v) is the set of neighbor nodes of node v in the drug molecule graph; E vm is the fusion information of node v and its neighbor node m; exp(·) represents the exponential function with base e.
[0081] S13. Identify the functional groups in the drug molecule, and combine their learnable prompt vectors with the updated representations of all atomic nodes of the drug molecule to form the final drug molecule representation. The detailed implementation process of step S13 may include the following sub-steps:
[0082] S131. Use the RDKit tool to identify the functional groups in the SMILES sequence of the drug molecule. The SMILES sequence is a cheminformatics standard for representing molecular structures using strings. For example, aspirin: CC(=O)OC1=CC=CC=C1C(=O)O.
[0083] S132. Perform a readout operation on the updated representations of all atomic nodes of the drug molecule using average pooling to generate a preliminary global representation of the drug molecule;
[0084] S133. Add the prompt vector of the identified functional group to the preliminary global representation to generate the final drug molecule representation:
[0085]
[0086] where, is the final drug molecule representation; READOUT(·) is the readout operation performed by average pooling on the representations of all atomic nodes of the drug molecule; is the updated representation of atomic node v; V is the set of all atomic nodes in the drug molecule; S(G) is the set of functional groups in the drug molecule graph G; p j is the learnable prompt vector of the j-th functional group in S(G).
[0087] Through the above steps, the model generates the final representation of the drug molecule This representation combines the atomic information and the characteristics of the functional groups of the drug molecule, and can more comprehensively reflect the structure and properties of the drug molecule, thereby improving the prediction performance of subsequent drug-target interactions.
[0088] S14. Use multiple drug molecule representations and vector representations of multiple target proteins to construct a bipartite graph of the drug-target relationship. The schematic of the bipartite graph can be referred to Figure 3 , Since the construction of the bipartite graph is a relatively mature technology in the prior art, it will not be elaborated here.
[0089] In step S2, use a graph neural network to aggregate the local neighborhood information of the bipartite graph to obtain the local embedding representation matrices of the drug molecule and the target protein respectively. Among them, when using a graph neural network to aggregate the local neighborhood information of the bipartite graph, the propagation process of the l-th layer is expressed as:
[0090]
[0091] where, and are respectively the local embedding representation matrices of the drug molecules and target proteins aggregated by the l-th propagation layer; and are respectively the local embedding representation matrices of the drug molecules and target proteins aggregated by the (l + 1)-th propagation layer; is the adjacency matrix after normalizing the interaction matrix A of the drug molecules and target proteins; is the drug embedding representation matrix containing M drug molecules; is the target embedding representation matrix containing N target proteins;
[0092] Finally, the graph neural network iterates through the above formula to the last L-th propagation layer to obtain the local embedding representation matrix of the drug molecules and the local embedding representation matrix
[0093] This solution is based on two learnable hyperedge connection matrices and which represent the drug-hyperedge matrix and target-hyperedge matrix respectively, and uses dynamic hypergraph learning to obtain their potential global dependencies, where K represents the number of hyperedges. Each hyperedge represents a potential relationship between drug nodes or target nodes in the bipartite graph. Assuming that the local structures of atomic nodes are similar, their hyperedge connections are more likely to be similar. Based on this assumption, the global embedding representation matrices of the drug molecules and target proteins in this solution are calculated.
[0094] In step S3, according to the local embedding representation matrices of the drug molecules and target proteins, use the dynamic hypergraph to obtain the global embedding representation matrices of the drug molecules and target proteins:
[0095]
[0096] where, is the global embedding representation matrix learned from the dynamic hypergraph at the (l + 1)-th layer; is the global embedding representation matrix of the target protein generated from the dynamic hypergraph at the (l + 1)-th layer; and are the drug-hyperedge matrix and target-hyperedge matrix respectively, and K is the number of hyperedges; and are the hyperedge embedding representations of the drug molecules and target proteins respectively; T is the transpose symbol; W (d) and represent the trainable parameter matrices of the drug-hyperedge matrix and target-hyperedge matrix respectively.
[0097] This solution performs contrastive learning by comparing the local embedded representation of the drug-target bipartite graph with the global embedded representation of the dynamic hypergraph; the local and global embedded representations of the same drug molecule (or target protein) are used as positive samples, and the local and global embedded representations of different drug molecules (or target proteins) are used as negative samples. The InfoNCE loss function is used to optimize the consistency, and the InfoNCE loss function used as the loss function of the dynamic hypergraph is as follows:
[0098]
[0099] where, and are the loss function values of drug molecules and target proteins in the dynamic hypergraph respectively; sim(·) is the cosine similarity; τ is the temperature hyperparameter used to control the penalty for positive and negative samples; is the global embedded representation matrix of drug molecule i obtained through hypergraph learning ′ at the l-th layer; is the global embedded representation matrix of target protein j obtained through hypergraph learning ′ at the l-th layer.
[0100] Through the above InfoNCE loss function, the distance of the positive sample pair can be minimized, while the distance of the negative sample pair can be maximized, that is, it can ensure that the local and global embedded representations are similar and maximize the discrimination between them.
[0101] In step S4, the local embedded representation matrix and the global embedded representation matrix of drug molecules and target proteins are subjected to contrastive learning and fusion to obtain the embedded representation of drug molecules and the embedded representation of target proteins:
[0102]
[0103] where, is the embedded representation of the i-th drug molecule; is the embedded representation of the j-th target protein; L is the total number of propagation layers of the graph neural network; is and are used to construct the fused embedded representation of drug molecule i by element-wise addition; is and are used to construct the fused embedded representation of target protein j by element-wise addition; is the local embedded representation matrix of drug molecule i aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedded representation matrix of drug molecule i obtained through hypergraph learning at the l-th layer; is the local embedding representation matrix of target protein j aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedding representation matrix of target protein j at the l-th layer obtained by hypergraph learning.
[0104] In step S5, a fully connected layer is used to concatenate the drug molecule embedding representation and the embedding representation of the target protein, and then the probability of the drug-target interaction relationship type is obtained by inputting into the prediction model. Among them, the expression of the probability of the drug-target interaction relationship type is:
[0105]
[0106] where, is the probability of the predicted drug-target interaction relationship type, C is the total number of relationship types; W1 and W2 are both learnable weight matrices, and b1 and b2 are both bias vectors; is the embedding representation of the i-th drug molecule; is the embedding representation of the j-th target protein; RELU(·) is the activation function; || is the concatenation operation of the fully connected layer;
[0107] The loss function of the prediction model is:
[0108]
[0109] where, is the loss function of the prediction model; y c ∈{0,1} is whether the class label c is the true label of the drug-target pair; is the probability that the predicted drug-target interaction relationship type is type c.
[0110] In summary, this solution can accurately model the molecular structure through the graph neural network, improving the accuracy of drug-target interaction prediction; based on the dynamic hypergraph, by contrastive learning and fusing local and global embedding representations, the modeling ability of the high-order complex interaction relationship between drugs and targets is enhanced.
Claims
1. A method for predicting drug-target interactions assisted by graph neural networks, characterized in that, Including the steps: S1. Obtain the drug molecule and target protein information in the dataset, and construct a bipartite graph of the drug-target relationship; S2. Use a graph neural network to aggregate the local neighborhood information of the bipartite graph, and respectively obtain the local embedding representation matrices of drug molecules and target proteins; S3. According to the local embedding representation matrices of drug molecules and target proteins, use a dynamic hypergraph to obtain the global embedding representation matrices of drug molecules and target proteins; S4. Perform contrastive learning and fusion on the local embedding representation matrices and global embedding representation matrices of drug molecules and target proteins to obtain the embedding representation of drug molecules and the embedding representation of target proteins; S5. Use a fully connected layer to concatenate the embedding representations of drug molecules and target proteins, and then input them into a prediction model to obtain the probability of the drug-target interaction relationship type.
2. The method for predicting drug-target interaction according to claim 1, wherein Step S1 further includes: S11. Obtain the drug molecule and target protein information in the dataset, use one-hot encoding to obtain the vector representation of the target protein, and use a pre-trained 3DGNN model to obtain the initial vector representation of the drug molecule graph; S12. Update the initial vector of the drug molecule according to the learnable vector related to the atom type and the attention mechanism to obtain the updated representation of each atomic node in the drug molecule; S13. Identify the functional groups in the drug molecule, and combine its learnable prompt vector with the updated representations of all atomic nodes of the drug molecule to form the final drug molecule representation; S14. Use multiple drug molecule representations and vector representations of multiple target proteins to construct a bipartite graph of the drug-target relationship.
3. The method for predicting drug-target interaction according to claim 1, wherein Step S12 further includes: S121. Calculate the atomic node representation of the atoms in the drug molecule according to the initial vector representation of the drug molecule and the learnable vector related to the atom type: Among them, represents the atomic node of node v in the drug molecule graph, and d is the dimension of E v ; is the initial vector representation of node v; T(v) is the atomic type of node v, is the learnable vector with atomic type T(v); S122. Represent the atomic node as E v and the atomic node representation E u of the neighbor node u of node v in the drug molecule graph are fused to obtain the fusion information of nodes v and u; S123. Aggregate the fusion information of node v and different neighbor nodes by using the attention mechanism to obtain an updated representation of atomic node v Among them, and are both parameters to be learned, and N(v) is the set of neighbor nodes of node v in the drug molecule graph; E vm is the fusion information between node v and its neighbor node m; exp(·) represents the exponential function with base e.
4. The drug-target interaction prediction method according to claim 2, wherein Step S13 further includes: S131. Use the RDKit tool to identify the functional groups in the SMILES sequence of the drug molecule; S132. Use average pooling to perform a readout operation on the updated representations of all atomic nodes of the drug molecule to generate a preliminary global representation of the drug molecule; S133. Add the prompt vector of the identified functional group to the preliminary global representation to generate the final drug molecule representation.
5. The method for predicting drug-target interaction according to claim 4, wherein The expression of the final drug molecule representation is: Among them, is the final representation of the drug molecule; READOUT(·) represents the readout operation of the average pooling on the representations of all atomic nodes of the drug molecule; is the updated representation of atomic node v; V is the set of all atomic nodes in the drug molecule; S(G) is the set of functional groups in the drug molecule graph G; p j is the learnable hint vector of the j-th functional group in S(G).
6. The method for predicting drug-target interaction according to claim 1, wherein Using a graph neural network to aggregate the local neighborhood information of the bipartite graph, the propagation process of the l-th layer is expressed as: Among them, and are the local embedding representation matrices of the drug molecules and target proteins aggregated by the l-th propagation layer, respectively; and are the local embedding representation matrices of the drug molecules and target proteins aggregated by the (l + 1)-th propagation layer, respectively; is the adjacency matrix after normalizing the interaction matrix A of the drug molecules and target proteins; is the drug embedding representation matrix containing M drug molecules; is the target embedding representation matrix containing N target proteins; The graph neural network is iterated to the last L-th propagation layer to obtain the local embedding representation matrix of the drug molecule and the local embedding representation matrix of the target protein 7. The method for predicting drug-target interaction according to claim 6, wherein The expression for obtaining the global embedding representation matrices of drug molecules and target proteins using a dynamic hypergraph is: Among them, is the global embedding representation matrix learned from the dynamic hypergraph at the (l + 1)-th layer; is the global embedding representation matrix of the target protein generated from the dynamic hypergraph at the (l + 1)-th layer; and are the drug-hyperedge matrix and the target-hyperedge matrix respectively, and K is the number of hyperedges; and are the hyperedge embedding representations of drug molecules and target proteins respectively; T is the transpose symbol; W (d) and represent the trainable parameter matrices of the drug-hyperedge matrix and the target-hyperedge matrix respectively.
8. The method for predicting drug-target interaction according to claim 7, wherein The expressions for the embedding representation of drug molecules and the embedding representation of target proteins are respectively: Among them, is the embedded representation of the i-th drug molecule; is the embedded representation of the j-th target protein; L is the total number of propagation layers of the graph neural network; is and the fused embedded representation of drug molecule i constructed by element-wise addition; is and the fused embedded representation of target protein j constructed by element-wise addition; is the local embedded representation matrix of drug molecule i aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedded representation matrix of drug molecule i obtained by hypergraph learning at the l-th layer; is the local embedded representation matrix of target protein j aggregated from itself and its bipartite graph neighbors at the l-th layer; is the global embedded representation matrix of target protein j obtained by hypergraph learning at the l-th layer.
9. The method for predicting drug-target interaction according to claim 8, wherein The loss function of the dynamic hypergraph is: wherein, and are the loss function values of the drug molecule and the target protein in the dynamic hypergraph, respectively; sim(·) is the cosine similarity; τ is the temperature hyperparameter; is the global embedding representation matrix of drug molecule i ′ at the l-th layer obtained by hypergraph learning; is the global embedding representation matrix of target protein j ′ at the l-th layer obtained by hypergraph learning.
10. The method for predicting drug-target interaction according to any one of claims 1-9, characterized in that, The expression for the probability of the drug-target interaction relationship type is: Among them, is the probability of the predicted drug-target interaction relationship type, C is the total number of relationship types; both W1 and W2 are learnable weight matrices, and both b1 and b2 are bias vectors; is the embedded representation of the i-th drug molecule; is the embedded representation of the j-th target protein; RELU(·) is the activation function; || is the concatenation operation of the fully connected layer; The loss function of the prediction model is: Among them, is the loss function of the prediction model; y c ∈ {0, 1} is whether the class label c is the true label of the drug-target pair; is the probability that the predicted drug-target interaction relationship type is type c.
Citation Information
Cited By
Drug-target effect prediction method and system based on multi-mode self-supervised learning
CN121884929A
Drug response prediction method based on multi-level interpretable hypergraph neural network
CN122417472A