Drug Interaction Prediction Method Based on Interactive Enhanced Graph Self-Attention Mechanism
By modeling drugs in local and global views, combining interaction-enhanced graph self-attention mechanism and multi-channel adaptive fusion, dynamically regulating the interaction between drug pairs, the problem of insufficient single perspective and interaction modeling in the existing methods is solved, and more accurate drug interaction prediction is achieved.
Patent Information
- Application Number
- CN202510552009.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing drug interaction prediction methods have limitations in exploring the interaction between drug pairs. A single drug perspective limits the deep mining of interactions. Multi-view information fusion is poor, and interaction modeling is insufficient, so it is impossible to effectively capture the intensity of dynamic interactions.
The method based on the interaction-enhanced graph self-attention mechanism is adopted, and the drugs are modeled as molecular graphs and motif graphs in local and global views respectively. The node embedding is updated through the interaction-enhanced attention mechanism, and the fusion characteristics are fusionized through the multi-channel adaptive fusion module, a dynamic interactive scaling prediction module is designed, and the interaction intensity is adjusted using Gumbel-Softmax distribution to construct a drug interaction prediction model.
It significantly improves the ability to capture the relationship between drug pairs and the accuracy and stability of DDI prediction, breaks through the limitations of a single drug perspective, fully integrates the characteristics of multi-view drug angle, dynamically regulates the interaction intensity, and improves the prediction effect of the model.
Smart Images

Figure CN120072354B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of drug interaction prediction, and particularly relates to a drug interaction prediction method based on an interactive enhanced graph self-attention mechanism. Background Art
[0002] With the continuous development of modern medicine, for patients with multiple complications or intractable diseases, combined drug use has become a quite common treatment mode in clinical practice. However, the combined intake of multiple drugs may cause them to interact with each other in the patient's body, thus triggering the phenomenon of drug-drug interaction (DDI). DDI may lead to weakened drug efficacy, increased toxicity or other adverse reactions, and even threaten life in severe cases. Therefore, studying and predicting drug interactions not only has important significance for improving the safety of drug treatment, but also provides a scientific basis for personalized medicine and precision drug use.
[0003] Traditional DDI identification mainly relies on limited sources such as professional databases, clinical trial reports and existing literature. Researchers and clinicians avoid common interaction risks by retrieving existing cases or using empirical knowledge. However, due to the large variety of drugs, complex medication scenarios, and the continuous emergence of new drugs or new medication combinations, it is difficult to effectively predict potential unknown interactions solely relying on traditional methods, which cannot meet the current clinical needs. In recent years, with the rapid development of computational methods, algorithm-based DDI prediction has gradually become a powerful tool. Compared with traditional experimental methods, computational methods can process a large amount of data in a shorter time and do not require high-cost and time-consuming experimental operations, thus providing a more efficient and accurate solution for DDI prediction and prevention.
[0004] At present, computational DDI recognition methods can be roughly divided into indirect methods (drug association-based) and direct methods (drug structure-based). Indirect methods mainly rely on constructing the association relationships between drugs, often representing drugs and their interaction relationships through heterogeneous networks or knowledge graphs. Drugs and their related biomarkers, targets, pathways and other entities are regarded as nodes in the network, while the relationships between them are represented by edges. By learning the potential association information between drugs, these methods can effectively capture the interactions between drugs and thus be used for DDI prediction. However, without stable data and financial support, indirect DDI prediction methods cannot guarantee effective data quality. Direct methods focus on encoding the molecular structure information of drugs, believing that the occurrence of drug interactions is mainly driven by the interactions between substructures or functional groups within the molecule. Therefore, by deeply analyzing the drug structure, its potential DDI characteristics can be revealed. As a deep learning method for processing graph structures, graph neural networks (GNNs) have achieved remarkable results in multiple fields in recent years, especially showing strong modeling capabilities in drug molecular graphs. Nyamabo AK et al. used graph attention networks (GATs) to extract substructure information at different levels in drugs, and then used the co-attention mechanism to measure the importance of the interactions between substructures for DDI prediction. Yang Z et al. proposed D-MPNN, which assigns different weights to substructures with different radii during the message passing process to obtain size-adaptive drug substructures for subsequent DDI prediction. Transformers were originally designed to process sequence data and have achieved great success especially in natural language processing (NLP) tasks. Their self-attention mechanism can capture the long-range dependencies between elements in the sequence, making them perform excellently in tasks such as translation and text generation. By adaptively adjusting the self-attention mechanism, graph transformers can effectively process the complex relationships between nodes and the structural information of graphs, and have gradually become an important tool in graph data analysis. Compared with traditional GNNs, graph transformers show stronger flexibility and efficiency when dealing with large-scale graph data and long-range dependencies. Niu D et al. proposed a substructure refined mechanism based on improved graph transformers to update the substructures of drugs and extract the most accurate substructures that determine drug properties. In addition, DAS-DDI proposed a DS encoder based on the graph transformer module to achieve synchronous update of nodes and keys to extract the substructures of drugs, demonstrating the feasibility of graph transformers in graph-structured data.
[0005] However, with the continuous advancement of the DDI prediction task, existing methods based on GNN and graph Transformers have several significant limitations when exploring the interactions between drug pairs. Existing drug encoders limit their perspectives to a single drug and only perform feature propagation between nodes within the graph structure of a single drug through self-attention mechanisms, which restricts the exploration of interactions between drug pairs. In addition, although many methods extract drug information by designing multiple drug views, most of these methods simply concatenate different views together and fail to effectively utilize the potential synergistic effects between different views, resulting in a significant reduction in the effect of multi-view information fusion. This not only leads to insufficient information utilization but also inevitably results in a certain degree of information redundancy. Moreover, in the pairwise interaction prediction task of drug pairs, existing methods generally assist in prediction by calculating the interaction relationships between substructures or functional groups between drug pairs. Although this can capture certain local structural information, its limitation lies in the insufficient dynamic modeling of interactions, ignoring that the interaction intensity between different substructures or functional groups in a specific drug pair is dynamically changing. Summary of the Invention
[0006] In view of the above problems existing in the prior art, the present invention proposes a drug interaction prediction method based on an interaction-enhanced graph self-attention mechanism, which solves the deficiencies of the prior art.
[0007] The drug interaction prediction method based on an interaction-enhanced graph self-attention mechanism includes the following steps:
[0008] S1. Model the drugs as molecular graphs and motif graphs respectively under local views and global views;
[0009] S2. Design a graph Transformer module for enhanced interaction, update the node embeddings respectively through an interaction-enhanced attention mechanism in two views of the molecular graph, and fuse them through a multi-channel adaptive fusion module;
[0010] S3. Obtain the final molecular representations of the molecular graph and the motif graph by passing the extracted features through a feed-forward neural network, a residual connection layer, and a normalization layer;
[0011] S4. Design a dynamic interaction scaling prediction module to calculate the DDI score;
[0012] S5. Construct a drug interaction prediction model with the graph Transformer module, the multi-channel adaptive fusion module, the feed-forward neural network, the residual connection layer, the normalization layer, and the dynamic interaction scaling prediction module, design a loss function, and optimize the model parameters.
[0013] Further, in S1, in the local view, each atom in the drug molecule is defined as a node, and the chemical bond is used as the edge connecting the nodes. The one-hot method is used to construct a feature vector to describe the atomic symbol and chemical properties, and the one-hot method is also used for the chemical bond to construct a feature vector to characterize the bond type and conjugation information; the local molecular graph represents the drug in the local view, is the node representation of the drug in the local view, is the representation of the chemical bond of the drug in the local view, is the adjacency information between nodes of the drug in the local view;
[0014] In the global view, a molecular graph is constructed with atoms as nodes and chemical bonds as edges. A relative distance matrix is introduced to describe the global positional relationship between atoms. The element values in the relative distance matrix are determined by calculating the shortest path between nodes, and this value reflects the relative distance between atoms in the molecular structure. The global molecular graph represents the drug in the global view, is the node representation of the drug in the global view, is the representation of the chemical bond of the drug in the global view, is the adjacency information between nodes of the drug in the global view, represents the relative distance matrix between each node in the drug;
[0015] Through a chemically inspired molecular fragmentation function, the complex molecular structure is decomposed into multiple chemically meaningful motifs. A vocabulary containing unique motifs is constructed using the ChEMBL29 database, and then the local molecular graph and the global molecular graph are respectively converted into a local motif graph and a global motif graph with motifs as nodes and shared atom connections as edges.
[0016] Further, in S2, the drug pair is divided into a head drug and a tail drug. For the head drug in the molecular graph, in the local view, the initial embedding of the head drug node is respectively mapped to the query space and the value space, and the initial embedding of the tail drug node is mapped to the key space, and then the interaction strength between the drug pair nodes is calculated and the features of the nodes are updated. The expression is:
[0017] ;(1)
[0018] ;(2)
[0019] ;(3)
[0020] ;(4)
[0021] Among them, , and are both learnable weight matrices, is the node embedding matrix of the head drug in the local view at the -th layer, which is the initial embedding of the head drug when = 0, is the embedding vector of the -th node in the local view of the head drug at the +1-th layer; is the softmax activation function, is 's dimension, is the number of neighbors of the -th node in the head drug, is the number of attention heads in the interactive enhanced attention; is the node embedding matrix of the tail drug in the local view at the -th layer, , , are respectively the query matrix of the -th node in the local view of the head drug, the key matrix of the -th node in the local view of the tail drug, and the value matrix of the -th node in the local view of the head drug;
[0022] In the global view, the initial embeddings of the head drug nodes are respectively mapped to the query space and the value space, and the initial embeddings of the tail drug nodes are mapped to the key space. At the same time, the relative distance matrix between the nodes in the head drug is constructed; given the node set of the head drug, for any two nodes and , calculate their relative distance in the head drug molecular graph, which is defined as the shortest path between the -th node and the -th node. Calculate the interactive attention between the drug pair nodes and update the node features through the following formula:
[0023] ; (5)
[0024] ; (6)
[0025] ; (7)
[0026] ; (8)
[0027] where, , and is a learnable weight matrix, is the number of nodes of the head drug, is 's dimension, is a hyperparameter, is the node embedding matrix of the head drug under the global view in the l-th layer, is the tail drug in the layer under the global view of the node embedding matrix, is the head and tail drugs in the +1 layer under the global view of the -th node embedding vector, , , are respectively the query matrix of the -th node in the head drug in the global view, the key matrix of the -th node in the tail drug in the global view, and the value matrix of the -th node in the head drug in the global view.
[0028] Furthermore, in the S2, a multi-channel adaptive fusion module is designed to fuse the local view and the global view corresponding to the molecular graph, and the expression is:
[0029] ; (9)
[0030] ; (10)
[0031] where is a learnable multi-channel operator for storing the attention map of the interaction under the two views, is the normalization operation on the embedding, and are both learnable weight matrices, represents the matrix multiplication operation, represents the element-wise multiplication operation, C is the number of channels in, is the bias vector, is the attention map between the -th node and the -th node in the c-th channel, is the fusion feature of the -th node of the head drug after passing through the multi-channel adaptive fusion module.
[0032] Furthermore, in the S3, the extracted features are passed through a feed-forward neural network, a residual connection layer, and a normalization layer to obtain the final molecular representation of the head drug, and the expression is:
[0033] ; (11)
[0034] ; (12)
[0035] ; (13)
[0036] Among them, is a learnable weight matrix, is the GELU activation function, is the layer normalization operation, is the th node embedding of the head drug output by the th layer, is the +1 th node embedding of the head drug output by the th layer;
[0037] The +1 th node embedding of the tail drug output by the th layer is , the tail drug molecule is represented as , is the number of tail drug nodes;
[0038] The obtained molecular graph embeddings and are used to update the initial embeddings of the motif graph of the head drug and the initial embedding of the motif graph of the tail drug respectively, and the expression is:
[0039] ; (14)
[0040] ; (15)
[0041] Among them, is the feature fusion function, is the initial embedding of the motif graph after the head drug is updated, is the initial embedding of the motif graph after the tail drug is updated;
[0042] and are used as the inputs of the interaction-enhanced graph Transformer module again to obtain the final hidden embeddings and of the motif graphs of the head drug and the tail drug.
[0043] Furthermore, in the above S4, first, a co-attention-based interaction layer is used to capture the interaction intensity between drugs, and a drug representation with context information is generated by weighted fusion of the features of drug pairs , the expression is:
[0044] ; (16)
[0045] Among them, and respectively represent the first elements corresponding to the final hidden embeddings of the head drug and the tail drug;
[0046] Next, calculate the interaction information between drug pairs, and the expression is:
[0047] ; (17)
[0048] ; (18)
[0049] Among them, and respectively represent the remaining parts of the final hidden embeddings of the head drug and the tail drug except the first elements, is the interaction information of the head drug, is the interaction information of the tail drug, is the trainable weight matrix;
[0050] In order to dynamically adjust the interaction strength between drug pairs, the Gumbel-Softmax distribution is used, and the expression is:
[0051] ; (19)
[0052] ; (20)
[0053] ; (21)
[0054] Among them, is the noise sampled from the uniform distribution is the probability distribution matrix sampled from the Gumbel-Softmax distribution of the head drug of the th element, represents the relative preference of the th functional group classification category in the head drug, is the temperature parameter, is the number of functional groups;
[0055] Finally, aggregate the features of the drug pair and input them into a multi-layer perceptron for DDI score calculation, and the expression is:
[0056] ; (22)
[0057] ;(23)
[0058] ;(24)
[0059] Among them, is the probability distribution matrix sampled from the Gumbel-Softmax distribution for the tail drug, is the prediction activation function, represents the vector concatenation operation, is the trainable weight matrix, is the DDI prediction probability.
[0060] Furthermore, in the step S5, by collecting various types of drug pairs in DrugBank, the drug interaction prediction model is tested using multi-classification tasks and binary classification tasks. The multi-classification task is to identify which DDI reaction the drug pair will have, and the binary classification task is to identify whether there is a DDI reaction between the drug pair;
[0061] For the multi-classification task, the following cross-entropy loss function is used to optimize the model parameters:
[0062] ;(25)
[0063] Among them, Z is the number of samples, is the number of classes, is the th sample's true label for the th class, is the model's predicted probability for the th sample in the th class;
[0064] For the binary classification task, binary cross-entropy is used as the loss function:
[0065] ;(26)
[0066] Among them, is the true label of the th sample, represents the class prediction probability of the th sample. When it means the model predicts the th sample as class 1, otherwise 0.
[0067] The beneficial technical effects brought by the present invention:
[0068] The present invention combines an interaction-enhanced graph Transformer module, a multi-channel adaptive fusion module, and an annealing algorithm-driven dynamic interaction scaling prediction module, successfully breaking through the limitations of the traditional single-drug perspective. At the same time, it fully integrates multi-perspective drug features, achieving accurate modeling of the interaction relationships between drug pairs. During the prediction process, the Gumbel-Softmax distribution and the annealing concept are introduced to dynamically adjust the interaction strength, thereby significantly improving the accuracy and stability of DDI prediction.
[0069] (1) By designing an interaction-enhanced attention mechanism in the graph Transformer module, the limitations of the traditional single-drug perspective are broken through, and the interaction between drug pairs can be explored more deeply. This mechanism models the interaction of drug pairs in the graph structure, thus effectively improving the ability to capture drug interaction relationships.
[0070] (2) To address the problem of fusing multi-view information of drugs, a module capable of interactive learning from different perspectives is proposed, which effectively integrates and utilizes the features and potential correlations of drugs in each view.
[0071] (3) To overcome the problem of insufficient interaction modeling in existing methods, the annealing concept and the Gumbel-Softmax distribution are introduced to explore and scale the interaction force between drug pairs, and the interaction strength is automatically adjusted during the training process. Description of the Drawings
[0072] Figure 1 It is a local view of the drug constructed in the present invention.
[0073] Figure 2 It is a global view of the drug constructed in the present invention.
[0074] Figure 3 It is a flowchart of the working process of the interaction-enhanced graph Transformer module in the present invention.
[0075] Figure 4 It is a flowchart of the multi-channel adaptive fusion module in the present invention.
[0076] Figure 5 It is a flowchart of the dynamic interaction scaling prediction module in the present invention.
[0077] Figure 6 It is a graph of the drug interaction prediction results of two groups of drug pairs in the present invention. Detailed Embodiments
[0078] The following further describes the specific embodiments of the present invention in conjunction with specific embodiments:
[0079] A method for predicting drug interactions based on an interaction-enhanced graph self-attention mechanism includes the following steps:
[0080] S1. Model the drug as a molecular graph and a motif graph respectively under the local view and the global view;
[0081] For the modeling of the drug, the present invention considers two representation methods of the drug. One is the drug based on the molecular graph (collectively referred to as the molecular graph), and the other is the drug based on the motif (i.e., the functional group that repeatedly appears in the drug) (collectively referred to as the motif graph). And the local view and the global view of the drug are respectively constructed based on the two methods, as Figure 1 and Figure 2 shown.
[0082] At the microscopic level, that is, the local view, specific functional groups serve as carriers of local information, and their chemical properties and spatial structures directly determine the affinity and specific binding ability between the drug molecule and the target. Taking the molecular graph as an example, in the local view, each atom in the drug molecule is defined as a node, and the chemical bond is used as the edge connecting the nodes. The one-hot is used to construct the feature vector to describe the atomic symbol and chemical properties, and the chemical bond also uses one-hot to construct the feature vector to characterize the bond type and conjugation information; the local molecular graph represents the drug in the local view, is the node representation of the drug in the local view, is the representation of the chemical bond of the drug in the local view, is the adjacency information between nodes of the drug in the local view;
[0083] At the macroscopic level, that is, the global view, the three-dimensional conformation, size scale of the whole molecule, and the interaction between long-distance regions cannot be ignored, and play a decisive role in regulating the biological activity, stability of the molecule, and the synergistic or antagonistic effects between drug molecules. In the global view, a molecular graph is also constructed with atoms as nodes and chemical bonds as edges, but a relative distance matrix is introduced to describe the global positional relationship between atoms. For a given set of nodes, the element values in the relative distance matrix are determined by calculating the shortest path between nodes, and this value reflects the relative position distance between atoms in the molecular structure. The global molecular graph represents the drug in the global view, is the node representation of the drug in the global view, is the representation of the chemical bond of the drug in the global view, is the adjacency information between nodes of the drug in the global view, represents the relative distance matrix between each node in the drug;
[0084] Secondly, the molecular graph provides high-resolution information for understanding the chemical environment of molecules. However, the complexity of drug molecules is not only reflected in the connections at the atomic level, but more importantly in the combination of substructures with specific chemical meanings. Therefore, in addition to the molecular graph, motif graphs are also introduced into the molecular characterization process. Through a chemically inspired molecular fragmentation function, complex molecular structures are decomposed into multiple motifs with chemical meanings. A vocabulary containing unique motifs is constructed using the ChEMBL29 database, and then the local molecular graph and the global molecular graph are respectively converted into a local motif graph and a global motif graph with motifs as nodes and shared atomic connections as edges, which can effectively capture key structural information such as functional groups in molecules, thereby characterizing molecular structures at a higher semantic level.
[0085] S2. Design a graph Transformer (self-attention mechanism) module that enhances interaction. Update the node embeddings separately through the interaction-enhanced attention mechanism in two views of the molecular graph, and fuse them through a multi-channel adaptive fusion module;
[0086] S3. Since the problem of gradient vanishing or gradient explosion may occur in deep networks, and at the same time, in order to eliminate the difference in feature distributions between different layers, avoid bias during model training, and improve the stability of the training process, the extracted features are passed through a feed-forward neural network, a residual connection layer, and a normalization layer to obtain the final molecular representations of the molecular graph and the motif graph;
[0087] The above graph Transformer module, multi-channel adaptive fusion module, feed-forward neural network, residual connection layer, and normalization layer have a total of L layers. The following calculates using the th layer as an example;
[0088] Traditional methods for predicting drug interactions based on graph model modeling are limited to a single drug space. In order to accurately capture the interactions between drug pairs and enhance the interaction modeling ability, the present invention designs a graph Transformer module that enhances interaction, as Figure 3 shown;
[0089] Taking the molecular graph as an example, drug pairs are divided into head drugs and tail drugs. Update the node embeddings separately through the interaction-enhanced attention mechanism in two views, and fuse them through a multi-channel adaptive fusion module. Taking the head drug as an example, first in the local view, the initial embeddings of the head drug nodes are respectively mapped to the query space and the value space, and the initial embeddings of the tail drug nodes are mapped to the key space, and then the interaction strength between the drug pair nodes is calculated and the features of the nodes are updated. The expression is:
[0090] ;(1)
[0091] ;(2)
[0092] ; (3)
[0093] ; (4)
[0094] Among them, , and are both learnable weight matrices, is the node embedding matrix of the head drug under the local view in the -th layer. When = 0, it is the initial embedding of the head drug, is the embedding vector of the +1-th layer of the head drug under the local view at the -th node; is the softmax activation function, is the dimension of, is the number of neighbors of the -th node in the head drug. Neighbors are adjacent nodes, is the number of attention heads in the interaction-enhanced attention; is the node embedding matrix of the tail drug under the local view in the -th layer, , , are respectively the query matrix of the -th node in the head drug under the local view, the key matrix of the -th node in the tail drug under the local view, and the value matrix of the -th node in the head drug under the local view;
[0095] In the global view, the initial embedding of the head drug nodes is mapped to the query space and the value space respectively, the initial embedding of the tail drug nodes is mapped to the key space, and at the same time, the relative distance matrix between the nodes in the head drug is constructed; given the node set of the head drug, for any two nodes and , calculate their relative distance in the head drug molecular graph, defined as the shortest path between the -th node and the -th node. For example, if the shortest path from the node to passes through , and the edge distances connecting to , to are both 1, then , calculate the interaction attention between drugs for nodes and update the features of the nodes through the following formula:
[0096] ; (5)
[0097] ; (6)
[0098] ; (7)
[0099] ; (8)
[0100] where , and are learnable weight matrices, is the number of nodes of the head drug, is 's dimension, is a hyperparameter, is the node embedding matrix of the head drug under the global view in the th layer, is the node embedding matrix of the tail drug under the global view in the th layer, is the embedding vector of the +1th node of the head drug under the global view in the th layer, , , are respectively the query matrix of the th node of the head drug in the global view, the key matrix of the th node of the tail drug in the global view, and the value matrix of the th node of the head drug in the global view.
[0101] Taking the tail drug as an example, obtain the embedding vector +1th node of the tail drug under the local view in the th layer, , and the embedding vector +1th node of the tail drug under the global view in the th layer .
[0102] In S2, a multi-channel adaptive fusion module is designed to fuse the local view and the global view corresponding to the molecular graph respectively. As Figure 4 shown, the expression is:
[0103] ; (9)
[0104] ; (10)
[0105] Among them, is a learnable multi-channel operator for storing the attention maps of the interaction under two views, is the normalization operation on the embedding, and are both learnable weight matrices, represents the matrix multiplication operation, represents the element-wise multiplication operation, C is the number of channels in is the bias vector, is the th node and the th attention map between the th node after the head drug passes through the multi-channel adaptive fusion module.
[0106] The local view and the global view corresponding to the molecular graph are fused for the tail drug to obtain the fused feature .
[0107] Up to this point, the multi-channel adaptive fusion module fuses the information of the local view and the global view. Specifically, the attention map calculation of the multi-channel adaptive fusion module enables the model to focus on the key information of the interaction between the two views, where the learnable multi-channel operator and the weight matrix enable the model to adaptively adjust the attention degree to the information of different views. This fusion method can not only mine the potential associations of drug features from different perspectives, avoiding the limitations of only focusing on the independent features of each view, but also fully utilize the collaborative information between views through interactive learning, greatly enhancing the effect of multi-view learning.
[0108] In S3, the extracted passes through the feed-forward neural network, the residual connection layer and the normalization layer to obtain the final molecular representation , and the expression is:
[0109] ; (11)
[0110] ; (12)
[0111] ; (13)
[0112] Among them, is the learnable weight matrix, is the GELU activation function, is the layer normalization operation, is the The first node embedding of the head drug output by the layer, is the first node embedding of the head drug output by the +1 layer;
[0113] The extracted passes through a feed-forward neural network, residual connection, and normalization to obtain the final molecular representation , , is the first node embedding of the tail drug output by the +1 layer, where
[0114] The obtained molecular graph embedding and are used to update the initial embeddings of the motif graphs of the head drug and the tail drug , respectively, and the expression is:
[0115] ; (1)
[0116] ; (2)
[0117] where is the feature fusion function, is the initial embedding of the motif graph after the head drug is updated, is the initial embedding of the motif graph after the tail drug is updated;
[0118] and are used as the inputs of the interactive enhanced graph Transformer module again, and the above process, i.e., formulas (1)~(12), is repeated to obtain the final hidden embeddings and of the motif graphs of the head drug and the tail drug.
[0119] S4. Design a dynamic interaction scaling prediction module to calculate the DDI score;
[0120] Aiming at the deficiencies of existing DDI prediction methods in interaction modeling, the present invention designs a dynamic interaction scaling prediction module. By introducing the annealing concept and the Gumbel-Softmax distribution, it explores and scales the interaction force between drug pairs and automatically adjusts the interaction strength between drug pairs during the training process, such as Figure 5As shown. Specifically, the interactive scaling factor is dynamically adjusted through the annealing algorithm, making the interaction stronger in the initial stage of training and gradually weakening in the later stage to avoid overfitting.
[0121] First, a co-attention-based interaction layer is used to capture the interaction intensity between drugs, and the features of drug pairs are weighted and fused to generate drug representations with context information , and the expression is:
[0122] ;(16)
[0123] where and respectively represent the first elements corresponding to the final hidden embeddings of the head drug and the tail drug;
[0124] Next, calculate the interaction information between drug pairs, and the expression is:
[0125] ;(17)
[0126] ;(18)
[0127] where and respectively represent the remaining parts of the final hidden embeddings of the head drug and the tail drug except the first element, is the interaction information of the head drug, is the interaction information of the tail drug, is the trainable weight matrix;
[0128] Taking the head drug as an example, in order to dynamically adjust the interaction intensity between drug pairs, the Gumbel-Softmax distribution is used, and the expression is:
[0129] ;(19)
[0130] ;(20)
[0131] ;(21)
[0132] where is the noise sampled from the uniform distribution to introduce randomness, making the shape of the probability distribution different each time of sampling. The noise will perturb the obtained interaction intensity, making the sample selection process exploratory rather than simply choosing the maximum value (greedy strategy), which makes Gumbel-Softmax a smooth approximation to the discrete sampling mechanism. is the probability distribution matrix sampled from the Gumbel-Softmax distribution for the head drug The th element, represents the relative preference for the th functional group classification category in the head drug, i.e., , Combining the sampling noise with to generate a new score, is the temperature parameter, which is used to control the smoothness of this process. When the temperature is high, the distribution becomes smoother; when the temperature is low, the distribution becomes sharper, approaching a hard decision (i.e., a choice close to the maximum value). At the beginning of model training, holds a high temperature, making the probability distribution tend to be uniform and the sampling results more random. As the model training process progresses, the temperature is gradually reduced, which means annealing, making the probability distribution tend to a single extreme value, so that a certain degree of randomness and exploration can be maintained during the sampling process, is the exponential function, which converts the new score into a positive value to form the original form of the probability distribution, is the number of functional groups;
[0133] Taking the tail drug as an example, the probability distribution matrix sampled from the Gumbel-Softmax distribution for the tail drug is .
[0134] Finally, the features of the drug pair are aggregated and input into a multi-layer perceptron for DDI score calculation. The expression is:
[0135] ;(22)
[0136] ;(23)
[0137] ;(24)
[0138] where, is the prediction activation function. Since two classification tasks are designed, when facing a multi-classification task, is set to the Softmax activation function. When facing a binary classification task, is set to the Sigmoid activation function, represents the vector concatenation operation, is the trainable weight matrix, is the DDI prediction probability.
[0139] S5. Construct a drug interaction prediction model by using the L-layer graph Transformer module, the multi-channel adaptive fusion module, the feed-forward neural network, the residual connection layer, the normalization layer, and the dynamic interaction scaling prediction module. Design a loss function and optimize the model parameters.
[0140] DrugBank is a classic drug database containing rich and reliable drug information. By collecting various types of drug pairs in DrugBank, the drug interaction prediction model is tested using multi-classification tasks and binary classification tasks. The multi-classification task is to identify which DDI reaction the drug pair will have, and the binary classification task is to identify whether there is a DDI reaction between the drug pair.
[0141] For the multi-classification task, use the following cross-entropy loss function to optimize the model parameters:
[0142] ;(25)
[0143] where Z is the number of samples, is the number of classes, is the -th sample's true label for the -th class, is the predicted probability of the model for the -th sample in the -th class;
[0144] For the binary classification task, use binary cross-entropy as the loss function:
[0145] ;(26)
[0146] where, is the true label of the -th sample, represents the class prediction probability of the -th sample. When , it means the model predicts the -th sample as class 1, otherwise 0.
[0147] The comparison results of the model (AMIE-DDI) proposed in this invention with other advanced models (MR-GNN, SA-DDI, DSN-DDI, PEB-DDI, MeTDDI) in two datasets. ZhongDDI is a multi-classification dataset, and ZhangDDI is a binary classification dataset. Using ACC, AUROC, and AUPR as the evaluation metrics of the model, as shown in Table 1. The model proposed in this invention has achieved performance improvement compared with the current advanced models in both classification scenarios.
[0148] Table 1 Comparison of evaluation indicators between the model of the present invention and other advanced models
[0149] ;
[0150] As Figure 6 shown, in the drug pair with label 0 composed of DB01558 and DB00277, the model proposed by the present invention labeled the pyridine ring (C1CN=CccN1) and the alkyl group (Cc) in DB01558, and these two motifs are closely related to the binding site of the metabolic enzyme. For DB00277, the model labeled the pyrimidine ring (c1cncnc1) and the methylamine (Cn). The pyrimidine ring usually competitively occupies the active site of the metabolic enzyme through coordination, inhibiting the metabolic activity of other drugs.
[0151] In the drug pair with label 1 composed of DB06266 and DB00454, the aminomethyl group (Cn) labeled by the model proposed by the present invention is the key functional group of DB06266 and is often used as the recognition and binding site for metabolic enzymes (such as the CYP450 family). The methylamine group (CN), as an electron-donating group, is prone to form a bond with the metabolic enzyme. DB00454 may bind to the active site of the metabolic enzyme (such as CYP450) through its methylamine group. In medicinal chemistry, compounds containing the methylamine group often undergo phase I reactions such as N-oxidation during metabolism, indicating that this group is prone to interact with the metabolic enzyme.
[0152] The above shows that the model proposed by the present invention can identify the key functional groups in the process of drug interaction, improving the interpretability of the model in the process of drug interaction.
[0153] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for predicting drug-drug interactions based on an interaction-enhanced graph self-attention mechanism, characterized in that It includes the following steps: Drug-drug interaction prediction method based on an interaction-enhanced graph self-attention mechanism S1. Model drugs as molecular graphs and motif graphs under local views and global views respectively; S2. Design a graph Transformer module for enhanced interaction, update node embeddings respectively through an interaction-enhanced attention mechanism in two views of the molecular graph, and fuse them through a multi-channel adaptive fusion module; S3. Obtain the molecular representations of the final molecular graph and motif graph by passing the extracted features through a feed-forward neural network, a residual connection layer, and a normalization layer; S4. Design a dynamic interaction scaling prediction module to calculate the DDI score; In S4, first, an interaction layer based on co-attention is used to capture the interaction intensity between drugs, and the features of drug pairs are weighted and fused to generate a drug representation with context information. , next, the interaction information between drug pairs is calculated. To dynamically adjust the interaction intensity between drug pairs, the Gumbel-Softmax distribution is used. Finally, the features of drug pairs are aggregated and input into a multi-layer perceptron for DDI score calculation. S5. Construct a drug-drug interaction prediction model with the graph Transformer module, the multi-channel adaptive fusion module, the feed-forward neural network, the residual connection layer, the normalization layer, and the dynamic interaction scaling prediction module, design a loss function, and optimize the model parameters.
2. The method for predicting drug-drug interactions based on the interactive enhanced graph self-attention mechanism according to claim 1, wherein In S1, in the local view, each atom in the drug molecule is defined as a node, and the chemical bond is used as the edge connecting the nodes. The one-hot method is used to construct the feature vector to describe the atomic symbol and chemical properties. The chemical bond also uses one-hot to construct the feature vector to characterize the bond type and conjugation information; the local molecular graph is used to represent the drug in the local view, which is the node representation of the drug in the local view, which is the representation of the chemical bond of the drug in the local view, which is the adjacency information between nodes of the drug in the local view; In the global view, a molecular graph is constructed with atoms as nodes and chemical bonds as edges. A relative distance matrix is introduced to describe the global positional relationship between atoms. The element values in the relative distance matrix are determined by calculating the shortest path between nodes, and this value reflects the relative distance between atoms in the molecular structure. The global molecular graph characterizes the drug in the global view, is the node representation of the drug in the global view, is the representation of the chemical bonds of the drug in the global view, is the adjacency information between nodes of the drug in the global view, represents the relative distance matrix between each node in the drug; Through a chemistry-inspired molecular fragmentation function, the complex molecular structure is decomposed into multiple chemically meaningful motifs, a vocabulary containing unique motifs is constructed using the ChEMBL29 database, and then the local molecular graph and the global molecular graph are respectively converted into a local motif graph and a global motif graph with motifs as nodes and shared atom connections as edges.
3. The method for predicting drug interactions based on an interaction-enhanced graph self-attention mechanism according to claim 2, wherein In S2, the drug pair is divided into a head drug and a tail drug. For the head drug in the molecular graph, in the local view, the initial embedding of the head drug node is respectively mapped to the query space and the value space, the initial embedding of the tail drug node is mapped to the key space, and then the interaction intensity between the drug pair nodes is calculated and the features of the nodes are updated. The expression is: ;(1) ;(2) ;(3) ;(4) Among them, , and are all learnable weight matrices, is the node embedding matrix of the head drug in the local view at the -th layer, and it is the initial embedding of the head drug when = 0. is the embedding vector of the +1-th node of the head drug in the local view at the -th layer; is the softmax activation function, is 's dimension, is the number of neighbors of the -th node in the head drug, is the number of attention heads in the interactive enhanced attention; is the node embedding matrix of the tail drug in the local view at the -th layer, , , are the query matrix of the -th node of the head drug in the local view, the key matrix of the -th node of the tail drug in the local view, and the value matrix of the -th node of the head drug in the local view, respectively; In the global view, the initial embeddings of the head drug nodes are respectively mapped to the query space and the value space, and the initial embeddings of the tail drug nodes are mapped to the key space, while constructing the relative distance matrix between the nodes in the head drug ; Given the set of nodes of the head drug , for any two nodes and , calculate their relative distance in the head drug molecular graph , defined as the shortest path between the -th node and the -th node, calculate the interactive attention between the drug pair nodes through the following formula and update the features of the nodes: ;(5) ;(6) ;(7) ;(8) Among them, , and are learnable weight matrices, is the number of nodes of the head drug, is 's dimension, is a hyperparameter, is the node embedding matrix of the head drug under the global view in the l-th layer, is the node embedding matrix of the tail drug under the global view in the -th layer, is the embedding vector of the +1-th node of the head and tail drugs under the global view in the -th position, , , are respectively the query matrix of the -th node in the head drug in the global view, the key matrix of the -th node in the tail drug in the global view, and the value matrix of the -th node in the head drug in the global view.
4. The method for predicting drug interactions based on the interactive enhanced graph self-attention mechanism according to claim 3, wherein In S2, a multi-channel adaptive fusion module is designed to fuse the local view and the global view corresponding to the molecular graph. The expression is: ;(9) ;(10) Among them, is a learnable multi-channel operator for storing the attention maps of interactions under two views. is the normalization operation on the embedding. and are both learnable weight matrices. represents the matrix multiplication operation. represents the element-wise multiplication operation. C is the number of channels in is the bias vector. is the attention map between the -th node and the -th node in the c-th channel. is the fusion feature of the -th node after the head drug passes through the multi-channel adaptive fusion module.
5. The method for predicting drug interactions based on an interaction-enhanced graph self-attention mechanism according to claim 4, wherein In S3, the extracted features are passed through a feedforward neural network, a residual connection layer, and a normalization layer to obtain the final representation of the lead drug molecule , and the expression is: ;(11) ;(12) ;(13) Among them, is a learnable weight matrix, is the GELU activation function, is the layer normalization operation, is the -th node embedding of the head drug output by the -th layer, is the +1-th node embedding of the head drug output by the -th layer; The first node embedding of the tail drug output at the +1 layer is , the tail drug molecule is represented as , where is the number of tail drug nodes; Embed the obtained molecular graph and respectively update the initial embeddings of the motif graphs of the head drug and the initial embeddings of the motif graphs of the tail drug , and the expression is: ;(14) ;(15) Among them, is the feature fusion function, is the initial embedding of the motif map after the head drug update, is the initial embedding of the motif map after the tail drug update; and are re - used as the input to the interactive enhanced graph Transformer module to obtain the final hidden embeddings of the motifs of the head drug and the tail drug and .
6. The method for predicting drug interactions based on an interaction-enhanced graph self-attention mechanism according to claim 5, wherein In S4, first, a co-attention-based interaction layer is used to capture the interaction intensity between drugs, and the features of drug pairs are weighted and fused to generate a drug representation with context information. , and the expression is: ;(16) Among them, and respectively represent the first elements corresponding to the final hidden embeddings of the head drug and the tail drug; Next, calculate the interaction information between drug pairs. The expression is: ;(17) ;(18) Among them, and respectively represent the remaining parts except the first element in the final hidden embeddings of the head drug and the tail drug, is the interaction information of the head drug, is the interaction information of the tail drug, is a trainable weight matrix; To dynamically adjust the interaction intensity between drug pairs, the Gumbel-Softmax distribution is used. The expression is: ;(19) ;(20) ;(21) wherein, is the noise sampled from a uniform distribution and is the probability distribution matrix sampled from the head drug Gumbel-Softmax distribution of the th element representing the relative preference for the th functional group classification category in the head drug is the temperature parameter and is the number of functional groups; Finally, aggregate the features of the drug pair and input them into a multi-layer perceptron for DDI score calculation. The expression is: ;(22) ;(23) ;(24) Among them, is the probability distribution matrix sampled by the tail drug from the Gumbel-Softmax distribution, is the prediction activation function, represents the vector concatenation operation, is the trainable weight matrix, is the DDI prediction probability.
7. The method for predicting drug interactions based on an interactive enhanced graph self-attention mechanism according to claim 6, wherein In S5, by collecting various types of drug pairs in DrugBank, the drug-drug interaction prediction model is tested using multi-classification tasks and binary-classification tasks. The multi-classification task is to identify which DDI reaction the drug pair will have, and the binary-classification task is to identify whether there is a DDI reaction between the drug pair; For multi-classification tasks, use the following cross-entropy loss function to optimize the model parameters: ;(25) where Z is the number of samples, is the number of classes, is the true label of the th sample in the th class, and is the predicted probability of the model for the th sample in the th class; For binary classification tasks, use binary cross-entropy as the loss function: ;(26) Among them, is the true label of the th sample, represents the class prediction probability of the th sample. When , it means that the model predicts the th sample as class 1, otherwise 0.
Citation Information
Patent Citations
Molecular interaction prediction method based on knowledge graph and multi-task learning
CN117037898A
Computer-implemented method, apparatus, computer-program product
WO2024113215A1