Asymmetric drug interaction prediction method based on diffusion diagram attention network
The drug interaction model is constructed through the diffusion map attention network, which solves the problem of insufficient directional features, feature extraction and multimodal fusion of DDI prediction in the prior art, and achieves efficient and accurate drug interaction prediction, which improves the robustness and interpretability of the model.
Patent Information
- Application Number
- CN202510771829.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing DDI prediction technology has shortcomings in dealing with the directional characteristics, feature extraction, model interpretability and multimodal feature fusion of drug interactions, and it is difficult to meet the needs of large-scale drug screening and clinical applications.
Using a diffusion graph attention network method, by constructing a drug molecular similarity matrix and a directed DDI matrix, combining a bidirectional graph attention network and a multi-head attention mechanism, a diffusion model is introduced for feature diffusion and denoising, and a deep neural network is used to fusion of multimodal features to achieve efficient modeling and accurate prediction of drug interactions.
It improves the accuracy and interpretability of DDI prediction, enhances the model's adaptability to sparse data and captures complex interaction modes, reduces the dependence on labeled data, and expands application scenarios, especially drug interaction prediction in cold start.
Smart Images

Figure CN120340685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug effect prediction, and particularly to an asymmetric drug-drug interaction prediction method based on a diffusion graph attention network. Background Art
[0002] In modern medicine, the combined use of drugs has become an important means for treating complex diseases. However, drug-drug interactions (DDIs) may lead to adverse consequences such as altered drug efficacy and increased toxicity, seriously threatening the health and safety of patients. Therefore, accurately predicting DDIs is of great significance for optimizing drug combination regimens and reducing medical risks.
[0003] Traditional DDI detection methods mainly rely on experimental means, such as cell experiments and animal experiments. Although the results of these methods are reliable, they have problems such as complex operations, long time consumption, and high costs, and it is difficult to meet the needs of large-scale drug screening and clinical applications. In recent years, with the development of computer technology, machine learning and deep learning methods have gradually been introduced into the field of DDI prediction to improve prediction efficiency and accuracy.
[0004] However, existing DDI prediction technologies still have many limitations. First, most existing methods focus on the prediction of non-directional DDIs and do not fully consider the directional characteristics of DDIs, that is, the influence of one drug on another may be different from the reverse influence. This directionality is crucial for understanding the mechanism of drug-drug interactions and optimizing the dosing sequence. Second, existing methods have deficiencies in feature extraction. The chemical structure, biological properties, and pharmacokinetic characteristics of drugs are all important factors affecting DDIs, but existing methods often can only capture some of these features, resulting in incomplete prediction results. In addition, many methods are inefficient in processing large-scale data and are highly dependent on the quality and quantity of data, making it difficult to achieve accurate prediction in the case of limited data.
[0005] In terms of model design, most existing DDI prediction methods are based on simple graph neural network architectures, such as graph convolutional networks (GCNs) and graph attention networks (GATs). Although these models can capture the interactions between drugs, they still struggle to handle complex DDI patterns. For example, they are difficult to effectively distinguish the source role (exerting influence) and target role (receiving influence) between drugs, and they cannot fully utilize the multi-modal features of drugs (such as chemical structure, biological properties, etc.) to improve prediction accuracy. In addition, existing models also have deficiencies in interpretability and are difficult to clearly show the basis and logic of prediction results, which to a certain extent limits their promotion in practical applications.
[0006] Generally speaking, existing DDI prediction technologies have constructed rich prediction models from multiple dimensions such as drug molecular characteristics, network structure, graph deep learning, and asymmetry prediction, and achieved certain results. However, there are still deficiencies in aspects such as the integrity of feature extraction, the accuracy of network construction, the interpretability of the model, data dependence, and the effectiveness of multi-modal data fusion. Therefore, developing a new method that can effectively capture DDI asymmetry, make full use of multi-modal features, and improve prediction accuracy and interpretability has important practical significance for promoting the development of DDI prediction technology.
[0007] In view of this, this application is proposed. Summary of the Invention
[0008] The present invention provides an asymmetric drug-drug interaction prediction method based on a diffusion graph attention network, which can at least partially improve the above problems.
[0009] To achieve the above object, the present invention adopts the following technical solutions: An asymmetric drug-drug interaction prediction method based on a diffusion graph attention network, which includes: Obtain the drug information to be predicted, collect relevant drug information and asymmetric DDI records from the preset DrugBank database based on the drug information to be predicted, and create a drug molecular similarity matrix and a directed DDI matrix according to the relevant drug information and asymmetric DDI records; Based on the drug molecular similarity matrix and the directed DDI matrix, combine a bidirectional graph attention network with a multi-head attention mechanism for extraction processing, and extract node features from the drug effect application view and the drug effect receiving view respectively; Use a preset diffusion model based on asymmetry to perform diffusion processing on the node features respectively to obtain the restored graph features. Among them, the diffusion processing includes: forward diffusion, introducing edge structure noise to simulate data degradation, and reverse diffusion, using a graph convolutional network to remove noise; Combine the restored graph features with the drug molecular similarity matrix to obtain the attribute features of the drug, and use the predictor module of the preset DiffGAT-DDI model to integrate and predict the multi-modal features of the drug to obtain the DDI probability prediction value.
[0010] In summary, the proposed method for predicting asymmetric drug-drug interactions based on the diffusion graph attention network realizes efficient modeling and accurate prediction of DDI asymmetry by integrating advanced graph attention mechanisms and diffusion models. First, the framework extracts Morgan fingerprints based on the chemical molecular structures of drugs and constructs a directed DDI network, leveraging the rich topological structure information in the network to provide context support for DDI prediction. This chemical structure-based feature extraction method ensures that we can capture the key characteristics of drug molecules, which are crucial for understanding drug interactions. Second, the model introduces the multi-head attention mechanism and learns the dual-view representations of drugs through an extended bidirectional graph attention network, modeling drug features from both the drug-effect receiving view and the drug-effect exerting view. This approach not only enhances the model's understanding of drugs but also improves its comprehension of complex interaction patterns through dual-view representations. Then, an asymmetry-based diffusion model is adopted to introduce edge structure noise by randomly deleting or adding edges in the graph, enhancing the model's adaptability to sparse graph structures. Subsequently, the reverse diffusion process gradually removes the noise and restores the original features of the drugs. Through this two-way mechanism, the model can more effectively capture the complex interaction patterns between drugs. Finally, the multi-modal embeddings are input into a deep neural network for feature fusion to accurately integrate the multi-modal features of drugs. This multi-modal feature integration method further improves the model's prediction ability, enabling it to perform more accurate DDI predictions considering multiple drug characteristics comprehensively.
[0011] Specifically, the core of this method lies in its unique bidirectional modeling strategy and feature extraction mechanism. By extracting features from both the drug-exerting and drug-receiving perspectives, the model can accurately capture the complex asymmetric relationships between drugs, providing richer semantic information for DDI prediction. In addition, the introduction of the diffusion model further optimizes the feature extraction process. By simulating data degradation and restoration, it significantly improves the model's adaptability to sparse data and the robustness of feature extraction. This innovative feature extraction method not only enhances the model's ability to handle complex drug molecular structures but also provides a more accurate feature basis for DDI prediction.
[0012] In terms of the model architecture, this method uses a deep neural network to fuse multi-modal features, achieving comprehensive consideration of multiple drug characteristics. This fusion mechanism not only improves the prediction accuracy of the model but also enhances its detection performance on class-imbalanced datasets. Compared with existing technologies, this method shows significant advantages in terms of prediction accuracy, model robustness, and dependence on labeled data. Especially in cold-start scenarios, the model can effectively predict the interactions between new drugs or drugs with limited information, greatly expanding its application scenarios.
[0013] The innovation of this method is not only reflected in the technical level, but also lies in its profound significance for practical applications. By accurately predicting the interactions between drugs, it can provide important decision-making support for clinicians, help optimize the combination drug therapy plan, and reduce the risk of adverse drug reactions. At the same time, its efficient feature extraction and prediction capabilities also provide a powerful tool for drug research and development, contributing to accelerating the new drug launch process and reducing the R & D cost. In addition, the design of this method fully considers the interpretability of the model, providing a new perspective for researchers to deeply understand the mechanism of drug interactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a schematic flowchart of an asymmetric drug-drug interaction prediction method based on a diffusion graph attention network provided by an embodiment of the present invention; Figure 2 is an overall framework diagram of a DiffGAT-DDI model provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a topological feature encoding module provided by an embodiment of the present invention; Figure 4 is a schematic flowchart of the working process of Diffusion Models provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] Refer to Figure 1 and Figure 2 As shown, the first embodiment of the present invention discloses an asymmetric drug-drug interaction prediction method based on a diffusion graph attention network, which can be executed by an asymmetric drug-drug interaction prediction device based on a diffusion graph attention network (hereinafter referred to as the prediction device), and particularly, executed by one or more processors in the prediction device to implement the following method: S1. Obtain the drug information to be predicted, collect relevant drug information and asymmetric DDI records from a preset DrugBank database based on the drug information to be predicted, and create a drug molecular similarity matrix and a directed DDI matrix according to the relevant drug information and asymmetric DDI records; Specifically, step S1 further includes: collecting relevant drug information from a preset DrugBank database according to the drug information to be predicted, and calculating the Tanimoto coefficient of the Morgan fingerprints between each pair of drugs using Morgan fingerprints according to the drug information to be predicted and the relevant drug information to generate a symmetric drug molecular similarity matrix , where is the number of drugs, is the number of predefined substructures in the Morgan fingerprint, is the set of real numbers, and its matrix elements represent the structural similarity between drug i and drug j; Extract the asymmetric DDI records between drugs from the preset DrugBank database, and construct a directed DDI matrix based on the asymmetric DDI records , where, if the element in the directed DDI matrix indicates that drug i has an interaction with drug j. If the element , there is no interaction between drug i and drug j.
[0017] In this embodiment, in order to construct a network that can accurately reflect the similarity between drug molecules, according to the drug information to be predicted, relevant drug information is extracted from the preset DrugBank database; subsequently, the Tanimoto coefficient of the Morgan fingerprint is calculated for each pair of drugs using the Morgan fingerprint. The Morgan fingerprint is a molecular fingerprint method that can effectively capture molecular structure information. It decomposes the molecular structure into multiple substructural features and encodes them into binary vectors, thereby effectively capturing the structure information of the molecule. The Tanimoto coefficient is used to measure the similarity between two vectors, and its value ranges from 0 to 1. The closer it is to 1, the higher the similarity. In this way, a symmetric drug molecule similarity matrix is generated, and the dimension of the matrix is N×p, where N represents the number of drugs, and each element in the matrix represents the structural similarity between the corresponding drugs. For example, the matrix element S ij represents the structural similarity between drug i and drug j. If S ij is close to 1, it indicates that drug i and drug j are relatively similar in molecular structure, otherwise it means that their structures are quite different.
[0018] Furthermore, in terms of constructing a directed DDI network, extract the asymmetric DDI records between drugs from the preset DrugBank database. These records specify in detail the one-way interaction relationship between drugs, that is, the direction of the influence of one drug on another drug; and construct a directed DDI matrix based on these records. The directed DDI matrix is also an N×N matrix, and the elements in it are used to represent the interaction relationship between drugs. If the element D ij in the matrix is 1, it indicates that drug i has an interaction with drug j, that is, drug i may affect the efficacy of drug j when used in combination with drug j; if the element D ijIf it is 0, it means that there is no interaction between drug i and drug j, that is, they will not affect each other when used in combination. This directed DDI matrix can clearly reflect the one-way interaction relationship between drugs, providing important topological structure information for subsequent DDI prediction.
[0019] Through the operation of step S1 above, this method can obtain relevant drug information and DDI records from the DrugBank database based on the drug information to be predicted, and construct a drug molecular similarity matrix and a directed DDI matrix, which can clearly represent the one-way interaction relationship between drugs. These two matrices provide basic data support for the subsequent training and prediction of the DDI prediction model, enabling the model to better understand and analyze the interaction relationship between drugs, thereby improving the accuracy and reliability of DDI prediction. This process not only provides rich drug molecular structure information for the model, but also lays a foundation for the model to capture the asymmetric interaction relationship between drugs, helping to improve the model's ability to identify complex DDI patterns and providing a solid data foundation and structural support for subsequent prediction analysis.
[0020] Please refer to Figure 3 , S2, based on the drug molecular similarity matrix and the directed DDI matrix, combine the bi-directional graph attention network and the multi-head attention mechanism for extraction processing, and extract node features from the drug effect application view and the drug effect receiving view respectively; Specifically, step S2 further includes: based on the drug molecular similarity matrix and the directed DDI matrix, for a given directed graph , define its out-neighborhood and in-neighborhood of node u, and for each of its perspectives , calculate the corresponding attention coefficient, and its formula is: , where is the drug effect receiving view, is the drug effect application view, is the concatenation operation, is the neighbor of node , is the weight matrix, is the dimension of the output of the GAT layer, is the dimension of drug characteristics, is the learnable attention weight vector, is the non-linear activation function, is the transpose operation, is the initial vector of node , is used to represent the initial vector of node , is the neighbor node Initial vector; Select to discard the self-loop of the central node, and define the transition forms of the drug effect receiving feature and the drug effect exerting feature, and their formulas are respectively: , , where is the attention coefficient between nodes, is the learnable weight matrix for feature transformation in the drug effect receiving view, In the drug effect exerting view, for node and node the attention coefficient between them, is the learnable weight matrix for feature transformation in the drug effect exerting view; Based on the multi-head attention weighted fusion strategy, perform average aggregation substitution splicing operation on the transition form, while maintaining the advantages of multi-view learning, reducing the parameter scale by a preset multiple, and extracting node features, and its formula is: , , where is the number of attention heads.
[0021] In the DDI network, there are significant directional features and asymmetric influence differences between interacting drugs. To effectively model this direction specificity and extract rich topological features, this method designs a bidirectional graph attention network, aggregates neighborhood information from two perspectives of drug effect exerting (Effector) and drug effect receiving (Receptor) respectively, and introduces an attention mechanism. Among them, the core idea of the attention mechanism is to assign different weights to the neighbors of each node, so that the model can more flexibly focus on important neighbor nodes, and then improve the modeling ability of complex relationships, such as Figure 3 shown.
[0022] Specifically, in this embodiment, first, based on the directed graph constructed by the drug molecule similarity matrix and the directed DDI matrix, define the out-neighborhood and in-neighborhood of each node u. The out-neighborhood represents the influence range of node u on other nodes, and the in-neighborhood represents the influence range of other nodes on node u. This definition method enables the model to model the interaction of drugs from two different perspectives - the drug effect exerting view and the drug effect receiving view. The drug effect exerting view focuses on the influence ability of a drug on other drugs, while the drug effect receiving view focuses on the degree to which a drug is affected by other drugs. This dual-view modeling method can effectively capture the complex asymmetric interactions between drugs and provide rich semantic information for subsequent feature extraction.
[0023] For each perspective, the model calculates the corresponding attention coefficients, where the non-linear activation function can adopt LeakyReLU. Through this calculation method, the model can assign different weights to the neighbors of each node, enabling the model to more flexibly focus on important neighbor nodes, thereby enhancing the ability to model complex relationships. The design of this attention mechanism not only enhances the model's ability to identify drug interactions but also improves the model's interpretability, enabling researchers to better understand the internal mechanism of drug interactions.
[0024] Considering that the source drug and the target drug have different influences and action mechanisms during the interaction process, in order to more accurately model this difference, a mechanism for distinguishing node features is specifically designed in the graph attention network. Specifically, during the process of calculating the attention coefficients, this method chooses to discard the self-loop of the central node. A self-loop usually refers to the connection from a node to itself, which is used in some graph neural networks to ensure that the node can aggregate its own feature information. However, in the DDI network, retaining the self-loop may cause the model to be unable to clearly distinguish the source role and the target role of the nodes. Because the existence of the self-loop will cause each node to automatically regard its own features as an important input when aggregating information, thus blurring the interaction relationship between drugs. In addition, discarding the self-loop can also avoid the model's over-reliance on the node's own features, thereby reducing the risk of model overfitting and enhancing the model's generalization ability.
[0025] Next, define the transition form of the pharmacodynamic receiving feature and the pharmacodynamic exerting feature, where the attention coefficient between node and node represents the importance weight of node to node . In this way, the model can extract the feature information of nodes from the pharmacodynamic receiving view and the pharmacodynamic exerting view respectively, further enhancing the model's ability to model drug interactions.
[0026] Finally, to enhance the model's robustness, a multi-head attention weighted fusion strategy is adopted, and the average aggregation is used to replace the concatenation operation for the transition form. Among them, K represents the number of attention heads, which is used to capture features in different subspaces. This design replaces the concatenation operation with average aggregation, while maintaining the advantages of multi-perspective learning, reducing the parameter scale by a preset multiple, and effectively alleviating the oscillation problem in multi-head attention. In this way, the model can better capture the feature information in different subspaces, further improving the model's prediction accuracy and robustness.
[0027] Furthermore, Figure 3 the attention mechanism for node feature update in GAT, where node updates through its neighbor nodes The eigenvectors are linearly transformed and processed by Softmax to calculate the attention coefficients, and then the neighbor features are weighted and summed to generate the new feature representation of the node to effectively capture the complex relationships between nodes in the graph.
[0028] Through the operations in step S2 above, based on the drug molecule similarity matrix and the directed DDI matrix, the bidirectional graph attention network and the multi-head attention mechanism are used to extract node features from the pharmacodynamic application view and the pharmacodynamic reception view respectively. This process not only enhances the model's ability to identify drug interactions, but also improves the model's interpretability and generalization ability, provides more accurate and comprehensive feature information for subsequent DDI prediction, and significantly improves the model's prediction performance.
[0029] S3. Use a preset diffusion model based on asymmetry to perform diffusion processing on the node features respectively to obtain the restored graph features, where the diffusion processing includes: forward diffusion, introducing edge structure noise to simulate data degradation, and reverse diffusion, using a graph convolutional network to remove the noise; Specifically, step S3 further includes: taking the node features as the original data, performing forward diffusion on it, randomly deleting a part of the nodes in the graph, and introducing noise to the connected edges to destroy the structure of the original graph to simulate the data degradation process, where the forward diffusion includes continuously adding noise to it until the noise is completed to obtain the noised data; Among them, the formula for adding noise is: , is the conditional probability distribution, which describes the diffusion process from time step t - 1 to time step t, , are both adjacency matrices, is the Bernoulli distribution, is the hyperparameter, is the noise scheduling parameter, is the preset base probability, is the value of the edge (i, j) at time step t, is the value of the edge (i, j) at time step t - 1.
[0030] After the forward diffusion is completed, reverse diffusion is performed to remove the noise to restore the structure of the original data; Among them, a graph convolutional network is used as the denoising model to aggregate the neighborhood features of the nodes, capture the complex dependence relationships between the nodes, and restore the local and global structures of the graph during the denoising process. The graph convolutional network adopts a double-layer GCN architecture, and the formula for the information propagation mechanism of the first layer of GCN is: , is also the adjacency matrix, is the degree matrix, is the weight matrix of the first-layer GCN; The formula for the information propagation mechanism of the second-layer GCN is: , is the weight matrix of the second-layer GCN; Obtain the mathematical expression of the entire reverse diffusion: , to obtain the restored graph features, where is the mathematical expression of the entire reverse diffusion, is a multivariate normal distribution.
[0031] In the process of drug feature extraction, although traditional multi-layer perceptrons can extract some basic features, they perform poorly in dealing with complex drug molecular structures and multi-dimensional attributes. To improve the accuracy and robustness of feature extraction, the diffusion model has unique advantages in drug feature extraction. The diffusion model can capture subtle patterns and relationships in a complex feature space by gradually denoising and restoring the original features. Its denoising process not only enhances the feature expression ability but also improves the model's ability to handle high-dimensional attributes and complex interactions. In addition, the non-linear modeling ability of the diffusion model enables it to adapt to the complex distribution of drug features, thereby achieving more accurate feature representation, as Figure 4 shown. In this embodiment, first, the node features extracted in step S2 are used as the original data and input into the diffusion model. Combining Figure 4 , it can be seen that for the workflow of Diffusion Models, Figure 4 in represents the original data, and through continuous steps noise is gradually added until is completely noised. Subsequently, the model starts to gradually denoise from the noisy data and attempts to restore to the original data distribution. The forward process follows the noise addition rule, while the reverse process removes noise according to the learning rule, and finally achieves the goal of generating clear data from noise.
[0032] Traditional noise generation methods, such as adding Gaussian noise, are usually applicable to continuous data (e.g., images), but are not ideal for graph data with discrete structures (composed of nodes and edges). This mismatch is mainly because the structural characteristics of graph data are essentially different from those of continuous data. To better simulate the noise process of graph data and improve the generalization ability of the model, this method adopts the method of edge structure noise. In the forward diffusion stage, the model introduces noise by randomly deleting a part of the nodes in the graph and their connected edges, destroying the structure of the original graph, thereby simulating the degradation process of the data. This way of introducing noise can not only prevent the model from overfitting to a specific graph structure, but also significantly improve the generalization ability of the model on sparse graphs or partially observed graphs. In addition, deleting nodes and their edges can significantly reduce the computational complexity, which is especially helpful for processing large-scale sparse graphs. More importantly, the task of this paper focuses on the prediction of drug links, so this way of generating noise can enable the model to show stronger ability in predicting connections.
[0033] The forward diffusion process is a process of gradually adding noise. By continuously adding noise to it until the noise is completed, the noise data is obtained. Among them, the Bernoulli distribution is used to describe the probability distribution of the existence or non-existence of edges under a given noise level; the noise scheduling parameter controls the intensity of noise introduction; the preset base probability is usually set to 0.5, indicating the random noise when there is no information; the value of edge (i, j) at time step t is usually 0 or 1.
[0034] After the forward diffusion is completed, the model enters the reverse diffusion stage, the purpose of which is to remove noise and restore the structure of the original data. The reverse diffusion process uses a graph convolutional network (GCN) as the denoising model. GCN can effectively utilize the topological information of the graph structure; by aggregating the neighborhood features of nodes, it captures the complex dependencies between nodes, thereby restoring the local and global structures of the graph during the denoising process. In addition, the flexibility of GCN enables it to adapt to the structural characteristics of different graph data, significantly improving the generalization ability and performance of the model. At the technical implementation level, GCN performs information propagation through convolutional operations with the adjacency matrix and the node feature matrix.
[0035] In this embodiment, a two-layer GCN architecture is adopted. The information propagation mechanism of the first layer of GCN. This aggregation mechanism enables each node to integrate the information of its neighbors. As the number of network layers increases, the node can capture more global graph structure information. On this basis, the information propagation mechanism of the second layer of GCN is introduced. Through this two-layer GCN architecture, the model can effectively recover the original graph structure from the noise data, and at the same time provide a richer and more accurate feature representation for the DDI prediction task.
[0036] During the entire reverse diffusion process, the multivariate normal distribution is used to describe Probability distribution. Through this reverse diffusion process, the model can gradually remove noise and finally recover the original graph features, providing high-quality feature inputs for subsequent DDI prediction. Based on this multi-layer GCN structure, the model can effectively recover the original graph structure from noisy data and at the same time provide a richer and more accurate feature representation for the DDI prediction task. On this basis, the model further combines the drug similarity matrix and fuses the recovered graph features with the drug similarity information through similar GCN layers to extract the attribute features of the drug. Finally, feature alignment is performed through MLP to ensure that the dimensions of different features are the same.
[0037] Based on this step, this method uses an asymmetry-based diffusion model to perform diffusion processing on node features, significantly enhancing the robustness and accuracy of feature extraction. Forward diffusion simulates data degradation by introducing edge structure noise, while reverse diffusion uses a graph convolutional network to remove noise and recover the structure of the original data. This process not only improves the model's adaptability to sparse graph structures but also enhances the model's ability to capture complex drug interaction patterns, providing a more accurate and reliable feature basis for DDI prediction.
[0038] In step S4, the recovered graph features are combined with the drug molecule similarity matrix to obtain the attribute features of the drug, and the predictor module of the preset DiffGAT-DDI model is used to integrate and predict the multi-modal features of the drug to obtain the DDI probability prediction value.
[0039] Specifically, step S4 further includes: combining the recovered graph features with the drug molecule similarity matrix to extract the attribute features of the drug; Performing feature alignment processing on the extracted attribute features to ensure feature space consistency.
[0040] Using the predictor module of the preset DiffGAT-DDI model to integrate the drug's pharmacodynamic receiving features, pharmacodynamic exerting features, and drug attribute features to obtain the final fused features , where is the drug 's drug attribute feature; Select a deep neural network as the predictor, input the final fused features into the deep neural network, and perform non-linear transformation and pattern recognition on the final fused features through its multi-layer neural network structure to output the probability prediction value of DDI occurring between drug and drug ; Among them, the mathematical expression of the prediction process is: , is the activation function, is the deep neural network, is the final fused feature, Represents the probability of DDI prediction.
[0041] Feature fusion is performed through the inner product operation between features. However, experiments show that although the inner product operation is computationally efficient, it is difficult to capture the non-linear interaction patterns between drugs (such as synergistic enhancement or antagonistic inhibition of drug effects) due to its linear compression characteristics. In contrast, directly concatenating the original embedding vectors can completely preserve the independence and integrity of each feature, and flexibly learn the weight allocation through non-linear activation functions (such as ReLU, LeakyReLU) in the deep neural network, significantly improving the model's ability to model complex directional dependence relationships. Further, repeating the drug attribute features to form a four-vector concatenation can further improve the performance. Strengthen the inherent attributes of the drug attribute features, enhance their weights in the feature space by repeating the drug attribute features, and force the network to pay more attention to the contributions of key chemical attributes. For example, static information such as the molecular structure and chemical fingerprint of a drug is amplified through feature redundancy, while implicit regularization alleviates the risk of overfitting.
[0042] Specifically, in this embodiment, the restored graph features are combined with the generated drug molecule similarity matrix. The purpose of this combination process is to fully utilize the molecular structure information of the drug and the graph feature information restored by the diffusion model, so as to extract the attribute features that can comprehensively reflect the drug characteristics. The attribute features of the drug are the embodiment of its inherent characteristics, and contain multi-faceted information such as the chemistry and biology of the drug. By combining the restored graph features with the drug molecule similarity matrix, the model can more accurately capture the complex interaction patterns between drugs, providing a richer feature basis for subsequent predictions.
[0043] Next, perform feature alignment processing on the extracted attribute features to ensure the consistency of the feature space. Feature alignment is an important step in data preprocessing. It can ensure the consistency of features from different sources in terms of dimension and scale, thereby improving the training efficiency and prediction performance of the model. Through feature alignment, the model can better integrate multi-modal features and avoid prediction biases caused by inconsistent features.
[0044] Subsequently, use the predictor module of the preset DiffGAT-DDI model to integrate the drug's pharmacodynamic receiving features, pharmacodynamic exerting features, and drug attribute features. This integration process is achieved by concatenating or fusing the above three features. Specifically, the model concatenates the pharmacodynamic receiving features, pharmacodynamic exerting features, and drug attribute features to form the final fused feature. This fusion method can preserve the independence and integrity of each feature, and at the same time flexibly learn the weight allocation through non-linear activation functions in the deep neural network, thus significantly improving the model's ability to model complex directional dependence relationships.
[0045] Such as Figure 3As shown, in order to predict whether there is an interaction between drugs and drugs A deep neural network is selected as the prediction factor, and the final fused features are input into the deep neural network; for binary DDI prediction, the output neuron of the DNN is set to 1. The deep neural network performs complex non-linear transformations and pattern recognition on the final fused features through its multi-layer neural network structure, and finally outputs the probability prediction value of DDI occurring between drugs.
[0046] Based on this, this method can effectively combine the recovered graph features with the drug molecule similarity matrix, extract the attribute features of drugs, and ensure the consistency of the feature space through feature alignment processing. Finally, the predictor module of the preset DiffGAT-DDI model is used to integrate and predict the multi-modal features of drugs to obtain the DDI probability prediction value. This process not only makes full use of the multi-modal features of drugs, but also further improves the accuracy and reliability of the prediction through the non-linear transformation ability of the deep neural network. In addition, this integrated prediction method also enhances the generalization ability of the model, enabling it to maintain stable prediction performance in different data sets and application scenarios. Preferably, it further includes: in the training stage of the DiffGAT-DDI model, the mean squared error loss is used as the loss function to calculate the average of the sum of squares of the differences between the predicted value and the true value, and its formula is: , where is the model prediction function, is the data set, is the number of samples, is the true value of the i-th sample, is the predicted value of the model for the i-th sample; The binary cross-entropy loss is used as the loss function to measure the gap between the probability distribution predicted by the model and the true label, and its formula is: .
[0047] In this embodiment, in the training stage of the DiffGAT-DDI model, the mean squared error loss (MSE, Mean Squared Error) is first used as one of the loss functions. The mean squared error loss measures the prediction error of the model by calculating the average of the sum of squares of the differences between the predicted value and the true value. The characteristics of the mean squared error loss are simple and intuitive, efficient in calculation, and the square operation on the error can amplify larger errors, making the model pay more attention to those parts with larger prediction deviations during the optimization process. This loss function is particularly suitable for regression tasks and can effectively guide the model to learn the main features and laws in the data, thereby improving the prediction accuracy of the model.
[0048] In addition, to better adapt to the binary classification task of DDI prediction, binary cross-entropy loss (BCE) is also used as the loss function. Binary cross-entropy loss is suitable for binary classification tasks and can measure the gap between the probability distribution predicted by the model and the true labels. Binary cross-entropy loss is particularly suitable for classification tasks and can effectively measure the accuracy and reliability of model predictions. By minimizing the binary cross-entropy loss, the model can better learn the classification boundaries in the data, thereby improving the prediction ability for DDI. This loss function can not only effectively handle the problem of class imbalance but also enhance the model's ability to identify positive samples, further improving the model's prediction performance.
[0049] During the training process, combining mean squared error loss and binary cross-entropy loss can give full play to the advantages of both loss functions. Mean squared error loss can quickly guide the model to learn the main features in the data, while binary cross-entropy loss can further optimize the model's ability to identify classification boundaries. Through this combined use, the model can not only converge quickly but also gradually improve its learning ability for complex drug interaction patterns during the training process, thus showing higher accuracy and reliability in the final prediction task.
[0050] Through the above optimization measures, the proposed method introduces mean squared error loss and binary cross-entropy loss in the training stage of the DiffGAT-DDI model, significantly improving the training effect and prediction performance of the model. This way of combining loss functions not only improves the convergence speed of the model but also enhances the model's learning ability for complex drug interaction patterns, providing more accurate and reliable model support for DDI prediction.
[0051] In summary, the asymmetric drug interaction prediction method based on the diffusion graph attention network aims to solve the problems in the existing technology, such as insufficient modeling of DDI asymmetry, incomplete feature extraction, and over-reliance on labeled data. By innovatively combining the graph attention mechanism, diffusion model, and deep neural network, it realizes the efficient prediction and accurate modeling of DDI.
[0052] Specifically, first, relevant drug information and asymmetric DDI records are collected from the preset DrugBank database, and a drug molecular similarity matrix and a directed DDI matrix are constructed based on this information. This process provides a rich data foundation for subsequent feature extraction and model training, ensuring that the model can learn based on accurate drug molecular structure information and interaction relationships. Next, a bidirectional graph attention network (GAT) combined with a multi-head attention mechanism is used to extract node features from the drug effect application view and the drug effect receiving view respectively. Through this dual-view modeling method, the model can accurately capture the asymmetric interactions between drugs, significantly enhancing the ability to identify the directionality of DDI. In addition, a diffusion model is introduced to perform diffusion processing on the node features. By introducing edge structure noise through forward diffusion to simulate data degradation, and then using a graph convolutional network (GCN) to remove the noise through reverse diffusion to restore the features of the original graph. This process not only improves the adaptability of the model to sparse graph structures, but also enhances the robustness and accuracy of feature extraction, enabling the model to more effectively capture the complex interaction patterns between drugs.
[0053] In the feature integration stage, the restored graph features are combined with the drug molecular similarity matrix to extract the attribute features of the drugs, and feature alignment processing is performed to ensure the consistency of the feature space. Finally, the predictor module of the preset DiffGAT-DDI model is used to integrate and predict the multi-modal features of the drugs to obtain the DDI probability prediction value. This process makes full use of the multi-modal features of the drugs and further improves the accuracy and reliability of the prediction through the non-linear transformation ability of the deep neural network.
[0054] In the model training stage, the mean squared error loss (MSE) and the binary cross-entropy loss (BCE) are used as loss functions in combination. The mean squared error loss can quickly guide the model to learn the main features in the data, while the binary cross-entropy loss can further optimize the model's ability to identify the classification boundary. Through this combined use method, the model can not only converge quickly, but also gradually improve its ability to learn complex drug interaction patterns during the training process, thus showing higher accuracy and reliability in the final prediction task.
[0055] Compared with the existing technology, the beneficial effects of this method are significant. By accurately modeling the asymmetric interactions between drugs, the model can provide an important reference for the order of administration in drug combination therapy and reduce the risk of treatment failure caused by adverse interactions. In addition, the introduction of the diffusion model optimizes the feature extraction process and improves the model's ability to model complex drug molecular structures and high-dimensional properties, enabling the model to make more accurate DDI predictions while comprehensively considering multiple characteristics of the drug. At the same time, the system only requires basic drug chemical structure information to achieve high-performance predictions, has low requirements for data availability and adaptability, and has wide applicability and promotion value. In addition, the introduction of the multi-head attention mechanism enhances the interpretability of the model, enabling researchers to better understand the intrinsic mechanisms and influencing factors of drug interactions. The fusion of multimodal features by deep neural networks further improves the generalization ability of the model, enabling it to comprehensively consider multiple characteristics and interaction patterns of drugs, further improving the accuracy and robustness of DDI predictions.
[0056] In general, this method has brought breakthrough progress in the field of DDI prediction through its unique technical architecture and innovative feature extraction mechanism. Its high efficiency, accuracy and wide applicability make it an important tool for future drug interaction research and clinical application, providing strong support for precision medicine and drug development.
[0057] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. An asymmetric drug interaction prediction method based on a diffusion graph attention network, characterized in that Including: Obtain the drug information to be predicted, collect relevant drug information and asymmetric DDI records from the preset DrugBank database based on the drug information to be predicted, and create a drug molecular similarity matrix and a directed DDI matrix according to the relevant drug information and asymmetric DDI records; Based on the drug molecular similarity matrix and the directed DDI matrix, combine the bidirectional graph attention network and the multi-head attention mechanism for extraction processing, and extract node features from the drug effect application view and the drug effect receiving view respectively; Use the preset diffusion model based on asymmetry to perform diffusion processing on the node features respectively to obtain the restored graph features, where the diffusion processing includes: forward diffusion, introducing edge structure noise to simulate data degradation, and reverse diffusion, using a graph convolutional network to remove noise; Combine the restored graph features with the drug molecular similarity matrix to obtain the attribute features of the drug, and use the predictor module of the preset DiffGAT-DDI model to integrate and predict the multi-modal features of the drug to obtain the DDI probability prediction value.
2. The asymmetric drug-drug interaction prediction method based on a diffusion graph attention network according to claim 1, wherein Obtain the drug information to be predicted, collect relevant drug information and asymmetric DDI records from the preset DrugBank database based on the drug information to be predicted, and create a drug molecular similarity matrix and a directed DDI matrix, specifically: Collect relevant drug information from the pre-set DrugBank database according to the drug information to be predicted. Calculate the Tanimoto coefficient of the Morgan fingerprints between each pair of drugs based on the drug information to be predicted and the relevant drug information, and generate a symmetric drug molecule similarity matrix , where is the number of drugs, is the number of predefined substructures in the Morgan fingerprint, is the set of real numbers, and its matrix element represents the structural similarity between drug i and drug j; Extract the asymmetric DDI records between drugs from the preset DrugBank database, and construct a directed DDI matrix based on the asymmetric DDI records , where, if the element in the directed DDI matrix , it indicates that drug i has an interaction with drug j. If the element , there is no interaction between drug i and drug j.
3. The method for predicting asymmetric drug-drug interactions based on a diffusion graph attention network according to claim 1, wherein Based on the drug molecular similarity matrix and the directed DDI matrix, combine the bidirectional graph attention network and the multi-head attention mechanism for extraction processing, and extract node features from the drug effect application view and the drug effect receiving view respectively, specifically: Based on the drug molecule similarity matrix and the directed DDI matrix, for a given directed graph , define the out-neighborhood and the in-neighborhood of its node u, and for each perspective of it, calculate the corresponding attention coefficient, and its formula is: , where is the pharmacodynamic reception view, is the pharmacodynamic application view, is the concatenation operation, is the neighbor of node , is the weight matrix, is the dimension of the output of the GAT layer, is the dimension of the drug property, is the learnable attention weight vector, is the non-linear activation function, is the transpose operation, is the initial vector of node , is the initial vector representing node , is the initial vector of neighbor node ; Selectively discard the self-loops of the central nodes, and define the transition forms of the drug effect receiving features and the drug effect applying features, and their formulas are respectively: , , where is the attention coefficient between nodes, is the learnable weight matrix for feature transformation in the drug effect receiving view, in the drug effect applying view, for node and node the attention coefficient between them, is the learnable weight matrix for feature transformation in the drug effect applying view; Based on the multi-head attention weighted fusion strategy, the average aggregation of the transition forms is used to replace the splicing operation. While maintaining the advantages of multi-perspective learning, the parameter scale is reduced by a preset multiple, and the node features are extracted. The formula is as follows: , , where is the number of attention heads.
4. The method for predicting asymmetric drug-drug interactions based on a diffusion graph attention network according to claim 3, wherein Use the preset diffusion model based on asymmetry to perform diffusion processing on the node features respectively to obtain the restored graph features, specifically: Take the node features as the original data, perform forward diffusion on it, randomly delete a part of the nodes in the graph, and introduce noise to the edges connected to them to destroy the structure of the original graph to simulate the data degradation process, where the forward diffusion includes continuously adding noise to it until the noiseization is completed to obtain the noiseized data; Among them, the formula for adding noise is: , is the conditional probability distribution, which describes the diffusion process from time step t - 1 to time step t, 、 are both adjacency matrices, is the Bernoulli distribution, is a hyperparameter, is the noise scheduling parameter, is the preset base probability, is the value of edge (i, j) at time step t, is the value of edge (i, j) at time step t - 1.
5. The method for predicting asymmetric drug-drug interactions based on a diffusion graph attention network according to claim 4, wherein Also including: After the forward diffusion is completed, perform reverse diffusion to remove the noise to restore the structure of the original data; Among them, a graph convolutional network is used as the denoising model to aggregate the neighborhood features of nodes, capture the complex dependencies between nodes, and recover the local and global structures of the graph during the denoising process. The graph convolutional network adopts a two-layer GCN architecture. The formula for the information propagation mechanism of the first layer of GCN is: , which is also the adjacency matrix, is the degree matrix, and is the weight matrix of the first layer of GCN; The formula for the information propagation mechanism of the second-layer GCN is as follows: , is the weight matrix of the second-layer GCN; Obtain the mathematical expression for the entire reverse diffusion: , and obtain the restored graph features, where is the mathematical expression for the entire reverse diffusion, is a multivariate normal distribution.
6. The asymmetric drug-drug interaction prediction method based on a diffusion graph attention network according to claim 1, wherein Combine the restored graph features with the drug molecular similarity matrix to obtain the attribute features of the drug, specifically: Combine the restored graph features with the drug molecular similarity matrix, and extract the attribute features of the drug from it; Perform feature alignment processing on the extracted attribute features to ensure the consistency of the feature space.
7. The method for predicting asymmetric drug-drug interactions based on a diffusion graph attention network according to claim 3, wherein Use the predictor module of the preset DiffGAT-DDI model to integrate and predict the multi-modal features of the drug to obtain the DDI probability prediction value, specifically: The predictor module of the preset DiffGAT-DDI model integrates the drug's pharmacodynamic receiving characteristics, pharmacodynamic exerting characteristics, and drug attribute characteristics to obtain the final integrated characteristics , where is the drug 's drug attribute characteristics; Select a deep neural network as the predictor, input the final fused features into the deep neural network, perform non-linear transformation and pattern recognition on the final fused features through its multi-layer neural network structure, and output the predicted value of the probability of DDI occurring between drug and drug ; Among them, the mathematical expression of the prediction process is: , is the activation function, is the deep neural network, is the final fused feature, represents the probability of DDI prediction.
8. The method for predicting asymmetric drug-drug interactions based on a diffusion graph attention network according to claim 1, wherein Also including: In the training phase of the DiffGAT-DDI model, the mean squared error loss is used as the loss function to calculate the average of the sum of squares of the differences between the predicted values and the true values. The formula is as follows: , where is the model prediction function, is the dataset, is the number of samples, is the true value of the i-th sample, is the predicted value of the model for the i-th sample; The binary cross-entropy loss is used as the loss function to measure the gap between the probability distribution predicted by the model and the true label, and its formula is: .
Citation Information
Patent Citations
Asymmetric drug interaction prediction method and system and storage medium
CN117976245A
Drug interaction event prediction method based on multi-modal diagram diffusion static subgraph
CN119517201A
Drug interaction prediction method based on double-view graph neural network
CN119694442A
Drug interaction prediction method based on interaction enhancement graph self-attention mechanism
CN120072354A
Cited By
Continuous time dynamics prediction method and system for fusing diffusion model and Figure ordinary differential equation, terminal and medium
CN121389530A
Drug interaction prediction method based on bidirectional cross-view attention network
CN122091273A
A drug interaction prediction method based on a bidirectional cross-view attention network
CN122091273B