A drug interaction prediction method based on a bidirectional cross-view attention network
Patent Information
- Application Number
- CN202610550041.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-04-24
AI Technical Summary
不同视图在语义粒度、噪声分布与信息密度上存在显著差异,简单融合容易引入冗余与冲突信息,削弱关键判别特征
[0012]Compared with existing technologies, this method solves the following problems: (1) how to construct multi-scale topological views to explicitly characterize the multi-hop dependencies and high-order neighborhood structures between drugs, and alleviate the sparsity problem of the original DDI network; (2) how to effectively compress and enhance drug similarity features to avoid interference from high-dimensional redundant information on learning; (3) how to design a fine-grained cross-view interaction mechanism to achieve bidirectional complementarity and alignment between structural views and attribute views, rather than simple post-fusion; (4) how to maintain the key semantics of each view and stabilize gradient propagation during the multi-view representation learning process to avoid over-smoothing.
Smart Images

Figure CN122091273B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drug action prediction technology, and more specifically to a drug interaction prediction method based on a bidirectional cross-perspective attention network. Background Technology
[0002] Drug-drug interactions (DDIs) refer to the phenomenon where the simultaneous or sequential use of two or more drugs alters their efficacy or toxicity due to changes in pharmacodynamics or pharmacokinetics. Beneficial DDIs can enhance therapeutic effects through synergistic effects; conversely, harmful DDIs may weaken therapeutic effects, cause serious adverse reactions, or even endanger the patient's life. For example, the combined use of the anticoagulant warfarin and the nonsteroidal anti-inflammatory drug ibuprofen significantly enhances anticoagulation and increases the risk of bleeding. With the accelerating aging of the population and the increasing prevalence of chronic disease comorbidities, multidrug combination therapy is becoming more common in clinical practice. Accurate identification and risk prediction of DDIs have become crucial for drug safety assessment and rational drug use in clinical practice. Traditional DDI detection relies primarily on clinical trials, post-marketing surveillance, and pharmacovigilance systems. These methods are not only time-consuming, costly, and have limited sample sizes, but also struggle to comprehensively cover potential drug combinations, failing to meet the need for large-scale, high-efficiency DDI screening. Therefore, developing efficient and scalable computational methods to achieve accurate prediction of unknown DDI has significant research value and clinical application prospects.
[0003] In the interdisciplinary field of artificial intelligence and bioinformatics, existing drug interaction identification (DDI) prediction methods can be mainly divided into three categories. The first category is similarity-based methods, whose core assumption is that "structurally similar drugs are more likely to exhibit similar interaction patterns." These methods extract multi-source attribute features of drugs, such as chemical structure, targets, enzymes, and pathways, calculate the similarity between drugs (e.g., Tanimoto coefficient, Dice coefficient, etc.), and then combine them with traditional machine learning models for prediction. Typical works include using drug fingerprint vectors to calculate similarity and generate DDI candidate pairs, or integrating multiple similarity networks such as chemical structure, targets, enzymes, pathways, and ATC encoding for high-order similarity mining. This type of method is simple to implement and highly interpretable, but it usually relies only on the similarity of the properties of the drugs themselves, ignoring the rich topological information contained in the known DDI networks.
[0004] The second category is network-based methods, which treat drugs as nodes and known interactions as edges, modeling DDI prediction as a link prediction problem on a graph, and using graph neural networks (GNNs) or graph representation learning techniques to encode drug nodes. For example, some studies transform DDI prediction into link prediction and then use graph representation learning frameworks to extract node embeddings, or use graph autoencoders to perform end-to-end predictions in heterogeneous networks such as drug-protein networks. This type of method can aggregate neighborhood information from drug interaction topology, significantly enhancing the ability to characterize topological semantics. However, most methods only model on the original interaction graph at a single scale, making it difficult to simultaneously cover low-order direct interactions and high-order structural dependencies. Furthermore, when the DDI network is sparse, oversmoothing can easily occur, leading to convergence in the representation of different drug nodes, affecting the model's generalization performance on low-degree nodes or rare drugs.
[0005] The third category is based on multi-feature fusion methods, which attempt to integrate complementary information such as structural attributes, sequence information, network topology, and knowledge graphs to improve prediction performance. Existing fusion strategies often employ coarse-grained methods such as feature concatenation or weighted averaging, failing to deeply explore the complementarity and consistency between different views. Different views exhibit significant differences in semantic granularity, noise distribution, and information density; simple fusion easily introduces redundant and conflicting information, weakening key discriminative features. Furthermore, existing attention or alignment mechanisms are mostly unidirectional information injections, making it difficult to explicitly model the bidirectional dependency relationship of "mutual completion between views," thus limiting the consistency and upper bound of multi-view representations. Meanwhile, similarity-based methods often face problems of high dimensionality, high noise, and information redundancy when constructing drug similarity matrices, lacking effective nonlinear compression and denoising mechanisms.
[0006] In view of the above, this application is hereby submitted. Summary of the Invention
[0007] This invention provides a drug interaction prediction method based on a bidirectional cross-perspective attention network, which can at least partially improve the above-mentioned problems.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A drug interaction prediction method based on a bidirectional cross-perspective attention network, comprising: The original DDI map was constructed based on drug data, and the information enhancement processing of the original DDI map was performed using PPR multi-scale diffusion to obtain a multi-scale diffusion map. A symmetric similarity matrix between drugs is constructed based on drug data. The symmetric similarity matrix is used as a weighted adjacency matrix to construct a drug similarity graph. A pre-set multilayer perceptron encoder is used to perform a nonlinear transformation on the symmetric similarity matrix to obtain the corresponding multi-scale feature representation. The drug similarity map, the original DDI map, and the multi-scale diffusion map are input into a pre-defined shared weight GCN network for encoding, resulting in topological embedding, diffusion embedding, and similarity embedding, respectively. Additive fusion of topological embedding and diffusion embedding is performed to obtain structural query representation. Cross-attention preprocessing is performed based on structural query representation and similarity embedding to obtain unified drug embedding. Based on unified drug embedding, binary classification prediction is performed on drug pairs to be predicted to obtain prediction probabilities.
[0010] In summary, to address the problems in existing technologies such as similarity-based methods ignoring the topological associations of DDI networks, graph learning-based methods failing to capture high-order interactions leading to incomplete features, the inability to mine feature complementarity through simple concatenation or weighted averaging in multi-source feature fusion, and the tendency of deep GNNs to suffer from convergent node embeddings, weak adaptability to sparse networks leading to oversmoothing and poor generalization ability, this invention aims to provide a drug interaction prediction method that can integrate multi-perspective features, capture high-order interaction topology, achieve deep feature fusion, and improve prediction generalization ability.
[0011] Specifically, this invention constructs a collaborative modeling framework of "similarity perspective, topology perspective, and diffusion perspective," integrating drug molecular structure features with multi-scale network topology features to achieve high-precision prediction of drug interactions. First, a multi-view feature extraction module is constructed: the Morgan fingerprint perspective calculates the chemical structural similarity between drugs using Dice similarity, constructs a similarity matrix, and utilizes AutoEncoder for nonlinear dimensionality reduction and reconstruction learning, combined with the KNN algorithm to construct a similarity network graph to capture molecular structural similarity; the multi-scale diffusion perspective, based on the original DDI data, uses a personalized PageRank (PPR) algorithm to generate multiple diffusion maps at different diffusion scales, mining long-distance associations and high-order topological information of non-adjacent nodes. Second, a shared-weight multilayer graph convolutional network (GCN) is used to collaboratively encode the similarity map, the original DDI map, and the multi-scale diffusion map, achieving alignment of different view representation spaces through a parameter sharing mechanism to generate multi-dimensional embedded features. Finally, a bidirectional cross-view attention mechanism is used to fuse multi-view embeddings. This mechanism includes two symmetrical and complementary interaction directions: using the original and diffuse structural representations as queries to selectively aggregate similarity information (structure to attribute), while using similarity representations as queries to reversely aggregate structural semantics (attribute to structure), thus achieving bidirectional completion and collaborative enhancement between views. The fused features are then fed into a multilayer perceptron after residual connections and global reweighting to complete the binary classification prediction of drug interactions. The model is then jointly optimized using autoencoder reconstruction loss and cross-entropy loss.
[0012] Compared with existing technologies, this method solves the following problems: (1) how to construct multi-scale topological views to explicitly characterize the multi-hop dependencies and high-order neighborhood structures between drugs, and alleviate the sparsity problem of the original DDI network; (2) how to effectively compress and enhance drug similarity features to avoid interference from high-dimensional redundant information on learning; (3) how to design a fine-grained cross-view interaction mechanism to achieve bidirectional complementarity and alignment between structural views and attribute views, rather than simple post-fusion; (4) how to maintain the key semantics of each view and stabilize gradient propagation during the multi-view representation learning process to avoid over-smoothing. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating the drug interaction prediction method based on a bidirectional cross-perspective attention network provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of the framework of the drug interaction prediction method based on a bidirectional cross-view attention network provided in an embodiment of the present invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0016] refer to Figure 1 , Figure 2 As shown, the first embodiment of the present invention discloses a drug interaction prediction method based on a bidirectional cross-perspective attention network, which can be executed by a drug interaction prediction device based on a bidirectional cross-perspective attention network (hereinafter referred to as the prediction device), specifically, by one or more processors within the prediction device, to implement the following method: S1. Construct the original DDI map based on drug data, and use PPR multi-scale diffusion to enhance the information of the original DDI map to obtain a multi-scale diffusion map. Specifically, step S1 further includes: representing the original drug interaction data in the drug data as a DDI original graph. and will As the adjacency matrix of the original graph of DDI, where, N represents the total number of drugs, V is the set of drug nodes, and E is the set of edges representing known drug interactions. A is a set of real numbers; since A only describes first-order interactions and is usually sparse, it is difficult to fully characterize the multi-order neighbors and long-range topological associations on which potential DDIs depend.
[0017] Multi-scale diffusion using PPR is employed to enhance the information of the original DDI graph. Specifically, the adjacency matrix of the original DDI graph is normalized to obtain a normalized propagation matrix. Based on the normalized propagation matrix, multiple scale parameters are obtained. PPR diffusion matrix Its formula is: , , , I is the identity matrix, and k is the scale parameter index. The PPR algorithm can effectively balance local and global structural information through a personalized random walk mechanism, while providing a controllable diffusion range. Compared with other diffusion methods, it is more suitable for the sparsity characteristics of drug interaction networks. Used to control the diffusion range: smaller (e.g., 0.1) emphasizes the interaction patterns of the local neighborhood, while larger values... (e.g., 0.5) can capture a more global community structure and long-range dependencies.
[0018] Based on multiple PPR diffusion matrices, a corresponding set of multi-scale diffusion matrices is generated. , For the k-th scale parameter The PPR diffusion matrix is given by K, where K is the number of scale parameters; for example, K=3 corresponds to the scale set {0.1,0.3,0.5}.
[0019] By treating each PPR diffusion matrix as a weighted diffusion view, a multi-scale diffusion map is constructed. The multi-scale diffusion matrix set is a weighted adjacency matrix representation of the multi-scale diffusion graph. This multi-scale diffusion graph set can inject multi-level neighborhood information at different structural scales, effectively alleviating the sparsity of the DDI network and providing richer multi-scale topological representations for subsequent feature fusion and DDI prediction. In this embodiment, a pre-trained BiCrossNet-DDI model is used to process the drug data to be predicted, and the prediction results can be obtained. A DDI original graph and multi-scale diffusion graph construction module is employed. This module is used to realize multi-scale topology enhancement of the drug interaction network and is a fundamental step in the high-order structure modeling of the entire system. Specifically, based on known drug interaction data, a basic "DDI original graph" is first constructed from all drugs and their interaction relationships. Drugs are treated as nodes in the graph, and known drug interactions are treated as edges connecting the nodes. Since this original graph is usually sparse and can only reflect first-order direct interactions between drugs, it is difficult to capture deeper and longer-term potential associations. Therefore, this invention further enhances its information by employing a personalized PageRank (PPR) multi-scale diffusion algorithm. This algorithm, by simulating multi-step random walks on the original drug relationship network, can diffuse the influence of each drug to its multi-hop neighbors, thereby generating a series of "multi-scale diffusion graphs" with different diffusion ranges (i.e., different scales). By selecting multiple different diffusion parameters (e.g., three scales), this step can simultaneously obtain multi-level topological structure information from local to global, providing the model with richer network neighborhood features and effectively alleviating the learning difficulties caused by the sparsity of the original data.
[0020] S2. Based on drug data, a symmetric similarity matrix between drugs is constructed. The symmetric similarity matrix is used as a weighted adjacency matrix to construct a drug similarity graph. A pre-set multilayer perceptron encoder is used to perform a nonlinear transformation on the symmetric similarity matrix to obtain the corresponding multi-scale feature representation. Specifically, step S2 further includes: using the RDKit toolkit to parse the SMILES string of each drug in the drug data into a molecular structure, and generating a 2048-bit binary vector using a Morgan circular fingerprint with a radius of 2 to obtain the Morgan fingerprint of the drug; this is used to encode the local substructure information of the drug molecule. The Morgan fingerprint, through its circular encoding mechanism, can effectively capture substructure information at different distances in the molecule, and compared with other molecular descriptors, it has stronger structural discrimination ability and richer chemical semantic expression ability.
[0021] The Dice coefficient is used to measure the chemical similarity between drug i and drug j, and a symmetric similarity matrix between drug i and drug j is constructed. , , , Morgan fingerprint of drug i Morgan fingerprint of drug j; The number of fingerprint values of 1 (i.e., the total number of feature substructures contained in drug i). similar), This represents the number of times that the fingerprints of the two drugs are both 1 at the same position (i.e., the number of common feature substructures). The larger the value, the more similar the chemical structures of the two drugs are.
[0022] A drug similarity graph is constructed by using the symmetric similarity matrix as a weighted adjacency matrix. , Drug similarity map The set of edges; To further enhance node representation and alleviate high-dimensional redundancy, a pre-defined multilayer perceptron encoder (such as an AutoEncoder structure) is used to perform a nonlinear transformation on the symmetric similarity matrix. The multilayer outputs of the multilayer perceptron encoder are cascaded and fused through the channel dimension to form a multi-scale feature representation of nodes in the drug similarity map. , , , Furthermore, multi-scale features are represented as node features in the drug similarity map, where h is the feature dimension, representing the dimensionality of each layer of node features. For splicing operations, For activation function, For batch normalization, For random deactivation regularization (dropout rate set to 0.1), X is the input of the multilayer perceptron encoder. For transpose, This is the weight matrix of the first layer of the multilayer perceptron encoder. This is the weight matrix of the second layer of the multilayer perceptron encoder. This is the weight matrix of the third layer of the multilayer perceptron encoder. These are the parameters of the first layer of the multilayer perceptron encoder. These are the parameters of the second layer of the multilayer perceptron encoder. These are the parameters of the third layer of the multilayer perceptron encoder. Among them, , , , , The resulting feature matrix serves as the node features of the similarity graph, and is input into subsequent graph neural network layers along with the adjacency matrix. This design enables chemical structure similarity information to be explicitly encoded through edge weights (similarity) and implicitly enhanced through node features, providing a robust feature foundation for multi-view DDI prediction.
[0023] In this embodiment, a drug similarity view representation module is employed. This module constructs a chemical structure similarity map as an auxiliary view and learns more discriminative similarity features through nonlinear compression and denoising. Specifically, the first step generates a "fingerprint" (i.e., Morgan fingerprint) for each drug based on its molecular structure information. This fingerprint is equivalent to a drug's chemical "identity card," describing the various substructure information contained in the drug molecule. Based on this, the chemical similarity between each pair of drug fingerprints is calculated (using the Dice coefficient), resulting in a "similarity matrix" describing the chemical structure similarity between all drugs. Using this matrix as a connection relationship, a "drug similarity map" can be constructed, where drugs are nodes and chemical similarity is the edge weight. Since the original similarity matrix has high dimensionality and information redundancy, this step further employs a multilayer perceptron encoder (i.e., an autoencoder) to perform nonlinear compression and feature enhancement on the matrix. This encoder maps high-dimensional similarity information to a lower-dimensional, more compact feature space through multilayer transformations and fuses features extracted from different levels to form a multi-scale feature representation for each drug. This processing effectively removes redundant information and noise interference, while preserving multi-level chemical semantic information from fine-grained to coarse-grained, providing high-quality attribute feature input for subsequent multi-view fusion.
[0024] It should be noted that, in terms of drug molecule structure feature extraction, GNN-type molecular graph encoders or Transformer-based models can be used to replace Morgan fingerprint encoding. These methods can directly model molecular bond connections at the atomic level, but may be slightly inferior to the Morgan fingerprint combined with autoencoder scheme in terms of feature interpretability and computational efficiency.
[0025] S3, input the drug similarity map, the original DDI map, and the multi-scale diffusion map into the preset shared weight GCN network for encoding to obtain topological embedding, diffusion embedding and similarity embedding respectively; Specifically, step S3 further includes: adaptively fusing the node conditions of each node representation in the multi-scale diffusion map to obtain the fused diffusion features. Specifically: The scale diffusion map corresponding to the kth scale parameter After processing by a single-image encoder, the corresponding node features are obtained. ,in, , For scale parameters The set of edges obtained after diffusion; here, the single graph encoder refers to the GCN encoder, which is a multi-layer graph convolutional coding structure composed of multiple concatenated graph convolutional layers. It takes a single-scale diffusion graph as input, and performs feature extraction and dimensionality transformation layer by layer through multi-layer graph convolution, encoding the topological information at this scale into fixed-dimensional node features, providing a single-scale node representation with unified dimension and unified semantic space for subsequent multi-scale adaptive fusion.
[0026] The features of each node are stacked in the scale dimension and converged in the feature dimension to generate a scale description; The scale description is processed using a learnable map and the Softmax activation function to obtain node-scale weights, generating fusion diffusion features, the formula of which is: ,in, This is the node-level weight vector corresponding to the k-th scale parameter, stored in the scale weight matrix W. This is element-wise multiplication; The formula for the scale weight matrix W is: , , For stacking operations in the scale dimension, For mean aggregation operation on the feature dimension, The node features corresponding to the first scale parameter. For the node features corresponding to the Kth scale parameter, this matrix satisfies , This is the node-level weight vector corresponding to the k-th scale parameter, stored in the g-th row of the scale weight matrix W.
[0027] The drug similarity map, the original DDI map, and the multi-scale diffusion map are jointly input into a pre-defined shared-weight GCN network for encoding, resulting in topological embeddings. Diffusion embedding Similarity embedding d represents the embedding dimension, specifically: Let the normalized adjacency matrices of the drug similarity map, the original DDI map, and the multiscale diffusion map be respectively... , , The initial node characteristics are as follows: , , , The initial node features of the original DDI graph; The drug similarity map, the original DDI map, and the multi-scale diffusion map are propagated and updated using a shared weight GCN network to obtain topological embedding, diffusion embedding, and similarity embedding. The update formula for the l-th layer of the shared-weight GCN network is as follows: , , ,in, The update formula for the (l-1)th layer is as follows: It is a non-linear activation function (i.e., the Sigmoid function). Let L be the weight parameter shared by the l-th layer among the three types of views, and L be the number of layers in the shared weight GCN network. For example, in this embodiment, the number of layers L=3 and the embedding dimension d=32 are set. Through the shared weight constraint, the representations of each view are uniformly projected onto a consistent representation space, thereby obtaining alignable similarity embeddings, topological embeddings, and diffusion embeddings, providing high-quality multi-path input representations for subsequent bidirectional cross-view attention fusion and DDI prediction.
[0028] In this embodiment, a multi-view shared encoding and feature extraction module is used. This module is used to uniformly encode the multi-scale diffusion map, the original DDI map, and the similarity map to achieve consistent representation learning across views.
[0029] Specifically, the generated multi-scale diffusion maps are first fused. Since different drug nodes have varying sensitivities to topological information at different scales (some rely more on local neighbors, others on the global structure), this step employs an adaptive weighting mechanism to automatically learn an optimal multi-scale combination weight for each drug node, fusing features from each scale into a unified "fused diffusion feature." Subsequently, the three views—the drug similarity map, the original DDI map, and the fused multi-scale diffusion map—are jointly input into a shared-weight graph convolutional network (GCN) for encoding. This GCN uses the same weight parameters across the three views, meaning that the information from the three views is projected into the same semantic space, achieving natural alignment of representations. After propagation updates through multiple layers of graph convolution, this step outputs three types of embeddings: topological embedding (encoding the direct interactions of the original DDI network), diffusion embedding (encoding higher-order topological information after multi-scale enhancement), and similarity embedding (encoding chemical structure similarity information). Through this collaborative coding design with shared weights, the present invention not only effectively integrates multi-source information, but also ensures that the representations of different views are semantically comparable and aligned, providing high-quality and well-aligned input features for subsequent bidirectional cross-view attention fusion.
[0030] It should be noted that, in addition to using GCN as a shared encoder, structures such as GAT or GraphTransformer can also be used. These variants differ in their neighbor aggregation methods, but the reason why this invention chooses GCN is that it has high parameter efficiency, stable training, and good compatibility with multi-scale diffusion.
[0031] S4. Additive fusion processing is performed on topological embedding and diffusion embedding to obtain structural query representation. Cross-attention preprocessing is performed based on structural query representation and similarity embedding to obtain unified drug embedding. Based on unified drug embedding, binary classification prediction is performed on drug pairs to be predicted to obtain prediction probabilities.
[0032] Specifically, step S4 further includes: performing additive fusion processing on the topological embedding and the diffuse embedding before entering the cross-view interaction to obtain the structural query representation. This additive fusion method, while maintaining consistency in feature dimensions, superimposes and combines the original topological structure information with diffusion propagation semantics, so that the two structural features form a complementary and enhanced joint representation at the query end.
[0033] Using multi-head cross-attention as a cross-view interaction operator, for any attention head Given a query, a key, and a value, the output of the m-th attention head is: , Let m be the learnable query projection matrix for the m-th attention head. Let be the learnable key projection matrix of the m-th attention head. Let be the learnable projection matrix of the m-th attention head. M represents the dimension of the attention heads, and M is the number of attention heads (e.g., setting the number of attention heads M=4). The outputs of multiple attention heads are concatenated and linearly mapped to obtain the final cross-attention result. , For concatenation functions, This is the output of the first attention head. For the output of the Mth attention head, The output projection matrix of the multi-head attention mechanism; The first approach (structure to attribute): Based on the formula of the final cross-attention result, using the structure query representation as the query and the similarity embedding as the key and value, selective retrieval of chemical priors from a structure perspective is performed to obtain the first output result. This approach enables the structural perspective to adaptively aggregate attribute cues carried by "structural neighbors" for different drug nodes, thereby enhancing the transferability and stability of node representations when the DDI network is sparse or observations are missing.
[0034] The second approach (attribute-to-structure): Based on the formula of the final cross-attention result, using similarity embedding as the query and structural query representation as the key and value, it absorbs and constrains the structural semantics through attribute priors to obtain the second output result. This approach injects higher-order dependencies inherent in network topology and diffusion propagation into the attribute perspective, which helps to suppress overgeneralization caused solely by chemical structural similarity and improves the distinguishability of drugs with "similar structures but different interaction mechanisms".
[0035] To preserve the core information from each perspective and stabilize gradient propagation, residual enhancement is performed on the first and second output results respectively to obtain the first enhanced representation. Second Enhancement Representation By stacking two layers of bidirectional cross-view attention modules, the model can progressively deepen view interaction.
[0036] By concatenating the first enhancement representation and the second enhancement representation, a unified drug embedding is obtained. , For the concatenation operation, it means concatenating the features together; Based on unified drug embedding, for any drug pair Construct paired features and perform binary classification prediction to obtain the predicted probabilities. , For multilayer perceptron classifiers, To unify the row vector corresponding to drug i in the drug embedding, To unify the row vector corresponding to drug j in the drug embedding.
[0037] In this embodiment, a bidirectional cross-perspective attention fusion and prediction optimization module is used to achieve fine-grained bidirectional complementarity and alignment between the structural view and the attribute view. Specifically, the topological embedding and diffusion embedding are first additively fused (i.e., element-wise added) to obtain a "structural query representation" that integrates the original topological information and multi-scale diffusion information. Subsequently, a bidirectional cross-perspective attention mechanism is used to achieve deep interaction between structural information and chemical attribute information. This mechanism includes two symmetrical directions: the first direction uses the structural query representation as the query and the similarity embedding as the key and value, allowing the structural information to actively retrieve and aggregate useful prior knowledge from chemical attributes; the second direction is the opposite, using the similarity embedding as the query and the structural query representation as the key and value, allowing the attribute information to absorb and constrain structural semantics. Through the complementary interaction of these two directions, the structural perspective and the attribute perspective achieve mutual completion and synergistic enhancement. To maintain the original information of each perspective and stabilize model training, the output of each direction is added to its original input with residuals to obtain the enhanced representation. Then, these two enhanced representations are concatenated to form the final "unified drug embedding," which simultaneously carries the enhanced structural information and attribute information. Finally, for any drug to be predicted, extract the vectors of both from the unified drug embedding, construct a pairing feature including concatenation, element-wise multiplication, and absolute difference, input it into the multilayer perceptron classifier, and output the predicted probability that the drug pair has an interaction.
[0038] It should be noted that in the multi-view fusion module, in addition to using a bidirectional cross-view attention mechanism, strategies such as cross-view gating fusion or tensor fusion can also be tried. These methods can also achieve weighted integration of multi-source information, but they may be slightly inferior to the collaborative attention structure adopted in this invention in terms of feature semantic alignment and bidirectional complementarity modeling.
[0039] Preferably, in this embodiment, the BiCrossNet-DDI model uses an overall loss function during training. Conduct training, The binary cross-entropy loss function is... To reconstruct the loss function.
[0040] The formula for the reconstruction loss function is: , The number of original input features. For the c-th original input feature, The c-th model represents the reconstructed output feature; the formula for the binary cross-entropy loss function is... , For drug response The label represents the actual interaction between them. A value of 0 indicates no interaction, and a value of 1 indicates interaction.
[0041] Specifically, in this embodiment, the model training uses a joint loss function to simultaneously optimize feature reconstruction error and classification error; the autoencoder part maintains the stability of the feature space by minimizing the mean square error, and the drug interaction prediction part uses binary cross-entropy loss.
[0042] Training settings: The model is trained for 1000 epochs, with an initial learning rate of 0.001 and the Adam optimizer used; node embedding dimension d=32; dropout rate is set to 0.1; and full-graph training mode is used.
[0043] Through a series of model optimization and training strategies, including multi-scale topology modeling, bidirectional cross-view attention fusion, and residual enhancement, this invention aims to construct an efficient, accurate, and highly generalizable drug interaction prediction model. By integrating drug chemical structure similarity with multi-scale network topology information, this model achieves fine-grained cross-view feature alignment and complementary enhancement, effectively addressing challenges such as data sparsity, network heterogeneity, and complex high-order dependencies in practical applications. This provides reliable technical support for interaction screening and clinical combination drug safety assessment during drug development.
[0044] In summary, this method first constructs a DDI network graph based on the drug relationship network, and then uses a multi-scale diffusion algorithm to diffuse the DDI network graph to obtain a multi-scale diffusion graph. Next, a similarity matrix is obtained based on the similarity relationship between drugs, and a drug similarity graph is obtained through an autoencoder. The three graphs are simultaneously input into a GCN to obtain three embeddings. Finally, the three different embeddings are input into a bidirectional cross-view attention network to obtain a fused embedding, which is then used for prediction by an MLP.
[0045] Compared with existing technologies, this method has the following advantages: 1. It achieves multi-dimensional feature integration, significantly improving the accuracy of drug interaction prediction: This invention integrates information on drugs at three levels—chemical structure, direct interaction, and potential topological association—by constructing a Morgan fingerprint similarity view, the original DDI network view, and a multi-scale diffusion view. Multi-view collaborative learning is achieved through shared weight GCN and a bidirectional cross-view attention mechanism, significantly improving the comprehensiveness and discriminative power of drug embedding representations. Experiments show that on the DeepDDI dataset, compared with the best baseline method MOTOR, the model of this invention achieves a 4.52% improvement on ACC (from 92.15% to 96.67%), an improvement on AUROC from 94.50% to 99.11%, an improvement on AUPRC from 97.39% to 99.58%, and an improvement on F1 from 94.84% to 97.76%. It also achieves leading performance on the ZhangDDI and ChCh-Miner datasets, validating the effectiveness of the method. 2. Effectively captures high-order topological relationships, addressing the problem of traditional methods neglecting indirect interactions: To address the issue that existing DDI prediction models can only identify direct drug-drug interactions and struggle to reveal potential high-order relationships, this invention introduces a PPR-based multi-scale diffusion mechanism. By setting different diffusion parameters, multiple diffusion maps with different propagation radii are generated, explicitly characterizing continuous structural information from local neighborhoods to the global community. This design enables the model to simultaneously capture local neighborhood patterns and long-range topological associations, effectively alleviating the sparsity problem of DDI networks and enhancing the predictive ability for unobserved drug pairs. 3. Achieves adaptive bidirectional fusion of features, enhancing model robustness and generalization performance: The multi-view fusion module employs a bidirectional cross-view attention mechanism, achieving fine-grained cross-view alignment and complementarity through two reciprocal information flows: "structure to attribute" and "attribute to structure." This mechanism allows the model to dynamically determine when to rely on structural information and when to trust attribute priors, avoiding information redundancy and conflicts caused by simple splicing methods. 4. Node-level adaptive fusion and residual connections improve representation quality: This invention employs a node-conditional adaptive scale fusion mechanism to learn the optimal multi-scale combination weights for each drug node, ensuring that drugs with different structural patterns can obtain a suitable topological receptive field. Simultaneously, the introduced residual connections effectively alleviate the oversmoothing problem of deep graph networks, maintaining stable gradient propagation. 5. Joint optimization mechanism enhances model stability and convergence performance: The model adopts a joint optimization strategy of autoencoder reconstruction loss and binary cross-entropy loss, ensuring a balance between classification accuracy and structure preservation in feature representation. This design effectively prevents performance degradation due to overfitting during training, improving the stability and interpretability of prediction results. 6. Good engineering scalability and application prospects: Each module of the method in this invention has a modular design, facilitating flexible migration across different drug databases and feature types.In the case study, this invention successfully identified multiple high-risk drug pairs (such as meloxicam-carbofen and celecoxib-digoxin) from the DeepDDI test set. The prediction results were consistent with records in the DrugBank database, demonstrating its practical application value in assisting rational drug use decisions and focusing on screening high-risk combinations. This technical solution is not only applicable to drug-drug interaction prediction tasks, but can also be extended to biological network association prediction scenarios such as drug-target and drug-disease, providing an efficient and scalable computational framework for drug reuse and clinical safety assessment.
[0046] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A drug interaction prediction method based on a bidirectional cross-perspective attention network, characterized in that, include: The pre-trained BiCrossNet-DDI model is used to process the drug data to be predicted, specifically as follows: A raw DDI map was constructed based on drug data, and multi-scale diffusion processing (PPR) was used to enhance the information in the raw DDI map, resulting in a multi-scale diffusion map, as follows: The raw drug interaction data in the drug data is represented as the DDI raw graph. and will As the adjacency matrix of the original graph of DDI, where, N represents the total number of drugs, V is the set of drug nodes, and E is the set of edges representing known drug interactions. It is the set of real numbers; Multi-scale diffusion using PPR is employed to enhance the information of the original DDI graph. Specifically, the adjacency matrix of the original DDI graph is normalized to obtain a normalized propagation matrix. Based on the normalized propagation matrix, multiple scale parameters are obtained. PPR diffusion matrix Its formula is: , , , I is the identity matrix, and k is the scale parameter index; Based on multiple PPR diffusion matrices, a corresponding set of multi-scale diffusion matrices is generated. , For the k-th scale parameter The PPR diffusion matrix, where K is the number of scale parameters; By treating each PPR diffusion matrix as a weighted diffusion view, a multi-scale diffusion map is constructed. The multi-scale diffusion matrix set is a weighted adjacency matrix representation of the multi-scale diffusion graph. A symmetric similarity matrix between drugs is constructed based on drug data. This symmetric similarity matrix is then used as a weighted adjacency matrix to construct a drug similarity graph. A pre-defined multilayer perceptron encoder is used to perform a nonlinear transformation on the symmetric similarity matrix to obtain the corresponding multi-scale feature representation, specifically: The RDKit toolkit is used to parse the SMILES string of each drug in the drug data into a molecular structure, and a 2048-bit binary vector is generated using a Morgan circular fingerprint with a radius of 2 to obtain the Morgan fingerprint of the drug. The Dice coefficient is used to measure the chemical similarity between drug i and drug j, and a symmetric similarity matrix between drug i and drug j is constructed. , , , Morgan fingerprint of drug i Morgan fingerprint of drug j; A drug similarity graph is constructed by using the symmetric similarity matrix as a weighted adjacency matrix. , Drug similarity map The set of edges; A pre-defined multilayer perceptron encoder is used to perform a nonlinear transformation on the symmetric similarity matrix. The multilayer outputs of the multilayer perceptron encoder are cascaded and fused along the channel dimension to form a multi-scale feature representation of the nodes in the drug similarity map. , , , Furthermore, multi-scale features are represented as node features in the drug similarity map, where h is the feature dimension, representing the dimensionality of each layer of node features. The concatenation operation represents concatenating the features output from each layer together. For activation function, For batch normalization, For random deactivation regularization, X is the input of the multilayer perceptron encoder. For transpose, This is the weight matrix of the first layer of the multilayer perceptron encoder. This is the weight matrix of the second layer of the multilayer perceptron encoder. This is the weight matrix of the third layer of the multilayer perceptron encoder. These are the parameters of the first layer of the multilayer perceptron encoder. These are the parameters of the second layer of the multilayer perceptron encoder. These are the parameters of the third layer of the multilayer perceptron encoder. The drug similarity map, the original DDI map, and the multi-scale diffusion map are input into a pre-defined shared weight GCN network for encoding, resulting in topological embedding, diffusion embedding, and similarity embedding, respectively. Additive fusion of topological embedding and diffusion embedding is performed to obtain structural query representation. Cross-attention preprocessing is performed based on structural query representation and similarity embedding to obtain unified drug embedding. Based on unified drug embedding, binary classification prediction is performed on drug pairs to be predicted to obtain prediction probabilities.
2. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 1, characterized in that, The drug similarity map, the original DDI map, and the multi-scale diffusion map are jointly input into a pre-defined shared-weight GCN network for encoding, resulting in topological embedding, diffusion embedding, and similarity embedding, respectively: Adaptive fusion of node conditions is performed on the node representations in the multi-scale diffusion map to obtain the fused diffusion features. ; The drug similarity map, the original DDI map, and the multi-scale diffusion map are jointly input into a pre-defined shared-weight GCN network for encoding, resulting in topological embeddings. Diffusion embedding Similarity embedding d is the embedding dimension.
3. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 2, characterized in that, Adaptive fusion of node conditions is performed on the node representations in the multi-scale diffusion map to obtain the fused diffusion features, specifically: The scale diffusion map corresponding to the kth scale parameter After processing by a single-image encoder, the corresponding node features are obtained. ,in, , For scale parameters The set of edges obtained after diffusion; The features of each node are stacked in the scale dimension and converged in the feature dimension to generate a scale description; The scale description is processed using a learnable map and the Softmax activation function to obtain node-scale weights, generating fusion diffusion features, the formula of which is: ,in, This is the node-level weight vector corresponding to the k-th scale parameter, stored in the scale weight matrix W. This is element-wise multiplication; The formula for the scale weight matrix W is: , , For stacking operations in the scale dimension, For mean aggregation operation on the feature dimension, The node features corresponding to the first scale parameter. This represents the node feature corresponding to the Kth scale parameter.
4. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 3, characterized in that, The drug similarity map, the original DDI map, and the multi-scale diffusion map are jointly input into a pre-defined shared-weight GCN network for encoding, specifically as follows: Let the normalized adjacency matrices of the drug similarity map, the original DDI map, and the multiscale diffusion map be respectively... , , The initial node characteristics are as follows: , , , The initial node features of the original DDI graph; The drug similarity map, the original DDI map, and the multi-scale diffusion map are propagated and updated using a shared weight GCN network to obtain topological embedding, diffusion embedding, and similarity embedding. The update formula for the l-th layer of the shared-weight GCN network is as follows: , , ,in, The update formula for the (l-1)th layer is as follows: It is a non-linear activation function. represents the weight parameters shared by layer l among the three types of views, and L is the number of layers in the shared weight GCN network.
5. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 4, characterized in that, Additive fusion of topological embeddings and diffusion embeddings yields a structural query representation. Cross-attention preprocessing is then performed on this structural query representation and similarity embeddings to obtain a unified drug embedding. Finally, based on this unified drug embedding, binary classification prediction is performed on the drug pairs to be predicted, yielding the prediction probabilities. Specifically: Additive fusion of topological embedding and diffuse embedding yields a structural query representation. ; Using multi-head cross-attention as a cross-view interaction operator, for any attention head Given a query, a key, and a value, the output of the m-th attention head is: , Let m be the learnable query projection matrix for the m-th attention head. Let be the learnable key projection matrix of the m-th attention head. Let be the learnable projection matrix of the m-th attention head. Let M be the dimension of the attention heads, and M be the number of attention heads. The outputs of multiple attention heads are concatenated and linearly mapped to obtain the final cross-attention result. , For concatenation functions, This is the output of the first attention head. For the output of the Mth attention head, This is the output projection matrix of the multi-head attention mechanism.
6. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 5, characterized in that, Also includes: Based on the formula derived from the final cross-attention result, using the structural query representation as the query and the similarity embedding as the key and value, a selective retrieval of chemical priors from a structural perspective is performed to obtain the first output result. ; Based on the formula derived from the final cross-attention result, similarity embedding is used as the query, and structural query representation is used as the key and value. Attribute priors are then used to absorb and constrain structural semantics, resulting in the second output. ; Residual augmentation is performed on the first and second output results respectively to obtain the first augmented representation. Second Enhancement Representation ; By concatenating the first enhancement representation and the second enhancement representation, a unified drug embedding is obtained. , For the concatenation operation, it means concatenating the features together; Based on unified drug embedding, for any drug pair Construct paired features and perform binary classification prediction to obtain the predicted probabilities. , For multilayer perceptron classifiers, To unify the row vector corresponding to drug i in the drug embedding, To unify the row vector corresponding to drug j in the drug embedding.
7. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 1, characterized in that, The BiCrossNet-DDI model uses an overall loss function during training. Conduct training, The binary cross-entropy loss function is... To reconstruct the loss function.
8. The drug interaction prediction method based on a bidirectional cross-perspective attention network according to claim 7, characterized in that, The formula for the reconstruction loss function is: , The number of original input features. For the c-th original input feature, The c-th model represents the reconstructed output feature; the formula for the binary cross-entropy loss function is... , For drug response The label represents the actual interaction between them. A value of 0 indicates no interaction, and a value of 1 indicates interaction.
Citation Information
Patent Citations
Multi-category drug interaction prediction method, device, equipment and medium
CN116936126A
Asymmetric drug interaction prediction method based on diffusion diagram attention network
CN120340685A