DDI prediction method and system based on dual-view robust causal substructure network
By employing a method based on a dual-view robust causal substructure network, confounding factors in drug molecule maps are eliminated, and causal-related substructures are identified and learned. This addresses the shortcomings in the accuracy and generalization ability of drug interaction prediction in existing technologies, achieving more efficient drug interaction prediction.
Patent Information
- Application Number
- CN202511261101.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing drug interaction prediction methods struggle to effectively distinguish between substructures that play a key role in drug interactions and irrelevant or interfering substructures when faced with high-dimensional sparse data and complex molecular structures. Furthermore, they fail to effectively identify and eliminate confounding factors, resulting in models that perform well on the training set but poorly on unseen data.
A method based on a dual-view robust causal substructure network is adopted. A robust causal graph representation learning network is used to remove hybrid substructures, and a dual-view DDI prediction network is combined for alternating optimization training to identify causal relationships and perform drug representation learning, thereby improving the accuracy and generalization ability of prediction.
It improves the accuracy and robustness of drug interaction prediction, especially in the face of unknown drugs or complex interaction patterns, and has a stronger generalization ability, enabling more accurate identification of causal relationships between drugs.
Smart Images

Figure CN120748555B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of drug interaction prediction technology, and particularly relates to a method and system for predicting drug interaction (DDI) based on a dual-view robust causal structure network. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Drug-drug interactions (DDIs) refer to the phenomenon where the efficacy or toxicity of one drug is affected by another when two or more drugs are used concurrently. These interactions can lead to reduced drug efficacy, enhanced side effects, or even serious adverse drug events, posing a threat to patient health. With the increasing prevalence of multidrug combination therapy in clinical practice, accurately predicting potential DDIs has become a key challenge in drug development and clinical medication safety. Traditional DDI detection relies on expensive and time-consuming clinical trials and pharmacological studies, making it difficult to cover the vast number of potential drug combinations. Therefore, developing efficient computational methods to predict DDIs is of significant practical importance.
[0004] Traditional methods primarily rely on pharmacological rules, chemical similarity calculations, or association rule mining. However, these methods have limitations when dealing with high-dimensional sparse data and complex molecular structures. Therefore, in recent years, a significant amount of work has introduced deep neural networks, especially graph neural networks (GNNs), into DDI modeling, achieving remarkable progress. Due to their powerful ability to process graph-structured data (such as drug molecule structures), GNNs have become one of the mainstream techniques for DDI prediction. These methods typically represent drug molecules as graphs, with atoms as nodes and chemical bonds as edges, then use GNNs to learn the representation vectors of the drugs and predict the interactions between drug pairs based on these vectors.
[0005] However, existing DDI prediction methods still face many challenges. First, many methods focus primarily on the complete molecular structure of individual drugs when learning drug representations, potentially neglecting detailed interatomic interactions or failing to effectively distinguish between substructures that play a crucial role in DDI and irrelevant or interfering substructures, resulting in incomplete and noisy representations. Second, and more critically, confounding factors are prevalent in real-world graph data (including drug molecule graphs). These confounding factors are substructures or patterns that are related to both drug features and DDI labels but are not the direct cause of DDI. Failure to identify and eliminate these confounding factors may lead to the learning of spurious statistical associations rather than genuine causal relationships. Consequently, the model may perform well on the training set but poorly on new, unseen, or out-of-distribution (OOD) data, impacting prediction accuracy and generalization ability. Summary of the Invention
[0006] To address at least one of the technical problems in the background art, the present invention provides a method and system for predicting adverse drug reactions (DDIs) based on a dual-view robust causal structure network. By learning a more robust and denoised drug representation, it improves the accuracy, robustness, and generalization ability of DDI prediction, especially when facing unknown drugs or complex interaction patterns.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A first aspect of the present invention provides a method for predicting adverse drug reactions based on a dual-view robust causal structure network, comprising the following steps:
[0009] Obtain the drug molecule diagrams for each drug, and construct individual drug internal views and inter-drug views;
[0010] The drug molecule diagrams of each drug are input into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing hybrid substructures.
[0011] The robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view are input into the trained dual-view DDI prediction network to obtain the DDI prediction results. When training the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the training of the robust causal graph representation learning network is guided by the DDI prediction results.
[0012] Furthermore, the construction process of a robust causal graph representation learning network includes:
[0013] Feature extraction is performed on the drug molecule graph, edge weights are predicted based on the extracted features, causal relationships related to DDI are determined based on the predicted edge weights, and hybrid substructures that are statistically correlated with the DDI label but have no causal relationship are eliminated.
[0014] The construction process of the dual-view DDI prediction network includes:
[0015] Extract the single-drug internal view and inter-drug view of the robust causal graph structure after removing hybrid substructures; further extract the global semantic drug representation by splicing and fusing the single-drug internal view and inter-drug view; calculate the probability of drug pairs having interactions by combining the global semantic drug representation.
[0016] Furthermore, the molecular structure of a single drug is shown in Figure [Figure Number]. In this context, the node set V represents the atoms in the drug, and the edge set E represents the chemical bonds between atoms; for a single drug internal view, the adjacency matrix is used... To represent the edge, ,in Indicates the number of atoms, Meaning atoms and There are no edges between them. For the view between drugs, construct a bipartite graph. The adjacency matrix , This represents the atomic interaction between two drugs. and These represent the number of nodes in the two drugs, respectively. Atoms in drug A With the atoms in drug B There exists an edge between them in the bipartite graph for each atom in drug A. This is compared with each atom in drug B. Connected.
[0017] Furthermore, based on the constraint elimination set by the instrumental variable generation network, hybrid substructures that are statistically correlated with but not causally related to DDI labels are eliminated. The constraint conditions are as follows:
[0018] ,
[0019] ,
[0020] in, Graph structure The causal substructure, Graph structure The hybrid substructure, Generate the network output for instrumental variables. Generate a network for instrumental variables.
[0021] Furthermore, the loss function for training a robust causal graph representation learning network is:
[0022]
[0023]
[0024]
[0025]
[0026]
[0027] in, For the overall loss function, The loss is used to make the model output more consistent representations for the same class. These are hyperparameters used to balance the influence weights of M. It is the decision to activate regular terms hyperparameters, The weights for each sample, The adjustable weighting parameters are generated based on the results of positive and negative samples. It is a hyperparameter that adjusts the weighting effect. It is the cross-entropy loss, where r(•) represents a function with no adoptions. It is a graph structure. It is a node feature representation. It is the average value of node features. For parameters Participating functions , Represents the similarity calculation function. and These represent the minimum and maximum similarity values, respectively. It is a globally distinguishable regularization term. It is the representation mean of all samples; the goal is to make each Maintain a distance from the overall mean to avoid falling into category collapse.
[0028] Furthermore, a single-drug internal view of a robust causal graph structure. Represented as:
[0029]
[0030]
[0031] View between drugs Represented as:
[0032]
[0033] in, This represents the edge weights based on the attention mechanism. For the first Layer node representation, For activation function, and For trainable matrices, and For bias, Given a trainable matrix, and The representation consists of a vector representation of the node itself and the representations of its neighboring nodes. For the set of all nodes in another drug, For inter-view attention weights, For the trainable matrix between molecules, This represents the adjacent nodes between molecules.
[0034] Furthermore, the loss function for training the dual-view DDI prediction network is:
[0035] ,
[0036]
[0037]
[0038]
[0039] in, To use the standard binary classification cross-entropy loss function, The loss function for parameter θ in the prediction model. For regularization terms, This is a hyperparameter responsible for controlling the degree of influence of regularization. For positive samples, The predicted probability of the negative sampled pair. For two drugs , A triple consisting of an interaction type R, Indicates a fixed parameter. It is a similarity calculation function. This indicates training using robust clipped subgraphs with fixed parameters. This indicates training using a complete graph and learnable parameters. For DDI tags.
[0040] A second aspect of the present invention provides a DDI prediction system based on a dual-view robust causal substructure network, comprising:
[0041] The graph construction module is used to obtain drug molecule graphs for each drug and construct single-drug internal views and inter-drug views.
[0042] The dual-view robust causal substructure network training module is used to input the drug molecule diagrams of each drug into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing hybrid substructures.
[0043] The robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view are input into the trained dual-view DDI prediction network to obtain the DDI prediction results. When training the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the training of the robust causal graph representation learning network is guided by the DDI prediction results.
[0044] A third aspect of the present invention provides a computer-readable storage medium.
[0045] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the drug adverse reaction prediction method based on a dual-view robust causal structure network as described above.
[0046] A fourth aspect of the present invention provides a computer device.
[0047] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the DDI prediction method based on a dual-view robust causal structure network as described above.
[0048] A fifth aspect of the present invention provides a program product.
[0049] A program product, which is a computer program product, includes a computer program that, when executed by a processor, implements the steps in the DDI prediction method based on a dual-view robust causal structure network as described above.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] This invention proposes a novel Drug-Drug Interaction (DDI) prediction framework, DRCS-DDI. It utilizes the Instrumental Variable (IV) module of a robust causal substructure network as a preprocessing module for drug-drug response prediction. This causal substructure module actively identifies and removes confounding substructures in drug molecule graphs and drug-pair bipartite graphs that contribute no causal contribution or even interfere with DDI prediction, generating purer and more causal features. Subsequently, the Dual-view Structure Network (DSN)-DDI module performs its unique intra-drug and inter-drug dual-view representation learning on these robustly pruned and denoised graph structures, and performs the final DDI prediction, improving the accuracy, robustness, and generalization ability of DDI prediction, especially when facing unknown drugs or complex interaction patterns.
[0052] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0053] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0054] Figure 1 This is a flowchart of the DDI prediction method based on a dual-view robust causal substructure network provided in this embodiment of the invention.
[0055] Figure 2 This is the node propagation method in the prediction module provided in this embodiment of the invention;
[0056] Figure 3 This is a learning demonstration of the inner view and intermediate view provided in the embodiments of the present invention. Detailed Implementation
[0057] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0058] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0059] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0060] Drug identification and treatment (DDI) prediction methods have evolved from traditional machine learning to deep learning. Early methods relied on hand-designed features, such as the drug's physicochemical properties, target information, and side effects, and combined them with traditional machine learning algorithms such as support vector machines (SVM) and logistic regression for prediction. However, these methods were often limited by the quality and coverage of feature engineering.
[0061] With the development of deep learning, especially the success of graph neural networks (GNNs) in molecular representation learning, DDI prediction methods based on GNNs have become a research hotspot. These methods can automatically learn effective representations from drug molecular structure diagrams. For example, some works directly use GNNs to encode individual drug molecule diagrams, and then concatenate or combine the representations of drug pairs through specific functions before inputting them into a classifier for prediction.
[0062] To capture the mechanisms of drug-induced drug dissociation (DDI) more precisely, researchers have begun to focus on drug substructures. Some methods divide drugs into functional groups or chemical substructures, arguing that these substructures collectively determine the overall pharmacological properties of the drug, and predict DDI by identifying pairwise interactions between different drug substructures. For example, SSI-DDI treats the hidden representations of nodes as substructures and computes the interactions between these substructures. GMPNN-CS learns chemical substructures of different sizes and models the interactions between these substructures. SA-DDI employs a substructure-aware GNN and designs a substructure attention mechanism. However, the authors of DSN-DDI point out that substructures in these methods are often learned only from the hidden representations of a single drug, neglecting the interaction information between two drugs, which could provide more valuable information for substructure learning.
[0063] Dual-view learning has also been introduced into DDI prediction, aiming to simultaneously consider information about the drug itself and the joint information of drug pairs. MHCADDI designs an external message-passing mechanism to integrate the joint information of drug pairs during the learning of individual drug representations. However, MHCADDI only uses the hidden representation of the last layer of the GNN for prediction, ignoring the hierarchical global representations from previous GNN modules that contain neighborhood information at different scales. DSN-DDI itself, on the other hand, learns drug substructures from both the intra-view (within a single drug) and inter-view (between drug pairs) through iterative local and global representation learning modules, and utilizes all hierarchical global representations for DDI prediction.
[0064] Despite significant progress in existing methods, most of them do not explicitly address the confounding factors prevalent in graph data, which may limit their predictive performance and generalization ability.
[0065] Causal inference aims to identify genuine causal relationships between variables, rather than just statistical correlations. In machine learning, introducing causal inference helps improve the robustness, interpretability, and generalization ability of models. In graph representation learning, confounders in the data are a long-standing but often overlooked problem. Confounders are variables or substructures that are related to both input features and target labels, but are not true mediators in the causal path between them. In graph representation learning, they can be defined as substructures that fail to provide stable specific guidance for label prediction. GNNs are easily affected by these confounders during the learning process, thus learning spurious associations and causing poor performance on out-of-distribution (OOD) data.
[0066] Example 1
[0067] like Figure 1 As shown, this embodiment provides a method for predicting adverse drug reactions based on a dual-view robust causal structure network, including the following steps:
[0068] Step 1: Obtain the drug molecule diagram for each drug and construct the individual drug internal view and inter-drug view; for the individual drug internal view, use the adjacency matrix... To represent the edge, ,in Indicates the number of atoms, Meaning atoms and There is no boundary between them.
[0069] For the drug-drug view, design a bipartite graph for each drug pair; given two drugs and a triplet of one interaction type, construct a bipartite graph for each drug combination. The adjacency matrix , This represents the atomic interaction between two drugs. and These represent the number of nodes in the two drugs, respectively. Atoms in drug A With the atoms in drug B There exists an edge between them. In the bipartite graph, for each atom in drug A... This is compared with each atom in drug B. Finally, the two-view substructure learning for drug-drug interaction (DDI) prediction is based on having internal connectivity A and external connectivity B. The drug-drug graph is constructed using the DDI prediction function. f Learning it as a binary classification problem, expressed as: The output 1 indicates the presence of two drugs. , A triple consisting of an interaction type R An output of 0 indicates that the interaction does not exist. The goal of drug interaction prediction is to learn a function. f It also predicts the probability of the existence of a triplet between two drugs for a given interaction type.
[0070] In this embodiment, the acquired dataset consists of two parts. The first part uses the public dataset from DrugBank. Data preprocessing involves selecting the drug serial number dataset from DrugBank and using the drug interaction relationships to form the core triples: drug ID 1, drug ID 2, and drug-drug relationship identifier (numerical). In the data processing part, this entire triple is used as a reference to generate corresponding non-existent negative samples; that is, content not present in the triple is considered a negative triple that cannot have a ddi (dual ID). The second part of the data is the SMILES data. By accessing the drug-specific ID identifier, the corresponding SMILES structure is queried in the SMILES dataset, which serves as graph content support, providing a data interface for subsequent GATv2conv.
[0071] It should be noted that the drug interaction prediction task in this work is treated as a binary classification task, rather than directly predicting the interaction type. For different types of interactions, the learned model can make predictions using the corresponding triples.
[0072] Step 2: Input the drug molecule diagrams of each drug into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing confounding substructures; input the robust causal graph structure after removing confounding substructures, the single-drug internal view, and the inter-drug view into the trained dual-view DDI prediction network to obtain the DDI prediction results; when training the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the training of the robust causal graph representation learning network is guided by the DDI prediction results.
[0073] This embodiment innovatively proposes a dual-module fusion framework: the Dual-view Robust Causal Substructure Network (DRCS-DDI). This framework enhances the model's robustness and generalization performance in modeling the DDI mechanism by introducing the concept of causal inference and combining it with the representational capabilities of graph neural networks. The DRCS-DDI framework consists of two core modules: 1) a robust causal graph representation learning module based on instrumental variables; and 2) a dual-view drug substructure modeling and interaction prediction module. The entire model is trained end-to-end, using a synergistic mechanism of "causal pruning and representation learning" to improve the accuracy and generalization ability of DDI prediction. The entire model employs an alternating training strategy for optimization: the IV generator uses the dual-view DDI prediction results as a back-end guide, continuously optimizing the subgraph selection strategy using prediction error information; the dual-view DDI prediction module performs substructure modeling and prediction on the causally pruned graph structure, gradually improving performance and providing stable feedback. This collaborative mechanism enables the dual-view DDI prediction module not only to eliminate spurious associations in graph data, but also to enhance the model's ability to capture true causal structures, resulting in stronger generalization when faced with new drugs or complex drug combinations.
[0074] Specifically, the steps include the following:
[0075] Step 201: Input the drug molecule diagrams of each drug into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing hybrid substructures.
[0076] The construction process of a robust causal graph representation learning network includes: generating the network using an instrumental variable (IV). qG (•) is used to identify and quantify the causal importance of different parts (such as edges) in the graph, and then a parameterless function r(•) is used to remove confounding factors based on these IVs, thereby making the main GNN model more accurate. fG (•) It is able to learn causal relationships on cleaner graphs.
[0077] Let the original graph structure be... It contains causal substructures. With hybrid substructure I want to learn a function Accurately predicts DDI labels And without interference. However, due to and The mixture, directly used study It can be affected by obfuscation terms. Therefore, this embodiment introduces an N-generation network. Output pseudo-instrument variables The following two constraints are used to eliminate confounding effects:
[0078] (1),
[0079] (2),
[0080] Formula (1) indicates that there are no confounding factors in IV, and Formula (2) indicates the causal relationship of DDI in Table 6 in IV.
[0081] Assume the final prediction function is generally defined by the parameters Participating functions ,pass Output forecast plus corresponding confounding bias Right now:
[0082] (3),
[0083] The ultimate goal then changes to:
[0084] (4),
[0085] The robust causal graph structure after removing hybrid substructures is obtained by applying the above two constraints. To retain only with Causal relationship graph structure subgraph.
[0086] To ensure that the cropped subgraph retains sufficient DDI prediction information while removing confounding terms to the maximum extent, a target pair instrumental variable generator network is used to maximize mutual information. Conduct training. The specific format is as follows:
[0087] (5),
[0088] in, For the set of objective parameters of instrumental variables, For a fixed DSN-DDI master model, Representing mutual information, this training method aims to maximize the mutual information between the prediction model and the labels after the prediction graph is modified, ensuring correct training.
[0089] To improve the discriminativeness and robustness of causal subgraphs, the RCGRL (Robust Causal Graph Representation Learning) module produces an optimization objective—Robustness-Emphasizing Loss. This loss enhances the feature aggregation of correctly predicted samples and imposes a discriminative penalty on misclassified samples. The aim is to increase the similarity between correctly classified samples and their corresponding centers, and to impose a discriminative penalty on misclassified samples to prevent model output collapse. It consists of a weighted cross-entropy loss and a global discriminative regularization term.
[0090] Step 202: Combine the robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view into the trained dual-view DDI prediction network to obtain the DDI prediction results.
[0091] After the robust causal graph representation learning module removes hybrid substructures, the resulting robust graph structure serves as input to the dual-view DDI prediction module for dual-view substructure modeling and DDI prediction. This module seamlessly integrates with the original DSN-DDI structure, leveraging graph attention and cross-attention mechanisms to model the structural features and interaction behaviors of drugs from two dimensions, achieving more accurate drug-drug interaction prediction. Based on the graph structure pruned by the RCGRL module, the DSN-DDI module is responsible for learning the structural representation of drugs from two complementary perspectives: the intra-view represents the atomic structure within a single drug, and the inter-view represents the interaction graph structure between drug pairs. This module consists of a multi-layer stack of improved DSN Encoders, each layer including a variant GAT attention mechanism, an information fusion layer, and a SAGPooling substructure selection mechanism.
[0092] After RCGRL processing, these two types of graphs are cropped into more causal substructures, thereby improving the input quality and semantic purity of DSN-DDI.
[0093] Please refer to 2 and Figure 3 , Figure 2 This refers to the node propagation method in the prediction module, which first randomly selects a node from the original drugs as the starting node. i In the first layer, this is taken as i The three adjacent light-colored nodes are the neighboring nodes. j This corresponds to the substructure of the first layer; the newly spreading nodes... j As new i This could lead to the spread of a new wave of " jThis corresponds to the second substructure; and so on, ultimately achieving complete propagation. This approach allows the model to have a better understanding of the relationships within the entire model. Figure 3 yes Figure 2 The relative macroscopic representation, in fact, the upper and lower parts of the internal view represent respectively Figure 2 The node access method in the middle view represents how to access nodes. Figure 2 This method is applied between two drugs, that is, each node in Drug a. i Spread to all nodes in Drug b, then to nodes in Drug a. i The neighboring nodes are then fully linked with Drug b, and so on.
[0094] The DSN-DDI core consists of multiple stacked DSN Encoder modules, each of which comprises the following three sub-layers:
[0095] Intra-view layer (single-drug GAT layer): Employs graph attention mechanism (GAT) to update the features of each drug's internal atomic nodes, capturing local chemical structures.
[0096] (6),
[0097] in, This represents the edge weights based on the attention mechanism. For the first Layer node representation, For activation function, For trainable matrices, For bias, For the first l +1 level node representation;
[0098] In the substructure modeling phase, DRCS-DDI uses a variant GAT module as the core encoder to achieve stronger substructure interaction modeling capabilities. Unlike the original GAT, the variant GAT first concatenates source-target node features during attention weight calculation, then performs a linear transformation for scoring, thus possessing greater expressive flexibility and stability. The explicit attention calculation formula is as follows:
[0099] (7),
[0100] in, Given a trainable matrix, and The representation consists of a vector representation of the node itself and the representations of its neighboring nodes.
[0101] Interview layer (bipartite graph GAT layer): Enables information exchange between drug pairs across drug nodes, and calculates the association strength between each pair of nodes using cross-attention.
[0102] (8),
[0103] in, This represents the node at layer l+1 during the graph learning process. For the set of all nodes in another drug, For inter-view attention weights, For the trainable matrix between molecules, This represents the adjacent nodes between molecules. For bias;
[0104] (9),
[0105] Information fusion layer: Combines intra-view and inter-view information using MLP for fusion.
[0106] (10)
[0107] The structure progressively enhances the nodes' understanding of both their own structure and the context of other drugs at each layer.
[0108] To extract a more globally semantic drug representation, the SAGPooling mechanism is used after each Encoder layer to score and filter the substructure importance of the current layer nodes. The drug representation for each layer is calculated as follows:
[0109] (11),
[0110] in, For learnable parameters, Measuring atoms Contribution to the global representation.
[0111] Ultimately, the model yields a representation of each substructure layer. ,common The layers form a hierarchical semantic sequence for subsequent prediction.
[0112] The DSN-DDI decoder employs a Co-Attention matching mechanism to perform pairwise matching and weighted prediction of the substructural representations of two drugs at different levels:
[0113] (12),
[0114] in To illustrate the importance of interaction between views, The relation matrix representing the interaction type R, and the final score. This indicates the probability of an interaction between the drug pairs.
[0115] The dual-view DDI prediction network is trained using the standard binary classification cross-entropy loss function.
[0116] (13)
[0117] in For positive samples, is the predicted probability of negative sample pairs, maintaining a balance between positive and negative samples, and T represents the number of triples.
[0118] It is the overall training objective function, on which we add overall training objective parameters related to the predictor parameters. The loss function is described by the core idea of improving the robustness of the predictor so that it can actively learn to ignore the confounding factors in the graph. It is based on the comparative learning idea of "two perspectives of the graph" and does not rely on negative samples, so training is simple and efficient.
[0119] Given each graph sample The robust cropping module produces two types of images: a clean view and a clean view. and the original, unprocessed image These two graphs are placed into two GNN-based prediction training functions with identical structures and shared parameters:
[0120] The clean view after processing mixed with fixed parameters represents a prediction;
[0121] The original view, after unprocessing and profanity, is used to represent the prediction during training.
[0122] (14)
[0123] Among them, the loss function for parameter θ in the prediction model;
[0124] Indicates a fixed parameter. For the prediction model, i.e., the invocation of the prediction module, a comparison method was designed as a refinement of the training objective. To prevent the predictor from undergoing erroneous training by detaching from or ignoring the original graph, a regularization term was introduced:
[0125] (15)
[0126] in It is a similarity calculation function. This indicates training using robust clipped subgraphs with fixed parameters. This indicates training using a complete graph and learnable parameters.
[0127] Here is the The parameters are fixed, that is Excluded from backpropagation, forced to be fed by input data containing confounding factors. Generation and Similar outputs, fed by graphs without confounding factors (through a confounding removal operation). Then, The learned parameters will be used to update This training paradigm encourages By ignoring information with a confounding effect, certain confounding factors no longer have a confounding effect on the model of this invention. The final integrated training is as follows:
[0128] (16)
[0129] in, This is a hyperparameter responsible for controlling the degree of influence of regularization.
[0130] Step 203: During the training of the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training. The encoded representation produced by the dual-view DDI prediction network is set as follows: , Representation diagram The representation encoded by the dual-view DDI prediction network serves as the medium through which the robust pruning module receives guidance from the prediction module. First, weights are set for the weighted cross-entropy loss. During sample learning and training, the correctness of a sample is judged. If the prediction is correct, deeper differential learning is performed; if the sample outcome is incorrect, the weights of the incorrect learning are reduced, decreasing their emphasis. In general, the further the output deviates from the optimal value, the greater its contribution to training.
[0131] Specifically, for Belongs to a specific The class first obtains a feature vector. It is The average of all feature vectors in the class. Then, a series of weights are obtained for the samples. The formula for performing the emphasis operation is as follows:
[0132] (17)
[0133] in, Represents the similarity calculation function. and These represent the minimum and maximum similarity values, respectively.
[0134] The final loss is defined as:
[0135] (18)
[0136] in, The adjustable weighting parameters are generated based on the results of positive and negative samples. It is a hyperparameter that adjusts the weighting effect. It is cross-entropy loss.
[0137] To prevent all charts from converging to a single average representation, a global discriminative regularization term is introduced:
[0138] (19)
[0139] in, It is the representation mean of all samples; the goal is to make each To avoid falling into category collapse, maintain a distance from the overall mean. It is used to control the overall value range of M. This is the total number of nodes;
[0140] Therefore, the overall robustness loss is defined as follows:
[0141] (20)
[0142] in The definition is as follows:
[0143] (twenty one),
[0144] It is about deciding when to activate regular terms. The hyperparameters. This design effectively prevents the model from falling into a situation where "all representations converge," while preserving causal robustness training for important samples.
[0145] Guided by the performance feedback from the subsequent prediction module, the causal pruning module continuously adjusts the instrumental variable generation strategy, thereby dynamically optimizing the "pruning-prediction" collaborative mechanism.
[0146] Finally, the drug molecule map is input into the trained instrumental variable generator network. In the process, feature extraction was performed on the drug molecule map to obtain... It predicts the "retention probability" or "importance score" for each edge and outputs a set of edge weights. This is used to evaluate whether an edge or subgraph is causally related to the DDI task. Based on the instrumental variable generator. The output edge weights are used to perform controlled culling operations on the graph structure, such as retaining the previous edge weights. Weighted edges, or probability sampling based on weights, are used to generate new subgraphs. This subgraph is the causal representation subgraph for "de-confounding".
[0147] In this embodiment, a lightweight GNN structure is used to extract features from the drug molecule map;
[0148] In real drug molecule graphs and drug pair bipartite graphs, there are numerous confounding substructures that are statistically correlated with DDI labels but have no causal relationship. These structures may interfere with the model's modeling of the true causal mechanism, leading to overfitting to the training set or performance degradation in OOD tests. To address this, the DRCS-DDI framework introduces an instrumental variable (IV)-based mechanism to identify and eliminate these confounding substructures, thereby improving the causality and robustness of graph representations. The DRCS-DDI model consists of two modules: RCGRL and DSN-DDI, responsible for "causal substructure generation" and "DDI predictive representation learning," respectively. To fully leverage their synergistic effect, this invention employs an alternating optimization strategy for joint training. This allows the subgraph selection strategy learned by the RCGRL module to be guided by the real-time feedback of the DSN-DDI's predictive performance, which in turn promotes the optimization of the instrumental variable generation network, forming an end-to-end robust learning loop.
[0149] Finally, RCGRL and DSN-DDI are jointly trained using an alternating optimization strategy: with the main model parameters fixed. Optimize under the premise This makes the output cropped image It is more in line with the causal constraint; conversely, it is not. use Update main model parameters This optimizes the accuracy of the final DDI prediction.
[0150] The RCGRL module does not directly participate in DDI prediction, but rather acts as a "pre-filter" or "graph causal cleaner" for DSN-DDI. By eliminating substructures without causal significance and retaining key causal subgraphs, RCGRL enables DSN-DDI to perform intra-view and inter-view substructure modeling on cleaner inputs. This combination not only improves prediction accuracy but also significantly enhances the model's generalization ability in OOD scenarios such as unknown drug combinations and unseen structural distributions.
[0151] Example 2
[0152] This embodiment provides a DDI prediction system based on a dual-view robust causal substructure network, including:
[0153] The graph construction module is used to obtain drug molecule graphs for each drug and construct single-drug internal views and inter-drug views.
[0154] The dual-view robust causal substructure network training module is used to input the drug molecule diagrams of each drug into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing hybrid substructures.
[0155] The robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view are input into the trained dual-view DDI prediction network to obtain the DDI prediction results. When training the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the training of the robust causal graph representation learning network is guided by the DDI prediction results.
[0156] It should be noted that the specific implementation of the DDI prediction system based on dual-view robust causal substructure network in this embodiment of the invention is similar to the specific implementation of the DDI prediction method based on dual-view robust causal substructure network in this embodiment of the invention. For details, please refer to the description in the method section. To reduce redundancy, it will not be repeated here.
[0157] Example 3
[0158] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the DDI prediction method based on a dual-view robust causal structure network as described above.
[0159] Example 4
[0160] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the DDI prediction method based on a dual-view robust causal structure network as described above.
[0161] Example 5
[0162] This embodiment provides a program product, which is a computer program product including a computer program. When the computer program is executed by a processor, it implements the steps in the DDI prediction method based on a dual-view robust causal structure network as described above.
[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0164] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0167] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0168] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A DDI prediction method based on a dual-view robust causal network structure, characterized in that, include: Obtain the drug molecule diagrams for each drug, and construct individual drug internal views and inter-drug views; The drug molecule diagrams of each drug are input into the trained robust causal graph representation learning network to obtain the robust causal graph structure after removing hybrid substructures. The robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view are input into the trained dual-view DDI prediction network to obtain DDI prediction results. When training the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the training of the robust causal graph representation learning network is guided by the DDI prediction results. The construction process of a robust causal graph representation learning network includes: Feature extraction is performed on the drug molecule graph, edge weights are predicted based on the extracted features, causal relationships related to DDI are determined based on the predicted edge weights, and hybrid substructures that are statistically correlated with the DDI label but have no causal relationship are eliminated. The construction process of the dual-view DDI prediction network includes: Extract the single-drug internal view and inter-drug view of the robust causal graph structure after removing hybrid substructures; further extract the global semantic drug representation by splicing and fusing the single-drug internal view and inter-drug view; calculate the probability of drug pairs having interactions by combining the global semantic drug representation.
2. The DDI prediction method based on a dual-view robust causal subnetwork as described in claim 1, characterized in that, The molecular structure of a single drug is shown in Figure 1. In this context, the node set V represents the atoms in the drug, and the edge set E represents the chemical bonds between atoms; for a single drug internal view, the adjacency matrix is used... To represent the edge, ,in Indicates the number of atoms, Meaning atoms and There are no edges between them. For the view between drugs, construct a bipartite graph. The adjacency matrix , This represents the atomic interaction between two drugs. and These represent the number of nodes in the two drugs, respectively. Indicates the atoms in drug A With the atoms in drug B There exists an edge between them in the bipartite graph for each atom in drug A. This is compared with each atom in drug B. Connected.
3. The DDI prediction method based on a dual-view robust causal subnetwork as described in claim 1, characterized in that, Based on the constraints set in the instrumental variable generation network, hybrid substructures that are statistically correlated with but not causally related to DDI labels are eliminated. The constraints are as follows: , , in, Graph structure The causal substructure, Graph structure hybrid substructure, Generate the network output for instrumental variables. Generate a network for instrumental variables.
4. The DDI prediction method based on a dual-view robust causal subnetwork as described in claim 1, characterized in that, The loss function for training a robust causal graph representation learning network is: in, For the overall loss function, The loss is used to make the model output more consistent representations for the same class. These are hyperparameters used to balance the influence weights of M. It is the decision to activate conventional terms hyperparameters, The weights for each sample, The adjustable weighting parameters are generated based on the results of positive and negative samples. It is a hyperparameter that adjusts the weighting effect. It is the cross-entropy loss, where r(•) represents a function with no adoptions. It is a graph structure. It is a node feature representation. It is the average value of node features. For parameters Participating functions , Represents the similarity calculation function. and These represent the minimum and maximum similarity values, respectively. It is a globally distinguishable regularization term. It is the representation mean of all samples; the goal is to make each Maintain a distance from the overall mean to avoid falling into category collapse.
5. The DDI prediction method based on a dual-view robust causal subnetwork as described in claim 1, characterized in that, Single-drug internal view of robust causal graph structure Represented as: View between drugs Represented as: in, This represents the edge weights based on the attention mechanism. For the first Layer node representation, For activation function, and For trainable matrices, and For bias, Given a trainable matrix, and The representation consists of a vector representation of the node itself and the representations of its neighboring nodes. For the set of all nodes in another drug, For inter-view attention weights, For the trainable matrix between molecules, This represents the adjacent nodes between molecules.
6. The DDI prediction method based on a dual-view robust causal substructure network as described in claim 1, characterized in that, The loss function for training the dual-view DDI prediction network is: , in, To use the standard binary classification cross-entropy loss function, The loss function for parameter θ in the prediction model. For regularization terms, This is a hyperparameter responsible for controlling the degree of influence of regularization. For positive samples, The predicted probability of the negative sampled pair. For two drugs , A triple consisting of an interaction type R, Indicates a fixed parameter. It is a similarity calculation function. This indicates training using robust clipped subgraphs with fixed parameters. This indicates training using a complete graph and learnable parameters. For DDI tags.
7. A DDI prediction system based on a dual-view robust causal network structure, characterized in that, include: The graph construction module is used to obtain drug molecule graphs for each drug and construct single-drug internal views and inter-drug views. A dual-view robust causal substructure network training module is used to input drug molecule diagrams of various drugs into a trained robust causal graph representation learning network to obtain a robust causal graph structure after removing confounding substructures. The construction process of the robust causal graph representation learning network includes: Feature extraction is performed on the drug molecule graph, edge weights are predicted based on the extracted features, causal relationships related to DDI are determined based on the predicted edge weights, and hybrid substructures that are statistically correlated with the DDI label but have no causal relationship are eliminated. The robust causal graph structure after removing hybrid substructures, the single-drug internal view, and the inter-drug view are input into the trained dual-view DDI prediction network to obtain DDI prediction results. During the training of the robust causal graph representation learning network and the dual-view DDI prediction network, an alternating optimization strategy is used for joint training, and the DDI prediction results are used to guide the training of the robust causal graph representation learning network. The construction process of the dual-view DDI prediction network includes: Extract the single-drug internal view and inter-drug view of the robust causal graph structure after removing hybrid substructures; further extract the global semantic drug representation by splicing and fusing the single-drug internal view and inter-drug view; calculate the probability of drug pairs having interactions by combining the global semantic drug representation.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the DDI prediction method based on a dual-view robust causal substructure network as described in any one of claims 1-7.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the DDI prediction method based on a dual-view robust causal substructure network as described in any one of claims 1-7.
Citation Information
Patent Citations
Drug recommendation method based on causal comparison and message passing
CN118969175A
Drug pair interaction prediction method and device based on soft mask double-view learning
CN119296636A