Drug symmetric element path-based adverse reaction prediction method and system
By using a model MPGNN-DSA based on a heterogeneous information network in drug side effects prediction, information transmission and feature fusion are carried out along the metapath, the problems of insufficient data utilization and semantic correlation capture in the prior art are solved, and high accuracy and stability of drug side effects prediction are achieved.
Patent Information
- Application Number
- CN202510139966.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Existing drug side effects association prediction technologies are difficult to make full use of a variety of data, cannot effectively capture the complex semantic associations between biological entities, and lack interpretability, affecting the accuracy and reliability of predictions.
Using the model MPGNN-DSA based on Heterogeneous Information Network (HIN), MPGNN-DSA is used to construct heterogeneous graphs and explicit information transmission and feature fusion along carefully designed metapaths, complex semantic associations are captured, and the quality of feature representation is improved through metapathic path slicing technology.
It significantly improves the accuracy and stability of drug side effects prediction, improves AUC and AUPR values, and has a lower standard deviation of prediction results, indicating stronger prediction stability on different data samples.
Smart Images

Figure CN120164638A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of drug adverse reaction prediction, and specifically relates to a drug symmetric metapath-based adverse reaction prediction method and system. Background Art
[0002] According to the authoritative definition given by the International Drug Testing Cooperation Center of the World Health Organization (WHO), adverse drug reactions (ADRs) refer specifically to a series of harmful conditions that suddenly arise from qualified drugs during the process of following doctor's orders and normal use, such as side effects that are contrary to the original intention of treatment, worrying toxic reactions, allergic reactions that are difficult to prevent, and the far-reaching "three-cause effects" (causing fetal malformations, inducing gene mutations, inducing cancer, etc.), as well as lingering sequelae and secondary reactions like chain reactions.
[0003] Due to the harmfulness and severity of adverse drug reactions, measures need to be taken to strictly monitor the safety of marketed drugs. Once a highly suspected adverse reaction signal is found, it must be reported to the relevant units for research, analysis and management to reduce hidden dangers in medication.
[0004] In the field of drug side effect association (DSAs) prediction, there are currently three key technical approaches: docking-based methods, machine learning-based methods, and network-based methods.
[0005] (1) Docking-based methods: predict drug side effect associations (DSAs) by identifying the binding of ligands to target binding sites, performing docking predictions through steps such as molecular shape matching, mechanical optimization, and energy minimization, or attempting to dock single or multiple small molecules to receptor sites to find potential ligands.
[0006] (2) Machine learning-based methods: DSAs prediction is considered as a classification problem, and known DSAs are used as training data to train the classifier. Traditional machine learning methods include using small molecule chemical structure information to train random forest classifiers; deep learning methods include combining outer product operations, multi-layer perceptrons and convolutional neural networks to learn the representation of drugs and side effects, using deep graph convolutional neural networks and bidirectional long short-term memory networks to extract features and then use fully connected layers for prediction, and using deep neural network models to extract drug features for prediction.
[0007] (3) Network-based methods: Treat DSAs-related biological entities as nodes and connections as edges to construct a network, and model DSA prediction as a link prediction problem. One type of method represents the network with an adjacency matrix and uses matrix factorization and diffusion methods to learn drug and side effect representations; a projection correlation matrix is then introduced to construct a graph with graph regularization, and linear neighborhood similarity is used to measure drug similarity and propagate known side effect information based on the similarity graph for prediction. Another type of method applies graph neural networks to learn features, such as using graph attention networks to integrate data for feature learning, connecting graph embeddings and node embeddings as features and predicting through matrix completion.
[0008] These existing methods all have certain drawbacks or deficiencies. For example, they fail to fully utilize various data that may affect the occurrence of drug side effects, resulting in the inability to learn high-quality features to improve DSA prediction performance; or it is difficult to capture the complex semantic associations between biological entities during the occurrence of drug side effects. Different biological entities have heterogeneity, which poses a challenge to capturing the complex semantic associations between drugs and side effects; or, the predicted DSAs lack interpretability, only providing prediction results without clearly explaining the prediction basis, while interpretability is important for doctors' prescription decisions.
[0009] Given that the current prediction technology for drug side effect associations (DSAs) still needs to be optimized, in this study, a model based on heterogeneous information network (HIN) (MPGNN-DSA) was innovatively proposed specifically for DSA prediction. When constructing this heterogeneous information network (HIN), multiple DSA-related datasets were comprehensively integrated, just like building a "data hub" that converges information flows from all parties, capturing all data from different sources, deeply mining and making full use of the complex semantic associations contained therein, laying a solid foundation for subsequent accurate prediction. Further, a meta-path-based feature learning mechanism was carefully designed. This mechanism is like a precise "scalpel" that can skillfully cut into the complex structure of the heterogeneous information network (HIN), efficiently capturing the complex semantics hidden deep in the network through meta-paths. Just like exploring precious treasures along the hidden veins on an ancient map, high-quality representations were customized for drugs and side effects, making their features clearly presented. Moreover, meta-path splitting technology was adopted and the display features were finely feature-fused. Through this cutting-edge method, it is like replacing the "telescope" for prediction with a high-power lens, capturing semantic information of longer paths for the predicted drug side effects, enabling different meta-paths to communicate with each other instead of "acting independently" during the process of learning drug and side effect representations, gathering more and more valuable information, and thus significantly improving the accuracy and reliability of prediction. In this way, the predicted drug side effects based on this model were sorted according to the likelihood, and the TOP K most likely drug side effects were taken as the key focus objects, effectively realizing the forward-looking prediction of potential drug side effects, injecting new impetus into the field of pharmaceutical safety, and providing a solid and reliable decision-making reference. Summary of the Invention
[0010] Based on the deficiencies of the prior art, the present invention constructs a heterogeneous graph based on the MPGNN-DSA algorithm, and accurately predicts drug adverse reactions by performing explicit information transfer and feature fusion along carefully designed meta-paths in the graph.
[0011] The specific technical solution of the present invention is as follows:
[0012] In the first aspect of the present invention, a method for predicting adverse reactions based on drug symmetric meta-paths is provided, including the steps of:
[0013] 1) Collect data on the relationships among drugs Dr, proteins P, diseases Di, and adverse reactions A.
[0014] 2) Construct a basic model for predicting drug adverse reactions based on symmetric meta-paths, including the steps of:
[0015] 2.1 Construct a heterogeneous information network: According to the collected data, construct a relationship matrix to form a heterogeneous information network.
[0016] 2.2 Model training: Select symmetric long meta-paths and segment them from the inside out. Perform feature fusion on the segmented meta-paths. Based on the updated node features after feature fusion, reconstruct the relationship matrix between nodes to obtain the trained basic model for predicting adverse drug reactions based on symmetric meta-paths.
[0017] 3) Use the trained basic model for predicting adverse drug reactions based on symmetric meta-paths in step 2) to predict the relationship between drugs and adverse reactions.
[0018] The core of this prediction method lies in explicit information transfer and feature fusion along meta-paths to accurately predict adverse drug reactions, and this process depends on constructing a reasonable heterogeneous graph structure, accurately selecting meta-paths, and effectively fusing node features.
[0019] In the prediction method of the present invention, the explicit feature vector of a drug is mainly constructed based on the indication features of the drug and the bioinformatics features of the drug.
[0020] In some preferred embodiments, in step 1), the indication features of drugs are extracted through SIDER (Side Effect Resource); the bioinformatics features of drugs are extracted from the DrugBank database; the association feature information of drug-to-drug is collected from the STITCH database; the information related to protein-protein interactions is collected from the STRING database; the relationship information of drug-to-disease and protein-to-disease is collected from the public Comparative Toxicogenomics Database.
[0021] The SIDER database (Drug and Side Effects Resource Database) (see http: / / sideeffects.embl.de / ) contains information on marketed drugs and their recorded adverse reactions and indication information, and much of this information is extracted from public documents and materials through text mining techniques, mainly including drug side effect frequencies, classifications of drugs and side effects, and related information. The drug indication features extracted by the present invention based on the SIDER database can be denoted as V A-Dr 。
[0022] The DrugBank database (see https: / / go.drugbank.com / ) integrates bioinformatics and chemoinformatics resources and provides detailed drug data and comprehensive molecular information on drug targets and their mechanisms, including drug chemistry, pharmacology, pharmacokinetics, ADME, and their interaction information. The drug bioinformatics features extracted by the present invention based on the DmgBank database can be denoted as VA-P.
[0023] The STITCH database (see https: / / ngdc.cncb.ac.cn / databasecommons / database / id / 208) integrates data resources from multiple fields, covering not only key knowledge at the biological and chemical levels but also deeply associating the interaction information of drugs in actual application scenarios. From this database, we can accurately capture the synergistic effects, antagonistic relationships between drugs, and the detailed dynamics of their cooperation or mutual restraint in different disease treatment systems. The present invention is based on the drug-to-drug characteristics collected from the STITCH database and can be labeled as V Dr-Dr 。
[0024] The STRING database (see https: / / cn.string-db.org) is constructed relying on public databases and literature information. It collects the content of multiple public databases such as UniProt, KEGG, NCBI, and Gene Ontology. By integrating these data, a protein-protein interaction network database with a wide coverage and comprehensiveness is generated. The present invention is based on the protein-protein interaction information collected from the STRING database and can be labeled as V P-P 。
[0025] The Comparative Toxicogenomics Database (CTD, http: / / ctdbase.org / ) closely links toxicological information related to chemicals, genes, phenotypes, diseases, and exposures. By integrating the interaction content derived from the literature and carefully curated manually, it creates a rich knowledge base that can properly coordinate key information such as chemical exposure and its biological effects in cross-species heterogeneous data. The present invention is based on the drug-to-disease and protein-to-disease relationship information collected from the Comparative Toxicogenomics Database and can be respectively labeled as V Dr-Di 、V P-Di 。
[0026] In some preferred embodiments, the heterogeneous graph network adopted is composed of the relationships among drugs, targets, diseases, and adverse reactions. Therefore, first, according to the collected data, several key matrices such as drug-drug (Dr_Dr), drug-target (Dr_P), protein-protein (P_P), drug-disease (Dr_Di), drug-side effect (A_Dr), and protein-disease (P_Di) need to be constructed.
[0027] In some preferred embodiments, each bit of these relationship matrices consists of {0, 1}. In some specific embodiments, they are as follows:
[0028]
[0029] Matrix[x][y] represents the relationship between nodes. If Matrix[i][j] = 1, it means there is a relationship between node i and node j. If Matrix[i][j] = 0, it means there is no relationship between node i and node j.
[0030] In the present invention, it is considered that the semantic information features of long meta-paths should be aggregated, while the previous aggregation methods cannot capture the features of long paths. Therefore, a method of splitting according to symmetric long meta-paths is proposed. The meta-paths after splitting are subjected to feature fusion, and the method of feature fusion is improved. The symmetric long meta-path is denoted as MP.
[0031] In some preferred embodiments, the symmetric long meta-path is selected as MP = [[A-Dr, Dr-A], [[Di-Dr, Dr-Di], [P-Dr, Dr-P]]]; where A-Dr is the relationship between adverse reactions and drugs, Dr-A is the relationship between drugs and adverse reactions, Di-Dr is the relationship between diseases and drugs, Dr-Di is the relationship between drugs and diseases, P-Dr is the relationship between targets and drugs, and Dr-P is the relationship between drugs and targets.
[0032] According to MP, the long meta-path is divided into short paths Dr-Di-Dr, Dr-P-Dr, A-Dr-A. Among them, A-Dr-A is denoted as Taking Dr as the intermediate node of the aggregated short path, denoted as Am, and A as the starting node, denoted as A.
[0033] Among them, the type of the edge uses a mapping function.
[0034] The short paths after division are denoted as If it is divided into K short paths, the short path set is denoted as
[0035] The splitting of the long path is from the inside to the outside, and the fusion of node features is also from the inside to the outside. Among them, the target node is defined as the adverse reaction, and the type is denoted as As.
[0036] The feature fusion of the short paths after splitting is to take an intermediate node j of type Am and perform feature fusion along the short path set for feature fusion.
[0037] In some preferred embodiments, the feature fusion formula of the short path is as follows:
[0038]
[0039] Among them, the relationship of each neighbor of the target node i (i.e., As) is One of the intermediate nodes is , used to receive along The features of all neighbor nodes of type As on the relationship path, and this process is denoted as The accumulated features are fused with the original feature embedding of itself, and finally a fully connected layer is performed along the r1 path. At this time, the intermediate node type is A m has already fused the features of the target node and updated its own features.
[0040] After the above process, the short-path features have been fused and the intermediate node feature representation has been updated. Finally, the features of all nodes along the r1 path are fused to update the features of the final target node.
[0041] In some preferred embodiments, the formula for fusing the features of the nodes along the r1 path is as follows:
[0042]
[0043] where the intermediate node type is A m Performs a mean accumulation along all r1 paths, and visibly fuses multiple features transmitted and updated along the r2 path, denoted as There are K short meta-paths in the initial target node type As. The features along the set of these K short meta-paths are accumulated to the intermediate node, denoted as After all intermediate nodes receive the returned information, they fuse the features of the original target node type A, and finally perform a single layer of fully connected layer. At this time, the starting node of type As completes the feature fusion of the long path.
[0044] In some preferred embodiments, after the above-mentioned long and short path feature fusions, the final updated node features are obtained through the following feature fusion formula:
[0045]
[0046] where represents the set of neighbor nodes of node i, h i represents the current node i to be updated, represents the current node i being updated.
[0047] After updating the features of nodes of types Di, Dr, A, and P, it is necessary to reconstruct the relationship matrix between nodes.
[0048] Each type of node multiplies the updated node feature representation by a learnable diagonal matrix and then multiplies the transpose of the updated node feature representation matrix. The formula is as follows:
[0049] Matrix = XΛX T ;
[0050] Among them, according to the feature dimensions of all nodes, a diagonal matrix is generated, denoted as Λ. The value of Λ is randomly generated and conforms to a normal distribution with a mean of 0 and a standard deviation of 0.1. The node feature representation of each type of node after feature fusion through the long meta-path is denoted as X.
[0051] After the relationship matrices among drugs, diseases, targets, and adverse reactions are reconstructed, they are respectively as follows:
[0052]
[0053] Calculate the loss value after training for the relationship matrices among drugs, diseases, targets, and adverse reactions. The loss function used is as follows:
[0054] Matrix_loss = SUM[(Matrix - re_Matrix) 2 ,
[0055] where Matrix is the original feature matrix of this type of node, and re_Matrix is the feature matrix of this type of node after training.
[0056] In some preferred embodiments of the present invention, the relationship between drugs and adverse reactions is judged according to the value of the feature matrix re_A_Dr[x][y] of drugs and adverse reactions after training. For example, in some embodiments, the value of the i-th adverse reaction and the j-th drug in the matrix is re_A_Dr[i][j]. If the value is greater than the set threshold of 0.5, it indicates that there is a relationship between the i-th adverse reaction and the j-th drug, otherwise there is no such relationship. And when the value is larger, it indicates that the corresponding relationship is stronger, and they are sorted according to the value, and the top K adverse reactions with larger association strengths are selected as the potential adverse reactions of the drug.
[0057] The present invention also provides an adverse reaction prediction system based on drug symmetric meta-paths. This system is constructed based on the above prediction method and includes:
[0058] A model training module, which is used to collect data on the mutual relationships among drugs Dr, proteins P, diseases Di, and adverse reactions A, and train a basic model for predicting drug adverse reactions based on symmetric meta-paths; and
[0059] A prediction module, which is used to use the trained basic model for predicting drug adverse reactions based on symmetric meta-paths to predict the relationship between drugs and adverse reactions.
[0060] The beneficial effects of the present invention are:
[0061] The adverse drug reaction prediction method of the present invention is based on the hypothesis that the interaction of drugs with other elements (such as targets, diseases, etc.) in a complex network can reveal potential patterns of adverse drug reactions. It constructs a heterogeneous graph using multi-dimensional information of drugs, and through explicit information transfer and feature fusion along carefully designed meta-paths in the graph, optimizes the consideration of factors related to adverse drug reactions, finds nodes and their relationship paths that are semantically closely related to the target drug in the heterogeneous graph structure, and uses these associations to predict the potential association degree of the target drug with the target adverse reaction. Finally, according to the ranking of the association degrees of the target drug with each adverse reaction obtained by prediction, the TOPK adverse reactions with higher association degrees are selected as the potential adverse reactions of the target drug.
[0062] This method designs an information aggregation mechanism based on meta-paths, fully utilizes the complex semantic information in the HIN. Meta-path level segmentation not only helps to learn high-quality long-path node representations, but also can learn more representative node features, effectively capture semantic information, and improve the accuracy and stability of prediction. The AUC value of the prediction method of the present invention can reach 0.977, and the AUPR value can reach 0.841, which is better than all comparative baseline models. Moreover, compared with the relatively good comparative model simpleHGN, the standard deviation of the prediction results of this model is lower, indicating stronger prediction stability on different data samples. Brief Description of the Drawings
[0063] Figure 1 is a schematic diagram of the meta-path selection of the prediction method of the present invention;
[0064] Figure 2 is a schematic diagram of the prediction method flow of the present invention. Detailed Embodiments
[0065] The following combines the accompanying drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0066] Embodiment 1
[0067] As Figure 2 shown, the adverse drug reaction prediction method based on symmetric meta-paths includes the following steps:
[0068] S1 Collect data on the mutual relationships of drug Dr, protein P, disease Di, and adverse reaction A.
[0069] Data collection can be extracted and collected through existing databases.
[0070] The data in this embodiment is from the SIDER (Side Effect Resource) database (Drug and Side Effect Resource Database) to extract the indication characteristics of each drug; the DrugBank database (Drug Database) extracts the biological information characteristics of each drug, including targets, enzymes, transporters, carriers, etc.; the STITCH database collects the association characteristic information of each drug to drug; the STRING database collects protein interaction-related information; the public Comparative Toxicogenomics Database (CTD) collects the relationship information of drug to disease and protein to disease.
[0071] Based on these data, a drug-adverse reaction association matrix Dr_A is constructed, where the rows represent drugs, the columns represent adverse reactions, and each element value Dr_Aij represents the association strength of drug i to the adverse reaction. Denote the number of drugs as N and the number of adverse reactions as M. It can be understood that in the drug-adverse reaction association matrix Dr_A, some of the element values are 1, indicating that there is a relationship between the drug and the corresponding adverse reaction. Conversely, if the element value is 0, it means that there is no relationship between the drug and the corresponding adverse reaction. The goal of the method of the present invention is to predict the association strength of unknown drug-adverse reactions based on the known association strength of drug-adverse reactions, and sort them in descending order according to the magnitude of the association strength, and take the TOPK adverse reactions with larger association strengths as the potential adverse reactions of the drug, so as to realize the prediction of the potential adverse reactions of the drug.
[0072] Preferably, in this embodiment, 680 drug nodes, 732 target nodes, 5797 disease nodes and 2600 side effect nodes are extracted based on the DrugBank database, SIDER database, STITCH database, STRING database and CTD public comparative toxicogenomics database.
[0073] Preferably, in this embodiment, 2954 drug-target relationships are collected based on the DrugBank database, and a relationship matrix is constructed according to the relationships. Each bit of the matrix consists of {0, 1}, indicating whether the drug has the carrier protein / enzyme / target protein / transporter represented by this dimension.
[0074] Preferably, in this embodiment, 10821 drug-drug relationships are collected based on the STITCH database, and a relationship matrix is constructed according to the relationships. Each bit of the matrix consists of {0, 1}, indicating whether the drug has the drug represented by this dimension.
[0075] Preferably, in this embodiment, 69,544 protein-protein relationships are collected based on the STRING database, and a relationship matrix is constructed according to the relationships. Each bit of the matrix consists of {0, 1}, indicating whether the carrier protein / enzyme / target protein / transporter protein has the carrier protein / enzyme / target protein / transporter protein represented by this dimension.
[0076] Preferably, in this embodiment, 87,291 drug-disease relationships are collected based on the CTD database, and a relationship matrix is constructed according to the relationships. Each bit of the matrix consists of {0, 1}, indicating whether the drug has the disease represented by this dimension.
[0077] Preferably, in this embodiment, 432,662 protein-disease relationships are collected based on the CTD database, and a relationship matrix is constructed according to the relationships. Each bit of the matrix consists of {0, 1}, indicating whether the carrier protein / enzyme / target protein / transporter protein has the disease represented by this dimension.
[0078] Preferably, in this embodiment, the drug indication features extracted based on the SIDER database are a one-hot vector containing 17,393 high-frequency indications. Each bit of the vector consists of {0, 1}, indicating whether the drug has the indication represented by this dimension.
[0079] S2 Construct a basic model for predicting adverse drug reactions based on symmetric meta-paths.
[0080] 2.1 Construct a heterogeneous information network: According to the collected data, construct a mutual relationship matrix to form a heterogeneous information network.
[0081] Specifically, according to the above-collected data, in this embodiment, several key matrices of drug-drug (Dr_Dr), drug-target (Dr_P), protein-protein (P_P), drug-disease (Dr_Di), drug-side effect (A_Dr), protein-disease (P_Di) are constructed, as shown below:
[0082]
[0083] Matrix[x][y] represents the relationship between nodes. If Matrix[i][j] = 1, it means there is a relationship between node i and node j. If Matrix[i][j] = 0, it means there is no relationship between node i and node j.
[0084] Based on these mutual relationship matrices, a heterogeneous information network is formed.
[0085] 2.2 Model training; Select symmetric long meta-paths and perform segmentation from the inside out. Fuse the features of the segmented meta-paths, and based on the updated node features after feature fusion, reconstruct the relationship matrix between nodes to obtain the trained basic model for predicting adverse drug reactions based on symmetric meta-paths.
[0086] In this embodiment, the constructed symmetric long meta-path is MP = [[A-Dr, Dr-A], [[Di-Dr, Dr-Di], [P-Dr, Dr-P]]], that is, (A-Dr-Di-Dr-A, A-Dr-P-Dr-A), where A-Dr is the relationship between adverse reactions and drugs, Dr-A is the relationship between drugs and adverse reactions, Di-Dr is the relationship between diseases and drugs, Dr-Di is the relationship between drugs and diseases, P-Dr is the relationship between targets and drugs, and Dr-P is the relationship between drugs and targets.
[0087] The schematic diagram of meta-path segmentation is as Figure 1 shown. According to MP, the long meta-path is divided into short paths Dr-Di-Dr, Dr-P-Dr, and A-Dr-A. Among them, A-Dr-A is denoted as Take Dr as the intermediate node of the aggregated short paths, denoted as Am, and A as the starting node, denoted as As; among them, The short path is denoted as If it is divided into K short paths, the short path set is denoted as
[0088] In this embodiment, the long symmetric meta-path is divided into 2 short meta-paths Dr-Di-Dr and Dr-P-Dr.
[0089] The segmentation of the long path is from the inside out, and the fusion of node features is also from the inside out. Among them, the target node is defined as the adverse reaction, and the type is As.
[0090] The feature fusion of the segmented short paths is to fuse the features of a middle node j of type Am along the short path set In this embodiment, that is, to fuse the features of a middle node Dr along the short path set {Dr-Di-Dr, Dr-P-Dr}.
[0091] The short path feature fusion formula is as follows:
[0092]
[0093] Among them, the relationship of each neighbor of the target node i (i.e., As) is Among them, one middle node is , which is used to receive the features of all neighbor nodes of type As along the relationship path. This process is denoted as The fused features after accumulation are combined with their original feature embeddings, and then a fully connected layer is performed along the r1 path to complete the short-path feature fusion, i.e., the update of the intermediate node feature representation. Finally, the features of all nodes along the r1 path are fused to update the features of the final target node.
[0094] The formula for fusing the features of nodes along the r1 path is as follows:
[0095]
[0096] Among them, the intermediate node type Dr accumulates the mean along all r1 paths, and explicitly fuses multiple features passed and updated along the r2 path, denoted as There are K short meta-paths in the initial target node type A. The features along this set of K short meta-paths are accumulated to the intermediate node, denoted as After all intermediate nodes receive the returned information, they fuse the features of the original target node type A, and finally perform a single layer of fully connected operation. At this time, the starting node of type As completes the feature fusion of the long path.
[0097] After the above long- and short-path feature fusions, the final updated node features are obtained through the following feature fusion formula:
[0098]
[0099] Among them, represents the set of neighbor nodes of node i, and h i represents the current node i to be updated, represents the current node i being updated.
[0100] After updating the features of nodes of types Di, Dr, A, and P, the relationship matrix between nodes is reconstructed.
[0101] Each type of node multiplies the updated node feature representation by a learnable diagonal matrix and then multiplies by the transpose of the updated node feature representation matrix. The formula is as follows:
[0102] Matrix = XΛX T ;
[0103] Among them, according to the feature dimensions of all nodes, a diagonal matrix is generated, denoted as Λ. The values of Λ are randomly generated and follow a normal distribution with a mean of 0 and a standard deviation of 0.1. The node feature representation of each type of node after the feature fusion of the long meta-path is denoted as X.
[0104] Then, after the relationship matrix of drugs, diseases, targets, and adverse reactions is reconstructed, they are shown as follows respectively:
[0105]
[0106] Thus, a trained basic model for predicting adverse drug reactions based on symmetric meta-paths is obtained.
[0107] Calculate the loss value after training for the relationship matrix between drugs, diseases, targets, and adverse reactions. The loss function used is as follows:
[0108] Matrix_loss = SUM[(Matrix - re_Matrix) 2 ,
[0109] where Matrix is the original feature matrix of this type of node, and re_Matrix is the feature matrix of this type of node after training.
[0110] S3 uses the trained basic model for predicting adverse drug reactions based on symmetric meta-paths to predict the relationship between drugs and adverse reactions.
[0111] Judge the relationship between drugs and adverse reactions according to the value of the trained feature matrix re_A_Dr[x][y] of drugs and adverse reactions. For example, in some embodiments, the value of the i-th adverse reaction and the j-th drug in the matrix is re_A_Dr[i][j]. If the value is greater than the set threshold of 0.5, it indicates that there is a relationship between the i-th adverse reaction and the j-th drug; otherwise, there is no relationship. And when the value is larger, it indicates that the corresponding relationship is stronger, and they are sorted according to the value, and the top K adverse reactions with larger association strength are selected as the potential adverse reactions of the drug.
[0112] According to the method execution process provided in the embodiments of the present application, it can run on devices such as personal computers, servers, and cloud computing platforms.
[0113] Embodiment 2
[0114] This embodiment provides an adverse reaction prediction system based on drug symmetric meta-paths. This system is constructed based on the prediction method of Embodiment 1 and includes:
[0115] A model training module, configured to collect data on the mutual relationships among drugs Dr, proteins P, diseases Di, and adverse reactions A, and train a basic model for predicting adverse drug reactions based on symmetric meta-paths; and
[0116] A prediction module, configured to use the trained basic model for predicting adverse drug reactions based on symmetric meta-paths to predict the relationship between drugs and adverse reactions.
[0117] For some specific prediction systems in this embodiment, the AUC value can reach 0.977 and the AUPR value can reach 0.841, which are better than all the comparative baseline models. Moreover, compared with the relatively better comparative model simpleHGN, the standard deviation of the prediction results of this model is lower, indicating stronger prediction stability on different data samples.
[0118] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.
Claims
1. A method for predicting adverse reactions based on drug symmetric metapaths, characterized in that: Includes steps: 1) Collect data on the relationship between drug Dr, protein P, disease Di and adverse reaction A, 2) Constructing a basic model for predicting adverse drug reactions based on symmetric meta-paths, including the following steps: 2.1 Constructing a heterogeneous information network: Based on the collected data, construct a mutual relationship matrix to form a heterogeneous information network; 2.2 Model training: Select a symmetrical long meta-path to split it from the inside to the outside, perform feature fusion on the meta-path after segmentation, reconstruct the relationship matrix between nodes based on the updated node features after feature fusion, and obtain the trained basic model for drug adverse reaction prediction based on symmetrical meta-path; 3) The basic model for predicting adverse drug reactions based on symmetric meta-paths trained in step 2) is used to predict the relationship between drugs and adverse reactions.
2. The prediction method according to claim 1, characterized in that: In step 1), the SIDER database is used to extract the indication characteristics of the drug; the DrugBank database is used to extract the bioinformatics characteristics of the drug; the STITCH database collects the association characteristic information between drugs; and the STRING database collects the information related to protein interactions; The public comparative toxicogenomics database collects information on drug-to-disease and protein-to-disease relationships.
3. The prediction method according to claim 1, characterized in that: In step 2.1, the relationship matrix includes: drug-drug relationship matrix Dr_Dr, drug-target relationship matrix Dr_P, protein-protein relationship matrix P_P, drug-disease relationship matrix Dr_Di, drug-adverse reaction relationship matrix A_Dr, protein-disease relationship matrix P_Di.
4. The prediction method according to claim 1, characterized in that: Each bit of the mutual relationship matrix consists of {0, 1}.
5. The prediction method according to claim 1, characterized in that: In step 2.2: Select the symmetric long element path as MP = [[A-Dr, Dr-A], [[Di-Dr, Dr-Di], [P-Dr, Dr-P]]]; where A-Dr is the relationship between adverse reaction and drug, Dr-A is the relationship between drug and adverse reaction, Di-Dr is the relationship between disease and drug, Dr-Di is the relationship between drug and disease, P-Dr is the relationship between target and drug, and Dr-P is the relationship between drug and target; According to MP, the long path is divided into short paths Dr-Di-Dr, Dr-P-Dr, and A-Dr-A, where A-Dr-A is recorded as Let Dr be the intermediate node of the aggregated short path, denoted as Am, and A be the starting node, denoted as As; The shortest path is recorded as If it is divided into K short paths, the short path set is recorded as The short path feature fusion after segmentation is to combine an intermediate node j of type Am along the short path set Perform feature fusion.
6. The prediction method according to claim 5, characterized in that: The short path feature fusion formula is as follows: The relationship between each neighbor of the target node i is One of the intermediate nodes is , To receive The characteristics of all neighbor nodes of type As on the relationship path. This process is recorded as The accumulated features are fused with their original features and embedded, and finally fully connected along the r1 path to complete the short path feature fusion, that is, the update of the intermediate node feature expression. Finally, the features of all nodes along the r1 path are fused and the features of the final target node are updated.
7. The prediction method according to claim 6, characterized in that: The formula for feature fusion of nodes along the r1 path is as follows: Among them, the intermediate node Dr accumulates the mean along all r1 paths, which explicitly integrates multiple updated features transmitted along the r2 path, recorded as There are K short meta-paths at the starting node A. The set of paths along these K short meta-paths The features on the node are accumulated to the intermediate node, recorded as After receiving the returned information, all intermediate nodes fuse the features of the original target node type A and finally perform a full connection. At this time, the starting node of type As completes the feature fusion of the long path.
8. The prediction method according to claim 1, characterized in that: The updated node features are finally obtained through the following feature fusion formula: in, represents the set of neighbor nodes of node i, h i Indicates the node i to be updated currently. Represents the current node i to be updated.
9. The prediction method according to claim 1, characterized in that: In the step 3), the relationship between the drug and the adverse reaction is determined according to the value of the trained drug and adverse reaction feature matrix re_A_Dr[x][y].
10. A drug symmetric metapath-based adverse reaction prediction system, characterized in that: Based on any one of the methods of claims 1 to 9, the system comprises: A model training module is used to collect data on the relationship between drug Dr, protein P, disease Di and adverse reaction A, and train a basic model for predicting adverse reactions of drugs based on symmetric meta-paths; and The prediction module is used to predict the relationship between drugs and adverse reactions by using the trained drug-adverse reaction prediction basic model based on symmetric meta-paths.
Citation Information
Patent Citations
Metal path based drug-drug interaction
CN115512761A
Drug side effect prediction model based on meta-path graph neural network
CN115512857A
Microbial drug association relationship prediction method and system
CN118538298A
Patient-level adverse drug reaction prediction method and system based on graph neural network
CN118899096A
Prediction and generation of hypotheses on relevant drug targets and mechanisms for adverse drug reactions
US20190050538A1
Cited By
Heterogeneous graph neural network-based traditional Chinese medicine adverse reaction risk prediction method and system
CN120784007A