Heterogeneous graph neural network-based traditional Chinese medicine adverse reaction risk prediction method and system

By constructing a heterogeneous graph of traditional Chinese medicine, target and adverse reaction, and using a multi-head attention mechanism to automatically fuse multi-type node and relationship features, the problem of insufficient information fusion in existing technologies is solved, and a highly accurate and explainable prediction of the risk of adverse reactions to traditional Chinese medicine is achieved.

CN120784007APending Publication Date: 2025-10-14GUANGDONG PHARMA UNIV

Patent Information

Application Number
CN202510878351.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing methods for predicting the risk of adverse reactions to traditional Chinese medicine fail to fully integrate multi-source heterogeneous information, have insufficient feature extraction capabilities, find it difficult to capture deep semantic associations, and have a single model structure and lack interpretability, which affects the accuracy and reliability of predictions.

Method used

The HAPM algorithm based on heterogeneous graph neural network is adopted to construct a heterogeneous graph of traditional Chinese medicine-target-adverse reaction, and use the multi-head attention mechanism to automatically aggregate and fuse multi-type node and relationship features to achieve end-to-end information fusion and capture of complex nonlinear interactive relationships.

Benefits of technology

It significantly improves the accuracy and interpretability of risk prediction for adverse reactions to traditional Chinese medicine, provides reliable clinical decision support, and enhances the generalization and replicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120784007A_ABST
    Figure CN120784007A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine adverse reaction risk prediction method based on a heterogeneous graph neural network, and the method comprises the following steps: obtaining data, obtaining a traditional Chinese medicine-target relation file from an ETCM database, and obtaining an adverse reaction-target relation file from an ADReCS database; data preprocessing: performing data cleaning on the obtained relation file to obtain a preprocessed file; constructing an isomeric graph, taking the traditional Chinese medicine herb, the target spot and the adverse reaction adverse as three types of nodes, and constructing the isomeric graph by utilizing the pre-processing file; node features are initialized, initial feature vectors are constructed for each type of nodes herb, target and adverse, and initial embedding of all the nodes is mapped to the same dimension space; carrying out HAPM feature fusion, inputting the node features into an HAPM heterogeneous graph attention network to carry out message passing and aggregation, and outputting a representation vector of a node level; predicting and outputting; and performing verification and feedback iteration. The invention further provides a system adopting the method. The method and the system can accurately predict the adverse reaction risk of the traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of traditional Chinese medicine risk prediction technology and machine learning technology, and specifically relates to a method for predicting the risk of adverse reactions of traditional Chinese medicine based on heterogeneous graph neural networks. Background Art

[0002] Traditional Chinese medicine (TCM) is characterized by complex ingredients, multiple targets, and multiple mechanisms. The difficulty of studying its adverse drug reactions (ADRs) is much higher than that of single chemical drugs. Predicting the risk of adverse drug reactions to TCM is due to the complexity of TCM adverse drug reactions. On the one hand, TCM itself has complex ingredients, some of which contain multiple active ingredients and react complexly in the human body. Some potential toxicities are still unclear. At the same time, individual differences (such as different ages, physical conditions, and those with liver and kidney dysfunction have different sensitivities to drugs) and various factors (such as improper drug use, drug interactions, environmental pollution and counterfeit and inferior drugs, and psychological reactions during medication) can affect the occurrence of adverse reactions. On the other hand, it is to ensure drug safety, reduce the occurrence of adverse reactions, and promptly address potential problems. It is also to promote the rational use and development of TCM, provide clinicians with a reference for rational drug use, and promote the modernization of TCM. Its role is reflected in many aspects. It can provide support for clinical decision-making, optimize treatment plans, and improve diagnostic accuracy; it can help drug regulatory authorities strengthen drug quality control and improve regulatory policies; it can increase patients' awareness of the risks of adverse reactions to traditional Chinese medicine and promote communication between doctors and patients; it also has scientific research and academic value, can promote the development of related disciplines, accumulate research data, and provide a basis for the safety research of traditional Chinese medicine.

[0003] Traditional QSAR toxicity prediction methods based on chemical components can complete the toxicity risk assessment of chemical components of traditional Chinese medicine in a relatively short period of time by collecting existing toxicity data of traditional Chinese medicine and Western medicine, extracting molecular 2D descriptors, and using software such as MOE and WeKa to establish machine learning models such as KNN, RF, and SVM. Network pharmacology methods rely on public databases such as GeneCards, SymMap, and STRING to construct protein-protein interaction networks of adverse reaction targets and drug targets, identify key intersection targets through network topology analysis, and construct risk prediction models to achieve visual monitoring and risk control of traditional Chinese medicine ADRs. In recent years, data-driven methods based on deep learning have begun to be applied to the prediction of traditional Chinese medicine ADRs. Using spontaneous reports of adverse reactions to traditional Chinese medicines as the data source, binary feature vectors are generated through one-hot encoding, and a feedforward deep neural network (DNN) model is constructed to predict the adverse reactions that may be caused by a single traditional Chinese medicine.

[0004] Specifically, existing methods for predicting the risk of adverse reactions of traditional Chinese medicine include: (1) TCM chemical component prediction method: With the chemical components of traditional Chinese medicine as the core, a chemical database is constructed by collecting data related to the toxicity of traditional Chinese medicine; molecular 2D descriptors are extracted and descriptors are screened and optimized using software such as MOE and WeKa; QSAR toxicity prediction models are established using machine learning algorithms such as KNN, RF, and SVM; the models are internally cross-validated and externally tested, and are ultimately used to predict the toxicity risks that may be caused by each chemical component of traditional Chinese medicine. (2) Network pharmacology prediction method: an adverse reaction target database is constructed using target information related to adverse reactions of traditional Chinese medicine and patient adverse reaction reports; positive / negative drug and target data are obtained by comparing literature and databases (such as GeneCards, SymMap, etc.); a protein interaction network is constructed, and key intersection targets are determined using network topology analysis; a risk prediction model is constructed based on the intersection target parameters to achieve visualization and monitoring of the adverse reaction risks of traditional Chinese medicine. (3) Data-driven prediction method based on deep learning: Using spontaneous reports of adverse reactions to traditional Chinese medicine as the data source, extract the characteristics of traditional Chinese medicine and adverse reactions in the reports; use one-hot encoding to generate binary feature vectors and construct a data matrix as training data; use deep neural network (DNN) to build a feedforward neural network model (including input layer, hidden layer and output layer); through hyperparameter debugging and model training, realize the prediction of possible adverse reactions caused by a single traditional Chinese medicine. (4) Drug adverse reaction prediction method and system based on deep learning heterogeneous network: mainly use drug molecular structure information and related physical and chemical properties, extract drug molecular features through molecular fingerprints or descriptors, and combine graph neural network to complete drug adverse reaction prediction; directly extract node embedding of drugs and adverse reactions through multi-layer GCN; mainly face molecular atomic graph (input is a single type of node: atom), calculate the attention weight of atom-atom adjacency, and complete the feature extraction of drug molecular structure; focus more on single type of node (such as atoms in drug molecules) in a relatively single graph structure for attention aggregation. (5) Drug adverse reaction prediction method based on review text information enhancement: focuses on using drug review text information and patient characteristics to build a review text enhancement model and perform autoregressive generation prediction on drug adverse reactions; converts drug review text into word vectors and combines TransformerDecoder for autoregressive prediction; based on the word vector representation of the review text, generative prediction is directly performed on the drug and adverse reaction word vectors through the TransformerDecoder module; TransformerDecoder is used to generate adverse reaction information through autoregressive prediction, which usually requires the adverse drug reaction to be constructed in the form of a "sentence" and relies on a predefined end word to terminate the generation process.

[0005] Although the above methods have improved the efficiency and accuracy of TCM ADR risk prediction to a certain extent, there are still several shortcomings. First, most existing technologies only model a single data type (such as chemical structure, target or spontaneous report), and fail to fully integrate multi-source heterogeneous information such as TCM, molecular targets and clinical adverse reactions, which limits the ability to express the complex mechanism of action of TCM. Second, these methods often rely on manually constructed descriptors or simple encodings in the feature extraction process, which makes it difficult to capture potential high-order semantic associations, resulting in insufficient feature representation capabilities. Finally, the traditional model structure is relatively simple, and the ability to model multi-level and nonlinear interactions is limited. In addition, deep learning models generally lack interpretability, which is not conducive to revealing the basis for prediction and cannot provide reliable mechanism support for clinical decision-making. Most existing methods fail to fully integrate multi-source heterogeneous data, resulting in deficiencies in modeling complex associations between TCM, targets and adverse reactions, which in turn affects the accuracy of prediction. In addition, existing methods often rely on manually constructed descriptors or simple encodings in feature extraction, which makes it difficult to capture the deep semantic information implicit in the data and cannot learn high-quality features. Moreover, the model structure of traditional methods is relatively simple, with limited ability to express multi-level, nonlinear interactions between heterogeneous biological entities, and insufficient support for the interpretability of prediction results.

[0006] Limitations of the drug adverse reaction prediction method and system based on deep learning heterogeneous network: (1) The model uses multi-layer GCN to directly extract the embedding of drug and adverse reaction nodes, but it is only oriented to the atomic graph at the molecular level (the input is a single type of atomic node). Therefore, it only calculates the attention weights between atoms and cannot reflect the complex interactions between the drug as a whole and other biological entities; (2) This method focuses more on the aggregation of homogeneous information and the attention calculation of a single type of node in a single graph structure, but fails to fully capture the multi-level and nonlinear interaction relationships between multiple heterogeneous entities, thereby limiting the model's predictive ability and generalization performance in complex environments.

[0007] Limitations of the adverse drug reaction prediction method based on enhanced review text information: (1) Review text data dependency and noise issues: This method focuses on using drug review text information and patient characteristics to build an enhanced prediction model. The core of this method is to convert the review text into a word vector and then use TransformerDecoder to generate predictions through autoregression. However, review text data often has problems such as high noise, strong subjectivity, and inconsistent format. It is difficult to completely filter out irrelevant information in the preprocessing stage, which affects the accuracy and stability of the generative prediction. (2) Limitations of the generative prediction method: The TransformerDecoder module performs autoregressive predictions, which requires the adverse drug reactions to be constructed in the form of "sentences" and relies on predefined end words to terminate the generation process. This limits the types and range of adverse reactions predicted by the model to a certain extent. In addition, this method lacks an effective dynamic feedback and iterative update mechanism, making it difficult to feed back the verification results to the model in a timely manner, thereby limiting the ability to capture new or unknown adverse reactions, and thus affecting the overall prediction performance and generalization ability. The model structure is to construct adverse drug reactions in the form of sentences and use TransformerDecoder for autoregressive generation. This method focuses on leveraging the generative capabilities of language models and capturing the dependencies between words through the attention mechanism, but it mainly focuses on homogeneous text information, and the generation and prediction process requires pre-definition of the end word, which limits the prediction scope.

[0008] Therefore, there is an urgent need for a method that can effectively predict the adverse reactions of traditional Chinese medicine. Summary of the Invention

[0009] The purpose of the present invention is to provide an accurate and reliable method for predicting the risks of adverse reactions to traditional Chinese medicines in response to the above technical problems.

[0010] In order to achieve the above invention objectives, the present invention provides a method for predicting the risk of adverse reactions of traditional Chinese medicine based on heterogeneous graph neural networks (abbreviated as HAPM algorithm, i.e. Heterogeneous Graph Attention Propagation Module). This method is based on the assumption that the deep semantic association between traditional Chinese medicine and various biological entities (such as targets, adverse reactions, etc.) in a heterogeneous graph can reveal its potential adverse reaction patterns. By integrating three types of nodes, namely traditional Chinese medicine, targets and adverse reactions and their various relationships, a knowledge graph containing heterogeneous node types and multiple edge types is constructed. The method adopts the heterogeneous graph attention network module in HAPM to automatically aggregate and fuse multiple types of nodes and their relationship features along predefined meta-paths in the constructed knowledge graph, and efficiently captures the deep semantic associations between traditional Chinese medicine, targets and adverse reactions through a multi-head attention mechanism. In the model inference stage, the target traditional Chinese medicine and each adverse reaction are ranked according to the prediction score, and the top-K items with the highest correlation are selected as candidate adverse reactions. During the validation phase, the prediction results were verified using the book "Common Adverse Drug Reactions and Treatments—Traditional Chinese Medicine" (edited by Lin Jun, Liu Naxin, and Ou Xiaolong, published by the Military Medical Science Press). The confirmed relationships were dynamically fed back into the knowledge graph to expand the training set, and training was iterated to continuously improve model performance. This invention can automatically integrate multi-source heterogeneous information end-to-end and capture complex nonlinear interactions, significantly improving the accuracy and interpretability of TCM adverse reaction risk predictions and providing reliable technical support for the safe use of TCM.

[0011] The present invention provides a method for predicting the risk of adverse reactions of traditional Chinese medicine based on the HAPM algorithm. The core of the method is to construct a heterogeneous graph of traditional Chinese medicine-target-adverse reaction, and use a multi-head attention mechanism guided by predefined meta-paths to automatically aggregate and fuse multi-type node and relationship features in the graph, thereby efficiently capturing the deep semantic associations between traditional Chinese medicine, targets and adverse reactions to achieve accurate risk prediction.

[0012] In this paper, the explicit feature vectors of TCM nodes are composed of two types of association information: one is derived from the known association between TCM and its target, which is used to reflect the molecular mechanism of action of TCM; the other is derived from the association between adverse reactions and biological targets, which is used to reveal the molecular basis of adverse reactions. By integrating these two types of information, the semantic expression ability of TCM nodes in graph neural networks can be effectively enhanced, thereby improving the accuracy and reliability of adverse reaction risk prediction, as follows:

[0013] Preferably, the present invention collects information on the association between traditional Chinese medicines and targets from the ETCM database. The ETCM (Encyclopedia of Traditional Chinese Medicine) database integrates a large amount of data on the interactions between traditional Chinese medicines and their targets, verified by literature and experiments, covering the targets of commonly used traditional Chinese medicines and their active ingredients in vivo. This database can provide accurate traditional Chinese medicine-target mapping relationships, providing basic data for constructing a heterogeneous map of traditional Chinese medicine-target-adverse reaction. The present invention is based on the traditional Chinese medicine-target relationships collected by the ETCM database.

[0014] Preferably, the present invention collects the association relationship information of adverse reactions to targets from the ADReCS database. The ADReCS (Adverse Drug Reaction Classification System) database brings together a large amount of adverse reaction-target association data obtained from clinical reports and literature mining, covering adverse reactions induced by different drugs and their molecular mechanisms. Through the target mapping provided by ADReCS, the present invention can accurately obtain the molecular target relationship corresponding to the adverse reaction. The present invention is based on the adverse reaction to target relationship collected by the ADReCS database.

[0015] The present invention provides a method for predicting the risk of adverse reactions of traditional Chinese medicine based on the HAPM algorithm, which mainly includes the following steps:

[0016] S1. Obtain data: obtain the relationship files of traditional Chinese medicines and targets from the ETCM database and the relationship files of adverse reactions and targets from the ADReCS database;

[0017] S2. Data preprocessing: cleaning the acquired relational files to obtain preprocessed files.

[0018] S3. Construct a heterogeneous graph, taking herb, target and adverse reaction as three types of nodes, and use the preprocessing file to construct the heterogeneous graph;

[0019] S4, node feature initialization, constructing the initial feature vector for each node herb, target, and adverse, and all node initial embeddings are mapped to the same dimensional space;

[0020] S5, HAPM feature fusion, inputs node features into the HAPM heterogeneous graph attention network for message passing and aggregation, and outputs node-level representation vectors;

[0021] S6. Prediction and output: Use the link prediction module to calculate the prediction scores for all (herb, adverse) relationship pairs, sort them in descending order, and select the top-K items with the highest scores as candidate adverse reactions;

[0022] S7. Verification and feedback iteration: The candidate TCM-adverse reaction pairs are compared and verified with the "Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume". If certain (herb, adverse) relationships are confirmed, they will be added to the training data to further improve the model performance in the next round of iteration.

[0023] Specifically, step S1, constructing a heterogeneous graph: taking the traditional Chinese medicine node (herb), target node (target) and adverse reaction node (adverse) as three heterogeneous node types, respectively obtaining the association information between traditional Chinese medicine and target from the ETCM database, and obtaining the association information between adverse reaction and target from the ADReCS database.

[0024] Preferably, in step S2, first, the traditional Chinese medicine-target relationship file provided in the ETCM database and the adverse reaction-target relationship file provided in the ADReCS database are preliminarily integrated, duplicate records are deleted, and data with missing key fields are eliminated to ensure the integrity and consistency of the original data; secondly, the field names in the two data sources are unified, and the data format of each field is standardized; thirdly, the integrated data is verified, the logical consistency of each record is checked, and abnormal data is eliminated or corrected to ensure that each relationship data meets the requirements for constructing a heterogeneous graph; finally, according to the preprocessing requirements, the cleaned and standardized data is converted into a preprocessing file in a unified format.

[0025] Preferably, in step S3, the heterogeneous graph is represented as in

[0026] V={herb,target,adverse},

[0027] E={(herb,ht,target),(target,ta,adverse)}

[0028] By combining two paths, a composite relationship (herb, ind, adverse) is automatically generated to represent the path from Chinese medicine to target to adverse reaction; where herb represents an important node, target represents a target node, and adverse represents an adverse reaction node. represents the constructed heterogeneous graph, V represents the node set, E represents the edge set, ht represents the known association between traditional Chinese medicine and target (herb→target), ta represents the known association between adverse reaction and target (target→adverse), and ind represents the compound association between traditional Chinese medicine and adverse reaction (herb→target→adverse).

[0029] Specifically, step S4, node feature initialization: construct an initial feature vector for each node v∈V={herb, target, adverse}, and define the following formula:

[0030]

[0031] Among them, EmbedInit means mapping the entity information corresponding to the node v to a real vector of fixed dimension d; for the herb node, EmbedInit extracts its Chinese medicine-target association features from ETCM; for the adverse node, EmbedInit extracts its adverse reaction-target association features from ADReCS; for the target node, EmbedInit integrates its protein function, pathway annotation and other biological information; all nodes are initially embedded in the same dimensional space. Here, h represents the feature vector of a node, the subscript v indicates that this is the vector for node v in the graph, and the superscript (0) represents the "initial" layer, that is, the embedding of the 0th layer. represents the set of real numbers, Represents a real vector space of length d, where d represents the dimension of the unified mapping.

[0032] Preferably, in step S4, the "verified Chinese medicine-adverse reaction relationship" or other sorted Chinese medicine adverse reaction information is incorporated into the EmbedInit(·) function; for the herb node, if certain adverse reactions have been confirmed in the previous steps or external literature, the corresponding dimension is added to its initial feature vector to record the number, type or severity of the verified adverse reactions of the Chinese medicine; for the adverse node, the historical frequency of the adverse reaction or the confirmed clinical hazard level is introduced.

[0033] Specifically, step S5, HAPM feature fusion: In the constructed heterogeneous graph, define the HAPM attention module to aggregate node features; for the lth layer, the representation of each target node t is determined by the messages transmitted by all its connected neighbor nodes s through the predefined meta-path p; the attention weight of the hth attention head is calculated as follows:

[0034]

[0035] in are the query and key mapping matrices of the h-th head, d h For single head dimension, and represent the neighbor and target node embeddings of the previous layer respectively.

[0036] The message vector is obtained by weighted summation:

[0037]

[0038] in is the value mapping matrix, “||” represents the head-level splicing operation, represents the set of neighbors reachable along the meta-path p = (herb, ind, adverse), represents the attention weight from neighbor node s to target node t calculated by the h-th attention head, represents the embedding vector (node ​​representation) of the neighbor node s in the (l-1)th layer, that is, the feature representation of node s after aggregation and activation in the previous layer, which is used for weighted message passing in the current layer. Finally, the target node representation is updated by fusion with ReLU activation through residual connection:

[0039]

[0040] in is the output mapping matrix, Indicates the final output embedding vector of the target node t in the lth layer after attention aggregation, residual connection and ReLU activation in this layer, represents the input embedding vector of the target node t in the (l-1)th layer, It represents the message vector (the weighted result of each head) obtained by weighted summation of the target node t from all neighbor nodes that meet the meta-path p in the lth layer, which is used to concatenate with the embedding of the previous layer and then perform mapping update.

[0041] Specifically, step S6, link prediction and score calculation: after the output of the Lth layer, focus on the association strength between the Chinese medicine node (herb) and the adverse reaction node (adverse), and embed the Chinese medicine output of the final layer into Embedded with adverse reactions After concatenation, input the two-layer fully connected network and calculate the association score:

[0042]

[0043] in is a trainable weight matrix, σ is a Sigmoid activation function, which maps the output to the interval [0,1], indicating the risk association between Chinese medicine and adverse reactions. When representing the output of the Lth layer, the final embedding vector of the Chinese medicine node h after all attention heads are aggregated, residual connections and ReLU activation is used to characterize the semantic representation of the Chinese medicine node in the heterogeneous graph. When representing the output of the Lth layer, the final embedding vector of the adverse reaction node a after the same attention aggregation, residual connection and ReLU activation is used to characterize the semantic representation of the adverse reaction node in the heterogeneous graph.

[0044] Specifically, in step S7, the predicted probability calculated for each pair of Chinese medicine-adverse reaction (h, a) is Compared with the preset threshold τ, when , match its candidate relationships with known records in Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume to construct a verification set:

[0045]

[0046] Where ClinicalRef represents a list of manually verified TCM-adverse reaction pairs. The verified relationships are then fed back into the training graph, updating the training edge set to:

[0047]

[0048] Thereby expanding the scale of training samples and improving the model representation; repeating steps S3 to S6 until the verification indicators (F1, AUROC, AUPR, etc.) converge and then stopping the iteration.

[0049] Preferably, the method of the present invention further includes a loss function and training optimization: in order to make the prediction score of the positive sample significantly higher than that of the negative sample, an interest boundary parameter ω is introduced into the loss function m (in Associated with the i-th Chinese medicine node), the loss function is defined as:

[0050]

[0051] where y i is the sample label (1 for positive sample, 0 for negative sample). This loss function is based on point-by-point loss fitting through the auxiliary parameter ω m Adjust the score boundary between positive and negative samples to take into account the effects of single sample classification and positive and negative sorting.

[0052] The whole model uses the AdamW optimizer to train all parameters including the HAPM module and the link predictor end to end, and combines the CosineAnnealingLR learning rate scheduler to gradually reduce the learning rate to achieve stable and efficient convergence. During the training process, the model constantly adjusts the hyperparameters using the validation set feedback, and finally outputs the risk score of traditional Chinese medicine-adverse reactions, providing support for the subsequent Top-K screening of candidate relationships and clinical verification.

[0053] In another aspect, the application also provides a traditional Chinese medicine adverse reaction risk prediction system based on a heterogeneous graph neural network, which uses the method of the application, comprising:

[0054] A data acquisition module is used to acquire traditional Chinese medicine-target relationship files from the ETCM database and adverse reaction-target relationship files from the ADReCS database;

[0055] A data preprocessing module is used to clean the acquired relationship files and obtain preprocessed files;

[0056] A heterogeneous graph construction module is used to construct a heterogeneous graph by taking traditional Chinese medicine herb, target and adverse reaction as three types of nodes and using the preprocessed files;

[0057] A node feature initialization module is used to construct an initial feature vector for each type of node herb, target and adverse, and all node initial embeddings are mapped to the same dimensional space;

[0058] A feature fusion module is used to input node features into the HAPM heterogeneous graph attention network for message passing and aggregation, and output node-level representation vectors;

[0059] A prediction module is used to calculate the prediction scores of all (herb, adverse) relationship pairs using the link prediction module, sort them in descending order of scores, and select the top-K items with the highest scores as candidate adverse reactions;

[0060] A result output module is used to output the predicted candidate adverse reaction results;

[0061] A verification and feedback iteration module is used to verify the candidate adverse reaction results, add the results verified as true to the training data, and perform the next round of iteration.

[0062] Compared with the prior art, the method of the application has the following beneficial effects:

[0063] (1) Deep fusion of multi-source heterogeneous information - simultaneous integration of two types of relationships: traditional Chinese medicine → target and adverse reaction → target. Based on a complete heterogeneous knowledge graph, it breaks through the limitation of a single data source and realizes the end-to-end fusion of information on traditional Chinese medicine, biological targets and adverse reactions.

[0064] (2) Strong ability to capture high-order semantic associations - Through the multi-head attention mechanism guided by predefined meta-paths, it automatically learns complex nonlinear interactions between different types of nodes, significantly improving the quality of feature representation and avoiding the limitations of relying on manual descriptors or simple encoding.

[0065] (3) End-to-end automated training framework: No need to manually design meta-paths or manually screen features, it can learn directly from the cleaned graph data, greatly reducing the dependence on the experience of domain experts and enhancing the replicability and scalability of the method.

[0066] (4) Excellent predictive performance and robustness - the results of five-fold cross-validation show that the present invention surpasses traditional QSAR, network pharmacology and simple DNN methods in many indicators such as Precision, Recall, F1, AUROC, AUPR, etc., and remains stable under different sampling depths and relationship type missing.

[0067] (5) Strong interpretability - the key association paths between traditional Chinese medicine, target and adverse reactions are intuitively revealed through attention weights, providing a transparent and credible mechanism basis for clinical decision-making.

[0068] (6) Dynamic iterative feedback mechanism - Feedback the verification results of "Common Adverse Drug Reactions and Treatment - Traditional Chinese Medicine Volume" to the training set, continuously expand the positive samples, and effectively improve the model's generalization ability and long-term prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a flow chart of the method for predicting the risk of adverse reactions to traditional Chinese medicine based on heterogeneous graph neural networks of the present invention.

[0070] Figure 2 It is a schematic diagram of the method for predicting the risk of adverse reactions to traditional Chinese medicine based on heterogeneous graph neural network of the present invention. DETAILED DESCRIPTION

[0071] The present invention will be further described below with reference to specific examples. It should be understood that the following examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0072] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting the risk of adverse reactions of traditional Chinese medicine based on a heterogeneous graph neural network, comprising:

[0073] S1. Obtain data: obtain the relationship files of traditional Chinese medicine-target from the ETCM database, and obtain the relationship files of adverse reaction-target from the ADReCS database.

[0074] Preferably, the present invention focuses on collecting the association relationship information between traditional Chinese medicine and targets from the ETCM (Encyclopedia of Traditional Chinese Medicine) database. The ETCM database integrates a large amount of data on the interaction between traditional Chinese medicine and its targets verified by literature and experiments, covering the targets of commonly used traditional Chinese medicines and their active ingredients in organisms. The database can provide accurate traditional Chinese medicine-target mapping relationships and provide basic data for constructing traditional Chinese medicine-target-adverse reaction heterogeneous graphs. The present invention is based on the traditional Chinese medicine-target relationship files collected by the ETCM database to support subsequent heterogeneous graph construction and feature extraction.

[0075] Preferably, the present invention collects the association relationship information between adverse reactions and targets from the ADReCS (Adverse Drug Reaction Classification System) database. The ADReCS database brings together a large amount of adverse reaction-target association data obtained from clinical reports and literature mining, covering adverse reactions induced by different drugs and their molecular mechanisms. Through the mapping provided by ADReCS, the present invention can accurately obtain the target information corresponding to each adverse reaction. The present invention is based on the adverse reaction-target relationship files collected by the ADReCS database, providing a key biological basis for the subsequent construction of heterogeneous graphs and risk prediction.

[0076] S2. Data preprocessing: performing data cleaning on the acquired relational files to obtain preprocessed files.

[0077] First, the TCM-target relationship files provided in the ETCM database and the adverse reaction-target relationship files provided in the ADReCS database were preliminarily integrated, duplicate records were deleted, and data with missing key fields (such as herb_id, target_id, adverse_id) were eliminated to ensure the integrity and consistency of the original data.

[0078] Secondly, the field names in the two data sources were unified. For example, the TCM identifier was standardized as "herb_id," the target identifier was standardized as "target_id," and the adverse reaction identifier was standardized as "adverse_id." At the same time, the data format of each field was standardized, including character encoding conversion, removal of redundant whitespace, and capitalization.

[0079] Again, the integrated data is verified to check the logical consistency of each record, and abnormal data is removed or corrected to ensure that each relational data meets the requirements for building a heterogeneous graph.

[0080] Finally, according to the preprocessing requirements, the cleaned and standardized data is converted into a preprocessing file in a unified format (such as CSV or JSON format) to provide high-quality data support for subsequent steps (such as heterogeneous graph construction, node feature initialization and model training).

[0081] S3. Construct a heterogeneous graph, taking herb, target and adverse reaction as three types of nodes, and use the preprocessing file to construct the heterogeneous graph.

[0082] The heterogeneous graph is represented as in

[0083] V={herb,target,adverse},

[0084] E={(herb,ht,target),(target,ta,adverse)}

[0085] By combining two paths, a composite relationship (herb, ind, adverse) is automatically generated to represent the path from Chinese medicine to target to adverse reaction. Among them, herb represents an important node, target represents a target node, and adverse represents an adverse reaction node. represents the constructed heterogeneous graph, V represents the node set, E represents the edge set, ht represents the known association between traditional Chinese medicine and target (herb→target), ta represents the known association between adverse reaction and target (target→adverse), and ind represents the compound association between traditional Chinese medicine and adverse reaction (herb→target→adverse).

[0086] S4. Node feature initialization: construct an initial feature vector for each node (herb, target, adverse), and all node initial embeddings are mapped to the same dimensional space.

[0087] Construct the initial feature vector for each node v∈V={herb,target,adverse} and define the following formula:

[0088]

[0089] Among them, EmbedInit means mapping the entity information corresponding to the node v to a real vector of fixed dimension d; for the herb node, EmbedInit extracts its Chinese medicine-target association features from ETCM; for the adverse node, EmbedInit extracts its adverse reaction-target association features from ADReCS; for the target node, EmbedInit integrates its protein function, pathway annotation and other biological information. All nodes are initially embedded in the same dimensional space. Here, h represents the feature vector of a node, the subscript v indicates that this is the vector for node v in the graph, and the superscript (0) represents the "initial" layer, that is, the embedding of the 0th layer. represents the set of real numbers, Represents a real vector space of length d, where d represents the dimension of the unified mapping.

[0090] Preferably, the present invention can also incorporate "verified traditional Chinese medicine-adverse reaction relationship" or other sorted traditional Chinese medicine adverse reaction information into the EmbedInit(·) function. For the herb node, if certain adverse reactions have been confirmed in the previous steps or external literature, the corresponding dimension can be added to its initial feature vector to record the number, type or severity of the verified adverse reactions of the traditional Chinese medicine; for the adverse node, the historical frequency of occurrence of the adverse reaction or the confirmed clinical hazard level and other statistics can be introduced. By splicing such additional information into the basic vector and then mapping it to a unified dimensional space, the ability of the node features to express real biological and clinical significance can be further enhanced.

[0091] Through this approach, the present invention ensures that EmbedInit(·) not only includes basic correlation features from the ETCM and ADReCS databases, but also incorporates supplementary information such as verified TCM-adverse reaction relationships, thereby providing a more complete node representation for the heterogeneous graph attention network. In subsequent steps, these node embeddings will participate in the message transmission and aggregation of the multi-head attention mechanism, achieving accurate prediction of TCM adverse reaction risks.

[0092] S5, HAPM feature fusion, inputs the node features into the HAPM heterogeneous graph attention network for message passing and aggregation, and outputs the node-level representation vector.

[0093] In the constructed heterogeneous graph, the HAPM attention module is defined to aggregate node features. For the lth layer, the representation of each target node t is determined by the messages transmitted by all its connected neighbor nodes s through the predefined meta-path p. The attention weight of the hth attention head is calculated as follows:

[0094]

[0095] in are the query and key mapping matrices of the h-th head, d h For single head dimension, and Represent the neighbor and target node embeddings of the previous layer respectively. The message vector is obtained by weighted summation:

[0096]

[0097] in is the value mapping matrix, “||” represents the head-level splicing operation, represents the set of neighbors reachable along the meta-path p = (herb, ind, adverse). represents the attention weight from neighbor node s to target node t calculated by the h-th attention head, represents the embedding vector (node ​​representation) of the neighbor node s in the (l-1)th layer, that is, the feature representation of node s after aggregation and activation in the previous layer, which is used for weighted message passing in the current layer. Finally, the target node representation is updated by fusion with ReLU activation through residual connection:

[0098]

[0099] in is the output mapping matrix, Indicates the final output embedding vector of the target node t in the lth layer after attention aggregation, residual connection and ReLU activation in this layer, represents the input embedding vector of the target node t in the (l-1)th layer, It represents the message vector (the weighted result of each head) obtained by weighted summation of the target node t from all neighbor nodes that meet the meta-path p in the lth layer, which is used to concatenate with the embedding of the previous layer and then perform mapping update.

[0100] S6. Prediction and output: Use the link prediction module to calculate the prediction scores for all (herb, adverse) relationship pairs, sort them in descending order, and select the top-K items with the highest scores as candidate adverse reactions.

[0101] After the Lth layer output, the present invention focuses on the association strength between the Chinese medicine node (herb) and the adverse reaction node (adverse), and embeds the Chinese medicine output of the final layer into Embedded with adverse reactions After concatenation, input the two-layer fully connected network and calculate the association score:

[0102]

[0103] in is a trainable weight matrix, σ is a Sigmoid activation function, which maps the output to the interval [0,1], indicating the risk association between Chinese medicine and adverse reactions. When representing the output of the Lth layer, the final embedding vector of the Chinese medicine node h after all attention heads are aggregated, residual connections and ReLU activation is used to characterize the semantic representation of the Chinese medicine node in the heterogeneous graph. When representing the output of the Lth layer, the final embedding vector of the adverse reaction node a after the same attention aggregation, residual connection and ReLU activation is used to characterize the semantic representation of the adverse reaction node in the heterogeneous graph.

[0104] S7. Verification and feedback iteration: The candidate TCM-adverse reaction pairs are compared and verified with the "Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume". If certain (herb, adverse) relationships are confirmed, they will be added to the training data to further improve the model performance in the next round of iteration.

[0105] The predicted probability calculated for each pair of Chinese medicine-adverse reaction (h, a) Compared with the preset threshold τ, when , match its candidate relationships with known records in Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume to construct a verification set:

[0106]

[0107] Where ClinicalRef represents a list of manually verified TCM-adverse reaction pairs. The verified relationships are then fed back into the training graph, updating the training edge set to:

[0108]

[0109] This will expand the size of the training sample and improve the model representation. Repeat steps S3 to S6 until the verification indicators (F1, AUROC, AUPR, etc.) converge and stop the iteration.

[0110] However, the existing point-by-point loss and pair-wise loss each have limitations when dealing with TCM-adverse reaction risk prediction: point-by-point loss can directly fit the label of a single sample, but often cannot fully utilize the relative difference between positive and negative samples; while pair-wise loss can highlight the ranking relationship between positive and negative samples, it places higher requirements on sample construction and model training. In order to take into account the advantages of both without increasing the training complexity too much, the present invention further proposes a hybrid loss function and introduces an auxiliary parameter ω into it. m (Interest boundary) is used to adaptively adjust the boundary between positive and negative sample scores within the model, thereby improving the ability to distinguish high-risk adverse reactions. Based on this, the present invention defines the loss function as:

[0111]

[0112] where y i is the sample label (1 for positive sample, 0 for negative sample). This loss function is based on point-by-point loss fitting through the auxiliary parameter ω m Adjust the score boundary between positive and negative samples to take into account the effects of single sample classification and positive and negative sorting.

[0113] The entire model uses the AdamW optimizer for end-to-end training of all parameters, including the HAPM module and link predictor. The CosineAnnealingLR learning rate scheduler is used to gradually reduce the learning rate to achieve stable and efficient convergence. During training, the model continuously uses validation set feedback to adjust hyperparameters and ultimately outputs a risk score for TCM-adverse reactions, supporting subsequent Top-K screening of candidate relationships and clinical validation.

[0114] On the other hand, the present invention also provides a Chinese medicine adverse reaction risk prediction system based on heterogeneous graph neural network, which adopts the method described in the present invention and may include the following modules:

[0115] A data acquisition module is used to obtain the relationship files of traditional Chinese medicines and targets from the ETCM database and the relationship files of adverse reactions and targets from the ADReCS database;

[0116] A data preprocessing module is used to clean the acquired relational files to obtain preprocessed files;

[0117] A heterogeneous graph construction module is used to construct a heterogeneous graph using preprocessing files, taking Chinese herb, target, and adverse reaction as three types of nodes;

[0118] Node feature initialization module, which is used to construct the initial feature vector for each node herb, target, and adverse. The initial embedding of all nodes is mapped to the same dimensional space;

[0119] The feature fusion module is used to input node features into the HAPM heterogeneous graph attention network for message transmission and aggregation, and output node-level representation vectors;

[0120] The prediction module is used to calculate the prediction scores of all (herb, adverse) relationship pairs using the link prediction module, sort them in descending order, and select the top-K items with the highest scores as candidate adverse reactions;

[0121] A result output module, which is used to output predicted candidate adverse reaction results;

[0122] The verification and feedback iteration module is used to verify the candidate adverse reaction results and add the verified true results to the training data for the next round of iteration.

[0123] like Figure 2 As shown, the embodiment of the present invention provides a method for predicting the risk of adverse reactions of traditional Chinese medicine based on heterogeneous graph neural network. In order to more clearly illustrate the technical solution of the exemplary embodiment of the present invention, Figure 2 Five main modules are shown for illustration, including heterogeneous graph construction, node feature initialization, mapping layer, HAMP module, and link prediction module.

[0124] (1) Constructing a heterogeneous graph: In this embodiment, the herb-target and adverse reaction-target association data are extracted from the ETCM and ADReCS databases, respectively. After data cleaning and standardization, herb, target, and adverse reaction are used as three types of nodes. A heterogeneous graph is constructed using known direct relationships (solid lines). At the same time, a composite (indirect) relationship between herb and adverse reaction (dashed lines) is automatically generated to form a complete heterogeneous graph structure, which serves as the basis for subsequent model input.

[0125] (2) Node Feature Initialization: This embodiment uses the EmbedInit function to generate initial feature vectors for various nodes in the heterogeneous graph. Specifically, for herb nodes, EmbedInit extracts TCM-target association features from the ETCM database and can be combined with verified adverse reaction data; for adverse nodes, EmbedInit extracts adverse reaction-target association features from the ADReCS database and can be combined with verified adverse reaction data; for target nodes, EmbedInit integrates biological information such as protein function and pathway annotation.

[0126] (3) Mapping layer: After the node features are initialized, the initial vectors of each type of node pass through a separate linear mapping (EmbedInit) module, and after ReLU activation and Dropout processing, the initial features of all nodes in the heterogeneous graph are uniformly mapped to the same dimension. within, that is This lays the foundation for the input of the subsequent HAPM module. Here, h represents the feature vector of the node, the subscript v indicates that this is the vector for node v in the graph, and the superscript (0) represents the "initial" layer, that is, the embedding of the 0th layer. represents the set of real numbers, Represents a real vector space of length d, where d represents the dimension of the unified mapping.

[0127] (4) HAPM module: Next, the node embeddings of uniform dimension are input into the HAPM module, which consists of multiple HAPMConv layers, each of which uses a multi-head attention mechanism to update the node representations in the heterogeneous graph in parallel.

[0128] Each layer of HAPMConv output Both are used for:

[0129] As the input of the next layer of HAPMConv, it completes the layer-by-layer update, and in the final output stage, the output of all layers is spliced ​​into a comprehensive feature vector through the Concat operation. The figure shows that after each layer output, the arrows converge to a "Concat" module, and then pass through a FinalLinear layer mapping to finally obtain a unified node representation.

[0130] (5) Link prediction module: Finally, for the Chinese medicine node and adverse reaction node, their final embeddings are extracted respectively. and The two are concatenated and fed into a two-layer fully connected layer (MLP) network. After processing with ReLU, BatchNorm, and Dropout, a Sigmoid activation function is used to output a risk score (RiskScore), which is used to predict the potential association between traditional Chinese medicine and adverse reactions. This section is simply represented in the figure as "MLP Layer 1 → MLP Layer 2 → Risk Score."

[0131] Optionally, in order to further improve the generalization ability of the model and the training convergence speed, in an embodiment of the present invention, after linear mapping of node features and activation through ReLU, regularization processing such as batch normalization (BatchNorm) and random dropout (Dropout) can be introduced. Figure 2 In the multi-layer perceptron (MLP) module shown, the feature vectors of the concatenated Chinese medicine node and the adverse reaction node are and Follow these steps in order:

[0132] The feature vectors of the traditional Chinese medicine node and the adverse reaction node are spliced ​​in the feature dimension to obtain:

[0133]

[0134] Perform a linear mapping on the concatenated vector x to obtain an intermediate representation:

[0135] z=W1x+b1

[0136] Where W1 is the weight matrix and b1 is the bias vector. Then batch normalization is performed on z, and the process is as follows:

[0137] Calculate the mean μ of the feature vector z output by the first layer of full connection according to the mini-batch B and variance

[0138]

[0139] Where m is the number of mini-batch samples

[0140] Then z is normalized:

[0141]

[0142] Where σ is a very small constant that prevents division by zero;

[0143] And the normalized output is obtained by linear transformation through the trainable scale parameter γ and offset parameter β:

[0144]

[0145] The ReLU activation function is used for the batch normalized results:

[0146] y′=ReLU(y).

[0147] The activated output is then randomly deactivated through the Dropout operation, which is mathematically expressed as:

[0148]

[0149] where r i ~Bernoulli(p), where ⊙ represents element-wise multiplication and p is the probability of neuron retention. This operation helps reduce the risk of model overfitting.

[0150] The BatchNorm and Dropout processing steps mentioned above are optional regularization measures, which are used to enhance the robustness of the model in the implementation and are not directly reflected in Figure 2 In the simplified “MLP Layer 1→MLP Layer 2→Risk Score” structure.

[0151] It should be noted that Figure 2 Only the core part of the main structural process of the present invention is shown, aiming to clearly reflect the main information flow from heterogeneous graph construction to risk prediction results, but the diagram does not cover the complete training and feedback mechanism.

[0152] For example, the diagram does not explicitly depict the verification and feedback iteration mechanism in step S7, which involves comparing the model's output of TCM-adverse reaction high-risk data against common adverse reaction treatment data and incorporating verified associations as new samples into the training set to support the next round of iterative optimization. This mechanism has been detailed previously and is omitted in the diagram to avoid clutter.

[0153] In addition, the design and optimization details of the loss function are not Figure 2 In actual implementation, the present invention introduces a hybrid interest boundary loss function to supervise the prediction results, combines the sorting boundary of the positive and negative sample differences for optimization training, and adopts the AdamW optimizer and the CosineAnnealingLR learning rate scheduler to improve training convergence and model generalization ability.

[0154] In summary, Figure 2 This is a structural core diagram of the method of the present invention, which assists in explaining the overall system structure and main computing modules. Although some detailed processes are not drawn, they are fully described in other parts of the specification. Those skilled in the art can understand the overall solution in conjunction with the text.

[0155] In order to verify the effectiveness of the present invention, this embodiment adopts a five-fold cross-validation method to systematically test the proposed method for predicting the risk of adverse reactions of traditional Chinese medicine based on heterogeneous graph neural networks.

[0156] In the experiment, the key parameters are set as follows:

[0157] Number of attention heads: Each layer of HAPMConv uses 8 attention heads (H=8), allowing each node to capture neighbor features from multiple angles during message transmission, improving the expressive power of the attention mechanism.

[0158] Embedding Dimensionality: The initial embeddings of all nodes (herb, target, adverse) are mapped to the same 64-dimensional vector space (d=64), i.e.:

[0159]

[0160] HAPMConv Layers and Structure: The entire model uses three HAPMConv layers (L = 3). The outputs of each layer are first fused using residual connections and nonlinear activations (ReLU). Finally, the outputs of all layers are concatenated using a Concat operation and mapped to a uniform dimension through a FinalLinear layer for subsequent link prediction.

[0161] Interest boundary parameter ω m Setting: Interest boundary parameter ω introduced in the loss function mIt is initialized to a small constant (such as 0.05) and automatically updated with iterations as a learnable parameter during training to adaptively adjust the boundaries of positive and negative sample scores.

[0162] Batch Normalization and Dropout: Batch normalization (BatchNorm) and random dropout (Dropout) are optionally applied to the linear mapping output of each HAPMConv layer and the final MLP prediction module for regularization. The momentum of Batch Normalization is set to 0.9; the dropout probability (p) of Dropout is set to 0.2 to prevent overfitting and enhance the generalization performance of the model.

[0163] Optimizer and learning rate scheduling: The optimizer uses AdamW, and the initial learning rate is set to 0.001; the regularization coefficient weightdecay defaults to 1e-4; the learning rate scheduler uses CosineAnnealingLR, and the period (T_max) is set to 50, that is, a cosine annealing cycle is completed every 50 epochs, and the learning rate periodically decays to the minimum value and then increases until the end of 300 epochs of training.

[0164] Number of training rounds and mini-batch size: The total number of training epochs is set to 300.

[0165] Under the above configuration, the HAPM module proposed in this invention was trained end-to-end, and the precision, recall, F1 score, ROC-AUC, and PR-AUC indicators were statistically analyzed in a 5-fold cross-validation to evaluate the overall performance and stability of the model in the task of predicting the risk of adverse reactions to traditional Chinese medicine.

[0166] The experimental results are shown in Table 1 below.

[0167] Table 1. Experimental results of various properties of the present invention

[0168] Accuracy Recall F1 score ROC-AUC PR-AUC 0.852±0.030 0.805±0.035 0.828±0.025 0.925±0.010 0.907±0.010

[0169] As can be seen from Table 1, the proposed method demonstrates significant advantages over traditional methods in predicting adverse reactions to traditional Chinese medicine (TCM) risk. Based on the synergistic effect of heterogeneous graph neural networks and attention mechanisms, the model is able to fully explore the complex and diverse relationships between drugs and adverse reactions, effectively capturing the complementary information between different types of data, and achieving efficient differentiation of positive and negative samples and accurate detection of potential risks. Overall, while maintaining a high level of risk capture, the model also demonstrates excellent predictive stability, providing strong support for achieving more comprehensive and reliable TCM adverse reaction risk prediction.

[0170] In order to clearly and effectively demonstrate the results of the present invention in terms of adverse reaction prediction, Xanthium sibiricum will be used as a typical case below to present the relevant prediction data in detail, and compared and verified with "Common Adverse Drug Reactions and Treatment - Traditional Chinese Medicine Volume" (edited by Lin Jun, Liu Naxin, and Ou Xiaolong, published by Military Medical Science Press).

[0171] The Common Adverse Drug Reactions and Treatments - Chinese Medicine Volume records the following three clinical manifestations of adverse reactions to Xanthium sibiricum:

[0172] The incubation period is from 4 hours to 3 days after ingestion. It varies depending on the individual who eats it. Those who eat raw Xanthium sibiricum will develop symptoms 4 to 8 hours after ingestion, those who eat Xanthium sibiricum cakes will develop symptoms 10 to 24 hours after ingestion, and those who eat seedlings will develop symptoms 1 to 5 days after ingestion. The severity of the poisoning reaction varies.

[0173] 2. In mild cases, symptoms include headache, dizziness, fatigue, loss of appetite, nausea, vomiting, abdominal pain, diarrhea or fever, facial flushing, conjunctival congestion, urticaria, etc.

[0174] 3. Severe cases may cause irritability, lethargy, gastrointestinal bleeding, elevated SGPT, and the presence of white blood cells in urine routine tests. Most patients can recover with timely treatment, but a few may die from repeated seizures, massive bleeding, extensive hepatocyte degeneration and necrosis, cerebral edema, renal failure, and pulmonary edema.

[0175] Table 2. Comparison of prediction results of adverse reactions of Xanthium sibiricum

[0176]

[0177]

[0178] These results demonstrate that the prediction model demonstrates high accuracy in identifying adverse reactions to Xanthium sibiricum, successfully capturing most key symptoms with a matching rate of 93.33%. This demonstrates the practical value of this method in risk assessment and safe medication management, and can provide support for early warning and clinical treatment of adverse reactions to traditional Chinese medicines.

[0179] The beneficial effects of the prediction method of the present invention are as follows:

[0180] (1) A prediction model based on heterogeneous graph neural network (HAPM) is proposed, which automatically integrates multi-source data such as traditional Chinese medicine, targets and adverse reactions through an end-to-end deep learning framework to fully explore their complex semantic associations.

[0181] (2) Design a multi-head attention mechanism and hierarchical feature fusion strategy to extract and express features based on the heterogeneity of different types of data, thereby improving the model's predictive performance on the risk of adverse reactions to traditional Chinese medicine.

[0182] (3) Construct a prediction framework with strong interpretability, reveal the prediction basis through information such as internal attention weights, and provide more reliable support for clinical decision-making.

[0183] (4) The problem of comprehensive fusion of multi-source heterogeneous data is solved. In order to make up for the shortcomings of existing methods that only rely on the molecular structure characteristics of drugs, ignore the overall information of traditional Chinese medicine, and have large noise in the comment text and unstable data, the present invention proposes a prediction method based on heterogeneous graph neural network. This method constructs a heterogeneous graph consisting of three types of nodes: traditional Chinese medicine, target, and adverse reaction. It automatically integrates ETCM traditional Chinese medicine-target data, ADReCS adverse reaction-target data, and traditional Chinese medicine adverse reaction information confirmed in the verification literature to achieve deep joint modeling of multi-source data, thereby obtaining more accurate and complete input features.

[0184] (5) It solves the problem of insufficient multi-head attention mechanism and hierarchical feature fusion capabilities. In terms of feature extraction, existing methods are mainly based on single-layer or multi-layer GCN for neighborhood aggregation, which lacks a detailed description of the deep, nonlinear interaction relationship between heterogeneous nodes; or focuses on text generation and fails to fully utilize graph structure information. To this end, the present invention designs a customized HAPM module, the core of which adopts a multi-head attention mechanism and a hierarchical feature fusion strategy. By independently processing the updates of each heterogeneous node at each layer, and finally splicing the outputs of all layers using the Concat operation, and then obtaining a unified node representation through FinalLinear mapping, this solution not only retains shallow features but also integrates deep semantics, thereby improving the overall feature expression capability and risk prediction accuracy.

[0185] (6) Solved the problem of insufficient dynamic feedback and iterative training mechanism. Existing methods usually rely on one-time generation and fixed prediction strategies, lacking a mechanism for continuous adaptive updating, making it difficult for the model to timely capture new adverse reactions that occur during actual medication. To this end, the present invention introduces a dynamic feedback iteration mechanism in the end-to-end model framework: in the prediction stage, the candidate Chinese medicine-adverse reaction relationship that has been clinically verified or confirmed by the literature is fed back to the training set, the positive sample data is continuously expanded, and the model parameters are further optimized through multiple rounds of iterations, thereby enhancing the generalization ability and real-time prediction performance of the model.

[0186] The above descriptions are only some preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for predicting the risk of adverse reactions to traditional Chinese medicine based on heterogeneous graph neural networks, comprising the following steps: S1. Obtain data: obtain the relationship files of traditional Chinese medicines and targets from the ETCM database and the relationship files of adverse reactions and targets from the ADReCS database; S2. Data preprocessing: performing data cleaning on the acquired relational files to obtain preprocessed files; S3. Construct a heterogeneous graph, taking herb, target and adverse reaction as three types of nodes, and use the preprocessing file to construct the heterogeneous graph; S4, node feature initialization, constructing the initial feature vector for each node herb, target, and adverse, and all node initial embeddings are mapped to the same dimensional space; S5, HAPM feature fusion, inputs node features into the HAPM heterogeneous graph attention network for message passing and aggregation, and outputs node-level representation vectors; S6. Prediction and output: Use the link prediction module to calculate the prediction scores for all (herb, adverse) relationship pairs, sort them in descending order, and select the top-K items with the highest scores as candidate adverse reactions; S7. Verification and feedback iteration: The candidate TCM-adverse reaction pairs are compared and verified with the "Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume". If certain (herb, adverse) relationships are confirmed, they will be added to the training data to further improve the model performance in the next round of iteration.

2. The method according to claim 1, wherein: In step S2, first, the TCM-target relationship file provided in the ETCM database and the adverse reaction-target relationship file provided in the ADReCS database are preliminarily integrated, duplicate records are deleted, and data with missing key fields are eliminated to ensure the integrity and consistency of the original data; secondly, the field names in the two data sources are unified, and the data format of each field is standardized; Secondly, the integrated data is verified to check the logical consistency of each record, and abnormal data is eliminated or corrected to ensure that each relational data meets the requirements for building a heterogeneous graph; finally, according to the preprocessing requirements, the cleaned and standardized data is converted into a preprocessing file in a unified format.

3. The method according to claim 1, wherein: In step S3, the heterogeneous graph is represented as in V={herb,target,adverse}, E={(herb,ht,target),(target,ta,adverse)} By combining two paths, a composite relationship (herb, ind, adverse) is automatically generated to represent the path from Chinese medicine to target to adverse reaction; where herb represents an important node, target represents a target node, and adverse represents an adverse reaction node. represents the constructed heterogeneous graph, V represents the node set, E represents the edge set, ht represents the known association between traditional Chinese medicine and target (herb→target), ta represents the known association between adverse reaction and target (target→adverse), and ind represents the compound association between traditional Chinese medicine and adverse reaction (herb→target→adverse).

4. The method according to claim 1, wherein: In step S4, an initial feature vector is constructed for each node v∈V={herb, target, adverse}, and the following formula is defined: Among them, EmbedInit means mapping the entity information corresponding to the node v to a real vector of fixed dimension d; for the herb node, EmbedInit extracts its Chinese medicine-target association features from ETCM; for the adverse node, EmbedInit extracts its adverse reaction-target association features from ADReCS; for the target node, EmbedInit integrates its protein function and pathway annotation biological information; all nodes are initially embedded in the same dimensional space. Here, h represents the feature vector of the node, the subscript v indicates that this is the vector for node v in the graph, and the superscript (0) represents the "initial" layer, that is, the embedding of the 0th layer. represents the set of real numbers, Represents a real vector space of length d, where d represents the dimension of the unified mapping.

5. The method according to claim 1, wherein: In step S4, the "verified Chinese medicine-adverse reaction relationship" or other sorted Chinese medicine adverse reaction information is incorporated into the EmbedInit(·) function; for the herb node, if certain adverse reactions have been confirmed in the previous steps or in external literature, the corresponding dimension is added to its initial feature vector to record the number, type or severity of the verified adverse reactions of the Chinese medicine; for the adverse node, the historical frequency of the adverse reaction or the confirmed clinical hazard level is introduced.

6. The method according to claim 1, wherein: In step S5, in the constructed heterogeneous graph, a HAPM attention module is defined to perform feature aggregation on nodes. For the lth layer, the representation of each target node t is determined by the messages transmitted by all its connected neighbor nodes s through the predefined meta-path p. The attention weight of the hth attention head is calculated as follows: in are the query and key mapping matrices of the h-th head, d h For single head dimension, and Represent the neighbor and target node embeddings of the previous layer respectively, and the message vector is obtained by weighted summation: in is the value mapping matrix, "||" represents the first-level splicing operation, represents the set of neighbors reachable along the meta-path p = (herb, ind, adverse), represents the attention weight from neighbor node s to target node t calculated by the h-th attention head, represents the embedding vector (node ​​representation) of the neighbor node s in the (l–1)th layer, that is, the feature representation of node s after aggregation and activation in the previous layer, which is used for weighted message passing in the current layer. Finally, the target node representation is updated by fusion with ReLU activation through residual connection: in is the output mapping matrix, Indicates the final output embedding vector of the target node t in the lth layer after attention aggregation, residual connection and ReLU activation in this layer, represents the input embedding vector of the target node t in the (l-1)th layer, It represents the message vector obtained by weighted summation of all neighboring nodes that meet the meta-path p of the target node t in the lth layer through each attention head, that is, the weighted results of each head are spliced ​​together and used for splicing with the embedding of the previous layer and then performing mapping update.

7. The method according to claim 1, wherein: In step S6, after the Lth layer output, the correlation strength between the Chinese medicine node herb and the adverse reaction node adverse is focused on, and the Chinese medicine output of the final layer is embedded into Embedded with adverse reactions After concatenation, input the two-layer fully connected network and calculate the association score: in, is a trainable weight matrix, σ is a Sigmoid activation function, which maps the output to the interval [0,1], indicating the risk association between Chinese medicine and adverse reactions. When representing the output of the Lth layer, the final embedding vector of the Chinese medicine node h after all attention heads are aggregated, residual connections and ReLU activation is used to characterize the semantic representation of the Chinese medicine node in the heterogeneous graph. When representing the output of the Lth layer, the final embedding vector of the adverse reaction node a after the same attention aggregation, residual connection and ReLU activation is used to characterize the semantic representation of the adverse reaction node in the heterogeneous graph.

8. The method according to claim 1, wherein: In step S7, the predicted probability calculated for each pair of Chinese medicine-adverse reaction (h, a) Compared with the preset threshold τ, when , match its candidate relationships with known records in Common Adverse Drug Reactions and Treatments - Traditional Chinese Medicine Volume to construct a verification set: Among them, ClinicalRef represents the list of manually verified Chinese medicine-adverse reaction pairs. The verified relationships are then fed back to the training graph, so that the training edge set is updated to: Thereby expanding the scale of training samples and improving the model representation; repeating steps S3 to S6 until the verification indicator converges and the iteration is stopped.

9. The method according to claim 1, wherein: It also includes loss function and training optimization, introducing the interest boundary parameter ω in the loss function m ,in Associated with the i-th TCM node, the loss function is defined as: Among them, y i is the sample label, 1 represents a positive sample and 0 represents a negative sample.

10. A system for predicting adverse reactions to traditional Chinese medicine based on heterogeneous graph neural networks, which uses the method according to any one of claims 1 to 9, comprising: A data acquisition module is used to obtain the relationship files of traditional Chinese medicines and targets from the ETCM database and the relationship files of adverse reactions and targets from the ADReCS database; A data preprocessing module is used to clean the acquired relational files to obtain preprocessed files; A heterogeneous graph construction module is used to construct a heterogeneous graph using preprocessing files, taking Chinese herb, target, and adverse reaction as three types of nodes; Node feature initialization module, which is used to construct the initial feature vector for each node herb, target, and adverse. The initial embedding of all nodes is mapped to the same dimensional space; The feature fusion module is used to input node features into the HAPM heterogeneous graph attention network for message transmission and aggregation, and output node-level representation vectors; The prediction module is used to calculate the prediction scores of all (herb, adverse) relationship pairs using the link prediction module, sort them in descending order, and select the top-K items with the highest scores as candidate adverse reactions; A result output module, which is used to output predicted candidate adverse reaction results; The verification and feedback iteration module is used to verify the candidate adverse reaction results and add the verified true results to the training data for the next round of iteration.

Citation Information

Patent Citations

  • Adverse drug reaction prediction method based on comment text information enhancement

    CN119170295A

  • Adverse drug reaction monitoring method and system based on big data

    CN119418951A

  • Intelligent adverse drug reaction prediction method and device based on multi-source data fusion

    CN119517219A

  • Data collection and drug safety signal mining methods and agents for PMS

    CN119786078A

  • Drug symmetric element path-based adverse reaction prediction method and system

    CN120164638A

Cited By

  • Adverse drug reaction prediction method and system based on graph neural network

    CN122091277A

  • Polycarboxylate superplasticizer performance prediction method based on heterogeneous graph neural network

    CN122314148A