A drug adverse reaction prediction method and system based on a graph neural network
By employing a two-layer molecular graph encoding and collaborative multi-task training framework, combined with a temperature scaling method, the problems of insufficient molecular feature extraction and confidence quantification in drug adverse reaction prediction in existing technologies are solved, achieving drug adverse reaction prediction with higher accuracy and reliability.
Patent Information
- Application Number
- CN202610561364.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-04-27
AI Technical Summary
Existing methods for predicting adverse drug reactions based on graph neural networks cannot effectively capture toxicology-related molecular substructure information at the molecular feature expression level, and the prediction results lack reliable confidence quantification, making it difficult to meet the needs of clinical safety assessment.
We employ a two-layer molecular graph encoding mechanism and a collaborative multi-task training framework. By constructing a supernode graph for functional group-level message passing, we introduce multidimensional drug safety attributes as auxiliary tasks and combine them with a temperature scaling method for confidence calibration, thereby improving the quality of molecular representation and the reliability of prediction results.
It significantly improves the accuracy and confidence level of adverse drug reaction prediction, and can better support clinical medication decisions.
Smart Images

Figure CN122091277B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of healthcare information technology, and in particular to a method and system for predicting adverse drug reactions based on graph neural networks. Background Technology
[0002] Adverse drug reactions (ADRs) refer to harmful and unexpected reactions to drugs in the human body under normal medication conditions. They are a major cause of new drug development failures and drug withdrawals after market launch. Accurate prediction of ADRs is crucial for improving clinical drug safety and reducing R&D costs. With the development of deep learning technology, computational ADR prediction has gradually become an important research direction for early drug safety screening. Compared with traditional high-throughput experimental screening and clinical data mining methods, computational prediction methods have significant advantages such as low cost, short cycle time, and coverage of a large-scale candidate compound space, demonstrating great application potential in the early safety assessment of new drug development.
[0003] Existing computational ADR prediction techniques suffer from the following shortcomings: First, at the molecular feature expression level, mainstream graph neural network methods only construct a single-layer message passing mechanism at the atomic level, failing to effectively capture molecular substructure information with clear toxicological semantics, such as metabolically sensitive groups and toxicity warning structures directly related to ADR occurrence, resulting in insufficient granularity of toxicology-related molecular features. Second, at the knowledge utilization level, there is a systematic biological association between the drug's ADMET properties (including CYP450 metabolic enzyme inhibition, hepatotoxicity, cardiotoxicity, etc.) and ADR, but existing methods typically train ADR prediction as an isolated task, failing to effectively incorporate the aforementioned knowledge base information into the molecular representation learning process, causing decoupling between the knowledge base information and the molecular representation learning process. Third, at the reliability level, the raw prediction probabilities output by existing deep learning models generally suffer from overconfidence, with inconsistencies between the predicted probabilities and the actual ADR incidence rate, lacking statistically significant confidence quantification, making it difficult to meet the reliability requirements of clinical safety assessments for prediction results.
[0004] Chinese patent application CN116403730A discloses a drug interaction prediction method based on graph attention networks. This method converts the SMILES structures of candidate drugs into molecular graphs, uses a bond-aware message-passing neural network to perform atomic-level encoding of the molecular graph to obtain molecular graph embedding vectors, and then combines an external drug-drug interaction network graph, employing a multi-layer graph attention network to update the molecular embedding vectors. Finally, a fully connected predictor outputs the interaction relationship between the two drugs. While this method makes valuable progress in the construction and message-passing encoding of drug molecular graphs, its ability to capture toxicologically significant local chemical patterns in drug molecular structures is limited, leading to insufficient prediction accuracy for adverse drug reactions. Furthermore, the prediction results lack reliable quantitative confidence characterization, making it difficult to support the assessment of prediction reliability in clinical decision-making scenarios. Summary of the Invention
[0005] In view of this, the present invention provides a method and system for predicting adverse drug reactions based on graph neural networks, in order to solve the problems that existing graph neural network-based adverse drug reaction prediction methods have insufficient ability to extract drug molecular toxicology-related structural features, resulting in the inability to meet the clinical safety assessment requirements for the accuracy of adverse drug reaction prediction, as well as the problem that the output results of existing prediction models lack reliable confidence quantification and are difficult to provide effective support for clinical medication decisions.
[0006] The technical solution of this invention is implemented as follows: On the one hand, the present invention provides a method for predicting adverse drug reactions based on graph neural networks, comprising the following steps: S1. Receive and verify the molecular structure information of the candidate drug to obtain the verified molecular structure information of the candidate drug. S2. A two-layer molecular graph encoding mechanism is used to encode the molecular structure information of the verified candidate drugs to obtain the molecular representation vector of the candidate drugs. The molecular graph encoder used in the two-layer molecular graph encoding mechanism is trained offline through a collaborative multi-task training framework with adverse drug reaction prediction as the main task and multi-dimensional safety attribute prediction as the auxiliary task. S3. Input the molecular representation vector of the candidate drug into the multi-label prediction network for adverse drug reactions to predict the probability of occurrence of the candidate drug in each adverse drug reaction category, and obtain the original logits vector of each adverse drug reaction category. S4. The original logits vectors of each adverse drug reaction category are calibrated for confidence using a temperature scaling method to obtain the calibration probability and corresponding confidence level of each adverse drug reaction category, and an adverse drug reaction prediction report is output.
[0007] Based on the above technical solutions, preferably, step S2 employs a two-layer molecular graph encoding mechanism to perform feature encoding on the verified candidate drug molecular structure information, specifically including: S21. Convert the verified candidate drug molecular structure information into a molecular graph, and use a message passing neural network to perform atomic-level message passing on the molecular graph to obtain the atomic-level implicit representation of each atom in the molecular graph. S22. Based on the preset toxicological semantic rules, functional groups are identified in the molecular graph. The sets of atoms that meet the identification conditions are aggregated into functional group supernodes. A supernode graph is constructed with each functional group supernode as a node. A message passing neural network is used to perform supernode-level message passing on the supernode graph to obtain the supernode-level hidden representation of each functional group supernode. S23. Atomic-level implicit representations are obtained by pooling operations, and supernode-level implicit representations are obtained by aggregation operations. A gated fusion mechanism is used to fuse the atomic-level molecular representations and the supernode-level molecular representations to obtain the molecular representation vector of the candidate drug.
[0008] Based on the above technical solution, preferably, in step S22, the functional group supernodes include P-type supernodes, M-type supernodes and A-type supernodes, wherein P-type supernodes are used to characterize the pharmacophore structure, M-type supernodes are used to characterize the structure of the metabolically sensitive group, and A-type supernodes are used to characterize the warning structure related to adverse drug reactions. When the same set of atoms simultaneously meets the preset recognition conditions for multiple types of functional groups supernodes, the set of atoms is uniquely classified using the priority classification rule of A-type supernode > M-type supernode > P-type supernode, and a supernode graph is constructed.
[0009] Based on the above technical solutions, preferably, the gating fusion mechanism described in step S23 satisfies: ; ; in, A vector representing the atomic-level molecules. This is the supernode-level molecular representation vector. This represents the fused molecular representation vector. For element-wise gated vectors, For the Sigmoid function, The hyperbolic tangent activation function is used. For element-wise Hadamard product, This is a vector concatenation operation. , and For learnable parameters, It is a learnable bias vector.
[0010] Based on the above technical solutions, preferably, the training process of the collaborative multi-task training framework in step S2 specifically includes: Step 1: Extract the labels of training sample drugs on the multidimensional safety attributes of each drug from the drug knowledge base, and use them as training supervision signals for auxiliary tasks; Step 2: Using a molecular graph encoder as a shared backbone network, feature encoding is performed on the molecular structure information of the training samples to obtain the molecular representation vector of the training samples; a main task prediction head is set for the main task of drug adverse reaction prediction, and independent auxiliary task prediction heads are set for each auxiliary task of drug multidimensional safety attribute prediction. Each prediction head takes the molecular representation vector of the training samples as input. Step 3: Based on the molecular representation vector of the training samples, calculate the dynamic task association weight of molecular perception independently for each auxiliary task, and weight the loss of each auxiliary task according to the dynamic task association weight. Step 4: Perform end-to-end optimization training on the shared backbone network using the joint loss of the main task loss and the weighted auxiliary task loss to obtain the optimized molecular graph encoder.
[0011] Based on the above technical solutions, preferably, the dynamic task association weights in step 3 satisfy the following: ; The joint loss in step 4 satisfies: ; in, This represents the fused molecular representation vector. For the first k Dynamic task association weights for each auxiliary task. For the Sigmoid function, For the first k Task-aware query vectors for each auxiliary task. For the first k The molecular feature projection matrices of each auxiliary task and, For the first k Bias vectors for each auxiliary task K To assist in the total number of tasks, For the first k The label availability mask coefficient for each auxiliary task. For the current batch The number of auxiliary tasks with a value of 1 varies with the batch size. , To predict the main task of adverse drug reaction, For the first k The loss of an auxiliary task This is used as an auxiliary task's overall weighting coefficient.
[0012] Based on the above technical solutions, preferably, in step S3, the multi-label prediction network for adverse drug reactions is a multilayer perceptron with two hidden layers. The multilayer perceptron includes a first hidden layer, a second hidden layer and an output layer in sequence. The activation function of the first hidden layer and the second hidden layer is ReLU. Each hidden layer is followed by a batch normalization layer and a random inactivation layer. The multilayer perceptron receives the molecular representation vector of the candidate drug, performs feature transformation layer by layer through the first hidden layer and the second hidden layer, and outputs the original logits vector corresponding to each adverse drug reaction category by the output layer. The nth component of the original logits vector corresponds to the original prediction score of the nth adverse drug reaction category.
[0013] Based on the above technical solutions, preferably, step S4 specifically includes: S41. After the molecular graph encoder and the multi-label prediction network for adverse drug reactions are trained, an independent calibration set that does not overlap with the training set is used as the basis. Based on the true labels and original logits vectors of each sample in each adverse drug reaction category in the calibration set, the temperature parameter is used as the only variable to be optimized. A joint negative log-likelihood loss function for the temperature parameter is constructed. The L-BFGS optimizer is used to minimize the joint negative log-likelihood loss function to obtain the calibrated temperature parameter. S42. For any candidate drug, use the calibrated temperature parameters to perform temperature scaling on each component in the original logits vector to obtain the calibrated probability of each adverse reaction category of the drug. S43. Based on the calibration probability of each adverse drug reaction category, perform confidence level classification to obtain the corresponding confidence level and output an adverse drug reaction prediction report.
[0014] Based on the above technical solutions, preferably, the calibration probability of each adverse drug reaction category in step S42 satisfies: ; in, The original logits vector The nth component, For the calibrated temperature parameters, For the Sigmoid function, This represents the calibration probability for the nth adverse drug reaction category.
[0015] In addition, the present invention also provides a drug adverse reaction prediction system based on graph neural networks to implement the above-mentioned method, comprising: The molecular structure verification module is used to receive the molecular structure information of candidate drugs, verify it, and output the verified molecular structure information of candidate drugs. A two-layer molecular graph encoding module is used to perform feature encoding on the verified candidate drug molecular structure information using a two-layer molecular graph encoding mechanism, and output the molecular representation vector of the candidate drug. The multi-label prediction module for adverse drug reactions is used to receive the molecular representation vector of candidate drugs and input it into the multi-label prediction network for adverse drug reactions. It predicts the probability of occurrence of candidate drugs in each adverse drug reaction category and outputs the original logits vector of each adverse drug reaction category. The confidence calibration module is used to perform confidence calibration on the original logits vector of each adverse drug reaction category using a temperature scaling method, so as to obtain the calibration probability and corresponding confidence level of each adverse drug reaction category. The report output module is used to output adverse drug reaction prediction reports based on the calibration probability and corresponding confidence level of each adverse drug reaction category.
[0016] The present invention has the following advantages over the prior art: (1) This invention constructs a complete adverse drug reaction prediction scheme from molecular structure feature extraction to the quantification of prediction result reliability through two-layer molecular graph encoding, collaborative multi-task training, and temperature-shrunk reliability calibration. Compared with existing prediction methods based on graph neural networks, this invention works synergistically on the final prediction result in three aspects: molecular representation quality, richness of encoder training supervision signals, and statistical reliability of prediction probability. Among them, two-layer molecular graph encoding and collaborative multi-task training improve the model's ability to identify and model toxicology-related molecular structure features from the two dimensions of feature expression and knowledge utilization, which significantly improves the prediction accuracy of adverse drug reactions. Temperature-shrunk reliability calibration statistically aligns the model output probability with the actual incidence of adverse drug reactions, enabling the prediction result to have reliable confidence quantification capability, effectively solving the problem that the existing model prediction results lack reliable confidence quantification and are difficult to support clinical medication decisions.
[0017] (2) This invention employs a two-layer molecular graph encoding mechanism. Based on atomic-level message passing, an additional supernode graph is constructed, consisting of P-type pharmacophore supernodes, M-type metabolically sensitive group supernodes, and A-type warning structure supernodes. Independent message passing is performed at the supernode level, and the two-level representations are ultimately integrated through a gating fusion mechanism. This technical solution enables the model to explicitly model molecular substructure information with toxicological semantics at the functional group granularity, compensating for the shortcomings of single-layer atomic-level encoding in expressing local chemical patterns such as metabolically sensitive groups and toxicity warning structures. This helps improve the model's ability to identify toxicity-related molecular structural features.
[0018] (3) This invention introduces a collaborative multi-task training framework with ADMET attribute prediction as an auxiliary task during the training phase of the molecular graph encoder, and designs a molecular perception dynamic task association weight mechanism so that the gradient contribution of each auxiliary task to the shared encoder is adaptively adjusted according to the current molecular structure characteristics. This training strategy actively introduces the ADR-ADMET association information contained in the drug safety knowledge base into the molecular representation learning process, so that the encoder parameters are additionally constrained by systematic drug safety prior knowledge in addition to the supervision of the main task, which helps to improve the encoding quality of molecular representation for ADR-related biological features.
[0019] (4) After the prediction network is trained, this invention optimizes the temperature parameters by minimizing the negative log-likelihood loss using an independent calibration set, and applies this optimization to the logits scaling during the inference phase. The temperature scaling method statistically aligns the original output probability of the model with the actual ADR incidence rate with extremely low computational overhead and parameter increment, effectively alleviating the overconfidence problem commonly found in deep neural network classifiers. This enables the output probability to have statistically significant confidence quantification capabilities, providing a more valuable probability estimate for clinical drug safety assessment. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of the adverse drug reaction prediction method based on graph neural networks of the present invention; Figure 2 This is a flowchart of the two-layer molecular graph encoding mechanism architecture of the present invention; Figure 3 This is a flowchart of the collaborative multi-task training framework of the present invention; Figure 4 This is a flowchart of the temperature compression placement reliability calibration process of the present invention; Figure 5 This is a diagram of the adverse drug reaction prediction system based on graph neural networks of the present invention. Detailed Implementation
[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0023] like Figure 1 As shown, this invention provides a method for predicting adverse drug reactions based on graph neural networks, comprising the following steps: S1. Receive and verify the molecular structure information of the candidate drug to obtain the verified molecular structure information of the candidate drug. S2. A two-layer molecular graph encoding mechanism is used to encode the molecular structure information of the verified candidate drugs to obtain the molecular representation vector of the candidate drugs. The molecular graph encoder used in the two-layer molecular graph encoding mechanism is trained offline through a collaborative multi-task training framework with adverse drug reaction prediction as the main task and multi-dimensional safety attribute prediction as the auxiliary task. S3. Input the molecular representation vector of the candidate drug into the multi-label prediction network for adverse drug reactions to predict the probability of occurrence of the candidate drug in each adverse drug reaction category, and obtain the original logits vector of each adverse drug reaction category. S4. The original logits vectors of each adverse drug reaction category are calibrated for confidence using a temperature scaling method to obtain the calibration probability and corresponding confidence level of each adverse drug reaction category, and an adverse drug reaction prediction report is output.
[0024] In one embodiment of the present invention, step S1 includes: The system receives molecular structure information of candidate drugs as input, in the standard SMILES (Simplified Molecular Input Line Entry System) string format. The system performs validity checks on the input SMILES string, specifically: calling the RDKit molecular parsing interface to check if the string conforms to the SMILES syntax specification; validating the parsed molecular objects, rejecting structures with molecular weights exceeding a reasonable range (less than 50 Da or greater than 2000 Da, this range can be adjusted according to actual conditions), containing unconventional elements, or unable to generate 3D coordinates; structures that pass the validation proceed to subsequent processing, while those that fail return error messages. The system supports input via a RESTful API interface or a web interface. The API interface accepts a JSON request body containing a `smiles` (string) field and an optional `drug_name` (string) field, returning prediction results in structured JSON format.
[0025] In one embodiment of the present invention, such as Figure 2 As shown, step S2, which employs a two-layer molecular graph encoding mechanism to feature-encode the verified candidate drug molecular structure information, specifically includes: S21. Convert the verified candidate drug molecular structure information into a molecular graph, and use a message-passing neural network to perform atomic-level message passing on the molecular graph to obtain the atomic-level implicit representation of each atom in the molecular graph.
[0026] Specifically, the RDKit tool is invoked to perform molecular parsing on the input SMILES string, converting it into a molecular diagram. The set of nodes V corresponds to the atoms in the molecule, the set of edges E corresponds to the chemical bonds, and each undirected edge is represented as a bidirectional directed edge pair to support subsequent message passing.
[0027] In terms of atomic feature construction, a 39-dimensional atomic feature vector was designed. It is composed of the following dimensions: atom type (10-dimensional one-hot encoding, covering C / N / O / S / F / Cl / Br / I / P / others), hybridization state (5-dimensional one-hot encoding, covering...). ), connectivity (6-dimensional One-hot encoding, coverage from 0 to 5), formal charge (5-dimensional One-hot encoding, coverage) The features include: number of hydrogen atoms (5-dimensional one-hot encoding, covering 0 to 4), aromaticity (1-dimensional binary encoding), whether it is in a ring (1-dimensional binary encoding), atomic mass / van der Waals radius / electronegativity (3-dimensional continuous values, normalized to the [0,1] interval according to their respective theoretical maximum values), Gasteiger bias charge (1-dimensional continuous value, ranging from [-1,1], calculated by the RDKit's ComputeGasteigerCharges interface, reflecting the polarization degree of the atom in the molecular electrostatic field), and chirality label (3-dimensional one-hot encoding, covering R / S / None). For bond feature construction, a 12-dimensional bond feature vector is designed. It is composed of the following dimensions: bond type (4-dimensional one-hot encoding, covering single / double / triple / aromatic bonds), whether it is in a ring (1-dimensional binary encoding), whether it is conjugated (1-dimensional binary encoding), stereochemistry (3-dimensional one-hot encoding, covering E / Z / None), bond length (1-dimensional continuous value, in Å, normalized to the typical C / C bond length of 1.54 Å), rotational property (1-dimensional binary encoding, indicating whether the bond is rotatable) and hydrogen bond connection property (1-dimensional binary encoding, indicating whether the bond is connected to a hydrogen bond donor or acceptor atom).
[0028] In molecular diagram A four-layer message-passing neural network is constructed, with each layer having a hidden dimension of 256. An eight-head multi-head attention mechanism is introduced to differentially aggregate neighbor node information. Residual connections are set in each layer to alleviate gradient vanishing. The dropout rate is set to 0.2, and the activation function is LeakyReLU (negative slope). ). No. l The message generation and node feature update process of the layer is as follows: For directed edges Its message vector It is calculated by combining the implicit representation of the source node and the edge features: ; in, For nodes u In the l The implicit representation of layers, Let be the key feature vector of edge (u,v). This represents a vector concatenation operation. For the first l The layer is a two-layer feedforward network with an input dimension of 268 and an output dimension of 256.
[0029] For the One attention point ( ),node For nodes The attention coefficient is: ; in, For the first l Layer i Nodes in each attention head u For nodes v Attention coefficient, representing the node u Under this attention head, the node v Updated contribution weights i For attention head index, l For message passing layer index, It has 4 floors in total; For the first The linear transformation matrix of each attention head is a learnable parameter; For nodes v In the l The implicit representation vector of the layer; For nodes u In the l The implicit representation vector of the layer, u for v The neighboring nodes; For point w In the l The implicit representation vector of the layer, w Traversing nodes v All neighboring nodes are used to normalize the attention coefficients in the denominator; For the first i The learnable attention vector of each attention head is used to map the concatenated 64-dimensional vector into a scalar attention score, which is a learnable parameter. This is a vector concatenation operation; For nodes in the atomic-level molecular diagram G=(V,E) v The set of neighboring nodes, that is, with v All atomic nodes are connected by chemical bonds. This is the activation function for a linear rectified circuit with leakage. (Node) The implicit representation update formula is: ; in, For nodes v After the first l The updated hidden representation vector after layer update; This is a random deactivation operation; This indicates that the outputs of eight attention heads are concatenated, each with an output dimension of 32, resulting in a 256-dimensional vector; residual connection term. Used to preserve original feature information; Dropout randomly sets nodes to zero with a probability of 0.2 during the training phase; the initial hidden representation of a node is obtained through... The calculation yielded, where and For learnable parameters, For atoms v The initial feature vector. After 4 layers of updates, the hidden representation of each atom is obtained. ,in .
[0030] S22. Based on the preset toxicological semantic rules, functional groups are identified in the molecular graph. The sets of atoms that meet the identification conditions are aggregated into functional group supernodes. A supernode graph is constructed with each functional group supernode as a node. A message-passing neural network is used to pass messages at the supernode level to obtain the supernode-level hidden representation of each functional group supernode.
[0031] In this invention, a functional group supernode refers to abstracting a set of atoms with specific toxicological semantics in a molecular graph into a single node, constructing a higher semantic level graph structure above the atomic layer graph, thereby enabling the message passing mechanism to capture toxicity-related structural information at the functional group granularity. Functional group supernodes include P-type supernodes, M-type supernodes, and A-type supernodes, where P-type supernodes are used to characterize pharmacophore structures, M-type supernodes are used to characterize metabolically sensitive group structures, and A-type supernodes are used to characterize warning structures related to adverse drug reactions.
[0032] The SMARTS matching engine in RDKit was used to identify functional groups in molecular graphs. P-type supernodes, following Harper's pharmacophore rules, identified 18 classes of universal pharmacophore structures, including hydrogen bond donors (NH, OH), hydrogen bond acceptors (N, O with lone pairs of electrons), positively ionized groups (aliphatic amines), negatively ionized groups (carboxylic acids, sulfonic acids), hydrophobic groups (aliphatic hydrocarbon chains, alicyclic rings), and aromatic rings (benzene rings and heteroaromatic rings). Each class of pharmacophore atoms aggregated into a P-type supernode, reflecting the pharmacological information of molecule binding to biological targets. M-type supernodes identified known CYP450 metabolically sensitive groups, including tertiary amines (R3N), benzene rings, thioethers (RS-R'), aldehydes (R-CHO), primary alcohols (R-CH2OH), and other known substrate structures of major metabolic enzymes such as CYP3A4 / CYP2D6 / CYP1A2. Each class of metabolic groups aggregated into an M-type supernode, reflecting the reactivity of the molecule during in vivo metabolism. The existing literature on A-type supernodes establishes a SMARTS rule library of toxic functional groups to identify ADR-specific structural alert structures in molecules, including but not limited to aromatic amines (Ar-NH2), nitroaromatic hydrocarbons (Ar-NO2), active esters (RC(=O)-O-R', where R' is a leaving group), and electrophilic Michael receptors (…). -Unsaturated carbonyl groups), epoxy groups, halomethyl groups (R-CH2X, X=Cl / Br / I), etc.; each type of warning structure atom set is aggregated into a type A supernode, and the warning structure type index (integer type ID, mapped to a d-dimensional vector through a learnable embedding matrix) is additionally encoded in the supernode features, so that the model can distinguish the differences in the toxicity mechanism of different warning structures.
[0033] When the same set of atoms simultaneously meets the preset recognition conditions for multiple types of functional group supernodes, the set of atoms is uniquely classified using a priority classification rule of A-type supernodes > M-type supernodes > P-type supernodes. The classified set of atoms will not participate in the construction of other types of supernodes again. This priority design is based on the fact that A-type warning structures have the most direct correlation with ADRs, followed by M-type metabolically sensitive groups, while P-type pharmacophores have the weakest correlation with ADRs. This ensures that structural information with stronger toxicological semantics is prioritized for modeling, avoiding duplicate counting of information.
[0034] The initial representation of each type of supernode is obtained by attention-weighted aggregation of the hidden representations of its constituent atoms. For supernode s, its initial hidden representation is... The calculation is as follows: ; in, Let be the set of atoms contained in supernode s. The final hidden representation vector of atom v after four layers of atomic-level message passing, with attention weights. Calculated by the following formula: ; in, The learnable parameter matrix for attention weights is calculated, and the 512-dimensional concatenated vector is mapped to a scalar score. For atoms u The final implicit representation vector after four layers of atomic-level message passing, u traversed All atoms in the formula are used for normalizing the denominator calculation; The set of atoms contained in supernode s, that is, all atomic nodes belonging to supernode s after matching by SMARTS rules; Let be the type embedding vector of the supernode s. Each of the P-type, M-type, and A-type maintains an independent set of learnable type embeddings, and the type embedding vector is formed by the embedding matrix. The corresponding row vectors are provided and optimized end-to-end during training, enabling the model to weight its contained atoms differentially based on the toxicological semantic category of the supernodes during aggregation.
[0035] In the supernode graph In this context, when there is at least one chemical bond connecting the sets of atoms contained in supernodes s and s', that is... Make Then an edge is established between the two supernodes. Edge feature vectors It is composed of the following dimensions: the number of atoms shared by the two supernodes (1 dimension, normalized), the one-hot encoding of the highest-order bond type among the chemical bond types connecting the two supernodes (4 dimensions), and the shortest path length between the central atoms of the two supernodes (1 dimension, normalized), for a total of 6 dimensions.
[0036] In the supernode graph The system performs two layers of message passing, with each layer having a hidden dimension of 256. The update method employs a gated circular unit (GRU) to maintain information stability across layers. ; in, Let be the hidden representation vector of supernode s after being updated at layer t; For gated recurrent units, the first parameter is the current input state, and the second parameter is the aggregated neighbor information. The hidden state dimension of the GRU is 256, and the update formula follows the standard GRU design. Let be the hidden representation vector of supernode s at layer t, which serves as the input state of GRU at the current time step; Let be the attention weight of neighboring supernode s' to supernode s in layer t; Let be the hidden representation vector of the neighbor supernode s' at level t; Let be the set of neighboring supernodes of supernode s in the supernode graph, that is, all supernodes connected to s by edges in the supernode graph. This is a two-layer feedforward network (input dimension 262, output dimension 256) for supernode message passing. The edge feature vector from supernode s' to supernode s, with dimension 6, is composed of the number of atoms shared between the two supernodes (1D), the one-hot encoding of the highest-level chemical bond type (4D), and the shortest path length of the central atom (1D). Attention weights between supernodes. Calculated by the following formula: ; in, The learnable parameter matrix for calculating attention weights maps the 512-dimensional concatenated vector to a scalar attention score; s'' is the summation traversal variable, traversing all neighboring supernodes of supernode s. This is used to normalize the denominator. After two layers of supernode message passing, the final implicit representation of each supernode is obtained. .
[0037] S23. Atomic-level implicit representations are obtained by pooling operations, and supernode-level implicit representations are obtained by aggregation operations. A gated fusion mechanism is used to fuse the atomic-level molecular representations and the supernode-level molecular representations to obtain the molecular representation vector of the candidate drug.
[0038] Atomic-level molecular representations were obtained using the Set2Set graph pooling method, with a processing step count of [number missing]. Set2Set uses LSTM to maintain a global query vector and gradually extracts global molecular information. The specific process is as follows: ; ; ; ; in, Let t be the LSTM global query vector at step t, with initial values... ; The global query vector at step t-1 is used as the hidden state input for the current step of the LSTM. The attention-weighted molecular aggregation representation at step t-1 is used as the sequence input for the current step of the LSTM, with initial values... ; The attention weight of atom v at the t-th iteration reflects the degree of attention the current query vector pays to each atom; For the t-th iteration, with V is the molecular aggregation representation obtained by weighting the implicit representations of all atoms; V is the atomic-level molecular diagram. The set of nodes, corresponding to all atoms in the molecule; This is an atomic-level molecular representation vector, obtained from the final query vector. With final aggregation representation obtained by piecing together; This represents the total number of iteration steps for Set2Set.
[0039] Supernode-level molecular representations are obtained through type-aware weighted aggregation: ; ; in, The supernode-level molecular vector is obtained by weighted aggregation of the final implicit representations of all supernodes in the supernode graph; The global importance weight of supernode s is calculated by the type-aware attention mechanism; V' is the final implicit representation vector of supernode s after message passing through two layers of the supernode graph; V' is the supernode graph. The set of nodes corresponds to all functional group supernodes identified in the molecule; For summation index; Supernode type The corresponding learnable query vectors, P-type, M-type and A-type each maintain an independent set of query vectors, so that the model can automatically learn the relative importance of different types of functional groups to ADR prediction during training; The type marker for the supernode s, with a value of One of them corresponds to the pharmacophore supernode, the metabolism-sensitive group supernode, and the ADR warning structure supernode, respectively.
[0040] The two-layer feature fusion employs a gating mechanism, which satisfies the following: ; ; in, A vector representing the atomic-level molecules. This is the supernode-level molecular representation vector. This represents the fused molecular representation vector. As an element-wise gated vector, the fusion ratio of each dimension is constrained to the range (0,1) by the Sigmoid function; This is an atomic-level characteristic transformation matrix. This is the supernode-level feature transformation matrix. The gating calculation parameter matrix has the following input: and The concatenated vector has a dimension of 512 + 256 = 768; A learnable bias vector; Use the Sigmoid activation function; This is an element-wise Hadamard product. Each dimension of the gating vector g independently controls the fusion ratio of atomic-level and supernode-level information in that dimension, enabling the model to adaptively balance the contributions of the two levels of semantic information to the structural characteristics of different molecules. The final output is... As the molecular representation vector, it is passed to the next step.
[0041] The aforementioned two-layer molecular graph encoding mechanism, based on atomic-level message passing, introduces three types of toxicological semantic supernodes to explicitly incorporate functional group-level toxicological structural information into the molecular representation learning process, effectively compensating for the shortcomings of single-layer atomic-level encoding in the expression of toxicity-related local chemical patterns. The gated fusion mechanism allows the contribution of the two-level representation to be dynamically adjusted according to the characteristics of the molecular structure, avoiding the limitations of fixed-weight fusion in adaptability to drugs with different structural types.
[0042] In one embodiment of the present invention, such as Figure 3As shown, step S2, which uses a collaborative multi-task training framework for training, specifically includes: (1) Extract the labels of training sample drugs on the multidimensional safety attributes of each drug from the drug knowledge base, and use them as training supervision signals for auxiliary tasks.
[0043] Specifically, the following nine ADMET auxiliary tasks were extracted from existing public knowledge bases to identify candidate drugs: CYP1A2 inhibition (binary classification), CYP2C9 inhibition (binary classification), CYP2C19 inhibition (binary classification), CYP2D6 inhibition (binary classification), CYP3A4 inhibition (binary classification), P-gp substrate (binary classification), P-gp inhibition (binary classification), hepatotoxicity (binary classification), and cardiotoxicity (binary classification), for a total of K=9 auxiliary tasks. ADMET refers to the collective properties of a drug's absorption, distribution, metabolism, excretion, and toxicity. Among these, CYP450 metabolic enzyme inhibition affects the in vivo concentration of drugs sharing metabolic pathways, P-gp-related properties affect drug distribution and accumulation, and hepatotoxicity and cardiotoxicity are the most common systemic ADR endpoints. A clear biological association exists between the above auxiliary tasks and ADRs.
[0044] (2) Using a molecular graph encoder as a shared backbone network, feature encoding is performed on the molecular structure information of the training samples to obtain the molecular representation vector of the training samples; a main task prediction head is set for the main task of drug adverse reaction prediction, and an independent auxiliary task prediction head is set for each auxiliary task of drug multidimensional safety attribute prediction. Each prediction head takes the molecular representation vector of the training samples as input.
[0045] Specifically, the molecular graph encoder in step S2 (including the atomic-level MPNN and the supernode two-layer encoder) serves as a shared backbone network, and its parameters are shared by all tasks and jointly optimized under the gradient of the total loss function. The main task (ADR prediction) has an independent main task prediction head (3-layer MLP, see step S3 for details); each auxiliary task k has an independent auxiliary prediction head. The structure is a two-layer MLP ( The activation function for each hidden layer is ReLU, and the output of the last layer is Sigmoid. The input to each auxiliary prediction head is a molecular representation vector. .
[0046] (3) Based on the molecular representation vector of the training samples, calculate the dynamic task association weight of molecular perception independently for each auxiliary task, and weight the loss of each auxiliary task according to the dynamic task association weight.
[0047] Dynamic task association weights satisfy: ; ; in, The dynamic contribution weight of the k-th auxiliary task to the current molecule is independently activated by the Sigmoid function. The weights of each auxiliary task are independent of each other and do not compete with each other. The affinity score of the current molecule to the k-th auxiliary task reflects the degree of correlation between the structural features of the molecule and the k-th auxiliary task. Let be the molecular feature projection matrix of the k-th auxiliary task; It is the bias vector; The task-aware query vector for the kth auxiliary task; For the Sigmoid function; , , All parameters are learnable and optimized end-to-end during training. An independent Sigmoid activation design, rather than Softmax normalization, is used to ensure that each... It can independently reflect the structural relevance of the current molecule to the auxiliary task, rather than being forcibly suppressed to meet the weight normalization constraint.
[0048] (4) The shared backbone network is optimized end-to-end by using the joint loss of the main task loss and the weighted auxiliary task loss to obtain the optimized molecular graph encoder.
[0049] The joint loss function satisfies: ; in, The total loss of the main task is to predict adverse drug reactions; For the first Binary cross-entropy loss for each auxiliary task; Let be the label availability mask coefficient for the k-th auxiliary task, when the label for the k-th auxiliary task is available. When the label is missing This prevents the auxiliary task from participating in the current backpropagation; This represents the number of auxiliary tasks with valid tags in the current batch, used to normalize the total auxiliary loss level and avoid fluctuations in the total loss due to differences in the number of valid tasks in each batch. The overall weight coefficients for the auxiliary tasks are set by AUROC optimization on the validation set for the main task, with a default value of 0.3; K represents the total number of auxiliary tasks, which is K=9 in this embodiment. The combination of the masking strategy and K' normalization ensures the numerical stability of the total loss function in scenarios with sparse labeled data.
[0050] After training, the auxiliary task prediction head and dynamic weight calculation module do not participate in the inference phase calculation; they only use the molecular graph encoder, which has been fully optimized through co-training, to directly output the results. The reasoning process is completely consistent with that without auxiliary tasks, without introducing additional delays.
[0051] The aforementioned collaborative multi-task training framework actively incorporates safety-related information from the drug ADMET knowledge base into the parameter learning process of the molecular graph encoder by using molecularly-aware dynamic task-associated weights. This allows the molecular representation output by the encoder to internalize the systematic association knowledge of ADR-ADMET, thereby enhancing the sensitivity of the molecular representation to ADR-related biological characteristics. Compared to fixed-weight multi-task schemes, molecularly-aware dynamic weights enable the contribution of each auxiliary task to drugs with different structural types to be adaptively adjusted, avoiding interference from structure-independent auxiliary tasks on the gradient of the shared encoder.
[0052] In one embodiment of the present invention, step S3 specifically includes: Molecular representation vector of candidate drugs Input Drug Adverse Reaction Multi-Label Prediction Network Predicting candidate drugs in The probability of occurrence for each ADR category. The multi-label drug adverse reaction prediction network is a multilayer perceptron with two hidden layers, and the network structure is as follows: ,in The total number of ADR categories, based on the SIDER dataset. This corresponds to the ADR categories of 27 organ systems. The first hidden layer has a dimension of 256, the second hidden layer has a dimension of 128, and both layers use ReLU activation function. Each hidden layer is followed by a batch normalization (BatchNorm) layer and a dropout layer with a dropout rate of 0.2. The last layer is the output layer with a dimension of... Without an activation function, output directly. Original logits vector The prediction process can be represented as: ; in, The nth component The original predicted logits value corresponding to the nth ADR category, that is, the original predicted score of the nth adverse drug reaction category corresponding to the nth component of the original logits vector.
[0053] During training, the binary cross-entropy loss function is used independently for each ADR category, and the total loss of the main task is... for Mean of loss for each category: ; in, is the total number of ADR categories; n is the ADR category index. ; This is the true label for the nth ADR category, where 1 indicates that the drug has this type of ADR, and 0 indicates that it does not. For the Sigmoid function; The original logits vector The nth component, i.e., the multi-label prediction network The unactivated predicted score output for the nth ADR category. Training data is provided using a publicly available single-drug ADR dataset, with each drug sample containing... Multi-hot label vectors. The optimizer uses Adam (learning rate). Weight decay The batch size is 256, the number of training rounds is up to 200, and early stopping is performed based on the AUROC mean of each ADR category in the validation set (patience=20 rounds). The optimal model weights in the validation set are saved for subsequent inference.
[0054] In one embodiment of the present invention, such as Figure 5 As shown, step S4 specifically includes: S41. After the molecular graph encoder and the multi-label prediction network for adverse drug reactions are trained, an independent calibration set that does not overlap with the training set is used as the basis. Based on the true labels and original logits vectors of each sample in each adverse drug reaction category in the calibration set, a joint negative log-likelihood loss function for the temperature parameter T is constructed with the temperature parameter T as the only variable to be optimized. The L-BFGS optimizer is used to minimize the joint negative log-likelihood loss function to obtain the calibrated temperature parameter.
[0055] The joint negative log-likelihood loss function satisfies: ; in, The joint negative log-likelihood loss, calculated on the calibration set with temperature parameter T as the only variable, is used for parameter optimization under temperature scaling. The total number of samples in the independent calibration set. For the first The true label of a sample in the nth ADR category For calibration set sample index; For the calibration set The original logits value of a sample in the nth ADR category is obtained by forward inference of the trained model and remains fixed during the calibration phase. For the Sigmoid function, Let T be the temperature parameter, the only scalar variable to be optimized in temperature scaling. When T > 1, the predicted probability distribution is softened (confidence is reduced), and when T < 1, the distribution is made sharper (confidence is increased). The L-BFGS optimizer is used to minimize the above equation with respect to T, with a maximum of 50 iterations and a convergence threshold. The initial value of T is set to 1.0, and the search range is limited to (0.5, 5.0) to ensure numerical stability. Finally, the calibrated temperature parameters are obtained. The temperature scaling method adjusts the scale of the output probability without changing the core parameters of the model. It is a post-processing calibration scheme that introduces very few parameters (only one scalar) and the calibration overhead is negligible.
[0056] S42. For any candidate drug, use the calibrated temperature parameters. Temperature scaling is applied to each component of the original logits vector to obtain the calibration probability of each adverse drug reaction category.
[0057] ; in, The original logits vector The nth component, For the calibrated temperature parameters, For the Sigmoid function, This represents the calibration probability for the nth adverse drug reaction (ADR) category. All ADR categories share the same temperature parameter. Temperature scaling does not change the ranking of predictions for each category; it only adjusts the absolute value of the predicted probability to make it statistically consistent with the actual occurrence rate.
[0058] S43. Based on the calibration probability of each adverse drug reaction category, perform confidence level classification to obtain the corresponding confidence level and output an adverse drug reaction prediction report.
[0059] Specifically, based on the calibration probability of each adverse drug reaction category. Perform confidence level grading to obtain the corresponding confidence level: when When the risk is high, it indicates a higher risk of this type of ADR, and it is recommended to prioritize clinical safety monitoring or experimental validation; when When the ADR is classified as medium confidence, it indicates that this ADR category carries some risk, and further evaluation based on other information is recommended; when When the threshold is reached, it is judged as low confidence, indicating that the risk of this ADR category is low. The three thresholds (0.95 and 0.75) are determined based on a comprehensive consideration of minimizing the expected calibration error (ECE) on the validation set and clinical practicality, and can be adjusted according to specific application scenarios through the configuration interface.
[0060] After completing the confidence level classification, the system... The prediction results for each ADR category were structured and organized according to calibration probability. The results are sorted in descending order, prioritizing high-risk ADR categories, and a statistical summary is generated, including the number and percentage of high / medium / low-risk ADR categories. Each result record includes the following fields: candidate drug identifier (drug name and DrugBankID or PubChem CID), ADR category name and its corresponding System Organ Class (SOC), and raw predicted logits value. Calibration probability The report includes a confidence level and its corresponding clinical implications, as well as a preliminary toxicological mechanism description automatically generated based on the M-type and A-type supernode information identified in step S22 (e.g., detection of a CYP3A4 metabolically sensitive group, which may produce hepatotoxic metabolites upon metabolic activation; or detection of an aromatic amine warning structure, indicating a potential risk of idiotoxicity). The final output supports JSON structured data format and can optionally be exported as a complete prediction report in Markdown, PDF, or Word format.
[0061] The temperature-reduced reliability calibration step corrects the overconfidence probability of the model's original output to a level consistent with the actual ADR incidence rate statistics. Without changing the model parameters and prediction ranking, it effectively improves the statistical reliability of the prediction results and provides a confidence quantification index with practical reference value for clinical drug safety assessment.
[0062] The present invention also provides a drug adverse reaction prediction system based on graph neural networks. The system is used to implement the above-mentioned drug adverse reaction prediction method. The system includes a molecular structure verification module, a two-layer molecular graph encoding module, a drug adverse reaction multi-label prediction module, a confidence calibration module, and a report output module.
[0063] The molecular structure verification module receives the SMILES string input of the candidate drug, performs syntactic validity verification and molecular validity verification, filters out molecular structures that do not meet the requirements, and transmits the molecular structure information that passes the verification to the two-layer molecular graph encoding module.
[0064] The two-layer molecular graph encoding module comprises an atomic-level message passing subunit, a functional group supernode construction subunit, and a gating fusion subunit. These subunits sequentially perform atomic-level feature encoding, functional group-level feature encoding, and two-layer representation fusion, outputting the molecular representation vector of the candidate drug. This module embeds the molecular graph encoder weights obtained offline through a collaborative multi-task training framework and performs only forward propagation during the inference phase.
[0065] The multi-label prediction module for adverse drug reactions receives molecular representation vectors, calculates the original logits vectors for each adverse drug reaction category through the multi-label prediction network, and transmits the results to the confidence calibration module.
[0066] The confidence calibration module performs a temperature scaling operation on the original logits vector based on the temperature parameters fitted during the offline calibration phase, maps the calibrated probability values to a preset confidence level range, and outputs the calibration probability and confidence level for each adverse drug reaction category.
[0067] The report output module summarizes the output results of the confidence calibration module and generates a structured adverse drug reaction prediction report according to the preset report template. The report content includes the predicted probability, confidence level and risk warning of each adverse drug reaction category. It supports JSON format interface return and report display on the visualization page.
[0068] The modules mentioned above are called sequentially through a unified data interface. Each module has independent functions and can be replaced or upgraded independently according to actual deployment needs.
[0069] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting adverse drug reactions based on graph neural networks, characterized in that, Includes the following steps: S1. Receive and verify the molecular structure information of the candidate drug to obtain the verified molecular structure information of the candidate drug. S2. A two-layer molecular graph encoding mechanism is used to encode the molecular structure information of the verified candidate drugs to obtain the molecular representation vector of the candidate drugs. The molecular graph encoder used in the two-layer molecular graph encoding mechanism is trained offline through a collaborative multi-task training framework with adverse drug reaction prediction as the main task and multi-dimensional safety attribute prediction as the auxiliary task. S3. Input the molecular representation vector of the candidate drug into the multi-label prediction network for adverse drug reactions to predict the probability of occurrence of the candidate drug in each adverse drug reaction category, and obtain the original logits vector of each adverse drug reaction category. S4. The original logits vectors of each adverse drug reaction category are calibrated for confidence using the temperature scaling method to obtain the calibration probability and corresponding confidence level of each adverse drug reaction category, and the adverse drug reaction prediction report is output. Step S2 employs a two-layer molecular graph encoding mechanism to feature-encode the verified candidate drug molecular structure information, specifically including: S21. Convert the verified candidate drug molecular structure information into a molecular graph, and use a message passing neural network to perform atomic-level message passing on the molecular graph to obtain the atomic-level implicit representation of each atom in the molecular graph. S22. Based on the preset toxicological semantic rules, functional groups are identified in the molecular graph. The sets of atoms that meet the identification conditions are aggregated into functional group supernodes. A supernode graph is constructed with each functional group supernode as a node. The edge between two supernodes in the supernode graph is determined by the chemical bond connection relationship between the sets of atoms they contain. A message passing neural network is used to perform supernode-level message passing on the supernode graph to obtain the supernode-level hidden representation of each functional group supernode. The functional supernodes include P-type supernodes, M-type supernodes and A-type supernodes. P-type supernodes are used to characterize pharmacophore structures, M-type supernodes are used to characterize metabolically sensitive group structures, and A-type supernodes are used to characterize warning structures related to adverse drug reactions. When the same set of atoms simultaneously meets the preset recognition conditions for multiple types of functional groups supernodes, the set of atoms is uniquely classified using the priority classification rule of A-type supernode > M-type supernode > P-type supernode, and a supernode graph is constructed. S23. Atomic-level implicit representation is obtained by pooling operation, and supernode-level implicit representation is obtained by aggregation operation. A gated fusion mechanism is used to fuse atomic-level molecular representation and supernode-level molecular representation to obtain molecular representation vector of candidate drug. The training process of the collaborative multi-task training framework in step S2 specifically includes: Step 1: Extract the labels of training sample drugs on the multidimensional safety attributes of each drug from the drug knowledge base, and use them as training supervision signals for auxiliary tasks; Step 2: Using a molecular graph encoder as a shared backbone network, feature encoding is performed on the molecular structure information of the training samples to obtain the molecular representation vector of the training samples; a main task prediction head is set for the main task of drug adverse reaction prediction, and independent auxiliary task prediction heads are set for each auxiliary task of drug multidimensional safety attribute prediction. Each prediction head takes the molecular representation vector of the training samples as input; the auxiliary task prediction heads include CYP1A2 inhibition, CYP2C9 inhibition, CYP2C19 inhibition, CYP2D6 inhibition, CYP3A4 inhibition, P-gp substrate, P-gp inhibition, hepatotoxicity, and cardiotoxicity extracted from the drug knowledge base; Step 3: Based on the molecular representation vector of the training samples, calculate the dynamic task association weight of molecular perception independently for each auxiliary task, and weight the loss of each auxiliary task according to the dynamic task association weight. Step 4: Perform end-to-end optimization training on the shared backbone network using the joint loss of the main task loss and the weighted auxiliary task loss to obtain the optimized molecular graph encoder.
2. The method for predicting adverse drug reactions based on graph neural networks as described in claim 1, characterized in that, The gating fusion mechanism described in step S23 satisfies: ; ; in, A vector representing the atomic-level molecules. This is the supernode-level molecular representation vector. This represents the fused molecular representation vector. For element-wise gated vectors, For the Sigmoid function, The hyperbolic tangent activation function is used. For element-wise Hadamard product, This is a vector concatenation operation. , and For learnable parameters, It is a learnable bias vector.
3. The method for predicting adverse drug reactions based on graph neural networks as described in claim 1, characterized in that, In step 3, the dynamic task association weights satisfy the following: ; The joint loss in step 4 satisfies: ; in, This represents the fused molecular representation vector. For the first k Dynamic task association weights for each auxiliary task. For the Sigmoid function, For the first k Task-aware query vectors for each auxiliary task. For the first k Molecular feature projection matrices for each auxiliary task For the first k Bias vectors for each auxiliary task K To assist in the total number of tasks, For the first k The label availability mask coefficient for each auxiliary task. For the current batch The number of auxiliary tasks with a value of 1 varies with the batch size. , To predict the main task of adverse drug reaction, For the first k The loss of an auxiliary task This is used as an auxiliary task's overall weighting coefficient.
4. The method for predicting adverse drug reactions based on graph neural networks as described in claim 1, characterized in that, In step S3, the multi-label prediction network for adverse drug reactions is a multilayer perceptron with two hidden layers. The multilayer perceptron includes a first hidden layer, a second hidden layer and an output layer in sequence. The activation function of the first hidden layer and the second hidden layer is ReLU. Each hidden layer is followed by a batch normalization layer and a random inactivation layer. The multilayer perceptron receives the molecular representation vector of the candidate drug, performs feature transformation layer by layer through the first hidden layer and the second hidden layer, and outputs the original logits vector corresponding to each adverse drug reaction category by the output layer. The nth component of the original logits vector corresponds to the original prediction score of the nth adverse drug reaction category.
5. The method for predicting adverse drug reactions based on graph neural networks as described in claim 1, characterized in that, Step S4 specifically includes: S41. After the molecular graph encoder and the multi-label prediction network for adverse drug reactions are trained, an independent calibration set that does not overlap with the training set is used as the basis. Based on the true labels and original logits vectors of each sample in each adverse drug reaction category in the calibration set, the temperature parameter is used as the only variable to be optimized. A joint negative log-likelihood loss function for the temperature parameter is constructed. The L-BFGS optimizer is used to minimize the joint negative log-likelihood loss function to obtain the calibrated temperature parameter. S42. For any candidate drug, use the calibrated temperature parameters to perform temperature scaling on each component in the original logits vector to obtain the calibrated probability of each adverse reaction category of the drug. S43. Based on the calibration probability of each adverse drug reaction category, perform confidence level classification to obtain the corresponding confidence level, and output an adverse drug reaction prediction report.
6. The method for predicting adverse drug reactions based on graph neural networks as described in claim 5, characterized in that, The calibration probabilities for each adverse drug reaction category in step S42 satisfy the following: ; in, The nth component of the original logits vector For the calibrated temperature parameters, For the Sigmoid function, This represents the calibration probability for the nth adverse drug reaction category.
7. A drug adverse reaction prediction system based on graph neural networks, characterized in that, The system is used to implement the adverse drug reaction prediction method as described in any one of claims 1 to 6, comprising: The molecular structure verification module is used to receive the molecular structure information of candidate drugs, verify it, and output the verified molecular structure information of candidate drugs. A two-layer molecular graph encoding module is used to perform feature encoding on the verified candidate drug molecular structure information using a two-layer molecular graph encoding mechanism, and output the molecular representation vector of the candidate drug. The multi-label prediction module for adverse drug reactions is used to receive the molecular representation vector of candidate drugs and input it into the multi-label prediction network for adverse drug reactions. It predicts the probability of occurrence of candidate drugs in each adverse drug reaction category and outputs the original logits vector of each adverse drug reaction category. The confidence calibration module is used to perform confidence calibration on the original logits vector of each adverse drug reaction category using a temperature scaling method, so as to obtain the calibration probability and corresponding confidence level of each adverse drug reaction category. The report output module is used to output adverse drug reaction prediction reports based on the calibration probability and corresponding confidence level of each adverse drug reaction category.
Citation Information
Patent Citations
Drug interaction prediction method and system based on graph neural network
CN116403730A