Drug toxicity prediction method based on metabolism guidance
By constructing a metabolic hierarchy map and using a multi-level feature learning method, the problem of existing drug toxicity prediction models failing to consider toxicity during metabolism is solved, improving the accuracy of drug toxicity prediction and the ability to capture molecular structure information, especially the feature learning of benzene ring structure.
Patent Information
- Application Number
- CN202510950157.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-31
AI Technical Summary
Existing drug toxicity prediction models fail to adequately account for the toxicity that may occur during metabolism, especially the impact of non-covalent physical interactions on the benzene ring structure. This results in drugs becoming toxic after entering animals or humans, and the molecular structure information is not fully revealed.
By constructing a metabolic hierarchy graph, molecules are decomposed using metabolic reaction rules, and functional group graphs are obtained by combining BRICS and motif decomposition methods. Aromatic structure subgraphs are introduced, and molecular graph features are learned using message passing networks and attention mechanisms. Different decomposition strategies are fused using cross-attention mechanisms to predict the potential toxicity of drugs during metabolism.
It improves the accuracy of drug toxicity prediction, enables earlier identification of potential toxicity, reduces drug development costs, enhances the characteristic learning of benzene ring structure, and captures multi-level features of molecular structure information.
Smart Images

Figure CN120877926A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of drug toxicity prediction technology, specifically a drug toxicity prediction method based on metabolism guidance. Background Technology
[0002] Drug toxicity prediction is a crucial step in drug discovery and development. Its purpose is to identify and prioritize compounds most likely to be safe and effective for human use, while mitigating the high-cost failure risk due to toxicity issues in later stages of drug development. According to the paper (Artificial Intelligence in Drug Toxicity Prediction: Recent Advances, Challenges, and Future Perspectives), over 30% of drug candidates may be abandoned due to toxicity problems. With the development of artificial intelligence technologies, particularly the application of machine learning and deep learning algorithms, the accuracy and efficiency of drug toxicity prediction have been significantly improved. AI technology can provide more efficient methods to identify the potential toxic effects of new compounds before they enter human clinical trials, thereby saving time and costs. In the paper "MolFPG (MolFPG: Multi-level fingerprint-based Graph Transformer for accurate and robust drug toxicity prediction)" by S. Teng, C. Yin, and Y. Wang (2023), the authors proposed an innovative molecular fingerprint graph transformer framework called MolFPG. This framework combines multiple molecular fingerprinting techniques and introduces a graph transformer for graph representation learning to improve the learning ability of drug molecule representations, enabling the model to better predict drug toxicity. Patent application CN202311457739.0 provides a structure-activity relationship (QSAR) model and method for predicting the acute oral toxicity of tetrazolium compounds in mice. This invention is the first to discover that six molecular descriptors have a strong correlation with the toxicity of tetrazolium compounds: nwHBa, ATSC1c, ATSC3v, VE3_dt, BCUTP-1h, and MATS4m, resulting in a QSAR model for predicting the toxicity of tetrazolium compounds, thus enabling the prediction of acute oral toxicity of tetrazolium compounds.
[0003] However, these toxicity predictions only consider the molecule itself and do not take into account the toxic relationships that may arise during metabolism. This leads to situations where a drug shows no toxicity during machine learning experiments and cell experiments, but becomes toxic after entering animals or humans, preventing it from being marketed. Furthermore, for structures like benzene rings, the learning methods used for other rings are primarily the same, which prevents the learning of covalent bond structures in benzene-like structures and fails to consider the impact of non-covalent physical interactions between benzene rings and between benzene rings and cations on drug toxicity. Previous models used only one method for subgraph partitioning, which cannot fully capture information about different functional groups and cannot adequately reveal the influence of drug structure information on drug toxicity. Summary of the Invention
[0004] The purpose of this invention is to provide a metabolism-guided hierarchical graph for drug toxicity prediction. This method aims to guide prodrugs or non-toxic substances through metabolic approaches, considering their toxicity during metabolic processes (e.g., the conversion of prodrugs into active ingredients during metabolism), thereby predicting the potential toxicity of prodrugs or non-toxic substances. By guiding metabolism, it is hoped that accurate predictions can be made regarding whether prodrugs will produce toxicity in vivo. Simultaneously, the introduction of benzene ring-like hypergraph structures facilitates the characterization and learning of the pi-bond features of benzene rings. By using different rules to divide the molecular graph, different functional group (subgraph) representations can be obtained, helping to reveal the different functional group representations of the molecule and enhancing the model's predictive ability for drug toxicity.
[0005] The technical solution of this invention is:
[0006] A metabolism-guided method for predicting drug toxicity includes the following steps:
[0007] S1. Construct training data, specifically: obtain the SMILES formula of drug molecules and the corresponding toxicity attributes from existing toxicology data to form training data, wherein the toxicity attributes include two categories: toxicity type and toxicity value.
[0008] S2. Preprocess the training data, specifically as follows:
[0009] a) Based on the SMILES formula of the drug molecule and predefined metabolic reaction rules, the molecule is decomposed to obtain a metabolic graph. Specifically, the reaction sites of the molecules represented by the SMILES formula are located according to the metabolic reaction rules. The SMILES formulas of the initial metabolites are generated by breaking the corresponding chemical bonds. Each SMILES formula is used as a metabolic node, and the metabolic reaction is used as a directed edge to construct a directed graph. The SMILES formulas of new products are continuously generated based on the SMILES formulas of the initial metabolites, and a directed graph is constructed until the SMILES formula of the new product can no longer match the corresponding metabolic rule subgraph, thus constructing a complete metabolic graph. This metabolic graph uses the SMILES formula of the parent molecule as the root node, the SMILES formulas of multi-level metabolites as child nodes, and the metabolic transformation relationships as directed edges. It can completely map the network path from the original molecule to all possible metabolites, thus constructing the metabolic graph. For each node in the constructed metabologram corresponding to the SMILES formula, the rdkit tool is used to convert it into a molecular diagram. This yields the graph structure corresponding to each metabolic node;
[0010] b) The molecular map was analyzed using the bond-based BRICS method and the motif-based decomposition method. The function group diagrams were split to obtain two different scales. ;
[0011] c) Molecular diagram obtained from the above In the process, all aromatic structure subgraphs are extracted. If a molecular graph contains multiple aromatic subgraphs, a new node is added for each subgraph. This new node represents the π-bond interactions between atomic nodes in the subgraph, and the newly added node is connected to all other nodes in the subgraph to obtain the final atomic graph. ;
[0012] d) Based on metabolic maps , sensual group diagram and atomic diagram Constructing a drug metabolism hierarchy map The structure consists of a metabolism diagram at the top layer, with each node representing a molecule; a functional group diagram at the middle layer, with functional groups as nodes; and an atomic diagram at the bottom layer. The three layers are connected by edges to represent the relationships between the layers. The top two layers connect the nodes of the metabolism layer with the functional group nodes corresponding to their molecules, while the bottom two layers construct their connections based on the inclusion relationship between functional groups and atomic layer nodes.
[0013] S3, Based on the constructed drug metabolism hierarchy map Studying the chemical and topological characterization of molecules, specifically including:
[0014] Features of atoms within an atomic graph are aggregated using a message passing network:
[0015]
[0016]
[0017] in For message passing functions, For vertex update functions, Represents atomic graph nodes After the t-th message passing step is completed, the message embedding is calculated based on the features of its neighboring nodes and edge attributes. This indicates the time after the t-th update. The hidden state of a node, where 'a' represents the atomic level. Atomic diagram In and nodes The set of adjacent nodes, This indicates the time after the t-th update. The hidden state of a node, express , The edge between;
[0018] The influence of the atomic graph on functional clique layer nodes is learned through an attention mechanism, and features of functional clique nodes within the functional clique graph are aggregated using a message-passing network, including:
[0019] The influence of each atom in the atomic diagram on the functional group diagram is learned through an attention mechanism, as follows:
[0020]
[0021]
[0022]
[0023]
[0024]
[0025] in The activation function is tanh. This represents the atomic graph node after the t-th update. The hidden state, , Representative Hierarchical Diagram Nodes in the functional group diagram Adjacent atomic layer nodes, For the projection matrix, This represents the atomic features after projection at the t-th update. This represents the functional group feature after projection at the t-th update. For the atomic node at the t-th update Functional group graph nodes Similarity of influence This represents the influence weight of each atom at the t-th update; This represents a splicing operation;
[0026] Characteristics of functional clique nodes within a functional clique graph are aggregated using a message-passing network:
[0027]
[0028]
[0029] in For message passing functions, For vertex update functions, Represents functional group graph nodes After the t-th message passing step is completed, the message embedding is calculated based on the features of its neighboring nodes and edge attributes. This represents the node after the t-th update. The hidden state; Represents sensual group diagram In and nodes Adjacent nodes, This indicates the time after the t-th update. The hidden state of a node, express , The edge between;
[0030] The influence of functional group graphs on metabolic layer molecules is learned through an attention mechanism. Molecular features of metabolic graphs under different functional group decomposition strategies are weighted and fused to obtain metabolic graph node features that fuse functional group graph node features, including:
[0031] The influence of functional group features at different splits on metabolic layer molecules is learned through an attention mechanism, as shown in the following formula:
[0032]
[0033]
[0034]
[0035]
[0036]
[0037] in The activation function is tanh. This represents the metabolic graph after the t-th update. Nodes in Polymer functional group diagram The hidden state of the middle node; Representative Hierarchical Diagram Middle and metabolic layer nodes The set of adjacent functional group layer nodes, For the projection matrix, This represents the functional group feature after projection at the t-th update. This represents the aggregated functional group diagram after projection at the t-th update. Metabolic graph node features. For the facultative community graph node at the t-th update Metabolic map nodes Similarity of influence This represents the influence weight of each functional group at the t-th update. This represents a splicing operation;
[0038] The similarity weights between the metabolic subgraph node features generated by BRICS splitting and motif construction algorithm are dynamically calculated using a cross-attention mechanism, thereby achieving adaptive weighted fusion of metabolic molecular features under the two splitting strategies:
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047] in and These represent the nodes of the metabolic layer molecular map obtained through two different splitting methods. Features These represent the learnable weight matrices; Represent and The dimension; This represents a fused feature formed by incorporating key splitting features into motif features using an attention method; This represents a fusion feature formed by integrating motif features into key splitting features through an attention mechanism;
[0048] After obtaining two fusion characterizations, the overall features of the molecule are obtained through weighted averaging:
[0049]
[0050] in For learnable weight matrix, This indicates the node after merging the two splitting methods. Structural features;
[0051] S4. Toxicity prediction based on molecular structure characteristics and metabolites, including:
[0052] Predicting molecular toxicity based on molecular structure characteristics:
[0053]
[0054] This indicates the predicted molecular toxicity obtained by considering only molecular structure features. According to the toxicity attributes defined in S1, there are two categories: toxicity type and toxicity value. Therefore, in classification tasks... This indicates the probability that a drug is toxic; in regression tasks, This indicates the predicted toxicity level;
[0055] Toxicity prediction based on molecular structural features and their metabolites:
[0056]
[0057] in This represents the toxicity probability or numerical value corresponding to node a, which points to node i in the metabolic graph. Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. This represents the set of nodes pointing to i in the metabolic graph. This indicates consideration of the molecular toxicity of polymerized metabolites;
[0058] Loss functions are defined as classification task loss and regression task loss:
[0059]
[0060]
[0061] Where N represents the number of training samples, Indicates the true toxicity label. This indicates that the molecular toxicity of the polymerized metabolites should be taken into consideration. This indicates that molecular toxicity was not taken into account when considering metabolites;
[0062] S5. Using the training data in S1, after preprocessing in S2, feature extraction is performed using the method in S3, and the loss value is calculated using the loss function in S4. Finally, the iteration stops when the set conditions are met, and the required network parameters are obtained.
[0063] S6. To predict the toxicity of the target drug, specifically:
[0064] For the target drug molecule, a metabolic map of the target drug is obtained according to method S2. Then, for molecules in the metabolic map that include the target drug molecule and all its metabolites, a hierarchical map of each molecular node is obtained. Finally, the structural representation of the drug molecule is extracted according to method S3. and structural characteristics of metabolites Finally, the network parameters obtained from training are used to obtain predicted values of drug toxicity based on drug molecular structure according to the S4 method. and predicted values based on metabolite structure Since the target drug molecule and its metabolites are not included in the metabolic layer weights obtained during training, the weight parameters of the existing metabolic pathways are transferred to the metabolic process of the new molecule. The specific process of weight transfer and the prediction results of the new drug molecule are as follows:
[0065]
[0066]
[0067] in, This indicates a predicted value for drug toxicity. For the structural characteristics of the drug to be inferred, For the metabolic diagram and Structural characteristics of drugs that share the same metabolite a; Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. Indicates metabolites after migration Treatment of speculative drugs Toxicity impact weight.
[0068] The beneficial effects of this invention are that, compared with existing calculations, the technical solution proposed in this invention uses different characterization methods to learn molecules from molecular maps, enabling us to capture different features at different levels and under different segmentation methods. Guided by metabolic maps, reactants can acquire product-related properties, which can be used to explain the toxicity changes they produce during metabolism. Attached Figure Description
[0069] Figure 1 This is the overall flowchart of the present invention.
[0070] Figure 2 This is a schematic diagram of a metabolic mapping-guided network.
[0071] Figure 3 This is a schematic diagram of a drug hierarchy network.
[0072] Figure 4 This is a schematic diagram of the drug structure feature learning process. Detailed Implementation
[0073] The technical principles and solutions of the present invention will now be described in detail with reference to the accompanying drawings:
[0074] Figure 1 This is the main flowchart of this technical solution. (For example...) Figure 1 As shown, the specific steps of the metabolism-guided hierarchical graph-based drug toxicity prediction method proposed in this invention are as follows:
[0075] 1. Data Acquisition and Preprocessing
[0076] This approach uses data from the TOXRIC dataset for attribute prediction. Each attribute record in the dataset consists of two parts: the first is the SMILES representation of the drug; the second is the drug attribute. Among these, attributes such as whether the drug has acute toxicity, carcinogenicity, or mutagenicity can be used for classification tasks; while attributes with specific numerical values, such as the acute toxicity LD50 value (median lethal dose) and LC50 value (median lethal concentration), are used for regression tasks.
[0077] 2. Construction of a drug metabolism hierarchy map
[0078] (1) Construction of metabolic maps and acquisition of molecular maps based on structural similarity
[0079] Based on the SMILES formulas of drug molecules and predefined metabolic reaction rules, a metabolic graph is obtained by decomposing the molecules. Specifically, the reaction sites of the molecules represented by the SMILES formulas are located according to the metabolic reaction rules. The SMILES formulas of the initial metabolites are generated by breaking the corresponding chemical bonds. Each SMILES formula is used as a metabolic node, and the metabolic reaction is used as a directed edge to construct a directed graph. This process continues, generating SMILES formulas of new products based on the SMILES formulas of the initial metabolites and constructing a directed graph, until the SMILES formula of a new product can no longer match the corresponding metabolic rule subgraph, thus constructing a complete metabolic graph. This metabolic graph uses the SMILES formula of the parent molecule as the root node, the SMILES formulas of multi-level metabolites as child nodes, and metabolic transformation relationships as directed edges, completely mapping the network path from the original molecule to all possible metabolites. For each node in the constructed metabolic graph, the rdkit tool is used to convert it into a molecular graph. This allows us to obtain the graph structure corresponding to each metabolic node.
[0080] (2) Acquisition of faculty group diagrams
[0081] For each node in the metabolic graph, the corresponding molecular graph Two types of splitting methods were employed to obtain functional group diagrams at different scales. One type is a bond-based splitting method, using the BRICS method, which is mainly used to split chemical molecules into synthetic sub-fragments. Specifically, based on a series of predefined chemical synthesis rules, bonds in the molecule that conform to specific chemical reaction types are first identified. These rules cover bond connections involved in common organic chemical reactions. Then, bonds conforming to the rules are cut, decomposing the molecule into multiple fragments. Each fragment is treated as a node, and the split bonds as edges, thus constructing a functional group diagram. Another type is the motif-based decomposition method, which uses a motif-building algorithm. This algorithm first creates a central set of atoms, treating each atom as an initial substructure. It iteratively expands the atoms starting from non-carbon adjacent atoms to obtain motifs, which serve as nodes in the motif graph. For two motifs with shared atoms, the shared atoms are used as connecting edges. For two motifs without shared atoms, all atoms in one motif are traversed, and atoms with adjacent atoms in the other motif are selected. These atoms and their adjacent atoms in the other motif form substructure edges, thus connecting the two motifs and constructing a functional group graph. .
[0082] (3) Molecular diagram generation
[0083] Molecular diagram obtained from the previous step In the process, all aromatic structure subgraphs are extracted. If a molecular graph contains multiple aromatic subgraphs, a new node is added for each subgraph. This new node represents the π-bond interactions between atomic nodes in the subgraph, and the newly added node is connected to all other nodes in the subgraph to obtain the final atomic graph. .
[0084] (4) Construct a hierarchical diagram based on the relationship between drugs and functional group atoms.
[0085] Based on the metabolic information obtained in step (1), the functional group information of the molecule obtained in step (2), and the structural information of the molecule obtained in step (3), a hierarchical molecular map is constructed. The top layer is the metabolic layer, where each node represents a molecule. The middle layer is the functional group layer, a molecular graph with molecular subgraphs as nodes and edges connecting them. The bottom layer is the atomic layer, corresponding to molecular graphs at the atomic level. The three layers are connected by edges to represent the relationships between the layers. The top two layers connect the nodes of the metabolic layer with the subgraph nodes corresponding to their molecules; the bottom two layers construct their connections based on the inclusion relationships between the functional group layer and the atomic layer nodes.
[0086] 3. Chemical and topological characterization of molecules based on hierarchical graph learning
[0087] (1) Characteristics of atoms in aggregated molecules through message passing networks
[0088] First, a message-passing neural network (MPNN) is used for learning, which enhances the atomic information of the node by aggregating neighborhood information. For nodes within the atomic layer, the message-passing neural network is used for learning:
[0089]
[0090]
[0091] in For message passing functions, For vertex update functions, This represents the state of the t-th layer after message passing. Represents the t-th layer The hidden state of a node, where 'a' represents the atomic level. Atomic diagram In and nodes The set of adjacent nodes, Layer t The hidden state of a node, express , The edges between them.
[0092] (2) Learning the influence of each atom in the atomic layer on the functional group layer through the attention mechanism.
[0093] The purpose of this step is to learn the weights between the atomic layer and the functional group layer, which are used to measure the influence of each atomic node on the functional group. This step takes the information of the aggregated node atoms from step (1) as input to obtain the structural features of the drug subgraph. The formula is as follows:
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] in The activation function is tanh. This represents the atomic graph node after the t-th update. The hidden state, , Representative Hierarchical Diagram Nodes in the functional group diagram Adjacent atomic layer nodes, For the projection matrix, This represents the atomic features after projection at the t-th update. This represents the functional group feature after projection at the t-th update. For the atomic node at the t-th update Functional group graph nodes Similarity of influence This represents the influence weight of each atom at the t-th update; This represents a splicing operation;
[0100] (3) Characteristics of atoms in the aggregate substructure (functional group) through message passing network
[0101] In the functional group layer, a message passing network (MPNN) is used to aggregate the structural features of the connected subgraphs obtained in step (2) to enhance the representation of the nodes in the subgraph. For nodes within the layer, a message passing network will be used for learning:
[0102]
[0103]
[0104] in For message passing functions, For vertex update functions, Represents functional group graph nodes After the t-th message passing step is completed, the message embedding is calculated based on the features of its neighboring nodes and edge attributes. This represents the node after the t-th update. The hidden state; Represents sensual group diagram In and nodes Adjacent nodes, This indicates the time after the t-th update. The hidden state of a node, express , The edge between;
[0105] (4) Learning the influence of functional group sub-diagrams on metabolic layer molecules through attention mechanisms
[0106] By learning the weights between molecular nodes in the metabolic layer and molecules in the substructure (functional group) layer, the weights indicate the influence of the substructure on the molecule. Considering that one molecule corresponds to multiple functional group subgraphs during the splitting process, this step needs to complete two tasks: one is to aggregate functional group features under the same splitting rules, and the other is to fuse the features of functional groups from different splitting methods.
[0107] The influence of functional group features at different splits on metabolic layer molecules is learned through an attention mechanism, as shown in the following formula:
[0108]
[0109]
[0110]
[0111]
[0112]
[0113] in The activation function is tanh. This represents the metabolic graph after the t-th update. Nodes in Polymer functional group diagram The hidden state of the middle node; Representative Hierarchical Diagram Middle and metabolic layer nodes The set of adjacent functional group layer nodes, For the projection matrix, This represents the functional group feature after projection at the t-th update. This represents the aggregated functional group diagram after projection at the t-th update. Metabolic graph node features. For the facultative community graph node at the t-th update Metabolic map nodes Similarity of influence This represents the influence weight of each functional group at the t-th update. This represents a splicing operation;
[0114] The similarity weights between the metabolic subgraph node features generated by the BRICS splitting and motif construction algorithm are dynamically calculated using a cross-attention mechanism, thereby achieving adaptive weighted fusion of metabolic molecular features under the two splitting strategies. The formula is as follows:
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123] in and These represent the nodes of the metabolic layer molecular map obtained through two different splitting methods. Features These represent the learnable weight matrices; Represent and The dimension; This represents a fused feature formed by incorporating key splitting features into motif features using an attention method; This represents a fusion feature formed by integrating motif features into bond splitting features through an attention mechanism.
[0124] After obtaining two fusion characterizations, the overall features of the molecule are obtained through weighted averaging:
[0125]
[0126] in For learnable weight matrix, This indicates the node after merging the two splitting methods. Structural features.
[0127] 4. Toxicity prediction based on molecular structure characteristics and metabolites
[0128] (1) Predicting molecular toxicity through molecular characteristics
[0129]
[0130] This represents the output from the molecular metabolic layer. Considering that this scheme will predict different toxicity prediction tasks, its specific meaning will vary. In classification tasks, This represents the probability that a drug is toxic (if predicting the presence of toxicity); in regression tasks, This indicates the magnitude of the predicted toxicity value (such as IC50).
[0131] (2) Toxicity prediction based on molecular characteristics and the potential toxicity of metabolites
[0132] This step primarily focuses on establishing the potential relationships between molecules within the metabolic layer, with the aim of interpreting their impact on drug response. In this module, relationships between molecules are established by masking some molecules and then reconstructing them using other molecules.
[0133]
[0134] in This represents the toxicity probability or numerical value corresponding to node a, which points to node i in the metabolic graph. Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. This represents the set of nodes pointing to i in the metabolism graph. This indicates that the molecular toxicity of metabolites after polymerization should be taken into account.
[0135] The influence of products on reactants is inferred using a metabolic layer network, and the weights of the metabolic layer data are modified during each update. The weights of this layer reflect the influence of metabolites on the original molecule. Considering that this scheme will predict different toxicity tasks, a loss function is designed according to different toxicity prediction tasks. Toxicity prediction tasks mainly include two categories: one is toxicity label prediction, i.e., whether toxicity exists, which is essentially a classification task; the other is predicting toxicity values such as IC50, which is essentially a regression task. Furthermore, considering that this patent introduces drug features and features fused with metabolism for toxicity prediction, the total loss of this model includes two parts: the first part is the feature specific to the task type, and the second part considers the impact of introducing metabolites on the drug feature-based toxicity prediction. The loss function of this model is as follows:
[0136] Classification task loss:
[0137]
[0138] Where N represents the number of training samples, This indicates the true toxicity label. This indicates that the molecular toxicity of metabolites after polymerization should be taken into account. This indicates that molecular toxicity was not taken into account when considering metabolites. The goal is to make the toxicity prediction results when only considering the structural features of the molecule as close as possible to the molecular toxicity when considering the polymerization of metabolites, so that the model takes into account the impact of potential changes in structural features on toxicity when considering molecular structural features.
[0139] Losses from the return mission:
[0140]
[0141] Where N represents the number of training samples, Indicates the true toxicity label
[0142] After calculating the loss function, this approach learns by iterating through the parameters. Considering that message graph and attention graph-based methods may cause oversmoothing after too many iterations, this approach uses early stopping to train the model.
[0143] 5. Metabolic toxicity inference based on drug properties and metabolite properties
[0144] During the inference phase, for a newly input drug molecule, the metabolic layer in the molecule's hierarchical diagram is obtained according to the process described in step 2. For each molecular node in the metabolic layer (containing both the drug molecule and its metabolites), the hierarchical diagram of each molecular node is obtained, and the structural representation of the drug molecule is obtained according to the process in step 3. and structural characteristics of metabolites The predicted toxicity values based on drug molecular structure are obtained through the process in step 4-1. and predicted values based on metabolite structure In step 4-2, since the newly input drug molecule and its generated metabolites are not included in the pre-trained metabolic layer weights, a weight transfer mechanism needs to be designed: based on the chemical structural similarity between metabolites, the weight parameters of the existing metabolic pathways are transferred to the metabolic process of the new molecule. The specific process of weight transfer and the prediction results of the new drug molecule are described in the following formula:
[0145]
[0146]
[0147] in, This indicates a predicted value for drug toxicity. For the structural characteristics of the drug to be inferred, For the metabolic diagram and Structural characteristics of drugs that share the same metabolite a; Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. Indicates metabolites after migration Treatment of speculative drugs Toxicity impact weight.
Claims
1. A method for predicting drug toxicity based on metabolism guidance, characterized in that, Includes the following steps: S1. Construct training data, specifically: obtain the SMILES formula of drug molecules and the corresponding toxicity attributes from existing toxicology data to form training data, wherein the toxicity attributes include two categories: toxicity type and toxicity value. S2. Preprocess the training data, specifically as follows: a) Based on the SMILES formula of the drug molecule and predefined metabolic reaction rules, the SMILES formula is decomposed to obtain a metabolic graph; specifically: the reaction sites of the molecules represented by the SMILES formula are located according to the metabolic reaction rules, the SMILES formula of the initial metabolite is generated by breaking the corresponding chemical bonds, and the SMILES formula corresponding to each molecule is used as a metabolic node, and the metabolic reaction is used as a directed edge to construct a directed graph. The process involves continuously generating SMILES formulas for new products based on the initial metabolite formulas and constructing a directed graph until the SMILES formula of a new product can no longer match the corresponding metabolic rule subgraph, thus constructing a complete metabolic graph. This metabolic graph uses the SMILES formula of the parent molecule as the root node, the SMILES formulas of multi-level metabolites as child nodes, and metabolic transformation relationships as directed edges, which can completely map the network path from the original molecule to all possible metabolites, thereby constructing the metabolic graph. For each node in the constructed metabologram corresponding to the SMILES formula, the rdkit tool is used to convert it into a molecular diagram. This yields the graph structure corresponding to each metabolic node; b) The molecular map was analyzed using the bond-based BRICS method and the motif-based decomposition method. The function group diagrams were split to obtain two different scales. ; c) Molecular diagram obtained from the above In the process, all aromatic structure subgraphs are extracted. If a molecular graph contains multiple aromatic subgraphs, a new node is added for each subgraph. This new node represents the π-bond interactions between atomic nodes in the subgraph, and the newly added node is connected to all other nodes in the subgraph to obtain the final atomic graph. ; d) Based on metabolic maps , sensual group diagram and atomic diagram Constructing a drug metabolism hierarchy map The structure consists of a metabolism diagram at the top layer, with each node representing a molecule; a functional group diagram at the middle layer, with functional groups as nodes; and an atomic diagram at the bottom layer. The three layers are connected by edges to represent the relationships between the layers. The top two layers connect the nodes of the metabolism layer with the functional group nodes corresponding to their molecules, while the bottom two layers construct their connections based on the inclusion relationship between functional groups and atomic layer nodes. S3, Based on the constructed drug metabolism hierarchy map Studying the chemical and topological characterization of molecules, specifically including: Features of atoms within an atomic graph are aggregated using a message passing network: , , in For message passing functions, For vertex update functions, Represents atomic graph nodes After the t-th message passing step is completed, the message embedding is calculated based on the features of its neighboring nodes and edge attributes. This indicates the time after the t-th update. The hidden state of a node, where 'a' represents the atomic level. Atomic diagram In and nodes The set of adjacent nodes, This indicates the time after the t-th update. The hidden state of a node, express , The edge between; The influence of the atomic graph on functional clique layer nodes is learned through an attention mechanism, and features of functional clique nodes within the functional clique graph are aggregated using a message-passing network, including: The influence of each atom in the atomic diagram on the functional group diagram is learned through an attention mechanism, as follows: , , , , , in The activation function is tanh. This represents the atomic graph node after the t-th update. The hidden state, , Representative Hierarchical Diagram Nodes in the functional group diagram Adjacent atomic layer nodes, For the projection matrix, This represents the atomic features after projection at the t-th update. This represents the functional group feature after projection at the t-th update. For the atomic node at the t-th update Functional group graph nodes Similarity of influence This represents the influence weight of each atom at the t-th update; This represents a splicing operation; Characteristics of functional clique nodes within a functional clique graph are aggregated using a message-passing network: , , in For message passing functions, For vertex update functions, Represents functional group graph nodes After the t-th message passing step is completed, the message embedding is calculated based on the features of its neighboring nodes and edge attributes. This represents the node after the t-th update. The hidden state; Represents sensual group diagram In and nodes Adjacent nodes, This indicates the time after the t-th update. The hidden state of a node, express , The edge between; The influence of functional group graphs on metabolic layer molecules is learned through an attention mechanism. Molecular features of metabolic graphs under different functional group decomposition strategies are weighted and fused to obtain metabolic graph node features that fuse functional group graph node features, including: The influence of functional group features at different splits on metabolic layer molecules is learned through an attention mechanism, and the specific formula is as follows: , , , , , in The activation function is tanh. This represents the metabolic graph after the t-th update. Nodes in Polymer functional group diagram The hidden state of the middle node; Representative Hierarchical Diagram Middle and metabolic layer nodes The set of adjacent functional group layer nodes, For the projection matrix, This represents the functional group feature after projection at the t-th update. This represents the aggregated functional group graph after projection at the t-th update. Metabolic graph node features. For the facultative community graph nodes at the t-th update Metabolic map nodes Similarity of influence This represents the influence weight of each functional group at the t-th update. This represents a splicing operation; The similarity weights between the metabolic subgraph node features generated by BRICS splitting and motif construction algorithm are dynamically calculated using a cross-attention mechanism, thereby achieving adaptive weighted fusion of metabolic molecular features under the two splitting strategies: , , , , , , , , in and These represent the nodes of the metabolic layer molecular map obtained through two different splitting methods. Features These represent the learnable weight matrices; Represent and The dimension; This represents a fused feature formed by incorporating key splitting features into motif features using an attention method; This represents a fusion feature formed by integrating motif features into key splitting features through an attention mechanism; After obtaining two fusion characterizations, the overall features of the molecule are obtained through weighted averaging: , in For learnable weight matrix, This indicates the node after merging the two splitting methods. Structural features; S4. Toxicity prediction based on molecular structure characteristics and metabolites, including: Predicting molecular toxicity based on molecular structure characteristics: , This indicates the predicted molecular toxicity obtained by considering only molecular structure features. According to the toxicity attributes defined in S1, there are two categories: toxicity type and toxicity value. Therefore, in classification tasks... This indicates the probability that a drug is toxic; in regression tasks, This indicates the predicted toxicity level; Toxicity prediction based on molecular structural features and their metabolites: , in This represents the toxicity probability or numerical value corresponding to node a, which points to node i in the metabolic graph. Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. This represents the set of nodes pointing to i in the metabolic graph. This indicates consideration of the molecular toxicity of polymerized metabolites; Loss functions are defined as classification task loss and regression task loss: , , Where N represents the number of training samples, Indicates the true toxicity label. This indicates that the molecular toxicity of the polymerized metabolites should be taken into consideration. This indicates that molecular toxicity was not taken into account when considering metabolites; S5. Using the training data in S1, after preprocessing in S2, feature extraction is performed using the method in S3, and the loss value is calculated using the loss function in S4. Finally, the iteration stops when the set conditions are met, and the required network parameters are obtained. S6. To predict the toxicity of the target drug, specifically: For the target drug molecule, a metabolic map of the target drug is obtained according to method S2. Then, for molecules in the metabolic map that include the target drug molecule and all its metabolites, a hierarchical map of each molecular node is obtained. Finally, the structural representation of the drug molecule is extracted according to method S3. and structural characteristics of metabolites Finally, the network parameters obtained from training are used to obtain predicted values of drug toxicity based on drug molecular structure according to the S4 method. and predicted values based on metabolite structure Since the target drug molecule and its metabolites are not included in the metabolic layer weights obtained during training, the weight parameters of the existing metabolic pathways are transferred to the metabolic process of the new molecule. The specific process of weight transfer and the prediction results of the new drug molecule are as follows: , , in, This indicates a predicted value for drug toxicity. To determine the structural characteristics of the drug to be inferred, For the metabolic diagram and Structural characteristics of drugs that have the same metabolite a; Metabolic products on a metabolic diagram For drugs Toxicity affects weighting. Indicates metabolites after migration Treatment of speculative drugs Toxicity impact weight.
Citation Information
Patent Citations
Structure-activity relationship model for predicting acute oral toxicity of tetrazole compound in mice and construction method
CN117672401A