Metabolic Kinetics and Toxicity Prediction Method Based on Graph Representation for Multi-Task Learning

The graph representation-based MTL-ADMETox model addresses performance degradation and interpretability issues in ADME-Tox prediction by establishing task-specific molecular representations and selecting optimal auxiliary tasks, enhancing the accuracy and interpretability of ADME-Tox predictions.

CN116343930BActive Publication Date: 2025-07-15NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310037757.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-15
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

The prior art has the inability to establish shared representations in ADME-Tox prediction in the early stages of drug development, resulting in performance degradation and insufficient explanatory performance on ADME endpoints, ignoring the impact of task relationship modeling.

Method used

The multi-task learning model MTL-ADMETox is adopted to build an inter-task correlation network, and use directed graph state theory and maximum flow strategy to select the best auxiliary task combination, and combine molecular feature modules, gated modules and task predictor modules for training to realize multi-task joint learning.

Benefits of technology

It improves the accuracy and interpretability of ADME-Tox prediction, improves the performance of drug absorption, distribution, metabolism, excretion and toxicity prediction, and provides better drug optimization tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343930B_ABST
    Figure CN116343930B_ABST
Patent Text Reader

Abstract

The present invention discloses a metabolic kinetics and toxicity prediction method based on graph representation multi-task learning, and proposes a multi-task graph learning framework, namely MTL-ADMETox, through relevant auxiliary tasks based on effective gates, so as to screen effective positive auxiliary tasks to jointly train the target task and optimize the contribution of the auxiliary tasks. MTL-ADMETox includes a task-specific molecular feature module, a gate module centered on the main task, and a task predictor module. Using the relationship based on effective auxiliary tasks, a multi-task based graph neural network is designed to study the prediction methods based on absorption, classification, metabolism, excretion and toxicity, and explore the association rules between compound substructures and ADME, which can promote the development of candidate drug screening or drug design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of graph representation learning for assisting lead compound optimization, and particularly relates to a method for predicting pharmacokinetics and toxicity based on graph representation multi-task learning. Background Art

[0002] The discovery and development of drugs, especially a large amount of data from animal laboratory tests and experiments, take a long time and a large amount of cost. Unfavorable pharmacokinetic (PK) properties or high toxicity are the main reasons for the failure of candidate drugs in the clinical trial stage. Therefore, conducting absorption, distribution, metabolism, excretion, and toxicity (ADME-Tox) studies through computational methods, especially artificial intelligence (AI) technology, in the early stage of drug development helps reduce the cost of new drug development and improve the success rate.

[0003] In recent years, great progress has been made in predicting compound ADME-Tox based on computational methods. Generally speaking, most methods, especially machine learning and deep learning models, have been proven to be able to effectively analyze the current large amount of ADME-Tox data. Traditional machine learning (ML) algorithm models, such as multiple linear regression (MLR) and random forest (RF), are commonly used to construct ADME-Tox predictions. Although previous machine learning-based methods required a large amount of reliable expertise to design features, expertise in molecular structure and ADME-Tox data is often insufficient and subjective.

[0004] With the increase in ADME-Tox data and the development of deep learning (DL) technology, especially the development of graph neural networks (GNNs), the performance of ADME-Tox prediction has been greatly improved compared to traditional machine learning (ML). Generally speaking, these methods separately model ADME-Tox endpoint data. In this case, multi-task learning (MTL) may be a good solution to utilize the information learned from related ADME-Tox endpoints. The purpose of MTL is to improve the prediction accuracy by jointly learning multiple related tasks.

[0005] In recent years, due to the rapid development of graph representation learning algorithms and their successful applications in other fields, the research accumulation in ADME-Tox has also promoted the application prospect of deep learning in lead compound optimization. Structural data such as drugs can be automatically feature-extracted by graph neural networks. These structured deep learning models combined with multi-layer neural networks have been successfully applied in the field of drug design. However, despite the great efforts and remarkable achievements made by researchers in ADME-Tox prediction, there are still significant challenges in actual work, mainly manifested in the following aspects:

[0006] 1) When multi-task learning fails to establish a shared representation that can be generalized to all tasks, it may also lead to a significant performance degradation. In addition, existing MTL-based ADME-Tox methods ignore the impact of task relationship modeling.

[0007] 2) The lack of interpretability of ADME (absorption, distribution, metabolism, excretion), current methods mainly focus on the interpretation of toxicity (Tox) endpoints, while rarely paying attention to the interpretation of ADME endpoints. In fact, most ADME endpoints are usually related to the presence of certain compound substructures.

[0008] In view of this, it is necessary to design a new prediction method. Summary of the Invention

[0009] The object of the present invention is to solve the deficiencies of the prior art, and provides a metabolic kinetics and toxicity prediction method based on graph representation multi-task learning.

[0010] The concept of the present invention:

[0011] A model based on graph representation multi-task learning is proposed, namely MTL-ADMETox. By training single and paired tasks, an inter-task association network is established; then, the state theory of the directed graph and the maximum flow strategy are used to collect approximate auxiliary tasks for each task; finally, in order to train the main task and its auxiliary tasks, the above-mentioned graph representation multi-task learning model is adopted, including a task-specific molecular feature module, a gating module centered on the main task, and a task predictor module.

[0012] In view of the above invention concept, the technical solution provided by the present invention to achieve the invention object is:

[0013] A metabolic kinetics and toxicity prediction method based on graph representation multi-task learning, characterized by comprising the following steps:

[0014] 1) Construct the ADME-Tox prediction model MTL-ADMETox

[0015] The ADME-Tox prediction model MTL-ADMETox includes a molecular feature module, a gating module (for feature fusion), and a task predictor module from input to output;

[0016] The molecular feature module is used to obtain molecular features, including two layers of graph convolutional network layers GCN and a task-specific attention layer (the attention layer is used to learn the task-specific molecular representation of the compound to generate different molecular representations from the same atomic representation);

[0017] The task predictor module adopts a fully connected neural network layer;

[0018] 2) Collect sample data and train the ADME-Tox model constructed in step 1).

[0019] 2.1) Collect the structural information of drug molecules and their corresponding ADME-Tox type information, and construct a training dataset, a validation dataset, and a test dataset.

[0020] 2.2) Convert the SMILES (Simplified molecular input line entry specification) sequence information of drug molecules involved in the data obtained in step 2.1) into a compound graph to obtain compound structure data.

[0021] 2.3) Use the ADME-Tox type information of the drug molecules collected in step 2.1) and the compound structure data obtained in step 2.2) as inputs, and perform the following steps:

[0022] First, perform single-task and dual-task training in sequence, calculate the mutual influence between tasks, obtain the interaction relationship between tasks (reflecting), use the performance difference degree between two single-task and dual-task graph learning methods as a measure of the inter-task effect, select respective auxiliary tasks for each task, and establish an inter-task association network graph.

[0023] Second, use the state theory of directed graphs and the maximum flow strategy to select an effective combination of auxiliary tasks for each task.

[0024] These two steps are mainly to select the best combination of auxiliary tasks for the main task, using the performance difference degree between two single-task and dual-task graph learning methods as a measure of the inter-task effect; among them, the best task combination of the main task needs to satisfy the structural balance between endpoints and needs to be passed to the main task in the transitive triangle. After that, collect all the promoting tasks of the main task to strengthen the stability of the main task and improve the performance. In addition, putting all the promoting tasks together may not optimize the performance of the main task. Therefore, it is necessary to select the best combination of auxiliary tasks according to the training results.

[0025] Finally, use the information and data corresponding to the auxiliary tasks in each main task and its auxiliary task combination as inputs to obtain the molecular features of the main task and the auxiliary tasks through the molecular feature module.

[0026] 2.4) Through the gating module, fuse the molecular features of the auxiliary tasks with the molecular features of the main task respectively and add them to obtain the final molecular features of the main task.

[0027] 2.5) Use a fully connected neural network layer to predict and output a feature vector for the auxiliary task molecular features obtained in step 2.3) and the final molecular features of the main task obtained in 2.4).

[0028] 2.6) Use the cross-entropy loss function to calculate the loss between the output feature vector obtained in step 2.5) and the original label, and then update the trainable parameters in the ADME-Tox prediction model through negative feedback regulation. After multiple trainings, the final ADME-Tox prediction model is obtained.

[0029] 3) Use the trained ADME-Tox prediction model in step 2) to predict the ADME-Tox of drug molecules.

[0030] Furthermore, step 2.2) is specifically as follows:

[0031] Use the open-source chemical toolbox RDKit to convert the SMILES sequence into an interaction graph between atoms; the compound graph is represented as G = (V, E), where V is a set of N nodes and E is a set of edges.

[0032] Here, each node is a multi-dimensional binary feature vector, expressing the information in the atomic symbol, degree, charge, aromaticity, and the number of adjacent hydrogens in the structure.

[0033] Furthermore, in step 2.3), the training of both single-task and dual-task refers to the training through the molecular feature module and the fully connected neural network layer in the ADME-Tox prediction model MTL-ADMETox.

[0034] The interaction between tasks is calculated through the difference between the training results of single-task and dual-task.

[0035] Furthermore, in step 2.3),

[0036] The graph convolutional network layer GCN is designed for semi-supervised node classification, and its basic idea is to update the representation of nodes through information propagation between nodes; the hierarchical propagation rules of the multi-layer graph convolutional network layer GCN are as follows:

[0037]

[0038] Among them, is the adjacency matrix of an undirected graph with self-connections added, A ∈ R N×N is the adjacency matrix representing E; I N is the identity matrix, σ(·) is the activation function, and W (l) are a layer of specific trainable weight matrices; the hierarchical convolution operation can be approximated as follows:

[0039]

[0040] Among them, Q is the filter or feature map, B is the coarse-grained category, is the node output;

[0041] The task-specific attention layer is represented as follows:

[0042] a z = σ(W z ·h c + b z )

[0043] Among them, W z is the weight matrix, b z is the bias vector in the attention layer, learned during model training, σ is the activation function (i.e., sigmoid); h c is the shared atomic feature matrix; therefore, the final feature of compound c in task t z can be calculated in the following form:

[0044]

[0045] For a specific main task t k and its best auxiliary task t z , t w ..., through multiple attention layers, understand which substructures are crucial for improving the main task. Finally, obtain the molecular features of the main task h k and the auxiliary tasks h z , h w ....

[0046] Furthermore, in step 2.4),

[0047] The main-task-centered gating module, based on a single-layer feed-forward network (randomly initialized input), uses SofiMax as the activation function; obtain the weighted sum of the main task t k and the auxiliary tasks {t z , t w ...}, which is the output of the target task in the gating network;

[0048] Specifically, the gating network in t z expresses the output of task t k as follows:

[0049]

[0050] Among them, is the weighting function, calculating the main task t k and the auxiliary task t zWeight vector;

[0051]

[0052] where \(h\in\mathbb{R}\) d , \(d\) is the dimension of the input representation, \(d'\) is the dimension of the output representation after the DNN in the gate, is a matrix composed of vectors, including the main task \(t\) k and the auxiliary task \(t\) z ,

[0053] Therefore, the final representation of the main task \(t\) k is:

[0054]

[0055] Furthermore, in step 2.6), each task has a unique predictor to better learn the non - linear representation of a specific task. The fully - connected neural network layer has two layers, specifying losses for classification and regression respectively;

[0056]

[0057] FC_ k is the fully - connected layer for the main task, FC_ w and FC_ z are the fully - connected layers for the auxiliary tasks.

[0058] Furthermore, in step 2.7), the cross - entropy loss for the classification task and the mean squared error loss for the regression task are both used, defined as follows:

[0059]

[0060] In the formula, \(y\) c and are the true label and the predicted value of the compound \(c\) n with respect to the classification task \(t\) c ), \(y\) r is the true attribute value of \(c\) n with respect to the regression task \(t\) r ), is the corresponding predicted value, \(C\) is the number of classification tasks, \(M\) c is the number of compounds in the classification task; \(R\) is the number of regression tasks, \(M\) r is the number of compounds in the regression task; To alleviate the imbalance between positive and negative samples in the classification task, a weight \(p\) c is used in the loss function, representing the ratio of the number of negative samples to the number of positive samples.

[0061] At the same time, the present invention provides a computer-readable storage medium on which a computer program is stored, and the special feature of the computer program is that when the computer program is executed by a processor, the steps of the above method are implemented.

[0062] An electronic device is special in that it includes a processor and a computer-readable storage medium; the computer-readable storage medium stores a computer program, and the computer program executes the steps of the above method when executed by the processor.

[0063] The advantages of the present invention are:

[0064] 1. The present invention proposes a prediction model for metabolic kinetics and toxicity prediction methods based on graph representation multi-task learning, namely MTL-ADMETox, which solves these problems by constructing absorption, distribution, metabolism, excretion and toxicity representations. MTL-ADMETox consists of a molecular feature module for a specific task, a gating module centered on the main task, and a task predictor module (i.e., a fully connected neural network layer), and predictions for each task are performed through a fully connected layer neural network. This model can improve the performance of the model by mining the direct potential relationship between each endpoint task, while also making the drug ADMETox classification and regression tasks interpretable. In addition, MTL-ADMETox provides an attention-based key feature selection to more accurately predict the ADMETox type. The evaluation of MTL-ADMETox on a benchmark data set shows that MTL-ADMETox has good drug absorption, distribution, metabolism, excretion and toxicity prediction performance. The present invention can provide a computational prediction tool to promote the optimization of lead compounds.

[0065] 2. This paper proposes a multi-task graph learning framework, MTL-ADMETox, through relevant auxiliary tasks based on effective gates to screen effective positive auxiliary tasks to jointly train the target task and optimize the contribution of auxiliary tasks. By using the relationship based on effective auxiliary tasks, a multi-task graph neural network is designed to study the prediction methods based on absorption, classification, metabolism, excretion and toxicity, and explore the association between compound substructure and ADME, which can promote the development of candidate drug screening or drug design.

[0066] 3. The present invention utilizes a gated network to obtain the characteristics of the main task and the characteristics of the auxiliary task. Due to the addition of the characteristics of the auxiliary task, the performance of the model can also be greatly improved, thereby improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is the overall architecture of the method MTL-ADMETox proposed in the present invention;

[0068] Figure 2It is the relationship between the important substructures of the compounds of the present invention and ADME tasks. Detailed implementation manners

[0069] The content of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:

[0070] An embodiment of the ADME-Tox prediction model based on graph representation and multi-task learning proposed according to the present invention is specifically as follows:

[0071] This embodiment uses the ADME-Tox data set from the literature: This data set has 13 end-point data: oral bioavailability, P-gp inhibitor and substrate, Caco_2 permeability, blood-brain barrier, plasma protein binding rate, 2 enzyme inhibitors related to CYP450, clearance rate, LD 50 , respiratory toxicity, cardiac toxicity, solubility and lipophilicity. The data of drug molecules are divided into a training set, a validation set and a test set according to the ratio of 8:1:1.

[0072] For the SMILES sequence information of drug molecules in the data set, the RDKit algorithm is used to convert the SMILES sequence of drug molecules into a compound graph (i.e., an atomic interaction graph).

[0073] Using the collected type information of ADME-Tox corresponding to drug molecules and the converted atomic interaction graph data, the feature vectors of all drug atoms are obtained through stacked graph convolutional network layers.

[0074] Using multiple attention layers to obtain molecular features for specific tasks, and inputting the molecular features for specific tasks into a gated network for weighting.

[0075] A two-layer fully connected layer neural network is used to standardize the feature vectors of each task of the obtained drug molecules.

[0076] The loss function is calculated using the feature vectors of drug molecules and their original labels, and the trainable parameters in the model (such as the weights of the two-layer neural network) are trained according to the loss residuals through negative feedback adjustment.

[0077] After training is completed, a classification and regression model for compound molecule ADMETox, that is, a prediction model, is obtained.

[0078] During this period, an effective combination of auxiliary tasks needs to be selected for each main task. The molecular feature module and the task predictor module are used for single-task and dual-task training, the mutual influence between tasks is calculated to obtain the interaction relationship between tasks, and the performance difference degree between two single-task and dual-task graph learning methods is used as a measure of the effect between tasks. An auxiliary task is selected for each task, and a network diagram of the association between tasks is established. After that, the state theory of the directed graph and the maximum flow strategy are used to select an effective combination of auxiliary tasks for each task.

[0079] To evaluate the prediction performance, the present invention selects ROC-AUC and concordance index (R 2 ) as the basic evaluation indicators to measure the classification task and the regression task respectively. The higher the values of these indicators, the better the performance.

[0080] The trained model is tested using the test set data, and the test results are shown in Table 1.

[0081] Table 1 Performance display of MTL-ADMETox on the ADMETox dataset

[0082]

[0083] Compounds with blood-brain barrier, clearance rate, and cardiotoxicity labels are selected, and the weights of different chemical bonds of the compounds are extracted through the attention layer as Figure 2 shown.

[0084] In summary, the present invention can be used for the prediction of drug absorption, distribution, metabolism, excretion, and toxicity. The well-known implementation methods and characteristic common sense in the above-described solutions are not described in detail herein. It should be pointed out that those skilled in the art can make several improvements without departing from the present invention, and these should also be regarded as the protection scope of the present invention, which will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application should be subject to the content of the claims, and the specific implementation manners in the specification are used to explain the content of the claims.

Claims

1. A method for predicting metabolic kinetics and toxicity based on graph representation multi-task learning, characterized in that Including the following steps: 1) Construct an ADME-Tox prediction model MTL-ADMETox The ADME-Tox prediction model MTL-ADMETox includes a molecular feature module, a gating module, and a task predictor module from input to output; The molecular feature module is used to obtain molecular features, including two graph convolutional network layers GCN and a task-specific attention layer; The task predictor module uses a fully connected neural network layer; 2) Collect sample data and train the ADME-Tox model constructed in step 1) 2.1) Collect the structural information of drug molecules and their corresponding ADME-Tox type information, and construct a training data set, a validation data set, and a test data set; 2.2) Convert the SMILES sequence information of the drug molecules involved in the data obtained in step 2.1) into a compound graph to obtain compound structure data; 2.3) Use the ADME-Tox type information of the drug molecules collected in step 2.1) and the compound structure data obtained in step 2.2) as inputs, and perform the following steps: First, perform single-task and dual-task training in sequence, calculate the mutual influence between tasks, obtain the interaction relationship between tasks, use the performance difference degree between two single-task and dual-task graph learning methods as a measure of the task effect, select respective auxiliary tasks for each task, and establish an inter-task association network graph; Second, use the state theory of the directed graph and the maximum flow strategy to select an effective auxiliary task combination for each task; Finally, use the information and data corresponding to the auxiliary tasks in each task and its auxiliary task combination as inputs to obtain the molecular features of the main task and the auxiliary tasks through the molecular feature module; 2.4) Through the gating module, fuse and add the molecular features of the auxiliary tasks with the molecular features of the main task respectively to obtain the final molecular features of the main task; 2.5) Use a fully connected layer neural network layer to predict and output a feature vector for the auxiliary task molecular features obtained in step 2.3) and the final molecular features of the main task obtained in 2.4); 2.6) Use the cross-entropy loss function to calculate the loss between the feature vector output in step 2.5) and the original label, and then update the trainable parameters in the ADME-Tox prediction model through negative feedback regulation. After multiple trainings, obtain the final ADME-Tox prediction model; 3) Use the ADME-Tox prediction model trained in step 2) to predict the ADME-Tox of drug molecules.

2. The metabolic kinetics and toxicity prediction method based on graph representation for multi-task learning according to claim 1, characterized in that Step 2.2) is specifically: Use the open-source chemistry toolbox RDKit to convert the SMILES sequence into an interaction graph between atoms; the compound graph is represented as G=(V, E), where V is a set of N nodes and E is a set of edges; Here, each node is a multi-dimensional binary feature vector, expressing the information in the atomic symbol, degree, charge, aromaticity, and the number of adjacent hydrogens structure.

3. According to the method for predicting metabolism kinetics and toxicity based on graph representation multi-task learning according to claim 1, wherein: In step 2.3), the training of both single-task and dual-task refers to the training conducted through the molecular feature module and the fully connected neural network layer in the ADME-Tox prediction model MTL-ADMETox; The interaction between tasks is calculated by the difference between the training results of single-task and dual-task.

4. The metabolic kinetics and toxicity prediction method based on graph representation multi-task learning according to claim 3, wherein: In step 2.3), the graph convolutional network layer GCN is designed for semi-supervised node classification, and it updates the representation of nodes through information propagation between nodes; the hierarchical propagation rule of the multi-layer graph convolutional network layer GCN is as follows: Among them, is the adjacency matrix of an undirected graph with self-connections added, \(A\in\mathbb{R}\) N×N is the adjacency matrix representing \(E\); \(I\) N is the identity matrix, \(\sigma(\cdot)\) is the activation function, and \(W\) (l) is a specific layer of trainable weight matrix; the hierarchical convolution operation can be approximated as follows: Among them, Q is a filter or feature map, B is a coarse-grained category, is the node output; The task-specific attention layer is represented as follows: a z = σ(W z ·h c + b z ) Among them, W z is the weight matrix, b z is the bias vector in the attention layer, learned during model training, and σ is the activation function; h c is the shared atomic feature matrix; thus, the final feature of compound c in task t z can be calculated in the following form: For a specific main task t k and its optimal auxiliary task t z ,t w ... After passing through multiple attention layers, finally, the molecular features of the main task h k and the auxiliary task h z ,h w ... are obtained.

5. The metabolic kinetics and toxicity prediction method based on graph representation multi-task learning according to claim 4, wherein: In step 2.4), the main-task centered gating module is based on a single-layer feedforward network and uses SoftMax as the activation function; the main task t is obtained k and the weighted sum of the auxiliary tasks {t z , t w ...}, which is the output of the target task in the gating network; Specifically, t z The gating network in k outputs for task t are as follows: Among them, is a weighting function that calculates the weight vectors of the main task t k and the auxiliary task t z through linear transformation and the SoftMax layer; where h ∈ R d , d is the dimension of the input representation, d’ is the dimension of the output representation after the DNN in the gate, is a matrix composed of vectors, including the main task t k and the auxiliary task t z , Therefore, the main task t k The final representation is: 。 6. The metabolic kinetics and toxicity prediction method based on graph representation multi-task learning according to claim 5, wherein: In step 2.6), the fully connected neural network layer has two layers, which specify losses for classification and regression respectively; FC_ k The fully connected layer for the main task, FC_ w and FC_ z are the fully connected layers for the auxiliary tasks.

7. The metabolic kinetics and toxicity prediction method based on graph representation multi-task learning according to claim 6, wherein: The loss function adopted in step 2.7) is: where y c and are the true label and predicted value of compound c n with respect to classification task t c ), y r is the true attribute value of c n with respect to regression task t r ), is the corresponding predicted value, C is the number of classification tasks, M c is the number of compounds in the classification task; R is the number of regression tasks, M r is the number of compounds in the regression task; to alleviate the imbalance between positive and negative samples in the classification task, a weight p c is used in the loss function, representing the ratio of the number of negative samples to the number of positive samples.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: It includes a processor and a computer-readable storage medium; A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it executes the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Small molecule representation learning method based on Transform and enhanced interactive MPNN neural network

    CN113299354A

  • Metabolic pathway prediction method based on label correlation and graph representation learning

    CN114927173A