Pharmacokinetics and Toxicity Prediction Method Based on Coarse-Grained and Fine-Grained Classification
By adopting the MCF-PT model based on coarse and fine particle classification in pharmacokinetics and toxicity prediction, the problems of insufficient consideration of multi-task learning characteristics and insufficient interpretation of functional groups in the prior art are solved, and better predictive performance and interpretability are achieved.
Patent Information
- Application Number
- CN202211698293.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-28
AI Technical Summary
In the prior art, in pharmacokinetics and toxicity prediction, there are problems such as insufficient consideration of common and unique features of multi-task learning endpoint data, as well as insufficient interpretation of functional groups.
A pharmacokinetic and toxicity prediction method based on coarse and fine-grained classification is proposed. The MCF-PT model is proposed. By constructing a two-layer soft parameter sharing multi-task framework, coarse and fine-grained task-specific modules are designed to capture the dependencies between coarse, fine-grained tasks and fine-grained tasks.
Revealing the potential mechanism of multi-particle drug-producing properties linkage of pharmacokinetic and toxic properties is achieved, improving predictive performance, and providing interpretability in drug design.
Smart Images

Figure CN116206704B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence-assisted drug research and development, and particularly relates to a method for predicting pharmacokinetics and toxicity based on coarse-grained and fine-grained classification. Background Art
[0002] Pharmacokinetic problems and safety issues are the reasons for the failure of most candidate drugs. Compared with the time-consuming and costly traditional in vivo and in vitro evaluation methods, artificial intelligence provides an efficient and low-cost method for predicting pharmacokinetics and safety. Although from the perspective of artificial intelligence, the prediction of molecular metabolic kinetic properties is similar in form to the prediction of molecular properties, in addition to being affected by molecular properties, its metabolic kinetic properties are largely determined by the complex metabolic system in vivo.
[0003] In recent years, great progress has been made in predicting the pharmacokinetics and toxicity of small molecule compounds based on artificial intelligence. Generally speaking, most methods, especially machine learning and deep learning models, have been proven to be able to effectively analyze the current large amount of pharmacokinetic and toxicity data and predict new compounds. Early methods used molecular fingerprints and shallow classifiers to conduct evaluations and predictions. The admetSAR model applied MACCS molecular fingerprints and support vector machines to predict 27 metabolic kinetic properties. Because molecular fingerprints have good interpretability, such methods have always had good competitiveness. Recently, the FP-ADMET model used 20 kinds of molecular fingerprints and random forests to predict 76 properties. However, these methods can only build prediction models for each property separately and cannot simultaneously achieve information sharing for multiple property prediction tasks. With the development of deep learning, GNN (Graph Neural Network), especially GCN (Graph Convolutional Network Layer), has gradually replaced molecular fingerprints as a new feature representation method, and DNN (Deep Neural Network Algorithm) has gradually replaced shallow classifiers such as random forests. In recent years, people have tried to use multi-task learning to capture the potential dependencies between metabolic kinetic properties. In 2020, Feinberg et al. used GCN with gated recurrent units (GRUs) to extract task-specific molecular feature representations, achieving multi-task pharmacokinetic and toxicity prediction. In 2021, Xiong et al. used relational GCN to extract molecular structure features, designed task-specific attention layers to capture task-specific features, and constructed a prediction model based on a multi-task graph attention network. These multi-task graph representation models combined with multi-layer neural networks have been successfully applied in the field of drug design. However, despite the great efforts made by researchers in predicting the pharmacokinetics and toxicity of small molecule drugs and the remarkable achievements obtained, there are still quite a few challenges in actual work, which are mainly manifested in the following aspects:
[0004] 1) Insufficient consideration is given to the common features and specific features of the end - point data in multi - task learning. In current methods, task - specific features are not considered.
[0005] 2) Insufficient interpretability of functional groups. The relationship between the functional groups of compounds and pharmacokinetics and toxicity is lacking, and it is impossible to explain why a drug belongs to a certain pharmacokinetic or toxicity end - point based on the functional groups of the drug itself.
[0006] In view of this, it is necessary to design a new prediction method. Summary of the Invention
[0007] The purpose of the present invention is to solve the deficiencies existing in the prior art, and provide a pharmacokinetic and toxicity prediction method based on coarse - grained and fine - grained classification.
[0008] Conception of the present invention:
[0009] The present invention proposes a pharmacokinetic and toxicity prediction model based on coarse - grained and fine - grained classification, namely MCF - PT. A PT model is constructed according to the biological connotations of the metabolic kinetic properties (P) and toxicity (T) of small molecules, realizing task classification at the coarse - grained (C) and fine - grained (F) levels. A two - layer soft - parameter (parameters involved in the coarse - grained embedding module and the fine - grained embedding module) sharing multi - task (M) framework is constructed, and a coarse - grained task - specific module and a fine - grained task - specific module are designed to capture the dependencies between coarse - grained and fine - grained tasks and between fine - grained tasks. A pharmacokinetic and toxicity prediction method based on coarse - grained and fine - grained classification is studied based on the multi - task learning framework.
[0010] The MCF - PT model is divided into two granularity levels: coarse - grained and fine - grained. The coarse - grained level focuses on the property classification related to different tissues and organs (such as absorption, distribution, metabolism, excretion, toxicity), and the fine - grained level focuses on the subtle differences between similar categories within the same organ (such as type I and type II liver toxicity). Considering that a small molecule can cause multiple PT property changes, a two - level soft - parameter sharing multi - task framework is adopted as the overall architecture of the model.
[0011] In view of the above invention conception, the technical solution provided by the present invention to achieve the invention purpose is:
[0012] A pharmacokinetic and toxicity prediction method based on coarse - grained and fine - grained classification, characterized in that it includes the following steps:
[0013] 1) Construct a pharmacokinetic and toxicity prediction model MCF - PT based on coarse - grained and fine - grained classification
[0014] The pharmacokinetic and toxicity prediction model MCF - PT includes a coarse - grained embedding module and a fine - grained embedding module, and the prediction of pharmacokinetics and toxicity is output from the coarse - grained embedding module to the fine - grained embedding module;
[0015] The coarse-grained embedding module includes multiple graph convolutional network layers GCN (i.e., graph convolutional network layers set in sequence) and a gating unit; among them, the GCN is used to extract and distinguish the unique representation features of each coarse-grained task and the shared representation features among the coarse-grained tasks; the gating unit is used to integrate the shared representation features into the specific unique representation features of each one.
[0016] The fine-grained embedding module includes multiple graph attention network layers, a gating unit, and a fully connected layer neural network layer set in sequence;
[0017] 2) Collect sample data and train the pharmacokinetics and toxicity prediction models constructed in step 1)
[0018] 2.1) Collect the structural information of drug molecules and their corresponding pharmacokinetic and toxicity information, and construct a training data set, a validation data set, and a test data set;
[0019] 2.2) Convert the SMILES (Simplified molecular input line entry specification) sequence information of compound molecules involved in each data obtained in step 2.1) into a compound graph to obtain compound structure data;
[0020] 2.3) Use the compound structure data obtained in step 2.2) to extract the unique representation features of each coarse-grained task (each coarse-grained task is respectively: absorption, distribution, metabolism, excretion, toxicity) and the shared representation features among the coarse-grained tasks through the multiple graph convolutional network layers GCN in the coarse-grained embedding module; subsequently, through the gating unit in the coarse-grained embedding module, the shared representation features among the coarse-grained tasks are respectively fused with the unique representation features of each coarse-grained task to obtain the final representation features y1, y2, y3, y4, y5 of each coarse-grained task, and use them as the input of the fine-grained embedding module;
[0021] 2.4) Obtain the specific representation features of each fine-grained task and the shared representation features among the fine-grained tasks under each coarse-grained task through the multiple graph attention network layers in the fine-grained embedding module; subsequently, through the gating unit in the fine-grained embedding module, the shared representation features among the fine-grained tasks under each coarse-grained task are respectively fused with the specific representation features of each fine-grained task to obtain the weighted representation features of each fine-grained task;
[0022] 2.5) Use the weighted representation features of each fine-grained task obtained in step 2.4) as the output feature vectors of the fine-grained tasks through the fully connected layer neural network layer for prediction in the fine-grained embedding module;
[0023] 2.6) Calculate the loss between the output feature vectors obtained in step 2.5) and the original labels collected in step 2.1) using the cross-entropy loss function, and update the trainable parameters in the pharmacokinetics and toxicity prediction model (i.e., adjust the parameters in the model) through negative feedback regulation according to the loss residuals. After multiple trainings, the final pharmacokinetics and toxicity prediction model is obtained;
[0024] 3) Use the trained pharmacokinetics and toxicity prediction model in step 2) to predict the pharmacokinetics and toxicity of drug molecules.
[0025] Furthermore, in step 2.2), the open-source chemistry toolbox RDKit is used to convert the SMILES sequence into an interaction graph between atoms (i.e., compound graph); the compound graph is represented as G=(V, E), where V is a set of N nodes and E is a set of edges;
[0026] Here, each node is a multi-dimensional binary feature vector, expressing the information in the atomic symbol, degree, charge, aromaticity, and the number of adjacent hydrogens structure.
[0027] Furthermore, in step 2.3), a two-layer graph convolutional network layer GCN is used for extraction;
[0028] Among them, the graph convolutional network layer GCN adopts a semi-supervised node classification design, and its basic idea is to update the representation of nodes through information propagation between nodes; the hierarchical propagation rules of the multi-layer graph convolutional network layer GCN are as follows:
[0029]
[0030] Among them, is the adjacency matrix of the undirected graph with self-connections added, A∈R N×N is the adjacency matrix representing E, I N is the identity matrix, σ(·) is the activation function, and W (l) are a layer of specific trainable weight matrices; the hierarchical convolution operation can be approximated as follows:
[0031]
[0032] Among them, Q is the filter or feature map, B is the coarse-grained category, is the node output;
[0033] The unique representation feature F i extracted by the graph convolutional network layer GCN for the coarse-grained task T i (x) and the shared representation feature F s (x) among the coarse-grained tasks, and the gating unit Gi (x) determines the unique weights W of the two types of feature representations for each coarse-grained task i and the shared weight w s ;
[0034] The final representation features of each coarse-grained task (absorption, distribution, metabolism, excretion, and toxicity) after being weighted by the gating unit are used as the input to the fine-grained embedding module
[0035] Furthermore, in step 2.4), two-layer graph attention network layers H i and h i,j are used to obtain the specific representation features f i of each fine-grained task T i,j under the coarse-grained task T i,j and the shared representation features f i,s between fine-grained tasks;
[0036] The gating units {G i,s , G i,j} are used to fuse the shared representation features between the fine-grained tasks under each coarse-grained task T i with the specific representation features of each fine-grained task T i,j respectively (i.e., capture the dependencies between the coarse-grained task and the fine-grained tasks and between the fine-grained tasks), where G i,s +∑ j=1 G i,j = 1.
[0037] Furthermore, in step 2.5), the weighted representation features of each fine-grained task obtained in step 2.4) pass through the fully connected neural network layer h for prediction in the fine-grained embedding module as the output y i,j = h i,j (G i,j , f i,j ).
[0038] Furthermore, the loss function adopted in step 2.6) is:
[0039]
[0040] where y c and are the true label and the predicted value of the compound c n with respect to the classification task t c ), y r is the true attribute value of c n with respect to the regression task t r ), is the corresponding predicted value, C is the number of classification tasks, Mc is the number of compounds in the classification task; R is the number of regression tasks, M r is the number of compounds in the regression task; in order to alleviate the imbalance of positive and negative samples in the classification task, this application uses a weight p in the loss function c , which represents the ratio of the number of negative samples to the number of positive samples.
[0041] In addition, the present invention further provides a computer-readable storage medium on which a computer program is stored, and the special feature of the computer program is that when the computer program is executed by a processor, the steps of the above method are implemented.
[0042] An electronic device is special in that it includes a processor and a computer-readable storage medium; the computer-readable storage medium stores a computer program, and the computer program executes the steps of the above method when executed by the processor.
[0043] The advantages of the present invention are:
[0044] 1. The present invention proposes a pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification, namely MCF-PT, which clarifies the shared features and exclusive features of small molecule PT properties and reveals the potential mechanism of multi-granularity drug properties linkage. MCF-PT consists of two granularity levels, coarse-grained and fine-grained. In the coarse-grained embedding module of the first layer, each coarse-grained task-specific representation feature and the shared representation feature between coarse-grained tasks are generated, which are weighted by the gating unit. The weighted feature representation is used as the input of the second-layer fine-grained embedding module, and finally each endpoint task is predicted through the fully connected neural network layer; these problems are solved by constructing a multi-granularity representation drug model. MCF-PT consists of two granularity levels, coarse-grained and fine-grained. The output of the coarse-grained module embedding is used as the input of the fine-grained module, and finally the prediction of each key data is performed through the fully connected neural network layer. This model can mine the hidden related features between PT data to improve the performance of the model, and also make the small molecule pharmacokinetics and toxicity prediction interpretable.
[0045] 2. The MCF-PT in the present invention provides an attention-based key feature selection to more accurately predict pharmacokinetics and toxicity. The evaluation of MCF-PT on the benchmark dataset shows that MCF-PT has good pharmacokinetic and toxicity prediction performance and can provide a computational prediction tool to promote the discovery of candidate drugs.
[0046] 3. The present invention utilizes multi-task learning and combines task-specific features and shared feature methods to design a graph neural network based on fine-grained and coarse-grained drug formation, study small molecule pharmacokinetics and toxicity prediction methods based on it, and explore the correlation laws between compound functional groups and their various pharmacokinetic and toxicity endpoints. Since it can learn and use task-specific features and shared features between tasks simultaneously, the performance of the model can also be improved well.
[0047] 4. In the MCF-PT model of the present invention, the main focused part at the coarse-grained level can be selected by analyzing the weights of the gating unit; at the fine-grained level, the focused sub-structure of its own task can be selected by analyzing the weights of the graph attention network, thereby enhancing the interpretability of the fine-grained task. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the overall architecture of the MCF-PT method proposed by the present invention;
[0049] Figure 2 is the relationship between compound functional groups, pharmacokinetics and toxicity endpoints of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0050] The following further describes the content of the present invention in detail with reference to the drawings and specific embodiments:
[0051] An embodiment of the pharmacokinetics and toxicity prediction method based on fine-grained and coarse-grained classification proposed by the present invention is specifically as follows:
[0052] This embodiment uses a pharmacokinetics and toxicity (PT) dataset from the literature: this dataset belongs to 26 endpoints: human intestinal absorption (HIA), oral availability (OB), P-gp inhibitor, Caco_2 permeability, blood-brain barrier (BBB), 2 enzyme inhibitors related to CYP450, carcinogenicity, Ames, respiratory system toxicity, eye toxicity, liver toxicity, ear toxicity, heart toxicity, 11 endpoint data of Tox21 and solubility data. The data of drug molecules are divided into a training set, a validation set and a test set according to the ratio of 8:1:1.
[0053] For the SMILES sequence information of drug molecules in the dataset, the RDKit algorithm is used to convert the SMILES sequence of drug molecules into a compound graph (i.e., an atomic interaction graph).
[0054] In the coarse-grained embedding module, the feature vectors of all drug molecules are obtained through two-layer graph convolutional network layers and a gating unit using the converted atomic interaction graph data, and are used as the input of the second layer.
[0055] In the fine-grained embedding module, the output of the coarse-grained embedding module is used to extract the data (prediction results) of each end point through two layers of graph attention network layers, a gated unit, and a fully connected neural network layer.
[0056] The loss function is calculated using the feature vector of the drug molecule and its original label, and the trainable parameters in the pharmacokinetics and toxicity prediction model are trained through negative feedback regulation based on the loss residual.
[0057] After training is completed, a compound molecule pharmacokinetics and toxicity prediction model, that is, a prediction model, is obtained.
[0058] To evaluate the prediction performance of the model, the present invention selects ROC-AUC and concordance index (R 2 ) as the basic evaluation indicators to measure the classification task and the regression task respectively. The higher the values of these indicators, the better the performance.
[0059] The trained model is tested using the test set data, and the test results are shown in Table 1.
[0060] Table 1 Performance display of pharmacokinetics and toxicity prediction of MCF-PT on the PT dataset
[0061]
[0062]
[0063] It can be seen from Table 1 that the method of the present invention has improved in the prediction performance of each indicator compared with the current best model, the MGA method.
[0064] Compounds with oral bioavailability, hepatotoxicity, and solubility labels are selected, and the weights of different chemical bonds of the compounds are extracted through the graph attention network layer as Figure 2 shown.
[0065] In summary, the present invention can be used for the prediction of pharmacokinetics and toxicity. The well-known implementation methods and common knowledge of the features in the above-mentioned solutions are not described in detail here. It should be pointed out that for those skilled in the art, several improvements can be made without departing from the present invention, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be subject to the content of the claims, and the specific implementation manners and the like in the specification are used to explain the content of the claims.
Claims
1. A method for predicting pharmacokinetics and toxicity based on coarse-grained and fine-grained classification, characterized in that Including the following steps: 1) Construct a pharmacokinetics and toxicity prediction model MCF-PT based on coarse-grained and fine-grained classification The pharmacokinetics and toxicity prediction model MCF-PT includes a coarse-grained embedding module and a fine-grained embedding module, and outputs from the coarse-grained embedding module to the fine-grained embedding module for the prediction of pharmacokinetics and toxicity; The coarse-grained embedding module includes multiple graph convolutional network layers GCN and a gated unit; The fine-grained embedding module includes multiple graph attention network layers, a gated unit, and a fully connected layer neural network layer arranged in sequence; 2) Collect sample data and train the pharmacokinetics and toxicity prediction model constructed in step 1 2.1) Collect the structural information of drug molecules and their corresponding pharmacokinetic and toxicity information, and construct a training data set, a validation data set, and a test data set; 2.2) Convert the SMILES sequence information of compound molecules involved in each data obtained in step 2.1) into a compound graph to obtain compound structure data; 2.3) Use the compound structure data obtained in step 2.2) to extract the unique representation features of each coarse-grained task and the shared representation features between coarse-grained tasks through the multiple graph convolutional network layers GCN in the coarse-grained embedding module; Subsequently, the shared representation features between coarse-grained tasks are fused with the unique representation features of each coarse-grained task respectively through the gated unit in the coarse-grained embedding module to obtain the final representation features y1, y2, y3, y4, y5 of each coarse-grained task, and use them as the input of the fine-grained embedding module; 2.4) Obtain the specific representation features of each fine-grained task and the shared representation features between fine-grained tasks under each coarse-grained task through the multiple graph attention network layers in the fine-grained embedding module; Subsequently, the shared representation features between fine-grained tasks under each coarse-grained task are fused with the specific representation features of each fine-grained task respectively through the gated unit in the fine-grained embedding module to obtain the weighted representation features of each fine-grained task; 2.5) Use the weighted representation features of each fine-grained task obtained in step 2.4) as the output feature vector of the fine-grained task through the fully connected layer neural network layer for prediction in the fine-grained embedding module; 2.6) Use the cross-entropy loss function to calculate the loss between the output feature vector obtained in step 2.5) and the original label, and then update the trainable parameters in the pharmacokinetics and toxicity prediction model through negative feedback adjustment. After multiple trainings, the final pharmacokinetics and toxicity prediction model is obtained; 3) Use the pharmacokinetics and toxicity prediction model trained in step 2) to predict the pharmacokinetics and toxicity of drug molecules.
2. According to the pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification described in claim 1, wherein: In step 2.2), use the open-source chemical toolbox RDKit to convert the SMILES sequence into an interaction graph between atoms; the compound graph is represented as G=(V, E), where V is a set of N nodes and E is a set of edges; Here, each node is a multi-dimensional binary feature vector, expressing the information in the atomic symbol, degree, charge, aromaticity, and the number of adjacent hydrogens structure.
3. The pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification according to claim 2, characterized in that: In step 2.3), two graph convolutional network layers GCN are used for extraction; Among them, the graph convolutional network layer GCN adopts a semi-supervised node classification design, which updates the representation of nodes through information propagation between nodes; the hierarchical propagation rules of the multi-layer graph convolutional network layer GCN are as follows: Among them, is the adjacency matrix of an undirected graph with self-connections added, \(A\in\mathbb{R}\) N×N is the adjacency matrix representing \(E\), \(I\) N is the identity matrix, \(\sigma(\cdot)\) is the activation function, and \(W\) (l) is a specific layer of trainable weight matrices; the hierarchical convolution operation is approximated as follows: Among them, Q is a filter or feature map, and B is a coarse-grained category, is the node output; The coarse-grained task T extracted by the graph convolutional network layer GCN i The unique representation feature F i (x) and the shared representation feature F s (x), the gating unit G i (x) determines the unique weights W i and the shared weights W s ; The final representation features of each coarse-grained task after being weighted by the gating unit are used as the input of the fine-grained embedding module 4. The pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification according to claim 3, characterized in that: In step 2.4), two graph attention network layers H i and h i,j are used to obtain the specific representation features f i of each fine-grained task T i,j under the coarse-grained task T i,j and the shared representation features f i,s between fine-grained tasks; Adopt a gating unit {G i,s , G i,j} to fuse the shared representation features among the fine-grained tasks under each coarse-grained task T i with the specific representation features of each fine-grained task T i,j respectively, where G i,s + Z j=1 G i,j = 1.
5. The pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification according to claim 4, characterized in that: In step 2.5), the weighted representation features of each fine-grained task obtained in step 2.4) are passed through the fully connected neural network layer h for prediction in the fine-grained embedding module as the output y of the fine-grained task i,j = h i,j (G i,j , f i,j ).
6. The pharmacokinetics and toxicity prediction method based on coarse-grained and fine-grained classification according to claim 1, characterized in that The loss function adopted in step 2.6) is: where y c and are the true label and predicted value of compound c n with respect to classification task t c ), y r is the true attribute value of c n with respect to regression task t r ), is the corresponding predicted value, C is the number of classification tasks, M c is the number of compounds in the classification task; R is the number of regression tasks, M r is the number of compounds in the regression task; to alleviate the imbalance between positive and negative samples in the classification task, a weight p c is used in the loss function, representing the ratio of the number of negative samples to the number of positive samples.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of any one of the methods 1 to 6.
8. An electronic device, characterized in that: Comprising a processor and a computer-readable storage medium; A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it executes the steps of any one of claims 1 to 6.
Citation Information
Patent Citations
Social network influence prediction method and device based on graph neural network
CN113792937A
Metabolic pathway prediction method based on label correlation and graph representation learning
CN114927173A