Training methods and models for predicting inhibitors of drug-metabolizing enzymes
By selecting important molecular descriptors and using multiple binding energies, the method addresses the inefficiencies of existing enzyme inhibition prediction methods, enhancing accuracy and speed in identifying inhibitors for CYP, SULT, and UGT enzymes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INST NAT DE LA SANTE & DE LA RECHERCHE MEDICALE (INSERM)
- Filing Date
- 2021-07-23
- Publication Date
- 2026-04-30
AI Technical Summary
Existing methods for predicting drug-metabolizing enzyme inhibition, such as those involving cytochrome P450 (CYP), sulfotransferases (SULT), and UDP-glucuronosyltransferases (UGT), face challenges including high computational complexity, slow training and prediction times, and limited understanding of inhibitor mechanisms due to the use of numerous descriptors and single conformation binding energies.
A method is developed to train classification models for predicting enzyme inhibitors by selecting a subset of molecular descriptors based on their relative importance, using physicochemical parameters and multiple binding energies across enzyme conformations, and employing random forest and support vector machine models for improved accuracy and speed.
The method reduces computational time and enhances understanding of inhibitor properties, achieving higher prediction accuracy for CYP, SULT, and UGT enzymes by limiting descriptors and utilizing multiple binding energies, thus improving the performance and efficiency of inhibitor identification.
Smart Images

Figure 0007853944000006 
Figure 0007853944000007 
Figure 0007853944000008
Abstract
Description
[Technical Field]
[0001] This disclosure relates to training and using a classification model to predict the inhibitory properties of molecules in determined drug metabolizing enzymes (DMEs), specifically enzymes belonging to the families of cytochromes P450 (CYP), sulfotransferases (SULT), and UDP-glucuronosyltransferases. [Background technology]
[0002] Drug-metabolizing enzymes play a crucial role in the metabolism of endogenous molecules, xenobiotics, and drugs taken into the human body. Their primary role is to detoxify the body by modifying endogenous and exogenous compounds, such as contaminants or drugs, to facilitate their rapid excretion. However, in some cases, drug-metabolizing enzymes can enhance the toxicity of their substrates, thereby inducing serious side effects and adverse drug reactions. Phase I DME catalyzes oxidation reactions that yield metabolites, which may then be excreted or modified by Phase II DME, which catalyzes conjugation reactions. In some cases, Phase II DME can directly modify compounds without going through Phase I DME. DME inhibition is a complex process, involving competitive inhibition at the active site, modification of the substrate or metabolic flux between the active site and the outside of the enzyme, or inhibition by the drug itself or its metabolites (time-dependent inhibition), which can lead to harmful drug interactions. Therefore, predicting potential DME inhibition is crucial for preventing drug toxicity.
[0003] Among DMEs, cytochrome P450 (CYP) is a superfamily of oxidases responsible for the metabolism of drugs, xenobiotics, and endogenous molecules. It is estimated that approximately 75% of commercially available drugs are metabolized by CYP, which has six major isoforms: 1A2, 2C8, 2C9, 2C19, 2D6, and 3A4. CYP inhibition leads to reduced drug elimination, which is a major cause of harmful drug interactions. In some cases, CYP oxidation can result in toxic metabolites. Therefore, identifying potential inhibition of CYP enzymes is essential for clinical drug therapy and early-stage drug discovery.
[0004] Among DMEs, another family of enzymes that metabolizes many drugs is sulfotransferase (SULT). SULT catalyzes sulfation of a substrate's hydroxyl or amino group from its cofactor, 3'-phosphoadenosine 5'-phosphosulfate (PAPS), by performing nucleophilic attack. At high concentrations, some substrates inhibit the enzyme, and a dead-end complex containing the bound, inactive cofactor PAP has been identified. Sulfation usually promotes excretion, but in some special cases, it can also increase the pharmacological activity of some drugs (for example, the blood pressure-lowering prodrug minoxidil becomes fully active after sulfation). Furthermore, SULT can denature some chemicals into carcinogens or promutagenic activators (e.g., 7,12-dimethylbenz(a)anthracene) by producing highly reactive sulfate esters that can covalently bind to DNA. SULTs, which are responsible for the metabolism of small endogenous compounds and xenobiotics, are localized in the cytosol. Four families of human SULTs have been identified to date: SULT1, SULT2, SULT4, and SULT6. Among the SULT1 family, SULT 1A1, which metabolizes a wide variety of compounds such as phenols, sex steroid hormones (estrogen), thyroid hormones, and drugs (e.g., minoxidil, paracetamol, 17α-ethinylestradiol), is the most expressed (found in the liver, intestines, kidneys, thyroid, and platelets).
[0005] Finally, UDP-glucuronosyltransferase (UGT) belongs to the uridine diphosphate glucuronyltransferase gene family. UGT catalyzes the covalent addition of the sugar moiety of glucuronic acid to numerous therapeutic and environmental toxins, as well as various endogenous steroids and other signaling molecules. UGT-catalyzed glucuronidation is thought to account for up to 35% of phase II drug metabolism reactions. The three major isoforms, UGT 2B7, UGT 1A4, and UGT 1A1, are responsible for 35%, 20%, and 15% of drug modifications, respectively, of drugs metabolized by UGT.
[0006] To predict the inhibitory properties of molecules mediated by DME, classification model-based methods have been proposed.
[0007] Specifically, in the publication "Integrated structure- and ligand-based in silico approach to predict inhibition of cytochrome P450 2D6" by VY Martiny et al., published on August 26, 2015, in Bioinformatics 31(24), 3930-3937, doi: 10.1093 / bioinformatics / btv486, an in silico method for predicting CYP2D6 is disclosed. Here, three learning algorithms—a support vector machine, a random forest, and a naive Bayesian—are trained and tested to predict cytochrome P450 2D6 inhibition.
[0008] More specifically, this paper discloses the identification of the binding site conformation that best predicts the binding energy by performing various molecular dynamics simulations (MD) of CYP2D6 in one apostructure PDB ID 2F9Q and one holostructure co-crystallized with a prinomast.
[0009] Next, each classification model was trained in a learning database containing 343 inhibitors of CYP2D6 and 3002 inhibitors, using a descriptor set including extended connectivity fingerprints (ECFPs) and protein-ligand binding energies calculated at the best MD receptor conformation as input descriptors for a given molecule.
[0010] These inhibition models were able to predict CYP2D6 inhibition with 78% accuracy on the training set and 75% accuracy on the external validation set. However, this method suffers from several limitations. First, the large number of descriptors used in this model (up to 2000) and the number of descriptor types (ECFP) make it difficult to understand why a molecule is predicted as an inhibitor or non-inhibitor, or the main reasons that influence the inhibitory properties of a molecule.
[0011] Furthermore, the large number of descriptors means that computation time is required to compute each descriptor for a given numerator, which slows down the training and use of the classification model.
[0012] Finally, this model uses a single binding energy calculated by computer for a single conformation of the enzyme as a descriptor. It is expected that using a larger number of binding energies corresponding to various conformations could improve the performance of the predictive model. [Prior art documents] [Non-patent literature]
[0013] [Non-Patent Document 1] On August 26, 2015, VY Martiny et al. published "Integrated structure- and ligand-based in silico approach to predict inhibition of cytochrome P450 2D6" in Bioinformatics 31(24), 3930-3937, doi: 10.1093 / bioinformatics / btv486. [Non-Patent Document 2] Louet, M., Labbe, CM, Fagnen, C., Aono, CM, Homem-de-Mello, P., Villoutreix, BO, Miteva, MA, Insights into molecular mechanisms of drug metabolism dysfunction of human CYP2C9*30., PLoS One 2018, 13 (5), e0197249 [Overview of the project] [Problems that the invention aims to solve]
[0014] In view of the above, the present invention aims to propose a model with improved performance for predicting the inhibition of at least one drug-metabolizing enzyme, specifically, type CYP, SULT, or UGT.
[0015] Another objective of this invention is to propose a model that accelerates training and use.
[0016] Another objective of the present invention is to propose a model that can help us better understand enzyme inhibitors using molecules. [Means for solving the problem]
[0017] For this purpose, a method for training a model for predicting inhibitors of determined CYP, SULT, or UGT enzymes, implemented by a training device, is disclosed, the training device comprising a computer and a memory storing a training data set comprising the number of molecules known to be inhibitors or non-inhibitors of the determined enzyme, the method comprising: - selecting a subset of descriptors based on the relative importance of the descriptors in predicting the inhibitory properties of the molecules from an initial set of molecular descriptors comprising physicochemical molecular descriptors and at least one binding energy in at least one conformation of the determined enzyme; - performing supervised training on the training data set of a classification model configured to receive as input a vector formed from a subset of computer-calculated molecular descriptors for a molecule and to output an indication of the inhibitory properties of the molecule with respect to the determined enzyme. comprises.
[0018] In embodiments, the determined enzyme is - CYP 2C9 - CYP 2D6 - SULT 1A1 - SULT 1A3, and - UGT 1A1 selected from the group consisting of.
[0019]
[0020] In embodiments, the step of selecting descriptors based on their relative importance comprises training a plurality of random forest models on the training data set, computationally calculating the Gini importance index of all descriptors in the set, and selecting the descriptors with the highest Gini importance.In some embodiments, the determination of the number of descriptors to select, based on their relative importance, includes the steps of: computing the average balanced accuracy of multiple sets of descriptors for multiple random forest models having varying numbers of descriptors; and selecting the number of descriptors that maximizes the balanced accuracy.
[0021] In some embodiments, the method, prior to the selection step, selects an initial set of descriptors. - Highly correlated descriptors, - Descriptors with missing or infinite values in the training dataset, - Descriptors with variance below a determined threshold for the training dataset Includes the step of removing.
[0022] In various embodiments, the classification model is a random forest model or a support vector machine model.
[0023] For another purpose, a classification model is disclosed which is configured to predict whether a molecule is an inhibitor of a given enzyme, and the classification model is obtained by training it on a training dataset in accordance with the method described above.
[0024] In various embodiments, the classification model is - A first classifier formed by a random forest model trained according to the above description, - A second classifier formed by a support vector machine model trained according to the above description, and - A third classifier indicating whether a molecule is an enzyme inhibitor, based on a comparison of the lowest binding energy calculated by computer for multiple conformations of the enzyme with at least one threshold. The model's output can be formed from these elements, and is a majority vote on the three classifiers.
[0025] A method for predicting whether a candidate molecule is an inhibitor of a given enzyme is also disclosed, and this method is: - A step of computer calculation of a set of molecular descriptors for the candidate molecule and at least one binding energy of the candidate molecule in at least one conformation of the enzyme, - A step of providing a computer-calculated molecular descriptor and each computer-calculated binding energy to a classification model trained to output an indication of whether a candidate molecule is an enzyme inhibitor or a non-inhibitor, based on the set of molecular descriptors and the at least one binding energy of the candidate molecule in the enzyme conformation. - A step of receiving an instruction output by classification regarding whether the candidate molecule is an enzyme inhibitor or a non-inhibitor. Includes.
[0026] In some embodiments, a method for predicting whether a candidate molecule is an inhibitor of a given enzyme further includes the step of training a classification model according to the training method disclosed above.
[0027] In various embodiments, the method includes the steps of providing a computer-calculated molecular descriptor and each computer-calculated binding energy to a first classifier formed by a random forest model and a second classifier formed by a support vector machine model, and receiving instructions from each classifier regarding whether a candidate molecule is an inhibitor or non-inhibitor of a given enzyme. The method includes, - A step of computer calculation to determine the binding energy of candidate molecules to each conformation of the enzyme for multiple conformations of the enzyme, The steps include comparing the lowest computer-calculated binding energy with two thresholds and inferring a third instruction from the comparison, The steps involve determining whether a candidate molecule is an enzyme inhibitor or non-inhibitor based on a majority vote on three instructions. It also includes.
[0028] In various embodiments, the candidate molecule is a candidate drug or a xenobiotic substance.
[0029] For another purpose, a computer program product is disclosed, which includes code instructions for implementing the training or prediction method disclosed above.
[0030] The claimed method for training a classification model comprises the step of selecting a subset of molecular descriptors, which are used as input to a classification model, wherein the molecular descriptors include physicochemical parameters of the molecule under consideration and at least one binding energy in at least one conformation of the determined enzyme. [Effects of the Invention]
[0031] The selection of descriptors is based on their relative importance in predicting enzyme inhibition, thus limiting the number of descriptors and reducing the computation time required to calculate the descriptors for a given molecule.
[0032] Furthermore, using the physicochemical parameters of the molecule instead of ECFP allows for a better understanding of the inhibitors of the molecule.
[0033] Finally, to improve performance, the model may consider various binding energies calculated by computer for different conformations of the enzyme. However, the binding energies are also used for descriptor selection, and only those energies of high importance for predicting inhibition are retained.
[0034] Other features and advantages of the present invention will become apparent from the following detailed description, given by non-limiting examples with reference to the accompanying drawings. [Brief explanation of the drawing]
[0035] [Figure 1a] This figure schematically shows the main steps of a training method according to one embodiment. [Figure 1b] This figure schematically illustrates a computing device configured to implement a training method according to one embodiment, and / or a method for predicting the inhibitory properties of a molecule using a trained classification model. [Figure 2] This graph shows representative values in percentage for the equilibrium accuracy of 100 random forests using multiple descriptor sets for CYP2C9, CYP2D6, SULT1A1, SULT1A3, and UGT1A1. [Modes for carrying out the invention]
[0036] Next, a method for training a model to predict the inhibitors of a determined drug-metabolizing enzyme (DME) is described. Referring to Figure 1b, the method is implemented by a training apparatus 1, which comprises a computer 10, for example, a processor, microprocessor, controller, or microcontroller, a learning dataset consisting of a training dataset and a validation dataset, and a memory 11 that stores code instructions for implementing the method described above when executed by the computer. The training dataset and validation dataset include a list of known inhibitors and non-inhibitors of the determined enzyme.
[0037] The DMEs under consideration belong to the families of cytochrome P450 (CYP), sulfotransferase (SULT), or UDP-glucuronosyltransferase (UGT). More preferably, the enzymes belong to the following group: - CYP 2C9, - CYP 2D6, - SULT 1A1, - SULT 1A3, and - UGT 1A1 It is one of them.
[0038] Therefore, the method disclosed below is performed on a given enzyme from this group, and the resulting model is specific to predicting the inhibition of said enzyme.
[0039] Referring to Figure 1a, the method may include a preliminary step 90 for preparing a training dataset, which includes collecting known inhibitors and non-inhibitors of the determined enzyme from literature or databases such as ChEMBL, PubChem, BRENDA, Aureus Sciences, or TOXNET. Furthermore, if the number of inhibitors or non-inhibitors found is large, a selection may be made to retain the most active inhibitors. The inhibitory properties of a molecule are given by an index corresponding to the concentration of the molecule that gives a particular percentage of enzyme inactivation. To select the most active inhibitors, only those molecules that give 50% inhibition of the enzyme at concentrations of 10 μM or less can be selected. This is expressed as AC50(IC) ≤ 10 μM. Alternatively, a selection may be made to retain only molecules with the least inhibition, such as molecules that show less than 10% inhibition at a concentration of 50 μM (AC10(IC) < 50 μM). Chemical diversity with a similarity cutoff of 0.8 was adopted. Training and test sets were then constructed using centroids.
[0040] Once this set is obtained, an external validation dataset can be constructed by randomly selecting 20% of both inhibitory and non-inhibitory molecules from the dataset, with the remaining 80% being retained as a training dataset for the model.
[0041] As will be disclosed in more detail below, the predictive model is a classification model configured to take as input a number of computer-calculated descriptors for a given molecule and to output a classification of the molecule as either an inhibitor or a non-inhibitor of a determined enzyme.
[0042] The method includes steps 100 of constructing an initial set of molecular descriptors and 200 of selecting a subset of molecular descriptors from this initial set based on their relative importance in predicting the inhibitory properties of molecules.
[0043] The initial set of molecular descriptors includes physicochemical molecular descriptors that describe molecular characteristics such as size, mass, bulk, volume, shape, structural symmetry and complexity, flexibility, element, electric field and bonding, bond strength, polarity, electronegativity, polarity, ionization potential, aromaticity, lipophilicity, surface area, and polar surface area.
[0044] In various embodiments, physicochemical descriptors include 2D physicochemical descriptors that are computer-calculated numerical properties derived from a linkage table representation of the molecule and are conformationally independent of the molecule. For example, these descriptors may be calculated using PaDEL software.
[0045] The initial set of molecular descriptors may include at least 100 physicochemical descriptors, for example, at least 500 physicochemical descriptors, for example, an initial number of physicochemical descriptors between 500 and 2000.
[0046] The initial set of molecular descriptors also includes at least one binding energy in at least one conformation of the determined enzyme. In some embodiments, at least one structure of the enzyme may be selected from a known database, including, for example, an apo structure and / or at least one holo-cocrystallized structure, and molecular dynamics simulations may be performed for each structure to generate different conformations of the enzyme. For conformation generation, CHARMM or NAMD software may be used, for example.
[0047] The binding energies of molecules in each conformation of an enzyme can be calculated by computer by performing docking operations on the molecules in each conformation. Software such as AutoDock Vina may be used for this purpose.
[0048] Preferably, the initial set of descriptors includes multiple binding energies at several conformations of the enzyme. For example, the initial set of descriptors may include binding energies between 1 and 20, preferably between 2 and 15, for example between 2 and 10. This allows for the consideration of various conformations of the enzyme under consideration within the same classification model. These binding energies are then used in the final descriptor selection by calculating Gini importance.
[0049] The method then involves selecting a subset of descriptors from this initial set.
[0050] Step 200, which involves selecting from an initial set of descriptors, - Descriptors that have missing or infinite values in all data in the training dataset, - Descriptors with a variance close to null with respect to the training dataset (a threshold may be set to remove descriptors, and descriptors have a variance below the said threshold) - Highly correlated descriptors (for example, those with an absolute value of 0.85 or greater, or 0.9 or greater) The preliminary step 210 may include removing the material.
[0051] Next, the descriptor selection 200 includes a step 220 of selecting descriptors based on their relative importance in predicting the inhibitory properties of the molecule.
[0052] In some embodiments, the selection of a subset of descriptors includes training multiple random forest models on a training dataset and selecting a subset of descriptors having the highest Gini importance.
[0053] Here, X1, ..., X correspond to the descriptors. p Considering a binary classification problem with as the independent variable and Y as the response variable, if Y ∈ [0,1] at a given node t of a given tree T that constitutes a random forest, the Gini exponent is: G(t) = 2p t (1- p t ) It is defined as, however, p t =P(Y=0|node=t).
[0054] The Gini index, also known as the Gini impurity index, is a measure of the probability of misclassifying randomly selected elements in a dataset, given that they would be randomly labeled according to the classification distribution within the dataset. When training a decision tree or random forest, the best branch at a given node is selected by maximizing the decrease in the Gini index at that node. Variable X j If node t is branched into two subnodes t1 and t2, the decrease in the Gini exponent at t is:
[0055]
number
[0056] It is defined as follows.
[0057] Here, n t n1 is the number of sample subjects at node t, n2 is the number of sample subjects at node t1, and n3 is the number of sample subjects at node t2. j The importance of Jini is,
[0058]
number
[0059] That is the case.
[0060] Therefore, this method may include computing the Gini importance of each descriptor in an initial set of descriptors (some of which have been removed at the end of step 210), ranking the descriptors according to their Gini importance, and selecting the number of descriptors with the highest Gini importance. In some embodiments, multiple random forests, for example between several hundred and a thousand, may be computed, and the Gini importance of each descriptor may be averaged, thereby eliminating the impact of differences in randomness between models on prediction accuracy.
[0061] It should be noted that if the subset of descriptors selected at the end of step 220 is less important than the physicochemical descriptors used to predict inhibition, it may no longer include binding energies. However, the results given below show that for each of the enzymes CYP2C9, CYP2D6, SULT1A1, SULT1A3, and UGT1A1, multiple binding energies remain at the end of step 220.
[0062] The subset of descriptors selected at the end of step 220 may include fewer than 100 descriptors, for example, between 50 and 100 descriptors. In some embodiments, determining the number of descriptors to retain at the end of step 220 may involve calculating the performance of a random forest calculated using multiple sets of descriptors, from the first top 10 to the first 100 descriptors. The calculated performance may be a representative value of the equilibrium accuracy, which is the average of sensitivity and specificity.
[0063] Referring to Figure 2, the x-axis shows the number of descriptors, and the y-axis shows representative equilibrium accuracy for 100 random forests using multiple sets of descriptors, ranging from the first 10 to the first 100 descriptors, for the enzymes CYP2C9, CYP2D6, SULT 1A1, SULT 1A3, and UGT1A1. The representative equilibrium accuracy is computer-calculated on the training dataset. It can be seen that the equilibrium accuracy for SULT 1A1 and CYP2D6 increases between 40 and 60 descriptors, then reaches a flat region, or even slightly decreases for CYP2D6. Therefore, the number of descriptors can be set to maximize equilibrium accuracy.
[0064] In addition to improving the accuracy of predictive models, reducing the number of descriptors plays a crucial role in the speed and accuracy of descriptor calculations. For example, starting with a set of 382 molecular descriptors means that it takes 8 minutes to compute all descriptors for one molecule. If a selection of 83 descriptors is used, the computation time can be reduced to 1 minute per molecule.
[0065] The same applies to the computation time required for binding energy calculations because the computation time for a single molecule is multiplied by the number of binding energies calculated for the classification model and the number of corresponding various protein conformations.
[0066] Once a subset of descriptors has been selected, the method includes training a classification model configured to take the selected subset of descriptors as input and output a classification of a given molecule as inhibiting or not inhibiting the enzyme under consideration.300
[0067] The classification model is either a random forest model or a support vector machine model. The model is trained, with respect to a learning database, i.e., for each molecule of the training data set, by supervised training, a selected subset of descriptors is computed for the molecule by computer, and an indication of the inhibitory or non-inhibitory properties of the molecule in the determined enzyme is provided to the classification model.
[0068] In embodiments where the classification model is a random forest model, a plurality of decision trees are constructed based on bootstrap samples from the training data set, a small subset of descriptors is randomly selected, and a decision is made at each node of each tree. The final classification of the random forest is obtained by incorporating the results of all the trees by majority vote.
[0069] The number of descriptors at each node of each tree may be equal to √p, which is widely accepted in the art, where p is the number of descriptors in the subset of descriptors selected at the end of step 200. Further, a plurality of random forest models may be trained, within this random forest there are trees with a variable number (for example, between 25 and 1024), and the classification model may be selected as the one that provides the best internal accuracy in terms of the number of trees.
[0070] In embodiments where the classification model is a support vector machine, the descriptors selected at the end of step 200 may be centered at an average of 0 and scaled to a variance equal to 1 in advance. In embodiments, the SVM model is based on a radial basis function kernel. Parameter tuning can be performed by grid search, and the cost parameter is optimized in the range of 2 -2 ~2 20 and the gamma parameter / sigma parameter varies in the range of 2 -20 ~2 2 of the range.
[0071] In some embodiments, step 300 includes training both a random forest model and a support vector machine model.
[0072] The parameters of the trained classification model can be stored in memory.
[0073] Once trained, the classification model can be used to predict the inhibitory properties of molecules that can be candidate drug molecules. Next, testing the molecules includes computationally calculating a subset of the selected descriptors for a given molecule and providing the computationally calculated descriptors to the trained model, whereby the trained model will classify the molecule as inhibitory or non-inhibitory.
[0074] In embodiments where both a random forest model and a support vector machine model are trained, the prediction of the inhibitory properties of a molecule is - the SVM model, - the random forest model, - the third classifier of the calculated binding energies that is computationally calculated for various protein conformations of the drug-metabolizing enzyme, where at least one, and preferably two, thresholds are compared with by taking a majority vote regarding.
[0075] According to this last classifier, a molecule can be assigned as a non-inhibitor if the corresponding binding energy (the lowest among various protein conformations) is greater than a first threshold T1, and as an inhibitor if the binding energy is less than a second threshold T2 (<T1). If the binding energy is between T1 and T2, no determination is made. Using the lowest binding energy for various enzyme conformations makes it possible to find the enzyme conformation that is most suitable for accommodating the ligand of interest. The generated docking position based on the best ranking score (binding energy) for this ligand enables the acquisition of information regarding the enzyme / ligand interaction at the atomic level.
[0076] Therefore, in such embodiments, the prediction method, in addition to providing computer-computed descriptors to the trained SVM model and random forest model, involves an additional step: - A step of computer calculation to determine the binding energy of candidate molecules to each conformation of the enzyme for multiple conformations of the enzyme, - A step of comparing the lowest computer-calculated binding energy with two thresholds and inferring a third indication from the comparison, - A step in which candidate molecules are determined to be either enzyme inhibitors or non-inhibitors according to a majority vote on three instructions. It can include...
[0077] The computer equipment used to compute descriptors and apply the trained model may be the same as or different from the training equipment described above.
[0078] (Examples) Known inhibitors of CYP2C9 were retrieved from a database, and only inhibitors with AC50(IC) ≤ 10 μM were retained, while non-inhibitors showing <20% inhibition at a 50 μM concentration were retained. The training dataset ultimately yielded 3811 inhibitors and 2468 non-inhibitors. For CYP2D6, the training dataset ultimately yielded 343 inhibitors and 3002 non-inhibitors. For SULT1A1, 87 inhibitors and 500 decoy non-inhibitors were retained. For SULT1A3, 76 inhibitors and 370 decoy non-inhibitors were retained. For UGT1A1, 71 inhibitors and 361 decoy non-inhibitors were retained.
[0079] Two X-ray CYP2C9 structures, namely 5XXI co-crystallized with losartan and 1R90 co-crystallized with flurbiprofen, were imported from the Protein Data Bank. Seven conformations, including two crystals and five protein centroid structures with diverse binding pocket conformations, were generated from previously performed MD simulations, which are available in Louet, M., Labbe, CM, Fagnen, C., Aono, CM, Homem-de-Mello, P., Villoutreix, BO, Miteva, MA, Insights into molecular mechanisms of drug metabolism dysfunction of human CYP2C9*30., PLoS One 2018, 13 (5), e0197249.
[0080] For CYP2D6, six conformations were generated. Two X-ray structures, namely 3QM4, co-crystallized with the prinomast, and the apo structure, 2F9Q, were obtained from the Protein Databank. Six conformations, including two crystal structures and four protein centroid structures containing diverse binding pocket conformations, were generated from previously performed MD simulations, which are available in Martiny VY, Carbonell P, Chevillard F, Moroy G, Nicot AB, Vayer P, Villoutreix BO, and Miteva MA., Integrated structure- and ligand-based in silico approach to predict inhibition of cytochrome P450 2D6. Bioinformatics. 2015, 31(24):3930~7.
[0081] For SULT1A1, nine conformations were generated. One X-ray structure 4GRA was imported from the Protein Databank. In addition, two protein centroid structures containing the cofactor PAP, and six protein centroid structures containing the cofactor PAPS were generated from previously performed MD simulations.
[0082] For SULT1A3, the 13 protein centroid structures, including the cofactor PAPS, were generated from the previously performed MD simulation, starting with the X-ray structure 2A3R obtained from the Protein Structure Databank.
[0083] For UGT1A1, 10 protein centroid structures containing the cofactor UDP-glucuronic acid were generated from previously performed MD simulations, starting with a homology model of the substrate and cofactor-binding domain.
[0084] 1050 2D physicochemical molecular descriptors were computed on the training and validation datasets using PaDEL software. Some descriptors were removed according to step 210, specifically by removing descriptors with an absolute Pearson correlation coefficient greater than 0.9, resulting in 382 physicochemical descriptors. Binding energies for each conformation were added to these descriptors. Then, according to step 220, the most important descriptors were selected. Figure 2 shows the progress of representative equilibrium accuracy for 100 random forests along with the number of descriptors for CYP2D6, CYP2C9, SULT1A1, SULT1A3, and UGT1A1. Table 1 below shows the representative equilibrium accuracy in percentage for 100 random forests, including all 382 physicochemical descriptors and binding energy descriptors, i.e., before descriptor selection.
[0085] [Table 1]
[0086] Ultimately, the top 88 descriptors containing five binding energies were selected for CYP2C9, the top 88 descriptors containing five binding energies for CYP2D6, the top 60 descriptors containing four binding energies for SULT1A1, the top 85 descriptors containing five binding energies for SULT1A3, and the top 86 descriptors containing six binding energies for UGT1A1.
[0087] Random forest classification was performed using the Random Forest R library in the statistical software package R. Random forest calculations were performed by scanning across ntrees (number of trees) ranging from 25 to 1024 and mtry (number of descriptors per node) ranging from 5 to 18. For each model, the combination of ntree and mtry parameters that yielded the best internal accuracy was selected. A second scan was performed across the samplesize parameter of RandomForest in the R software, which allows selecting the number of positive / negative numerators to include for each tree when the training dataset is non-equilibrium.
[0088] Table 2 shows the final random forest model prediction accuracy in percentage and its corresponding parameters. Equilibrium accuracy is the average of sensitivity and specificity.
[0089] [Table 2]
[0090] The support vector machine model was also constructed using the radiating kernel implemented in the R package with the e1071 and Caret libraries. Parameter tuning was performed using a grid search with 10 validations repeated 5 times. The cost parameter was within the range of 2. -2 ~2 20 Optimized with gamma / sigma = 2 -20 From 2 2We varied the values up to that point. To highly compensate for the non-equilibrium of the dataset, we used weight parameters and penalized misclassified observables.
[0091] Table 3 shows the final SVM model prediction accuracy in percentage and its corresponding parameters.
[0092] [Table 3]
[0093] It can be observed that both models yield higher equilibrium accuracy for CYP2D6 than the methods discussed in the section on conventional techniques. Specifically, the selected descriptors allow for higher accuracy and increased computer computation speed for testing the inhibitory properties of the molecule.
[0094] As shown above, in addition to the final SVM and RF models, the lowest of the calculated binding energies for various protein conformations of DME can be used as a third classifier. This allows for a final determination of which molecules will be assigned as inhibitors or non-inhibitors of DME when a majority vote is taken on the SVM model, RF model, and energy decisions.
[0095] Using the calculated binding energy as a third classifier, - The thresholds for CYP2C9 and CYP2D6 can be -7.0 kcal / mol and -8.5 kcal / mol, respectively. Therefore, according to this classifier, a molecule is considered a non-inhibitor if its binding energy is >-7.0 kcal / mol, and an inhibitor if its binding energy is <-8.5 kcal / mol. - The thresholds for SULT1A1 and SULT1A3 may be -5.0 kcal / mol and -7.5 kcal / mol, respectively. Therefore, according to this classifier, a molecule is determined to be a non-inhibitor if its binding energy is >-5.0 kcal / mol, and an inhibitor if its binding energy is <-7.5 kcal / mol. - The thresholds for UGT1A1 may be -6.5 kcal / mol and -8.0 kcal / mol. Therefore, according to this classifier, a molecule is judged to be a non-inhibitor if its binding energy is >-6.5 kcal / mol, and an inhibitor if its binding energy is <-8.0 kcal / mol. [Explanation of Symbols]
[0096] 1 training device 10 Computers 11 memory 90 Preliminary steps to prepare the training dataset Steps to build an initial set of 100 molecular descriptors Selection of a subset of 200 molecular descriptors 210 Preliminary step to remove descriptors 220 Steps to select descriptors 300 Steps to train both a Random Forest model and a Support Vector Machine model
Claims
1. A method for training a random forest model or support vector machine classification model for predicting inhibitors of determined CYP, SULT, or UGT enzymes, which is implemented by a training device (1), wherein the training device (1) comprises a computer (10) and a memory (11) that stores a list of molecules that are known inhibitors or non-inhibitors of the determined enzyme, and the method is - For each molecule, the steps (220) include selecting a subset of the descriptors from an initial set of molecular descriptors in the list of molecules, which include a physicochemical molecular descriptor and at least one binding energy in at least one conformation of the enzyme determined, based on the relative importance of the descriptors in predicting the inhibitory properties of the molecule. - Step (300) of supervising training a random forest model or a support vector machine classification model configured to take a vector formed from the subset of molecular descriptors computed for each molecule as input, and output the indication of the inhibitory or non-inhibitory properties of the molecule for the determined enzyme, on a training dataset that includes for each molecule the subset of selected descriptors and an indication of the inhibitory or non-inhibitory properties of the molecule for the determined enzyme, and output an indication of the inhibitory properties of the molecule for the determined enzyme. Methods that include...
2. The enzyme determined above is - CYP 2C9 - CYP 2D6 - SULT 1A1 - SULT 1A3, and - UGT 1A1 The method according to claim 1, selected from the group consisting of the following.
3. The method according to claim 1 or 2, wherein the step of selecting descriptors based on the relative importance of the descriptors includes the steps of training multiple random forest models on a training dataset, computing the Gini importance index of all descriptors in the set, and selecting the descriptor having the highest Gini importance.
4. The method according to claim 3, wherein the step of determining the number of descriptors to select based on the relative importance of the descriptors includes the step of computing a representative value of the equilibrium accuracy of multiple random forest models having varying numbers of descriptors for multiple sets of descriptors, and the step of selecting the number of descriptors that maximizes the equilibrium accuracy.
5. Prior to the selection step (220) above, from the initial set of descriptors, - Highly correlated descriptors, - Descriptors having missing values or infinite values in the data of the training dataset, - The method according to claim 3 or 4, comprising the step (210) of removing descriptors having a variance below a determined threshold with respect to the training dataset.
6. A computer-assisted method for predicting whether a candidate molecule is an inhibitor of a predetermined enzyme selected from CYP, SULT, or UGT enzymes, - A step of computer-calculating a set of molecular descriptors for the candidate molecule and at least one binding energy of the candidate molecule in at least one conformation of the enzyme, - A step of providing the computer-calculated molecular descriptors and each computer-calculated binding energy to at least one random forest model or support vector machine classification model trained by implementing the method of any one of claims 1 to 5, to output an indication of whether the candidate molecule is an inhibitor or a non-inhibitor of the enzyme, based on the set of molecular descriptors and the at least one binding energy of the candidate molecule in the conformation of the enzyme. - A step in which the computer receives instructions from the random forest model or support vector machine classification model regarding whether the candidate molecule is an inhibitor or non-inhibitor of an enzyme selected from CYP, SULT, or UGT enzymes. Methods that include...
7. The steps include providing the computer-calculated molecular descriptor and each computer-calculated binding energy to a first classifier formed by a random forest model and a second classifier formed by a support vector machine model, and receiving instructions from each classifier regarding whether the candidate molecule is an inhibitor or non-inhibitor of a predetermined enzyme, - A step of using a computer to calculate the binding energy of the candidate molecule to each conformation of the enzyme for multiple conformations of the enzyme, - A step of comparing the lowest computer-calculated binding energy by the computer with two thresholds, and inferring a third instruction from the comparison, - A step in which a computer determines whether the candidate molecule is an inhibitor or non-inhibitor of the enzyme according to a majority vote on the three instructions. A method performed by a computer according to claim 6, further comprising:
8. The computer-based method according to any one of claims 6 to 7, wherein the candidate molecule is a candidate drug or a xenobiotic.
9. A computer program product, when executed by a computer, comprising code instructions for implementing the method described in any one of claims 1 to 8.