A method for constructing a drug hepatotoxicity prediction model and use thereof

The drug hepatotoxicity prediction system, built using high-content analysis and machine learning models, solves the problem of insufficient accuracy in drug hepatotoxicity prediction in existing technologies, achieving efficient and accurate drug hepatotoxicity screening and improving the safety and efficiency of new drug development.

CN114694770BActive Publication Date: 2025-12-19ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011616758.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-12-19
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

Existing methods for predicting drug hepatotoxicity have insufficient accuracy in animal models, especially for idiosyncratic drug-induced liver injury (DILI), and the causal relationship between traditional cytotoxicity test indicators and clinical DILI is unclear, making it difficult to efficiently screen for drugs with hepatotoxicity.

Method used

Changes in specific cellular phenotypic parameters of human hepatocytes were measured using high-content analysis techniques. Combined with drug exposure levels, machine learning models were used to construct drug hepatotoxicity prediction models, including Fisher linear discriminant analysis or Naive Bayes classification models, to identify mechanistic patterns related to drug-induced hepatotoxicity (DILI).

Benefits of technology

It achieves efficient early identification of drug hepatotoxicity with an accuracy of 87%, a sensitivity of 84%, and a specificity of 94%, significantly improving the safety and efficiency of new drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694770B_ABST
    Figure CN114694770B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of pharmacy and medicine, in particular to a method for constructing a drug hepatotoxicity prediction model and its application in predicting drug hepatotoxicity, further relates to a method, system and device for predicting drug hepatotoxicity, and further relates to the use of a combination of cell phenotype parameters in predicting drug hepatotoxicity. The method for predicting drug hepatotoxicity provided by the present disclosure can identify the toxicity of a drug to be tested on hepatocytes by performing high-content analysis test experiments on hepatocytes in at least 3 groups (7 parameters) and at most 6 groups (13 parameters), combined with the drug human body exposure C max value, i.e. the toxicity of the drug to be tested on hepatocytes can be identified, and the accuracy of the method can be up to 87%, and the highest sensitivity and specificity are 84% and 94% respectively, which are significantly higher than the reported prediction methods based on HCA or other technologies.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of pharmacy and medicine, in particular to a method for constructing a drug hepatotoxicity prediction model and application thereof in predicting drug hepatotoxicity, and further relates to a method, system and device for predicting drug hepatotoxicity, and further relates to the use of a combination of cell phenotype parameters in predicting drug hepatotoxicity. BACKGROUND

[0002] Drug induced liver injury (DILI) is a term used to describe liver injury caused by prescription drugs, over-the-counter drugs or Chinese herbal medicines. In clinical practice, it manifests as various acute and chronic liver diseases, and in severe cases, it can lead to acute liver failure and liver transplantation. It is the most common adverse drug reaction (Hamilton LA, Collins-Yoder A, Collins RE, AACN advanced critical care 2016, 27(4):430-40). DILI not only seriously endangers the health of drug users, but also is the main reason for the failure of drug development in the late stage, the restriction of use and the withdrawal of drugs, resulting in huge losses of drug development investment. According to statistics, nearly one-fourth of the candidate drugs entering clinical development between 2013 and 2015 failed due to safety reasons (Segall MD, Chris BJDDT. Addressing toxicity risk when designing and selecting compounds in early drug discovery[J]. 2014;19(5):688-93; Harrison RK. Phase II and phase III failures: 2013-2015[J]. Nature reviews Drug discovery 2016;15(12):817-8), and DILI is one of the important reasons. As of 2015, drugs withdrawn due to liver toxicity accounted for 21% of all withdrawn drugs (Watkins PB, Merz M, Avigan MI, Kaplowitz N, Regev A, Senior JRJDS. The Clinical Liver Safety Assessment Best Practices Workshop: Rationale, Goals, Accomplishments and the Future[J]. 2014;37(1):1-7; Siramshetty VB, Nickel J, Omieczynski C, Gohlke B-O, Drwal MN, Preissner R. WITHDRAWN-a resource for withdrawn and discontinued drugs[J]. Nucleic Acids Research 2015;44(D1):D1080-D6).Therefore, DILI has become a serious concern for new drug development companies, regulatory agencies, and clinicians (Hoofnagle JH, Serrano J, Knoben JE, Navarro VJ. LiverTox: a website on drug-induced liver injury [J]. Hepatology 2013; 57(3): 873-4; Bjornsson ES. Hepatotoxicity by Drugs: The Most Common Implicated Agents [J]. Int J Mol Sci 2016; 17(2): 224). How to efficiently identify the liver toxicity of candidate compounds at an early stage of new drug development and establish a more efficient drug liver toxicity prediction method for clinical transformation has become an urgent need for the pharmaceutical industry.

[0003] Traditionally, the prediction of DILI has been mainly based on the evaluation of safety in animals developed decades ago, by observing the changes in the liver of rodents and non-rodents after long-term drug administration, and then further evaluated in clinical phase I-III trials. However, the species differences between animals and humans limit the use of animal models to predict human clinical responses (Martignoni M, Groothuis GM, de Kanter R. Species differences between mouse, rat, dog, monkey and human CYP-mediated drug metabolism, inhibition and induction [J]. Expert Opin Drug Metab Toxicol 2006; 2(6): 875-94). Previous studies have shown that the accuracy of non-rodent and rodent safety evaluation results in predicting human toxicity events is only 63% and 43%, respectively, and for DILI, the accuracy of animal model evaluation results is about 55% (Olson H, Betton G, Robinson D, et al. Concordance of the toxicity of pharmaceuticals in humans and in animals [J]. Regulatory toxicology and pharmacology: RTP 2000; 32(1): 56-67). Moreover, the predicted liver toxicity is basically intrinsic DILI (Intrinsic drug induced liver injury), also known as non-specific DILI, and the prediction rate for idiosyncratic DILI (Idiosyncratic drug induced liver injury, IDILI) is even lower (Funk C, Roth A. Current limitations and future opportunities for prediction of DILI from in vitro [J]. Arch Toxicol 2017; 91(1): 131-42).Even if the candidate drug passes the clinical phase III experiment, however, due to the limited exposure of the human population, most IDILI is still difficult to discover, eventually leading to the withdrawal of the drug after its marketing (Seung-Hyun K, Naisbitt DJ. Update on Advances in Research on Idiosyncratic Drug-Induced Liver Injury [J]. Allergy Asthma & Immunology Research 2016; 8(1): 3-11). Given the limitations of traditional liver toxicity prediction methods, and the long time-consuming, high-cost, large-dose animal safety evaluation method, which is not suitable for high-throughput screening in the early stage of modern drug research, combined with the increasing emphasis on animal welfare internationally, advocating the 3R principle (reduce, replace, optimize), therefore, looking for animal alternative toxicity evaluation method has become the strategy of the development of drug hepatotoxicity evaluation method (Macdonald JS, Robertson RTJTS. Toxicity testing in the 21st century: a view from the pharmaceutical industry [J]. 2009; 110(1): 40-6; Krewski D, Andersen ME, Tyshenko MG, et al. Toxicity testing in the 21st century: progress in the past decade and future perspectives [J]. Archives of Toxicology 2020; 94(1): 1-58).

[0004] In the past decade, with the rapid development of modern life sciences, computer, bioinformatics and other technologies, the development of toxicology has also entered a new stage. Based on the existing understanding of the mechanism of DILI, the development of in vitro test methods using human cells as models has become the direction of DILI drug prediction technology. Scientists have made various attempts and applied life science technologies including genomics, proteomics, and metabolomics to establish various in vitro liver toxicity test prediction systems, including 2D or 3D cells, liver tissue, micro-liver chips, liver slices, and computer prediction methods integrating various information of the compound. However, on the one hand, due to the complexity of DILI mechanism, the understanding of DILI, especially the mechanism of idiosyncratic DILI (iDILI), is still very limited, on the other hand, there are still many technical bottlenecks to be broken through in some physiological in vitro models. So far, there is no drug liver toxicity screening method that is widely accepted and permitted by the industry and regulatory agencies.

[0005] High Content Analysis (HCA) technology is an analysis technology based on quantitative determination of cell image developed to meet the needs of efficient new drug screening, which can obtain multi-dimensional information related to the mechanism and effect of drug toxicity in a single experiment. Currently, only the HCA drug liver toxicity test method based on various human liver cells has been verified by a large number of drugs, and it is also suitable for the screening of liver toxicity of early-stage candidate drugs in new drug development. However, the existing methods are mainly based on the test of several simple and non-specific cytotoxic parameters, and these indicators are not only unclearly related to the causal relationship of clinical DILI, cannot reflect the diverse and complex characteristics of DILI drugs, and have not been systematically screened. Moreover, due to the different drug selection and classification standards used by different researchers in the modeling process, it directly affects the consistency and credibility of the existing method prediction, and it is impossible to predict drugs with iDILI. Therefore, a new drug liver toxicity prediction method based on 2D cell HCA technology needs to be discovered and developed.

[0006] DISCLOSURE

[0007] The present disclosure uses machine learning methods to identify and verify the mechanism pattern (combination of specific cell phenotypes) with the strongest correlation with DILI by HCA determination of the liver cell toxicity phenotype spectrum of a large number of known DILI drugs combined with the human body exposure level of the drug, thereby constructing a new 2D human cell HCA-based DILI drug identification and prediction system suitable for early-stage drug discovery and clinical research, which is efficient. This system can provide technical support for reducing the cost of new drug research and development, improving the efficiency of new drug research and development, and ensuring the safety of clinical drug use.

[0008] The present disclosure provides a method for constructing a drug liver toxicity prediction model, comprising:

[0009] n known drugs with severe DILI (sDILI) and non-DILI (nDILI) are collected, C max information of the drugs is collected;

[0010] cells are treated with different concentrations of the drugs, the rate of change of the parameters of specific cell phenotypes of the cells is determined by HCA, the lowest effective concentration (LEC) of the drugs on each cell phenotype is determined, and the TI LEC value of each cell phenotype is calculated using the formula TI max = LEC / C LEC ;

[0011] The TI LEC values of specific cell phenotypes of sDILI and nDILI drugs are used to train a machine learning model to build a drug liver toxicity prediction model,

[0012] wherein n is an integer greater than or equal to 10 (e.g., greater than or equal to 50, greater than or equal to 60, such as 55, 60, 65, 70, 80),

[0013] LEC is the concentration of the drug that causes the rate of change of the parameters of the cell phenotype to be greater than or equal to 25%,

[0014] The specific cell phenotypes include: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h, or

[0015] The specific cell phenotypes include: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h,

[0016] Preferably, the machine learning model is a Fisher linear discriminant analysis model or a Naive Bayes (Bayes) classification model.

[0017] In one implementation, the method of building a drug liver toxicity prediction model, wherein the TI LEC values of specific cell phenotypes of sDILI and nDILI drugs are used to train a machine learning model includes: ​

[0018] Randomly extract 55% to 95% (for example, 60%, 65%, 75%, 85%, or 90%) of the n drug samples as a training set, train the machine learning model using the training set, and construct a drug liver toxicity prediction model.

[0019] In one implementation, the method of constructing a drug liver toxicity prediction model, wherein the training of the machine learning model using the TI values of the specific cell phenotypes of the sDILI and nDILI drugs comprises: LEC

[0020] Randomly extract 55% to 95% (for example, 60%, 65%, 75%, 85%, or 90%) of the n drug samples as a training set, randomly extract m times to obtain m training sets, train the machine learning model using the m training sets, and construct m drug liver toxicity prediction models, and integrate the m drug liver toxicity prediction models into a drug liver toxicity integrated prediction model,

[0021] wherein m is an integer greater than or equal to 200, for example, 300, 400, 600, 800, 1000, 1200, 1500, 2000,

[0022] Preferably, the machine learning model is a Fisher linear discriminant analysis model or a Naive Bayes (Bayes) classification model.

[0023] In one implementation, the method of constructing a drug liver toxicity prediction model, wherein the cells are selected from the group consisting of hepatocytes (for example, HepG2 cells), P65-EGFP_CHO, Hif1a-EGFP_CHO, ATF6-EGFP_U2OS, Nrf2-EGFP_A549, or any combination thereof,

[0024] Preferably, the cells are hepatocytes (for example, HepG2 cells), the parameter change rate of the cell phenotype = (fluorescence intensity of the drug treatment group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (fluorescence intensity of the test drug treatment group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug treatment group - fluorescence intensity of the solvent control group) x 100%;

[0025] ​​Preferably, when the cell is P65-EGFP_CHO and / or ATF6-EGFP_A549, the parameter change rate of the cell phenotype = (R value of the drug treatment group - R value of the solvent control group) / R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (R value of the test drug treatment group - fluorescence intensity of the solvent control group) / (R value of the positive drug treatment group - R value of the solvent control group) x 100%, wherein R value = (fluorescence intensity of the fluorescent protein in the nucleus of the test drug treatment cell - background fluorescence intensity) / (fluorescence intensity of the fluorescent protein in the cytoplasm of the test drug treatment cell - background fluorescence intensity);

[0026] Preferably, when the cell is Hif1a-EGFP_CHO and / or Nrf2-EGFP_A549, the parameter change rate of the cell phenotype = (R value of the drug treatment group - R value of the solvent control group) / R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (R value of the test drug treatment group - fluorescence intensity of the solvent control group) / (R value of the positive drug treatment group - R value of the solvent control group) x 100%, wherein R value = fluorescence intensity of the fluorescent protein in the cell / nucleus of the test drug treatment - background fluorescence intensity.

[0027] The present disclosure also provides a drug hepatotoxicity prediction model constructed by the method for constructing a drug hepatotoxicity prediction model according to the present disclosure.

[0028] The present disclosure also provides the use of the constructed drug hepatotoxicity prediction model in predicting drug hepatotoxicity.

[0029] The present disclosure also provides a method for predicting drug hepatotoxicity, comprising:

[0030] treating cells with different concentrations of the test drug, determining the parameter change rate of the specific cell phenotype of the cells by HCA, determining the lowest effective concentration (LEC) of the drug on each cell phenotype, and calculating the TI LEC value of each cell phenotype using the formula TI max = LEC / C LEC ;

[0031] inputting the TI LEC value of the specific cell phenotype of the test drug into the drug hepatotoxicity prediction model constructed according to the present disclosure to determine whether the test drug has hepatotoxicity, or

[0032] inputting the TI LECinputting the drug hepatotoxicity integrated prediction model constructed by the present disclosure, and obtaining m prediction values (i.e., determining the hepatotoxicity of the drug as positive or negative) according to the results output by the m drug hepatotoxicity prediction models, if more than 50% of the m prediction values are determined as positive hepatotoxicity, it is judged that the drug to be tested has hepatotoxicity,

[0033] wherein LEC is the drug concentration causing a change rate of the parameter of the cell phenotype greater than or equal to 25%,

[0034] The specific cell phenotypes include: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h, or

[0035] The specific cell phenotypes include: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h.

[0036] The present disclosure also provides a system for predicting drug hepatotoxicity, comprising: a high content analysis (HCA) instrument, a calculation module, and a prediction module,

[0037] The high content analysis (HCA) instrument is used to determine the change rate of the parameter of the cell phenotype of the cells after being treated with different concentrations of the drug to be tested;

[0038] The calculation module is used to determine the lowest effective concentration (LEC) of the drug, and the formula TI LEC = LEC / C max The TI LEC value is calculated;

[0039] The prediction module comprises a drug hepatotoxicity prediction model or a drug hepatotoxicity integrated prediction model constructed by the present disclosure, and is used to predict the hepatotoxicity of the drug to be tested.

[0040] The present disclosure also provides a device for predicting drug hepatotoxicity, comprising:

[0041] A memory configured to store instructions;

[0042] A processor coupled to the memory, the processor being configured to execute the method for predicting drug hepatotoxicity as described in the present disclosure based on the instructions stored in the memory.

[0043] The present disclosure also provides a computer readable storage medium, wherein the computer readable storage medium stores computer instructions, and the instructions, when executed by a processor, implement the method for predicting drug-induced liver toxicity as described in the present disclosure.

[0044] The present disclosure also provides the use of a specific cell phenotype combination in predicting drug-induced liver toxicity, wherein the specific cell phenotype combination comprises: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h, or

[0045] The specific cell phenotype combination comprises: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h.

[0046] In one embodiment, the specific cell phenotype combination is a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h.

[0047] In one embodiment, the specific cell phenotype combination is a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h.

[0048] In one embodiment, the use of the specific cell phenotype combination in predicting drug-induced liver toxicity, wherein in predicting drug-induced liver toxicity, cells are treated with different concentrations of drugs, the rate of change of the parameters of each cell phenotype in the specific cell phenotype combination is determined by high content analysis (HCA), the lowest effective concentration (LEC) of each cell phenotype is determined, the TI LEC value of each cell phenotype is calculated using the formula TI max = LEC / C LEC , and the liver toxicity of the drug is predicted using the TI LEC value of each cell phenotype in the specific cell phenotype combination.

[0049] Definitions of terms

[0050] In the present disclosure, unless otherwise stated, the scientific and technical terms used herein have the meanings that would be generally understood by a person skilled in the art. Also, the cell culture, molecular genetics, nucleic acid chemistry, immunological laboratory procedures used herein are in accordance with conventional steps as generally used in the respective fields. Meanwhile, in order to better understand the present disclosure, the definitions and explanations of the relevant terms are provided as follows.

[0051] As used herein, the term "LC3B_16h" is the rate of change of the fluorescence intensity of the fluorescently labeled microtubule-associated protein 1 light chain 3B protein output by high content analysis after the test drug incubates HepG2 cells for 16h, and the calculation method of the rate of change of the fluorescence intensity is as follows:

[0052] (Fluorescence intensity of the test drug group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group;

[0053] (Fluorescence intensity of the test drug group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug group - fluorescence intensity of the solvent control group).

[0054] As used herein, the term "MMP_72h" is the rate of change of the fluorescence intensity of the mitochondrial membrane potential output by high content analysis after the test drug incubates HepG2 cells for 72h, and the calculation method of the rate of change of the fluorescence intensity is as follows:

[0055] (Emission value of the test drug group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group;

[0056] (Emission value of the test drug group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug group - fluorescence intensity of the solvent control group).

[0057] As used herein, the term "MnSOD_24h" is the rate of change of the fluorescence intensity of the fluorescently labeled manganese superoxide dismutase output by high content analysis after the test drug incubates HepG2 cells for 24h, and the calculation method of the rate of change of the fluorescence intensity is as follows:

[0058] (Fluorescence intensity of the test drug group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group;

[0059] (Fluorescence intensity of the test drug group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug group - fluorescence intensity of the solvent control group).

[0060] As used herein, the term "Nuclear_24h" is the rate of change of the parameter value reflecting the morphology of the nucleus output by high content analysis of fluorescently labeled nuclei after 24h incubation of HepG2 cells with the test drug, calculated as follows:

[0061] (Test drug group value - solvent control group value) / solvent control group value;

[0062] (Test drug group value - solvent control group value) / (positive drug group value - solvent control group value).

[0063] As used herein, the term "a-Tubulin_24h" is the rate of change of the fluorescence intensity output by high content analysis of fluorescently labeled β-tubulin after 24h incubation of HepG2 cells with the test drug, calculated as follows:

[0064] (Fluorescence intensity of test drug group - fluorescence intensity of solvent control group) / fluorescence intensity of solvent control group;

[0065] (Fluorescence intensity of test drug group - fluorescence intensity of solvent control group) / (fluorescence intensity of positive drug group - fluorescence intensity of solvent control group).

[0066] As used herein, the term "Nuclear_72h" is the rate of change of the parameter value reflecting the morphology of the nucleus output by high content analysis of fluorescently labeled nuclei after 72h incubation of HepG2 cells with the test drug, calculated as follows:

[0067] (Test drug group value - solvent control group value) / solvent control group value;

[0068] (Test drug group value - solvent control group value) / (positive drug group value - solvent control group value).

[0069] As used herein, the term "GSH_24h" is the rate of change of the fluorescence intensity output by high content analysis of fluorescently labeled glutathione after 24h incubation of HepG2 cells with the test drug, calculated as follows:

[0070] (Fluorescence intensity of test drug group - fluorescence intensity of solvent control group) / fluorescence intensity of solvent control group;

[0071] (Fluorescence intensity of test drug group - fluorescence intensity of solvent control group) / (fluorescence intensity of positive drug group - fluorescence intensity of solvent control group).

[0072] As used herein, the term "pH2AX_16h" is the rate of change of the fluorescent intensity of the fluorescently labeled phosphorylated histone H2AX output by high content analysis after incubating HepG2 cells with the test drug for 16h, which is calculated as follows:

[0073] (Fluorescent intensity of the test drug group - fluorescent intensity of the solvent control group) / fluorescent intensity of the solvent control group;

[0074] (Fluorescent intensity of the test drug group - fluorescent intensity of the solvent control group) / (fluorescent intensity of the positive drug group - fluorescent intensity of the solvent control group).

[0075] As used herein, the term "ATF6_5h" is the degree of endoplasmic reticulum stress (R) after incubating ATF6-EGFP_U2OS cells with the test drug for 5h. ATF6-EGFP_U2OS cells are an activated transcription factor 6 nuclear translocation activation response model, in which the fluorescently labeled transcription factor protein is distributed in the cytoplasm in a resting state of the cell, and is translocated to the nucleus when activated, so that the degree of activation of the pathway (R) can be characterized by quantitatively analyzing the number of nuclear translocation of the fluorescent protein. The degree of endoplasmic reticulum stress (R) can be calculated from the fluorescent intensity of the fluorescent protein in the nucleus and the cytoplasmic fluorescent protein output by high content analysis, and the specific calculation method is as follows:

[0076] R = (fluorescent intensity of the fluorescent protein in the nucleus of the test drug treated cells - background fluorescent intensity) / (fluorescent intensity of the cytoplasmic fluorescent protein of the test drug treated cells - background fluorescent intensity).

[0077] As used herein, the term "Hif1α_24h" is the degree of activation of the hypoxic stress response pathway (R) after incubating Hif1a-EGFP_CHO cells with the test drug for 24h. Hif1a-EGFP_CHO cells are a low oxygen-induced factor-1α intracellular accumulation activation response model, in which the fluorescently labeled protein is lowly expressed in a resting state of the cell, and the expression amount of the fluorescently labeled protein increases or even translocates to the nucleus when activated, which is manifested as an increase in the fluorescent intensity in the cell, so that the degree of activation of the corresponding pathway (R) can be characterized by analyzing the content of the accumulated fluorescent protein in the nucleus or the cell. The degree of activation of the hypoxic stress response pathway (R) can be calculated from the fluorescent intensity of the fluorescent protein in the cell and / or the nucleus output by high content analysis, and the specific calculation method is as follows:

[0078] R = (fluorescent intensity of the fluorescent protein in the cell treated with the test drug and / or the nucleus of the cell treated with the test drug) - background fluorescent intensity.

[0079] As used herein, the term "Hif1a_3h" is the degree of activation of the hypoxic stress response pathway (R) after the Hif1a-EGFP_CHO cells are incubated with the test drug for 3h. The Hif1a-EGFP_CHO cell is an activated response pattern of Hif1a accumulation in the cell. In a resting state, the fluorescently labeled protein is expressed at a low level, and when activated, the expression level of the fluorescently labeled protein increases and even moves to the nucleus, which is manifested as an increase in the fluorescence intensity in the cell. Therefore, the degree of activation (R) of the corresponding pathway is characterized by analyzing the accumulation of fluorescent protein in the nucleus or in the cell. The degree of activation (R) of the hypoxic stress response pathway can be calculated by the fluorescence intensity of the fluorescent protein in the cell and / or the nucleus output by high-content analysis, and the specific calculation method is as follows: R = fluorescence intensity of fluorescent protein in the cell and / or nucleus of the cell treated with the test drug - background fluorescence intensity.

[0080] As used herein, the term "Nrf2_6h" is the degree of activation of the oxidative stress pathway (R) after the Nrf2-EGFP_A549 cells are incubated with the test drug for 6h. The Nrf2-EGFP_A549 cell is an activated response pattern of Nrf2 (nuclear factor erythroid-2 related factor 2) accumulation in the cell. In a resting state, the fluorescently labeled protein is expressed at a low level, and when activated, the expression level of the fluorescently labeled protein increases and even moves to the nucleus, which is manifested as an increase in the fluorescence intensity in the cell. Therefore, the degree of activation (R) of the corresponding pathway is characterized by analyzing the accumulation of fluorescent protein in the nucleus or in the cell. The degree of activation (R) of the oxidative stress pathway can be calculated by the fluorescence intensity of the fluorescent protein in the cell and / or the nucleus output by high-content analysis, and the specific calculation method is as follows:

[0081] R = fluorescence intensity of fluorescent protein in the cell and / or nucleus of the cell treated with the test drug - background fluorescence intensity.

[0082] As used herein, the term "NF-κB_0.67h" is the degree of activation of the inflammatory stress NF-κB pathway (R) after the P65-EGFP_CHO cells are incubated with the test drug for 0.67h. The P65-EGFP_CHO cell is an activated response pattern of P65 transcription factor nuclear translocation. In a resting state, the fluorescently labeled P65 protein is distributed in the cytoplasm, and when activated, the fluorescently labeled P65 translocates to the nucleus. Therefore, the degree of activation (R) of the pathway can be characterized by quantitatively analyzing the number of nuclear translocation of the fluorescent protein in the cell. The degree of activation of the inflammatory stress NF-κB pathway (R) can be calculated by the fluorescence intensity of the fluorescent protein in the nucleus and the cytoplasmic fluorescent protein output by high-content analysis, and the specific calculation method is as follows:

[0083] R = (fluorescence intensity of fluorescent protein in the nucleus of the test drug treated cell - background fluorescence intensity) / (fluorescence intensity of fluorescent protein in the cytoplasm of the test drug treated cell - background fluorescence intensity).

[0084] As used herein, the term "C max Cmax is the maximum concentration of a drug exposed in the human body.

[0085] Advantages of the present disclosure

[0086] The method for predicting drug hepatotoxicity provided by the present disclosure can identify the toxicity of the drug to be tested to hepatocytes by performing high-content analysis test experiments on hepatocytes in at least 3 groups (7 parameters) and at most 6 groups (13 parameters), combined with the C max max value of the drug exposed in the human body, and the accuracy of the method can be as high as 87%, and the sensitivity and specificity can be as high as 84% and 94%, respectively, which are significantly higher than the reported prediction methods based on HCA or other technologies. BRIEF DESCRIPTION OF DRAWINGS

[0087] Figure 1 A flowchart showing one embodiment of the method for constructing a drug hepatotoxicity prediction model of the present disclosure is shown;

[0088] Figure 2 A flowchart showing one embodiment of the method for predicting drug hepatotoxicity of the present disclosure is shown;

[0089] Figure 3 A structural diagram showing one embodiment of the system for predicting drug hepatotoxicity of the present disclosure is shown;

[0090] Figure 4 A structural diagram showing one embodiment of the device for predicting drug hepatotoxicity of the present disclosure is shown;

[0091] Figure 5 A flowchart showing the establishment of the method for predicting drug hepatotoxicity of the present disclosure is shown;

[0092] Figure 6 The grouping characteristics of the test drug are shown. Wherein, A shows the types and quantities of indications covered by the test drug; B shows the liver damage types, logP values and daily dosages of the test set drugs, C max max value distribution diagram; C shows the liver damage types, logP values and daily dosages of the validation set drugs, C max max value distribution diagram; CAD represents cardiovasculardiseases drug; GID represents Gastrointestinal diseases drug.

[0093] Figure 7 A flow chart showing the procedure of constructing drug hepatotoxicity prediction model based on specific combination of cell phenotypes using Tclass classification system.

[0094] Figure 8 A display of the preliminary screening results of test set drug HepG2 cell phenotype parameters;

[0095] Figure 9 A display of the preliminary screening results of test set drug cell stress response phenotype parameters;

[0096] Figure 10 A display of the determination results of test set drug cell phenotype parameter effects. Among them, A shows the drug cell phenotype parameter EC 50 and LEC value heat map; B shows the TI 50 and TI LEC value heat map based on drug cell phenotype parameters; C shows the TI 50 value distribution of drugs of different DILI damage types; Figure D shows the TI LEC value distribution of drugs of different DILI damage types; TI 50 = EC 50 / C max , TI LEC = LEC / C max ; *P<0.05, **P<0.01, ***P<0.001 vs nDILI class drugs;

[0097] Figure 11 A display of the DILI determination of test set drugs based on drug phenotype TI 50 and TI LEC values. Among them, A shows the sDILI, mDILI and nDILI drug hepatotoxicity positive parameter heat map; B and E show the number and comparison of hepatotoxicity positive parameters of the three types of drugs; C and F show the sensitivity and specificity of DILI determination based on TI 50 and TI LEC values; D and G show the sensitivity of sDILI, mDILI and nDILI class drugs based on drug TI 50 and TI LEC values, single cell phenotype parameters; *P<0.05, **P<0.01, ***P<0.001 vs nDILI class drugs;

[0098] Figure 12 A display of ROC curve analysis based on TI 50 and TI LEC values of drug affecting cell phenotype parameters, among which the used cell phenotype parameters are all 23 parameters when based on TI 50 prediction, and 23 parameters when based on TI LECPredictive, the used cell phenotype parameters are the remaining 20 parameters after removing IR_72h, F-actin_24h, MMP_24h;

[0099] Figure 13 The Tclass classification system determines the optimized test combination plate and characteristics. A shows the parameter composition of different test combination plates; B shows the ROC curve analysis of different test combination plates;

[0100] Figure 14 The parameter screening results of the optimized test combination plate of the validation set drug are shown;

[0101] Figure 15 The optimized phenotype test combination parameter results of the validation set drug are shown. A shows the LEC value heat map of the drug optimization test combination plate parameters; B shows the TI LEC value heat map of the drug optimization test combination plate parameters; C shows the sDILI, mDILI, aDILI and nDILI class drug hepatotoxicity positive parameter heat map;

[0102] Figure 16 The distribution diagram of the positive parameters of the test drug based on the cell phenotype parameter combination 1 and combination 4 is shown;

[0103] Figure 17 The sensitivity of the prediction method to different types of liver damage of the test drug is shown. C is cholestasis type damage, H is hepatocyte type damage, and M is mixed type damage. DETAILED DESCRIPTION

[0104] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. The description of the at least one exemplary embodiment is actually only illustrative, but not as any limitation on the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present disclosure.

[0105] Unless otherwise specifically indicated, the relative arrangement of parts and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.

[0106] Techniques, methods, and equipment known to those of ordinary skill in the relevant art can not be discussed in detail, but should be considered as part of the authorized specification where appropriate.

[0107] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.

[0108] It should be noted that like reference numerals and letters refer to like items throughout the attached drawings, and therefore once an item is defined in one drawing, it is not necessary to discuss it further in subsequent drawings.

[0109] Figure 1 A flowchart of one embodiment of the method of constructing a drug hepatotoxicity prediction model of the present disclosure, wherein:

[0110] Step 1, collect n known drugs with severe DILI (sDILI) and non-DILI (nDILI), collect the C max information of the drugs. Wherein n is an integer greater than or equal to 10 (for example greater than or equal to 50, greater than or equal to 60, for example 55, 60, 65, 70, 80).

[0111] Step 2, treat cells with different concentrations of drugs, determine the lowest effective concentration (LEC) of each drug on each cell phenotype by HCA to determine the rate of change of the parameters of the specific cell phenotype of the cells, and calculate the TI LEC value of each cell phenotype using the formula TI max = LEC / C LEC . Wherein LEC is the concentration of the drug that causes the rate of change of the parameters of the cell phenotype to be greater than or equal to 25%, and the specific cell phenotype includes LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h and GSH_24h, or the specific cell phenotype includes LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h and Nrf2_24h.

[0112] Step 3, use the TI LEC values of the specific cell phenotypes of the sDILI and nDILI drugs to train a machine learning model to construct a drug hepatotoxicity prediction model.

[0113] The machine learning model in step 3 can be a Fisher linear discriminant analysis model, or a Naive Bayes Bayes) classification model.

[0114] In step 3, the TI values of specific cell phenotypes of sDILI and nDILI drugs are used to train the machine learning model. LEC When training the machine learning model using the TI values of specific cell phenotypes of sDILI and nDILI drugs, 55% to 95% (e.g., 60%, 65%, 75%, 85%, or 90%) of the samples from the n drug samples can be randomly extracted as a training set, and the resulting training set can be used to train the machine model, thereby constructing a drug liver toxicity prediction model.

[0115] In step 3, the TI values of specific cell phenotypes of sDILI and nDILI drugs are used to train the machine learning model. LEC When training the machine learning model using the TI values of specific cell phenotypes of sDILI and nDILI drugs, 55% to 95% (e.g., 60%, 65%, 75%, 85%, or 90%) of the samples from the n drug samples can be randomly extracted as a training set, and the resulting training set can be used to train the machine model, thereby constructing a drug liver toxicity prediction model.

[0116] In step 2, the cells can be selected from liver cells (e.g., HepG2 cells), P65-EGFP_CHO, Hif1a-EGFP_CHO, ATF6-EGFP_U2OS, Nrf2-EGFP_A549, or any combination thereof.

[0117] In step 2, the cells can be liver cells (e.g., HepG2 cells), in which case the parameter change rate of the cell phenotype = (fluorescence intensity of the drug treatment group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (fluorescence intensity of the test drug treatment group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug treatment group - fluorescence intensity of the solvent control group) x 100%, wherein the cell phenotype parameter is selected from LC3B, MnSOD, a-Tubulin, GSH, pH2AX, and MMP.

[0118] In step 2, the cell can be a hepatocyte (e.g., HepG2 cell), the parameter change rate of the cell phenotype = (the value of the response of the cell nucleus morphology of the drug treatment group - the value of the response of the cell nucleus morphology of the solvent control group) / the value of the response of the cell nucleus morphology of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (the value of the response of the cell nucleus morphology of the test drug treatment group - the value of the response of the cell nucleus morphology of the solvent control group) / (the value of the response of the cell nucleus morphology of the positive drug treatment group - the value of the response of the cell nucleus morphology of the solvent control group) x 100%, wherein the cell phenotype parameter is the cell nucleus.

[0119] In step 2, the cell can be P65-EGFP_CHO and / or ATF6-EGFP_U2OS, the parameter change rate of the cell phenotype = (the R value of the drug treatment group - the R value of the solvent control group) / the R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (the R value of the test drug treatment group - the fluorescence intensity of the solvent control group) / (the R value of the positive drug treatment group - the R value of the solvent control group) x 100%, wherein the R value = (the fluorescence intensity of the fluorescent protein in the cell nucleus of the test drug treatment - the background fluorescence intensity) / (the fluorescence intensity of the fluorescent protein in the cytoplasm of the test drug treatment - the background fluorescence intensity).

[0120] In step 2, the cell can be Hif1a-EGFP_CHO and / or Nrf2-EGFP_A549, the parameter change rate of the cell phenotype = (the R value of the drug treatment group - the R value of the solvent control group) / the R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (the R value of the test drug treatment group - the fluorescence intensity of the solvent control group) / (the R value of the positive drug treatment group - the R value of the solvent control group) x 100%, wherein the R value = the fluorescence intensity of the fluorescent protein in the test drug treatment cell or cell nucleus - the background fluorescence intensity.

[0121] The drug hepatotoxicity prediction model or the drug hepatotoxicity integrated prediction model constructed by the above-mentioned method for constructing a drug hepatotoxicity prediction model can be used to predict the hepatotoxicity of a new drug.

[0122] Figure 2 The flowchart of one embodiment of the method for predicting drug hepatotoxicity of the present disclosure. Among them

[0123] Step 1, treat the cells with different concentrations of the drug to be tested, determine the lowest effective concentration (LEC) of the drug on each cell phenotype by HCA to determine the parameter change rate of the specific cell phenotype of the cell, and calculate the TI LEC value of each cell phenotype using the formula TI max = LEC / C LEC .

[0124] The LEC in step 1 is a drug concentration that causes a parameter change rate of a cell phenotype to be greater than or equal to 25%.

[0125] The specific cell phenotype in step 1 includes LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, and GSH_24h, or

[0126] The specific cell phenotype includes LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1a_3h, Nrf2_6h, Hif1a_24h, and Nrf2_24h.

[0127] In step 2, the TI value of each cell phenotype of the drug to be tested is input into the drug liver toxicity prediction model constructed by the present disclosure to determine whether the drug to be tested has liver toxicity, or LEC In step 3, the TI value of each cell phenotype of the drug to be tested is input into the drug liver toxicity integrated prediction model constructed by the present disclosure, and m prediction values (liver toxicity positive or negative) are obtained according to the results output by the m drug liver toxicity prediction models. If more than 50% of the m prediction values are determined to be liver toxicity positive, it is determined that the drug to be tested has liver toxicity.

[0128] In step 2, the TI value of each cell phenotype of the drug to be tested is input into the drug liver toxicity prediction model constructed by the present disclosure to determine whether the drug to be tested has liver toxicity, or LEC In step 3, the TI value of each cell phenotype of the drug to be tested is input into the drug liver toxicity integrated prediction model constructed by the present disclosure, and m prediction values (liver toxicity positive or negative) are obtained according to the results output by the m drug liver toxicity prediction models. If more than 50% of the m prediction values are determined to be liver toxicity positive, it is determined that the drug to be tested has liver toxicity.

[0129] Figure 3 A structural schematic diagram of one embodiment of the system for predicting drug liver toxicity of the present disclosure. As shown in Figure 3 The system for predicting drug liver toxicity includes a high content analysis (HCA) instrument, a calculation module, and a prediction module. Wherein:

[0130] The high content analysis (HCA) instrument is used to determine the parameter change rate of the cell phenotype of the cells after being treated with different concentrations of the drug to be tested;

[0131] The calculation module is used to determine the lowest effective concentration (LEC) of the drug, and calculate the TI value using the formula TI LEC = LEC / C max LEC

[0132] The prediction module includes the drug liver toxicity prediction model or the drug liver toxicity integrated prediction model constructed by the present disclosure, which is used to predict the liver toxicity of the drug to be tested.

[0133] ​​Figure 4 Structure diagram of one embodiment of the device for predicting drug-induced liver injury. As shown in Figure 4 the device for predicting drug-induced liver injury comprises a memory 41 and a processor 42.

[0134] The memory 41 is configured to store instructions, and the processor 42 is coupled to the memory 41, and the processor 42 is configured to execute the method according to any one of the embodiments of the device for predicting drug-induced liver injury based on the instructions stored in the memory. Figure 2

[0135] As shown in Figure 4 the device for predicting drug-induced liver injury further comprises a communication interface 43 configured to interact with other devices. Meanwhile, the device for predicting drug-induced liver injury further comprises a bus 44, and the processor 42, the communication interface 43, and the memory 41 complete communication with each other through the bus 44.

[0136] The memory 41 can include a high-speed RAM memory, and can further include a non-volatile memory, such as at least one disk memory. The memory 41 can also be a memory array. The memory 41 can also be divided into blocks, and the blocks can be combined into a virtual volume according to a certain rule.

[0137] In addition, the processor 42 can be a central processing unit CPU, or can be an application-specific integrated circuit ASIC, or one or more integrated circuits configured to implement the embodiments of the present disclosure.

[0138] The present disclosure also relates to a computer readable storage medium, wherein the computer readable storage medium stores computer instructions, and the instructions are executed by a processor to implement the method according to any one of the embodiments of the device for predicting drug-induced liver injury. Figure 2

[0139] The present disclosure also relates to the use of a specific cell phenotype combination in predicting drug-induced liver injury, wherein the specific cell phenotype combination comprises: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h, or

[0140] The specific cell phenotype combination comprises: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h.

[0141] ​​For example, the specific combination of cell phenotypes can be a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, and GSH_24h.

[0142] For another example, the specific combination of cell phenotypes can be a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1a_3h, Nrf2_6h, Hif1a_24h, and Nrf2_24h.

[0143] In predicting drug hepatotoxicity, cells are treated with different concentrations of drugs, the rate of change of each cell phenotype parameter in the specific combination of cell phenotypes is determined by high content analysis (HCA), the lowest effective concentration (LEC) of each cell phenotype is determined, the TI LEC = LEC / C max value of each cell phenotype is calculated, and the hepatotoxicity of a drug is predicted using the TI LEC value of each cell phenotype parameter in the specific combination of cell phenotypes. LEC

[0144] The present disclosure is described in detail below with reference to a specific embodiment.

[0145] Unless otherwise specifically indicated, the molecular biology experimental methods and immunological detection methods used in the present disclosure are performed essentially according to the methods described in J. Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, 1989, and F. M. Ausubel et al., Short Protocols in Molecular Biology, 3rd Ed., John Wiley & Sons, Inc., 1995; the use of restriction enzymes is in accordance with the conditions recommended by the product manufacturer. Unless otherwise specified in the examples, the procedures are performed under conventional conditions or under conditions recommended by the manufacturer. Unless otherwise specified, the reagents or instruments used are conventional products that can be obtained commercially. Those skilled in the art know that the examples describe the present disclosure by way of example and are not intended to limit the scope of the present disclosure as claimed.

[0146] Example: Construction of a drug hepatotoxicity prediction model

[0147] ​The present embodiment is based on a high-content technology analysis platform of cell phenotype parameters, HepG2 cells are used as experimental test models, 223 marketed drugs of four different liver injury degrees (severe, moderate, undefined, non-toxic) defined by standard liver injury severity are used as research objects, which are divided into two groups of test set (120 drugs) and verification set (103 drugs), the influence of drug treatment on hepatocyte at different time points on whole cell toxicity related phenotype (a total of 23) characteristics is quantitatively determined, a drug hepatocyte toxicity phenotype spectrum database is established, and then combined with the maximum blood drug concentration (maximum total concentration, Cmax), a DILI related hepatocyte phenotype feature spectrum is constructed. The Tclass classification recognition software is used, combined with the clinical liver injury degree of the test set drugs and the in vitro hepatocyte phenotype spectrum feature data, the stepwise workflow and ROC regression analysis are used to identify and verify the best cell phenotype combination for predicting sDILI drugs. The best cell phenotype combination is used to construct a DILI integrated prediction model. Finally, the cell phenotype of the verification set drugs is used for verification, and the DILI prediction combination and prediction model are finally determined. The specific workflow is shown in Figure 5

[0148] I. Experimental materials

[0149] 1.1 Main drugs and reagents:

[0150] 1) Drugs

[0151] 223 test drugs were purchased from Selleck Company in the United States and EFEBIO Company in China (FDA-approved Drug Library, L1300). Tool drugs 2,2'-bipyridyl (2,2'-Bipyridyl, BP), tunicamycin (tunicamycin, TM), tertiary butylhydroquinone (Tertiary butylhydroquinone, TBHQ), hydroxychloroquine (Hydroxychloroquine, HCQ), etoposide (Etoposide, Ept) were purchased from Sigma Company in the United States; Tool drug interleukin-1β (Interleukin-1β, IL-1β) was purchased from PEPROTECH Company in the United States.

[0152] 2) Fluorescent dyes and antibodies

[0153] ​Dyes: Cell nucleus dye Hoechst 33342, microfilament F-actin dye Alexa Flour 488-phalloidin, mitochondria membrane potential dye Mito Tracker Red CMXRos, lysosome pH dye Lyso Tracker deep Red, GSH dye CM-H2DCFDA, dead cell nucleus dye TOTO-3, glutathione dye mBCI were purchased from Life Technologies, USA.

[0154] Antibodies: Rabbit anti-LC3B monoclonal antibody, mouse anti-a-tubulin monoclonal antibody, mouse anti-MnSOD monoclonal antibody, mouse anti-pH2AX monoclonal antibody, and Alexa Fluor 488-labeled donkey anti-mouse IgG secondary antibody, Alexa Fluor 549-labeled donkey anti-mouse IgG secondary antibody, Alexa Fluor 549-labeled donkey anti-rabbit IgG secondary antibody were purchased from Life Technologies, USA.

[0155] 3) Cell culture related reagents

[0156] DMEM high glucose medium, RPMI1640 basal medium, F12 medium, fetal bovine serum (FBS) and Hank's Balanced Salt Solution (HBSS) were purchased from Thermo Scientific HyClone, USA; bovine serum albumin (BSA) was purchased from Sigma-Aldrich, USA; G418, L-Glutamine and HEPES buffer were purchased from Life Technologies, USA. Trypsin was purchased from Merk, USA; DMSO solution, formaldehyde solution, Triton X-100, penicillin, streptomycin and other conventional chemical reagents were domestic reagents.

[0157] 1.2 Main instruments and consumables

[0158] High content imaging system (In Cell Analyzer 1000 or In Cell Analyzer 2000) and analysis workstation (In Cell Analyzer Workstation 3.7.2) were purchased from GE healthcare lifescience, USA.

[0159] Black-walled 96-well plates were purchased from Corning, USA.

[0160] 1.3 Cell lines

[0161] HepG2 cells (provided by the Academy of Military Medical Sciences of the Chinese People's Liberation Army), Hif1a-EGFP_CHO (CHO cell line stably expressing HiF1a-EGFP fluorescent protein), P65-EGFP_CHO (CHO cell line stably expressing NF-κB-EGFP fluorescent protein), and ATF6-EGFP_U2OS (U2OS cell line stably expressing ATF6-EGFP fluorescent protein) cell strains were purchased from Thermo Fisher Scientific, USA; Nrf2-EGFP_A549 (A549 cell line stably expressing Nrf2-EGFP fluorescent protein) cell strain was constructed according to the literature method (Angela Schoolmeesters, Daniel D Brown, Yuriy Fedorov. Kinome-wide functional genomics screen reveals a novel mechanism of TNFa-induced nuclear accumulation of the HIF-1a transcription factor in cancer cells. PLoS One. 2012; 7(2): e31270. doi: 10.1371 / journal.pone.0031270).

[0162] II. Experimental methods

[0163] 2.1 Selection, classification, and grouping of test drugs

[0164] 223 test drugs were selected from the NIH Pubmed database, the liver toxicity website (http: / / www.livertox.nih.gov) livertox database, the Chinese liver toxicity professional website (http: / / www.hepatox.org / ) hepatox database, and the laboratory drug entity sub-library. The SMILES expression, liposolubility coefficient logP value, indications, daily drug dosage, C max , drug target, and other information of the test drugs were collected from the drug bank database (https: / / www.drugbank.ca / ).

[0165] Currently, there are two major systems for DILI injury classification. One is the livertox database, which classifies the degree of drug-induced liver injury into A to E according to the "likelihood score" of liver injury. The levels are defined as definitely possible, highly possible, very possible, possible and not possible, respectively. The other is the DILIrank database established by Professor Minjun Chen et al. based on the drug usage instructions published by FDA. The drugs in the database are divided into four categories: DILI most relevant (Most-DILI concern), DILI less relevant (Less-DILI concern), DILI not relevant (No-DILI concern) and ambiguous (Ambiguous-DILI concern) (Chen M, Suzuki A, Thakkar S, Yu K, Hu C, Tong W. DILIrank: the largest reference drug list ranked by the risk for developing drug-induced liver injury in humans [J]. Drug Discovery Today 2016; 21(4): 648-53; Chen M, Vijay V, Shi Q, Liu Z, Fang H, Tong W. FDA-approved drug labeling for the study of drug-induced liver injury [J]. Drug Discov Today 2011; 16(15-16): 697-703.). This application takes into account the advantages of the above two classification systems, and divides the selected test drugs into four categories: severe DILI (sDILI), moderate DILI (mDILI), ambiguous DILI (aDILI) and non-DILI (nDILI). The specific classification criteria are shown in Table 1. Among the selected drugs, there are 81 sDILI, 81 mDILI, 30 aDILI and 31 nDILI drugs.

[0166] In addition, according to the modeling and verification requirements, this study randomly divides 223 test drugs into test set and validation set. The test set contains 120 drugs, of which sDILI, mDILI and nDILI are 50, 50 and 20, respectively; the validation set contains 31 sDILI, 31 mDILI, 30 aDILI and 11 nDILI drugs. The drug names in the test set and the validation set are shown in Tables 2 and 3.

[0167] The test drugs involve antibiotics, non-steroidal anti-inflammatory drugs, antipsychotic drugs (mainly anti-depression and anti-mania), anti-tumor drugs, anti-viral drugs, cardiovascular disease drugs, lipid-lowering drugs, anti-fungal drugs and hormone drugs, etc. The proportion of each type of drug is shown in Table 1. Figure 6 The types of liver injury, logP value, dosage and C max value distribution of the two groups of drugs are shown in Table 2, Table 3 and Table 4, respectively. Figure 6 It can be seen from Table B and C that the above characteristics of the two groups of drugs in the experiment are basically balanced, and there is no obvious bias.

[0168] Table 1. Classification and criteria of liver injury degree of test drugs

[0169]

[0170] Table 2. List of test set drugs

[0171]

[0172]

[0173] Table 3. List of validation set drugs

[0174]

[0175]

[0176] 2.2 Cell culture and drug treatment

[0177] 2.2.1 Preparation of drug solution

[0178] The drug dry powder was dissolved in DMSO solution to prepare a 10-30 mM storage solution, which was aliquoted and stored in a -20°C refrigerator for standby. In each experiment, the corresponding drug storage solution was diluted with culture medium to prepare a 3x final concentration working solution.

[0179] 2.2.2 Cell line culture

[0180] Hif1a-EGFP_CHO and P65-EGFP_CHO cells were cultured in F12 complete medium (containing 10% fetal bovine serum (FBS), 100 kU / L penicillin and streptomycin, ampicillin and 0.5 mg / ml geneticin (G418)), ATF6-EGFP_U2OS and Nrf2-EGFP_A549 cells were cultured in DMEM complete medium (containing 10% FBS, 100 kU / L penicillin and streptomycin, ampicillin, 0.5 mg / ml G418 and 2 mM L-glutamine), HepG2 cells were cultured in RPMI 1640 complete medium (containing 10% FBS, 100 kU / L penicillin and streptomycin, ampicillin) in a cell incubator at 37 °C, 5% CO2. Analytical medium was used for ATF6-EGFP_U2OS and P65-EGFP_CHO cell lines when they were subjected to drug testing, which were DMEM medium containing 1% FBS and F12 medium containing 0.1% FBS, 10 mM HEPES, respectively.

[0181] 2.2.3 Drug incubation treatment of HepG2 cells

[0182] HepG2 cells were cultured in RPMI 1640 complete medium (containing 10% FBS, 100 kU / L penicillin, streptomycin and ampicillin) in a cell incubator at 37 °C, 5% CO2. Before the experiment, the HepG2 cells growing close to confluence were digested into single cells, diluted with complete medium, and seeded in black-sided bottom transparent 96-well plates at an appropriate density. The seeding densities for 16, 24, 72 h drug treatment were 8 x 10 3 , 8 x 10 3 , 4 x 10 3 / 100 μL / well, respectively. After 18-24 h of cell seeding culture, the HepG2 cells were replaced with 100, 100, 150 μL / well of fresh medium at 16, 24, 72 h before the addition of the test drug, respectively, and then 50, 50, 75 μL / well of 3x final concentration of the test drug working solution was added, respectively, for incubation in the incubator. The incubation time of the test drug and the positive control drug used in different test combinations are shown in Table 4.

[0183] Table 4. Experimental detection conditions of HepG2 cell phenotype parameters

[0184]

[0185] 2.2.4 Drug incubation treatment of stress response cells

[0186] Hif1a-EGFP_CHO and P65-EGFP_CHO cells were cultured in F12 complete medium (containing 10% fetal bovine serum (FBS), 100 kU / L penicillin / streptomycin and 0.5 mg / ml G418) and ATF6-EGFP_U2OS and Nrf2-EGFP_A549 cells were cultured in DMEM complete medium (containing 10% FBS, 100 kU / L penicillin / streptomycin, 0.5 mg / ml G418 and 2 mM L-glutamine) at 37 °C in a 5% CO2 cell incubator. Before the experiment, the cells were seeded in black-walled bottom transparent 96-well cell culture plates at appropriate densities according to the experimental purpose. In the short-term drug treatment experiment, the seeding densities of P65-EGFP_CHO, Hif1a-EGFP_CHO, ATF6-EGFP_U2OS and Nrf2-EGFP_A549 cells were 1 x 10 4 , 1 x 10 4 , 8 x 10 3 , 8 x 10 3 / 100 μL / well, respectively. In the 24 h drug treatment experiment, the seeding densities of the cells were 8 x 10 4 , 8 x 10 4 , 6 x 10 3 / 100 μL / well, respectively. After 18-24 h of cell culture, the Hif1a-EGFP_CHO and Nrf2-EGFP_A549 cells were replaced with 100 μL / well of fresh culture medium, the ATF6-EGFP_U2OS cells and P65-EGFP_CHO cells were replaced with 100 μL / well of assay medium, and 50 μL / well of 3x final concentration of test drug / positive control drug working solution and solvent (containing culture medium with the same concentration of DMSO as the working solution) were added according to the test requirements of each cell line, respectively, and then the cells were incubated in the cell incubator for the corresponding time. The specific experimental conditions of different cell lines are shown in Table 5.

[0187] Table 5. Experimental conditions of stress response pathway cell lines

[0188]

[0189] 2.3 Fluorescent labeling of cell phenotype parameters

[0190] 2.3.1 Multi-parameter fluorescent labeling of HepG2 cells

[0191] Directly after drug treatment of cells, fluorescence labeling method can refer to Long L, Li W, Chen W, Li FF, Li H, Wang L. Dynamic Cytotoxic Profiles of Sulfur Mustard in Human Dermal Cells Determined with Multiparametric High-Content Analysis [J]. Toxicology Research 2016; 5(2): 583-93.

[0192] 1) Fluorescence labeling of nucleus, microtubule-associated protein 1 light chain 3B (LC3B) protein, lysosome and phosphorylated histone (pH2AX)

[0193] After drug treatment of cells for 16 h, 50 L / well of culture solution containing nucleus (4 M Hoechst 33342) and lysosome (200 nM LysoTracker deep Red) dyes was directly added, and incubated in a cell incubator for 30 min; 100 L / well of fixing solution (PBS containing 12% formaldehyde) was directly added for room temperature fixation for 20 min; the solution was discarded, 200 L / well of membrane permeabilization solution (PBS solution containing 0.1% Triton X-100) was added for membrane permeabilization for 30 min; the solution was discarded, 200 L / well of PBS solution was used for washing once, and 200 L / well of blocking solution (PBS solution containing 5% BSA) was added for room temperature incubation for 1 h; the solution was discarded, 40 L / well of blocking solution containing mouse anti-pH2AX monoclonal antibody (1:1000 dilution of blocking solution) was added, and incubated at 4°C overnight in the dark; the solution was discarded, 100 L / well of blocking solution was used for washing three times, and 40 L / well of blocking solution containing rabbit anti-LC3B monoclonal antibody (1:1000 dilution of blocking solution) was added, and incubated at room temperature for 1 h; the solution was discarded, 50 L / well of blocking solution containing Alexa Flour 488-labeled donkey anti-mouse IgG secondary antibody and Alexa Flour 549-labeled donkey anti-rabbit IgG secondary antibody (both were 1:500 dilution of blocking solution) was added, and incubated at room temperature in the dark for 1 h; the solution was discarded, 200 L / well of PBS solution was used for washing three times, and finally 200 L / well of PBS solution was added, and detected and analyzed on a machine.

[0194] 2) Fluorescence labeling of nucleus, mitochondrial membrane potential (MMP), manganese superoxide dismutase (MnSOD) and nuclear membrane permeability (NMP)

[0195] The specific experimental operation steps are as described in 1). After the drug exposure treatment of the cells for 24 h, 50 L / well of culture solution containing cell nuclei, mitochondria (2 M MitoTracker Red CMXRos) and dead cell nuclei (0.4 M TOTO-3) dyes was added, and the cells were incubated in a cell incubator for 30 min; fixation was performed for 20 min, membrane permeation was performed for 30 min, and blocking was performed for 1 h; mouse anti-MnSOD monoclonal antibody-containing blocking solution (1:500 blocking solution dilution) and Alexa Flour 488-labeled donkey anti-mouse IgG secondary antibody-containing blocking solution (1:500 blocking solution dilution) were used for labeling of MnSOD, respectively; and finally, the labeled cells were placed in 200 L of PBS solution for machine detection and analysis.

[0196] 3) Cell nucleus, glutathione (GSH) and cytoskeleton fluorescence labeling

[0197] After drug incubation for 24 h, 100 L / well of HBSS solution was used for washing once, and then 100 L / well of HBSS solution containing cell nuclei and GSH dyes (cell nucleus dye 10 mM Hoechst 33342, 1 L per 10 mL of HBSS solution, and GSH dye 1 mM CM-H2DCFDA, 10 L per 10 mL of HBSS solution) was added, and the cells were incubated at 37°C for 45 min; room temperature fixation solution (HBSS solution containing 0.1% Triton X-100) was used for fixation for 20 min, and the cells were washed once with HBSS solution, membrane permeation was performed for 30 min, 50 L / well of PBS solution containing 488-labeled phalloidin microfilament dye (27.5 L of Alexa 488-Phalloidin methanol stock solution was dissolved in 5.5 mL of PBS solution) was added, and the cells were incubated at room temperature in the dark for 1 h; blocking was performed for 1 h, and mouse a-anti-tubulin monoclonal antibody-containing blocking solution (1:500 blocking solution dilution) and Alexa Flour 549-labeled donkey anti-mouse IgG secondary antibody-containing blocking solution (1:500 blocking solution dilution) were used for labeling of a-tubulin; and finally, the labeled cells were placed in 200 L of PBS solution for machine detection and analysis.

[0198] 4) Cell nucleus, MMP, lysosome and pH2AX detection labeling

[0199] The specific experimental operation steps are as described in 1). After drug exposure treatment for 72 h, 75 L / well of culture solution containing cell nuclei, mitochondria and lysosome dyes was added, and the cells were incubated in a cell incubator for 30 min; fixation was performed for 20 min, membrane permeation was performed for 30 min, and blocking was performed for 1 h; mouse anti-pH2AX monoclonal antibody-containing blocking solution and Alexa Flour 488-labeled donkey anti-mouse IgG secondary antibody-containing blocking solution were used for labeling of pH2AX; and finally, the labeled cells were placed in 200 L of PBS solution for machine detection and analysis.

[0200] 2.3.2 Nuclear labeling of cells in stress response pathway

[0201] After drug treatment, 75 μL / well of room temperature pre-warmed fixation solution (containing 12% formaldehyde in PBS) was added to the four stress response pathway cells, and incubated at room temperature for 20 min. After washing with PBS buffer solution, 1 μM Hochst 33342 in PBS was added, and incubated at room temperature for 1 h, and then subjected to machine detection and analysis.

[0202] 2.4 Image acquisition and analysis

[0203] The high-content imaging system In Cell Analyzer 1000 / 2000 was used to collect cell fluorescence images, and the fluorescence detection channel settings corresponding to each cell phenotype are shown in Table 6; a 20x objective lens was used, and 9 fields of view per well were collected. The Multi Target Analysis module in IN Cell Analyzer Workstation was used to analyze the collected cell images, 3 wells (27 fields of view) per test were analyzed, and not less than 200 cells were analyzed, and the output parameters of each phenotype are shown in Table 6. The specific analysis and expression methods are as follows:

[0204] The test values of each parameter of HepG2 cells were obtained by analyzing the fluorescence intensity, area and number of specific phenotypes in cell fluorescence imaging. The parameter change rate (%) of each cell phenotype affected by the test drug = (drug treatment group - solvent control group) / solvent control group x 100% or (test drug treatment group - solvent control group) / (positive drug treatment group - solvent control group) x 100%.

[0205] P65-EGFP_CHO and ATF6-EGFP_U2OS cells are the activation response mode of transcription factor nuclear translocation. In the resting state of cells, the fluorescently labeled transcription factor protein is distributed in the cytoplasm, and when activated, the fluorescently labeled marker molecule translocates to the nucleus. Therefore, the degree of activation of the pathway can be characterized by quantitatively analyzing the nuclear translocation number of cell fluorescent proteins. The formula for calculating the nuclear translocation coefficient of fluorescent proteins is: R = (fluorescent intensity of fluorescent proteins in the nucleus of test drug treated cells - background fluorescent intensity) / (fluorescent intensity of cytoplasmic fluorescent proteins of test drug treated cells - background fluorescent intensity).

[0206] Hif1a-EGFP_CHO and Nrf2-EGFP_A549 cells are the activation response mode of transcription factor accumulation in cells. In the resting state of cells, the expression of fluorescently labeled proteins is low, and when activated, the expression of fluorescently labeled proteins increases and even moves to the nucleus, showing an increase in fluorescence intensity in cells. Therefore, the degree of activation of the corresponding pathway can be characterized by analyzing the content of nuclear and / or intracellular accumulated fluorescent proteins, and the calculation formula of the accumulation coefficient of nuclear or cellular proteins is: R = fluorescence intensity of fluorescent proteins in the nucleus and / or cells of test drug treated cells - background fluorescence intensity.

[0207] Table 6. Experimental detection conditions of each phenotype parameter

[0208]

[0209]

[0210] 2.5 Data processing and statistical analysis

[0211] 2.5.1 Data processing and main analysis methods

[0212] Excel software was used for data processing, and the experimental results were expressed as the average value and standard deviation (Means ± sd) of 3 replicate holes, and the data were normalized by taking the negative and positive control groups of each 96-well plate as the standard. Cell inhibition rate IR (%) = (control group cell number - test drug treated group cell number) / control group cell number x 100%; Test drug fluorescent protein activation rate (%) = (test drug treated group R value - solvent control group R value) / (positive drug treated group R value - solvent control group R value) x 100%. Cell phenotype parameter change rate (%) = (drug treated group - solvent control group) / solvent control group x 100% or (test drug treated group - solvent control group) / (positive drug treated group - solvent control group) x 100%.

[0213] The EC 50 / IC 50 values were calculated using Origin 6.1 software (Sigmoidal Fit method), and the ROC curve (receiver operating characteristic curve) was drawn to investigate and evaluate the correlation of phenotype parameters and their combinations with sDILI.

[0214] The reliability of the phenotypic parameter high content test method was evaluated by Z' factor (Z' factor = 1 - (3 x standard deviation of positive control - 3 x standard deviation of negative control) / (mean of control - mean of negative control) (see Zhang XD, Espeseth AS, Johnson EN, et al. Integrating experimental and analytic approaches to improve data quality in genome-wide RNAi screens [J]. J Biomol Screen 2008; 13(5): 378-89). In addition, GraphPad Prism 6.0, R language, and Mev software were used to draw scatter plots and columnar charts, heat maps, and cluster maps of experimental results, and One-way ANOVA was used for statistical testing, with p < 0.05 being considered to have a significant difference.

[0215] 2.5.2 Sensitivity, specificity, and accuracy calculation methods

[0216] Two methods were used in the present study to evaluate the constructed liver toxicity prediction method. The first method obtained the sensitivity and specificity of the method through ROC (receiver operating characteristic curve) curve analysis. The closer the curve obtained by ROC curve analysis to the upper left corner of the coordinate axis, the higher the sensitivity and specificity of the method. Another method used a formula to calculate the sensitivity, specificity, and accuracy of the method (Parikh R, Mathai A, Parikh S, Chandra Sekhar G, Thomas R. Understanding and using sensitivity, specificity and predictive values [J]. Indian J Ophthalmol 2008, 56(1): 45-50):

[0217] Sensitivity = number of DILI positive drugs detected / total number of DILI positive drugs x 100%;

[0218] Specificity = number of DILI negative drugs detected / total number of DILI negative drugs x 100%;

[0219] Accuracy = 100% x (number of DILI positive drugs detected + number of DILI negative drugs) / (total number of DILI positive drugs + total number of DILI negative drugs).

[0220] 2.6 Identification and Validation of the Optimal Cell Phenotype Combination for Predicting sDILI Drugs

[0221] The Tclass system is a classification system that integrates Fisher's linear discriminant analysis and feature forward selection (Wuju L, Momiao X. Tclass: tumor classification system based on gene expression profile[J]. Bioinformatics (Oxford, England) 2002; 18(2):325-6). The Fisher's linear discriminant analysis in the Tclass system can also be replaced by Naive Bayes. Bayes classification method. To obtain the optimal phenotypic combination for predicting sDILI drugs, the LEC / Cmax values ​​of cell phenotypes for sDILI and nDILI drugs in the test set were first used as the training dataset. The T-class system was used to find one or more feature combinations with the strongest classification ability. Then, the test set data was randomly divided into two parts in a 3:1 ratio. The larger part of the data was used as the training data to build a classifier, and the smaller part of the data was used as the test data to evaluate the model. This process was repeated 1000 times. By selecting the candidate feature combination with the highest stability index, a DILI ensemble prediction model with 1000 classifiers was established. Finally, the model was validated using the LEC / Cmax values ​​of cell phenotypes of drugs in the validation set to determine the final DILI prediction combination. The specific operation process is as follows: Figure 7 As shown.

[0222] III. Experimental Results

[0223] 3.1 Reliability Analysis of HCA Test Method Based on Cell Phenotype

[0224] HCA assay of cell phenotypes included 11 cell phenotypes (cell count, nuclear, a-tubulin, F-actin, GSH, MnSOD, pH2AX, MMP, NMP, Lysosome, LC3B) and 4 pathways (Hifl a, Nrf2, ATF6, NF-κB) of cell stress response. Since some cell phenotypes might reflect different cell toxicity mechanisms at different time points, we also measured the changes of cell count, nuclear, pH2AX, MMP, Lysosome, Hifl a, Nrf2, NF-κB at different time points. To evaluate whether the HCA method of cell phenotypes is suitable for HTS (high throughput screening), we calculated the Z' factor of each test method, as shown in Table 7.

[0225] The results showed that the Z' factor of each test method was greater than 0.5 except for Nrf2_6h which was 0.43, indicating that the test method based on HCA cell phenotype spectrum used in this experiment was reliable.

[0226] Table 7. Z' factor values of HCA-based cell phenotype test methods

[0227]

[0228]

[0229] 3.2 Determination of cell phenotype spectrum of test set drugs

[0230] First, according to drug solubility and cytotoxicity, 10, 100 μΜ or 3, 30 μΜ drug concentrations were used for primary screening on HepG2 cells and 4 reporter cell lines, respectively. The scatter plot of drug primary screening results was drawn with the cell inhibition rate of the drug in the same test as the horizontal coordinate and the change rate of the phenotype parameter as the vertical coordinate, and the results were as follows Figure 8 、 Figure 9The positive reaction was determined by the rate of change of cell phenotype parameter ≥ 25% (ordinate), and the toxicity threshold was the inhibition rate (IR) (ordinate) ≥ 15%. The results showed that in each phenotype effect, there were two types of positive effects caused by drugs: one was cell toxicity-independent, i.e., specific reaction, that is, the positive change of the phenotype parameter was not accompanied by the appearance of cell toxicity effect (IR < 15%); the other was cell toxicity-dependent, i.e., non-specific reaction, that is, the positive change of the phenotype parameter was accompanied by the appearance of cell toxicity effect (IR ≥ 15%). The screening results of the test set drugs on 20 cell phenotype parameters are shown in Table 8. It can be seen that the number of drugs affecting the same phenotype is different under different cell states (toxic and non-toxic). Some phenotypes have more drugs that cause changes in the early stage of drug treatment, such as Hif1, Actin, and LC3B, suggesting that these phenotype parameters represent the cell disturbance or adaptive response mechanism induced by drugs. In addition, with the extension of the drug treatment time, the number of cell toxicity-dependent cell phenotype changes increases, which is consistent with the result that the cell toxicity increases with the increase of drug treatment time. However, the drugs that cause cell toxicity-dependent phenotype changes are not completely the same as those that cause cell toxicity-independent phenotype changes, indicating that the mechanisms of cell toxicity represented by the phenotype changes at different time points are not completely the same.

[0231] The re-screening standard was the rate of change of cell phenotype parameter ≥ 45%, and the number of drugs re-screened for each phenotype is shown in Table 8, which accounts for more than 50% of the positive drugs in the primary screening. These drugs were selected for re-screening at 5-7 concentrations, and the EC 50 or IC 50 values were calculated. Considering that the number of drugs not entering the re-screening accounts for a large proportion, the lowest effective concentration (LEC) of the drugs was determined based on the primary and re-screening results. In this experiment, the LEC was the concentration at which the rate of change of the phenotype parameter was about 25% (while the cell inhibition rate was 15%). The heat map of the negative logarithm values of EC 50 and LEC (see FIG. A in the Figure 10 ) can be seen. In the EC 50 / IC 50 graph, there were 493 data points that could obtain EC 50 or IC 50 , accounting for only 17.86% of the total data points; and there were 823 data points that could measure LEC, accounting for 29.82% of the total data points. In both graphs, the measurable rates of sDILI, mDILI, and nDILI drugs were comparable, slightly higher than those of nDILI drugs, but the difference was limited, indicating that different types of DILI drugs, including nDILI drugs, can affect multiple cell phenotypes in vitro, and also indicating that it is not possible to distinguish DILI and non-DILI drugs based on the effects of drugs on cell phenotypes in vitro.

[0232] Table 8. Summary of primary screening results of drug cell phenotype profiles of test set

[0233]

[0234] 3.3 Selection of cell phenotype parameters for modeling

[0235] There is a process of in vitro effect extrapolation to in vivo effect for predicting drug human effect in vitro experiment. It is generally believed that the drug in vitro cytotoxicity concentration less than 100 times of the maximum concentration (C max ) exposed in vivo is the safe concentration (Falgun S, Louis L, Barton HA, et al. Setting Clinical Exposure Levels of Concern for Drug-Induced Liver Injury (DILI) Using Mechanistic in vitro Assays [J]. Toxicological Sciences An Official Journal of the Society of Toxicology 2015, 147(2): 500-14; O'Brien PJ, Irwin W, Diaz D, et al. High concordance of drug-induced human hepatotoxicity with in vitro cytotoxicity measured in a novel cell-based model using high content screening [J]. Arch Toxicol 2006, 80(9): 580-604). Therefore, we use C max to correct the in vitro phenotype effect concentration, which is expressed as TI (in vivo Toxicity Index), TI = in vitro phenotype effect concentration / C max . Correspondingly, the TI values corresponding to EC 50 and LEC are expressed as TI 50 and TI LEC , respectively, and the data are plotted as a heatmap as shown in B of Figure 10 . It can be seen that the TI 50 and TI LEC values of sDILI, mDILI and nDILI drugs have strong regularity, the TI values of sDILI drugs are the lowest (mainly yellow), followed by mDILI, and the values of nDILI drugs are higher (mainly blue); the TI 50 and TI LECValue distribution (see Figure 10 C and D) and statistical analysis also confirmed the above facts, and there was a statistical difference between sDILI, mDILI drugs and nDILI drugs. Thus, the in vitro cell phenotype effects combined with C max had a good correlation with the degree of clinical DILI, indicating that these data could be used as the basis for DILI prediction modeling.

[0236] To improve the quality of the basis data for modeling, we further investigated the TI 50 , TI LEC and the correlation of each cell phenotype with DILI. With TI≥100 as the standard for in vivo liver toxicity negative and TI<100 as the standard for in vivo liver toxicity positive, the qualitative determination of each cell phenotype effect of the test set drugs was performed, and the results are shown in Figure 11 A, with red indicating DILI positive. This result is basically consistent with the TI heat map Figure 10 C and D), whether TI 50 or TI LEC is used as the basis data to determine in vivo liver toxicity positive, sDILI drugs are the most, mDILI is the second, and nDILI is the least, and sDILI and mDILI classes have a statistical difference compared with nDILI class (see Figure 11 B and E); on this basis, first, we compared the specificity of predicting clinical DILI based on TI 50 , TI LEC data using all phenotype data, and the results are shown in Figure 11 D and G, it can be seen that based on TI 50 data, the sensitivity for predicting sDILI and mDILI classes is low, and the specificity is high (80%), while based on TI LEC data, the sensitivity is high, and for sDILI class drugs, it can be as high as 90%, but the specificity is significantly reduced to only 39%; further investigation of the specificity of all parameters found that based on TI LEC data for prediction, except for the specificity of phenotype parameters IR_72h, F-actin_24h and MMP_24h, which were 70%, 60% and 75% respectively, the false positive rate was high, and the specificity of other parameters was 95% and above (see Figure 11 G). It is suggested that these three parameters are too sensitive when TI LEC data is used for prediction. Accordingly, we removed these three parameters and used the data of the remaining parameters to perform ROC curve analysis of the phenotype and DILI, and the results are shown in Figure 12 and Table 9, compared with TI 50 data, the optimized TI LECThe ROC prediction method has a larger AUC value, indicating higher sensitivity and specificity. Therefore, in this experiment, the three parameters IR_72h, F-actin_24h, and MMP_24h will be removed, and the TI of the other 20 cell phenotypic parameters will be used. LEC The values ​​were used as modeling data. The 20 cell phenotypic parameters were: 1, nuclear_72h; 2, MMP_72h; 3, lysosome_72h; 4, pH2AX_72h; 5, nuclear_24h; 6, α-tubulin_24h; 7, GSH_24h; 8, MnSOD_24h; 9, NMP_24h; 10, NF-κB_24h; 11, Hif1α_24h; 12, Nrf2_24h; 13, nuclear_16h; 14, LC3B_16h; 15, Lysosome_16h; 16, pH2AX_16h; 17, NF-κB_0.67h; 18, Hif1α_3h; 19, ATF6_5h; 20, Nrf2_6h.

[0237] Table 9. Cell phenotypic parameters TI 50 and TI LEC ROC analysis results of values ​​and clinical DILI

[0238]

[0239] 3.4 Identification of the optimal phenotypic test combination and construction of the prediction model

[0240] To obtain a combination of cell phenotypic test results and predictive models with optimal predictive efficacy, practicality, and convenience, we used TI data from 20 phenotypic parameters of two classes of drugs with the most clearly defined DILI injury effects: sDILI and nDILI (negative liver injury). LEC The values ​​were used as training data for modeling (including 50 sDILI samples and 20 nDILI samples), and a T-class classification system was adopted, using Fisher linear discriminant analysis or Naive Bayes. Bayesian classification, combined with feature-forward selection, is used to identify and classify cell phenotypic parameters. Samples are randomly divided into a training set (e.g., 75% of the samples) and a test set. The number of correctly classified samples in each set is calculated, along with accuracy (Train_ac) and stability (Test_ac). Accuracy is defined as the percentage of correctly classified samples in the training set; stability is defined as the percentage of correctly classified samples in the test set. A series of combinations with the highest training accuracy (Train_ac) and detection stability (Test_ac) are selected, such as... Figure 13Test_ac values of Test Combination 1 consisting of 13 parameters were the highest, and then the accuracy decreased with the decrease of the number of parameters, and the stability remained the highest level with 5 parameters, but the accuracy reached nearly 84% with 7 parameters or more; ROC and statistical analysis (see Figure 13 B) showed consistent trends (Table 10), and the AUC curves of Test Combination 1 and 2 consisting of 13 and 11 parameters were the highest and almost the same, and the sensitivity and specificity of both were the same, being 86% and 90%, respectively; and the specificity of 9, 7, 5, and 3 parameter combinations could all reach 90%, but the sensitivity gradually decreased, being 84%, 82%, 78%, and 74%, respectively. In addition, according to the convenience of multi-parameter testing, as shown in Table 10A, the same color test indicators were implemented in the same test, and Combination 4 (7 parameters) consisted of only three independent test experiments, while Combination 3 (9 parameters) required 4 experiments, and Combination 1 (13 parameters) and Combination 2 (11 parameters) both consisted of 6 test experiments; therefore, we determined Combination 1 and Combination 4 as the best DILI prediction test combinations to meet different sensitivity requirements. Figure 13

[0241] From the samples, 75% were randomly extracted as the training set, and the remaining 25% were randomly extracted as the test set, and 1000 times of random extraction were performed to obtain 1000 different distributions of the “training” and “test” sets, and 1000 classifiers were constructed by using Fisher linear discriminant analysis model or naive Bayes (Bayes) classification model, and the 1000 classifiers were integrated into a DILI integrated prediction model, and the experimental results obtained by the unknown drug on the above determined test combination were input into the model, and according to the results output by the 1000 classifiers, the DILI potential of the unknown drug could be predicted. According to the results output by the 1000 classifiers, 1000 prediction values (positive or negative) were obtained, if there were 500 or more positive values in the 1000 prediction values, the sample was judged to be positive, and the probability was P / 1000 (P value was the number of prediction values being positive), otherwise it was negative.

[0242] Table 10. ROC analysis results of different cell phenotype parameter combinations

[0243]

[0244] Note: CI is confidence interval.

[0245] 3.5 Validation of the prediction model

[0246] ​​To verify the effectiveness of the above constructed model, the preferred phenotype test combination was used to determine the validation set drugs (103), and the primary screening and secondary screening were carried out, respectively, and the results are shown in Table 11 and Figure 14 , and the corresponding EC 50 , LEC were obtained, and the human exposure C max was combined to obtain the TI LEC value. The LEC and TI LEC value heat map of different damage types of drugs in the validation set is shown in A and B of Figure 15 , and the TI LEC value is correlated with the DILI damage degree. We input the TI LEC value of the validation set drugs into the DILI integrated prediction model to judge the DILI positive parameters, and the results are shown in C of Figure 15 , and the sDILI drug liver toxicity positive parameters are significantly more than the mDILI and aDILI drugs, and the nDILI drug has no liver toxicity positive parameters. Further, the sensitivity, specificity and accuracy of the validation set drugs tested by the 2 best test combination discs 1 and 4 were calculated, and the results are shown in Table 12. The specificity of test combination discs 1 and 4 is 100%, and the accuracy is 88.1% and 85.7%, respectively; in addition, the sensitivity of test combination disc 1 for sDILI, mDILI and aDILI drugs in the validation group is 83.87%, 54.84% and 66.67%, respectively, while combination disc 4 also reaches 80.65%, 51.61% and 63.33%, respectively; the prediction performance is comparable to that of the test combination used in modeling. It shows that the method established in the present disclosure has good prediction ability and repeatability, and is successful.

[0247] Table 11. Summary of primary screening of validation set drugs in the best test combination disc

[0248]

[0249]

[0250] Table 12. Comparison of prediction performance of the best test combination for test set and validation set drugs

[0251]

[0252] 3.6 Comprehensive evaluation of the prediction model

[0253] The accuracy of the test combination 1 of the present disclosure for predicting the sDILI and mDILI classes of the test set and the validation set is 85.7-88.1%, the specificity is 90-100%, and the sensitivity is between 83.87-84% and 54-54.84%, respectively; the combination 4 is slightly lower, but the specificity is the same as the combination 1, the accuracy of the test combination 1 of the present disclosure for predicting the sDILI and mDILI classes of the test set and the validation set is 82.86-85.7%, the sensitivity is between 80-80.65% and 48-51.61%, respectively. In order to comprehensively evaluate the ability of the prediction model, we combined all the test drugs and investigated the sensitivity of the prediction method for different types of drugs. Figure 16 The distribution of DILI positive parameters based on the test data of combination 1 and combination 4 is shown, and it can be seen that the present prediction method has strong correlation with the DILI damage type. The statistical results show that when predicting based on combination 1 (13 parameters), the sensitivity of sDILI, mDILI and aDILI drugs is 83.95%, 54.32% and 66.67%, respectively, the specificity is 93.55%, the accuracy is 86.61%, and the sensitivity of iDILI category drugs is 70.97%; when predicting based on combination 4 (7 parameters), the sensitivity of sDILI, mDILI and aDILI drugs is 80.25%, 49.38% and 63.33%, respectively, the specificity is 93.55%, the accuracy is 83.95%, and the sensitivity of iDILI category drugs is 61.29%, which is only slightly lower than the result of combination 1.

[0254] The RO2 principle refers to a daily dose ≥ 100 mg / day and logP ≥ 3 (Chen M, Tung CW, Shi Q, et al. A testing strategy to predict risk for drug-induced liver injury in humans using high-content screen assays and the 'rule-of-two' model[J]. Arch Toxicol 2014;88(7):1439-49). To examine whether RO2 helps improve the prediction accuracy of this model, RO2 was included as an independent parameter in the prediction method, and the impact of combining RO2 on prediction performance was compared. The results are shown in Table 13. The RO2 combined with test combination 1 slightly improved the sensitivity of prediction for various drugs, such as 1.24%, 8.64%, and 3.33% for sDILI, mDILI, and aDILI, respectively. However, the effect on iDILI was more significant, increasing by 12.9%, but the specificity and accuracy decreased. The results of the RO2 combined with test combination 4 were similar. It is evident that including RO2 as an independent factor in the prediction parameter combination has little impact on the accuracy of the prediction, suggesting that the optimal test combination for DILI drugs obtained based on the cytotoxic full phenotypic profile assay already covers the RO2 characteristics of DILI drugs. Given the extreme harm of iDILI, and the fact that the inclusion of RO2 can significantly improve the sensitivity of the prediction method when evaluating iDILI drugs, the RO2 value of the drug can be incorporated into the prediction of iDILI drugs.

[0255] Table 13. Comparison of prediction methods based on optimal test combinations 1 and 4 and those combined with RO2.

[0256]

[0257] Furthermore, this disclosure also examines the sensitivity of a combination 1 test-based predictive model to drugs for different types of liver injury. The results are as follows: Figure 17 As shown, the drug sensitivity was highest (83.7%) for cholestasis, hepatocellular injury, and mixed injury types, followed by hepatocellular injury type (75.9%) and unknown type (73.1%). The drug sensitivities for cholestasis and mixed injury types, cholestasis type, and cholestasis and hepatocellular injury type were 60%, 58.3%, and 54.8%, respectively. These results indicate that the prediction method established in this experiment has strong predictive ability for cholestasis, hepatocellular injury and mixed injury types, and hepatocellular injury type.

[0258] 3.7 Conclusion:

[0259] In summary, the present disclosure, through the analysis of 223 drug cytotoxic full-phenotype profiles of basically clear clinical liver injury types, combined with human exposure C max , using HCA, machine learning-based TCLASS classification recognition system, ROC analysis and other technologies and methods, through modeling, verification and comprehensive analysis, innovatively established an in vitro liver toxicity prediction system based on specific cell phenotype combination (specific cell phenotype pattern) analysis, specifically constructed an integrated classification model for drug liver toxicity prediction composed of 7 or 13 cell phenotype parameter test combinations and 1000 classifiers. The sensitivity, specificity and accuracy of the method are 84%, 94%, 87% respectively. The prediction method can be completed at most by 3 groups of cell phenotype parameters (LC3B + pH2AX, Nuclear + MnSOD + GSH + α-tubulin, Nuclear + MMP) and 3 stress pathways (Hif1a, ATF6 and Nrf2) detection; when combined with the RO2 principle, the sensitivity of the prediction of IDILI drugs can be as high as 84%, and the method of the present disclosure realizes the breakthrough of in vitro cell test-based prediction of DILI, especially iDILI prediction.

[0260] Although the specific embodiments of the present disclosure have been described in detail, those skilled in the art will understand that various modifications and changes can be made to the details in accordance with all the teachings disclosed, and these changes are all within the protection scope of the present disclosure. The entire disclosure of the present disclosure is given by the appended claims and any equivalents thereof.

Claims

1. A method for constructing a drug hepatotoxicity prediction model, comprising: Collect n known drugs with severe DILI (sDILI) and non-DILI (nDILI), collect the C max information of the drugs; The cells are treated with different concentrations of the drug, the rate of change of the parameter of the specific cell phenotype of the cells is determined by HCA, the lowest effective concentration (LEC) of the drug for each cell phenotype is determined, and the TI LEC value for each cell phenotype is calculated using the formula TI max = LEC / C LEC ; TI of specific cell phenotypes with sDILI and nDILI drugs LEC The machine learning model is trained on values to build a drug hepatotoxicity prediction model, wherein n is an integer greater than or equal to 50, LEC is a drug concentration causing a parameter change rate of a cell phenotype greater than or equal to 25%, and the specific cell phenotype comprises LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, and GSH_24h, or the specific cell phenotype comprises LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, α-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1α_3h, Nrf2_6h, Hif1α_24h, and Nrf2_24h. wherein the machine learning model is a Fisher linear discriminant analysis model or a naïve Bayes classification model.

2. The method of claim 1, wherein, n is 55, 60, 65, 70, or 80.

3. The method of claim 1, wherein the TI values for specific cell phenotypes of sDILI and nDILI drugs are utilized to train a machine learning model. LEC Training the machine learning model includes: 55% to 95% of the n drug samples are randomly extracted as a training set, the machine learning model is trained using the training set, and a drug hepatotoxicity prediction model is constructed.

4. The method of claim 3, wherein the TI values for specific cell phenotypes of sDILI and nDILI drugs are utilized to train a machine learning model. LEC Training the machine learning model includes: 55% to 95% of the n drug samples are randomly extracted as a training set, and m training sets are obtained by randomly extracting m times, the machine learning model is trained using the m training sets, and m drug hepatotoxicity prediction models are constructed, and the m drug hepatotoxicity prediction models are integrated into a drug hepatotoxicity integrated prediction model, wherein m is an integer greater than or equal to 200; and wherein the machine learning model is a Fisher linear discriminant analysis model or a naïve Bayes classification model. 5.The method of any one of claims 1-4, wherein the cells are selected from the group consisting of hepatocytes, P65-EGFP_CHO, Hif1a-EGFP_CHO, ATF6-EGFP_U2OS, Nrf2-EGFP_A549, or any combination thereof.

6. The method of claim 5, wherein, The method has one or more features selected from the group consisting of: (1) the cells are hepatocytes, and the parameter change rate of the cell phenotype = (fluorescence intensity of the drug treatment group - fluorescence intensity of the solvent control group) / fluorescence intensity of the solvent control group × 100%, or the parameter change rate of the cell phenotype = (fluorescence intensity of the test drug treatment group - fluorescence intensity of the solvent control group) / (fluorescence intensity of the positive drug treatment group - fluorescence intensity of the solvent control group) × 100%, wherein the cell phenotype parameter is selected from the group consisting of LC3B, MnSOD, α-Tubulin, GSH, pH2AX, and MMP. (2) the cell is a hepatocyte, and the parameter change rate of the cell phenotype = (the value of the drug treatment group response cell nucleus morphology - the value of the solvent control group response cell nucleus morphology) / the value of the solvent control group response cell nucleus morphology x 100%, or the parameter change rate of the cell phenotype = (the value of the test drug treatment group response cell nucleus morphology - the value of the solvent control group response cell nucleus morphology) / (the value of the positive drug treatment group response cell nucleus morphology - the value of the solvent control group response cell nucleus morphology) x 100%, wherein the cell phenotype parameter is the cell nucleus; (3) the cell is P65-EGFP_CHO and / or ATF6-EGFP_U2OS, and the parameter change rate of the cell phenotype = (the R value of the drug treatment group - the R value of the solvent control group) / the R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (the R value of the test drug treatment group - the fluorescence intensity of the solvent control group) / (the R value of the positive drug treatment group - the R value of the solvent control group) x 100%, wherein R value = (the fluorescence intensity of the fluorescent protein in the cell nucleus of the test drug treatment - the background fluorescence intensity) / (the fluorescence intensity of the fluorescent protein in the cytoplasm of the test drug treatment - the background fluorescence intensity); (4) the cell is Hif1a-EGFP_CHO and / or Nrf2-EGFP_A549, and the parameter change rate of the cell phenotype = (the R value of the drug treatment group - the R value of the solvent control group) / the R value of the solvent control group x 100%, or the parameter change rate of the cell phenotype = (the R value of the test drug treatment group - the fluorescence intensity of the solvent control group) / (the R value of the positive drug treatment group - the R value of the solvent control group) x 100%, wherein R value = the fluorescence intensity of the fluorescent protein in the test drug treatment cell or cell nucleus - the background fluorescence intensity.

7. The method of claim 5, wherein, The cell is a HepG2 cell.

8. A drug liver toxicity prediction model or a drug liver toxicity integrated prediction model, wherein: the drug liver toxicity prediction model is constructed by the method for constructing a drug liver toxicity prediction model according to any one of claims 1-3; the drug liver toxicity integrated prediction model is constructed by the method for constructing a drug liver toxicity prediction model according to claim 4.

9. Use of the drug liver toxicity prediction model or the drug liver toxicity integrated prediction model according to claim 8 in predicting drug liver toxicity.

10. A method for predicting drug liver toxicity, comprising: The cells are treated with different concentrations of the drug to be tested, the rate of change of the parameter of the specific cell phenotype of the cells is determined by HCA, the lowest effective concentration (LEC) of the drug for each cell phenotype is determined, and the TI LEC value of each cell phenotype is calculated using the formula TI max = LEC / C LEC . TI of specific cell phenotypes of the drug under test LEC Input the drug hepatotoxicity prediction model constructed according to any one of claims 1-3, and determine whether the drug to be tested has hepatotoxicity, or TI of specific cell phenotypes of the drug under test LEC The input value is taken into the integrated drug hepatotoxicity prediction model constructed according to claim 4. Based on the output results of m drug hepatotoxicity prediction models, m predicted values ​​are obtained to determine whether the drug's hepatotoxicity is positive or negative. If more than 50% of the m predicted values ​​are determined to be positive for hepatotoxicity, then the drug to be tested is determined to be hepatotoxic. wherein LEC is the drug concentration causing the parameter change rate of the cell phenotype to be greater than or equal to 25%, the specific cell phenotype includes LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, alpha-tubulin_24h, nuclear_72h, and GSH_24h, or The specific cell phenotypes include: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1a_3h, Nrf2_6h, Hif1a_24h, and Nrf2_24h.

11. A system for predicting drug hepatotoxicity, comprising: a high content analysis (HCA) instrument, a calculation module, and a prediction module, The high content analysis (HCA) instrument is used to determine the parameter change rate of the cell phenotype of the cells after being treated with different concentrations of the drug to be tested. The calculation module is configured to determine a lowest effective concentration (LEC) of the drug and to utilize the formula TI LEC = LEC / C max to calculate a value for TI LEC . The prediction module comprises the drug liver toxicity prediction model or the drug liver toxicity integrated prediction model constructed according to any one of claims 1-4, and is used for predicting liver toxicity of the drug to be tested.

12. An apparatus for predicting drug liver toxicity, comprising: a memory configured to store instructions; a processor coupled to the memory, the processor configured to execute the method of claim 10 based on the instructions stored in the memory.

13. A computer readable storage medium, wherein, A computer readable storage medium stores computer instructions, and the instructions are executed by a processor to implement the method of claim 10.

14. Use of a specific combination of cell phenotypes in predicting drug hepatotoxicity, wherein the specific combination of cell phenotypes comprises: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, and GSH_24h, or The specific cell phenotype combination includes: LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1a_3h, Nrf2_6h, Hif1a_24h, and Nrf2_24h.

15. The use of claim 14, wherein, The specific cell phenotype combination is a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, and GSH_24h, or The specific cell phenotype combination is a combination of LC3B_16h, MMP_72h, MnSOD_24h, nuclear_24h, a-tubulin_24h, nuclear_72h, GSH_24h, pH2AX_16h, ATF6_5h, Hif1a_3h, Nrf2_6h, Hif1a_24h, and Nrf2_24h.

16. The use of claim 14, wherein in predicting drug hepatotoxicity, cells are treated with different concentrations of the drug, the rate of change of the parameter of each cell phenotype in the specific combination of cell phenotypes is determined by HCA, the lowest effective concentration (LEC) of the drug for each cell phenotype is determined, the TI value of each cell phenotype is calculated using the formula TI = LEC / C, the TI values of each cell phenotype in the specific combination of cell phenotypes are used to predict hepatotoxicity of the drug. LEC max LEC LEC ​​​​

Citation Information

Patent Citations

  • Predicting toxicity of a compound over a range of concentrations

    US20120219204A1

  • Method of predicting toxicity for chemical compounds

    US20140278130A1