Systems and methods for predicting drug-induced liver injury

By acquiring cellular component datasets and using machine learning models to assess the liver injury risk of drug compounds, the accuracy of DILI risk assessment in the drug discovery phase was addressed, enabling effective assessment of drug safety and optimization of clinical trials.

CN121605480APending Publication Date: 2026-03-03SERALITI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480048961.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-25
Filing Date
2024-07-24
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately assess the risk of drug-induced liver injury (DILI) in human subjects, particularly in the early stages of drug discovery.

Method used

By acquiring cellular component datasets, we determine the pathway values ​​of multiple pathway modules and use logistic regression and neural network models to assess the liver injury risk of compounds. Combined with safety margins and clinical trial design, we provide treatment options and safety assessments for compounds.

Benefits of technology

It improves the predictive accuracy of the risk of drug compounds inducing liver injury in human subjects, guides drug discovery and clinical trial design, and reduces the risk of drug-induced liver injury (DILI).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121605480A_ABST
    Figure CN121605480A_ABST
Patent Text Reader

Abstract

Assays for determining the risk that a compound will induce liver injury in a subject are provided. A data set comprising a plurality of cellular component abundance values for a plurality of cellular components for each of a plurality of cells exposed to a test chemical compound for a first period of time is obtained. The data set is used to determine a pass value for each of a plurality of pass modules. Each module contains an independent subset of the plurality of cellular components. The pass values are input into one or more first models, thereby obtaining one or more first values. The abundance value is input into one or more second models, thereby obtaining one or more second values. The one or more first values and the one or more second values determine the risk that the compound will induce the liver injury.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 515,536, filed July 25, 2023, entitled "Systems and Methods for Predicting Drug-Induced Liver Injury," which is hereby incorporated by reference. Technical Field

[0003] This invention generally relates to systems and methods for identifying compounds that may be associated with drug-induced liver injury. Background Technology

[0004] Drug-induced liver injury (DILI) can lead to acute liver failure and is a common cause of clinical depletion of drug candidates, precautionary warnings, and post-marketing withdrawal of approved drugs. See Kaplowitz, 2004, “Drug-induced liver injury,” *Clin. Infect. Dis.* 38, S44-S48, doi: 10.1086 / 381445; and Senior, 2007, “Drug hepatotoxicity from a regulatory perspective,” *Clin. Liver Dis.* 11, pp. 507-524, doi:10.1016 / j.cld.2007.06.002, each of which is hereby incorporated by reference. Early assessment of the risk of drug-induced liver injury (DILI) is an important aspect of drug development, but it is challenging to perform this assessment prior to clinical trials due to the many complex factors that contribute to liver injury, and the metabolic differences between animals and humans further complicate the process. In practice, attributing the predicted risk of DILI to any given compound during drug discovery is considered arbitrary and difficult. See, for example, Hoofnagle and Björnsson, 2019, “Drug-induced liver injury - Types and phenotypes,” *The New England Journal of Medicine* 381, 264–273. doi:10.1056 / nejmra1816149e, which is hereby incorporated by reference.

[0005] Given the importance of DILI to drug development and the difficulty in attributing DILI-related risks to drug candidates, there is a need in the art for improved systems, methods, and assays to more accurately assess the risk of investigational compounds used for drug discovery potentially inducing liver injury in human subjects. Summary of the Invention

[0006] This disclosure addresses the shortcomings of the aforementioned identification methods. It provides a method for assessing the risk that compounds used in drug discovery may induce liver injury in human subjects.

[0007] In some embodiments, the test chemical compound is an organic compound, particularly a small molecule organic compound with a molecular weight of less than 2000 Daltons. In some embodiments, the test chemical compound is an organic compound with a molecular weight of less than 1000 Daltons. In some embodiments, the test chemical compound is an organic compound that meets each of the criteria in the Lipinski rule of five criteria. In some embodiments, the test chemical compound is an organic compound that meets at least three of the criteria in the Lipinski rule of five criteria. In some embodiments, the test chemical compound is a polypeptide, protein, oligonucleotide, RNA-based therapeutic agent, antibody, or antibody-drug conjugate. In some embodiments, the liver injury is hepatitis, hepatotoxicity, jaundice, liver fibrosis, cirrhosis, hepatomegaly, alcoholic and non-alcoholic fatty liver disease, and / or cholangiocarcinoma. In some embodiments, the liver injury is characterized by elevated one or more liver enzymes, cholestasis, and / or alcoholic stools.

[0008] A cellular component dataset was obtained in electronic form. The dataset contains multiple cellular component abundance values ​​for various cellular components in a first plurality of cells that have been exposed to the test chemical compound for a first duration.

[0009] In some embodiments, each of the plurality of cellular components is a specific gene, a specific mRNA associated with the gene, a carbohydrate, a lipid, an epigenetic feature, a metabolite, a protein, or a combination thereof.

[0010] In some embodiments, the first plurality of cells are primary human hepatocytes. In some embodiments, the abundance values ​​of the plurality of cellular components are determined using single-cell RNA sequencing (scRNA-seq) data of the first plurality of cells. In some embodiments, the abundance values ​​of the plurality of cellular components are determined by batch RNA sequencing. In some embodiments, the abundance values ​​of the plurality of cellular components are determined using single-cell assays of transposase-accessible chromatin (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof. In some embodiments, the plurality of cellular components consists of 100 to 35,000 cellular components.

[0011] The cellular component dataset is used to determine the corresponding pathway value for each of the plurality of pathway modules, thereby obtaining a plurality of pathway values. Each corresponding pathway module contains an independent subset of the plurality of cellular components. Each independent subset of the plurality of cellular components in each corresponding pathway module contains two or more cellular components. Each of the plurality of cellular components is located in at least one independent subset of the plurality of cellular components associated with the pathway module in the plurality of pathway modules. The plurality of cellular components contains 10 or more cellular components. In some embodiments, the plurality of cellular component abundance values ​​contain 5000 or more cellular component abundance values. In some embodiments, the plurality of cellular component abundance values ​​contain 5000, 10,000, 15,000, 20,000, 25,000, or 30,000 or more cellular component abundance values. In some embodiments, each cellular component is a different gene.

[0012] In some embodiments, each independent subset of the multiple cellular components of each corresponding pathway module comprises two to three hundred cellular components. In some embodiments, each corresponding pathway value of each corresponding pathway module is a measure of the central tendency of the abundance of each cellular component in the independent subset of the cellular components corresponding to the corresponding subset of the cellular components of the corresponding pathway across the cellular component dataset. In some embodiments, each corresponding pathway value of each corresponding pathway module is a negative logarithm to base 10 of the false discovery rate, which is determined based on the distribution of the abundance of each cellular component in the independent subset of the cellular components corresponding to the corresponding subset of the cellular components of the corresponding pathway module across the cellular component dataset relative to the abundance of cellular components in the cellular component dataset.

[0013] In some embodiments, the method further includes identifying the plurality of pathway modules by comprising the following process: obtaining one or more first datasets in electronic form, the one or more first datasets comprising: for each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes and collectively represents a plurality of annotated cell states; for each corresponding cell component in the plurality of cell components: the corresponding abundance of the corresponding cell component in the corresponding primary human hepatocyte; identifying the plurality of cell components in the one or more first datasets as a plurality of differentially expressed cell components across the one or more first datasets; and associating each independent subset of the cell components in the plurality of cell components with a corresponding pathway of a pathway module in the plurality of pathway modules.

[0014] In some embodiments, the annotated cell states of the plurality of annotated cell states are cells in the plurality of primary human hepatocytes exposed to the training compound under exposure conditions (e.g., exposure duration, concentration of the training compound, or a combination of exposure duration and concentration of the training compound).

[0015] In some embodiments, the plurality of annotated cell states include those at a variety of different concentrations (e.g., the variety of different concentrations at Ci of the first training compound). max Or estimated C max (ii) IC of the first training compound 10 (within the range) exposed to the first training compound.

[0016] In some embodiments, the plurality of annotated cell states include exposure to the first training compound for multiple different exposure durations. In some embodiments, the plurality of annotated cell states include each being exposed to the same first training compound for the same duration. In some such embodiments, the optimal duration is determined experimentally.

[0017] In some embodiments, the plurality of annotated cell states include corresponding plurality of different concentrations (e.g., the corresponding plurality of different concentrations in (i) the C of the corresponding training compound). max (ii) IC of the corresponding training compound 10 (within the range) exposed to each corresponding training compound in a variety of training compounds.

[0018] In some embodiments, the plurality of annotated cell states include each of the plurality of training compounds being exposed to a plurality of different exposure durations.

[0019] In some embodiments, the plurality of pathway modules consists of 10 to 2000 pathway modules.

[0020] The multiple pathway values ​​are input into one or more first models, thereby obtaining one or more first values ​​as outputs from the one or more first models.

[0021] In some embodiments, each of the one or more first models is a logistic regression model. In some embodiments, each of the one or more first models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

[0022] In some embodiments, the one or more first models include 10 or more first modules, 20 or more first modules, 30 or more first modules, 40 or more first modules, 50 or more first modules, 60 or more first modules, 70 or more first modules, 80 or more first modules, or 90 or more first modules.

[0023] The abundance values ​​of the multiple cell components are input into one or more second models, thereby obtaining one or more second values ​​as outputs from the one or more second models.

[0024] In some embodiments, the one or more second models are a single logistic regression model. In some embodiments, the one or more second models are multiple logistic regression models. In some such embodiments, a first subset of the logistic regression models in the multiple second models is multiple L1 regression models, and a second subset of the logistic regression models in the multiple second models is multiple L2 regression models. In some embodiments, the multiple second models are two to 1000 logistic regression models. In some embodiments, the one or more second models include 10 or more second modules, 20 or more second modules, 30 or more second modules, 40 or more second modules, 50 or more second modules, 60 or more second modules, 70 or more second modules, 80 or more second modules, or 90 or more second modules.

[0025] In some embodiments, each of the one or more second models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

[0026] The first value and one or more of the second values ​​are used to determine the risk that the test chemical compound will induce liver damage in the human subjects.

[0027] In some embodiments, the risk is in the form of the probability that the test chemical compound will induce liver damage in the human subject.

[0028] In some embodiments, the risk is one of a set of listed risk levels at which the test chemical compound will induce liver injury in the human subject. In some such embodiments, the set of listed risk levels comprises low-risk, intermediate-risk, and high-risk levels at which the test chemical compound will induce liver injury in the human subject.

[0029] In some embodiments, the risk is in the form of the lowest concentration of the test chemical compound that induces the liver injury in the human subject, as predicted by the one or more first models and the one or more second models.

[0030] In some embodiments, the method further includes determining a safety margin for the test chemical compound using the lowest concentration of the test chemical compound predicted by the one or more first models and the one or more second models to induce the liver injury in the human subject (e.g., determining the safety margin includes dividing the risk by the Cmax of the test chemical compound).

[0031] In some embodiments, the method informs the selection of one or more human subjects to be treated with the test chemical compound and / or the selection of one or more human subjects to continue or discontinue treatment with the test chemical compound.

[0032] In some embodiments, the method informs one or more human subjects undergoing treatment of the dosage, duration, and / or frequency of administration of the test chemical compound.

[0033] In some embodiments, the method informs the design of a clinical trial that includes the use of the test chemical compound.

[0034] In some embodiments, the method informs the clinical trial of the type and / or level of acceptable adverse events, one or more inclusion criteria, and / or one or more exclusion criteria.

[0035] In some embodiments, the method informs of one or more modifications to the clinical trial while it is in progress.

[0036] In some embodiments, the method informs the design of an adaptive clinical trial, which includes the use of the test chemical compound.

[0037] In some embodiments, the method informs the type and / or level of acceptable adverse events in the adaptive clinical trial, one or more inclusion criteria and / or one or more exclusion criteria.

[0038] In some embodiments, the method informs one or more modifications to the clinical trial while the adaptive clinical trial is in progress.

[0039] In some embodiments, the method informs the classification of the test chemical compound as either a compound suitable for human therapy and / or testing or a compound unsuitable for human therapy and / or testing.

[0040] In some embodiments, the method further includes formulating the test chemical compound for use in a therapy.

[0041] Another aspect of this disclosure provides a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors, the one or more programs containing instructions for performing any of the methods and / or embodiments disclosed herein.

[0042] Another aspect of this disclosure provides a non-transitory computer-readable storage medium whose storage is configured to be executed by a computer, the one or more programs containing instructions for performing any of the methods and / or embodiments disclosed herein.

[0043] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the disclosure are shown and described. As will be appreciated, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious ways, all without departing from this disclosure. Therefore, the drawings and description are to be regarded in an illustrative rather than restrictive manner. Attached Figure Description

[0044] In the accompanying figures, embodiments disclosed herein are shown by way of example and not limitation. Similar reference numerals refer to corresponding parts throughout the figures.

[0045] Figure 1A and Figure 1B Block diagrams of exemplary systems and computing devices according to embodiments of the present disclosure are shown.

[0046] Figure 2A , Figure 2B , Figure 2C , Figure 2D , Figure 2E , Figure 2F , Figure 2G and Figure 2H Together, flowcharts are provided of the processes and features of exemplary methods for assessing the risk of a test chemical compound inducing liver injury in human subjects, according to various embodiments of the present disclosure.

[0047] Figure 3 Transcriptomics from primary human hepatocytes according to embodiments of the present disclosure are demonstrated as being used to predict DILI safety margins to guide hit selection and lead optimization in drug discovery procedures.

[0048] Figure 4 The corresponding pathway values ​​of multiple pathway modules formed from training data according to embodiments of the present disclosure are shown in the form of color-coded heatmaps.

[0049] Figure 5 The performance of a trained ensemble model according to embodiments of the present disclosure in predicting DILI toxicity is illustrated.

[0050] Figure 6 The predicted safety margin of amiodarone calculated using a trained ensemble model according to embodiments of the present disclosure is shown.

[0051] Figure 7 The gene expression characteristics of selected test compounds, measured at each of four different exposure concentrations according to embodiments of the present disclosure, are shown.

[0052] Figure 8A and Figure 8B The pathway values ​​of the pathway modules of the selected test compound, measured at each of four different exposure concentrations according to embodiments of the present disclosure, are shown.

[0053] Figure 9 This demonstrates how erythromycin and labetalol failed to elicit the expression of many DILI markers.

[0054] Figure 10 A graph was plotted showing the relationship between the predicted safety margin (Y-axis) and the actual DILI toxicity (X-axis) of 30 test compounds determined by the integrated model of the present disclosure for toxicity according to embodiments of the present disclosure.

[0055] Figure 11 A graph was plotted showing the relationship between the predicted safety margin (X-axis) of the toxicity of 30 test compounds according to embodiments of the present disclosure, determined by an integrated model, and a known clinical therapeutic index (Y-axis). Detailed Implementation

[0056] introduce.

[0057] In light of the foregoing background, this disclosure provides a method for determining the risk that a compound will induce liver injury in a subject. Multiple cellular component abundance values ​​are obtained from multiple cells that have been exposed to a test chemical compound for a first time period. This data is used to determine a pathway value for each of multiple pathway modules. Each such module contains an independent subset of the multiple cellular components. The pathway values ​​are input into one or more first models, thereby obtaining one or more first values. The multiple cellular component abundance values ​​are input into one or more second models, thereby obtaining one or more second values. The one or more first values ​​and the one or more second values ​​are used to determine the risk that the compound will induce the liver injury.

[0058] Detailed reference will now be made to the embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0059] Multiple instances can be provided for the components, operations, or structures described herein as a single instance. Finally, the boundaries between individual components, operations, and data stores are somewhat arbitrary, and specific operations are described within the context of a particular illustrative configuration. Other forms of functionality are envisioned, and they may fall within the scope of the implementation. Generally, structures and functions presented as separate components in the example configuration can be implemented as composite structures or components. Similarly, structures and functions presented as single components can be implemented as separate components. These structures and functions, as well as other variations, modifications, additions, and improvements, fall within the scope of the implementation.

[0060] It should also be understood that although the terms "first," "second," etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the invention, a first dataset may be referred to as a second dataset, and similarly, a second dataset may be referred to as a first dataset. Both the first dataset and the second dataset are datasets, but the first dataset and the second dataset are not the same dataset.

[0061] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and the appended claims, the singular forms “a / an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that, as used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the relevant enumerated items. It will be further understood that, when used in this specification, the terms “comprises” and / or “comprising” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0062] As used herein, depending on the context, the term "if" can be interpreted as meaning "when the stated prerequisite is true" or "when the stated prerequisite is true" or "in response to determining that the stated prerequisite is true" or "based on the determination that the stated prerequisite is true" or "in response to detecting that the stated prerequisite is true". Similarly, depending on the context, the phrases "if it is determined (that the stated prerequisite is true)" or "if (that the stated prerequisite is true)" or "when (that the stated prerequisite is true)" can be interpreted as meaning "when the stated prerequisite is determined to be true" or "in response to determining that the stated prerequisite is true" or "based on the determination that the stated prerequisite is true" or "in response to detecting that the stated prerequisite is true".

[0063] Furthermore, when the reference numerals are assigned "the first" When "individual" is used, the reference numerals refer to general components, collections, or embodiments. For example, "cellular component" is referred to as... "Cellular component" refers to the first of multiple cell components. Cellular components.

[0064] For clarity, not all conventional features of the embodiments described herein are shown or described. It will be understood that in the development of any such practical embodiment, numerous implementation-specific decisions are made to achieve the designer's specific goals, such as compliance with constraints related to use cases and business, and these specific goals will vary depending on the implementation and the designer. Furthermore, it will be understood that such design work can be complex and time-consuming, but remains a routine engineering task for those skilled in the art who benefit from this disclosure.

[0065] Some parts of this description describe embodiments of the invention based on algorithms and symbolic representations of informational operations. These algorithmic descriptions and representations are typically used by those skilled in the art of data processing to effectively convey the essence of their work to others skilled in the art. While these operations are described functionally, computationally, or logically, they should be understood to be implemented via computer programs or equivalent circuits, microcode, etc.

[0066] The language used in this specification has been chosen primarily for readability and guidance purposes, and not for describing or limiting the subject matter of the invention. Therefore, the scope of the invention is intended to be limited not by this specific embodiment, but by any of the claims published in the application based thereon. Consequently, the disclosure of embodiments of the invention is intended to be illustrative and not to limit the scope of the invention.

[0067] Generally, the terms used in the claims and specification are intended to be interpreted in the ordinary sense as understood by one of ordinary skill in the art. Certain terms are defined below to make this clearer. If any ordinary sense conflicts with the provided definitions, the provided definitions will prevail.

[0068] Any term not directly defined herein shall be understood to have the meaning generally associated with it as understood within the art of this invention. Certain terms are discussed herein to provide additional guidance to those skilled in the art in describing compositions, apparatuses, methods, etc., and how said compositions, apparatuses, methods, etc., and in how they are made or used. It will be understood that the same thing can be described in more than one way. Therefore, alternative languages ​​and synonyms may be used to represent any one or more terms discussed herein. It is not important whether a particular term is elaborated or discussed herein. Several synonyms or alternative methods, materials, etc., are provided. Unless explicitly stated, stating one or more synonyms or equivalents does not preclude the use of other synonyms or equivalents. Examples (including examples of terms) are for illustrative purposes only and do not limit the scope and meaning of the various aspects of the invention herein.

[0069] definition.

[0070] As used herein, the term "about" or "approximately" means within an acceptable margin of error for a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value was measured or determined, for example, limitations of the measurement system. For example, according to practice in the art, in some embodiments, "about" means within one or more standard deviations. In some embodiments, "about" means a range of ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term "about" or "approximately" means within an order of magnitude of the value, within five times the value, or within twice the value. When specific values ​​are described in this application and claims, unless otherwise stated, it may be assumed that the term "about" means within an acceptable margin of error for the specific value. All numerical values ​​in the specific embodiments herein are modified by the value indicated by "about" and take into account experimental errors and variations that can be expected by one of ordinary skill in the art. The term "about" may have a meaning commonly understood by one of ordinary skill in the art. In some embodiments, the term "about" means ±10%. In some embodiments, the term "about" means ±5%.

[0071] As used herein, the terms “abundance,” “abundance level,” or “expression level” refer to the amount of a cellular component (e.g., a gene product, such as an RNA species (e.g., mRNA or miRNA)) or a protein molecule) present in one or more cells, or the average amount of a cellular component present across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a specific genomic locus (e.g., a specific gene). However, in some embodiments, abundance may refer to the amount of a specific isotype of mRNA or protein corresponding to a specific gene that produces multiple mRNA or protein isotypes. Genomic loci can be identified using gene names, chromosomal locations, or any other genetic mapping measure.

[0072] As used interchangeably herein, “cell state” or “biological state” refers to the state or phenotype of a cell or population of cells. For example, a cell state can be healthy or diseased. A cell state can be one of many diseases. A cell state can be a response to compound treatment and / or a differentiated cell lineage. A cell state can be characterized by measurements of one or more cellular components, including but not limited to one or more genes, one or more proteins, and / or one or more biological pathways.

[0073] As used herein, "cell state transition" or "cell transformation" refers to the change in the state of a cell from a first cell state to a second cell state. In some embodiments, the second cell state is an altered cell state (e.g., from a healthy cell state to a diseased cell state). In some embodiments, one of the corresponding first and second cell states is an undisturbed state, and the other of the corresponding first and second cell states is a perturbed state caused by cell exposure to a certain condition. The perturbed state can be caused by cell exposure to a compound. Cell state transitions can be marked by changes in the abundance of cellular components within the cell, and thus by the identity and quantity (e.g., perturbation signature) of cellular components produced by the cell (e.g., mRNA, transcription factors).

[0074] As used herein, in some contexts, the term "dataset" for the measurement of cellular component abundance in one or more cells can refer to a high-dimensional dataset collected from a single cell (e.g., a single-cell cellular component abundance dataset). In other contexts, the term "dataset" can refer to multiple high-dimensional datasets collected from a single cell (e.g., multiple single-cell cellular component abundance datasets), each of which is collected from one cell among multiple cells.

[0075] As used herein, the term "differential abundance" or "differential expression" refers to the difference in quantity and / or frequency of a cellular component present in a first entity (e.g., a first cell, multiple cells, and / or a sample) compared to a second entity (e.g., a second cell, multiple cells, and / or a sample). In some embodiments, the first entity is a sample characterized by a first cellular state (e.g., a diseased phenotype), and the second entity is a sample characterized by a second cellular state (e.g., a normal or healthy phenotype). For example, a cellular component may be a polynucleotide (e.g., an mRNA transcript) present at an elevated or decreased level in an entity characterized by a first cellular state compared to an entity characterized by a second cellular state. In some embodiments, a cellular component may be a polynucleotide detected at a higher or lower frequency in an entity characterized by a first cellular state compared to an entity characterized by a second cellular state. A cellular component may be differentially abundant in terms of quantity, frequency, or both. In some cases, a cellular component is differentially abundant between two entities if the amount of a cellular component in one entity is statistically significantly different from the amount of a cellular component in another entity. For example, if a cellular component present in one entity is at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% higher than a cellular component present in another entity, or if the cellular component is detectable in one entity but not in another, then the cellular component is differentially abundant in the two entities. In some cases, if the frequency of detecting a cellular component in a first subset of the entity (e.g., cells representing a first subset of annotated cell states) is statistically significantly higher or lower than the frequency of detecting a cellular component in a second subset of the entity (e.g., cells representing a second subset of annotated cell states), then the cellular component is differentially expressed in the two subsets of the entity. For example, if at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% of cellular components are observed to be detected more frequently or less frequently in a first set of entities than in another set of entities, then the cellular components are differentially expressed in the two sets of entities.

[0076] As used herein, the terms “sample,” “biological sample,” or “patient sample” refer to any sample taken from a subject that reflects the biological state associated with the subject. Examples of samples include, but are not limited to, a subject’s blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural effusion, pericardial fluid, or peritoneal fluid. Samples may include any tissue or material derived from a living or deceased subject. Samples may be cell-free. Samples may contain one or more cellular components. For example, a sample may contain nucleic acids (e.g., DNA or RNA) or fragments thereof, or proteins. The term “nucleic acid” may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. Nucleic acids in a sample may be free nucleic acids. Samples may be liquid or solid samples (e.g., cell or tissue samples). Samples may be bodily fluids. Samples may be fecal samples. Samples can be processed to physically disrupt tissue or cell structures (e.g., centrifugation and / or cell lysis) to release intracellular components into a solution, which may further contain enzymes, buffers, salts, detergents, etc., that can be used to prepare samples for analysis.

[0077] As used interchangeably in this article, the terms “classifier,” “model,” “algorithm,” “regressor,” and / or “classifier” refer to machine learning models or algorithms.

[0078] In some embodiments, the model used in this disclosure is a supervised machine learning model. Non-limiting examples of supervised learning algorithms that can be used as models in this disclosure include, but are not limited to, logistic regression, neural networks, support vector machines, Naive Bayes, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosting tree algorithms, multinomial logistic regression algorithms, linear models, linear regression, gradient boosting, mixture models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, or any combination thereof. In some embodiments, the model is a multinomial classifier algorithm. In some embodiments, the model is a 2-stage stochastic gradient descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a deep and broad sample-level model). In some embodiments, the classifier or model of this disclosure has 25 or more, 100 or more, 1000 or more, 10,000 or more, 100,000 or more, or 1 x 10^64 elements. 6 There are one or more parameters, and therefore, in some embodiments, the model cannot be calculated by the brain.

[0079] Furthermore, as used herein, the term "parameter" refers to any coefficient or similarly internal or external element (e.g., weights and / or hyperparameters) in an algorithm, model, regressor, and / or classifier that may affect (e.g., modify, customize, and / or adjust) one or more inputs, outputs, and / or functions of the algorithm, model, regressor, and / or classifier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, customize, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, parameters are used to increase or decrease the influence of inputs (e.g., features) on an algorithm, model, regressor, and / or classifier. As a non-limiting example, in some embodiments, parameters are used to increase or decrease the influence of nodes (e.g., nodes in a neural network), where nodes include one or more activation functions. Assigning parameters to specific inputs, outputs, and / or functions for a given algorithm, model, regressor, and / or classifier is not limited to any particular paradigm and can be used for any suitable algorithm, model, regressor, and / or classifier architecture to achieve the desired performance. In some embodiments, parameters have fixed values. In some embodiments, the values ​​of the parameters may be adjusted manually and / or automatically. In some embodiments, the values ​​of the parameters are modified through the validation and / or training process of the algorithm, model, regressor, and / or classifier (e.g., through error minimization and / or backpropagation methods). In some embodiments, the algorithms, models, regressors, and / or classifiers of this disclosure include multiple parameters. In some embodiments, the plurality of parameters is n parameters, where: n ≥ 2; n ≥ 5; n ≥ 10; n ≥ 25; n ≥ 40; n ≥ 50; n ≥ 75; n ≥ 100; n ≥ 125; n ≥ 150; n ≥ 200; n ≥ 225; n ≥ 250; n ≥ 350; n ≥ 500; n ≥ 600; n ≥ 750; n ≥ 1,000; n ≥ 2,000; n ≥ 4,000; n ≥ 5,000; n ≥ 7,500; n ≥ 10,000; n ≥ 20,000; n ≥ 40,000; n ≥ 75,000; n ≥ 100,000; n ≥ 200,000; n ≥ 500,000, n ≥ 1x 10 6 n ≥ 5 x 10 6 Or n ≥ 1 x 10 7 Therefore, in some embodiments, the algorithms, models, regressors, and / or classifiers of this disclosure cannot be processed by the brain. In some embodiments, n is between 20 and 100, between 50 and 1000, between 100 and 5000, and between 10,000 and 1 x 10^25. 7 Between 100,000 and 5 x 106 Between or between 500,000 and 1 x 10 6 In some embodiments, the algorithms, models, regressors, and / or classifiers of this disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). Therefore, the algorithms, models, regressors, and / or classifiers of this disclosure cannot be generated by the brain.

[0080] Neural Networks. In some embodiments, the models used in this disclosure are neural networks (e.g., convolutional neural networks and / or residual neural networks). Neural network models are also called artificial neural networks (ANNs), including convolutional and / or residual neural network models (deep learning models). A neural network can be a machine learning model that can be trained to map an input dataset to an output dataset, wherein the neural network contains a set of interconnected nodes organized into multiple layers. For example, a neural network architecture can contain at least an input layer, one or more hidden layers, and an output layer. A neural network can contain any total number of layers and any number of hidden layers, wherein the hidden layers act as trainable feature extractors that allow mapping of an input dataset to an output value or a set of output values. As used herein, a deep learning model (DNN) can be a neural network containing multiple hidden layers, such as two or more hidden layers. Each layer of a neural network can contain multiple nodes (or “neurons”). Nodes can receive input directly from input data or the outputs of nodes in previous layers and perform specific operations, such as summation. In some embodiments, the connections from the input to the node are associated with parameters (e.g., weights and / or weighting factors). In some embodiments, a node can pair all inputs with x i The product of the weighted sum and its associated parameters is added. In some embodiments, the weighted sum cancels out the bias b. In some embodiments, the output of a node or neuron can be gated using a threshold or an activation function f, which can be a linear or nonlinear function. The activation function can be, for example, a modified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as a saturated hyperbolic tangent function, an identity function, a binary step function, a logic function, an arcTan function, a softsign function, a parameter-modified linear unit function, an exponential linear unit function, a softPlus function, a curved identity function, a softExponential function, a Sinusoid function, a Sine function, a Gaussian function, or a sigmoid function, or any combination thereof.

[0081] The weighting factors, biases, and / or thresholds, or other computational parameters of a neural network can be "taught" or "learned" during the training phase using one or more training datasets. For example, the parameters can be trained using input data from the training dataset and gradient descent or backpropagation methods so that the output values ​​of the ANN are consistent with the instances included in the training dataset. The parameters can also be obtained during the backpropagation neural network training process.

[0082] Any type of neural network is suitable for this disclosure. Examples may include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, and any combination thereof. In some embodiments, machine learning utilizes pre-trained and / or transfer learning ANN or deep learning architectures. Convolutional and / or residual neural networks can be used as models in this disclosure.

[0083] For example, a deep neural network model comprises an input layer, multiple individually parameterized (e.g., weighted) convolutional layers, and an output scorer. The parameters (e.g., weights) of each of the convolutional layers and the input layer contribute to a plurality of parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Therefore, deep neural network models require a computer to operate because they cannot be solved by the brain. In other words, given the model's inputs, the model's output needs to be determined using a computer, not by the brain in such embodiments. See, for example, Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems, 2, edited by Pereira, Burges, Bottou, and Weinberger, pp. 1097-1105, Curran Associates, Inc.; Zeiler, 2012, “ADADELTA: an adaptive learning rate method,” CoRR, Vol. abs / 1212.5701; and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” Chapter: Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press. (MIT Press), each of the references in this document is hereby incorporated by reference.

[0084] Neural network models (including convolutional neural network models) suitable for use as models are disclosed in the following literature: for example, Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in adeep network with a local denoising criterion”, Journal of Machine Learning Research 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks”, Journal of Machine Learning Research 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is hereby incorporated by reference. Other example neural networks suitable for use as models are disclosed in the following literature: Duda et al., 2001, Pattern Classification, 2nd ed., John Wiley & Sons, Inc., New York; and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is hereby incorporated in its entirety by reference.Other example neural networks suitable for use as models are described in the following literature: Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC; and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is hereby incorporated in its entirety by reference.

[0085] In some embodiments, the neural network is a fully connected neural network with ReLU activation. For example, in some embodiments, the model is a neural network comprising: corresponding to one or more inputs, wherein each of the corresponding one or more inputs is used to test the chemical structure of a chemical compound; corresponding to a first hidden layer, the corresponding first hidden layer comprising a plurality of hidden neurons, wherein each of the plurality of hidden neurons is (i) fully connected to each of the plurality of inputs, (ii) associated with a first activation function type, and (iii) associated with a corresponding parameter (e.g., weight) of a plurality of parameters of the neural network; and one or more corresponding neural network outputs, wherein each corresponding neural network output (i) directly or indirectly receives the output of each of the plurality of hidden neurons as input, and (ii) is associated with a second activation function type. In some such embodiments, the neural network is a fully connected network.

[0086] In some embodiments, a neural network comprises multiple hidden layers. As described above, hidden layers are located between the input layer and the output layer (e.g., to capture additional complexity). In some embodiments, where multiple hidden layers exist, each hidden layer may have the same or different corresponding numbers of neurons.

[0087] In some embodiments, each hidden neuron (e.g., in a corresponding hidden layer in a neural network) is associated with an activation function that performs a function (e.g., a linear or nonlinear function) on the input data. Generally, the purpose of an activation function is to introduce nonlinearity into the data, such that the neural network is trained on a representation of the original data and can subsequently “fit” or generate additional representations of new (e.g., previously unseen) data. The choice of activation function (e.g., a first and / or a second activation function) depends on the use of the neural network, as some activation functions may cause saturation at the extremes of the dataset (e.g., the tanh function and / or the sigmoid function). For example, in some embodiments, the activation function (e.g., the first and / or the second activation function) is selected from any suitable activation function known in the art, including but not limited to any activation function disclosed herein.

[0088] In some embodiments, each hidden neuron is further associated with parameters (e.g., weights and / or biases) that contribute to the output of the neural network, determined based on an activation function. In some embodiments, the hidden neuron is initialized with arbitrary parameters (e.g., randomized weights). In some alternative embodiments, the hidden neuron is initialized with a predetermined set of parameters.

[0089] In some embodiments, the number of hidden neurons in the neural network (e.g., across one or more hidden layers) is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 neurons. In some embodiments, the plurality of hidden neurons are at least 100, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 15,000, at least 20,000, or at least 30,000 neurons. In some embodiments, the number of hidden neurons is no more than 30,000, no more than 20,000, no more than 15,000, no more than 10,000, no more than 9,000, no more than 8,000, no more than 7,000, no more than 6,000, no more than 5,000, no more than 4,000, no more than 3,000, no more than 2,000, no more than 1,000, no more than 900, no more than 800, no more than 700, no more than 600, no more than 500, no more than 400, no more than 300, no more than 200, no more than 100, or no more than 50 neurons. In some embodiments, the plurality of hidden neurons is 2 to 20, 2 to 200, 2 to 1000, 10 to 50, 10 to 200, 20 to 500, 100 to 800, 50 to 1000, 500 to 2000, 1000 to 5000, 5000 to 10,000, 10,000 to 15,000, 15,000 to 20,000, or 20,000 to 30,000 neurons. In some embodiments, the plurality of hidden neurons falls into another range that begins with at least 2 neurons and ends with at least 30,000 neurons.

[0090] In some embodiments, the neural network includes 1 to 50 hidden layers. In some embodiments, the neural network includes 1 to 20 hidden layers. In some embodiments, the neural network includes at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 hidden layers. In some embodiments, the neural network includes no more than 100, no more than 90, no more than 80, no more than 70, no more than 60, no more than 50, no more than 40, no more than 30, no more than 20, no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, or no more than 5 hidden layers. In some embodiments, the neural network includes 1 to 5, 1 to 10, 1 to 20, 10 to 50, 2 to 80, 5 to 100, 10 to 100, 50 to 100, or 3 to 30 hidden layers. In some embodiments, the neural network includes a plurality of hidden layers falling into another range that begins with at least one layer and ends with at least 100 layers.

[0091] In some embodiments, the neural network comprises a shallow neural network. A shallow neural network is a neural network with a small number of hidden layers. In some embodiments, such neural network architectures improve the efficiency of neural network training and save computational power because the number of layers involved in training is reduced. In some embodiments, the neural network comprises one hidden layer. In some embodiments, the neural network comprises two, three, four, or five hidden layers.

[0092] In some embodiments, the neural network is a message-passing neural network. A message-passing neural network refers to a framework for supervised learning on graphs (e.g., a graph-based representation of a chemical structure), where nodes represent atoms and edges represent bonds between atoms. Typically, a message-passing neural network comprises two phases in the forward pass: a message-passing phase and a readout phase. The message-passing phase runs for T time intervals and includes the readout phase based on a message function M. t and vertex update function U tThe hidden state at each node in the graph is updated. The readout phase uses the readout function R to compute the eigenvectors of the graph. In some embodiments, the message-passing neural network comprises a convolutional network (e.g., a spatial graph convolutional network and / or a spectral graph convolutional network), a gated graph neural network (GG-NN), an interaction network, a molecular graph convolution, a deep tensor neural network, and / or a Laplacian-based method. See, for example, Gilmer et al., 2017, “Neural Message Passing for Quantum Chemistry,” arXiv:1704.01212v2, which is incorporated herein by reference in its entirety.

[0093] Support Vector Machine. In some embodiments, the model used in this disclosure is a Support Vector Machine (SVM). Suitable SVM models are described in the following literature: Cristianini and Shaw-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin models,” Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, *Statistical Learning Theory*, Wiley, New York; Mount, 2001, *Bioinformatics: Sequence and Genome Analysis*, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York; Duda, *Pattern Classification*, 2nd edition. John Wiley & Sons, 2001, pp. 259, 262-265; and Hastie, 2001, Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated in its entirety by reference. When used for classification, SVM separates a given binary-labeled dataset from a hyperplane that is maximally farthest from the labeled data. For cases where linear separation is not possible, SVM can be operated in conjunction with a 'kernel' technique that automatically implements a nonlinear mapping of the feature space. The hyperplane discovered by SVM in the feature space can correspond to a nonlinear decision boundary in the input space. In some embodiments, multiple parameters (e.g., weights) associated with the SVM define the hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and the SVM model requires computation by a computer because it cannot be solved by the brain.

[0094] Naive Bayes Model. In some embodiments, the model used in this disclosure is a Naive Bayes model. Suitable Naive Bayes models for use as models are disclosed in, for example, Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes”, Advances in Neural Information Processing Systems, 14, which is hereby incorporated by reference. A Naive Bayes classifier is any classifier in the family of “probabilistic classifiers” that applies the Bayes theorem and the strong (naive) independence assumption between features. In some embodiments, it is combined with kernel density estimation. See, for example, Hastie et al., 2001, The elements of statistical learning: data mining, inference, and prediction (eds.: Tibshirani and Friedman, Springer Publishers, New York), which is hereby incorporated by reference.

[0095] Nearest Neighbor Model. In some embodiments, the model used in this disclosure is a nearest neighbor model. The nearest neighbor model may be memory-based and does not include the model to be fitted. For nearest neighbors, given a query point x0 (test subject), the k nearest training points to x0 can be identified. That is, r, ..., k (here, training subjects), and then classification is performed using the k nearest neighbor pairs x0. Here, the distance to these neighbors is a function of the abundance value of the discriminant gene set. In some embodiments, the distance is determined using Euclidean distance in the feature space. Typically, when using nearest neighbor models, the abundance data used to compute linear discriminants is standardized to have a mean of zero and a variance of 1. The nearest neighbor rule can be refined to address issues of unequal classification priors, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting on neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, 2nd Edition, 2001, John Wiley & Sons; and Hastie, Elements of Statistical Learning, 2001, Springer, New York, each of which is hereby incorporated by reference.

[0096] The k-nearest neighbor model is a nonparametric machine learning method where the input consists of the k nearest training instances in the feature space. The output is a class membership. An object is classified by multiple votes from its neighbors, whereby the object is assigned to the most common class among its k nearest neighbors (k is a positive integer, typically small). If k = 1, the object is simply assigned to the class of a single nearest neighbor. See Duda et al., 2001, Pattern Classification, 2nd Edition, John Willie & Sons, which is hereby incorporated by reference. In some embodiments, the amount of distance computation required to solve the k-nearest neighbor model makes it necessary to use a computer to solve the model for a given input, as it cannot be done by the brain.

[0097] Random forest, decision tree, and boosting tree models. In some embodiments, the model used in this disclosure is a decision tree. Decision trees suitable for use as models are described in: Duda, 2001, Pattern Classification, John Willie & Sons, New York, pp. 395-396, which is hereby incorporated by reference. Tree-based methods divide the feature space into a set of rectangles and then fit a model (e.g., a constant) to each rectangle. In some embodiments, the decision tree is a random forest regression. One particular model that can be used is the Classification and Regression Tree (CART). Other specific decision tree models include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in: Duda, 2001, Pattern Classification, John Willie & Sons, New York, pp. 396-408 and 411-412, which is hereby incorporated by reference. CART, MART, and C4.5 are described in: Hastie et al., 2001, Elements of Statistical Learning, Springer, New York, Chapter 9, which is hereby incorporated in its entirety by reference. Random forests are described in: Breiman, 1999, “Random Forests—Random Features,” Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is hereby incorporated in its entirety by reference. In some embodiments, decision tree models include at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions) and require computation by a computer because they cannot be solved by the brain.

[0098] Regression. In some embodiments, the model used in this disclosure is a regression algorithm. The regression algorithm can be any type of regression. For example, in some embodiments, the regression is logistic regression. In some embodiments, the regression is logistic regression regularized using lasso, L2, or elastic net regularization. In some embodiments, the regression is logistic regression regularized using L1 elastic net regularization. In some embodiments, extracted features with corresponding regression coefficients that fail to meet a threshold are removed (not considered). In some embodiments, a generalization of a logistic regression model that handles multi-class responses is used as the model. Logistic regression is disclosed in: Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Willie & Sons, New York, which is hereby incorporated by reference. In some embodiments, the model utilizes a regression model disclosed in: Hastie et al., Elements of Statistical Learning, Springer, New York, 2001. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights). In some embodiments, the logistic regression model requires computation by a computer because it cannot be done by the brain.

[0099] Linear discriminant analysis. Linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis can be a generalization of Fisher's linear discriminant, a method used in statistics, pattern recognition, and machine learning to find linear combinations of features that characterize or distinguish two or more classes of objects or events. The resulting combinations can be used as models (linear models) in some embodiments of this disclosure.

[0100] Hybrid Models and Hidden Markov Models. In some embodiments, the models used in this disclosure are hybrid models, such as those described in McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those including a time component, the models are hidden Markov models, such as those described in Schliep et al., Bioinformatics 19(1):i255-i263, 2003.

[0101] Model ensembles and boosting. In some embodiments, a model ensemble (two or more models) is used. In some embodiments, boosting techniques such as AdaBoost are combined with many other types of learning algorithms to improve model performance. In this approach, the outputs of any model disclosed herein, or their equivalents, are combined into a weighted sum, which represents the final output of the boosted model. In some embodiments, multiple outputs from the models are combined using any measure of central tendency known in the art, including but not limited to mean, median, mode, weighted average, weighted median, weighted mode, etc. In some embodiments, a voting method is used to combine multiple outputs. In some embodiments, the respective models in the model ensemble are weighted or unweighted.

[0102] I. Exemplary System Implementation

[0103] An overview of some aspects of this disclosure and some definitions used in this disclosure have now been provided, and details of the exemplary system are described in conjunction with Figure 1.

[0104] Figure 1 provides a block diagram illustrating a system 100 according to some embodiments of the present disclosure. System 100 assesses the risk that a tested chemical compound will induce liver injury in human subjects. In Figure 1, system 100 is shown as a computing device. Other topologies of the computer system 100 are possible. For example, in some embodiments, system 100 may actually constitute several computer systems linked together in a network, or virtual machines or containers in a cloud computing environment. Therefore, the exemplary topologies shown in Figure 1 are only used to describe features of embodiments of the present disclosure in a manner readily understood by those skilled in the art.

[0105] Referring to FIG1, in some embodiments, computer system 100 (e.g., computing device) includes network interface 104. In some embodiments, network interface 104 interconnects the computing devices of system 100 within the system with each other and with optional external systems and devices via one or more communication networks. In some embodiments, network interface 104 optionally provides communication via the Internet, one or more local area networks (LANs), one or more wide area networks (WANs), other types of networks, or combinations of such networks.

[0106] Examples of networks include the World Wide Web (WWW), intranets and / or wireless networks (such as cellular telephone networks), wireless local area networks (LANs) and / or metropolitan area networks (MANs), and other devices that communicate wirelessly. Wireless communication optionally uses any of a variety of communication standards, protocols, and technologies, including Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolved Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE...). 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)) and / or Short Message Service (SMS) or any other suitable communication protocol, including communication protocols that have not been developed as of the date of submission of this document.

[0107] In some embodiments, system 100 includes one or more processing units (CPUs) 102 (e.g., processors, processing cores, etc.), one or more network interfaces 104, a user interface 106 for user use including (optionally) a display 108 and an input system 105 (e.g., input / output interfaces, keyboards, mice, etc.), memory (e.g., non-persistent memory 107, persistent memory), and one or more communication buses 103 for interconnecting the aforementioned components. The one or more communication buses 103 optionally include a circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Non-persistent memory 107 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while persistent memory (not shown) typically includes CD-ROMs, digital versatile optical discs (DVDs) or other optical storage devices, magnetic tape cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, disk storage devices, optical disc storage devices, flash memory devices or other non-volatile solid-state storage devices. Persistent memory 109 optionally includes one or more storage devices located remotely from CPU 102. Non-volatile memory devices within permanent and non-permanent memory include non-transitory computer-readable storage media. In some embodiments, sometimes in conjunction with permanent memory, non-permanent memory 107 or alternatively, non-transitory computer-readable storage media stores programs, modules, and data structures or subsets thereof:

[0108] ● Optional operating system (not shown) (e.g., embedded operating systems such as ANDROID, iOS, DARWIN, RTXC, LINUX, UNIX, OSX, WINDOWS, or VxWorks), which includes programs for handling various basic system services and for performing hardware-related tasks;

[0109] ● Optional network communication module (not shown) (or instruction) for connecting system 100 to other devices and / or communication network 104.

[0110] ● For the tested chemical compound 114, for each concentration 116 / exposure time 118 pair (e.g., 116-1 / 118-1 to 116-M / 118-M, where M is a positive integer of two or greater), a cell component dataset 114 containing measured cell component abundance values ​​122 (e.g., 122-1-1, 122-1-2, ..., 120-1-N, where N is a positive integer of 3 or greater) for each cell component 120 (e.g., 120-1-1, 120-1-2, ..., 120-1-N);

[0111] ● One or more first models 130 that contain multiple model parameters 132 (e.g., 132-1, ..., 132-K, where K is a positive integer) and are configured to output one or more first model outputs 134 for evaluating the risk of DILI;

[0112] ● One or more second models 136 (e.g., second models 136-1, 136-2, ..., 136-T) used to assess the risk of DILI, each model 136 containing a corresponding plurality of model parameters 138 (e.g., 138-1-1, ..., 138-1-Q, where Q is a positive integer), and configured to output second model 2 output 140 (e.g., 140-1); and

[0113] ● Multiple pathway modules 142 (e.g., 142-1, 142-2, ..., 142-X), each associated with a subset of cellular components 144 (e.g., 144-1-1, 144-1-2, 144-1-P, where P is a positive integer; 144-2-1, 144-2-2, 144-2-W, where W is a positive integer; 144-X-1, 144-X-2, 144-XR, where R is a positive integer).

[0114] In various embodiments, one or more of the aforementioned elements are stored in one or more previously mentioned memory devices and correspond to an instruction set for performing the functions described above. The aforementioned modules, data, or programs (e.g., instruction sets) do not need to be implemented as separate software programs, programs, datasets, or modules, and therefore, various subsets of these modules and data can be combined or otherwise rearranged in various embodiments. In some embodiments, non-persistent memory 107 optionally stores a subset of the aforementioned modules and data structures. Furthermore, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the aforementioned elements are stored in a computer system outside the computer system of system 100, which is addressable by system 100, such that system 100 can access all or part of such data when needed.

[0115] Although Figure 1 depicts "System 100," the figure is intended more as a functional description of the various features that may exist in a computer system than as a structural schematic of the implementation described herein. In practice, and as those skilled in the art will recognize, items shown individually may be combined, and some items may be separated. Furthermore, although Figure 1 depicts some data and modules in non-persistent memory 107, some or all of these data and modules may be stored in persistent memory or in more than one memory. For example, in some embodiments, at least the cellular component dataset 112 may be stored in a remote storage device that may be part of a cloud-based infrastructure.

[0116] II. Methods for assessing the risk of liver injury induced by the tested chemical compound in subjects.

[0117] Reference box 200, in some embodiments, provides a method for assessing the risk that a test chemical compound will induce liver injury in human subjects.

[0118] Referring to box 202, in some embodiments, the test chemical compound is an organic compound with a molecular weight less than 2000 Daltons. In some embodiments, the test chemical compound is inorganic or organic. In some embodiments, the test chemical compound is an organic compound with a molecular weight less than 2000 Daltons (Da). In some embodiments, the molecular weight of the test chemical compound is at least 10 Da, at least 20 Da, at least 50 Da, at least 100 Da, at least 200 Da, at least 500 Da, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 5 kDa, at least 10 kDa, at least 20 kDa, at least 30 kDa, at least 50 kDa, at least 100 kDa, or at least 500 kDa. In some embodiments, the molecular weight of the test chemical compound is not more than 1000 kDa, not more than 500 kDa, not more than 100 kDa, not more than 50 kDa, not more than 10 kDa, not more than 5 kDa, not more than 2 kDa, not more than 1 kDa, not more than 500 Da, not more than 300 Da, not more than 100 Da, or not more than 50 Da. In some embodiments, the molecular weight of the tested chemical compound is 10 Da to 900 Da, 50 Da to 1000 Da, 100 Da to 2000 Da, 1 kDa to 10 kDa, 5 kDa to 500 kDa, or 100 kDa to 1000 kDa. In some embodiments, the molecular weight of the tested chemical compound falls into another range that begins at or above 10 Daltons and ends at or above 1000 kDa.

[0119] Boxes 204 and 206. Referring to box 204, in some embodiments, the test chemical compound is an organic compound that satisfies each of the criteria in the Lipinski Five-Rate Rule. Referring to box 206, in some embodiments, the test chemical compound is an organic compound that satisfies at least three of the criteria in the Lipinski Five-Rate Rule. In some embodiments, the test chemical compound is an organic compound that satisfies two or more, three or more, or all four of the Lipinski Five-Rate Rule: (i) no more than five hydrogen bond donors (e.g., OH and NH groups), (ii) no more than ten hydrogen bond acceptors (e.g., N and O), (iii) a molecular weight less than 500 Daltons, and (iv) a LogP less than 5. It is called the "Five-Rate Rule" because three of the four criteria involve the number five. See Lipinski, 1997, *Advanced Drug Del. Rev.* 23, 3, which is hereby incorporated herein by reference in its entirety. In some embodiments, in addition to the Lipinski Five-Rate Rule, the compounds of this disclosure also satisfy one or more criteria. For example, in some embodiments, the compounds disclosed herein have five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings.

[0120] Referring to box 208, in some embodiments, liver injury is hepatitis, hepatotoxicity, jaundice, liver fibrosis, cirrhosis, hepatomegaly, alcoholic and non-alcoholic fatty liver disease, and / or cholangiorrhea.

[0121] Drug-induced liver disease can be classified into three injury modalities: hepatocellular, cholestatic, and mixed hepatocellular-cholestatic. These names refer to the histological features of the injury, but are often defined based on patterns of elevated serum enzymes. See LiverTox: Clinical and research information on drug-induced liver injury [Internet]. Bethesda (MD): National Institute of Diabetes and Digestive and Kidney Diseases; 2012–. Clinical Course and Diagnosis of Drug Induced Liver Disease. [Updated May 4, 2019]. Available at: https: / / www.ncbi.nlm.nih.gov / books / NBK548733 / , which is hereby incorporated by reference.

[0122] Hepatitis, hepatotoxicity. Drug-induced liver disease, similar to acute viral hepatitis, is characterized by a prominent pattern of hepatocellular damage. Liver biopsy (if available) typically reveals significant hepatocellular necrosis and inflammation, with only mild cholestasis at least in the early stages. If present, fatigue and weakness are predominant symptoms. Serum alanine and aspartate aminotransferase (ALT and AST) levels are usually significantly elevated (typically >10-fold), while alkaline phosphatase or gamma-glutamyl transferase (GGT) are only moderately elevated. An ALT-to-ALT "R" ratio (both expressed as multiples of the upper limit of the normal range) of 5 or greater is generally used to define the pattern of hepatocellular damage. Ibid.

[0123] Cholestatic injury. Drug-induced liver injury presents with cholestatic features similar to bile duct obstruction or common bile duct stones. Liver biopsy typically reveals cholestasis, portal vein inflammation, and proliferation or damage to bile ducts and small ducts. Clinically, it is predominantly characterized by jaundice and pruritus. (Ibid.)

[0124] Mixed hepatocellular-cholestatic injury. A mixture of hepatocellular and cholestatic injury is characteristic of many drugs and, in fact, is the most characteristic pattern of drug-induced liver injury, rarely occurring in other forms of acute liver disease. In cases of mixed injury, liver biopsy reveals prominent hepatocellular necrosis and inflammation, accompanied by significant cholestasis. Symptoms may include both fatigue and itching, and laboratory tests show similar elevations in serum ALT and alkaline phosphatase. An R ratio of ALT to alkaline phosphatase between 2 and 5 (both expressed as multiples of the upper limit of normal) is used to define the mixed injury pattern. Drugs that cause mixed hepatocellular-cholestatic injury include sulfonamides, phenytoin, and enalapril. Ibid.

[0125] In addition to the injury pattern, drug-induced liver injury can be classified into at least twelve “phenotypes” based on the overall clinical pattern and injury process. These phenotypes overlap to some extent with the injury patterns (hepatocellular, cholestatic, and mixed) and include: acute hepatic necrosis, acute (hepatocellular) hepatitis, cholestatic hepatitis, mixed hepatitis, elevated serum enzymes without jaundice, mild cholestasis, acute fatty liver with lactic acidosis, non-alcoholic fatty liver, chronic hepatitis, sinusoidal obstruction syndrome, nodular regenerative hyperplasia, and liver tumors such as hepatic adenoma and hepatocellular carcinoma. Furthermore, characteristic features or outcomes are often assigned to each phenotype, such as immune hypersensitivity features, autoimmune features, acute liver failure, bile duct disappearance syndrome, and cirrhosis. (Ibid.)

[0126] Referring to box 210, in some embodiments, liver injury is characterized by elevated levels of one or more liver enzymes, cholestasis, and / or alcoholic stools. For example, in some embodiments, serum alanine and aspartate aminotransferase (ALT and AST) levels are elevated more than 10-fold, while alkaline phosphatase or gamma-glutamyl transferase (GGT) is moderately increased. In some embodiments, liver injury is characterized by liver function tests such as ALP (alkaline phosphatase), ALT (alanine aminotransferase), AST (aspartate aminotransferase), or gamma-glutamyl transferase (GGT). In some embodiments, liver injury is characterized by histological findings (e.g., liver biopsy). In some embodiments, liver injury is characterized by transient elastography (fiber scanning). See, for example, Foucher et al., 2006, “Diagnosis of cirrhosis by transient elastography (FibroScan): a prospective study,” *Gut* 55(3), pp. 403-408, which is hereby incorporated by reference. In some embodiments, liver injury is characterized by imaging such as ultrasound, computed tomography (CT), or magnetic resonance imaging (MRI). In some embodiments, liver injury is characterized by a Child-Plugh score to determine cirrhosis (e.g., using five clinical parameters: bilirubin, albumin, INR, ascites, hepatic lesions, and encephalopathy). See, for example, Tsoris A, Marlar CA., “Use of the Child Pugh Score in Liver Disease,” [Updated March 13, 2023]. In StatPearls [Internet]. Treasure Island (Florida, FL): StatPearls Publishing; January 2023. Available online at ncbi.nlm.nih.gov / books / NBK542308. In some embodiments, liver injury is characterized by viral load (e.g., viral hepatitis).See, for example, Oliveira et al., “Advanced Liver Injury in Patients with Chronic Hepatitis B and Viral Load Below 2,000 Iu / Ml,” Rev Inst Med Trop Sao Paulo., Sep 22, 2016; 58:65. doi:10.1590 / S1678-9946201658065. PMID: 27680170; PMCID: PMC5048636, which is hereby incorporated by reference.

[0127] In cholestasis with no or only moderate hepatocellular damage, itching and jaundice are prominent; however, in other cases, patients usually feel well. Serum alkaline phosphatase and ALT levels may be normal or slightly (less than twice) elevated, especially at the peak of symptoms and jaundice. At onset, ALT levels may be significantly elevated (3 to 10 times), and in more severe cases, serum alkaline phosphatase may be more than twice as high. Simple or mild cholestasis is a characteristic form of liver injury caused by estrogen and anabolic steroids and may occur with the illicit use of muscle-building regimens, and thioguanine (such as azathioprine and mercaptopurine) is rarer. Ibid.

[0128] Boxes 212-214. Referring to box 212, in some embodiments, a cellular component dataset 112 is obtained. In some embodiments, the cellular component dataset 112 is in electronic form. The dataset contains multiple cellular component abundance values ​​for multiple cellular components in a first plurality of cells that have been exposed to a test chemical compound for a first duration. In some embodiments, multiple cellular components are measured at a single time point. In some embodiments, multiple cellular components are measured at multiple time points. In embodiments where multiple cellular components are measured at multiple time points, a first subset of the plurality of cells is exposed to the test chemical compound for a first duration (118-1), a second subset of the plurality of cells is exposed to the test chemical compound for a second duration (118-2), ..., and an Mth subset of cells is exposed to the test chemical compound for an Mth duration, where M is a positive integer. In some embodiments, M is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or greater, 20 or greater, or 30 or greater. For each of these M subsets, the exposure time may be the same or different. For example, in some embodiments, the exposure time is different, such as ranging from 4 hours at one time point to 24 hours at another time point.

[0129] In some embodiments, each cell component value among multiple cell component abundance values ​​of multiple cell components in a first plurality of cells that have been exposed to the test chemical compound for a first duration of time is presented in the form of a fold change in cell component expression (DMSO relative to the compound perturbation). For example, in some embodiments, the cell component is a gene, and the measured gene expression readings are measured by RNA-seq to obtain transcriptome abundance (e.g., by SMART-Seq DE) from the perturbation (cells exposed to the test composition at a specific concentration) and DMSO (cells exposed to DMSO only). In some embodiments, the fold change is determined using a procedure such as DESeq2 (… https: / / genomebiology. biomedcentral.com / articles / 10.1186 / s13059-014-0550-8Love et al., 2014, “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2,” *Genome Biology*, 15, 550, available online at doi.org / 10.1186 / s13059-014-0550-8, normalize the logarithmic fold change estimate for the perturbation spectrum against DMSO (cells exposed only to DMSO) (cells exposed to the test compound) and provide the logarithmic fold change estimate. DESeq2 first estimates the size factor for each sample. The size factor represents the scaling factor that should be applied to each sample to account for differences in sequencing depth. DESeq2 calculates this as the median ratio of gene counts across all samples (e.g., three copies, etc.) to their geometric mean. DESeq2 then uses these size factors to normalize the raw counts by dividing the count of each gene in the sample by the corresponding size factor of the sample. This normalization step adjusts the counts for sequencing depth to make the samples comparable. The logarithmic fold change is then determined by taking the difference between the perturbation of each gene (cells exposed to the test compound at a given concentration for a given time) and DMSO (cells exposed to DMSO for a given time) as log2 - the transformed normalized count (log2 - count). DESeq2 employs a shrinking procedure called "regularized logarithmic transformation" to moderate the logarithmic fold change. This helps improve the accuracy of the logarithmic fold change for genes with low counts and high dispersion. Finally, DESeq2 performs a statistical test (Wald test) to determine whether the observed logarithmic fold change is significantly different from zero. If so, it is used as the logarithmic fold change, which serves as a cellular component abundance value for multiple cellular components in a first plurality of cells that have been exposed to the test chemical compound for a first duration.

[0130] In some embodiments, the exposure time is the same for each subset of the M subsets, and the concentration 116 at each subset of the M subsets varies. In some embodiments, each subset of multiple cells is exposed to a different concentration of the test chemical compound. For example, in some embodiments, separate PK studies are optionally used to estimate the C of the test chemical compound. max Then select C at an exposure concentration of 116 that at least covers the tested chemical compound. maxThe range of exposure concentrations allows the calculation of predicted safety margins using a cellular component dataset. In some embodiments, the exposure concentration range is 0.1 µM to 1 µM. In some embodiments, the exposure concentration range is 0.01 µM to 1 µM. In some embodiments, the exposure concentration range is 0.1 µM to 10 µM. In some embodiments, the lowest concentration used in the exposure concentration range is 0.001 µM, 0.01 µM, 0.01 µM, or 1 µM. In some embodiments, the highest concentration used in the exposure concentration range is 100 µM, 50 µM, 10 µM, 1 µM, 0.5 µM, or 0.1 µM. In some embodiments, the cellular component data provides separate subsets of cellular component abundance values ​​for multiple cellular components from one, two, three, four, five, six, seven, eight, nine, or ten different subsets of multiple cells, where each different subset of multiple cells is exposed to the test chemical compound at a different exposure concentration. In some embodiments, such measurements are performed in duplicate, triplicate, or with a higher order of redundancy. For example, in some embodiments, such measurements are performed in triplicate. As an example, for a given exposure concentration 116 / exposure time 118 pair, three distinct subsets of multiple cells are independently exposed to the test chemical composition at the given exposure concentration 116 (e.g., in separate wells, in separate cell aliquots) for exposure time 118, and the cellular component abundance values ​​for multiple cellular components are obtained separately from each separate well / separate cell aliquot. For example, the cellular component abundance value for each of the multiple cellular components is measured in the cells of the first well / cell aliquot, the cellular component abundance value for each of the multiple cellular components is measured in the cells of the second well / cell aliquot, and the cellular component abundance value for each of the multiple cellular components is measured in the cells of the third well / cell aliquot. These measurements are then combined to obtain the cellular component abundance values ​​for the multiple cellular components for the given exposure concentration 116 / exposure time 118 pair.

[0131] Reference box 214, in some embodiments, each of the multiple cellular components is a specific gene, a specific mRNA associated with the gene, a carbohydrate, a lipid, an epigenetic feature, a metabolite, a protein, or a combination thereof.

[0132] In some embodiments, cellular components are genes, gene products (e.g., mRNA and / or proteins), carbohydrates, lipids, epigenetic traits, metabolites, and / or combinations thereof. In some embodiments, each of the multiple cellular components is a specific gene, a specific mRNA associated with the gene, carbohydrates, lipids, epigenetic traits, metabolites, proteins, or combinations thereof. In some embodiments, the multiple cellular components include nucleic acids, which include DNA, modified (e.g., methylated) DNA, RNA including encoding (e.g., mRNA) or non-coding RNA (e.g., sncRNA), proteins including post-transcriptionally modified proteins (e.g., phosphorylated, glycosylated, myristized, etc. proteins), lipids, carbohydrates, nucleotides including cyclic nucleotides (e.g., cyclic adenosine monophosphate (cAMP) and cyclic guanosine monophosphate (cGMP)) (e.g., adenosine triphosphate (ATP), adenosine diphosphate (ADP), and adenosine monophosphate (AMP)), other small molecule cellular components, such as oxidized and reduced forms of nicotinamide adenine dinucleotide (NADP / NADPH), and any combination thereof.

[0133] In some embodiments, the first plurality of cells comprises at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 500,000, or at least 1 million cells. In some embodiments, the first plurality of cells comprises no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 500, no more than 200, no more than 100, or no more than 50 cells. In some embodiments, the first plurality of cells comprises 5 to 100, 10 to 50, 20 to 500, 200 to 10,000, 1,000 to 100,000, 50,000 to 500,000, or 10,000 to 1 million cells. In some embodiments, the first plurality of cells falls into another range that begins with no less than 5 cells and ends with no more than 5 million cells.

[0134] Reference box 216, in some embodiments, the first plurality of cells are primary human hepatocytes. In some embodiments, the primary human hepatocytes are dissolved in a certain percentage of DMSO.

[0135] Boxes 218-224. Referring to box 218, in some embodiments, the abundance values ​​of multiple cellular components are determined by single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data of the first plurality of cells.

[0136] Reference box 220, in some embodiments, multiple cellular component abundance values ​​are determined by batch RNA sequencing. For example, in some embodiments, for a given concentration 116 / exposure time pair 118, the corresponding cellular component abundance values ​​measured for each cell in cells exposed to a given concentration 116 for a given exposure time are averaged together.

[0137] Reference box 222, in some embodiments, the abundance values ​​of multiple cellular components are determined by scTag-seq. Reference box 224, in some embodiments, the abundance values ​​of multiple cellular components are determined by single-cell assays of transposase-accessible chromatin (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.

[0138] In some embodiments, the abundance of multiple cellular components in a first plurality of cells is preprocessed. In some embodiments, preprocessing includes one or more of the following: filtering, normalization, mapping (e.g., mapping to a reference sequence), quantification, scaling, deconvolution, cleaning, dimensionality reduction, transformation, statistical analysis, and / or aggregation.

[0139] For example, in some embodiments, multiple cellular components are filtered based on desired quality (e.g., the size and / or quality of nucleic acid sequences) or minimum and / or maximum abundance values ​​of the corresponding cellular components. In some embodiments, filtering is performed partially or entirely by various software tools (such as Skewer). See Jiang, H. et al., BMC Bioinformatics 15(182):1-12 (2014). In some embodiments, for example, sequencing data QC software (such as AfterQC, Kraken, RNA-SeQC, FastQC) or another similar software program is used to filter multiple cellular components for quality control. In some embodiments, multiple cellular components are normalized, for example, to account for pull-downs, amplification, and / or sequencing biases (e.g., plottibility, GC bias, etc.). See, for example, Schwartz et al., PLoS ONE 6(1):e16685 (2011) and Benjamini and Speed, Nucleic Acids Research 40(10):e72 (2012), the contents of which are hereby incorporated by reference for all purposes. In some embodiments, pretreatment removes a subset of cellular components from a first plurality of cellular components. In some embodiments, pretreatment of the corresponding abundance of the first plurality of cellular components improves (e.g., reduces) the high signal-to-noise ratio.

[0140] Therefore, in some embodiments, the corresponding abundance of a cellular component in a cellular component dataset may include any of a variety of forms, including but not limited to raw abundance values, absolute abundance values ​​(e.g., transcript quantity), relative abundance values ​​(e.g., relative fluorescence units, transcriptome analysis, and / or gene set expression analysis (GSEA)), complex or aggregated abundance values, transformed abundance values ​​(e.g., log2 and / or logarithmic transformation), variations relative to a reference (e.g., normal samples, matched samples, reference datasets, housekeeping genes, and / or reference standards) (e.g., fold change or logarithmic change), normalized abundance values, measures of central tendency (e.g., mean, median, mode, weighted mean, weighted median, and / or weighted mode), measures of dispersion (e.g., variance, standard deviation, and / or standard error), or adjusted abundance values ​​(e.g., normalized, scaled, and / or error-corrected).

[0141] Abundance values ​​for each cellular component in a cellular component dataset can be obtained using any of a variety of abundance counting techniques (e.g., cellular component measurement techniques).

[0142] In some embodiments, one or more methods are used to determine the abundance of each cellular component in a dataset of multiple cellular components, including microarray analysis of fluorescence, chemiluminescence, electrical signal detection, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), digital droplet PCR (ddPCR), solid-state nanopore detection, RNA switch activation, Northern blotting, and / or serial analysis of gene expression (SAGE).

[0143] In some embodiments, the abundance of each of the multiple cell components in a cell component dataset is determined by colorimetric measurement, fluorescence measurement, luminescence measurement, or resonance energy transfer (FRET) measurement.

[0144] In some embodiments, cellular component abundance data are gene expression data. In some embodiments, gene expression in a given number of cells can be measured by sequencing the RNA transcripts of the cells and then counting the amount of each gene transcript identified during sequencing. In some embodiments, the sequenced and quantified gene transcripts include RNA, such as mRNA. In some embodiments, the sequenced and quantified gene transcripts include downstream products of mRNA, such as proteins (e.g., transcription factors). Generally, as used herein, the term “gene transcript” can be used to refer to any downstream product of gene transcription or translation, including post-translational modifications, and “gene expression” and “cellular component abundance” can be used interchangeably to refer to any measurement of gene transcripts in general.

[0145] In some embodiments, the abundance of a corresponding cellular component among multiple cellular components is RNA abundance (e.g., gene expression), and the abundance of a corresponding cellular component is determined by measuring the polynucleotide levels of one or more nucleic acid molecules corresponding to a corresponding gene in multiple cells. The transcript level of the corresponding gene can be determined by the amount of mRNA or polynucleotides derived therefrom present in multiple cells. Polynucleotides can be detected and quantified by a variety of methods, including but not limited to microarray analysis, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), Northern blotting, serial analysis of gene expression (SAGE), RNA switch, RNA fingerprinting, ligase chain reaction, Qβ replicase, isothermal amplification methods, strand displacement amplification, transcription-based amplification systems, nuclease protection assays (Si nuclease or RNase protection assays), and / or solid-state nanopore detection.See, for example, Draghici, *Data Analysis Tools for DNA Microarrays*, Chapman & Hall / CRC Press, 2003; Simon et al., *Design and Analysis of DNA Microarray Investigations*, Springer Press, 2004; *Real-Time PCR: Current Technology and Applications*, edited by Logan, Edwards, and Saunders, Caister Academic Press, 2009; *Bustin AZ of Quantitative PCR* (IUL Biotechnology, Volume 5), International University Line, 2004; Velculescu et al., (1995) *Science* 270: 484-487; Matsumura et al., (2005) *Cell. Microbiol.* 7: 11-18; Serial Analysis of Gene Expression (SAGE): Methods and Protocols (Methods in Molecular Biology), Humana Press, 2008; Each of the references cited herein is hereby incorporated in its entirety by reference.

[0146] In some embodiments, the abundance of each of the multiple cellular components in the cellular component dataset 112 is obtained from expressed RNA or nucleic acids derived therefrom (e.g., cDNA or amplified RNA derived from cDNA incorporated into an RNA polymerase promoter) from multiple cells, including naturally occurring nucleic acid molecules and synthetic nucleic acid molecules. Thus, in some embodiments, the abundance of each of the multiple cellular components in the cellular component dataset is obtained from such non-limiting sources as total cellular RNA, poly(A)+ messenger RNA (mRNA) or fragments thereof, cytoplasmic mRNA, or RNA transcribed from cDNA (e.g., cRNA)). Methods for preparing total RNA and poly(A)+ RNA are well known in the art and are generally described, for example, in Sambrook et al., *Molecular Cloning: A Laboratory Manual* (3rd edition, 2001). RNA can be extracted from the cells of interest using guanidine thiocyanate lysis followed by CsCl centrifugation (see, for example, Chirgwin et al., 1979, Biochemistry 18:5294-5299), silica-based columns (e.g., RNeasy (Qiagen, Valencia, California) or StrataPrep (Stratagene, La Jolla, California), or phenol and chloroform, as described in Ausubel et al., ed., Current Protocols in Molecular Biology, Volume III, Green Publishing Associates, Inc., John Wiley & Sons, Inc., New York, pp. 13.12.1-13.12.5). Poly(A)+ can also be used. RNA can be selected, for example, by selection with oligomeric dT cellulose, or alternatively by reverse transcription initiated by oligomeric dT of total cellular RNA. RNA can be fragmented using methods known in the art, for example, by incubation with ZnCl2 to produce RNA fragments.

[0147] In some embodiments, the abundance of each cellular component in the multiple cellular components of the cellular component dataset 112 is determined by sequencing transcripts from multiple cells. In some embodiments, the abundance of each cellular component in the multiple cellular components of the cellular component dataset 112 is determined by single-cell RNA sequencing (scRNA-seq), scTag-seq, single-cell assay using transposase-accessible chromatin via sequencing (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, and any combination thereof.

[0148] Cellular component abundance measurement techniques can be selected based on the desired cellular component to be measured. For example, scRNA-seq, scTag-seq, and miRNA-seq can be used to measure RNA expression. Specifically, scRNA-seq measures the expression of RNA transcripts, scTag-seq allows the detection of rare mRNA species, and miRNA-seq measures the expression of microRNAs. CyTOF / SCoP and E-MS / Abseq can be used to measure protein expression in cells. CITE-seq simultaneously measures both gene and protein expression in cells, and scATAC-seq measures chromatin conformation in cells. Table 1 below provides example protocols for performing each of the cellular component abundance measurement techniques described above.

[0149] Table 1 - Example Measurement Scheme

[0150] Box 226. Referring to box 226, in some embodiments, the multiple cellular components consist of 100 to 35,000 cellular components. In some embodiments, the plurality of cellular components comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, or more than 50,000 cellular components. In some embodiments, the plurality of cellular components comprises no more than 70,000, no more than 50,000, no more than 30,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 500, no more than 200, no more than 100, no more than 90, no more than 80, no more than 70, no more than 60, no more than 50, or no more than 40 cellular components. In some embodiments, the plurality of cellular components consists of twenty to 10,000 cellular components. In some embodiments, the plurality of cellular components consists of 100 to 35,000 cellular components. In some embodiments, the plurality of cellular components comprises 5 to 20, 20 to 50, 50 to 100, 100 to 200, 200 to 500, 500 to 1000, 1000 to 5000, 5000 to 10,000, or 10,000 to 50,000 cellular components. In some embodiments, the plurality of cellular components falls into another range that begins with at least 5 cellular components and ends with at least 70,000 cellular components.

[0151] As an example, in some embodiments, multiple cellular components comprise multiple genes, optionally measured at the RNA level. In some embodiments, the multiple genes comprise at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 genes. In some embodiments, the multiple genes comprise at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 30,000, at least 50,000, or more than 50,000 genes. In some embodiments, the plurality of genes comprises 5 to 20, 20 to 50, 50 to 100, 100 to 200, 200 to 500, 500 to 1000, 1000 to 5000, 5000 to 10,000, or 10,000 to 50,000 genes.

[0152] As another example, in some embodiments, the multiple cellular components comprise multiple proteins. In some embodiments, the multiple proteins comprise at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 proteins. In some embodiments, the multiple proteins comprise at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 30,000, at least 50,000, or more than 50,000 proteins. In some embodiments, the plurality of proteins comprises 5 to 20, 20 to 50, 50 to 100, 100 to 200, 200 to 500, 500 to 1000, 1000 to 5000, 5000 to 10,000, or 10,000 to 50,000 proteins.

[0153] In some embodiments, the cell component dataset 112 includes cell component abundance values ​​122 for multiple instances of multiple cell components. For example, in some embodiments, the cell component dataset 112 includes a separate set of cell component abundance values ​​122 for multiple cell components in each of multiple cells.

[0154] In some embodiments, the cell component dataset 112 includes a separate set of cell component abundance values ​​122 for multiple cell components in each subset of cells from a plurality of cells having the same concentration 116 / exposure time pair 118. In some such embodiments, for each corresponding cell component, the measured cell component abundance values ​​of the corresponding cell component in each cell within the corresponding subset of cells having the same concentration 116 / exposure time pair 118 are averaged together to form the cell component abundance value for the corresponding cell component.

[0155] In some embodiments, the cell component dataset 112 comprises a separate set of cell component abundance values ​​122 for multiple cell components from each subset of cells in a plurality of cells from identical replicas (e.g., identical wells and / or cell aliquots) having the same concentration 116 / exposure time pair 118. In some such embodiments, for each corresponding cell component, the measured cell component abundance values ​​for the corresponding cell component in each cell from the corresponding subset of cells in the same replica having the same concentration 116 / exposure time pair 118 are averaged together to form the cell component abundance value for the corresponding cell component. Thus, if three replicas are present, this will result in three different measurements of the same cell component for a given concentration 116 / exposure time pair 118.

[0156] Box 228. Referring to box 228, in some embodiments, a cell component dataset 112 is used to determine the corresponding pathway value for each of the plurality of pathway modules 142, thereby obtaining a plurality of pathway values.

[0157] Each corresponding pathway module 142 in the multiple pathway modules contains an independent subset of various cellular components. For example, reference Figure 1B Independent subsets of the multiple cellular components of pathway module 142-1 consist of cellular components 144-1-1, 144-1-2, ..., 144-1-P, while independent subsets of the multiple cellular components of pathway module 142-2 consist of cellular components 144-2-1, 144-2-2, ..., 144-1-W, where W and P are the same or different positive integers. Therefore, each independent subset of the multiple cellular components of each corresponding pathway module contains two or more cellular components. Furthermore, each cellular component in the multiple cellular components (of cellular component dataset 112) is located in at least one independent subset of the multiple cellular components associated with pathway module 142 in the plurality of pathway modules. In some embodiments, the cellular components in the multiple cellular components are located in 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, or 15 or more pathway modules in the plurality of pathway modules.

[0158] In some embodiments, the plurality of cellular components comprises 10 or more cellular components. According to various embodiments of this disclosure, block 226 above provides an additional suitable number of cellular components among the plurality of cellular components.

[0159] In some embodiments, the plurality of cellular component abundance values ​​comprises 5000 or more cellular component abundance values. In some embodiments, the plurality of cellular component abundance values ​​comprises 10 or more, 50 or more, 100 or more, 250 or more, 500 or more, 1000 or more, 2000 or more, 5000 or more, 10,000 or more, or 15,000 or more cellular component abundance values. In some such embodiments, each cellular component is a different gene.

[0160] It is not required that each cellular component in a pathway module be unique. For example, consider the case where pathway module A contains cellular components 1, 3, and 10. Other pathway modules among multiple pathway modules can also contain these cellular components. Here, the term "independent" means that a subset of the multiple cellular components in a particular pathway module is unique as a whole. Therefore, considering the example pathway module A above, another pathway module among multiple pathway modules can contain cellular components 1, 3, and 10, provided that it further contains other cellular components not present in pathway module A. Further considering the example pathway module A above, another pathway module among multiple pathway modules can have cellular components 1, 3, and 10, provided that it further contains other cellular components not present in pathway module A.

[0161] In some embodiments, each of the plurality of pathway modules contains the same or different numbers of cellular components. In some embodiments, each corresponding independent subset of cellular components corresponding to each respective pathway module is a unique subset of cellular components (e.g., non-overlapping, wherein each of the plurality of cellular components is grouped into no more than one module). In some embodiments, a first pathway module has a first subset of cellular components that overlaps with a second subset of cellular components corresponding to a second pathway module (e.g., overlapping, wherein at least one of the plurality of cellular components is common to two or more different pathway modules).

[0162] In some embodiments, one or more pathway modules 142 have Wikipathway reference numbers. See Martens et al., 2021, “WikiPathway: connecting communities,” Nucleic Acids Research 8:49(D1) D613-D621, and Pico et al., 2008, “WikiPathway: Pathway Editing for the People,” PLoS Biol, each of which is hereby incorporated by reference. Such Wikipathway reference numbers are in Figure 8A and Figure 8B As shown in [the image]. For example, in [the image]. Figure 8A In the WikiPathway section, one pathway module in Module 142 is the "Metapathological Biotransformation Phases I and II" biopathway, with the reference number WP702. Independent subsets of various cellular components constituting the WP702 pathway were found in: https: / / www.wikipathways.org / pathways / WP702.html. They are: GSTA5, CYP4B1, NDST4, GSTA4, AKR7A2, NAT12, NAT1, SULT1C2, UGT1A5, NAT2, UGT1A4, CYP39A1, CYP27A1, CYP2D6, NDST2, CYP2U1, CYP450, FMO5, NDST3, HS6ST1, GSTT2, UGT1A7, CYP4F3, AKR1C2, CYP4X1, HS3ST1, MGST2, HS3ST3A1, CYP2A7, CYP4F22, GPX1, CHST7, TPMT, GLYAT, HNMT, BAAT, FMO3, CYP2B6, CYP2C19, CHST1, GSTA1, CYP24A1, CHST3, UGT1A10, CYP4F12, UGT2B4, SULT1B1, UGT1A9, CYP27C1, CYP2A6, SULT1C4, CYP46A1, CYP2C9, CYP4V2, NAT6, GSTM4, GPX2, CYP11B1, CYP2C18, CYP4F11, GLYATL1, UGT1A1, CYP1A2, AKR1B1, HS3ST2, NAT14, GPX3, CYP1B1, GLYATL2, GSTO1, CYP20A1, CYP2S1, SULT1E1, CYP19A1, CYP21A2, AKR7A3, CYP26C1, SULT2A1, HS6ST2, CYP27B1, CYP26B1, GSTT1, AKR1A1, FMO1, UGT2B7, NAT5, UGT2A2, INMT, NNMT, GSTM2, AKR1C3, CYP2W1, NAT10, CHST11, CHST9, SULT1C3, GAL3ST4, CYP3A7, UGT2B15, UGT2B11, CYP7A1, GSTZ1, CYP4F2, CHST2, CYP2F1, GAL3ST1, HS3ST5, GSTK1, GAL3ST3, HS2ST1, CYP8B1, CYP1A1, GSR, SULT4A1, HS6ST3, CYP11B2, GSTM5, CYP4Z1, AKR1C4, UGT2A1, NDST1, AKR1C1, KCNAB1, GSTCD, CHST13, CYP7B1, CYP11A1, CYP2C8, UGT2A3, CHST10, SULT1C1, EPHX1, NAT9, NAT8L, SULT1A2, NAT13, HS3ST4, UGT1A6, GPX4, CYP2R1, HS3ST3B1, SULT2B1, GSS, GPX5, glutathione transferase, CYP4F8,EPHX2, CHST5, UGT2B28, CYP3A4, GSTM1, KCNAB3, FMO2, GSTA3, KCNAB2, NAT11, FMO4, CYP26A1, SULT1A1, CYP2E1, CHST14, SULT6B1, HS3ST6, CYP17A1, CHST4, MGST1, UGT2B17, SULT1A4 , KR1B10, NAT8, CYP3A43, SULT1A3, UGT1A8, CYP3A5, GSTA2, GSTP1, COMT, CHST6, UGT1A3, CYP2J2, MGST3, AKR1D1, GSTO2, GSTM3, GAL3ST2, GSTT2B, CHST8, CYP2A13, CYP51A1, and CHST12. ,

[0163] In another instance, in Figure 8AIn Wikipathways, one of the pathway modules in Module 142 is the "Tamoxifen metabolism" biological pathway, and its reference number is WP691. Independent subsets of the various cellular components that make up the WP691 pathway are found at: https: / / www.wikipathways.org / pathways / WP691.html. They are: CYP2D6, CYP2A6, SULT2A1, CYP2C9, CYP3A5, CYP2C19, CYP3A4, CYP3A4, UGT1A10, CYP3A4, SULT2A1, SULT2A1, CYP1A2, CYP3A4, UGT2B7, CYP2D6, SULT1E1, SULT1A1, FMO1, CYP2C9, CYP2E1, SULT1A1, FMO3, SULT1E1, CYP3A4, UGT1A4, CYP3A5, CYP2C19, UGT1A4, CYP1B1, UGT1A4, UGT1A8, CYP2C19, CYP3A4, CYP2D6, CYP1A1, CYP2D6, CYP2C8, UGT2B15, and CYP3A4. In addition to these genes, the WP691 pathway has numerous metabolites, such as amoxifen-N-oxide, cis-4-hydroxytamoxifen, and α-hydroxy-N-demethyltamoxifen. In some embodiments of the Wikipathway that include both genes and metabolites, a separate subset of the cellular components of the corresponding pathway module 142 includes both metabolites and genes. In some embodiments, a separate subset of the cellular components of the corresponding pathway module 142 includes only genes. In some embodiments, the corresponding pathway module 142 of the Wikipathway consists only of a subset of the genes in the corresponding Wikipathway (e.g., five or more, ten or more, or three to 30 genes in the corresponding Wikipathway). In some embodiments, the corresponding pathway module 142 of the Wikipathway consists only of those genes in the corresponding Wikipathway that are differentially expressed across multiple annotated cell states. Such annotated cell states and differential expression are described in box 236 below.

[0164] In some embodiments, the plurality of pathway modules include Figure 8A and Figure 8B The list includes 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more Wikipathways.

[0165] In some embodiments, one or more pathway modules in pathway module 142 are KEGG pathways. For example, in some embodiments, a subset of the cellular components of such pathway modules consists of cellular components in a corresponding KEGG pathway. In some embodiments, one or more pathway modules in pathway module 142 are human KEGG pathways. Such pathways are disclosed in https: / / www.genome.jp / kegg / pathway.html In some embodiments, one or more pathway modules in pathway module 142 are Reactome pathways. For example, in some embodiments, a subset of the cellular components of such pathway modules consists of cellular components in the corresponding Reactome pathway. In some embodiments, one or more pathway modules in pathway module 142 are human Reactome pathways. In some embodiments, the corresponding pathway module 142 of a Reactome pathway consists only of a subset of genes in the corresponding Reactome pathway (e.g., five or more, ten or more, or three to 30 genes of the Reactome pathway). In some embodiments, the corresponding pathway module 142 of a Reactome pathway consists only of those genes in the corresponding Reactome pathway that are differentially expressed across multiple annotated cellular states. Such differential expression and annotated cellular states are described in box 236 below. Reactome pathways are disclosed in https: / / reactome.org / See also Gillespie et al., 2022, “The reactome pathway knowledgebase 2022”, Nucleic Acid Research 50(D1):D687-D692, which is hereby incorporated by reference.

[0166] In some embodiments, one or more pathway modules in pathway module 142 are Pathway Commons pathways. For example, in some embodiments, a subset of the cellular components of such a pathway module consists of cellular components in the corresponding Pathway Commons pathway. In some embodiments, one or more pathway modules in pathway module 142 are human Pathway Commons pathways. In some embodiments, the corresponding pathway module 142 of a Pathway Commons pathway consists only of a subset of genes in the corresponding Pathway Commons pathway (e.g., five or more, ten or more, or three to 30 genes of a Pathway Commons pathway). In some embodiments, the corresponding pathway module 142 of a Pathway Commons pathway consists only of those genes differentially expressed across multiple annotated cellular states in the corresponding Pathway Commons. Such differential expression and annotated cellular states are described in box 236 below. Pathway Commons pathways are disclosed in https: / / www.pathwaycommons.org / See also Rodchenkov et al., 2020, “Pathway Commons 2019 Update: integration, analysis and exploration of pathway data,” Nucleic Acid Research, Vol. 48, No. D1, January 8, 2020, pp. D489–D497, which are hereby incorporated by reference.

[0167] Referring to box 230, in some embodiments, each independent subset of the multiple cellular components in each corresponding pathway module comprises two to three hundred cellular components. In some embodiments, the independent subset of the multiple cellular components in a corresponding pathway module comprises three, four, five, or more cellular components. In some embodiments, the independent subset of the multiple cellular components in a corresponding pathway module in multiple pathway modules comprises at least 2, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 cellular components. In some embodiments, the independent subset of the multiple cellular components comprises no more than 500, no more than 200, no more than 100, no more than 90, no more than 80, no more than 70, no more than 60, or no more than 50 cellular components. In some embodiments, the independent subset of the multiple cellular components comprises 5 to 100, 2 to 300, 20 to 500, or 200 to 1000 cellular components. In some embodiments, the independent subset of the multiple cellular components falls into another range that begins with at least 2 cellular components and ends with at least 2000 cellular components.

[0168] In some embodiments, an independent subset of multiple cellular components in a corresponding pathway module comprises cellular components from cellular processes (e.g., molecular pathways). For example, in some embodiments, the independent subset of multiple cellular components in a corresponding pathway module comprises two to 20 cellular components from molecular pathways associated with DILI. According to some such embodiments, Figure 8A and Figure 8B Some of these molecular pathways were named.

[0169] Referring to reference box 232, in some embodiments, each corresponding pathway value of each corresponding pathway module 142 is a measure of the central tendency of the abundance of each cellular component in an independent subset of cellular components corresponding to a corresponding subset of cellular components in the cross-cellular component dataset. As used herein, the term "cross-cellular component dataset" means the measured abundance of an independent subset of cellular components in cells exposed to a specific concentration 116 of a compound for a specific exposure time 118. If the cellular component dataset includes data at the same exposure time at four different concentrations, then in such embodiments, the corresponding pathway module 142 will have four distinct pathway values, one for each of the four different exposure times. If the cellular component dataset includes data at four different exposure times, each at two different concentrations, then in such embodiments, the corresponding pathway module 142 will have eight distinct pathway values, one for each of eight unique combinations of exposure time and concentration.

[0170] In some embodiments, the measure of central tendency is the mean, median, mode, weighted mean, weighted mode, and weighted mode of the measured abundance of each cellular component in an independent subset of cellular components corresponding to a corresponding subset of cellular components of a respective pathway module. For example, in the case where there are five cellular components in an independent subset of cellular components of a given pathway module, in such embodiments, the measure of central tendency of the measured cellular component abundance of the five cellular components (at a given concentration 116 and exposure time 118) is used as the pathway value of the corresponding pathway module (at a given concentration 116 and exposure time 118).

[0171] Referring to reference box 234, in alternative embodiments, each corresponding pathway value for each corresponding pathway module is based on a p-value. In some such embodiments, multiple cellular components are genes. Furthermore, in some embodiments where multiple cellular components are genes and pathway enrichment analysis is used to obtain the corresponding pathway value for each of multiple pathway modules. Pathway enrichment analysis is a statistical method for determining whether a particular pathway module is overexpressed or enriched with differentially expressed genes (DEGs) in RNA-seq data. In such embodiments, the pathway value for a corresponding pathway module is determined by first using a statistical test to determine the p-value of the corresponding pathway module through pathway enrichment analysis. This statistical test determines the significance of the enrichment of DEGs observed within the corresponding pathway module. In some embodiments, a hypergeometric test is used to calculate the p-value of the pathway module. The hypergeometric test determines the probability of incidentally obtaining the number of DEGs observed in a particular pathway module given the total number of genes in a given pathway and the total number of DEGs in the entire dataset. To convert the p-value of a pathway module into a pathway value for the pathway module, in some embodiments, the negative logarithm or false discovery rate (FDR) of the adjusted p-value is calculated and used as the pathway value. FDR corrects for multiple testing in pathway enrichment analysis to control for false positives. Therefore, in more detail, in some embodiments using the hypergeometric test, the following parameters are defined:

[0172] ●N: The number of genes in various cellular components (e.g., approximately 30,000 in some embodiments).

[0173] ●n: The number of genes annotated to a specific pathway module (e.g., in some embodiments, a specific pathway module contains genes for a single pathway of interest).

[0174] ●K: The total number of DEGs across multiple cellular components (e.g., genome / dataset), and

[0175] ●k: The total number of DEGs annotated to a specific pathway module t.

[0176] In some embodiments, the hypergeometric probability formula is used to calculate the probability of precisely observing k DEGs in a pathway module out of a total of n genes in the pathway module and the probability of precisely observing K DEGs out of a total of N genes in the dataset.

[0177] In some embodiments, multiple test correction is applied to the p-values ​​of pathway modules. In some embodiments, multiple test correction is a Benjamini-Hochberg procedure. Multiple test correction adjusts the p-values ​​of multiple comparisons. The result is a set of adjusted p-values ​​or FDR (false discovery rate) values, where each adjusted p-value represents the significance of the corresponding pathway module across multiple pathway modules, taking into account the number of pathway modules tested.

[0178] In some embodiments, to convert FDR values ​​to pathway values, the negative logarithm (base 10) of each FDR value is taken. In some embodiments, this conversion helps to highlight significant enrichment while linearizing the data. In some embodiments, the formula -log10(FDR) is used to calculate the negative logarithm (base 10). The resulting -log10(FDR) value represents the pathway value of multiple pathway modules. A higher -log10(FDR) value indicates a more significant enrichment of DEG in the pathway module.

[0179] Referring to reference box 236, in some embodiments, the method further includes identifying multiple pathway modules. That is, in some embodiments of this disclosure, the pathway modules consist of cellular components of pathways not recorded in publicly managed databases such as Wikipathway, KEGG, Reactome, or Pathway Commons.

[0180] In some embodiments, the pathway module comprises at least 20, at least 30, at least 40, at least 50, or all 64 genes: RNF187, CHAC1, CARS, GPD1L, AARS, CSTA, TRIB3, LONP1, RHBDD1, ANKRD30BL, FGF21, PIR, FXYD5, ARL6IP6, ASNS, PHGDH, SLC31A1, SLC7A1, NARS, MTHFD2, SARS, JAK1, XPOT, PSAT1, EIF4EBP1, MOCOS, GARS, SLC 7A11, SESN2, RHOF, SLFN5, PCK2, MKX, AGPAT9, LCN2, STC2, MKNK2, TOM1L1, DDIT4, LOC100128822, RSL24D1, YARS, M6PR, CTH, ANO10, C19orf80, ECI1, FBXO25, NCF2, FGF2, TMEM206, PRNP, TUBE1, MARS, EFHD1, LRRC8B, PPM1E, PVR, SLC43A1, EMP1, SLC7A5, ME1, UNC5B, and ATF5. This set of genes does not correspond to any pathways documented in Wikipathway, KEGG, Reactome, or Pathway Commons.

[0181] In some embodiments, pathway modules are identified by obtaining one or more first datasets (e.g., in electronic form). For each corresponding primary human hepatocyte among a plurality of primary human hepatocytes, for each corresponding cellular component among a plurality of cellular components, the one or more first datasets contain: the corresponding abundance of the corresponding cellular component in the corresponding primary human hepatocyte. In some such embodiments, the plurality of primary human hepatocytes comprises 20, 30, 40, 50, 100, or 100 or more primary human hepatocytes and collectively represents a plurality of annotated cell states. Based on this data, the plurality of cellular components from one or more first datasets are identified as a plurality of differentially expressed cellular components across one or more first datasets. Differentially expressed cellular components are those that express differently across a plurality of annotated cell states. Differentially expressed cellular components are sought because they provide more information for determining DILI toxicity than cellular components that express unchanged across a plurality of annotated cell states.

[0182] Several measurement and statistical methods can be used to assess the differential expression of cellular components.

[0183] One such metric is the fold change. The fold change measures the ratio of the expression levels of a cellular component between two conditions. It is calculated as the ratio of the expression level of a cellular component under one condition to the expression level under another. A fold change of 1 indicates no differential expression, while values ​​greater than 1 or less than 1 indicate upregulation or downregulation, respectively. When testing more than two conditions, a measure of the central tendency of the fold changes for each unique pair of conditions can be used to determine differential expression of a given cellular component. The cellular components can then be ranked according to the absolute values ​​of these fold changes, and the top N cellular components can be considered differentially expressed based on their absolute fold changes, where N is a positive integer such as 50, 100, 500, or a number between 10 and 2000.

[0184] Another such measure is the logarithmic fold change. The logarithmic fold change measure is the logarithm of the ratio of the expression levels of a cellular component between two conditions. It is calculated as the ratio of the expression level of a cellular component under one condition to its expression level under another. When testing more than two conditions, a measure of the central tendency of the logarithmic fold change for each unique pair of conditions can be used to determine the differential expression of a given cellular component. Cellular components can then be ranked according to the absolute values ​​of these logarithmic fold changes, and the top N cellular components can be considered differentially expressed based on their absolute logarithmic fold changes, where N is a positive integer such as 50, 100, 500, or a number between 10 and 2000.

[0185] Another such metric is the p-value. The p-value is a statistical measure of the probability that differential expression of a cellular component observed only by chance across multiple annotated cell states is true. It can be calculated, for example, using statistical tests such as t-tests, ANOVA, or nonparametric tests. A lower p-value indicates a higher level of statistical significance and suggests that differential expression observed across multiple annotated cell states is unlikely to be due to random chance. Cellular components can then be ranked according to their p-values, and the top N cell components (indicating the N cell components with the smallest p-values) can be considered differentially expressed cell components, where N is a positive integer, such as 50, 100, 500, or a number between 10 and 2000. In some embodiments, a adjusted p-value is used instead of a p-value. For example, in some embodiments, the Benjamini-Hochberg p-value is used to control for the false discovery rate. In other embodiments, the Bonferroni-corrected p-value is used to control for the family error rate (FWER).

[0186] Once differentially expressed cellular components are identified, they can be assigned to pathway modules. In some such embodiments, the differentially expressed cellular components are compared to a set of pathway-associated cellular components in publicly managed databases such as Wikipathway, KEGG, Reactome, or Pathway Commons. When a subset of differentially expressed cellular components (e.g., at least 3, 4, 5, 6, 7, 8, 9, or 10 differentially expressed cellular components) is found to be present in such a pathway, the pathway is selected as one of pathway modules 142.

[0187] In some embodiments, each pathway module in the pathway module represents a cluster of differentially expressed cellular components. In some embodiments, such clusters are identified by obtaining one or more first datasets in electronic form, the one or more first datasets comprising: for each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes, and collectively representing a plurality of annotated cellular states: for each corresponding cellular component among a plurality of cellular components, the corresponding abundance of the corresponding cellular component in the corresponding primary human hepatocyte. In this manner, a plurality of vectors are obtained or formed. Each corresponding vector (i) corresponds to a corresponding cellular component among a plurality of (differentially expressed) cellular components, and (ii) contains a plurality of corresponding elements, each of the plurality of corresponding elements representing the abundance of the corresponding cellular component in the corresponding primary human hepatocyte among the plurality of primary human hepatocytes. The plurality of vectors are clustered, thereby forming a plurality of clusters. Each of the plurality of clusters consists of an independent subset of the plurality of cellular components of the corresponding cellular component module in the plurality of cellular component modules.

[0188] In some embodiments, each pathway module in the pathway module represents a cluster of differentially expressed cellular components. In some embodiments, such clusters are identified by obtaining one or more first datasets in electronic form, the one or more first datasets comprising: for each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes, and collectively representing multiple annotated cell states; for each corresponding cellular component in a plurality of cellular components; and the corresponding abundance of the corresponding cellular component in the corresponding primary human hepatocyte. Thus, a shared cellular component module matrix is ​​generated using one or more first datasets using shared nonnegative matrix factorization. Then, using multiple least squares regression of the matrix relative to the shared GEP, independent subsets of the plurality of cellular components in each corresponding cellular component module in the plurality of cellular component modules are identified from the shared matrix.

[0189] In some such embodiments, each independent subset of cellular components from a plurality of cellular components is associated with a corresponding pathway from a plurality of pathway modules. For example, in one embodiment, the independent subset of cellular components identified using the above-described common matrix method is: ARSE, ARG1, ANKS4B, ABHD2, KANK1, TACSTD2, SALL1, LPIN2, SEC16B, CEBPD, NAV2, ERBB3, C1orf115, PLA2G12B, ABCB4, CYP39A1, TAT, METTL7A, RPRD1B, PPAP2A, CYP8B1, MTCP1NB, DEPDC7, DUSP16, RCL1, ANKRD35, KL HDC2, ABCG2, CIDEC, C11orf54, C11orf75, RAB17, PEBP1, PET117, AKAP12, TBX15, GALM, LEAP2, CROCCP2, RNASE4, IRF2, SORBS1, TMEM37, CYP27A1, CLRN3, RAPGEF5, AQP3, C2orf72, PCCA, HSD17B11, NSMCE2, RHOB, FAM134B, AGTR1, AADAC, HMGB3, PANK1, and HPGD. Analysis of this subset of cellular components revealed that some of these components are present in the KEGG primary bile acid biosynthesis pathway. Therefore, the pathway modules of this subset of cellular components are associated with the KEGG primary bile acid biosynthesis pathway. The importance of this allocation is to identify the biological basis of the clusters of cellular components allocated to pathway modules in such embodiments. However, in embodiments that identify independent subsets of cellular components through clustering, it is not required that such subsets be associated with known molecular pathways. In fact, such clusters may represent undiscovered molecular mechanisms associated with DILI.

[0190] Referring to reference box 238, in some embodiments, the annotated cell states of a plurality of annotated cell states are cells in a plurality of primary human hepatocytes exposed to the training compound under exposure conditions (e.g., exposure duration, concentration of the training compound, or a combination of exposure duration and concentration of the training compound). The Examples section below describes the exposure of cells to 120 different such training compounds.

[0191] Reference box 240, in some embodiments, multiple annotated cell states include multiple different concentrations (e.g., the multiple different concentrations in (i) the C of the first training compound). max Or estimated C max (ii) IC of the first training compound 10(within the range of) exposure to the first training compound. In some embodiments, the multiple different concentrations are 2, 3, 4, 5, 6, 7, 8, 9, or 10 different concentrations in the range of 0.1 µM to 1 µM. In some embodiments, the multiple different concentrations are 2, 3, 4, 5, 6, 7, 8, 9, or 10 different concentrations in the range of 0.01 µM to 1 µM. In some embodiments, the multiple different concentrations are 2, 3, 4, 5, 6, 7, 8, 9, or 10 different concentrations in the range of 0.1 µM to 10 µM. In some embodiments, the lowest concentration among the multiple different concentrations is 0.001 µM, 0.01 µM, 0.01 µM, or 1 µM. In some embodiments, the highest concentration used among the multiple different concentrations is 100 µM, 50 µM, 10 µM, 1 µM, 0.5 µM, or 0.1 µM.

[0192] Referring to box 242, in some embodiments, multiple annotated cell states include exposure to a first training compound for multiple different exposure durations. In some embodiments, multiple annotated cell states include exposure to multiple different training compounds at the same concentration and for multiple different exposure durations. In some embodiments, multiple annotated cell states include exposure to multiple different training compounds at several different concentrations and for several different exposure durations.

[0193] In some embodiments, the plurality of annotated states includes 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 100 or more, 200 or more, or 1000 or more annotated cell states.

[0194] Reference box 244, in some embodiments, multiple annotated cell states include corresponding multiple different concentrations (e.g., the corresponding multiple different concentrations in (i) the C of the corresponding training compound). max (ii) IC of the corresponding training compound 10 (within the range) exposed to each corresponding training compound in a variety of training compounds.

[0195] In some embodiments, the annotated cell state is the cell being exposed to the training compound under exposure conditions (e.g., exposure duration, compound concentration, or a combination of exposure duration and compound concentration).

[0196] Reference box 246, in some embodiments, multiple annotated cell states include each training compound exposed to multiple training compounds for multiple different exposure durations (e.g., 4 hours, 8 hours, 16 hours, 24 hours or some subset thereof).

[0197] Referring to block 248, in some embodiments, the plurality of pathway modules comprises 10 to 2000 pathway modules. In some embodiments, the plurality of pathway modules includes five or more pathway modules. In some embodiments, the plurality of pathway modules includes 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more pathway modules.

[0198] In some embodiments, the plurality of pathway modules comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, or at least 3000 pathway modules. In some embodiments, the plurality of pathway modules comprises no more than 2000, no more than 1000, no more than 500, no more than 300, no more than 200, no more than 100, no more than 90, no more than 80, no more than 70, no more than 60, or no more than 50 pathway modules. In some embodiments, the plurality of pathway modules consists of 50 to 500 pathway modules. In some embodiments, the plurality of pathway modules comprises 5 to 20, 20 to 50, 50 to 100, 100 to 200, 200 to 500, 500 to 1000, or 1000 to 5000 pathway modules. In some embodiments, the plurality of cellular component modules fall into another range that begins with at least 5 cellular component modules and ends with at least 10,000 pathway modules.

[0199] Referring to box 249, in some embodiments, multiple pathway values ​​are input into each of one or more first models 130, thereby obtaining one or more first values ​​as first model outputs 134 from each of the one or more first models. That is, each first model outputs a first value, and pathway values ​​calculated for each of the multiple pathway modules 142 are simultaneously input into each of the one or more first models.

[0200] In some embodiments, each of the one or more first models is a logistic regression model trained using a training dataset of compounds, wherein the DILI risk for each compound in the training dataset is known. In some such embodiments, the logistic regression model includes different weights for each pathway module value among the pathway module values.

[0201] For example, the Examples section below describes training a first logistic regression model using gene expression data and toxicological information of 120 compounds from the Japan Toxicogenomics Project (TGP). This data includes in vitro gene expression data, toxicological information, and pathological data for these compounds. This data is available on the internet from... http: / / camda2021. bioinf.jku.at / doku.php / tgp_prepro Obtained from above.

[0202] In some embodiments, each of the first models 130 trained in one or more first models further applies regularization to each model parameter. For example, in some embodiments, regularization is performed by adding a penalty to the loss function, wherein the penalty is proportional to the value of the parameter in the first model. Such practices can produce more general models and reduce overfitting of the data. In some embodiments, regularization includes L1 or L2 penalty. For example, in some preferred embodiments, regularization includes L2 penalty on the lower and upper bound parameters. In some embodiments, regularization includes spatial regularization (e.g., determined based on prior and / or experimental knowledge) or dropout regularization. In some embodiments, regularization includes independently optimized penalty.

[0203] Generally, training one or more first models involves updating multiple parameters (e.g., weights 132-1, 132-2, ..., 132-K, where K is a positive integer) of a given first model 130 in response to such training data. In some embodiments, this is accomplished via backpropagation (e.g., gradient descent). First, forward propagation is performed, where input data (multiple pathway values ​​for each of the various training compounds) is received one by one into the first model 130, and a first model output 134 for each of these training compounds is computed based on an initial set of model parameters (e.g., weights 132). In some embodiments, model parameters (e.g., weights 132 and / or hyperparameters) are randomly assigned (e.g., initialized) for a given first model prior to training. In some embodiments, the values ​​of these parameters are transferred from previously saved parameters or from a pre-trained model (e.g., through transfer learning).

[0204] In some embodiments, backpropagation is then performed by computing the error gradient of a given first model, wherein the error is determined by computing a loss (e.g., error) based on the given first model output 134 (e.g., the predicted DILI risk) and the actual known DILI risk for each training compound. The given first model is then trained by updating the parameters (e.g., weights) by adjusting their values ​​based on the computed loss.

[0205] For example, in some general embodiments of machine learning, backpropagation is a method for training multiple weights of a model. Given multiple pathway values ​​as input, the output of an untrained model (e.g., the predicted DILI risk) is first generated using an arbitrarily chosen set of initial weights. An error is then calculated by evaluating an error function (e.g., using a loss function), comparing the output (predicted DILI risk) of each training compound with the original input (e.g., the actual DILI risk). The weights are then updated such that the error is minimized (e.g., according to the loss function). In some embodiments, any of various backpropagation algorithms and / or methods are used to update multiple weights, as will be apparent to those skilled in the art.

[0206] In some embodiments, the loss function is mean squared error, quadratic loss, mean absolute error, mean bias error, hinge, multi-class support vector machine, and / or cross-entropy. In some embodiments, training a given first model includes calculating errors according to a gradient descent algorithm and / or a minimization function. In some embodiments, training a given first model includes calculating multiple errors using multiple loss functions. In some embodiments, each of the multiple loss functions receives the same or different weighting factors.

[0207] In some embodiments, an error function is used to train a given first model by updating one or more parameters (e.g., weights) in a given first model by adjusting the values ​​of one or more parameters by an amount proportional to a calculated loss. In some embodiments, the amount of parameter adjustment is measured by a learning rate hyperparameter, which indicates the extent or severity of the parameter update (e.g., a small or large adjustment). Therefore, in some embodiments, training updates all or a subset of multiple parameters based on the learning rate. In some embodiments, the learning rate is a differential learning rate.

[0208] In some embodiments, for each of a plurality of training instances, the training process is repeated, the training process including adjusting a plurality of parameters associated with a given first model (e.g., in response to the difference between the predicted DILI risk and the actual DILI risk).

[0209] In some embodiments, the plurality of training instances comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or at least 7500 training instances. In some embodiments, the plurality of training instances comprises no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, or no more than 50 training instances. In some embodiments, the plurality of training instances comprises 3 to 10, 5 to 100, 100 to 5000, or 1000 to 10,000 training instances. In some embodiments, the plurality of training instances falls into another range that begins with at least 3 training instances and ends with no more than 10,000 training instances.

[0210] In some such embodiments, training includes repeatedly adjusting the parameters 132 of a given first model 130 on multiple training instances (e.g., via backpropagation), thereby increasing the accuracy of the model in determining the risk of compound DILI.

[0211] In some embodiments, model training is performed on a given first model after the first evaluation of the error function. In some such embodiments, model training is performed on a given first model after updating one or more parameters for the first time based on the first evaluation of the error function. In some alternative embodiments, model training is performed on a given first model by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 10,000, 50,000, 100,000, 200,000, 500,000, or 1 million evaluations of the error function. In some such embodiments, model training is performed using at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 10,000, 50,000, 100,000, 200,000, 500,000, or 1 million iterations based on an error function. The given first model is trained by evaluating at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1 million updates of one or more parameters.

[0212] In some embodiments, model training is repeated until a given first model meets minimum performance requirements. For example, in some embodiments, after evaluating an error function (e.g., the difference between the predicted DILI risk and the actual DILI risk), model training is repeated until the error calculated for the trained model meets an error threshold. In some embodiments, the error calculated by the error function meets the error threshold when the error is less than 20%, less than 18%, less than 15%, less than 10%, less than 5%, or less than 3%.

[0213] In one example embodiment, a given first model is trained using the classification cross-entropy loss in a multi-task formulation, wherein each of the plurality of covariates corresponds to a cost function among the plurality of cost functions, and each corresponding cost function among the plurality of cost functions has a common weighting factor.

[0214] Reference box 250, in some embodiments, the given first model is a logistic regression model. The logistic regression model is disclosed in the following literature: Agresti, Introduction to Classification Data Analysis, 1996, Chapter 5, pp. 103-144, John Willie & Son Publishing, New York, NY, which is hereby incorporated by reference.

[0215] Box 252. Referring to box 252, in some embodiments, each of one or more first models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model. Details of such models are disclosed in the definitions section above.

[0216] In some embodiments, each of the first models 130 in one or more first models includes a plurality of parameters 132 (e.g., 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 200 or more, 300 or more, 500 or more, 1000 or more, or 10,000 or more).

[0217] In some embodiments, each first model 130 in one or more first models includes a plurality of parameters 132 (e.g., weights and / or hyperparameters). In some embodiments, the plurality of parameters of each first model in one or more first models includes at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, or at least 5 million parameters. In some embodiments, the plurality of parameters of each of the one or more first models comprises no more than 8 million, no more than 5 million, no more than 4 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5,000, no more than 1,000, or no more than 500. In some embodiments, the plurality of parameters of each of the one or more first models comprises 10 to 5,000, 500 to 10,000, 10,000 to 500,000, 20,000 to 1 million, or 1 million to 5 million parameters. In some embodiments, the plurality of parameters of each of the one or more first models falls into another range that begins with no less than 10 parameters and ends with no more than 8 million parameters.

[0218] In some embodiments, the training of each of the one or more first models is further characterized by one or more hyperparameters (e.g., one or more values ​​that can be tuned during training). In some embodiments, the hyperparameter values ​​are tuned (e.g., adjusted) during training. In some embodiments, the hyperparameter values ​​are determined based on specific features of the training dataset and / or one or more inputs (e.g., cells, cellular component modules, covariates, etc.). In some embodiments, experimental optimization is used to determine the hyperparameter values. In some embodiments, hyperparameter scans are used to determine the hyperparameter values. In some embodiments, hyperparameter values ​​are assigned based on previous templates or default values.

[0219] In some embodiments, a corresponding hyperparameter of one or more hyperparameters includes a learning rate. In some embodiments, the learning rate is at least 0.0001, at least 0.0005, at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1. In some embodiments, the learning rate is no more than 1, no more than 0.9, no more than 0.8, no more than 0.7, no more than 0.6, no more than 0.5, no more than 0.4, no more than 0.3, no more than 0.2, no more than 0.1, no more than 0.05, no more than 0.01, or less. In some embodiments, the learning rate is 0.0001 to 0.01, 0.001 to 0.5, 0.001 to 0.01, 0.005 to 0.8, or 0.005 to 1. In some embodiments, the learning rate falls into another range starting at a value not lower than 0.0001 and ending at a value not higher than 1. In some embodiments, one or more hyperparameters further include regularization strength (e.g., L2 weight penalty, dropout rate, etc.). For example, in some embodiments, each of one or more first models (e.g., a neural network) is trained using regularization of the corresponding parameters (e.g., weights) of each of a plurality of hidden neurons. In some embodiments, regularization includes L1 or L2 penalty.

[0220] In some embodiments, the corresponding hyperparameters of one or more hyperparameters are the identity (type) of the loss function. In some embodiments, the loss function is mean squared error, flat mean squared error, quadratic loss, mean absolute error, mean bias error, hinge, multi-class support vector machine, and / or cross-entropy. In some embodiments, the loss function is a gradient descent algorithm and / or a minimization function.

[0221] In some embodiments, each of the one or more first models is associated with one or more activation functions (e.g., in the case that the first model is a neural network). In some embodiments, the activation functions among the one or more activation functions are tanh function, sigmoid function, softmax function, Gaussian function, Boltzmann-weighted averaging function, absolute value function, linear function, modified linear unit (ReLU) function, bounded modified linear function, soft modified linear function, parameterized modified linear function, averaging function, max function, min function, sign function, square function, square root function, multivariate quadratic function, inverse quadratic function, inverse multivariate quadratic function, multiple harmonic spline function, swish function, mish function, Gaussian error linear unit (GeLU) function, and / or thin plate spline function. In some embodiments, each of the one or more first models outputs the predicted DILI risk in response to inputting a path value into the model.

[0222] Boxes 254-260. Referring to box 254, in some embodiments, multiple cellular component abundance values ​​122 are input into each of one or more second models 136 (e.g., 136-1, 136-2, ..., 136-T, where T is a positive integer), thereby obtaining one or more second values ​​(e.g., 140-1, 140-2, ..., 140-T) as second model outputs from one or more second models. Referring to box 256, in some embodiments, the one or more second models are a single logistic regression model. Referring to box 258, in some embodiments, the one or more second models are multiple logistic regression models. In some such embodiments, the multiple logistic regression models are four to 100 logistic regression models. Referring to box 260, in some embodiments, a first subset of the multiple logistic regression models in the multiple second models is a multiple L1 regression model, and a second subset of the multiple logistic regression models in the multiple second models is a multiple L2 regression model. In some such embodiments, the first subset of the multiple logistic regression models is two to fifty logistic regression models. In some such embodiments, a second subset of the multiple logistic regression models is two to fifty logistic regression models.

[0223] In some embodiments, each second model 136 is a logistic regression model trained using a training dataset of compounds, wherein the DILI risk of each compound in the training dataset is known. For example, the Examples section below describes training multiple second logistic regression models using gene expression data and toxicological information of 120 compounds from the Japan Toxicogenomics Project (TGP), said data including gene expression data, toxicological information, and pathological data of these 120 compounds. This data is available on the Internet from http: / / camda2021.bioinf.jku.at / doku.php / tgp_ prepro Obtained from above.

[0224] In some embodiments, training each of the one or more second models 136 further applies regularization to each model parameter 138 in each of the one or more second models. For example, in some embodiments, regularization is performed by adding a penalty to the loss function, wherein the penalty is proportional to the value of the parameter 138 in the second model 136. Such practices can produce more general models and reduce overfitting of the data. In some embodiments, regularization includes L1 or L2 penalty. For example, in some embodiments, the one or more second models 136 are multiple logistic regression models, wherein a first subset of the multiple logistic regression models is trained using regularization including L1 penalty on the lower bound and upper bound parameters, and a second subset of the multiple logistic regression models is trained using regularization including L2 penalty on the lower bound and upper bound parameters.

[0225] In some embodiments, model training of one or more second models includes regularization, which may include spatial regularization (e.g., determined based on prior and / or experimental knowledge) or dropout regularization. In some embodiments, regularization may include penalties for independent optimization.

[0226] Generally, training one or more second models involves updating multiple parameters (e.g., weights) of each of the one or more second models in response to such training data. In some embodiments, this is accomplished via backpropagation (e.g., gradient descent). First, forward propagation is performed, where input data (multiple cellular component abundance values ​​for each of the corresponding training compounds in a variety of training compounds) is received one by one into each of the one or more second models, and the model output for each of these training compounds from each of the one or more second models is computed based on an initial set of model parameters (e.g., weights, parameters 132) for each of the one or more second models. In some embodiments, before training, the model parameters (e.g., weights and / or hyperparameters) of each of the one or more second models are randomly assigned (e.g., initialized). In some embodiments, the model parameters are transferred from a previously saved set of parameters or from a pre-trained model (e.g., through transfer learning).

[0227] In some embodiments, backpropagation is then performed by computing the error gradient of each of the one or more second models, wherein the error is determined by computing a loss (e.g., error) based on the model output (e.g., the predicted DILI risk) of each of the one or more second models and the actual known DILI risk of each training compound. Then, the parameters 138 (e.g., weights) of each of the one or more second models 136 are updated by adjusting the parameter values ​​based on the computed loss, thereby training the one or more second models.

[0228] For example, in some general embodiments of machine learning, backpropagation is a method for training multiple weights of a model. Given multiple cellular component abundance values ​​as input, an untrained model output (e.g., the predicted DILI risk) is first generated using an arbitrarily chosen set of initial weights. An error is then calculated by evaluating an error function (e.g., using a loss function), comparing the output (predicted DILI risk) for each training compound with the original input (e.g., the actual DILI risk). The model weights are then updated to minimize the error (e.g., according to the loss function). In some embodiments, any of various backpropagation algorithms and / or methods are used to update multiple weights of each second model in a second model, as will be apparent to those skilled in the art.

[0229] In some embodiments, the loss function is mean squared error, quadratic loss, mean absolute error, mean bias error, hinge, multi-class support vector machine, and / or cross-entropy. In some embodiments, training one or more second models includes calculating errors according to a gradient descent algorithm and / or a minimization function. In some embodiments, training one or more second models includes calculating multiple errors using multiple loss functions. In some embodiments, each of the multiple loss functions receives the same or different weighting factors.

[0230] In some embodiments, an error function is used to train one or more second models by updating one or more parameters (e.g., weights) in each of the one or more second models by adjusting the values ​​of one or more parameters in each of the one or more second models by an amount proportional to a calculated loss. In some embodiments, the amount of parameter adjustment is measured by a learning rate hyperparameter, which indicates the extent or severity of the parameter update (e.g., small or large adjustment). Therefore, in some embodiments, training updates all or a subset of multiple parameters of each of the one or more second models based on the learning rate. In some embodiments, the learning rate is a differential learning rate.

[0231] In some embodiments, for each of a plurality of training instances, the training process is repeated, the training process including adjusting a plurality of parameters associated with each of one or more second models (e.g., in response to the difference between the predicted DILI risk and the actual DILI risk).

[0232] In some embodiments, the plurality of training instances comprises at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or at least 7500 training instances. In some embodiments, the plurality of training instances comprises no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 100, or no more than 50 training instances. In some embodiments, the plurality of training instances comprises 3 to 10, 5 to 100, 100 to 5000, or 1000 to 10,000 training instances. In some embodiments, the plurality of training instances falls into another range that begins with at least 3 training instances and ends with no more than 10,000 training instances.

[0233] In some such embodiments, training includes repeatedly adjusting the parameters of each of one or more second models on multiple training instances (e.g., via backpropagation), thereby increasing the accuracy of each of the one or more second models in determining the risk of compound DILI.

[0234] In some embodiments, model training trains one or more second models based on a first evaluation of the error function. In some such embodiments, model training trains one or more second models after updating one or more parameters for the first time based on the first evaluation of the error function. In some alternative embodiments, model training is performed after each of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, or at least 1 million evaluations of the error function. In some such embodiments, model training is performed through at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 10,000, 50,000, 100,000, 200,000, 500,000, or 1 million evaluations based on an error function. One or more parameters are updated at least once, at least twice, at least three times, at least four times, at least five times, at least six times, at least seven times, at least eight times, at least nine times, at least ten times, at least twenty times, at least thirty times, at least fourty times, at least fifty times, at least one hundred times, at least five hundred times, at least one thousand times, at least one hundred thousand times, at least one hundred thousand times, at least two hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, at least one hundred thousand times, or at least one million times to train one or more second models.

[0235] In some embodiments, model training is repeated until one or more second models, individually, jointly, or in combination with one or more first models, meet the minimum performance requirements. For example, in some embodiments, after evaluating the error function (e.g., the difference between the predicted DILI risk and the actual DILI risk), model training is repeated until the error calculated for one or more second models (individually, jointly, or in combination with one or more first models) meets an error threshold. In some embodiments, the error calculated by the error function meets the error threshold when the error is less than 20%, less than 18%, less than 15%, less than 10%, less than 5%, or less than 3%.

[0236] In one example embodiment, one or more second models are trained using the classification cross-entropy loss in a multi-task formulation, wherein each of the plurality of covariates corresponds to a cost function among the plurality of cost functions, and each of the plurality of cost functions has a common weighting factor.

[0237] Reference box 262, in some embodiments, the multiple second models are two to 1000 logistic regression models. Logistic regression models are disclosed in: Agresti, Introduction to Classification Data Analysis, 1996, Chapter 5, pp. 103-144, John Willie & Son Publishing, New York, NY, which is hereby incorporated by reference.

[0238] Reference box 264, in some embodiments, one or more second models include 10 or more second modules, 20 or more second modules, 30 or more second modules, 40 or more second modules, 50 or more second modules, 60 or more second modules, 70 or more second modules, 80 or more second modules, or 90 or more second modules.

[0239] Referring to box 266, in some embodiments, each of one or more second models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model. Details of such models are disclosed in the definitions section above.

[0240] In some embodiments, each second model includes multiple (e.g., 10 or more, 20 or more, 30 or more, 50 or more, 100 or more, 200 or more, 300 or more, 500 or more, 1000 or more, or 10,000 or more) parameters.

[0241] In some embodiments, each second model includes multiple parameters (e.g., weights and / or hyperparameters). In some embodiments, the multiple parameters of each second model include at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 200,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, or at least 5 million parameters. In some embodiments, the multiple parameters of each second model include no more than 8 million, no more than 5 million, no more than 4 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5,000, no more than 1000, or no more than 500 parameters. In some embodiments, the plurality of parameters of each second model comprises 10 to 5,000, 500 to 10,000, 10,000 to 500,000, 20,000 to 1 million, or 1 million to 5 million parameters. In some embodiments, the plurality of parameters of each second model falls into another range that begins with no fewer than 10 parameters and ends with no more than 8 million parameters.

[0242] In some embodiments, the training of each second model is further characterized by one or more hyperparameters (e.g., one or more values ​​that can be tuned during training). In some embodiments, the hyperparameter values ​​are tuned (e.g., adjusted) during training. In some embodiments, the hyperparameter values ​​are determined based on specific features of the training dataset and / or one or more inputs (e.g., cells, cellular component modules, covariates, etc.). In some embodiments, experimental optimization is used to determine the hyperparameter values. In some embodiments, hyperparameter scans are used to determine the hyperparameter values. In some embodiments, hyperparameter values ​​are assigned based on previous templates or default values.

[0243] In some embodiments, a corresponding hyperparameter of one or more hyperparameters includes a learning rate. In some embodiments, the learning rate is at least 0.0001, at least 0.0005, at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1. In some embodiments, the learning rate is no more than 1, no more than 0.9, no more than 0.8, no more than 0.7, no more than 0.6, no more than 0.5, no more than 0.4, no more than 0.3, no more than 0.2, no more than 0.1, no more than 0.05, no more than 0.01, or less. In some embodiments, the learning rate is 0.0001 to 0.01, 0.001 to 0.5, 0.001 to 0.01, 0.005 to 0.8, or 0.005 to 1. In some embodiments, the learning rate falls into another range starting at a value not lower than 0.0001 and ending at a value not higher than 1. In some embodiments, one or more hyperparameters further include regularization strength (e.g., L2 weight penalty, dropout rate, etc.). For example, in some embodiments, each second model (e.g., a neural network) is trained using regularization of the corresponding parameters (e.g., weights) of each of the plurality of hidden neurons. In some embodiments, regularization includes L1 or L2 penalty.

[0244] In some embodiments, the corresponding hyperparameters of one or more hyperparameters of the second model are the identity of the loss function. In some embodiments, the loss function is mean squared error, flat mean squared error, quadratic loss, mean absolute error, mean bias error, hinge, multi-class support vector machine, and / or cross-entropy. In some embodiments, the loss function is a gradient descent algorithm and / or a minimization function.

[0245] In some embodiments, each second model is associated with one or more activation functions (e.g., in the case of such models being neural networks). In some embodiments, the activation functions among the one or more activation functions are tanh, sigmoid, softmax, Gaussian, Boltzmann weighted average, absolute value, linear, modified linear unit (ReLU), bounded modified linear, soft modified linear, parameterized modified linear, average, max, min, sign, square, square root, multivariate quadratic, inverse quadratic, inverse multivariate quadratic, multiple harmonic spline, swish, mish, Gaussian error linear unit (GeLU), and / or thin-plate spline. In some embodiments, each second model outputs the predicted DILI risk in response to inputting the abundance value of a cellular component at a specific concentration 116 with a specific exposure time 118 into the model.

[0246] Referring to box 268, in some embodiments, a first value and one or more second values ​​are used to determine the risk that a test chemical compound will induce liver injury in human subjects. In such embodiments, one or more first models and one or more second models together serve as an integrated model, wherein one or more first values ​​from one or more first models and one or more second values ​​forming one or more second models determine the risk that a test chemical compound will induce liver injury in human subjects.

[0247] In some embodiments, this ensemble model comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 component models (e.g., the remaining independently trained instances of one or more first and second models).

[0248] In some embodiments, the ensemble model comprises no more than 500, 400, 300, 200, or 100 component models (e.g., the remaining independently trained instances of one or more first and second models). In some embodiments, the ensemble model comprises no more than 100, 50, 40, 30, or 20 component models. In some embodiments, the ensemble model comprises 1 to 50, 2 to 20, 5 to 50, 10 to 80, 5 to 15, 3 to 30, 10 to 500, 2 to 100, or 50 to 100 component models. In some embodiments, the ensemble model comprises another range of component models, starting with no fewer than 2 component models and ending with no more than 500 component models.

[0249] In some embodiments, the output of the ensemble model is formed by combining multiple outputs obtained from multiple component models (e.g., one or more first values ​​from one or more first models and one or more second values ​​from one or more second models). In some embodiments, the multiple outputs from one or more first models and one or more second models are combined using any measure of central tendency known in the art, including but not limited to mean, median, mode, weighted average, weighted median, weighted mode, arithmetic mean, number of columns, number of axes, three means, and / or Winsorized mean. For example, the final determination from the ensemble model can be obtained based on the average of the outputs of all component models across the ensemble model. In some embodiments, the output of each of one or more first models and one or more second models is a DILI probability with a value from 0 to 1.0. In some such embodiments, the DILI probabilities output by each of the one or more first models and each of the one or more second models are combined using any measure of central tendency known in the art.

[0250] In some embodiments, a voting method is used to combine multiple outputs. For example, in some embodiments, multiple outputs are combined by recording the number of outputs from each component model in the ensemble model, the outputs indicating that the test compound has a DILI risk at a specific concentration. In some embodiments, a majority vote at a specific concentration is used to combine multiple outputs from the component models. In some embodiments, each component model in the ensemble model is unweighted (e.g., each component model has one vote in the ensemble model). In some embodiments, one or more component models in the ensemble model are further weighted (e.g., have more than one vote in the ensemble model).

[0251] In some embodiments, the method described above in conjunction with Figure 2 is used to evaluate a variety of test compounds for DILI. In such embodiments, each of the multiple test compounds is tested by... Figure 3 The method is run. Therefore, if there are 100 test compounds, in such an embodiment, method 300 is run 100 times, with each of the 100 instances used to test a different test compound among the compounds.

[0252] In some embodiments, the plurality of test compounds includes at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 8000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 200,000, at least 500,000, at least 800,000, at least 1 million, or at least 2 million test compounds.

[0253] In some embodiments, the plurality of compounds includes no more than 10 million, no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 8,000, no more than 5,000, no more than 2,000, no more than 1,000, no more than 800, no more than 500, no more than 200, or no more than 100 test compounds. In some embodiments, the plurality of compounds consists of 10 to 500, 100 to 10,000, 5,000 to 200,000, or 10,000 to 1 million test compounds.

[0254] In some embodiments, the number of test compounds is 10 to 1 x 103 6 Multiple test compounds. In some embodiments, the multiple test compounds are 100 to 100,000. In some embodiments, the multiple test compounds are 1,000 to 100,000.

[0255] Referring to reference box 270, in some embodiments, the risk is presented as the probability that the test chemical compound will induce liver injury in human subjects. In some embodiments, the risk is a value between 0 and 1, where zero indicates that the test chemical compound has almost no chance of inducing liver injury in human subjects, and one indicates that the test chemical compound has a 100% chance of inducing liver injury in human subjects.

[0256] Boxes 272-276. Referring to box 272, in some embodiments, the risk is one of a set of listed risk levels for the test chemical compound to induce liver injury in human subjects. Referring to box 274, in some such embodiments, the set of listed risk levels comprises low-risk, intermediate-risk, and high-risk levels for the test chemical compound to induce liver injury in human subjects.

[0257] Reference box 276, in some embodiments, the risk is in the form of the lowest concentration of the test chemical compound that induces liver injury in human subjects, as predicted by one or more first models and one or more second models.

[0258] Box 278. Referring to box 278, in some embodiments, the method further includes determining a safety margin for the test chemical compound using the lowest concentration of the test chemical compound predicted by one or more first models and one or more second models that induces liver injury in human subjects (e.g., determining the safety margin includes dividing the risk by the C of the test chemical compound). max or LD 50 In some embodiments according to box 278, a first value and one or more second values ​​from an integrated model are used to determine the risk that the test chemical compound will induce liver injury in human subjects at each of several different concentrations, thereby generating a predicted DILI probability dose-response curve, such as... Figure 6 The curve shown is used to predict the dose-response curve. This predicted dose-response curve is then used to calculate the safety margin of the test chemical compound. For example, the predicted dose-response curve is used to estimate the LD50 of the test chemical compound. 50 This refers to the dose at which the predicted probability of DILI exceeds 50%. Additionally, separate pharmacokinetic studies are used to determine the C60 of the test chemical compound. max Concentration. The safety margin for DILI is based on LD50. 50 With C max The safety margin is calculated as a multiple of the difference between the values. In some embodiments, a test chemical compound is considered not a DILI risk chemical when the estimated safety margin is greater than 50. In some embodiments, a test chemical compound is considered a DILI risk chemical when the calculated safety margin is less than 50. In some embodiments, a threshold other than 50 is used to determine whether a compound has an appropriate safety margin. For example, in some embodiments, a threshold of 10 or 25 is used. In other embodiments, thresholds of 30, 35, 40, 45, 55, 60, 65, or 70 are used instead.

[0259] Referring to box 280, in some embodiments, the method informs the selection of one or more human subjects for treatment with a test chemical compound and / or the selection of one or more human subjects for continuing or discontinuing treatment with the test chemical compound. For example, in some embodiments, cells from each subject in a subject population are exposed to the test chemical compound at one or more concentrations 116 / exposure time ratios 118. Cell component abundance values ​​122 and pathway module values ​​from such exposed cells are then input into an integrated model (comprising one or more first models 130 and one or more second models 136) to determine the safety margin of the test chemical compound in such subjects using the techniques disclosed above with respect to box 278, and to treat only those subjects with a safety margin greater than 50 with the test chemical compound. In some embodiments, a threshold other than 50 is used to determine whether a particular subject has an appropriate safety margin. For example, in some embodiments, instead of thresholds of 30, 35, 40, 45, 55, 60, 65, or 70, thresholds are used to determine whether a particular subject can use the test chemical compound to treat a disease afflicting the subject.

[0260] Referring to box 282, in some embodiments, the method informs one or more human subjects of the dosage, duration, and / or frequency of administration of a test chemical compound. For example, in some embodiments, cells from a single subject or a population of subjects are exposed to the test chemical compound at one or more concentrations 116 / exposure time ratios 118. Cellular component abundance values ​​122 and pathway module values ​​from such exposed cells are then input into an integrated model (comprising one or more first models 130 and one or more second models 136) to determine the safety margin of the test chemical compound in a single subject or population of subjects using the techniques disclosed above with respect to box 278, and to treat only those subjects with a safety margin greater than 50 with the test chemical compound. In some embodiments, a threshold other than 50 is used to determine whether a subject has an appropriate safety margin. For example, in some embodiments, instead of thresholds of 30, 35, 40, 45, 55, 60, 65, or 70, thresholds are used to determine whether a subject can use the test chemical compound to treat a disease afflicting a human subject.

[0261] Boxes 284-288. Referring to box 284, in some embodiments, the method informs the design of a clinical trial that includes the use of a test chemical compound. Referring to box 286, in some such embodiments, the method informs the type and / or level of acceptable adverse events for the clinical trial, one or more inclusion criteria, and / or one or more exclusion criteria. Referring to box 288, in some such embodiments, the method informs one or more modifications to the clinical trial while it is ongoing. In one instance, cells from candidate subjects of a clinical trial are exposed to a test chemical compound at one or more concentrations 116 / exposure time to 118. Cell component abundance values ​​122 and pathway module values ​​from such exposed cells are then input into an integrated model (comprising one or more first models 130 and one or more second models 136) to determine the appropriate dose of the test chemical compound for the trial. In another instance, cells from candidate subjects of a clinical trial are exposed to a test chemical compound at one or more concentrations 116 / exposure time to 118 to determine whether such candidate subjects can tolerate the test chemical compound and therefore can participate in the trial. In another instance, cells from candidate subjects in clinical trials are exposed to a test chemical compound at one or more concentrations of 116 / exposure time to determine whether such candidate subjects can remain in the trial.

[0262] Boxes 290-294. Referring to box 290, in some embodiments, the method informs the design of an adaptive clinical trial that includes the use of a test chemical compound. Referring to box 292, in some embodiments, the method informs the type and / or level of acceptable adverse events for the adaptive clinical trial, one or more inclusion criteria, and / or one or more exclusion criteria. Referring to box 294, in some embodiments, the method informs one or more modifications to the clinical trial while the adaptive clinical trial is in progress. In one instance, cells from candidate subjects of an adaptive clinical trial are exposed to a test chemical compound at one or more concentrations 116 / exposure time to 118. Cell component abundance values ​​122 and pathway module values ​​from such exposed cells are then input into an integrated model (comprising one or more first models 130 and one or more second models 136) to determine an appropriate dose of the test chemical compound for the adaptive trial. In another instance, cells from candidate subjects of a clinical trial are exposed to a test chemical compound at one or more concentrations 116 / exposure time to 118 to determine whether such candidate subjects can tolerate the test chemical compound and therefore can participate in the adaptive trial. In another instance, cells from candidate subjects in clinical trials are exposed to a test chemical compound at one or more concentrations of 116 / exposure time to determine whether such candidate subjects can remain in the adaptation trial.

[0263] Referring to box 296, in some embodiments, the method informs the classification of test chemical compounds as either compounds suitable for human therapy and / or testing or compounds unsuitable for human therapy and / or testing. For example, in some embodiments, cells from a subject population are exposed to the test chemical compound at one or more concentrations 116 / exposure time ratios 118. Cellular component abundance values ​​122 and pathway module values ​​from such exposed cells are then input into an integrated model (comprising one or more first models 130 and one or more second models 136) to determine the safety margin of the test chemical compound for the subject population using the techniques disclosed above with respect to box 278. When a compound is determined to have a safety margin greater than 50 across the tested subject population (e.g., a measure of central tendency of safety margins across subject populations), the test chemical compound is determined to be suitable for human therapy and / or testing. When a compound is determined to have a safety margin less than 50 across the tested subject population (e.g., a measure of central tendency of safety margins across subject populations), the test chemical compound is determined to be unsuitable for human therapy and / or testing. In some embodiments, thresholds other than 50 are used to determine whether a subject has an appropriate safety margin. For example, in some embodiments, instead of using thresholds of 30, 35, 40, 45, 55, 60, 65, or 70, thresholds are used to determine whether a chemical compound is suitable for human therapy and / or testing.

[0264] Reference box 298, in some embodiments, the method further includes preparing a test chemical compound for use in a therapy. In some embodiments, the therapy is to alleviate symptoms, such as inflammation. In some embodiments, the treatment is to alleviate or treat a disease or condition. In some embodiments, the disease or condition is cancer, a hematological disorder, an autoimmune disease, an inflammatory disease, an immune disorder, a metabolic disorder, a neurological disorder, a genetic disorder, a mental disorder, a gastrointestinal disorder, a kidney disorder, a cardiovascular disorder, a dermatological disorder, a respiratory disorder, a viral infection, or another disease or condition.

[0265] In some embodiments, formulating a test chemical compound for a therapy comprises preparing a composition comprising the test chemical compound and one or more excipients and / or one or more pharmaceutically acceptable carriers and / or one or more diluents.

[0266] Such excipients and / or carriers include all conventional solvents, dispersion media, fillers, solid carriers, coatings, antifungal and antimicrobial agents, skin penetrants, surfactants, isotonic agents, and absorbents. It should be understood that the compositions disclosed herein may also include other supplementary physiologically active agents.

[0267] The exemplary carrier is pharmaceutically "acceptable" and harmless to the subject in the sense of compatibility with other components of the composition (e.g., a composition containing a test chemical compound). The composition can be conveniently presented in unit dosage form and can be prepared by any method well known in the pharmaceutical industry. Such methods involve the step of associating the active ingredient with a carrier constituting one or more auxiliary components. Generally, the composition is prepared by homogeneously and tightly associating the active ingredient with a liquid carrier or a finely dispersed solid carrier, or both, and subsequently shaping the product if desired.

[0268] Exemplary compounds, compositions, or combinations thereof (e.g., test chemical compounds) or their pharmaceutically acceptable salts, solvates, or prodrugs prepared for intravenous, intramuscular, or intraperitoneal administration may be administered by injection or infusion.

[0269] Injectable formulations for such applications can be prepared in conventional forms as liquid solutions or suspensions, or as solids suitable for dissolution or suspension in a liquid prior to injection, or as emulsions. Carriers may include, for example, water, saline solutions (e.g., normal saline (NS), phosphate-buffered saline (PBS), balanced salt solutions (BSS)), sodium lactate Ringer's solution, glucose, glycerol, ethanol, etc.; and, if desired, small amounts of excipients such as wetting agents or emulsifiers, buffers, etc., may be added. For example, in the case of dispersions, the desired particle size can be maintained by using coatings such as lecithin, and appropriate flowability can be maintained by using surfactants.

[0270] The compounds, compositions, or combinations disclosed herein (e.g., test chemical compounds) may also be suitable for oral administration and may be presented as discrete units, such as capsules, pouches, or tablets, each containing a predetermined amount of the active ingredient; powders or granules; solutions or suspensions in aqueous or non-aqueous liquid form; or oil-in-water or water-in-oil liquid emulsions. The active ingredient may also be presented as a bolus, saccharin, or paste.

[0271] Tablets can be prepared by compression or molding, optionally with one or more excipients. Compressed tablets can be prepared by compressing an active ingredient (e.g., a test chemical compound) in a free-flowing form (such as powder or granules) in a suitable machine, optionally mixed with a binder (e.g., an inert diluent, a preservative disintegrant (e.g., sodium starch glycolate, croscarmellose, croscarmellose sodium carboxymethyl cellulose), a surfactant, or a dispersant). Molded tablets can be prepared by molding a mixture of powdered compounds wetted with an inert liquid diluent in a suitable machine. Tablets can optionally be coated or scored and can be formulated such that varying proportions of hydroxypropyl methylcellulose are used to provide sustained or controlled release of the active ingredient to provide a desired release profile. Tablets can optionally be provided with an enteric coating to provide release in the intestinal portion other than the stomach.

[0272] The compounds, compositions, or combinations disclosed herein (e.g., test chemical compounds) may be suitable for topical oral application, including lozenges containing active ingredients in the form of a flavoring matrix, typically sucrose and arabinose or gum arabic; soft lozenges containing active ingredients in the form of an inert matrix, such as gelatin and glycerin, or sucrose and gum arabic; and mouthwashes containing active ingredients in the form of a suitable liquid carrier.

[0273] The compounds, compositions, or combinations thereof (e.g., test chemical compounds) disclosed herein are suitable for topical application to the skin, may comprise compounds dissolved or suspended in any suitable carrier or matrix, and may be in the form of lotions, gels, creams, pastes, ointments, etc. Suitable carriers include mineral oil, propylene glycol, polyoxyethylene, polyoxypropylene, emulsified waxes, sorbitan monostearate, polysorbate 60, hexadecyl ester wax, cetearyl alcohol, 2-octyldodecyl alcohol, benzyl alcohol, and water. Transdermal patches may also be used to administer the compounds of the present invention.

[0274] Compounds, compositions, or combinations of the present disclosure suitable for parenteral administration (e.g., test chemical compounds) include aqueous and non-aqueous isotonic sterile injectable solutions, which may contain antioxidants, buffers, bactericides, and solutes (to make the compound, composition, or combination isotonic with the blood of the intended recipient); and aqueous and non-aqueous sterile suspensions, which may include suspending agents and thickeners. The compounds, compositions, or combinations may be present in single-dose or multi-dose sealed containers, such as ampoules and vials, and may be stored under lyophilized (freeze-dried) conditions where a sterile liquid carrier (e.g., water for injection) needs to be added immediately prior to use. Temporary injectable solutions and suspensions can be prepared from sterile powders, granules, and tablets of the previously described types.

[0275] It should be understood that, in addition to the active ingredients specifically mentioned above, the compositions or combinations disclosed herein (e.g., test chemical compounds) may include other conventional pharmaceutical agents of the type discussed, such as agents suitable for oral administration, including binders, sweeteners, thickeners, flavoring agents, disintegrants, coating agents, preservatives, lubricants, and / or delay agents. Suitable sweeteners include sucrose, lactose, glucose, aspartame, or saccharin. Suitable disintegrants include corn starch, methylcellulose, polyvinylpyrrolidone, xanthan gum, bentonite, alginate, or agar. Suitable flavoring agents include peppermint oil, wintergreen oil, cherry, orange, or raspberry flavorings. Suitable coating agents include polymers or copolymers of acrylic acid and / or methacrylic acid and / or their esters, waxes, fatty alcohols, corn gluten, shellac, or gluten. Suitable preservatives include sodium benzoate, vitamin E, α-tocopherol, ascorbic acid, methylparaben, propylparaben, or sodium bisulfite. Suitable lubricants include magnesium stearate, stearic acid, sodium oleate, sodium chloride, or talc. Suitable delay agents include glyceryl monostearate or glyceryl distearate.

[0276] III. Other embodiments

[0277] Another aspect of this disclosure provides a method for assessing the risk that a test chemical compound will induce liver injury in human subjects. The method includes obtaining a cellular component dataset 112 in electronic form, the cellular component dataset comprising multiple cell abundance values ​​122 of multiple cellular components 120 in a first plurality of cells that have been exposed to the test chemical compound for a first sustained period of time. Details of such a cellular component dataset are described above in conjunction with box 212.

[0278] A cellular component dataset is used to determine the corresponding activation value for each of a plurality of cellular component modules, thereby obtaining a plurality of activation values. Each corresponding cellular component module contains an independent subset of a plurality of cellular components. Each corresponding activation value for each corresponding cellular component module is a measure of the central tendency of the abundance of each cellular component in the independent subset of cellular components corresponding to the corresponding subset of cellular components across the cellular component dataset. Each independent subset of the plurality of cellular components in each corresponding cellular component module contains two or more cellular components. Each cellular component in the plurality of cellular components is located in at least one independent subset of the plurality of cellular components associated with the cellular component module in the plurality of cellular component modules. In some embodiments, the plurality of cellular components contains 10 or more cellular components, and the plurality of cellular component abundance values ​​contains 1000 or more cellular component abundance values.

[0279] Multiple activation values ​​are input into the model to obtain values ​​that indicate the risk of the tested chemical compound inducing liver damage in human subjects, which are then used as the model's output.

[0280] In some embodiments, the first plurality of cells are primary human hepatocytes.

[0281] In some embodiments, multiple cell abundance values ​​are determined using single-cell RNA sequencing (scRNA-seq) data from a first plurality of cells. In some embodiments, multiple cell abundance values ​​are determined using batch RNA sequencing. In some embodiments, multiple cell abundance values ​​are determined using scTag-seq. In some embodiments, multiple cell abundance values ​​are determined using single-cell assays of sequenced transposase-accessible chromatin (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof. Details of suitable cell component abundance measurement techniques are disclosed above in conjunction with boxes 218 to 224.

[0282] In some embodiments, the model is a logistic regression model. In alternative embodiments, the model is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

[0283] In some embodiments, the plurality of cell component modules includes 10 or more cell component modules, 20 or more cell component modules, 30 or more cell component modules, 40 or more cell component modules, 50 or more cell component modules, 60 or more cell component modules, 70 or more cell component modules, 80 or more cell component modules, or 90 or more cell component modules.

[0284] In some embodiments, each of the plurality of cellular components is a specific gene, a specific mRNA associated with the gene, a carbohydrate, a lipid, an epigenetic feature, a metabolite, a protein, or a combination thereof. Further description of suitable cellular components is given in conjunction with box 214 above.

[0285] In some embodiments, the method further includes identifying multiple cellular component modules by obtaining one or more first datasets in electronic form. The one or more first datasets include, for each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes and collectively represents multiple annotated cell states, for each corresponding cellular component in a plurality of cellular components: the corresponding abundance of the corresponding cellular component in the corresponding primary human hepatocyte, thereby obtaining or forming multiple vectors. Each corresponding vector in the plurality of vectors (i) corresponds to a corresponding cellular component in the plurality of components, and (ii) contains a corresponding plurality of elements, each of the corresponding plurality of elements representing the abundance of the corresponding cellular component in the corresponding primary human hepatocyte in the plurality of primary human hepatocytes. The multiple vectors are clustered to form multiple clusters, wherein each of the multiple clusters consists of an independent subset of the multiple cellular components of the corresponding cellular component module in the plurality of cellular component modules.

[0286] In some alternative embodiments, the method further includes identifying multiple cellular component modules by comprising a process of obtaining one or more first datasets in electronic form. The one or more first datasets comprise, for each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes and collectively represents a plurality of annotated cell states, for each corresponding cellular component in a plurality of cellular components: the corresponding abundance of the corresponding cellular component in the corresponding primary human hepatocyte. A shared cellular component module matrix is ​​generated using the one or more first datasets using shared nonnegative matrix factorization. In this embodiment, multiple least squares regression of the matrix is ​​used with respect to the corresponding abundance of cellular components in the one or more first datasets to identify independent subsets of multiple cellular components from the shared matrix for each corresponding cellular component module in the plurality of cellular component modules.

[0287] In some embodiments, an annotated cell state among a plurality of annotated cell states is a cell in a second plurality of cells exposed to a training compound under exposure conditions. In some embodiments, exposure conditions are exposure duration, concentration of the training compound, or a combination of exposure duration and concentration of the training compound.

[0288] In some embodiments, multiple annotated cell states include exposure to a first training compound at multiple different concentrations. In some embodiments, multiple different concentrations of the first training compound at C max IC to the first training compound 10 Within the range. In some such embodiments, multiple annotated cell states include exposure to the first training compound for multiple different exposure durations.

[0289] In some embodiments, multiple annotated cell states comprise each corresponding training compound exposed to multiple training compounds at corresponding multiple different concentrations. In some such embodiments, the corresponding multiple different concentrations range from the Cmax to the IC50 of each corresponding training compound. 10 Within the range.

[0290] In some embodiments, the plurality of annotated cell states include each of the plurality of training compounds being exposed to a plurality of different exposure durations.

[0291] In some embodiments, the clustering is in the form of graph clustering. In some such embodiments, graph clustering is Leiden clustering based on Pearson-related distance metrics. In alternative embodiments, graph clustering is Louvain clustering.

[0292] In some embodiments, the plurality of cell component modules consists of 10 to 2000 cell component modules.

[0293] In some embodiments, the plurality of cellular components consists of 100 to 35,000 cellular components.

[0294] DILI predicts the ranking of features.

[0295] As discussed herein, fold changes in single-cell transcriptomic abundance data of various cellular components were obtained from (i) cells exposed to a control (e.g., DMSO) relative to (ii) cells already exposed to the compound. These fold changes were used to determine pathway scores for multiple modules. Both the fold changes in cellular components and the pathway scores served as input features to an integrated model. Specifically, multiple pathway values ​​were input into one or more first models of the integrated model to obtain one or more first model values. The fold changes in single-cell transcriptomic abundance data were input into each of one or more second models of the integrated model to obtain one or more second values ​​as outputs from the one or more second models. The integrated model used one or more first values ​​and one or more second values ​​to determine the risk that the tested chemical compound would induce liver injury in human subjects. Thus, in some embodiments of this disclosure, the following input features were used: cellular component features (logarithmic fold change in cellular component abundance (DMSO relative to compound perturbation)) and pathway module features: pathway module scores (e.g., -log10(FDR)).

[0296] In some embodiments, each model of one or more first models and each model of one or more second models is a logistic regression. In some embodiments, one or more first models and one or more second models together consist of 70 to 150 models. In some embodiments, one or more first models and one or more second models together consist of 90 to 120 models.

[0297] In some embodiments, ranking the importance of these input features is of interest. In some such embodiments, this is done by training an ensemble of models using a bagging (bootstrapping aggregation) technique, which is an ensemble learning technique that independently trains multiple models and combines them with equal weights. This helps reduce variance and prevent overfitting, potentially leading to better generalization. In some embodiments, the ensemble consists of: a) multiple base models independent of different subsets of the training data, b) multiple base models with different hyperparameters, and c) individual gene-level and pathway-level features. Once trained, the feature importance of each feature across the ensemble is based on its coefficient (weight). Features with larger absolute coefficients are considered more important in determining the output. The ensemble feature importance is calculated as the average of the coefficients of the ensemble models.

[0298] In some embodiments, permutation importance is considered. For the first K features, in addition to feature importance, the values ​​of one feature are shuffled one at a time, and the change in model performance (AUC) is measured. The greater the performance degradation, the more important the feature.

[0299] IV. Examples

[0300] This article provides an example of assessing the risk of a tested chemical compound inducing liver injury in human subjects.

[0301] The Japan Toxicogenomics Project (TGP) includes gene expression data, toxicological information, and pathological data on the human toxicity of 131 compounds screened in vitro. This data is available online from [website / source]. http: / / camda2021.bioinf.jku.at / doku.php / tgp_prepro The data were obtained online. For each compound, the training data contained gene expression values ​​(pretreated FARMS) at different time points (2-hour exposure, 8-hour exposure, and 24-hour exposure) and dose levels (low, medium, and high), drug-induced liver injury (DILI), and label categories ("-1", "+1").

[0302] For model training, 120 high-concentration compounds were selected from the human TGP dataset. The training data was... Figure 4 As shown in the image. Figure 4Each column represents one of the selected training compounds. The training compounds are arranged from the left (DILI with the highest toxicity) to the right (DILI with the lowest toxicity, toxicity not measured).

[0303] Abundance data of differentially expressed genes in the TGP dataset enriched known gene networks. Therefore, Figure 4 Each row at the bottom of the table represents a different gene network (pathway module).

[0304] Table 2 below provides example pathway modules and independent subsets of genes associated with each such pathway module. It is noteworthy that some genes are present in multiple pathways. In Table 2, OS represents oxidative stress, LP represents lipid peroxidation, MD represents mitochondrial dysfunction, and AD represents antioxidant defense.

[0305] For any given compound, the measurement of differentially expressed genes associated with a pathway module will be statistically significant, as measured against the distribution of expression values ​​for randomly selected genes from the TGP dataset. This statistical significance can be expressed as a p-value. For a specific compound, p-values ​​less than p < 0.05 are marked with a single asterisk (*), p-values ​​less than p < 0.01 are marked with a double asterisk (**), and p-values ​​less than p < 0.001 are marked with a triplet (***). The statistical significance (p-value) for each pathway module will vary from compound to compound. Table 2 shows the p-values ​​for each pathway enrichment in the pathway enrichment of predicted differentially expressed genes (DEGs) for at least one of 120 compounds evaluated during training based on a statistical hypergeometric test to assess the significance of observed pathway enrichments with DEGs.

[0306] Table 2. Example pathways and independent subsets of genes associated with each pathway.

[0307] In this example, each corresponding pathway module provides a corresponding pathway value for each corresponding compound, which is a weighted average of microarray expression data of independent subsets of genes associated with the corresponding pathway module when TGP cell cultures are exposed to the corresponding compound. In an alternative embodiment, the corresponding pathway value for a given compound is a p-value for the pathway module of the given compound.

[0308] The ensemble model was trained using pathway values ​​derived from TGP data and the abundance data of the genes themselves within the TGP data. The ensemble model consisted of a first model and multiple second models. Pathway values ​​for each of the 120 training compounds were fed into the first model. In this example, the first model was an independent logistic regression model (Model 1). After inputting the pathway values, the output DILI predictions of the first model were trained against the known toxicities of the training compounds.

[0309] Individually, gene expression data of 7254 genes from cells exposed to each of a set of 120 training compounds were input into each of several secondary models. In this example, each secondary model was a separate logistic regression model. After inputting gene expression values ​​of the 7254 genes for each training compound, the output DILI prediction of each secondary model was trained against known toxicity values ​​of the training compounds. A subset of the multiple logistic regression models was trained on an L1 basis, and another subset was trained on an L2 basis.

[0310] For any given training compound at any given training composition, the output of the ensemble model is a range of DILI predicted probabilities, referred to in this example as a confidence interval. Contributing to this confidence interval are the DILI predicted probabilities of the first model and the DILI predicted probabilities of each of the multiple second models. Therefore, for any given training compound, at any given concentration, the output of the ensemble model is a confidence interval of the DILI predicted probabilities. Figure 5 The performance of this ensemble model (a first model together with multiple second models) is shown. In the computation... Figure 5 When assessing performance, the predicted DILI probability of any given composition at any given composition exposure concentration is used as the center of the measured confidence interval. For Figure 5 The two graphs show that the Y-axis represents sensitivity (true positive rate, TPR, calculated as TP / (TP+FN), where TP represents true positives and FN represents false negatives). Figure 5 In the left figure, the X-axis represents specificity (false positive rate, FPR, which is calculated as FP / (FP+TN, where FP represents false positive and TN represents true negative). Figure 5 The left figure shows the trained ensemble model with an area under the curve (AUC) of 0.9. A model with an AUC of 1 will perfectly determine whether a compound has DILI toxicity. A model with an AUC of 0.5 means that the model does not have a significant ability to determine whether a compound has DILI toxicity.

[0311] Figure 5The right-hand figure shows the cut-point of the same ROC analysis. In this right-hand figure, the X-axis represents the predicted probability. Figure 5 The tangent point in the right-hand figure is the point where the sensitivity and specificity curves intersect. The cutoff point determines the point at which the ensemble model classifies a drug as either DILI risk or non-DILI risk. By adjusting the cutoff point, a trade-off can be struck between the true positive rate and the false positive rate. Lower cutoff values ​​(on the X-axis) classify more cases as DILI risk, resulting in higher sensitivity but potentially higher false positives, while higher thresholds classify fewer cases as positive, resulting in lower sensitivity but potentially lower false positives. After the classifier assigns predicted probabilities to compounds, it is necessary to decide which category to assign to each compound. This decision is typically made by comparing the predicted probabilities to a predetermined threshold (e.g., the ROC cutoff point). If the predicted probability is above the threshold, the compound is classified as positive (i.e., the compound is considered to have DILI risk); otherwise, the compound is classified as negative (i.e., the compound is considered to have no DILI risk).

[0312] Validation was performed using hepatotoxicity assays.

[0313] Then, the ensemble model trained using TGP data was tested against internal testing (validation) data for thirty compounds. This internal validation data consisted of each of the 30 compounds in three replicas, with Cmax to IC50 values. 10 SMART-Seq DE data of primary human hepatocytes after 24 hours of exposure at four different concentrations within the range. Primary human hepatocytes were used because of their physiological relevance and predictive cell type for DILI.

[0314] Cell selection. Cryopreserved primary human hepatocytes, purchased from LifeNet Health, were selected using viability and plating criteria to meet the requirements for 96-well compatibility for culture. Specific donors were selected for long-term plating compatibility (plating ability at 10–15 days), 96-well compatibility, and to demonstrate Grade A quality. Primary human hepatocytes were cultured according to the supplier's protocols and proprietary culture media.

[0315] Experiment setup.

[0316] On day zero, cells were seeded at a density of 500,000 cells / mL in collagen I-coated 96-well plates using the supplier’s thaw and plate-laying medium.

[0317] On day 1, cells were covered with 0.5 mg / mL matrix gel until day 2 to allow the cells to grow under “ECM sandwich culture” conditions to promote the three-dimensional matrix of hepatocytes.

[0318] On the second day, the cells were maintained by changing the culture medium to stabilize the culture.

[0319] On day three, cells were treated with the test compound. For toxicology runs, the compound was dissolved in DMSO at doses of 1000 mM, 500 mM, or 250 mM (depending on solubility), according to the protocol described below. The compound was further diluted 200-fold to bring it down to the final working concentration in the culture medium, following a 6-point curve (complete logarithmic decrease).

[0320] Test the solubility of the compound in DMSO

[0321] Relatively high concentrations of compounds in DMSO were used for hepatotoxicity assays. To obtain optimal solubility of the compounds at these high concentrations, a suitable compound dissolution protocol was developed, under which the compounds are dissolved on the same day as the experiment to minimize potential "solution precipitation" problems that might otherwise arise from freeze-thaw cycles.

[0322] The stock solution was ordered in powder form and dissolved in sterile DMSO at 200 mM. If the compound was visually insoluble at 200 mM, it was further diluted to a stock concentration of 100 mM, and if it was still visually insoluble at 100 mM, it was further diluted to 50 mM. The compound was not used if it was insoluble at 50 mM. According to two protocols (used to determine IC50),... 10 (Toxicity assays and toxicity assays for obtaining RNA samples) further process the stock compounds.

[0323] Used to determine IC 10 Toxicity testing.

[0324] The stock compound solution was vortexed and placed in a bead bath to ensure solubility was maintained.

[0325] The stock compound solutions were further processed using automated equipment (Hamilton machine), titrated 5 to 10 times to test a 6-point toxicity profile for each compound, with each run performed in triplicate.

[0326] The stock compound solution was titrated 10-fold, and the sample was further diluted 200-fold in hepatocyte maintenance medium to achieve final concentrations of the compound, starting from initial concentrations of 1000 µM, 500 µM, or 250 µM, respectively. This final dilution of the compound ensured a “safe” 0.5% DMSO concentration, which had previously been shown not to significantly induce toxicity or significant transcriptional activity in cells.

[0327] Before adding the compound, maintenance medium was aspirated from the cells, and a diluted solution of the test compound was added to the cells in 100 µL volumes through a 96-well plate.

[0328] Used for toxicity assays to obtain RNA samples.

[0329] During sequencing, follow the guidelines used to determine IC. 10 The toxicity assay procedure is similar to the assay procedure for obtaining RNA samples, with the only substantial difference being that the initial concentration of the test compound is based on the method mentioned above for determining the IC50. 10 The results of the toxicity assays are predetermined. The IC50 value determined by the above assays will be... 10 The concentration is used as the starting concentration for sequencing runs to obtain RNA samples.

[0330] On the fourth day, cell viability was assessed using an LDH assay.

[0331] For sequencing runs: Select the concentration in each toxicology run by defining the maximum tolerated dose (MTD). Define the MTD as the concentration at which <10% activity is observed. Redissolve the compound in 200 mM DMSO (or 100 mM or 50 mM stock solution if it does not dissolve at 200 mM). Then dilute the compound 200-fold in culture medium to reach the final maximum tolerated dose and plot a 4-point curve (half-logarithmic decrease) in culture medium.

[0332] On day four, cell viability was assessed using an LDH assay, and the cells were euthanized using lysis buffer to produce lysates. The lysates were then processed to extract RNA, which was then used to run SMART-SeqDE (a commercially available Smart-seq2-based kit compatible with multiplex cDNA library preparation and allowing for increased throughput) (Takara Bio).

[0333] Gene expression readings measured by SMART-Seq DE from perturbation (cells exposed to each composition) and DMSO (cells exposed to DMSO only) were processed using DESeq2. https: / / genomebiology.biomedcentral.com / articles / 10.1186 / s13059-014-0550-8Love et al., 2014, “Moderate estimation of fold change and dispersion of RNA-seq data using DESeq2,” *Genome Biology* 15, 550, available online at doi.org / 10.1186 / s13059-014-0550-8, used for normalization and log fold change estimation of perturbation spectra against DMSO. DESeq2 is first used to estimate the size factor for each sample. The size factor represents the scaling factor that should be applied to each sample to account for differences in sequencing depth. It is calculated as the median ratio of gene counts across all samples to their geometric mean. DESeq2 then uses these size factors to normalize the raw counts by dividing the count of each gene in the sample by the corresponding size factor for that sample. This normalization step adjusts the counts for sequencing depth to make the samples comparable. The log fold change is then determined by taking the difference between the perturbation of each gene and the log2-transformed normalized count (log2-count) between DMSOs. DESeq2 employs a shrinkage procedure known as “regularized logarithmic transformation” to moderate the fold change. This helps improve the accuracy of fold change for genes with low counts and high dispersion. Differential Gene Expression: Finally, DESeq2 performs statistical tests (Wald test) to determine whether the observed fold change is significantly different from zero. Therefore, multiple gene abundance values ​​for multiple genes that have been exposed to the test chemical compound for a sustained first time period are presented as fold changes in gene expression (DMSO relative to the compound perturbation). These fold changes are used to calculate pathway values ​​for one or more first models and are also directly input into one or more second models to determine the risk that the test chemical compound will induce liver injury. Based on this analysis, 22 out of 30 compounds were identified as non-DILI-preserving compounds, while 8 compounds were identified as DILI-safe compounds.

[0334] Based on this validation data, in one validation study, the ensemble model showed 86% sensitivity (19 / 22 DILI) and 100% specificity (8 / 8 non-DILI). In a replication of this validation study using data from the same 30 compounds in different laboratory settings, the ensemble model showed 89% sensitivity (17 / 19 DILI) and 100% specificity (9 / 9 non-DILI).

[0335] Of the 30 compounds, only three were misclassified: labetalol, acetaminophen, and isoniazid. In comparison, the industry standard based on QSAR is 62% sensitivity and 75% specificity. Table 3 provides a list of the compound name, CMAX, toxic dose, safety margin, and known DILI status in the literature for each of the 30 validation compounds.

[0336] Table 3.

[0337] Figure 6 The validation results for amiodarone are presented. As shown in Table 3, the toxic dose of amiodarone is 8 µM. The Cmax of amiodarone is 0.8 µM, where Cmax is a standard pharmacokinetic measurement (the maximum or peak serum concentration of the drug reached in a specific compartment or test area of ​​the body after the initial administration and before the second dose). Therefore, the safety margin for amiodarone is 10, and it is considered a DILI risk. The recommended safety margin for DILI risk is greater than 50. See Leishman et al., 2020 “Revisiting the hERG safety margin after 20 years of routine hERG screening,” *J Pharmaocol Toxicol Methods* 105, 106900, which is hereby incorporated by reference.

[0338] Figure 6 The ensemble model shows the predictions made by primary human hepatocytes 24 hours after exposure at each of four different concentrations of amiodarone. Each of the four different concentrations is represented by a black dot. Because the ensemble model comprises multiple models, each predicting the DILI probability, each point is centered on a confidence interval defined by the minimum and maximum probabilities identified by any model within the ensemble model at a given concentration. These confidence intervals are represented by vertical lines passing through all DILI prediction probabilities of each model in the ensemble model. The average DILI prediction probability at each tested concentration is used to define curve 602. Using curve 602, LD... 50 Concentrations derived as having a DILI probability greater than 0.5. Therefore, Figure 6 The figure shows the LD given by the DILI probability of amiodarone when it is 0.5. 50The concentration is 7.9 uM, while the maximum concentration (Cmax) is 1.34 uM. Therefore, the safety margin for DILI predicted by the ensemble model is 7.9 / 1.34 = 5.9, which is less than 50. Thus, the ensemble model classifies amiodarone as a DILI risk, consistent with the known data for this compound provided in Table 3.

[0339] Figure 7 The gene expression characteristics measured for 12 out of 30 tested compounds are presented. For each of the 12 compounds, four lists of expression data are provided, arranged from the lowest to the highest concentration of the tested compound exposed. Figure 7 In the model, gene expression is presented as a heatmap ranging from low gene expression (dark blue) to high gene expression (dark red). The ensemble model could not provide predictions of DILI toxicity for erythromycin or labetalol. Figure 7 The gene expression signatures of these compounds are shown to be unsignatured, indicating that the ensemble model cannot read signals from these gene signatures in order to predict the DILI toxicity of these compounds.

[0340] Figure 8A and Figure 8B The pathway values ​​for 12 out of 30 test compounds across 16 different pathway modules are presented. For each of the 12 compounds, there are four columns arranged from the lowest to the highest test compound exposure concentration. In Figure 8, for each corresponding pathway module and for each corresponding compound, the corresponding pathway values ​​are presented as a heatmap from low pathway values ​​(dark blue) to high pathway values ​​(dark red). Figure 8 also highlights two compounds for which the ensemble model could not provide DILI toxicity predictions: erythromycin or labetalol. Figure 8 shows that the pathway values ​​for exposure to these compounds were also featureless, indicating that the ensemble model could not read the signal from these pathway module values ​​to predict the DILI toxicity of these compounds.

[0341] Figure 9 This further explains why the ensemble model failed to predict erythromycin and labetalol. Exposure of primary human hepatocytes to any of these compounds did not cause [the following symptoms]. Figure 9 The expression of known DILI markers is listed in the table.

[0342] Figure 10A graph was plotted showing the predicted safety margin versus actual DILI toxicity (X-axis) for the toxicity of 30 tested compounds, determined by an integrated model (Y-axis). The integrated model correctly labeled all compounds except three that should have a safety margin below 50, as all of these compounds are known to have DILI toxicity. A safety margin below 50 is considered a risk of DILI toxicity. The integrated model mislabeled only three compounds with a safety margin greater than 50 (DILI safe). The integrated model should have predicted a safety margin below 50 (DILI toxicity) for these three compounds. These three mislabeled compounds are labetalol, acetaminophen, and isoniazid.

[0343] Figure 11 This demonstrates how the preclinical DILI safety margin predicted by the ensemble model correlates with known clinical treatment indices. Figure 11 A plot was created showing the relationship between the predicted safety margin (X-axis) of the toxicity determined by the ensemble model for 30 test compounds and the known clinical treatment index (Y-axis). The correlation coefficient between (i) the predicted safety margin (X-axis) of the toxicity determined by the ensemble model for the 30 test compounds and (ii) the known clinical treatment index (Y-axis) was 0.78. Figure 11 As can be seen, the compounds labetalol, acetaminophen, and isoniazid, which were incorrectly labeled with DILI by the ensemble model, have relatively large therapeutic indices that are more similar to those of compounds that are safe in DILI.

[0344] This example demonstrates that trained ensemble models, by predicting DILI safety margins from in vitro transcriptomics of primary hepatocytes, can guide compound selection and optimization by mitigating the risk of DILI (drug-induced liver injury). This example shows that trained ensemble models exhibit stronger predictive power than industry-standard QSAR models. Furthermore, the nature of the data used by the ensemble model (i.e., in vitro transcriptomics of primary hepatocytes at the time of compound exposure) is scalable.

[0345] Referenced literature and alternative embodiments

[0346] All references cited in this article are incorporated herein by reference in their entirety and for all purposes as if each individual publication, patent, or patent application were specifically or individually indicated to be incorporated herein by reference in their entirety for all purposes.

[0347] This invention can be implemented as a computer program product comprising a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, the computer program product may contain program modules shown in any combination of Figures 1 and 2. These program modules may be stored on a CD-ROM, DVD, disk storage product, or any other non-transitory computer-readable data or program storage product.

[0348] This specification includes example systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments. Numerous specific details are set forth for purposes of explanation in order to provide an understanding of various embodiments of the subject matter of this invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of this invention can be practiced without these specific details. Generally, well-known illustrative examples, solutions, structures, and techniques are not shown in detail.

[0349] Many modifications and variations can be made to this invention without departing from its spirit and scope, as will be apparent to those skilled in the art. The specific embodiments described herein are given by way of example only. These embodiments were chosen and described to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention and its various embodiments with modifications suitable for the contemplated particular purpose. This invention is defined only by the full scope of the appended claims together with their equivalents.

Claims

1. A method for assessing the risk of a test chemical compound inducing liver injury in human subjects, the method comprising: (A) Obtain a cell component dataset in electronic form, the cell component dataset containing multiple cell component abundance values ​​of multiple cell components in a first plurality of cells that have been exposed to the test chemical compound for a first time period; (B) Using the aforementioned cellular component dataset, determine the corresponding pathway value for each of the multiple pathway modules, thereby obtaining multiple pathway values, wherein... Each of the plurality of pathway modules comprises an independent subset of the plurality of cellular components. Each independent subset of the multiple cellular components in each corresponding pathway module contains two or more cellular components. Each of the multiple cellular components is located in at least one independent subset of the multiple cellular components associated with a pathway module among the multiple pathway modules. The multiple cellular components include 10 or more cellular components, and Multiple differential cellular component abundance values ​​include 500 or more cellular component abundance values; (C) Input the plurality of path values ​​into one or more first models, thereby obtaining one or more first values ​​as outputs from the first models; (D) Input the abundance values ​​of the plurality of cellular components into one or more second models, thereby obtaining one or more second values ​​as outputs from the one or more second models; as well as (E) Use the one or more first values ​​and the one or more second values ​​to determine the risk that the test chemical compound will induce the liver injury in the human subject.

2. The method according to claim 1, wherein the first plurality of cells are primary human hepatocytes.

3. The method according to claim 1 or 2, wherein the abundance of each of the plurality of cellular components in each of the plurality of cells is determined by single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data of the plurality of cells.

4. The method of claim 1 or 2, wherein the abundance of each of the plurality of cellular components in each of the first plurality of cells is determined by batch RNA sequencing.

5. The method of claim 1 or 2, wherein the abundance of each of the plurality of cellular components in each of the first plurality of cells is determined by scTag-seq.

6. The method of claim 1 or 2, wherein the abundance of each of the plurality of cellular components in each of the first plurality of cells is determined by using single-cell assays of transposase-accessible chromatin (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.

7. The method according to any one of claims 1 to 6, wherein each of the one or more first models is a logistic regression model.

8. The method of claim 7, wherein each corresponding pathway value of each corresponding pathway module is a measure of the central tendency of the cellular component abundance values ​​of each cellular component in the independent subset of the cellular components corresponding to the corresponding subset of the cellular components of the corresponding pathway.

9. The method of claim 7, wherein the corresponding pathway value of the corresponding pathway module in the plurality of pathway modules is a negative logarithm of the false discovery rate to base 10, the false discovery rate being determined based on the distribution of the cellular component abundance of each cellular component in the independent subset of cellular components corresponding to the corresponding subset of cellular components of the corresponding pathway module relative to the cellular component abundance of cellular components in the cellular component dataset.

10. The method according to any one of claims 1 to 9, wherein the one or more second models are a single logistic regression model.

11. The method according to any one of claims 1 to 9, wherein the one or more second models are multiple logistic regression models.

12. The method of claim 11, wherein The first subset of the logistic regression models in the plurality of second models is a plurality of L1 regression models, and The second subset of the logistic regression models in the plurality of second models is a plurality of L2 regression models.

13. The method according to any one of claims 1 to 12, wherein the plurality of second models are two to 1000 logistic regression models.

14. The method of claim 1, wherein the one or more second models comprise 10 or more second modules, 20 or more second modules, 30 or more second modules, 40 or more second modules, 50 or more second modules, 60 or more second modules, 70 or more second modules, 80 or more second modules, or 90 or more second modules.

15. The method of claim 1, wherein each of the one or more first models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

16. The method of claim 1, wherein each of the one or more second models is a regression module, a neural network model, a support vector machine model, a Naive Bayes model, a nearest neighbor model, a boosting tree model, a random forest model, a decision tree model, a multinomial logistic regression model, a linear model, or a linear regression model.

17. The method according to any one of claims 1 to 16, wherein each of the plurality of cellular components is a specific gene, a specific mRNA associated with the gene, a carbohydrate, a lipid, an epigenetic feature, a metabolite, a protein, or a combination thereof.

18. The method according to any one of claims 1 to 17, the method further comprising identifying the plurality of pathway modules by including a process of: Obtain one or more first datasets in electronic form, said one or more first datasets comprising: For each corresponding primary human hepatocyte in a plurality of primary human hepatocytes, wherein the plurality of primary human hepatocytes comprises twenty or more primary human hepatocytes and collectively represents a plurality of annotated cell states: For each of the various cellular components: The corresponding abundance of the corresponding cellular components in the corresponding primary human hepatocytes. The multiple cellular components in the one or more first datasets were identified as multiple differentially expressed cellular components across the one or more first datasets; and Each independent subset of the cellular components in the plurality of cellular components is associated with a corresponding pathway of a pathway module in the plurality of pathway modules.

19. The method of 18, wherein the annotated cell state of the plurality of annotated cell states is a cell in the plurality of primary human hepatocytes exposed to a training compound under exposure conditions.

20. The method of claim 19, wherein the exposure condition is an exposure duration, a concentration of the training compound, or a combination of an exposure duration and a concentration of the training compound.

21. The method according to any one of claims 18 to 20, wherein the plurality of annotated cell states comprise exposure to a plurality of different concentrations of the first training compound.

22. The method of claim 21, wherein the plurality of different concentrations are (i) the Cmax or estimated Cmax of the first training compound to (ii) the IC50 of the first training compound. 10 Within the range.

23. The method of any one of claims 18 to 20, wherein the plurality of annotated cell states comprise exposure to the first training compound for a plurality of different exposure durations.

24. The method according to any one of claims 18 to 20, wherein the plurality of annotated cell states comprise each corresponding training compound exposed to a plurality of training compounds at correspondingly different concentrations.

25. The method of claim 24, wherein the corresponding plurality of different concentrations are at (i) the Cmax of the corresponding training compound to (ii) the IC50 of the corresponding training compound. 10 Within the range.

26. The method of any one of claims 18 to 19, wherein the plurality of annotated cell states comprise each of the plurality of training compounds exposed to the training compounds for a plurality of different exposure durations.

27. The method according to any one of claims 1 to 26, wherein the plurality of pathway modules comprises 10 to 2000 pathway modules.

28. The method according to any one of claims 1 to 27, wherein the plurality of cellular components comprises 100 to 35,000 cellular components.

29. The method according to any one of claims 1 to 28, wherein each independent subset of the multiple cellular components of each respective pathway module comprises two to three hundred cellular components.

30. The method according to any one of claims 1 to 29, wherein the test chemical compound is an organic compound with a molecular weight of less than 2000 Daltons.

31. The method according to any one of claims 1 to 30, wherein the test chemical compound is an organic compound that satisfies each of the criteria in the Lipinski five-fold rule.

32. The method according to any one of claims 1 to 31, wherein the test chemical compound is an organic compound that satisfies at least three of the criteria in the Lipinski five-fold rule.

33. The method according to any one of claims 1 to 32, wherein the risk is in the form of the probability that the test chemical compound will induce the liver injury in the human subject.

34. The method according to any one of claims 1 to 32, wherein the risk is one of a set of listed risk levels of the test chemical compound inducing the liver injury in the human subject.

35. The method of claim 34, wherein the enumerated set of risk levels comprises low-risk, intermediate-risk, and high-risk levels at which the test chemical compound will induce liver injury in the human subject.

36. The method according to any one of claims 1 to 32, wherein the risk is in the form of the lowest concentration of the test chemical compound that induces the liver injury in the human subject, as predicted by the first model and the one or more second models.

37. The method of claim 36, further comprising: The safety margin of the test chemical compound is determined using the lowest concentration of the test chemical compound predicted by the first model and the one or more second models to induce liver injury in the human subjects.

38. The method of claim 37, wherein determining the safety margin comprises dividing the risk by the Cmax of the test chemical compound.

39. The method according to any one of claims 1 to 38, wherein the liver injury is hepatitis, hepatotoxicity, jaundice, liver fibrosis, cirrhosis, hepatomegaly, alcoholic and non-alcoholic fatty liver disease and / or cholangiorrhea.

40. The method according to any one of claims 1 to 38, wherein the liver injury is characterized by elevated one or more liver enzymes, cholestasis, and / or alcoholic stools.

41. The method according to any one of claims 1 to 40, wherein the method informs the selection of one or more human subjects for treatment with the test chemical compound and / or the selection of one or more human subjects for continuing or discontinuing treatment with the test chemical compound.

42. The method according to any one of claims 1 to 41, wherein the method informs one or more human subjects undergoing treatment of the dosage, duration and / or frequency of administration of the test chemical compound.

43. The method according to any one of claims 1 to 42, wherein the method informs the design of a clinical trial, the clinical trial comprising the use of the test chemical compound.

44. The method according to any one of claims 1 to 42, wherein the method informs the design of an adaptive clinical trial, the adaptive clinical trial comprising the use of the test chemical compound.

45. The method of claim 43, wherein the method informs the clinical trial of the type and / or level of acceptable adverse events, one or more inclusion criteria and / or one or more exclusion criteria.

46. ​​The method of claim 44, wherein the method informs the type and / or level of acceptable adverse events in the adaptive clinical trial, one or more inclusion criteria and / or one or more exclusion criteria.

47. The method of claim 43, wherein the method informs one or more modifications to the clinical trial while the clinical trial is in progress.

48. The method of claim 44, wherein the method informs one or more modifications to the clinical trial while the adaptive clinical trial is in progress.

49. The method according to any one of claims 1 to 40, wherein the method informs the classification of the test chemical compound as a compound suitable for human therapy and / or testing or a compound unsuitable for human therapy and / or testing.

50. The method according to any one of claims 1 to 49, wherein the method further comprises formulating the test chemical compound for use in a therapy.

51. The method according to any one of claims 1 to 50, wherein each cell component abundance value in the cell component dataset represents the difference in cell component abundance between cells exposed to the test compound and cells exposed to a control.

52. The method of claim 51, wherein the control is dimethyl sulfoxide.

53. The method according to any one of claims 1 to 51, wherein the one or more first models consist of a single first model.

54. The method according to any one of claims 1 to 52, wherein the abundance value of each corresponding cell component is in the form of a logarithmic fold change in the abundance of the corresponding cell component relative to the abundance of the corresponding cell component in a second plurality of cells exposed to the control solution for the first time period.

55. The method of claim 54, wherein the control solution is dimethyl sulfoxide.

56. A computer system comprising one or more processors and a memory storing instructions for performing a method for assessing the risk that a test chemical compound will induce liver or heart damage in a human subject, the method comprising: (A) Obtain a cell component dataset in electronic form, the cell component dataset containing multiple cell component abundance values ​​of multiple cell components in a first plurality of cells that have been exposed to the test chemical compound for a first time period; (B) Using the aforementioned cellular component dataset, determine the corresponding pathway value for each of the multiple pathway modules, thereby obtaining multiple pathway values, wherein... Each of the plurality of pathway modules comprises an independent subset of the plurality of cellular components. Each independent subset of the multiple cellular components in each corresponding pathway module contains two or more cellular components. Each of the multiple cellular components is located in at least one independent subset of the multiple cellular components associated with a pathway module among the multiple pathway modules. The multiple cellular components include 10 or more cellular components, and Multiple differential cellular component abundance values ​​include 500 or more cellular component abundance values; (C) Input the plurality of path values ​​into one or more first models, thereby obtaining one or more first values ​​as outputs from the first models; (D) Input the abundance values ​​of the plurality of cellular components into one or more second models, thereby obtaining one or more second values ​​as outputs from the one or more second models; as well as (E) Use the one or more first values ​​and the one or more second values ​​to determine the risk that the test chemical compound will induce the liver injury in the human subject.

57. A computer system comprising one or more processors and a memory, the memory storing instructions for performing a method according to any one of claims 1 to 55 for assessing the risk that a test chemical compound will induce liver or heart damage in human subjects.

58. A non-transitory computer-readable medium storing one or more computer programs executable by a computer to assess the risk of a test chemical compound inducing liver or heart damage in human subjects, said computer comprising one or more processors and memory, said one or more computer programs collectively encoding computer-executable instructions for performing methods comprising: (A) Obtain a cell component dataset in electronic form, the cell component dataset containing multiple cell component abundance values ​​of multiple cell components in a first plurality of cells that have been exposed to the test chemical compound for a first time period; (B) Using the aforementioned cellular component dataset, determine the corresponding pathway value for each of the multiple pathway modules, thereby obtaining multiple pathway values, wherein... Each of the plurality of pathway modules comprises an independent subset of the plurality of cellular components. Each independent subset of the multiple cellular components in each corresponding pathway module contains two or more cellular components. Each of the multiple cellular components is located in at least one independent subset of the multiple cellular components associated with a pathway module among the multiple pathway modules. The multiple cellular components include 10 or more cellular components, and Multiple differential cellular component abundance values ​​include 500 or more cellular component abundance values; (C) Input the plurality of path values ​​into one or more first models, thereby obtaining one or more first values ​​as outputs from the first models; (D) Input the abundance values ​​of the plurality of cellular components into one or more second models, thereby obtaining one or more second values ​​as outputs from the one or more second models; as well as (E) Use the one or more first values ​​and the one or more second values ​​to determine the risk that the test chemical compound will induce the liver injury in the human subject.

59. A non-transitory computer-readable medium storing one or more computer programs executable by a computer to assess the risk that a test chemical compound will induce liver or heart damage in human subjects, said computer comprising one or more processors and memory, said one or more computer programs collectively encoding computer-executable instructions for performing the method according to any one of claims 1 to 55.