Disease detection systems and methods
The human hepatic abnormality detection system uses single-nucleus or single-cell transcriptome data to segment hepatic abnormalities, addressing the limitations of current detection methods by enabling accurate identification of genetic markers and pathways for targeted therapeutics.
Patent Information
- Application Number
- PCT/US2024/061132
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for detecting human hepatic abnormalities, such as nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH), are limited by the need for liver biopsies and lack of accurate systematic screening tools.
A method for manufacturing a human hepatic abnormality detection system that utilizes single-nucleus or single-cell transcriptome data from liver tissue samples to segment hepatic abnormalities into meaningful subcategories based on transcriptome expression differences, allowing for the identification of genetic pathways and specific genes associated with each subcategory.
Enables reliable and accurate detection of hepatic abnormalities, facilitating the identification of genetic markers and pathways, which can serve as a basis for developing targeted therapeutics for specific hepatic states.
Smart Images

Figure US2024061132_26062025_PF_FP_ABST
Abstract
Description
DISEASE DETECTION SYSTEMS AND METHODSTECHNICAL FIELD
[0001] The present disclosure is directed to systems, methods, and devices for manufacturing human hepatic abnormality detection systems.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 612,320, filed on December 19, 2023, the content of which is hereby incorporated by reference in its entirety.BACKGROUND
[0003] Nonalcoholic fatty liver disease (NAFLD) is a common form of liver disease and abnormal liver function tests and its prevalence is increasing due to the rise in obesity. See, Marchesini et al., 2008 “Obesity-associated liver disease,” J Clin Endocrinol Metab. 93(11 Suppl l):S74-S80. Progression of NAFLD may lead to nonalcoholic steatohepatitis (NASH), marked by inflammation of the liver, and can further progress to fibrosis and eventual cirrhosis. Although simple NAFLD is usually benign and does not frequently progress to more advanced stages of liver disease, because of its high prevalence it is an increasing public health concern and a leading cause of cirrhosis. See, Dyson et al., 2014, “Non-alcoholic fatty liver disease - a practical approach to diagnosis and staging,” Frontline Gastroenterol 5:211-218; and Williams et al., 2014, “Addressing liver disease in the UK: a blueprint for attaining excellence in health care and reducing premature mortality from lifestyle issues of excess consumption of alcohol, obesity, and viral hepatitis,” Lancet 384: 1953-1997. The estimated prevalence of NAFLD and NASH in the general population varies based on diagnostic method: NAFLD prevalence is estimated to be between 6.3% and 33% and NASH around 3%-5%. See, Chalasani et al., 2012 “The diagnosis and management of non-alcoholic fatty liver disease: practice Guideline by the American Association for the Study of Liver Diseases, American College of Gastroenterology, and the American Gastroenterological Association,” Hepatology 55:2005-2023. However, accurate prevalence of NASH is difficult to estimate because liver biopsy is required under conventional practice and thus systematic screening or prospective study in the general population is not possible. See, Gaidos et al., 2008, “A decision analysis study of the value of a liver biopsy innonalcoholic steatohepatitis,” Liver Int. 28:650-658. There are currently no drugs approved specifically for NASH or liver fibrosis.
[0004] Given the above background, what is needed in the art are human hepatic abnormality detection systems, based on improved human hepatic atlases, that can be used to reliably segment human hepatic abnormalities into meaningful subcategories (hepatic states) based on relevant and actual differences in transcriptome expression between the different subcategories. Such segmentation can be used to identify genetic pathways and particular genes that are associated with each such subcategory and thus be used as new basis for developing therapeutics that address particular hepatic states.SUMMARY
[0005] The present disclosure addresses the above-identified shortcomings. A method for manufacturing a human hepatic abnormality detection system is provided at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors. The method obtains, in electronic form, first information comprising single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells.
[0006] Each nucleus or cell in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples.
[0007] Each liver tissue sample in the plurality of liver tissue samples is from a different subject in a cohort of subjects.
[0008] The first plurality of nuclei or cells includes a different subset of nuclei or cells from a liver tissue sample from each subject in the cohort of subjects.
[0009] The first plurality of nuclei or cells comprises at least 1000 nuclei or cells. In some embodiments each liver tissue sample is frozen.
[0010] In some embodiments at least some of the liver tissue samples are not frozen. In some embodiments at least some of the liver tissue samples are frozen.
[0011] In some embodiments, the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes. In some embodiments, the plurality of genes comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 of the genes listed in Table 2 below.
[0012] In some embodiments, the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
[0013] In some embodiments, the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
[0014] In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50 liver tissue samples, or 100 or more liver tissue samples.
[0015] In some embodiments, the first information further comprises first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state. At least a first subset of subjects in the cohort of subjects have the first hepatic state and a second subset of subjects in the cohort of subjects have the second hepatic state.
[0016] In some embodiments, the first hepatic state or the second hepatic state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis (NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
[0017] In some embodiments, the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
[0018] In some embodiments, the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
[0019] In some embodiments, the first metadata for each respective subject in the cohort of subjects, indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state, comprises a histologically graded disease status for the respective subject.
[0020] In some embodiments, the first metadata for each respective subject in the cohort of subjects comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each liver tissue in the plurality of liver tissue samples.
[0021] In some embodiments, the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
[0022] In some embodiments, sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
[0023] In some embodiments, the single-nucleus or single-cell transcriptome data is filtered to remove counts of ambient RNA molecules, doublets and / or empty droplets.
[0024] In some embodiments, the first hepatic state is fibrosis and the second hepatic state is absence of fibrosis.
[0025] In some embodiments, the first hepatic state is absence of fibrosis and the second hepatic state is fibrosis.
[0026] In some embodiments, the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
[0027] In some embodiments, the first hepatic state is presence of liver inflammation and the second hepatic state is absence of liver inflammation.
[0028] In some embodiments, the first hepatic state is absence of liver inflammation and the second hepatic state is presence of liver inflammation.
[0029] In some embodiments, the first hepatic state is a first stage of liver inflammation and the second hepatic state is a second stage of liver inflammation.
[0030] In some embodiments, the first hepatic state is presence of liver steatosis and the second hepatic state is absence of liver steatosis.
[0031] In some embodiments, the first hepatic state is absence of liver steatosis and the second hepatic state is presence of liver steatosis.
[0032] In some embodiments, the first hepatic state is a first stage of liver steatosis and the second hepatic state is a second stage of liver steatosis.
[0033] In some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex. The secondinformation is used to prune the cohort of subjects based on age, body mass index, or sex. This causes the cohort of subjects to be free of confounding for age, body mass index, or sex. The pruning causes a subset of subjects to be removed from the cohort of subjects.
[0034] In some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex. The second information is used to prune the cohort of subjects based on age, body mass index, and sex. This causes the cohort of subjects to be free of confounding for age, body mass index, and sex. The pruning causes a subset of subjects to be removed from the cohort of subjects.
[0035] The respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data.
[0036] In accordance with the method, the first plurality of nuclei or cells is clustered into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function.
[0037] In some such embodiments, the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells. Each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei or cells and (ii) a respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function.
[0038] In accordance with the method, the first metadata is used to identify a first cluster in the plurality of clusters with the first hepatic state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first hepatic state.
[0039] In some such embodiments, the first cluster is used to determine an extent to which race is a covariate with respect to the first hepatic state.
[0040] In some such embodiments, the first cluster is used to determine an extent to which sex is a covariate with respect to the first hepatic state.
[0041] In some such embodiments, the first cluster is used to determine an extent to which age is a covariate with respect to the first hepatic state.
[0042] In some such embodiments, the first cluster is used to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first hepatic state.
[0043] In some embodiments, the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage, and the first cluster is used to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first hepatic state.
[0044] In some embodiments, genotype data for each subject in the plurality of subjects is obtained. The genotype data for each subject represented in the first cluster is overlayed with a hepatic state of each subject in the first cluster.
[0045] In some embodiments, genotype data for each subject in the plurality of subjects is obtained. An extent to which a genotype is a covariate for the first hepatic state is determined using the genotype data for each subject represented in the first cluster.
[0046] In some embodiments, the hepatic abnormality detection system is used to associate a test subject with the first hepatic state by a procedure comprising obtaining, in electronic form, second information comprising single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells. Each nucleus or cell in the second plurality of nuclei or cells is obtained from a liver tissue sample obtained from the test subject. The first plurality of nuclei or cells and the second plurality of nuclei or cells are co-clustered into the plurality of clusters. The test subject is identified as having the first hepatic state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
[0047] In some embodiments, the method informs a response to a drug compound in a patient or in a plurality of patients.
[0048] In some embodiments, the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
[0049] In some embodiments, the first metadata is used to identify a second cluster in the plurality of clusters with the second hepatic state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second hepatic state.
[0050] In some embodiments, the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
[0051] In some embodiments, a determination is made that the first cluster comprises nuclei or cells of quiescent hepatic stellate cells and the second cluster comprises nuclei or cells of activated hepatic stellate cells.
[0052] In some embodiments, a determination is made that the first cluster comprises nuclei or cells of a first type of activated hepatic stellate cells and the second cluster comprises nuclei or cells of a second type of activated hepatic stellate cells.
[0053] In some embodiments, a determination is made that the first cluster comprises nuclei or cells of activated hepatic stellate cells and the second cluster comprises nuclei or cells of RGS5+ stellate cells.
[0054] In some embodiments, a determination is made that the first cluster comprises nuclei or cells of activated hepatic stellate cells and the second cluster comprises nuclei or cells of vascular smooth muscle cells.
[0055] In some embodiments, the first hepatic state is absence of fibrosis or inflammation and the second hepatic state is presence of fibrosis or inflammation.
[0056] In some embodiments, a further determination is made that the first cluster comprises quiescent stellate cells or nuclei or quiescent stellate cells and the second cluster comprises activated stellate cells or nuclei of activated stellate cells.
[0057] In some embodiments, the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes. One or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
[0058] In some embodiments, the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state. In some embodiments a plurality of compound-specific differential transcriptional signatures is accessed in electronic form. Each respective compound-specific differential transcriptional signature is a difference between (i)a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set. The baseline transcriptional signature data set is from a control sample of one or more cells of a cell type. Each respective compound-treated transcriptional signature is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds. A test differential transcription signature is determined by differential comparison of the transcriptional signature of the nuclei or cells of the first cluster and the second cluster. The test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
[0059] In some embodiments, the hepatic disease state is selected from an acute stage, a chronic stage, a clinical stage, a flare-up, a remission, a progressive stage, a refractory, a subclinical stage, and a terminal phase of a hepatic disease.
[0060] In some embodiments, the control sample and each corresponding compound-treated sample is exposed to a solvent. The solvent is the same solvent for the control sample and each corresponding compound-treated sample. Optionally, the solvent is dimethyl sulfoxide.
[0061] In some embodiments, the control sample and each corresponding compound-treated sample is exposed to a polar aprotic solvent. The polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
[0062] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data. Optionally, the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
[0063] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of singlecell RNA sequencing (scRNA-seq) data.
[0064] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises scRNA-seq data.
[0065] In some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of scRNA-seq data.
[0066] In some such embodiments, each corresponding compound-treated sample of one or more cells comprises hepatic stellate cells.
[0067] In some embodiments, each corresponding compound-treated sample of one or more cells comprises a first type of activated hepatic stellate cells.
[0068] In some embodiments, each corresponding compound-treated sample of one or more cells comprises a second type of activated hepatic stellate cells.
[0069] In some embodiments, each corresponding compound-treated sample of one or more cells consists of hepatic stellate cells.
[0070] In some embodiments, each corresponding compound-treated sample of one or more cells consists of a first type of activated hepatic stellate cells.
[0071] In some embodiments, each corresponding compound-treated sample of one or more cells consists of a second type of activated hepatic stellate cells.
[0072] In some embodiments, each corresponding compound-treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
[0073] In some embodiments, each corresponding compound-treated sample of the one or more cells is a frozen sample. In some embodiments, each corresponding compound-treated sample of one or more cells is an unfrozen sample.
[0074] In some embodiments, each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
[0075] In some embodiments, each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
[0076] Another aspect of the present disclosure provides a computer system having one or more processors, and memory storing one or more programs for execution by the one or moreprocessors, the one or more programs comprising instructions for performing any of the methods and / or embodiments disclosed herein.
[0077] Another aspect of the present disclosure provides a non-transitory computer readable storage medium storing one or more programs configured for execution by a computer, the one or more programs comprising instructions for carrying out any of the methods and / or embodiments disclosed herein.
[0078] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS
[0079] The embodiments disclosed herein are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. Like reference numerals refer to corresponding parts throughout the drawings.
[0080] Figures 1 illustrates a block diagram of an exemplary system and computing device for manufacturing a human hepatic abnormality detection system, in accordance with an embodiment of the present disclosure.
[0081] Figures 2A, 2B, 2C, 2D, 2E, 2F, 2G, 2H, 21, 2J, 2K, and 2L collectively provide a flow chart of processes and features of an example method manufacturing a human hepatic abnormality detection system, in accordance with various embodiments of the present disclosure.
[0082] Figure 3 illustrates how the high dimensional single-nucleus transcriptome data for a plurality of genes for a plurality of nuclei has been reduced two dimensions for visualization. Each point represents a different nuclei in the transcriptome data of 832,000 nuclei from 103 different subjects. The cluster assignment of the various nuclei represented in Figure 3 is shown by color coding and labeling, in accordance with an embodiment of the present disclosure.
[0083] Figure 4 illustrates nuclei of stellate cells that are quiescent in cluster 402 (left panel) and nuclei of stellate cells that are activated in cluster 404 (right panel). From first metadata it is known that the nuclei of the quiescent stellate cells of cluster 402 are from subjects that are F0 on the fibrotic scale whereas the nuclei of stellate cells that are activated of cluster 404 are from subjects that are F3 on the fibrotic scale.
[0084] Figure 5 illustrates how a first standardized set of biomarkers that each define a unique cell type can be used to identify which clusters of Figure 4 specific cell types are found, in accordance with an embodiment of the present disclosure.
[0085] Figure 6 illustrates a summary of the cell specific type clustering identified in Figure 5, in accordance with an embodiment of the present disclosure.
[0086] Figure 7 illustrates an enlarged view of cluster 404 of Figure 4, in accordance with an embodiment of the present disclosure.
[0087] Figure 8 compares the relative expression of a muscle associated signature (genes ACTA2, CARMN, LM0D1, MY01E, SLIT3, CRIM1, MRV1, MAG11, COL4A1, and MYOF) in a first type of activated hepatic stellate cells (HSC activated l) versus a second type of activated hepatic stellate cells (HSC_activated_2), in accordance with an embodiment of the present disclosure.
[0088] Figure 9 illustrates summary statistics for certain first metadata for a cohort of subjects in accordance with an embodiment of the present disclosure.
[0089] Figure 10 provides an analysis of the transcriptome data for each of the nuclei in a first plurality of nuclei of a cohort of subjects including the various cell types represented (left hand column), how many cells of each of these cell types is represented in the first plurality of nuclei (middle column), and the average number per subject in the cohort of subject per cell type (right-hand column) in accordance with an embodiment of the present disclosure.
[0090] Figure 11 illustrates how ambient RNA correction reduced hepatocyte marker score in non-hepatocyte cell types without affecting cell type specific marker scores in accordance with an embodiment of the present disclosure. In Figure 11, “pre-CB” refers to prior to ambient RNA correction whereas “CB” refers to after ambient RNA correction.
[0091] Figure 12 illustrates a UMAP of the clustering of transcriptome data limited to those nuclei in a plurality of nuclei that have the transcriptional cell type signature of macrophages in accordance with an embodiment of the present disclosure.
[0092] Figure 13 illustrates a difference within the clustering of Figure 12 between healthy cells (left panel) and F3 fibrotic cells (right panel) in accordance with an embodiment of the present disclosure.
[0093] Figure 14 illustrates a UMAP of the clustering of transcriptome data limited to those nuclei in the plurality of nuclei that have the transcriptional cell type signature of T cells.
[0094] Figure 15 illustrates the difference within the clustering of Figure 14 between healthy cells (left panel) and fibrotic cells (right panel).
[0095] Figure 16 illustrates a transition from quiescent to activated stellate cells and its association with fibrosis in accordance with an embodiment of the present disclosure.
[0096] Figures 17A and 17B illustrate a method of identifying a compound that transitions a hepatic disease state to a healthy state using a NASH atlas in accordance with an embodiment of the present disclosure.
[0097] Figure 18 illustrates fractional presence of cell type and expression of marker genes in a transition from quiescent to activated stellate cells in accordance with an embodiment of the present disclosure.
[0098] Figure 19 illustrates the stimulation of LX-2 cells with TGF-bl in accordance with an embodiment of the present disclosure.
[0099] Figure 20 illustrates a schematic approach to identifying activated stellate cell line and then using the activated stellate cell line to search for modulators (e.g., compounds) that will revert this activated state in accordance with an embodiment of the present disclosure.
[0100] Figure 21 illustrates how canonical polarization states can be assigned to the macrophage and Kupffer cell clusters of Figure 12 in accordance with an embodiment of the present disclosure.
[0101] Figures 22A and 22B illustrates how there is a significant shift towards fibrosis, steatosis, and inflammation towards the macrophage M4>M3 state and that the KCM2lowstate is more associated with a healthy state in accordance with an embodiment of the present disclosure.
[0102] Figures 23 A and 23B illustrate a method of identifying a compound that transitions a hepatic disease state to a healthy state using a NASH atlas in accordance with an embodiment of the present disclosure.
[0103] Figure 24 illustrates an alternative cell state transition that assume that most or all M3 macrophage behaviors help to control NASH progression in accordance with an embodiment of the present disclosure.
[0104] Figure 25 illustrates how peripheral blood mononuclear cells (PBMCs) can be primed with GM-CSF to activate the M3-like signature in accordance with an embodiment of the present disclosure.
[0105] Figure 26 illustrates how biological pathway analysis identified biological pathways associated with NASH M3 macrophages in accordance with an embodiment of the present disclosure.DETAILED DESCRIPTION
[0106] Given the above background, the present disclosure provides systems and methods for manufacturing a human hepatic abnormality detection system that obtain single-nucleus or single-cell transcriptome data for a plurality of genes for each of a plurality of nuclei or cells. Each nucleus or cell is obtained from liver tissue samples. Each liver sample is from a different subject in a cohort of subjects. Metadata for each subject in the cohort is also obtained, indicating for each respective subject whether they have a first or second hepatic state. Optionally, the transcriptome data for the plurality of genes for each nucleus or cell is barcoded with the corresponding subject in the cohort. The nuclei or cells are clustered into clusters by computing distances with the transcriptome data for the genes for each unique pair of nuclei or cells in the plurality of nuclei or cells and by evaluating the distances with a criterion function. Each distance represents a different pair of nuclei or cells in the plurality of nuclei or cells and quantifies a distance between (i) a vector formed by the transcriptome data for the plurality of genes for a respective first nucleus or cell and (ii) a vector formed by the transcriptome data for the plurality of genes for a second nucleus or cell. Each cluster represents a subset of nuclei or cells of the plurality of nuclei or cells clustered together based on evaluation of distances with the criterion function. The metadata is used to identify a cluster in the plurality of clusters with the first hepatic state by determining that the cluster includes nuclei or cells from subjects in the cohort having the first hepatic state.
[0107] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-knownmethods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0108] Plural instances may be provided for components, operations or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other forms of functionality are envisioned and may fall within the scope of the implementation(s). In general, structures and functionality presented as separate components in the example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the implementation(s).
[0109] It will also be understood that, although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first dataset could be termed a second dataset, and, similarly, a second dataset could be termed a first dataset, without departing from the scope of the present invention. The first dataset and the second dataset are both datasets, but they are not the same dataset.
[0110] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the claims. As used in the description of the implementations and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0111] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in accordance with a determination” or “in response to detecting,” that a stated condition precedent is true, depending on the context. Similarly, the phrase “if it is determined (that a stated condition precedent is true)” or “if (a stated condition precedent is true)” or “when (a stated condition precedent is true)” may be construed to mean “upon determining” or “in response to determining” or “in accordance with a determination”or “upon detecting” or “in response to detecting” that the stated condition precedent is true, depending on the context.
[0112] Furthermore, when a reference number is given an “zth” denotation, the reference number refers to a generic component, set, or embodiment. For instance, a cellular- component termed “cellular-component z” refers to the zthcellular-component in a plurality of cellular-components.
[0113] In the interest of clarity, not all of the routine features of the implementations described herein are shown and described. It will be appreciated that, in the development of any such actual implementation, numerous implementation-specific decisions are made in order to achieve the designer’s specific goals, such as compliance with use case- and business-related constraints, and that these specific goals will vary from one implementation to another and from one designer to another. Moreover, it will be appreciated that such a design effort might be complex and time-consuming, but nevertheless be a routine undertaking of engineering for those of ordering skill in the art having the benefit of the present disclosure.
[0114] Some portions of this description describe the embodiments of the invention in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like.
[0115] The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based hereon. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention.
[0116] In general, terms used in the claims and the specification are intended to be construed as having the plain meaning understood by a person of ordinary skill in the art. Certain terms are defined below to provide additional clarity. In case of conflict between the plain meaning and the provided definitions, the provided definitions are to be used.
[0117] Any terms not directly defined herein shall be understood to have the meanings commonly associated with them as understood within the art of the invention. Certain terms are discussed herein to provide additional guidance to the practitioner in describing the compositions, devices, methods and the like of aspects of the invention, and how to make or use them. It will be appreciated that the same thing may be said in more than one way. Consequently, alternative language and synonyms may be used for any one or more of the terms discussed herein. No significance is to be placed upon whether or not a term is elaborated or discussed herein. Some synonyms or substitutable methods, materials and the like are provided. Recital of one or a few synonyms or equivalents does not exclude use of other synonyms or equivalents, unless it is explicitly stated. Use of examples, including examples of terms, is for illustrative purposes only and does not limit the scope and meaning of the aspects of the invention herein.
[0118] Definitions.
[0119] As used herein, the term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, in some embodiments “about” means within 1 or more than 1 standard deviation, per the practice in the art. In some embodiments, “about” means a range of ±20%, ±10%, ±5%, or ±1% of a given value. In some embodiments, the term “about” or “approximately” means within an order of magnitude, within 5-fold, or within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value can be assumed. All numerical values within the detailed description herein are modified by “about” the indicated value, and consider experimental error and variations that would be expected by a person having ordinary skill in the art. The term “about” can have the meaning as commonly understood by one of ordinary skill in the art. In some embodiments, the term “about” refers to ±10%. In some embodiments, the term “about” refers to ±5%.
[0120] As used herein, the terms “abundance,” “abundance level,” or “expression level” refers to an amount of a cellular constituent (e.g., a gene product such as an RNA species, e.g., mRNA or miRNA, or a protein molecule) present in one or more cells, or an average amount of a cellular constituent present across multiple cells. When referring to mRNA or protein expression, the term generally refers to the amount of any RNA or protein species corresponding to a particular genomic locus, e.g, a particular gene. However, in someembodiments, an abundance can refer to the amount of a particular isoform of an mRNA or protein corresponding to a particular gene that gives rise to multiple mRNA or protein isoforms. The genomic locus can be identified using a gene name, a chromosomal location, or any other genetic mapping metric.
[0121] As used interchangeably herein, a “cell state” or “biological state” refers to a state or phenotype of a cell or a population of cells. For example, a cell state can be healthy or diseased. A cell state can be one of a plurality of diseases. A cell state can be a response to a compound treatment and / or a differentiated cell lineage. A cell state can be characterized by a measure (e.g., an activation, expression, and / or measure of abundance) of one or more cellular constituents, including but not limited to one or more genes, one or more proteins, and / or one or more biological pathways.
[0122] As used herein, a “cell state transition” or “cellular transition” refers to a transition in a cell’s state from a first cell state to a second cell state. In some embodiments, the second cell state is an altered cell state (e.g., a healthy cell state to a diseased cell state). In some embodiments, one of the respective first cell state and second cell state is an unperturbed state and the other of the respective first cell state and second cell state is a perturbed state caused by an exposure of the cell to a condition. The perturbed state can be caused by exposure of the cell to a compound. A cell state transition can be marked by a change in cellular constituent abundance in the cell, and thus by the identity and quantity of cellular constituents (e.g., mRNA, transcription factors) produced by the cell (e.g., a perturbation signature).
[0123] As used herein, the term “dataset” in reference to cellular constituent abundance measurements for a cell or a plurality of cells can refer to a high-dimensional set of data collected from a single cell (e.g., a single-cell cellular constituent abundance dataset) in some contexts. In other contexts, the term “dataset” can refer to a plurality of high-dimensional sets of data collected from single cells (e.g., a plurality of single-cell cellular constituent abundance datasets), each set of data of the plurality collected from one cell of a plurality of cells.
[0124] As used herein, the term “differential abundance” or “differential expression” refers to differences in the quantity and / or the frequency of a cellular constituent present in a first entity (e.g., a first cell, plurality of cells, and / or sample) as compared to a second entity e.g., a second cell, plurality of cells, and / or sample). In some embodiments, a first entity is a sample characterized by a first cell state (e.g., a diseased phenotype) and a second entity is a sample characterized by a second cell state (e.g., a normal or healthy phenotype). Forexample, a cellular constituent can be a polynucleotide (e.g., an mRNA transcript) which is present at an elevated level or at a decreased level in entities characterized by a first cell state compared to entities characterized by a second cell state. In some embodiments, a cellular constituent can be a polynucleotide which is detected at a higher frequency or at a lower frequency in entities characterized by a first cell state compared to entities characterized by a second cell state. A cellular constituent can be differentially abundant in terms of quantity, frequency or both. In some instances, a cellular constituent is differentially abundant between two entities if the amount of the cellular constituent in one entity is statistically significantly different from the amount of the cellular constituent in the other entity. For example, a cellular constituent is differentially abundant in two entities if it is present at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% greater in one entity than it is present in the other entity, or if it is detectable in one entity and not detectable in the other. In some instances, a cellular constituent is differentially expressed in two sets of entities if the frequency of detecting the cellular constituent in a first subset of entities (e.g., cells representing a first subset of annotated cell states) is statistically significantly higher or lower than in a second subset of entities (e.g., cells representing a second subset of annotated cell states). For example, a cellular constituent is differentially expressed in two sets of entities if it is detected at least about 120%, at least about 130%, at least about 150%, at least about 180%, at least about 200%, at least about 300%, at least about 500%, at least about 700%, at least about 900%, or at least about 1000% more frequently or less frequently observed in one set of entities than the other set of entities.
[0125] As used herein, the term “sample,” “biological sample,” or “patient sample,” refers to any sample taken from a subject, which can reflect a biological state associated with the subject. Examples of samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, fecal, saliva, sweat, tears, pleural fluid, pericardial fluid, or peritoneal fluid of the subject. A sample can include any tissue or material derived from a living or dead subject. A sample can be a cell-free sample. A sample can comprise one or more cellular constituents. For instance, a sample can comprise a nucleic acid (e.g., DNA or RNA) or a fragment thereof, or a protein. The term “nucleic acid” can refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) or any hybrid or fragment thereof. The nucleic acid in the sample can be a cell-free nucleic acid. A sample can be a liquid sample or a solid sample (e.g., a cell or tissue sample). A sample can be a bodily fluid. Asample can be a stool sample. A sample can be treated to physically disrupt tissue or cell structure (e.g., centrifugation and / or cell lysis), thus releasing intracellular components into a solution which can further contain enzymes, buffers, salts, detergents, and the like which can be used to prepare the sample for analysis.
[0126] The terms “sequence reads” or “reads,” used interchangeably herein, refer to nucleotide sequences produced by any sequencing process described herein or known in the art. Reads can be generated from one end of nucleic acid fragments (“single-end reads”), and sometimes are generated from both ends of nucleic acids (e.g., paired-end reads, double-end reads). The length of the sequence read is often associated with the particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of a mean, median or average length of about 15 bp to 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In some embodiments, the sequence reads are of a mean, median or average length of about 1000 bp or more. Nanopore sequencing, for example, can provide sequence reads that vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads vary to a lesser extent (e.g, where most sequence reads are of a length of about 200 bp or less). A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g, a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from part of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. A sequence read can be obtained in a variety of ways, e.g., using sequencing techniques or using probes (e.g., in hybridization arrays or capture probes) or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.
[0127] As disclosed herein, the terms “sequencing,” “sequence determination,” and the like refer generally to any and all biochemical processes that may be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as a DNA fragment.
[0128] As used herein, the term “tissue” corresponds to a group of cells that group together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells or blood cells), but also can correspond to tissue from different organisms (mother vs. fetus) or to healthy cells vs. tumor cells. The term “tissue” can generally refer to any group of cells found in the human body (e.g., heart tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some aspects, the term “tissue” or “tissue type” can be used to refer to a tissue from which a cell-free nucleic acid originates. In one example, viral nucleic acid fragments can be derived from blood tissue. In another example, viral nucleic acid fragments can be derived from tumor tissue.
[0129] I. Exemplary System Embodiments
[0130] Now that an overview of some aspects of the present disclosure and some definitions used in the present disclosure have been provided, details of an exemplary system are described in conjunction with Figure 1.
[0131] Figure 1 illustrates a computer system 100 for manufacturing a human hepatic abnormality detection system. In typical embodiments, computer system 100 comprises one or more computers. For purposes of illustration in Figure 1, the computer system 100 is represented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not so limited. The functionality of the computer system 100 may be spread across any number of networked computers and / or reside on each of several networked computers and / or virtual machines. One of skill in the art will appreciate that a wide array of different computer topologies is possible for the computer system 100 and all such topologies are within the scope of the present disclosure.
[0132] Turning to Figure 1 with the foregoing in mind, the computer system 100 comprises one or more processing units (CPUs) 52, a network or other communications interface 54, a user interface 56 (e.g., including an optional display 58 and optional input 60 (e.g. keyboard or other form of input device)), a memory 92 (e.g, random access memory, persistent memory, or combination thereof), and one or more communication busses 94 for interconnecting the aforementioned components. To the extent that components of memory 92 are not persistent, data in memory 92 can be seamlessly shared with non-volatile memory (not shown) or portions of memory 92 that are non-volatile / persistent using known computing techniques such as caching. Memory 92 can include mass storage that is remotely located with respect to the central processing unit(s) 52. In other words, some data stored inmemory 92 may in fact be hosted on computers that are external to computer system 100 but that can be electronically accessed by the computer system 100 over an Internet, intranet, or other form of network or electronic cable using network interface 54. In some embodiments, the computer system 100 makes use of models that are run from the memory associated with one or more graphical processing units in order to improve the speed and performance of the system. In some alternative embodiments, the computer system 100 makes use of models that are run from memory 92 rather than memory associated with a graphical processing unit.
[0133] The memory 92 of the computer system 100 stores:• an optional operating system 102 that includes procedures for handling various basic system services;• an analysis module 103 for manufacturing a human hepatic abnormality detection system;• first information 104 comprising below referenced single-nucleus or single-cell transcriptome data 106 and first metadata 116;• single-nucleus or single-cell transcriptome data 106 for a plurality of genes 112 (112- 1-1, ..., 112-1-Q; 112-P-l, ...., 112-P-Q) for each nucleus or cell 108 (108-1, ..., 108- P) in a first plurality of nuclei or cells, each nucleus or cell including a barcode 110 (110-1, ..., 110-P);• first metadata 116 for each respective subject 118 (118-1, ..., 118-X) in a cohort of subjects indicating their hepatic state 120 (e.g., 120-1), their barcode 110 (e.g., 110- 1), sex 124 (e.g., 124-1), race 126 (e.g., 126-1), age 128 (e.g., 128-1), and one or more physical conditions 130 (e.g., 130-1, e.g., body mass index);• Clusters 132, each cluster 134 (134-1, ..., 134-Z) representing a subset 136 (136-1, ... 136-Z) of the plurality of nuclei or cells found in the single-nucleus or single-cell transcriptome data 106 that cluster together based on their respective transcriptome data.
[0134] In some implementations, one or more of the above identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices, and correspond to a set of instructions for performing a function described above. The above identified data, modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules may be combined or otherwise re-arranged in various implementations. In some implementations, the memory 92 optionally stores a subset of the modules and datastructures identified above. Furthermore, in some embodiments the memory 92 stores additional modules and data structures not described above.
[0135] II. Methods for manufacturing a hepatic abnormality detection system
[0136] Referring to block 200 of Fig. 2A, a method for manufacturing a human hepatic abnormality detection system is provided at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors.
[0137] Referring to block 202 first information is obtained that comprises single-nucleus, or single-cell, transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells. Each nucleus or cell in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples. Each liver tissue sample in the plurality of liver tissue samples is from a different subject in a cohort of subjects. The first plurality of nuclei or cells includes a different subset of nuclei or cells from a liver tissue sample from each subject in the cohort of subjects. In some embodiments, the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, at least 2,000 nuclei or cells, at least 3,000 nuclei or cells, at least 4,000 nuclei or cells, at least 5,000 nuclei or cells, or at least 10,000 nuclei or cells. In some embodiments the first plurality of nuclei or cells includes at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nuclei or cells from each liver tissue sample and each such liver tissue sample is from a different subject in the cohort of subjects. Throughout the present disclosure it will be understood that while certain advantages may be incurred by using single-nucleus data, single-cell data can be used. As such, each metric or range that is referenced in relation to the plurality of nuclei (such as the number of nuclei) is equally applicable to embodiments that make use of singlecell transcriptome data.
[0138] In a non-limiting example, the liver sample is harvested and / or obtained from a deceased liver donor and then frozen (e.g., treated with and / or frozen in liquid nitrogen), and nuclei of cells from the samples are isolated and analyzed (e.g., sequenced) in single-nucleus assay experiments (e.g., snRNA-seq). In embodiments, and without wishing to be bound by theory, single-nucleus assay experiments provide improvements over single-cell assay experiments because single-nucleus assay experiments are more amendable to scaling and / or performing within short timeframes, and obtaining samples can be more facile than obtaining fresh samples. In embodiments, fresh samples are samples that were harvested and / or obtained from a donor less than about 1 week, about 3 days, about 2 days, about 1 day, orabout 12 hours before analysis, and not subjected to any freezing procedures prior to analysis) samples. In embodiments, the single-nucleus data is single-nucleus ribonucleic acid (RNA) sequencing (snRNA-seq) data.
[0139] In some embodiments, each liver sample is approximately 3cm long, and is obtained through a biopsy with a 16-18 gauge needle using 11 portal tracks. See Pandey et al., “Liver Biopsy,” [Updated 2023 Jul 24], In: StatPearls [Internet], Treasure Island (FL): StatPearls Publishing; 2023 Jan-. Available on the Internet at ncbi.nlm.nih.gov / books / NBK470567 / . In some embodiments the needle is a cutting needled. In some embodiments the needle is an aspiration needle. In some embodiments the needle is an aspiration needle and suction with a syringe is applied to obtain liver core tissue. In some embodiments the liver biopsy is a percutaneous liver biopsy. In some embodiments the percutaneous liver biopsy is performed by a palpation / percussion method, an imaging-guided method, or a real-time image-guided method. In some embodiments the liver biopsy is performed by a transvenous method. In some embodiments the liver biopsy is performed by a laparoscopic method. In some embodiments the liver biopsy is performed by a plugged biopsy method. See Pandey et al., Id.
[0140] In some embodiments single-nucleus sequence or single cell sequencing is used to obtain a plurality of sequence reads from each nucleus or cell in a plurality of nuclei or cells in the liver sample. See Jiang et al., February 23, 2023, “Isolated nuclei from frozen tissue are the superior source for single cell RNA-seq compared with whole cells,” available on the Internet at doi.org / 10.1101 / 2023.02.19.529150, which is hereby incorporated by reference.
[0141] In some embodiments, first information comprises a respective plurality of sequence reads for each respective nuclei or cell in the first plurality of nuclei or cells. In some embodiments each respective plurality of sequence reads comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 6 million, at least 7 million, at least 8 million, at least 9 million, or more sequence reads. In some embodiments, each respective plurality of sequence reads comprises at least 1 x 107, at least 2 x 107, at least 3 x 107, at least 4 x 107, at least 5 x 107, at least 6 x 107, at least 7 x 107, at least 8 x 107, at least 9 x 107, at least 1 x 108, at least 2 x 108, at least 3 x 108, at least 4 x 108, at least 5 x 108, at least 6 x 108, at least 7 x 108, at least 8 x 108, at least 9 x 108, at least 1 x 109, or more sequence reads. In some embodiments, each respective plurality of sequence reads consists of no more than 5 x 107, no more than 1 x 107, no more than 5 x 106, no morethan 4 x 106, no more than 3 x 106, no more than 2 x 106, no more than 1 x 106, no more than 500,000, no more than 100,000, no more than 50,000, no more than 30,000, no more than 20,000, no more than 10,000, no more than 9000, no more than 8000, no more than 7000, no more than 6000, no more than 5000, no more than 4000, no more than 3000, no more than 2000, no more than 1000, or less sequence reads.
[0142] In some embodiments, each respective plurality of sequence reads consists of between 1000 to 5000, from 1000 to 10,000, from 2000 to 20,000, from 5000 to 50,000, from 10,000 to 100,000, from 100,000 to 500,000 from 10,000 to 500,000, from 500,000 to 1 million, from 1 million to 30 million, from 30 million to 80 million, or from 10 million to 500 million sequence reads. In some embodiments, the respective plurality of sequence reads falls within another range starting no lower than 1000 sequence reads and ending no higher than 1 x 109sequence reads.
[0143] In some embodiments, the Cell Ranger Single Cell Software Suite (10X Genomics, Inc.) is used for processing each respective plurality of sequence reads. In some such embodiments this includes library demultiplexing, fastq file generation, read alignment and unique molecular identification quantification against a lOx Genomics, Inc. pre-built human genome. This produces a unique molecular identifier (UMI) count for each gene for each nuclei or cell that is used in such embodiments as a basis for determining gene expression in each nuclei or cell. In some embodiments these UMI counts are log normalized.
[0144] In some embodiments the sequence reads are obtained by a whole-genome sequencing. In some embodiments the sequence reads are obtained by targeted DNA sequencing using a plurality of nucleic acid probes. In some embodiments, the plurality of sequence reads is determined by scTag-seq. In some embodiments, the plurality of sequence reads is determined by single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.
[0145] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells are preprocessed. In some embodiments, the preprocessing includes one or more of filtering, normalization, mapping (e.g., to a reference sequence), quantification, scaling, deconvolution, cleaning, dimension reduction, transformation, statistical analysis, and / or aggregation.
[0146] For example, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is filtered based on adesired quality, e.g., size and / or quality of a nucleic acid sequence, or a minimum and / or maximum abundance value for a respective gene in a particular nucleus or cell. In some embodiments, filtering is performed in part or in its entirety by various software tools, such as Skewer. See, Jiang, H. etal., BMC Bioinformatics 15(182): 1-12 (2014). In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is filtered for quality control, for example, using a sequencing data QC software such as AfterQC, Kraken, RNA-SeQC, FastQC, or another similar software program. In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is normalized, e.g., to account for pull-down, amplification, and / or sequencing bias (e.g., mappability, GC bias etc.). See, for example, Schwartz et al., PLoS ONE 6(l):el6685 (2011) and Benjamini and Speed, Nucleic Acids Research 40(10):e72 (2012), the contents of which are hereby incorporated by reference, in their entireties, for all purposes. For instance, in some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells is scaled to have a mean value of zero and a standard of deviation of 1. In some embodiments, the preprocessing the single-nucleus or single-cell transcriptome data for the plurality of genes in the first plurality of nuclei or cells improves (e.g., lowers) a high signal -to-noise ratio.
[0147] Thus, in some embodiments, the single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells comprises any one of a variety of forms, including, without limitation, raw abundance values, absolute abundance values (e.g., transcript number), relative abundance values (e.g., relative fluorescent units, transcriptome analysis, and / or gene set expression analysis (GSEA)) for each gene in each nucleus or cell, compound or aggregated abundance value for each gene in each nucleus or cell, transformed abundance value (e.g., Iog2 and / or log transformed) for each gene in each nucleus or cell, a change (e.g., fold- or log-change) relative to a reference (e.g., a normal sample, matched sample, reference dataset, housekeeping gene, and / or reference standard) for each gene in each nucleus or cell, a standardized abundance value for each gene in each nucleus or cell, a measure of central tendency (e.g., mean, median, mode, weighted mean, weighted median, and / or weighted mode) for each gene in each nucleus or cell, a measure of dispersion (e.g., variance, standard deviation, and / or standard error) for each gene in each nucleus or cell, or an adjusted abundance value (e.g., normalized, scaled, and / or error- corrected) for each gene in each nucleus or cell.
[0148] Any one of a number of counting techniques may be used to obtain the abundance value for each gene in each nucleus or cell in the first plurality of nuclei or cells.
[0149] In some embodiments, the single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells is determined using one or more methods including microarray analysis via fluorescence, chemiluminescence, electric signal detection, polymerase chain reaction (PCR), reverse transcriptase polymerase chain reaction (RT-PCR), digital droplet PCR (ddPCR), solid-state nanopore detection, RNA switch activation, a Northern blot, and / or a serial analysis of gene expression (SAGE).
[0150] In some embodiments, single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by a colorimetric measurement, a fluorescence measurement, a luminescence measurement, or a resonance energy transfer (FRET) measurement.
[0151] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is gene expression data. In some embodiments, gene expression in a respective nucleus or cell in the first plurality of nuclei or cells can be measured by sequencing the RNA transcripts of the nuclei or cells and then counting the quantity of each gene transcript mapping to a gene in the first plurality of genes, identified during the sequencing. In some embodiments, the gene transcripts sequenced and quantified include RNA, such as mRNA. In some embodiments, the gene transcripts sequenced and quantified include a downstream product of mRNA, such as a protein (e.g., a transcription factor). In general, as used herein, the term “gene transcript” may be used to denote any downstream product of gene transcription or translation, including post-translational modification, and “gene expression” and “cellular constituent abundance” may interchangeably be used to refer generally to any measure of gene transcripts.
[0152] In some embodiments, the abundance of a respective gene in the first plurality of cellular constituents is RNA abundance (e.g., gene expression), and the abundance of the respective gene in a nucleus or cell in the plurality of nuclei or cells is determined by measuring polynucleotide levels of one or more nucleic acid molecules corresponding to the respective gene in a nucleus or cell. The transcript levels of the respective gene can be determined from the amount of mRNA, or polynucleotides derived therefrom, present in the nucleus or cell. Polynucleotides can be detected and quantitated by a variety of methods including, but not limited to, microarray analysis, polymerase chain reaction (PCR), reversetranscriptase polymerase chain reaction (RT-PCR), Northern blot, serial analysis of gene expression (SAGE), RNA switches, RNA fingerprinting, ligase chain reaction, Qbeta replicase, isothermal amplification method, strand displacement amplification, transcription based amplification systems, nuclease protection assays (Si nuclease or RNAse protection assays), and / or solid-state nanopore detection. See, e.g., Draghi ci, Data Analysis Tools for DNA Microarrays, Chapman and Hall / CRC, 2003; Simon et al., Design and Analysis of DNA Microarray Investigations, Springer, 2004; Real-Time PCR: Current Technology and Applications, Logan, Edwards, and Saunders eds., Caister Academic Press, 2009; Bustin A-Z of Quantitative PCR (IUL Biotechnology, No. 5), International University Line, 2004; Velculescu etal., (1995) Science 270: 484-487; Matsumura et al, (2005) Cell. Microbiol. 7: 11-18; Serial Analysis of Gene Expression (SAGE): Methods and Protocols (Methods in Molecular Biology), Humana Press, 2008; each of which is hereby incorporated herein by reference in its entirety.
[0153] In some embodiments, single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained from expressed RNA or a nucleic acid derived therefrom (e.g., cDNA or amplified RNA derived from cDNA that incorporates an RNA polymerase promoter) from the first plurality of nuclei or cells, including naturally occurring nucleic acid molecules, as well as synthetic nucleic acid molecules. Thus, in some embodiments, the abundance of each gene in the first plurality of genes of the single-nucleus or single-cell transcriptome data for each nucleus or cell in the first plurality of nuclei or cells is obtained from such non-limiting sources as total cellular RNA, poly(A)+ messenger RNA (mRNA) or a fraction thereof, cytoplasmic mRNA, or RNA transcribed from cDNA (e.g., cRNA). Methods for preparing total and poly(A)+ RNA are well known in the art, and are described generally, e.g., in Sambrook, et al., Molecular Cloning: A Laboratory Manual (3rd Edition, 2001). RNA can be extracted from a cell or nucleus of interest using guanidinium thiocyanate lysis followed by CsCl centrifugation (see, e.g., Chirgwin et al., 1979, Biochemistry 18:5294-5299), a silica gelbased column (e.g., RNeasy (Qiagen, Valencia, Calif.) or StrataPrep (Stratagene, La Jolla, Calif.)), or using phenol and chloroform, as described in Ausubel et al., eds., 1989, Current Protocols In Molecular Biology, Vol. Ill, Green Publishing Associates, Inc., John Wiley & Sons, Inc., New York, at pp. 13.12.1-13.12.5). Poly(A)+ RNA can be selected, e.g., by selection with oligo-dT cellulose or, alternatively, by oligo-dT primed reverse transcription of total cellular RNA. RNA can be fragmented by methods known in the art, e.g., by incubation with ZnCh, to generate fragments of RNA.
[0154] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by sequencing transcripts from the plurality of cells. In some embodiments, the singlenucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is determined by single-cell ribonucleic acid (RNA) sequencing (scRNA-seq), scTag-seq, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq), CyTOF / SCoP, E-MS / Abseq, miRNA-seq, CITE-seq, or any combination thereof.
[0155] In some embodiments, scRNA-seq, scTag-seq, and miRNA-seq are used to measure RNA expression in order to obtain the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells. scRNA-seq measures expression of RNA transcripts, scTag-seq allows detection of rare mRNA species, and miRNA-seq measures expression of micro-RNAs. CyTOF / SCoP and E-MS / Abseq can be used to measure protein expression in the cell. CITE-seq simultaneously measures both gene expression and protein expression in the cell, and scATAC-seq measures chromatin conformation in the cell. Table 1 below provides example protocols for performing each of the measurement techniques described above. In some embodiments, any of the protocols described in Table 1 of Shen et al.,“Recent advances in high-throughput single-cell transcriptomics and spatial transcriptomics,” Lab Chip 22, p. 4774, is used to measure the abundance of cellular constituents, such as genes, in order to obtain the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells.
[0156] Table 1 - Example Measurement Protocols
[0157] Referring to block 204, in some embodiments, the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes. In some embodiments, the plurality of genes consists of between 50 and 15,000 genes. In some embodiments, the plurality of genes consists of between 100 and 15,000 genes. In some embodiments, the plurality of genes consists of between 200 and 12,000 genes. In some embodiments, the plurality of genes consists of between 250 and 10,000 genes. In some embodiments, the plurality of genes consists of between 300 and 5,000 genes. In embodiments, the plurality of genes comprises at least 2, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 genes. In embodiments, the plurality of genes comprises at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, or at least 1000 genes. In embodiments, the plurality of genes comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 30,000, at least 50,000, or more than 50,000 genes. In embodiments, the plurality of genes comprises between 2 and 20, between 20 and 50, between 50 and 100, between 100 and 200, between200 and 500, between 500 and 1000, between 1000 and 5000, between 5000 and 10,000 genes, or between 10,000 and 50,000 genes.
[0158] In some embodiments, the first plurality of genes is limited to those genes that are highly variable in their expression across the first plurality of nuclei or cells. In some embodiments these highly variable genes are identified using the Seurat function FindVariableFeatures. See, Stuart el al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference. In some embodiments, the most highly variable genes selected in this manner was limited to between 50 genes and 15,000 genes. In some embodiments. In some embodiments, the most highly variable genes selected in this manner was limited to between 100 genes and 15,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 200 genes and 12,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 250 genes and 10,000 genes. In some embodiments, the most highly variable genes selected in this manner was limited to between 300 genes and 5,000 genes.
[0159] In some embodiments, only genes expressed in at least 10, at least 20, at least 30, at least 40, or at least 50 nuclei or cells are kept in the first plurality of genes.
[0160] Referring to block 206, in some embodiments, the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 1000 nuclei or cells and 100,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 2000 nuclei or cells and 200,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 3000 nuclei or cells and 250,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 4000 nuclei or cells and 500,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 5000 nuclei or cells and 600,000 nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 6000 nuclei or cells and 1 x 106nuclei or cells. In some embodiments, the first plurality of nuclei or cells consists of between 7000 nuclei or cells and 2 x 106nuclei or cells.
[0161] In some embodiments, the first plurality of nuclei or cells comprises at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, atleast 300, at least 400, at least 500, at least 1000, at least at least 2000, at least 3000, at least 4000, at least 5000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 500,000, or at least 1 million nuclei or cells. In some embodiments, the first plurality of nuclei or cells comprises no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 5000, no more than 1000, no more than 500, no more than 200, no more than 100, or no more than 50 nuclei or cells. In some embodiments, the first plurality of nuclei or cells comprises from 5 to 100, from 10 to 50, from 20 to 500, from 200 to 10,000, from 1000 to 100,000, from 50,000 to 500,000, or from 10,000 to 1 million nuclei or cells. In some embodiments, the first plurality of nuclei or cells falls within another range starting no lower than 5 nuclei or cells and ending no higher than 10 million nuclei or cells.
[0162] In some embodiments, outlier nuclei or cells are removed from the first plurality of nuclei or cells using the Scater R isOutlier function. See, McCarthy et al., 2017, “Scater: preprocessing, quality control, normalization and visualization of single-cell RNA-seq data in R,” Bioinformatics, 33, 1179-1186. which is hereby incorporated by reference.
[0163] In some embodiments, only those nuclei or cells that express at least 20, at least 40, at least 50, at least 100, or at least 200 genes are kept in the first plurality of nuclei or cells.
[0164] Referring to block 208, in some embodiments, the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects. In some embodiments, the plurality of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects, 20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the plurality of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consists of between 5 subjects and 1000 subjects. In some embodiments, the plurality of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the plurality of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the plurality of subjects consists of between 20 subjects and 3000 subjects.
[0165] Referring to block 210, in some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50liver tissue samples, or 100 or more liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 3 or more liver tissue samples, 5 or more liver tissue samples, 10 or more liver tissue samples, 15 or more liver tissue samples, 20 or more liver tissue samples, 25 or more liver tissue samples, 30 or more liver tissue samples, 35 or more liver tissue samples, 40 or more liver tissue samples, 45 or more liver tissue samples, 50 or more liver tissue samples, 55 or more liver tissue samples, 60 or more liver tissue samples, 65 or more liver tissue samples, 70 or more liver tissue samples, 75 or more liver tissue samples, 80 or more liver tissue samples, 85 or more liver tissue samples, 90 or more liver tissue samples, 95 or more liver tissue samples, or 100 or more liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples consists of between 3 liver tissue samples and 500 liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples consists of between 5 liver tissue samples and 1000 liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples consists of between 10 liver tissue samples and 2000 liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples consists of between 15 liver tissue samples and 2500 liver tissue samples. In some embodiments, each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples consists of between 20 liver tissue samples and 3000 liver tissue samples. In some embodiments, each of the liver tissue samples was flash frozen after being biopsied or removed from the corresponding subject.
[0166] Referring to block 212, the first information obtained further comprises first metadata for each respective subject in the cohort of subjects. The first information indicates at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state. At least a first subset of subjects in the cohort of subjects have the first hepatic state and a second subset of subjects in the cohort of subjects have the second hepatic state.
[0167] Referring to block 214, in some embodiments, the first hepatic state or the second hepatic state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-AlcoholicSteatohepatitis (NASH), also known as metabolic dysfunction-associated steatohepatitis (MASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension. In some embodiments, the first hepatic state or the second hepatic state is Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis (NASH), also known as metabolic dysfunction-associated steatohepatitis (MASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, prediabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
[0168] In some embodiments, the first hepatic state or the second hepatic state is associated with hepatitis, hepatotoxicityjaundice, hepatic fibrosis, cirrhosis, hepatomegaly, alcoholic or non-alcoholic liver fatty liver disease and / or choluria.
[0169] In some embodiments, the first hepatic state or the second hepatic state is associated with a form of drug induced liver disease. Drug induced liver disease can be classified into three injury patterns: hepatocellular, cholestatic, and mixed hepatocellular-cholestatic. These designations refer to histologic features of injury, but are usually defined based upon the pattern of serum enzyme elevations. See, LiverTox: Clinical and Research Information on Drug-Induced Liver Injury [Internet], Bethesda (MD): National Institute of Diabetes and Digestive and Kidney Diseases; 2012-. Clinical Course and Diagnosis of Drug Induced Liver Disease. [Updated 2019 May 4], Available from: ncbi.nlm.nih.gov / books / NBK548733 / , which is hereby incorporated by reference.
[0170] Hepatitis, hepatotoxicity. Drug induced liver disease that resembles acute viral hepatitis is typified by a prominent hepatocellular pattern of injury. Liver biopsy, if available, usually shows marked liver cell necrosis and inflammation with only mild bile stasis, at least in the early stages. If present, symptoms of fatigue and weakness predominate. Serum alanine and aspartate aminotransferase (ALT and AST) levels typically are markedly elevated (usually > 10-fold), while the alkaline phosphatase or gamma glutamyl transpeptidase (GGT) are only modestly increased. An “R” ratio of ALT to alkaline phosphatase (both expressed as multiples of the upper limit of the normal range) of 5 or more is often used to define a hepatocellular pattern of injury. Id.
[0171] Cholestatic injury. A cholestatic picture of drug induced liver injury resembles bile duct obstruction or choledocholithiasis. The liver biopsy findings are generally of bile stasis,portal inflammation and proliferation or injury of bile ducts and ductules. Clinically, symptoms of jaundice and itching predominate. Id.
[0172] Mixed hepatocellular-cholestatic injury. A mixture of hepatocellular and cholestatic injury is typical of many drugs and, indeed, is the pattern that is most characteristic of drug induced liver injury, occurring rarely in other forms of acute liver disease. In cases of mixed injury, liver biopsy shows prominent hepatocyte necrosis and inflammation accompanied by marked bile stasis. Symptoms may include both fatigue and itching, and laboratory tests show similar elevations in serum ALT and alkaline phosphatase. An R ratio of ALT to alkaline phosphatase (both expressed as multiples of the upper limit of the normal range) between 2 and 5 is used to define a mixed pattern of injury. Drugs that cause a mixed hepatocellular-cholestatic pattern of injury include the sulfonamides, phenytoin and enalapril. Id.
[0173] In addition, to the injury patterns, drug induced liver injury can be categorized by an overall clinical pattern and course of injury into at least twelve “phenotypes”. These phenotypes overlap to some degree with the pattern of injury (hepatocellular, cholestatic, mixed) and include: acute hepatic necrosis, acute (hepatocellular) hepatitis, cholestatic hepatitis, mixed hepatitis, serum enzyme elevations without jaundice, bland cholestasis, acute fatty liver with lactic acidosis, nonalcoholic fatty liver, chronic hepatitis, sinusoidal obstruction syndrome, nodular regenerative hyperplasia, and liver tumors, such as hepatic adenoma and hepatocellular carcinoma. In addition, characteristic features or outcomes are often assigned to each phenotype such as immunoallergic features, autoimmune features, acute liver failure, vanishing bile duct syndrome and cirrhosis. Id.
[0174] In some embodiments, the liver injury is characterized by an elevation of one or more liver enzymes, cholestasis, and / or alcoholic stools. For example, in some embodiments, serum alanine and aspartate aminotransferase (ALT and AST) levels are elevated more than 10-fold, while the alkaline phosphatase or gamma glutamyl transpeptidase (GGT) are modestly increased. In some embodiments liver injury is characterized by a liver function tests, such as ALP (alkaline phosphatase), ALT (alanine transaminase), AST (aspartate aminotransferase), or gamma-glutamyl transferase (GGT). In some embodiments liver injury is characterized by a Histology (e.g., a liver biopsy). In some embodiments liver injury is characterized by transient elastography (Fibroscan). See for example Foucher et al., 2006, “Diagnosis of cirrhosis by transient elastography (FibroScan): a prospective study,” Gut 55(3), pp. 403-408, which is hereby incorporated by reference. In some embodiments liver injury is characterized by imaging such as ultrasound, computed tomography (CT) scan, ormagnetic resonance imaging (MRI). In some embodiments, liver injury is characterized by a Child-Plugh score to determine Cirrhosis (e.g., using five clinical parameters: bilirubin, albumin, INR, ascites, hepatic, encephalopathy). See, for example, Tsoris and Marlar, “Use Of The Child Pugh Score In Liver Disease,” [Updated 2023 Mar 13], In: StatPearls [Internet], Treasure Island (FL): StatPearls Publishing; 2023 Jan-. Available on the Internet at ncbi.nlm.nih.gov / books / NBK542308. In some embodiments, liver injury is characterized by viral load (e.g., viral hepatitis). See for example, Oliveira et al., “Advanced Liver Injury in Patients with Chronic Hepatitis B And Viral Load Below 2,000 lu / Ml,” Rev Inst Med Trop Sao Paulo. 2016 Sep 22;58:65. doi: 10.1590 / S1678-9946201658065. PMID: 27680170; PMCID: PMC5048636, which is hereby incorporated by reference.
[0175] In cholestasis without or with only modest liver cell injury itching and jaundice are prominent, but, otherwise, patients typically feel well. Serum alkaline phosphatase and ALT levels may be normal or minimally (less that two-fold) elevated, particularly at the peak of symptoms and jaundice. At the onset, ALT levels may be more prominently elevated (3 to 10 fold) and in more severe instances, serum alkaline phosphatase may rise above 2 fold. Pure or bland cholestasis is the characteristic form of liver injury caused by estrogenic and anabolic steroids and can occur with illicit use of muscle building regimens and more rarely with the thioguanines such as azathioprine and mercaptopurine. Id.
[0176] Referring to block 216, in some embodiments, the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects. In some embodiments, the first subset of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects, 20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the first subset of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consists of between 5 subjects and 1000 subjects. In some embodiments, the first subset of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the first subset of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the first subset of subjects consists of between 20 subjects and 3000 subjects. In some embodiments, the second subset of subjects comprises 3 or more subjects, 5 or more subjects, 10 or more subjects, 15 or more subjects,20 or more subjects, 25 or more subjects, 30 or more subjects, 35 or more subjects, 40 or more subjects, 45 or more subjects, 50 or more subjects, 55 or more subjects, 60 or more subjects, 65 or more subjects, 70 or more subjects, 75 or more subjects, 80 or more subjects, 85 or more subjects, 90 or more subjects, 95 or more subjects, or 100 or more subjects. In some embodiments, the second subset of subjects consists of between 3 subjects and 500 subjects. In some embodiments, the plurality of subjects consists of between 5 subjects and 1000 subjects. In some embodiments, the second subset of subjects consists of between 10 subjects and 2000 subjects. In some embodiments, the second subset of subjects consists of between 15 subjects and 2500 subjects. In some embodiments, the second subset of subjects consists of between 20 subjects and 3000 subjects.
[0177] Referring to block 218, in some embodiments, the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition. In embodiments, the physical condition comprises one or more features selected from sex, body mass index (BMI), age, ethnicity, cause of death, serology (CMV, EBV status), final lab profile, alcohol consumption, tobacco use, illicit drug use, medical history, family medical history, disease diagnostic information, medication profile, etc.
[0178] In some embodiments, the first metadata is any combination of sex, BMI, age, ethnicity, cause of death, serology (CMV, EBV status), final lab profile, alcohol use, tobacco use, illicit drug use, fibrosis score, NAS score, medical history, and / or medications.
[0179] In embodiments, the disease diagnostic information comprises one or more features selected from pathology notes and key features.
[0180] In embodiments, the disease diagnostic information comprises one or more features selected from NASH activity score (NAS) with steatosis, NAS score with inflammation, NAS score with ballooning, fibrosis score F0-F4, cirrhosis status, and pathology notes and key features.
[0181] In embodiments, the disease diagnostic information comprises one or more features selected from NASH activity score (NAS) with steatosis, NAS score with inflammation, NAS score with ballooning, fibrosis score F0-F4, cirrhosis status, aspartate aminotransferase-to- platelet ratio index (APRI), FIB-4 score, NAFLD fibrosis score (NFS), hepascore, FibroTest result, FibroMeter result, enhanced liver fibrosis (ELF) score, shear-wave elastography (SWE) measurement, vibration-controlled transient elastography (VCTE) results, acousticradiation force impulse (ARFI) elastography, magnetic resonance elastography (MRE), and pathology notes and key features.
[0182] In embodiments, the disease diagnostic information comprises disease status. In embodiments, the disease status comprises histologically graded disease status. In some embodiments, disease status is measured using any method described in Kleiner etal., 2005, “Nonalcoholic Steatohepatitis Clinical Research Network. Design and validation of a histological scoring system for nonalcoholic fatty liver disease,” Hepatology Jun;41(6):1313- 21, incorporated herein by reference in its entirety.
[0183] Referring to block 220, in some embodiments, the first metadata for each respective subject in the cohort of subjects indicting at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state comprises a histologically graded disease status for the respective subject. In some such embodiments, the disease status is measured using any method described in Kleiner et al., 2005, “Nonalcoholic Steatohepatitis Clinical Research Network. Design and validation of a histological scoring system for nonalcoholic fatty liver disease,” Hepatology, Jun;41(6): 1313- 21, incorporated herein by reference in its entirety.
[0184] In embodiments, the disease status comprises a graded nonalcoholic steatohepatitis (NASH) status. NASH is a type of liver disease characterized by inflammation and liver cell damage associated with the accumulation of fat in the liver. In some embodiments the histological grading and staging of NASH is assessed using the (i) Nonalcoholic Fatty Liver Disease Activity (NAS) score and (ii) fibrosis staging. These scores provide information about the severity of the disease and the degree of fibrosis (scarring) in the liver. See, Sheka et al., 2020, “Nonalcoholic Steatohepatitis: A Review,” JAMA 232(12), pp. 1175-1183, which is hereby incorporated by reference.
[0185] In some embodiments, the NAS score is used to assess the activity or inflammation in NASH. In some embodiments is based on the evaluation of three histological features in a liver biopsy: steatosis, lobular inflammation, and hepatocellular ballooning. Steatosis is the degree of fat accumulation in liver cells and is graded on a scale from 0 to 3 in some embodiments, with 0 meaning no steatosis and 3 indicating severe steatosis. Lobular inflammation is the level of inflammation in the liver lobules and is also graded from 0 to 3 in some embodiments, with 0 indicating no inflammation and 3 indicating severe inflammation. Hepatocellular ballooning is the presence and severity of ballooned hepatocytes (swollen liver cells) are graded from 0 to 2 in some embodiments, with 0 being none and 2 indicatingmany ballooned cells. In some embodiments the NAS score is the sum of these individual scores, (e.g., in some embodiments) resulting in a total score ranging from 0 to 8. A higher NAS score indicates more severe inflammation and activity in the liver. See Juluri et al., 2011, “Generalizability of the NASH CRN Histological Scoring System for Nonalcoholic Fatty Liver Disease,” J. Clin. Gastroenterol 45(1), pp. 55-58, which is hereby incorporated by reference.
[0186] In addition to the NAS score, the degree of fibrosis or scarring in the liver is assessed. In some embodiments this is done using a fibrosis staging system that evaluates the extent of fibrosis on a scale from F0 to F4, where F0 indicates no fibrosis, Fl indicates mild fibrosis, F2 indicates moderate fibrosis, F3 indicates severe fibrosis, and F4 indicates cirrhosis (advanced scarring).
[0187] In some embodiments the combination of NAS score and fibrosis stage provides a comprehensive assessment of the severity of NASH and is included in the first metadata for each subject in the cohort of subjects. Subjects with higher NAS scores and more advanced fibrosis are at greater risk of liver-related complications and may require closer monitoring and more aggressive treatment strategies in some embodiments.
[0188] Referring to block 222, in some embodiments, the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each liver tissue in the plurality of liver tissue samples.
[0189] Referring to block 224, in some embodiments, the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers (e.g., genes, proteins, etc.) that each define a unique cell type. Examples of such first biomarker annotations are given below in block 256 in conjunction with Figures 4 and 5 as well as Table 2 below.
[0190] Referring to block 226, in some embodiments, sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects. For instance, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in batches due to possible throughput limitations in obtaining such data. In such embodiments, each batch of the single-nucleus or single-cell transcriptome data is single-nucleus transcriptome orsingle-cell data for the plurality of genes for each nucleus or cell in a subset of the plurality of nuclei or cells.
[0191] In some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, between 21 and 100, or between 100 and 1000 batches where each respective batch is single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a corresponding subset of the plurality of nuclei or cells from a corresponding subset of the subjects in the cohort of subjects. In some embodiments, each respective batch is examined relative to all other batches to ensure that the subjects represented in the respective batch are balanced for sex, race, age, and / or physical condition. For instance, in some such embodiments, this is done using one-way analysis of variance (ANOVA). That is, the ANOVA is used to determine whether there are any statistically significant differences between the means in sex, race, age, and / or physical condition (e.g., BMI) of the plurality of batches. The one-way ANOVA compares the means between the k batches and determines whether any of those means are statistically significantly different from each other by testing the null hypothesis:HQ - Pi=P2=P3= =Pk where p is the respective group mean for sex, race, age, and / or physical condition and k is the number of batches. If, however, the one-way ANOVA returns a statistically significant result, the alternative hypothesis HA), which is that there are at least two batches whose means are statistically significantly different from each other. In some embodiments, this is calculated using the F-statistic: n(variance of batch means) (mean of batch variances)In this equation, n is number of subjects represented in each of the batches. Each respective batch mean is the mean value for the property under study (sex, race, age, and / or physical condition) within the respective batch. Each respective batch variance is the variance in the value for the property under study (sex, race, age, and / or physical condition) within the respective batch. This F value can then be used to compute a p-value for the null hypothesis. See, Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 16, which is hereby incorporated by reference. In some embodiments, a p-value of less than 0.05 is used to accept the null hypothesis. That is, in some embodiments, ANOVA of sex, race, age, or physical condition across the batches yields a p-value of 0.05 or less. In some suchembodiments, a p-value of less than 0.05 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0192] In some embodiments, ANOVA of sex, race, age, and / or physical condition across the batches yields a p-value of 0.10 or less. In some such embodiments, a p-value of less than 0.10 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0193] In some embodiments, ANOVA of sex, race, age, and / or physical condition across the batches yields a p-value of 0.01 or less. In some such embodiments, a p-value of less than 0.01 is required to accept the null hypothesis (that the batches are all balanced for the sex, race, age, and / or physical condition).
[0194] In some embodiments, an alternative statistical test is used to ensure sex, race, age, and / or physical condition is balanced across the batches. For instance, in some embodiments a t-test is used. See Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 9, which is hereby incorporated by reference. In some embodiments, a calculation of a p-value of 0.15 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.10 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.05 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced. In some embodiments, a calculation of a p-value of 0.01 or less for the null hypothesis that each batch is balanced for sex, race, age, and / or physical condition using the alternative statistical test is used to ensure that the batches are balanced.
[0195] In some embodiments a nonparametric test such as a sign test for median, Wilcoxon- Mann-Whitney Rank Sum test, Rank Correlation Test, or Runs test is used to ensure that the batches are balanced for sex, race, age, and / or physical condition. See Smith, 1991, Statistical Reasoning, Third Edition, Allyn and Bacon, Boston, Chapter 17, which is hereby incorporated by reference.
[0196] In some embodiments, each respective batch is examined relative to the singlenucleus or single-cell transcriptome data for each nucleus or cell in the first plurality of nuclei or cells for the entire cohort of subjects to ensure that the subjects represented in therespective batch are balanced for sex, race, age, and / or physical condition relative to the entire cohort.
[0197] In some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is obtained in a single batch and no balancing is performed for sex, race, age, and / or physical condition.
[0198] In some embodiments, sex is balanced additionally or alternatively in the cohort of subjects by ensuring that between 40 percent and 60 percent of the subjects are female. In some embodiments, sex is balanced additionally or alternatively in the cohort of subjects by ensuring that between 45 percent and 55 percent of the subjects are female. In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that between 47.5 percent and 52.5 percent of the subjects are female. In some embodiments, the cohort of subjects is not balanced for sex.
[0199] In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that the participation to prevalence ratio (PPR) for woman (across the entire cohort) is between 0.80 and 1.20, where PPR is defined as:Percentage of woman in the cohort of subjects Percentage of woman among disease population ’ (e. g. , subjects with first hepatic state)
[0200] In some embodiments, sex is balanced in the cohort of subjects additionally or alternatively by ensuring that the PPR for woman (across the entire cohort) is between 0.90 and 1.10.
[0201] In some embodiments, race is represented in a balanced manner additionally or alternatively by ensuring that the respective PPR of white subjects, black subjects, and Asian subjects (across the entire cohort) is between 0.80 and 1.20. The PPR of a given race in such embodiments is calculated as:Percentage of subjects in the cohort of subjects that are of race X Percentage of subjects among disease population (e. g. , have the first hepatic state)' that are of race X
[0202] In some embodiments, race is represented in a balanced manner additionally or alternatively in the cohort by ensuring that the respective PPR of white subjects, black subjects, and Asian subjects (across the entire cohort) is between 0.90 and 1.10.
[0203] In some embodiments, age is represented in a balanced manner additionally or alternatively by ensuring that the respective PPR of particular age groups (across the entirecohort) is between 0.80 and 1.20. The PPR of a given age group in such embodiments is calculated as:Percentage of subjects in the cohort in age group XPercentage of subjects among disease population (e. g. , have the first hepatic state)' in age group X
[0204] In some embodiments, age is represented in a balanced manner additionally or alternative by ensuring that the respective PPR of each respective age group in a particular set of age groups (across the entire cohort) is between 0.90 and 1.10. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being over 65 years in age. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being aged 45 to 64 years. In some embodiments one of the age groups that is balanced in the cohort is the age group defined as being aged 18-44 years.
[0205] In some embodiments, physical condition is balanced additionally or alternatively in the cohort of subjects by ensuring that the participation to prevalence ratio (PPR) for subjects with the physical condition is between 0.80 and 1.20, where PPR is defined as:Percentage of subjects in the cohort having the physical condition Percentage of subject among disease population (e. ^., have the first hepatic state)' with the physical condition
[0206] In some embodiments, physical condition is balanced in the cohort of subjects additionally or alternatively by ensuring that the participation to prevalence ratio (PPR) for subjects with the physical condition (across the entire cohort) is between 0.90 and 1.10. In some embodiments the physical condition is body mass index. In some embodiments the physical condition is smoking status.
[0207] For further discussion of the use of PPR to analyze whether a cohort is balanced, see Varma et cd.. 2021, “Reporting of Study Participant Demographic Characteristics and Demographic Representation in Premarketing and Postmarketing Studies of Novel Cancer Therapeutics,” JAMA Netw Open Apr; 4(4): e217063, which is hereby incorporated by reference.
[0208] Referring to block 228, in some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is filtered to remove counts of ambient RNA molecules, doublets and / or empty droplets. Ambient RNA is the pool of mRNA molecules that have been released in a cell suspension, likely from cells that are stressed or have undergone apoptosis during singlenucleus or single-cell sequencing. Cross-contamination occurs when the ambient RNA getsincorporated into droplets representing single nuclei or single cells, which did not originate the ambient RNA, and is barcoded and amplified along with the nucleus’ or cell’s native mRNA. Liver biopsies often comprise 90% hepatocytes and the ambient RNA from the hepatocytes tends to drown out other less prevalent cell types in the biopsies through such ambient RNA contamination. Accordingly, in some embodiments the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells are filtered to remove counts of ambient RNA molecules. Contamination from ambient RNA is evident when highly expressed cell type-specific genes are observed at low levels in other cell populations. Different proportions of contamination can be found in different droplets depending on the amount of ambient and native mRNA present.
[0209] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules by assuming that the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells in fact represents a mixture of counts from two multinomial distributions: (1) a distribution of native transcript counts from the nuclei’s or cell’s actual population and (2) a distribution of contaminating transcript counts from all other nuclei or cell populations captured in the assay. In some such embodiments, a program such as DecontX, Yang et al., 2020, “Decontamination of ambient RNA in single-cell RNA-seq with DecontX”, Genome Biology 21 :57, is used to deconvolute a gene-by-nuclei or gene-by-cell count matrix and a vector of nuclei or cell population labels into a matrix of contamination counts and a matrix of native counts that can be used in downstream analyses.
[0210] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using Cellbender. See Fleming et al., 2019, “CellBender remove-background: a deep generative model for unsupervised removal of background noise from scRNA-seq datasets,” bioRxiv 791699, doi: 10.1101 / 791699, which is hereby incorporated by reference.
[0211] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using SoupX. See, Young and Behjati, 2020, “SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data,” Gigascience 9, doi: 10.1093 / gigascience / giaal51, which is hereby incorporated by reference.
[0212] In some embodiments, ambient RNA is filtered to remove counts of ambient RNA molecules using any known method. Additional examples of such methods are disclosed in Caglayan et al., l^ , “Ambient RNA analysis reveals misinterpreted and masked cell typesin brain single-nuclei datasets,” Neuron doi: 10.1016 / j. neuron.2022.09.010, which is hereby incorporated by reference.
[0213] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells are filtered to remove doublets. This arises when more than one cell or nucleus is captured in a droplet, also known as a “doublet” or “multiplet.” In microfluidic systems, the occurrence of doublets is proportional to the concentration of cells or nuclei in the suspension and capture rate of the device. In some embodiments Scrublet, Wolock etal., 2019, “Scrublet: computational identification of cell doublets in single-cell transcriptomic data,” Cell Syst. 8(4):281-91, or DoubletFinder, McGinnis et al., 2019, “Doubletfinder: doublet detection in single-cell rna sequencing data using artificial nearest neighbors,” Cell Syst. 8(4):329-37, is used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data. These programs simulate artificial doublets from the original data coordinates in a reduced-dimensional representation, then create doublet score for each barcode by calculating the similarity of its representation with artificial doublets. In some embodiments demuxlet, Kang et al., 2018, “Multiplexed droplet single-cell rna-sequencing using natural genetic variation,” Nat Biotechnol. 36(1):89, or scds, Bais and Kostka, 2019, “scds: computational annotation of doublets in single cell RNA sequencing data,” bioRxiv. 2019564021 available on the Internet at biorxiv.org / content / 10.1101 / 564021vl, is used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data. These programs model gene expression from the original data, then assign doublet to barcodes that have observed expression from genes that are likely to not occur simultaneously. In some embodiments any known method is used to remove doublets or multiplets in the single-nucleus or single-cell transcriptome data.
[0214] In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cell are filtered to remove empty droplets or droplets representing damaged nuclei or cells. See Lun et al., 2019, “EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data,” Genome Biol 20, 63; and Heiser et al., 2021, “Automated quality control and cell identification of droplet-based single-cell data using dropkick,” Genome Res 31, 1742-1752. In some embodiments, the single-cell or single-nucleus RNA sequencing utilizes abundantly more cell barcodes than the number of targeted cells to ensure single-cell capture per droplet. This results in a relatively small number of droplets that capture real cells or nuclei as most droplets are “empty droplets” and do not contain real cells in suchembodiments. However, ambient material created during sample processing can also contain transcripts called “ambient RNA” that can be captured by empty droplets, making them appear non-empty in data analysis. As a result, separation of empty droplets from real cells becomes a significant problem in single-cell RNA-seq analysis. Since real cells or nuclei contain more transcripts (and more unique molecular identifiers -UMIs- that denote unique reads) than ambient RNAs captured in empty droplets, it is possible to apply a cutoff based on number of UMIs to retain real cells or nuclei. However, this can be inaccurate as UMI distribution of cell barcodes is not completely discrete which makes a hard cutoff arbitrary. Moreover, certain cell types may be transcriptomically more silent than others which could lead to filtering out by a UMI-based cutoff. In some embodiments this problem of distinguishing real cells or nuclei and empty droplets is addressed by using other metrics such as expression profile and nuclear fraction.2,4, 5 In some such embodiments DropletQC, Muskovic and Powel, 2021, “DropletQC: improved identification of empty droplets and damaged cells in single-cell RNA-seq data.,” Genome Biol. 22, 329, is used to remove empty droplets or droplets representing damaged nuclei or cells in the single-nucleus or single-cell transcriptome data. In some such embodiments and known method is used to remove empty droplets or droplets representing damaged nuclei or cells in the single-nucleus or single-cell transcriptome data.
[0215] Referring to block 230, in some embodiments, the first hepatic state is fibrosis and the second hepatic state is absence of fibrosis. In some embodiments, the first hepatic state is stage Fl, F2, F3, or F4 fibrosis, where Fl indicates mild fibrosis, F2 indicates moderate fibrosis, F3 indicates severe fibrosis, and F4 indicates cirrhosis (advanced scarring). In some such embodiments the second hepatic state is stage F0, indicating no fibrosis.
[0216] In some embodiments, the first hepatic state is stage 1 (perisinusoidal or periportal), 1A (mild, zone 3, perisinusoidal), IB (moderate, zone 3, perisinusoidal), 1C (portal / periportal), 2 (perisinusoidal and portal / periportal), 3 (bridging fibrosis), or 4 (cirrhosis), while the second hepatic state is stage F0, indicating no fibrosis. See Kleiner et al., “Design and Validation of a Histological Scoring System for Nonalcoholic Fatty Liver Disease” Hepatology 41(6), pp. 1313-132, which is hereby incorporated by reference.
[0217] Referring to block 232, in some embodiments, the first hepatic state is absence of fibrosis and the second hepatic state is fibrosis. In some embodiments, the second hepatic state is stage Fl, F2, F3, or F4 fibrosis, where Fl indicates mild fibrosis, F2 indicates moderate fibrosis, F3 indicates severe fibrosis, and F4 indicates cirrhosis (advanced scarring). In some such embodiments the first hepatic state is stage F0, indicating no fibrosis.
[0218] In some embodiments, the second hepatic state is stage 1 (perisinusoidal or periportal), 1A (mild, zone 3, perisinusoidal), IB (moderate, zone 3, perisinusoidal), 1C (portal / periportal), 2 (perisinusoidal and portal / periportal), 3 (bridging fibrosis), or 4 (cirrhosis), while the first hepatic state is stage F0, indicating no fibrosis. See Kleiner et al., “Design and Validation of a Histological Scoring System for Nonalcoholic Fatty Liver Disease” Hepatology 41(6), pp. 1313-132, which is hereby incorporated by reference.
[0219] Referring to block 234, in some embodiments, the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis. In some such embodiments, the first hepatic state is stage F0, Fl, F2, F3, or F4 fibrosis, and the second hepatic state is a F0, Fl, F2, F3, or F4 fibrosis other than the first hepatic state. In some such embodiments the first hepatic state is stage 0 (no fibrosis), stage 1 (perisinusoidal or periportal), 1 A (mild, zone 3, perisinusoidal), IB (moderate, zone 3, perisinusoidal), 1C (portal / periportal), 2 (perisinusoidal and portal / periportal), 3 (bridging fibrosis), or 4 (cirrhosis), and the second hepatic state is a 0, 1, 1 A, IB, 1C, 2, 3 or 4 fibrosis other than the first hepatic state.
[0220] Referring to block 236, in some embodiments, the first hepatic state is presence of liver inflammation and the second hepatic state is absence of liver inflammation. Referring to block 240, in some embodiments, the first hepatic state is a first stage of liver inflammation and the second hepatic state is a second stage of liver inflammation. Dysregulated swelling and inflammation of the liver (liver inflammation), defined as hepatitis, is characterized by the presence of excess inflammatory cells. See, Stauffer et al., 2012, “Chronic inflammation, immune escape, and oncogenesis in the liver: A unique neighborhood for novel intersections,” Hepatology 56(4), pp. 1567-1574, which is hereby incorporated by reference.
[0221] Referring to block 242, in some embodiments, the first hepatic state is presence of liver steatosis and the second hepatic state is absence of liver steatosis. Referring to block 244, in some embodiments, the first hepatic state is absence of liver steatosis and the second hepatic state is presence of liver steatosis. As used herein, “liver steatosis” is the distinct necroinflammatory lesions and fibrosis associated with NASH. See, Kleiner et al., 2005, “Nonalcoholic Steatohepatitis Clinical Research Network. Design and validation of a histological scoring system for nonalcoholic fatty liver disease,” Hepatology Jun;41(6):1313- 21, incorporated herein by reference in its entirety. In some embodiments, a subject is considered to have liver steatosis if they have a steatosis grade of 1, 2, or 3. See, Kleiner, Id. In some embodiments, a subject is considered to not have liver steatosis if they have a steatosis grade of 0. See, Kleiner, Id.
[0222] Referring to block 246, in some embodiments, the first hepatic state is a first stage of liver steatosis and the second hepatic state is a second stage of liver steatosis. In some embodiments, the first stage of liver steatosis is grade 0, 1, 2, or 3, see, Kleiner, Id. for definitions of these grades, while the second stage of liver steatosis is grade 0, 1, 2, or 3 other than the first stage of liver steatosis.
[0223] Referring to block 248, in some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and / or sex. The second information is used to prune the cohort of subjects based on age, body mass index, and / or sex. This causes the cohort of subjects to be free of confounding for age, body mass index, and / or sex. Referring to block 250, in some embodiments, the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex. The second information is used to prune the cohort of subjects based on age, body mass index, and sex. This causes the cohort of subjects to be free of confounding for age, body mass index, and sex. The pruning causes a subset of subjects to be removed from the cohort of subjects. In some embodiments the pruning balances the cohort of subjects for age, body mass index, and / or sex as discussed above in conjunction with block 226.
[0224] In some embodiments, the pruning of block 248 or 250 causes a subset of subjects to be removed from the cohort of subjects. In some embodiments the pruning balances the cohort of subjects for age, body mass index (BMI), and / or sex as discussed above in conjunction with block 226. In some embodiments, the cohort is pruned for the following BMI categories [15,20), [20,22.5), [22.5, 25), [25, 27.5), [27.5, 30), [30, 32.5), [32.5, 35), [35, 37.5), [37.5, 40), and [40,60], where all ranges are in units of kg / m2, square bracket “[“ means included in the interval and rounded bracket “)” means not included in the interval. See Loomis, 2016, “Body Mass Index and Risk of Nonalcoholic Fatty Liver Disease: Two Electronic Health Record Prospective Studies,” J. Clin. Endocrinol Metab 101(3), pp. 945- 952, which is hereby incorporated by reference. In some embodiments, the cohort is pruned for the following BMI categories more than or equal to 35 kg / m2, between 23 and 35 kg / m2, and less than 23 kg / m2. In some embodiments, the cohort is pruned for the following BMI categories more than or equal to 35 kg / m2, between 23 and 35 kg / m2, and less than 22 kg / m2. In some embodiments, a particular BMI group is additionally or alternatively represented in a balanced manner by ensuring that the respective PPR of the particular BMI (across the entire cohort) is between 0.80 and 1.20. The PPR of a given BMI group in such embodiments is calculated as:ks Percentage of subjects in the cohort in BMI group X (e. g., between 23 and 35 ^) Percentage of subjects among disease population (e. g. , have the first hepatic state)' in BMI group XIn some embodiments, each respective BMI group is additionally or alternatively represented in a balanced manner by ensuring that the respective PPR of each respective BMI group in a particular set of BMI groups (across the entire cohort) is between 0.90 and 1.10. In some embodiments one of the BMI groups that is balanced in the cohort is the BMI group defined as being more than or equal to 35 kg / m2.
[0225] Referring to block 252, in some embodiments, the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subject originating the respective single-nucleus or single-cell transcriptome data.
[0226] In some embodiments a unique barcode for a respective single nucleus or cell is provided in the form of an oligonucleotide that comprises a nucleic acid barcode sequence that is attached to the nucleic acid molecules originating from the respective single nucleus or cell. The oligonucleotide is partitioned such that as between the nucleic acid molecules originating from the respective single nucleus or cell, the nucleic acid barcode sequence is the same, but as between nucleic acids originating from other nuclei or cells, the barcode can, and preferably have differing sequences. In preferred embodiments, only one nucleic acid barcode sequence is associated with a given nuclei or cell, although in some embodiments, two or more different barcode sequences are associated with a given nucleus or cell.
[0227] The nucleic acid barcode sequences will typically include from 6 to about 20 or more nucleotides within the sequence of the barcode. In some embodiments, these nucleotides are completely contiguous, e.g., in a single stretch of adjacent nucleotides. In alternative embodiments, they are separated into two or more separate subsequences that are separated by one or more nucleotides. Typically, separated subsequences are separated by about 4 to about 16 intervening nucleotides.
[0228] In some embodiments, each barcode for each nucleus or cell encodes a unique predetermined value selected from the set { 1, ..., 1024}, { 1, ..., 4096}, { 1, ..., 16384}, { 1, ..., 65536}, { 1, ..., 262144}, { 1, ..., 1048576}, { 1, ..., 4194304}, { 1, ..., 16777216}, { 1, ..., 67108864}, or { 1, ..., 1 x 1012}. For instance, consider the case in which the barcode sequence is represented by a set of five nucleotide positions. In this instance, each nucleotide position contributes four possibilities (A, T, C or G), giving rise, when all five positions areconsidered, to 4 x 4 x 4 x 4 x 4 = 1024 possibilities. As such, the five nucleotide positions form the basis of the set { 1,..., 1024}. In other words, when the barcode sequence is a 5-mer, the barcode encodes a unique predetermined value selected from the set { 1,..., 1024}. Likewise, when the barcode sequence is represented by a set of six nucleotide positions, the six nucleotide positions collectively contribute 4 x 4 x 4 x 4 x 4 x 4 = 4096 possibilities. As such, the six nucleotide positions form the basis of the set { 1, ... , 4096} . In other words, when the barcode sequence is a 6-mer, the barcode encodes a unique predetermined value selected from the set { 1,..., 4096}.
[0229] By contrast, in some embodiments, the barcode of a sequence read in the plurality of sequence reads is localized to a noncontiguous set of oligonucleotides within the sequence read. In one such exemplary embodiment, the predetermined noncontiguous set of nucleotides collectively consists of N nucleotides, where N is an integer in the set {4, . . ., 20}. As an example, in some embodiments, a barcode sequence comprises a first set of contiguous nucleotide positions at a first position in an oligonucleotide tag and a second set of contiguous nucleotide positions at a second position in an oligonucleotide tag, that is displaced from the first set of contiguous nucleotide positions by a spacer. In one specific example, the barcode sequence comprises (Xl)nYz(X2)m, where XI is n contiguous nucleotide positions, Y is a constant predetermined set of z contiguous nucleotide positions, and X2 is m contiguous nucleotide positions. In this example, the barcode in the second portion of a sequence read produced by a schema invoking this exemplary barcode is localized to a noncontiguous set of oligonucleotides, namely (Xl)nand (X2)m. This is just one of many examples of noncontiguous formats for a barcode.
[0230] Further discussion of the use of barcodes in single-nucleus RNA-seq is discussed in Jiang et al., February 23, 2023, “Isolated nuclei from frozen tissue are the superior source for single cell RNA-seq compared with whole cells,” available on the Internet at doi.org / 10.1101 / 2023.02.19.529150, which is hereby incorporated by reference. In some embodiments a 10X Genomics Chromium 3’ gene expression assay is used, and the barcoding is performed as part of this assay. See Jiang et al., Id, which is hereby incorporated by reference.
[0231] Referring to block 254, in some embodiments, the first plurality of nuclei or cells is clustered into a plurality of clusters by (i) computing a plurality of distances using the singlenucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function.
[0232] Referring to block 256, in some embodiments, the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells. Each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei or cells and (ii) a respective second vector formed by the single-nucleus transcriptome data for the plurality of genes for a respective second nucleus or cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function.
[0233] Clustering is described at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter “Duda 1973”) which is hereby incorporated by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a dataset. To identify natural groupings, two issues are addressed. First, a way to measure similarity (or dissimilarity) between two nuclei or cells is determined. This metric (similarity measure) is used to ensure that the nuclei or cells in one cluster are more like one another than they are to nuclei or cells in other clusters based on their gene expression. Second, a mechanism for partitioning the nuclei or cells into clusters using the similarity measure is determined.
[0234] Similarity measures are discussed in Section 6.7 of Duda 1973, where it is stated that one way to begin a clustering investigation is to define a distance function and to compute the matrix of distances between all pairs of nuclei or cells in a training set. If distance is a good measure of similarity, then the distance between reference nuclei or cells in the same cluster will be significantly less than the distance between the reference nuclei or cells in different clusters. However, as stated on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a nonmetric similarity function s(x, x') can be used to compare two vectors x and x'. Conventionally, s(x, x') is a symmetric function whose value is large when x and x' are somehow “similar.” An example of a nonmetric similarity function s(x, x') is provided on page 218 of Duda 1973.
[0235] Once a method for measuring “similarity” or “dissimilarity” between vectors in a dataset has been selected, clustering requires a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterionfunction are used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in Section 6.8 of Duda 1973.
[0236] More recently, Duda et al., Pattern Classification, 2nd edition, John Wiley & Sons, Inc. New York, has been published. Pages 537-563 describe clustering in detail. More information on clustering techniques can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, N.Y.; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, N.Y.; and Backer, 1995, Computer- Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is hereby incorporated by reference. Particular exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest-neighbor algorithm, farthest- neighbor algorithm, the average linkage algorithm, the centroid algorithm, or the sum-of- squares algorithm), k-means clustering, fuzzy k-means clustering algorithm, Jarvis-Patrick clustering, Louvain clustering, or Leiden clustering. See, Blondel et al., July 25, 2008, “Fast unfolding of communities in large networks,” arXiv:0803.0476v2 [physical. coc-ph]; and Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572, each of which is hereby incorporated by reference. Such clustering can be on the vectors formed directly from the single-nucleus or single-cell transcriptome data for the plurality of genes or vectors of principal components derived from such single-nuclei or single-cell transcriptome data. In some embodiments, the clustering comprises unsupervised clustering where no preconceived notion of what clusters should form when the training set is clustered are imposed.
[0237] In some embodiments before each respective first and second vector is generated, the single-nucleus or single-cell transcriptome data is subjected to principal component transformation to produce principal components. In such embodiments, the respective first vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei or cells is than the set of principal components derived for the respective first nucleus or cell. Likewise, the respective second vector formed by the single-nucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus or cell in the different pair of nuclei or cells is the set of principal components derived for the respective second nucleus or cell. Principal component analysis (PCA) algorithms that may be used to transform vectors of gene expression data to vectors of principal components are described in Jolliffe, 1986, Principal Component Analysis, Springer, New York, which is hereby incorporated byreference. PCA is also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, which is hereby incorporated by reference. Principal components (PCs) are uncorrelated and are ordered such that the kth PC has the kthlargest variance among PCs. The kthPC can be interpreted as the direction that maximizes the variation of the projections of the data points such that it is orthogonal to the first k-1 PCs. The first few PCs capture most of the variation in dataset. In contrast, the last few PCs are often assumed to capture only the residual 'noise' in the dataset. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between four and five hundred principal components. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between five and six hundred principal components. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains between six and second hundred principal components. In some embodiments, each respective vector that is clustered in accordance with block 254 or 256 corresponds to a nucleus or cell in the plurality of nuclei or cells and contains at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 principal components. In some embodiments, the principal component analysis is performed on the first plurality of genes using the Seurat function RunPCA. In some embodiments, clusters are identified with Seurat function FindClusters, optionally using the shared nearest neighbor modular optimization based on the first 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or between 20 and 500 principal components identified by principal component analysis. In some such embodiments the clustering resolution is set to a value between 0.2 and 0.6, such as 0.4. See, Stuart et al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference.
[0238] In some embodiments the clustering comprises k-means clustering of the nuclei or cells into a predetermined number of clusters. The goal of k-means clustering is to cluster the nuclei or cells based upon either the original transcriptome data or the principal components derived from the original transcriptome data for the plurality of nuclei or cells into K partitions. In some embodiments, the k-means algorithm computes like clusters of entities from the higher dimensional data (where each dimension is either a different gene or a different principal component) and then after some resolution, the k-means clustering tries to minimize error. In this way, the k-means clustering provides cluster assignments.
[0239] In some embodiments, K is a number between 2 and 50 inclusive. In some embodiments, the number K is set to a predetermined number such as 10. In some embodiments, the number K is optimized. In some embodiments, the number K is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more than 30. In some embodiments, the number K is at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100.
[0240] In some embodiments, the clustering provides a suitable two or three dimensional clustering for visualization. Community based clustering, such as Louvain or Leiden clustering, are examples of such clustering. See, Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572, which is hereby incorporated by reference. In some embodiments, the single-nucleus or single-cell transcriptome data for the plurality of genes (where the number of dimensions are either the number of genes in the plurality of genes or, alternatively, in embodiments in which principal components are used, the number of principal components) for the first plurality of nuclei or cells are subjected to dimension reduction, e.g., in accordance with a manifold, into two or three dimensions and the resulting points in the dimension reduced plot, each representing a respective nucleus or cell in the plurality of nuclei or cells are coded by their cluster assignments. Dimension reduction programs such as UMAP, t-SNE, and PHATE can be used for this purpose. See, Heumos et al., 2023, “Best practices for single-cell analysis across modalities,” Nature Review Genetics 24, 550-572. In either approach, the high dimensional nature of the single-nucleus or single-cell transcriptome data is reduced to either a two-dimensional or three-dimensional plot in which each point represents a different nucleus or cell in the plurality of nuclei or cells and the identity of which cluster each nucleus or cell is in can be further emphasized by color coding the nucleus or cell by a unique color associated with each cluster. Figure 3 illustrates labeling. In this figure the high dimensional single-nucleus or single-cell transcriptome data for the plurality of genes for the first plurality of nuclei or cells has been reduced two dimensions for visualization. In Figure 3, each point represents a different nucleus. In fact, Figure 3 was constructed from the transcriptome data of 832,000 nuclei from 103 different subjects. The cluster assignment of the various nuclei represented in Figure 3 is shown by color coding and labeling.
[0241] Referring to block 256, in some embodiments, the first metadata is then used to identify a first cluster in the plurality of clusters with the first hepatic state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first hepatic state. Figure 4 illustrates. Both the left and right panel of Figure 4 illustratesthe clustering of nuclei from 47,500 mesenchymal cells (fibroblasts, stellate cells, smooth muscle cells and RGS5-high stellate cells) from frozen liver samples across 103 subjects of a cohort of subjects. Further illustrated in Figure 4 is that, in the left panel, nuclei of stellate cells that are quiescent are in cluster 402 whereas, in the right panel, nuclei of stellate cells that are activated are in cluster 404. From the first metadata it is known that the nuclei of quiescent stellate cells of cluster 402 are from subjects that are F0 on the fibrotic scale discussed in block 220 above, whereas the nuclei of stellate cells that are activated of cluster 404 are from subjects that are F3 on the fibrotic scale discussed in block 220 above. Thus, Figure 4 details the discovery that stellate cell transition from a quiescent to an activated state is significantly associated with worsening of fibrosis in NASH patients in accordance with the disclosure of block 256. Note that in some embodiments the cluster identifications made in Figure 4 (of which nuclei are associated with F0 and which are associated with F3) is made possible because of the barcoding discussed above in conjunction with block 252, which maps the single-nucleus transcriptome data of nuclei to the first metadata for such nuclei thus indicating the hepatic state (in this example, degree of fibrosis) of the respective source subject for each nuclei in clusters 402 and 404.
[0242] Figure 5 provides a further illustration of the clustered 47,500 mesenchymal cells of Figure 4. In Figure 5, first biomarker annotations in the first standardized set of biomarkers (e.g., genes) that each define a unique cell type are used to identify the nuclei originating from such cell types and to see which clusters these cell types are found. Determination that particular clusters are enriched for particular cell types is particularly illuminating in instances where clusters that are associated with particular hepatic state (e.g., first hepatic state, second hepatic state, etc.) are enriched for such cell types. As illustrated in Figure 5, panel 502, expression of genes FBLN2, OSR1, and CD34 indicate nuclei originating from fibroblasts, and panel 502 also shows which cluster fibroblasts are found. Panel 504 of Figure 5 illustrates how expression of genes SPARC, LAMA2, COL6A3, MYO10, and DCN indicate nuclei originating from stellate cells and panel 504 also shows which cluster stellate cells are found. Panel 506 of Figure 5 illustrates how expression of genes RGS5, PDGFB, PDGFRB, MCAM, PEC AMI, and CSPG4 indicate nuclei originating from RGS5-high stellate cells and panel 506 also shows which cluster RGS5-high (RGS5+) stellate cells are found. Panel 508 of Figure 5 illustrates how expression of genes MYH11, CNN1, MYL9, RYR2, ERBB4, and CACNA2D3 indicate nuclei originating from vascular smooth muscle cells (VSMC) and panel 508 also shows which cluster vascular smooth muscle cells are found. Figure 6 illustrates a summary of the cell specific type clustering identified in Figure5. In some embodiments gene markers set forth in Table 2 below are used to identify cell types set forth in Table 2 below. In some embodiments, for a respective cell type set forth in Table 2, at least 1, 2, 3, 4, 5, 6, 7, or 8 of the gene markers corresponding to the respective cell type set forth in Table 2 below are used to identify the respective cell type.
[0243] Table 2. Cell Types and associated Gene Markers
[0244] Identification of a cluster associated with a first hepatic state, such as example cluster 404 of Figure 4, allows for advantageous segmentation of human hepatic abnormalities into meaningful subcategories (hepatic states) based on relevant and actual differences in transcriptome expression between the different subcategories, among other practical applications.
[0245] Moreover, the identification of such a cluster can be used to determine whether any of the information in the first metadata is a covariate with respect to the first hepatic state. For instance, the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the identified cluster, or for the entire first plurality of nuclei or cells, can be fitted to a varying coefficient model to determine to examine the effect various covariates have on the first hepatic state. See, Hastie el al., 2001, The Elements of Statistical Learning, Data Mining, Inference, and Prediction,, Springer-Verlag, New York, New York, Section 6.4.2 beginning at p. 177, which is hereby incorporated by reference.
[0246] As an example, referring to block 258, in some such embodiments, the first cluster is used to determine the extent to which race affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0247] As another example, referring to block 260, in some such embodiments, the first cluster is used to determine the extent to which sex affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0248] As still another example, referring to block 262, in some such embodiments, the first cluster is used to determine the extent to which age affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0249] As still another example, referring to block 264, in some such embodiments, the first cluster is used to determine the extent to which a particular biomarker, such as the expression of a biomarker in the first standardized set of biomarkers of block 224, affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state. For instance, whether an abundance (expression level) of the particular biomarker affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0250] As still another example, referring to block 266, in some embodiments, the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage. For instance, as illustrated in Figure 16, stellate cells are defined by the expression of the gene signature SPARC, LAMA2, COL6A3, MYO10, and DCN. As another example as discussed in conjunction with Figure 12 in Example 1, Kupffer cells are characterized by expression of the gene signature comprising genes CD163, MARCO, TIMD4, CD5L, ADGRE1 (F4 / 80). Scar-associated macrophages (SAM) are characterized by expression of the gene signature comprising genes TREM2, CD9, SPPl(OPN), IL1B, CCR2, and LGALS3. Monocyte-derived macrophages are characterized by expression of the gene signature comprising genes CD14, FCGR3A (CD16), MNDA, ITGAM (CD1 IB), and CX3CR1.
[0251] In some such embodiments the first cluster is used to determine an extent to which a second biomarker, in the second standardized set of biomarkers, affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state. For instance, whether an abundance (expression level) of the second biomarker affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0252] Block 268 details an additional practical application for the discovery of the first cluster as detailed in block 257. In some such embodiments, genotype data for each subject in the plurality of subjects is obtained and this genotype data is overlay ed in each subject represented in the first cluster with a hepatic state of each subject in the first cluster. Alternatively or additionally, such genotype data case is overlay ed with each subjectrepresented in any cluster. Since each point in a cluster represents a nuclei or a cell from a particular subject on the cohort, and the identity of the subject is known because of the barcoding, each point in the cluster can be coded by a genotypic status of the corresponding subject. For instance, in some embodiments the genotypic status is the absence or presence of a major allele for a single nucleotide polymorphism (SNP) at a particular genetic locus in a genome. In one such embodiment, the nuclei or cell in the cluster that are from subjects that have the major allele for the SNP are indicated with a first color and the nuclei or cells in the cluster that are from subjects that have the minor allele for the SNP are indicated with a second color. In another example, in some embodiments the genotypic status is the absence or presence of a particular restriction fragment length polymorphism (RFLP) at a particular genetic locus in a genome. In one such embodiment, the nuclei or cells in the cluster that are from subjects that have the RFLP (or have a first allele for the RFLP) are indicated with a first color and the nuclei or cells in the cluster that are from subjects that do not have the RFLP (or have a second allele for the RFLP) are indicated with a second color. In alternative embodiments, rather than being a SNP or RFLP, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for particular copy number variations, insertions, or deletions. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the allelic status of a particular copy number variation, a particular genetic insertion, or a particular genetic deletion that arises in the genome of the species of the cohort of subjects. In alternative embodiments, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for a particular haplotype. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the haplotype of a particular haplotype arising in the genome of the species of the cohort of subjects. In alternative embodiments, the genotypic data overlayed on the first cluster is the allelic status of the cohort of subjects for a particular microsatellite or short tandem repeat. That is, for each subject having a respective nucleus or cell in the first cluster, the respective nucleus or cell is labeled by the allelic status of a particular microsatellite or short tandem repeat arising in the genome of the species of the cohort of subjects.
[0253] Referring to block 270, in some such embodiments, genotype data is obtained for each subject in the plurality of subjects, and upon overlay on the clustered transcriptome data, the clustered transcriptome data (e.g., the first cluster, 2 or more of the clusters, etc.) is used to determine an extent to which a genotype affects (is a covariate for) whether or not a subject incurs the first hepatic state and / or the degree they incur the first hepatic state.
[0254] Referring to block 272, in some embodiments, the hepatic abnormality detection system is used to associate a test subject with the first hepatic state. In some such embodiments, second information is obtained that comprises single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells. Each nucleus or cell in the second plurality of nuclei or cells is obtained from a liver tissue sample obtained from the test subject. The transcriptome data is co-clustered with the transcriptome data of the first plurality of nuclei or cells into the plurality of clusters. In instances where at least some of the nuclei or cells cluster into a particular cluster that is known from the metadata for the first plurality of nuclei or cells to be associated with a first hepatic state, some embodiments of the present disclosure associate the test subject with this first hepatic state. For instance, consider an example where the clustering of the transcriptome data for the first plurality of nuclei or cells identifies a first cluster associated with a first hepatic state. In this example, it is determined that the nuclei or cells in the first cluster are from activated stellate cells from subjects having an F3 fibrosis status. Coclustering of the transcriptome data for the second plurality of nuclei or cells with the transcriptome data for the second plurality of nuclei or cells reveals that nuclei or cells from activated stellate cells from the test subject also cluster into this first cluster. From this, it is deduced that the test subject also has an F3 fibrosis statis.
[0255] Referring to block 274, in some embodiments, the method informs a response to a drug compound in a patient or in a plurality of patients. Referring to block 276, in some embodiments, the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients. For instance, clusters 404 and 402 of Figure 4 can be used to screen for drugs that cause nuclei to shift from cluster 404 (activated stellate cells associated with F3 fibrosis) to cluster 402 (quiescent stellate cells associated with a health non-fibrotic state). For example consider the transcriptome data from the second plurality of nuclei of block 272 in which the nuclei of activated stellate cells in the second plurality of nuclei test of the test subject colocalize into cluster of Figure 4. A particular drug compound, a dosing amount, duration, and / or frequency of a drug can be recommended to the subject in order to push the state of the activated stellate cells towards cluster 402.
[0256] Referring to block 278, in some embodiments, the first metadata is used to identify a second cluster in the plurality of clusters with the second hepatic state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second hepatic state. Referring to block 280, in some embodiments, the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis. Clusters 404 and402 of Figure 4 illustrate such an embodiment. Cluster 404 includes nuclei from activated stellate cells from subjects having an F3 fibrotic state (first cluster, first hepatic state, first stage of fibrosis) whereas cluster 402 includes nuclei from quiescent stellate cells from subjects that do not have fibrosis (F0; first cluster, first hepatic state, second state of fibrosis). In some such embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the first cluster are cells or nuclei from subjects in the cohort of subjects that have the first hepatic state and at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are cells or nuclei from subjects in the cohort of subjects that have the second hepatic state.
[0257] Referring to block 282, in some embodiments, a determination is made that the first cluster comprises nuclei or cells of quiescent hepatic stellate cells and the second cluster comprises nuclei or cells of activated hepatic stellate cells. Clusters 404 and 402 of Figure 4 illustrate such an embodiment. Cluster 402 includes nuclei from quiescent stellate cells (first cluster, nuclei of quiescent hepatic stellate cells) whereas cluster 404 includes nuclei from activated stellate cells (second cluster, nuclei of activated hepatic stellate cells). In some such embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are nuclei of quiescent hepatic stellate or quiescent hepatic stellate cells and at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are activated hepatic stellate cells or are nuclei from activated hepatic stellate cells.
[0258] Referring to block 284, in some embodiments, a determination is made that the first cluster comprises nuclei or cells of a first type of activated hepatic stellate cells and the second cluster comprises nuclei or cells of a second type of activated hepatic stellate cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are nuclei or cells of the first type of activated hepatic stellate cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are nuclei or cells of the second type of activated hepatic stellate cells. Referring to Figure 4, two types of activated hepatic stellate cells are found by examination of the transcriptome data of the nuclei in cluster 404. Figure 7 is an enlarged view of cluster 404 of Figure 4. In Figure 7 it isseen that cluster 404 contains two types of activated stellate cells, type 1 found in subcluster 404A and type 2, found in subcluster 404B. Figure 8 compares the relative expression of a muscle associated signature (genes ACTA2, CARMN, LM0D1, MY01E, SLIT3, CRIM1, MRV1, MAG11, COL4A1, and MYOF) in the first type of activated hepatic stellate cells (HSC activated l) versus the second type of activated hepatic stellate cells (HSC_activated_2). It is seen in Figure 8 that the mean expression of each of these genes is greater in the second type of activated hepatic stellate cells than the first type of activated hepatic stellate cells. The expression of the muscle associated signature in the second type of activated hepatic stellate cells illustrated in Figure 8 is consistent with that of myofibroblasts. Myofibroblasts are absent from normal liver and are derived from hepatic stellate cells (HSCs) and portal mesenchymal cells in an injured liver. See Lemoinne et al., 2013, “Origins and functions of liver myofibroblasts,” Biochima et Biophysica Acta - Molecular Basis for Disease 1832(7), pp. 948-954, which is hereby incorporated by reference. Thus, in some embodiments of block 284, the second type of activated hepatic stellate cells are in fact myofibroblasts.
[0259] Referring to block 286, in some embodiments, a determination is made that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises RGS5+ stellate cells or nuclei of RGS5+ stellate cells. Cluster 510 and panel 506 of Figure 5 illustrate a cluster of RGS5+ stellate cells whereas a first cluster comprising nuclei of activated hepatic stellate cells is detailed previously. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are RGS5+ stellate cells or nuclei of RGS5+ stellate cells. As illustrated in Figure 5, the RGS5+ stellate cells exhibit high expression of the gene RGS5. The RGS5+ stellate cells of cluster 510 of Figure 5 also exhibit high expression of the genes PDGFB, PDGFRB, MCAM, PECAM1, and CSPG4.
[0260] Referring to block 288, in some embodiments, a determination is made that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises vascular smooth muscle cells or nuclei of vascular smooth muscle cells. Cluster 512 and panel 508 of Figure 5 illustrate a cluster of vascular smooth muscle cells whereas a first cluster comprising nuclei of activated hepatic stellate cells is detailed previously. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the cells or nuclei in the second cluster are vascular smooth muscle cells or nuclei of vascular smoothmuscle cells. As illustrated in Figure 5, the vascular smooth muscle cells of cluster 512 exhibit high expression of the genes MYH11, CNN1, MYL9, RYR2, ERBB4, and CACNA2D3.
[0261] Referring to block 290, in some embodiments, the first hepatic state is absence of fibrosis or inflammation and the second hepatic state is presence of fibrosis or inflammation. Clusters 402 and 404 of Figure 4 illustrate such an embodiment. Cluster 402 includes nuclei from quiescent stellate cells from subjects that do not have fibrosis (first cluster, first hepatic state, absence of fibrosis) whereas cluster 404 includes nuclei of cells from subjects having an F3 fibrotic state (second cluster, second hepatic state, presence of fibrosis). In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are nuclei or cells from subjects in the cohort of subjects that have the first hepatic state. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are nuclei or cells from subjects in the cohort of subjects that have fibrosis or inflammation.
[0262] Referring to block 292, in some embodiments, a further determination is made that the first cluster comprises quiescent stellate cells or nuclei of quiescent stellate cells and the second cluster comprises activated stellate cells or nuclei of activated stellate cells. Clusters 402 and 404 of Figure 4 also illustrate such an embodiment. Cluster 402 includes nuclei of quiescent stellate cells whereas cluster 404 includes nuclei of activated stellate cells. In some embodiments, at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the first cluster are quiescent stellate cells or nuclei of quiescent stellate cells. In some embodiments at least 10 percent, 20 percent, 30 percent, 40 percent, 50 percent, 60 percent, 70 percent, 80 percent or at least 90 percent of the nuclei or cells in the second cluster are activated stellate cells or nuclei of activated stellate cells.
[0263] Referring to block 294, in some embodiments, the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes. Figure 4 illustrates a first cluster representing a hepatic disease state (cluster 404, F3 fibrosis) and a second cluster representing a hepatic healthy state (cluster 402, no fibrosis). One or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster. Any art excepted definition of overexpression and under-expression can be used to identify genes that are overexpressed or under-expressed in the first cluster relative to the secondcluster. In some embodiments a measure of central tendency of each gene in the plurality of genes is determined across the nuclei or cells in the first cluster as well as in the nuclei or cells in the second cluster. Examples of a measure of central tendency include e.g., mean, median, mode, weighted mean, weighted median, and / or weighted mode. Then the foldchange for each respective gene is determined between the measure of central tendency for the respective gene across the first cluster versus across the second cluster. A fold change represents how many times the expression level of a gene has increased in the first cluster compared to the second cluster. In some embodiments a gene is considered overexpressed in the first cluster relative to the second cluster if its fold change (measure of central tendency across the first cluster divided by measure of central tendency across the second cluster) exceeds a predetermined threshold (e.g., 2 times higher, 2.5 times higher, 3 times higher, etc.). In some embodiments a gene is considered under-expressed in the first cluster relative to the second cluster if its fold change (measure of central tendency across the first cluster divided by measure of central tendency across the second cluster) is below a predetermined threshold (e.g., below 0.5 times, below .25 times, below 0.10 times).
[0264] In some embodiments a gene is considered overexpressed in the first cluster relative to the second cluster if a t-test or analysis of variance (ANOVA) test determines that differences in expression levels of the gene are statistically significant between the first cluster and the second cluster. Genes whose expression level are higher in the first cluster relative to the second cluster where a t-test or ANOVA indicates this higher expression has a p-value below a first threshold value are considered over-expressed in the first cluster relative to the second cluster in some embodiments. In some embodiments the first threshold is 0.10, 0.05, or 0.01. Genes whose expression level are lower in the first cluster relative to the second cluster where a t-test or ANOVA indicates this lower expression has a p-value below a second threshold value are considered under-expressed in the first cluster relative to the second cluster in some embodiments. In some embodiments the second threshold is 0.10, 0.05, or 0.01.
[0265] In some embodiments one or more genes are overexpressed in the first cluster relative to the second cluster while, at the same time one or more genes are under-expressed in the first cluster relative to the second cluster.
[0266] In some embodiments, the first and second clusters are further filtered to select a single cell type (e.g., stellate cells) in the first cluster and a single cell type in the second cluster (e.g., vascular smooth muscle cells, RGS5+ stellate cells, myofibroblasts) before performing differential expression analysis to find one or more genes that are overexpressedin the first cluster relative to the second cluster and / or one or more genes that are underexpressed in the first cluster relative to the second cluster In some embodiments selection of such cell types within the cluster is done by looking for expression of a respective gene signature in such cells that is characteristic of such cell types. The gene signature for many such cell types are known in the art and thus the filtering of the clusters for many specific cell types is done using conventional methods in such embodiments.
[0267] In some embodiments, the differential expression analysis identifies, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more genes that are over-expressed in the first cluster relative to the second cluster. In some embodiments, the differential expression analysis identifies, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more genes that are under-expressed in the first cluster relative to the second cluster. For instance, in some embodiments, using thresholds for differential expression, gene sets are derived that represent the most significant over-expressed and under-expressed genes.
[0268] In some embodiments, differentially expressed genes are identified using the Seurat FindAllMarkers function. See, Stuart et al., 2019, “Comprehensive integration of single-cell data,” Cell 177, 1888-1902, which is hereby incorporated by reference.
[0269] In some embodiments, these over-expressed and under-expressed genes (differential gene set expression signatures) are then characterized using algorithms that measure statistical enrichment for genes in particular pathways, with particular functions or with particular structural characteristics attained from publicly available databases. In some embodiments the statistical significance of enrichment is determined using a hypergeometric distribution or equivalently a one-tailed version of Fisher’s exact test. This and other methods for determining over-expressed and under-expressed genes in the differential expression analysis are disclosed in Plaisier et al., 2010, “Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures,” Nucleic Acids Research 38(17): el69, which is hereby incorporated by reference.
[0270] In some embodiments an examination of publicly curated databases is performed with each of the genes that are over-expressed and / or under-expressed in the first cluster relative to the second cluster to identify a metabolic pathway comprising a set of genes. Examples of such publicly curated databases is provided in Table 3, below.
[0271] Table 3 - publicly curated databases:
[0272] Each of the libraries listed in Table 3 is downloadable from the Internet at maayanlab.cloud / Enrichr / index.jsp#libraries. Another database where such pathways is found is MsigDB (v7.4). See Liberzon et al., 2015, The Molecular Signatures Database (MSigDB) hallmark gene set collection,” Cell Syst 1 : 417-425; and Liberzon et al., 2011, “Molecular signatures database (MSigDB) 3.0,” Bioinformatics 27: 1739-1740, each of which is hereby incorporated by reference.
[0273] Figure 26 illustrates how such analysis identified biological pathways associated with NASH M3 macrophages. As discussed in further detail below in conjunction with Figures 22A and 22B, NASH M3 macrophages are implicated in fibrosis, inflammation, and steatosis.
[0274] Referring to block 296, in some embodiments, the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state. In some such embodiments a plurality of compound-specific differential transcriptional signatures is accessed. Each respective compound-specific differential transcriptional signature is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set. The baseline transcriptional signature data set is from a control sample of one or more cells of a cell type (e.g., a cell type whose expression signature serves as a good proxy for the cells found in the first cluster or the second cluster). Each respective compound- treated transcriptional signature is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds. An example of compound-specific differential transcriptional signaturesis disclosed in United States Patent Application No. 63 / 580,612, entitled “Cellular Data Library and Methods of Using Same,” filed September 5, 2023, which is hereby incorporated by reference. A test differential transcription signature is generated from a differential comparison of the transcriptional signature of the nuclei or cells of the first cluster and the second cluster. In some embodiments, the differential comparison requires that the baseline transcriptional signature data set and each respective compound-treated transcriptional signature be run two, three, four times or more. In some embodiments, DESeq2, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Love et al, 2014, “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2,” which is hereby incorporated by reference. In some embodiments, edgeR, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Robinson and Smyth, 2007, “Moderated statistical tests for assessing differences in tag abundance,” Bioinformatics. 2007, 23: 2881-2887; and McCarthy et al., 2012, “Differential expression analysis of multifactor RNA-seq experiments with respect to biological variation,” Nucleic Acids Res. 40: 4288-4297, each of which is hereby incorporated by reference. In some embodiments, BBSeq, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Zhou, 2011, “powerful and flexible approach to the analysis of RNA sequence count data,” Bioinformatics 27, pp. 2672-2678, which is hereby incorporated by reference. In some embodiments, DSS, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Wang, 2013, “A new shrinkage estimator for dispersion improves differential expression detection in RNA-seq data,” Biostatistics 14: 232-243, which is hereby incorporated by reference. In some embodiments, baySeq, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Hardcastle and Kelly, 2010, “baySeq: empirical Bayesian methods for identifying differential expression in sequence count data,” BMC Bioinformatics 11 : 422, which is hereby incorporated by reference. In some embodiments, ShrinkBayes [, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Van De Wiel, 2013, “Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors,” Biostatistics 14: 113-128, which is hereby incorporated by reference.
[0275] The test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.For instance, consider a test differential transcription signature between cluster 404 (F3 activated) and cluster 402 (heathy quiescent) of Figure 4. What is desired is a compoundspecific differential transcriptional signature that matches the test differential transcription signature. In other words, identification of a compound that would push cells from having the transcriptional signature of cluster 404 to having one like cluster 402. In some embodiments, each of the compound-specific differential transcriptional signatures is compared to the test differential transcription signature using a program such as Rank-rank Hypergeometric Overlap (RRHO), or an equivalent algorithm. See Plaiser, 2010, “Rankrank hypergeometric overlap: identification of statistically significant overlap between geneexpression signatures,” Nucleic Acids Research 38(17), el 69, which is hereby incorporated by reference. In some embodiments, a compound-specific differential transcriptional signatures is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top N matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature, where N is a positive integer (e.g., 1, 5, 10, 20, 100, etc.). For instance, in embodiments where N is 5, a compound-specific differential transcriptional signature is considered to match the test differential transcription signature when a program such as RRHO identifies the compoundspecific differential transcriptional signatures as being upon the top 5 matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature. In some embodiments, a statistical test (e.g., t-test, ANOVA) is used to compare the plurality of compound-treated transcriptional signatures to the test differential transcription signature and each compound-treated transcriptional signature that has statistically significant similarity (e.g., P-value less than 0.15, less than 0.10, less than 0.05, less than 0.01) to the test differential transcription signature is considered to match the test differential transcription signature.
[0276] In some embodiments the plurality of compound-treated transcriptional signatures includes at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 8000, at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 80,000, at least 100,000, at least 200,000, at least 500,000, at least 800,000, at least 1 million, or at least 2 million transcriptional signatures, each representing a different compound.
[0277] In some embodiments, the plurality of compound-treated transcriptional signatures includes no more than 10 million, no more than 5 million, no more than 1 million, no more than 500,000, no more than 100,000, no more than 50,000, no more than 10,000, no more than 8000, no more than 5000, no more than 2000, no more than 1000, no more than 800, no more than 500, no more than 200, or no more than 100 compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures consists of from 10 to 500, from 100 to 10,000, from 5000 to 200,000, or from 10,000 to 1 million compound-treated transcriptional signatures, each representing a different compound.
[0278] In some embodiments, the plurality of compound-treated transcriptional signatures is between 10 and 1 x 106compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures is between 100 and 100,000 compound-treated transcriptional signatures, each representing a different compound. In some embodiments, the plurality of compound-treated transcriptional signatures is between 1000 and 100,000 compound-treated transcriptional signatures, each representing a different compound.
[0279] In some embodiments, the method further comprises formulating the first compound for use alleviating the hepatic disease state. In some embodiments formulating the first compound for use in a therapy comprises manufacturing a composition comprising the first compound and one or more excipients and / or one or more pharmaceutically acceptable carriers and / or one or more diluents.
[0280] Such excipients and / or carriers include all conventional solvents, dispersion media, fillers, solid carriers, coatings, antifungal and antibacterial agents, dermal penetration agents, surfactants, isotonic and absorption agents and the like. It will be understood that the compositions of the present disclosure may also include other supplementary physiologically active agents.
[0281] An exemplary carrier is pharmaceutically “acceptable” in the sense of being compatible with the other ingredients of the composition (e.g., the composition comprising the test chemical compound) and not injurious to a subject. The compositions may conveniently be presented in unit dosage form and may be prepared by any methods well known in the art of pharmacy. Such methods include the step of bringing into association the active ingredient with the carrier that constitutes one or more accessory ingredients. In general, the compositions are prepared by uniformly and intimately bringing into associationthe active ingredient with liquid carriers or finely divided solid carriers or both, and then, if necessary, shaping the product.
[0282] Exemplary compounds, compositions or combinations of the present disclosure (e.g., the first compound) formulated for intravenous, intramuscular or intraperitoneal administration, or a pharmaceutically acceptable salt, solvate or prodrug thereof may be administered by injection or infusion.
[0283] Injectables for such use can be prepared in conventional forms, either as a liquid solution or suspension or in a solid form suitable for preparation as a solution or suspension in a liquid prior to injection, or as an emulsion. Carriers can include, for example, water, saline (e.g., normal saline (NS), phosphate-buffered saline (PBS), balanced saline solution (BSS)), sodium lactate Ringer's solution, dextrose, glycerol, ethanol, and the like; and if desired, minor amounts of auxiliary substances, such as wetting or emulsifying agents, buffers, and the like can be added. Proper fluidity can be maintained, for example, by using a coating such as lecithin, by maintaining the required particle size in the case of dispersion and by using surfactants.
[0284] The compound, composition or combinations of the present disclosure (e.g., the first compound) may also be suitable for oral administration and may be presented as discrete units such as capsules, sachets or tablets each containing a predetermined amount of the active ingredient; as a powder or granules; as a solution or a suspension in an aqueous or nonaqueous liquid; or as an oil-in-water liquid emulsion or a water-in-oil liquid emulsion. The active ingredient may also be presented as a bolus, electuary or paste.
[0285] A tablet may be made by compression or molding, optionally with one or more accessory ingredients. Compressed tablets may be prepared by compressing in a suitable machine the active ingredient (e.g., the first compound) in a free-flowing form such as a powder or granules, optionally mixed with a binder (e.g., inert diluent, preservative disintegrant (e.g. sodium starch glycolate, cross-linked polyvinyl pyrrolidone, cross-linked sodium carboxymethyl cellulose) surface-active or dispersing agent). Molded tablets may be made by molding in a suitable machine a mixture of the powdered compound moistened with an inert liquid diluent. The tablets may optionally be coated or scored and may be formulated so as to provide slow or controlled release of the active ingredient therein using, for example, hydroxypropylmethyl cellulose in varying proportions to provide the desired release profile. Tablets may optionally be provided with an enteric coating, to provide release in parts of the gut other than the stomach.
[0286] The compound, composition or combinations of the present disclosure (e.g., the first compound) may be suitable for topical administration in the mouth including lozenges comprising the active ingredient in a flavored base, usually sucrose and acacia or tragacanth gum; pastilles comprising the active ingredient in an inert basis such as gelatine and glycerin, or sucrose and acacia gum; and mouthwashes comprising the active ingredient in a suitable liquid carrier.
[0287] The compound, composition or combinations of the present disclosure (e.g., the first compound) may be suitable for topical administration to the skin may comprise the compounds dissolved or suspended in any suitable carrier or base and may be in the form of lotions, gel, creams, pastes, ointments and the like. Suitable carriers include mineral oil, propylene glycol, polyoxyethylene, polyoxypropylene, emulsifying wax, sorbitan monostearate, polysorbate 60, cetyl esters wax, cetearyl alcohol, 2-octyldodecanol, benzyl alcohol and water. Transdermal patches may also be used to administer the compounds of the invention.
[0288] The compound, composition or combination of the present disclosure (e.g., the first compound) may be suitable for parenteral administration include aqueous and non-aqueous isotonic sterile injection solutions which may contain anti-oxidants, buffers, bactericides and solutes which render the compound, composition or combination isotonic with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions which may include suspending agents and thickening agents. The compound, composition or combination may be presented in unit-dose or multi-dose sealed containers, for example, ampoules and vials, and may be stored in a freeze-dried (lyophilized) condition requiring only the addition of the sterile liquid carrier, for example water for injections, immediately prior to use.Extemporaneous injection solutions and suspensions may be prepared from sterile powders, granules and tablets of the kind previously described.
[0289] It should be understood that in addition to the active ingredients particularly mentioned above, the composition or combination of this present disclosure (e.g., the first compound) may include other agents conventional in the art having regard to the type of composition or combination in question, for example, those suitable for oral administration may include such further agents as binders, sweeteners, thickeners, flavoring agents disintegrating agents, coating agents, preservatives, lubricants and / or time delay agents. Suitable sweeteners include sucrose, lactose, glucose, aspartame or saccharine. Suitable disintegrating agents include cornstarch, methylcellulose, polyvinylpyrrolidone, xanthan gum, bentonite, alginic acid or agar. Suitable flavoring agents include peppermint oil, oil ofwintergreen, cherry, orange or raspberry flavoring. Suitable coating agents include polymers or copolymers of acrylic acid and / or methacrylic acid and / or their esters, waxes, fatty alcohols, zein, shellac or gluten. Suitable preservatives include sodium benzoate, vitamin E, alpha-tocopherol, ascorbic acid, methyl paraben, propyl paraben or sodium bisulphite. Suitable lubricants include magnesium stearate, stearic acid, sodium oleate, sodium chloride or talc. Suitable time delay agents include glyceryl monostearate or glyceryl distearate.
[0290] Referring to block 298, in some embodiments, the hepatic disease state is selected from an acute stage, a chronic stage, a clinical stage, a flare-up, a remission, a progressive stage, a refractory, a subclinical stage, and a terminal phase of a hepatic disease.
[0291] Referring to block 300, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a solvent. In the case of the control sample, the cells are exposed to the solvent for a particular period of time and the solvent does not include a compound. In the case of the corresponding compound-treated sample, the ells are exposed to the solvent for a particular period of time with the compound dissolved in the solvent. Thus, in typical embodiments the solvent is the same solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Nonlimiting examples of solvents include water or aqueous-based solvents, alcohol-based solvents (e.g., ethanol or isopropyl alcohol), an ether such as bi s(2-methoxy ethyl) ether, P- cyclodextrin, or a polar aprotic solvent such as dimethylsulfoxide (DMSO), dimethylformamide (DMF), N,N-dimethylacetamide (DMAc), tetrahydrofuran (THF) acetonitrile (CEECN), or mixtures thereof. In some embodiments, the solvent is a polar aprotic solvent. In embodiments, the solvent is dimethylsulfoxide (DMSO). In some embodiments, the solvent comprises DMSO.
[0292] Referring to block 302, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a polar aprotic solvent. In the case of the control sample, the cells are exposed to the polar aprotic solvent for a particular period of time and the polar aprotic solvent does not include a compound. In the case of the corresponding compound-treated sample, the compound is exposed to the polar aprotic solvent with the corresponding compound dissolved in it. Thus, in typical embodiments, the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Optionally, the polar aprotic solvent comprises a mixture of a polar aprotic solvent and water. In some suchembodiments, the polar aprotic solvent concentration ranges from about 0.01 pM to about 10 pM, about 0.1 pM to about 5 pM, or about 0.5 pM to about 2 pM. In some embodiments, the polar aprotic solvent concentration (in water) is about 0.01 pM, about 0.1 pM, about 0.5 pM, about 1.0 pM, about 5 pM, or about 10 pM. In some embodiments, the polar aprotic solvent (in water) is about 1 pM.
[0293] Referring to block 304, in some embodiments, the control sample and each corresponding compound-treated sample is exposed to a DMSO solvent. Optionally, the DMSO solvent comprises a mixture of DMSO and water. In the case of the control sample, the cells are exposed to the DMSO solvent for a particular period of time and the DMSO solvent does not include a compound. In the case of the corresponding compound-treated sample, the compound is exposed to the DMSO solvent with the corresponding compound dissolved in it. Thus, in typical embodiments, the DMSO solvent is the same DMSO solvent for the control sample and each corresponding compound-treated sample with the exception that the solvent for each corresponding compound-treated sample includes a particular compound. Optionally, the DMSO solvent comprises a mixture of a DMSO and water. In some such embodiments, the DMSO concentration in the DMSO solvent ranges from about 0.01 pM to about 10 pM, about 0.1 pM to about 5 pM, or about 0.5 pM to about 2 pM. In some embodiments, the DMSO concentration in the DMSO solvent is about 0.01 pM, about 0.1 pM, about 0.5 pM, about 1.0 pM, about 5 pM, or about 10 pM. In some embodiments, the DMSO concentration in the DMSO solvent is about 1 pM.
[0294] Referring to block 306, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus or single-cell assay and / or single-cell assay data. Optionally, the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
[0295] Referring to block 308, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
[0296] Referring to block 310, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises scRNA-seq data.
[0297] Referring to block 312, in some embodiments, each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of scRNA-seq data.
[0298] Referring to block 314, in some embodiments, each corresponding compound- treated sample of one or more cells comprises hepatic stellate cells. Referring to block 316, in some embodiments, each corresponding compound-treated sample of one or more cells comprises a first type of activated hepatic stellate cells. Referring to block 318, in some embodiments, each corresponding compound-treated sample of one or more cells comprises a second type of activated hepatic stellate cells. Referring to block 320, in some embodiments, each corresponding compound-treated sample of one or more cells consists of hepatic stellate cells. Referring to block 322, in some embodiments, each corresponding compound-treated sample of one or more cells consists of a first type activated hepatic stellate cells. Referring to block 324, in some embodiments, each corresponding compound-treated sample of one or more cells consists of a second type of activated hepatic stellate cells. Referring to Figure 4, two types of activated hepatic stellate cells are found by examination of the transcriptome data of the nuclei in cluster 404. Figure 7 is an enlarged view of cluster 404 of Figure 4. In Figure 7 it is seen that cluster 404 contains two types of activated stellate cells, type 1 found in subcluster 404A and type 2, found in subcluster 404B. Figure 8 compares the relative expression of a muscle associated signature (genes ACTA2, CARMN, LM0D1, MY01E, SLIT3, CRIM1, MRV1, MAG11, COL4A1, and MYOF) in the first type of activated hepatic stellate cells (HSC activated l) versus the second type of activated hepatic stellate cells (HSC_activated_2). It is seen in Figure 8 that the mean expression of each of these genes is greater in the second type of activated hepatic stellate cells than the first type of activated hepatic stellate cells. The expression of the muscle associated signature in the second type of activated hepatic stellate cells illustrated in Figure 8 is consistent with that of myofibroblasts. Myofibroblasts are absent from normal liver and are derived from hepatic stellate cells (HSCs) and portal mesenchymal cells in an injured liver. See Lemoinne et al., 2013, “Origins and functions of liver myofibroblasts,” Biochima et Biophysica Acta - Molecular Basis for Disease 1832(7), pp. 948-954, which is hereby incorporated by reference. Thus, in some embodiments of block 284, the second type of activated hepatic stellate cells are in fact myofibroblasts.
[0299] Referring to block 326, in some embodiments, each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells, and / or cells from a cell line.
[0300] Referring to block 328, in some embodiments, each corresponding compound- treated sample of one or more cells is a frozen sample. In some embodiments, each corresponding compound-treated sample of one or more cells is an unfrozen sample.
[0301] Referring to block 330, in some embodiments, each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons. In some embodiments, each respective compound in the plurality of compounds is inorganic or organic. In some embodiments, each respective compound in the plurality of compounds is an organic compound having a molecular weight of less than 2000 Daltons (Da). In some embodiments, each respective compound in the plurality of compounds has a molecular weight of at least 10 Da, at least 20 Da, at least 50 Da, at least 100 Da, at least 200 Da, at least 500 Da, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 5 kDa, at least 10 kDa, at least 20 kDa, at least 30 kDa, at least 50 kDa, at least 100 kDa, or at least 500 kDa. In some embodiments, each respective compound in the plurality of compounds has a molecular weight of no more than 1000 kDa, no more than 500 kDa, no more than 100 kDa, no more than 50 kDa, no more than 10 kDa, no more than 5 kDa, no more than 2 kDa, no more than 1 kDa, no more than 500 Da, no more than 300 Da, no more than 100 Da, or no more than 50 Da. In some embodiments, each respective compound in the plurality of compounds has a molecular weight of from 10 Da to 900 Da, from 50 Da to 1000 Da, from 100 Da to 2000 Da, from 1 kDa to 10 kDa, from 5 kDa to 500 kDa, or from 100 kDa to 1000 kDa. In some embodiments, each respective compound in the plurality of compounds has a molecular weight that falls within another range starting no lower than 10 Daltons and ending no higher than 1000 kDa.
[0302] Referring to block 332, in some embodiments, each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria. In some each respective compound in the plurality of compounds is an organic compound that satisfies two or more rules, three or more rules, or all four rules of the Lipinski's Rule of Five: (i) not more than five hydrogen bond donors (e.g., OH and NH groups), (ii) not more than ten hydrogen bond acceptors (e.g. N and O), (iii) a molecular weight under 500 Daltons, and (iv) a LogP under 5. The “Rule of Five”is so called because three of the four criteria involve the number five. See, Lipinski, 1997, Adv. Drug Del. Rev. 23, 3, which is hereby incorporated herein by reference in its entirety. In some embodiments, a respective compound of the present disclosure satisfies one or more criteria in addition to Lipinski's Rule of Five. For example, in some embodiments, a compound of the present disclosure has five or fewer aromatic rings, four or fewer aromatic rings, three or fewer aromatic rings, or two or fewer aromatic rings.
[0303] Running differential stellate markers discovered using a NASH Atlas of Example 1 against an Intervention Library
[0304] Figures 17-20 in conjunction with Figure 16 and block 296 of Figure 21, illustrate a practical application of the disclosed NASH atlas.
[0305] In accordance with block 1700 of Figure 17A, a method for identifying a compound that transitions a hepatic disease state to a healthy state is provided.
[0306] In accordance with block 1702, single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a plurality of nuclei or cells is obtained, as described for instance in Example 1, where each nucleus or cell in the plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples from a cohort of subjects. As in Example 1, at least a first subset of subjects in the cohort of subjects have the hepatic disease state and a second subset of subjects in the cohort of subjects have the healthy state.
[0307] In accordance with block 1704, a second plurality of nuclei or cell is selected from the first plurality of nuclei or cells on the basis that each nuclei in the second plurality of nuclei or cells are a stellate cell. For instance, as illustrated in Figure 16, this can be done using the transcriptional signature of stellate cells which includes the markers SPARC, LAMA2, COL6A3, MYO 10, and DCN.
[0308] In accordance with block 1706, the second plurality of nuclei or cells is then clustered, based on the transcriptome data obtained for each nuclei or cells in the respective nuclei or cells into a first cluster representing a quiescent state and a second cluster representing an activated state. The UMAP projection of Figure 16, reproduced from Figure 4, illustrates the first cluster (cluster 402) and the second cluster (cluster 404).
[0309] In accordance with block 1708, a test differential transcription signature is then determined by differential comparison of expression of a second plurality of genes between the first cluster and the second cluster. In some embodiments, DESeq2, or an analogousalgorithm, is used to determine the differential transcriptional signature. See, Love et al, 2014, “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2,” which is hereby incorporated by reference. More discussion on suitable differential signature algorithms is disclosed above in conjunction with block 296.
[0310] In some embodiments, the second plurality of genes comprises at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 genes in the group consisting of LRAT, LHX2, NCAM1, NES, HGF, HAND2, RELN, ECM1, BAMBI, ETS1, NOTCH1, COLECI 1, ETS2 PLIN2, PPARG, RBP1, SPARCL1, TCF21, GATA4, GATA6, ACTA2, COL1A1, COL3A1, PDGFRB, PDGFRA, TIMP1, VCL, CCL2, MMP2, LAMC3, and LXN.
[0311] In some embodiments, the second plurality of genes consists of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or all 31 genes in the group consisting of LRAT, LHX2, NCAM1, NES, HGF, HAND2, RELN, ECM1, BAMBI, ETS1, NOTCH1, COLECI 1, ETS2 PLIN2, PPARG, RBP1, SPARCL1, TCF21, GATA4, GATA6, ACTA2, COL1A1, COL3A1, PDGFRB, PDGFRA, TIMP1, VCL, CCL2, MMP2, LAMC3, and LXN.
[0312] As noted in Figure 16, the quiescent healthy state is noted by expression of LRAT, LHX2, NCAM1, NES, HGF, HAND2, RELN, ECM1, BAMBI, ETS1, NOTCH1, COLECI 1, ETS2 PLIN2, PPARG, RBP1, SPARCL1, TCF21, GATA4, and GATA6, whereas the activated fibrotic state is noted by expression of ACTA2, COL1 Al, COL3A1, PDGFRB, PDGFRA, TIMP1, VCL, CCL2, MMP2, LAMC3, and LXN.
[0313] Figure 18 illustrates the fraction of cells in cluster 402 (quiescence) and cluster 404 (activated) by cell type and the expression of some of the above noted genes in these cell types.
[0314] In accordance with block 1710, in the screening method, a plurality of compoundspecific differential transcriptional signatures is accessed. Each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature of the second plurality of genes in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set of the second plurality of genes. The baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separatelytreated with a different compound in a plurality of compounds. In some embodiments the cell type is the LX-2 cell line. LX-2 is commercially available from sources such as Millipore Sigma. Such compound-specific differential transcriptional signatures are further disclosed in block 296 and in United States Patent Application No. 63 / 580,612, entitled “Cellular Data Library and Methods of Using Same,” filed September 5, 2023, which is hereby incorporated by reference.
[0315] In some embodiments the baseline transcriptional signature data set of the second plurality of genes is from the cell type after it has been activated. For example, in some embodiments the cell type is the LX-2 cell line that has been activated to represent the state of cluster 404 of Figure 4. In some embodiments this is done by exposing the cells to TGF- bl. In some embodiments each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds as well as TGF-bl. Figure 19 illustrates how LX-2 treated with TGF-bl showed enhanced production of pro-fibrotic endpoints Fibronectin 1 and pro-Collal. Figure 19 illustrates how LX-2 showed dose responsive increase in Fibronectin & Collagen lai (pro-Collal) in response to TGF-bl. Figure 19 illustrates how signal is observed across gene expression, protein content within cell lysate & secreted protein endpoints.
[0316] The goal of identifying an activated cell line that mimic cluster 404 of Figure 4 and then is used to search for modulators (compounds) that will revert this activated state is illustrated in Figure 20.
[0317] In accordance with block 1711 of Figure 17B, the test differential transcription signature is compared to each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures. This comparison identifies a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature. In some embodiments, each of the compound-specific differential transcriptional signatures is compared to the test differential transcription signature using a program such as Rank-rank Hypergeometric Overlap (RRHO), or an equivalent algorithm. See Plaiser, 2010, “Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures,” Nucleic Acids Research 38(17), el69, which is hereby incorporated by reference. In some embodiments, a compound-specific differential transcriptional signatures isconsidered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top N matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature, where N is a positive integer (e.g., 1, 5, 10, 20, 100, etc . For instance in embodiments where N is 5, a compound-specific differential transcriptional signatures is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top 5 matches among the plurality of compound- treated transcriptional signatures compared to the test differential transcription signature. In some embodiments, a statistical test (e.g., t-test, ANOVA) is used to compare the plurality of compound-treated transcriptional signatures to the test differential transcription signature and each compound-treated transcriptional signature that has statistically significant similarity (e.g., P-value less than 0.15, less than 0.10, less than 0.05, less than 0.01) to the test differential transcription signature is considered to match the test differential transcription signature.
[0318] In accordance with block 1712, through the identification of the first compound, a compound that transitions the hepatic disease state to the healthy state is identified.
[0319] Running differential macrophage markers discovered using the NASH Atlas of Example 1 against an Intervention Library
[0320] Figures 12, 13, 21, 22A, 22B, 23 A, 23B and 24, in conjunction block 296 of Figure 21, illustrate another practical application of the disclosed NASH atlas.
[0321] In accordance with block 2300 of Figure 23 A, a method for identifying a compound that transitions a hepatic disease state to a healthy state is provided.
[0322] In accordance with block 2302, single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a plurality of nuclei or cells is obtained, as described for instance in Example 1, where each nucleus or cell in the plurality of nuclei is obtained from a liver tissue sample in a plurality of liver tissue samples from a cohort of subjects. As in Example 1, at least a first subset of subjects in the cohort of subjects have the hepatic disease state and a second subset of subjects in the cohort of subjects have the healthy state.
[0323] In accordance with block 2304, a second plurality of nuclei or cell is selected from the first plurality of nuclei or cells on the basis that each nuclei or cell in the second plurality of nuclei or cells is a macrophage.
[0324] In accordance with block 2306, the second plurality of nuclei or cells is then clustered, based on the transcriptome data obtained for each nuclei or cells in the respective nuclei or cells into a first cluster representing a first macrophage state and a second cluster representing a second macrophage state. The UMAP projection of Figure 12 illustrates an example of this cluster (cluster 1204, M3) and the second cluster (clusters 1206 and or cluster 1202). Figure 22B also illustrates possible clustering, where the cells denoted M3 in Figure 22B are the first cluster and the second cluster are the cells denoted Ml and / or M2. Figure 24 illustrates an alternative clustering option.
[0325] In accordance with block 2308, a test differential transcription signature is then determined by differential comparison of expression of a second plurality of genes between the first cluster and the second cluster. In some embodiments, DESeq2, or an analogous algorithm, is used to determine the differential transcriptional signature. See, Love el al, 2014, “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2,” which is hereby incorporated by reference. More discussion on suitable differential signature algorithms is disclosed above in conjunction with block 296.
[0326] In some embodiments, the second plurality of genes comprises at least 2, 3, 4, 5, 6, 7, or 8 genes in the group consisting of include ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB.
[0327] In some embodiments, the second plurality of genes consists of 2, 3, 4, 5, 6, 7, or 8genes in the group consisting of include ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB.
[0328] In some embodiments, the second plurality of genes consists of Trem2 and CD9 expression.
[0329] In some embodiments, the second plurality of genes consists of Trem2, CD9, and SPP1 expression.
[0330] In accordance with block 2310, in the screening method, a plurality of compoundspecific differential transcriptional signatures is accessed. Each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature of the second plurality of genes in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set of the second plurality of genes. The baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptionalsignature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds. In some embodiments the cell type is a PBMC-derived cell line. In some embodiments the cell type is a THP-1 cell line. Such compound-specific differential transcriptional signatures are further disclosed in block 296 and in United States Patent Application No. 63 / 580,612, entitled “Cellular Data Library and Methods of Using Same,” filed September 5, 2023, which is hereby incorporated by reference.
[0331] In some embodiments the baseline transcriptional signature data set of the second plurality of genes is from the cell type after it has been activated. For example, referring to Figure 25, in some embodiments the cell type is a peripheral blood mononuclear cell (PBMC) derived cell line that has been exposed to GM-CSF in order to active the M3-like signature. In some embodiments this is done by exposing the cells to GM-CSF. In some embodiments each respective compound-treated transcriptional signature in the plurality of compound- treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds as well as GM-CSF.
[0332] In some embodiments the cell type is a peripheral blood mononuclear cell (PBMC) derived cell line that has been exposed to M-CSF in order to active the M3-like signature. In some embodiments this is done by exposing the cells to M-CSF. In some embodiments each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds as well as M-CSF.
[0333] In some embodiments the cell type is a peripheral blood mononuclear cell (PBMC) derived cell line that has been exposed to GM-CSF and M-CSF in order to active the M3-like signature. In some embodiments this is done by exposing the cells to GM-CSF and M-CSF. In some embodiments each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds as well as GM-CSF and M-CSF.
[0334] In accordance with block 2311 of Figure 22B, the test differential transcription signature is compared to each respective compound-specific differential transcriptionalsignature in the plurality of compound-specific differential transcriptional signatures. This comparison identifies a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature. In some embodiments, each of the compound-specific differential transcriptional signatures is compared to the test differential transcription signature using a program such as Rank-rank Hypergeometric Overlap (RRHO), or an equivalent algorithm. See Plaiser, 2010, “Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures,” Nucleic Acids Research 38(17), el69, which is hereby incorporated by reference. In some embodiments, a compound-specific differential transcriptional signatures is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top N matches among the plurality of compound-treated transcriptional signatures compared to the test differential transcription signature, where N is a positive integer (e.g., 1, 5, 10, 20, 100, etc . For instance in embodiments where N is 5, a compound-specific differential transcriptional signatures is considered to match the test differential transcription signature when a program such as RRHO identifies the compound-specific differential transcriptional signatures as being upon the top 5 matches among the plurality of compound- treated transcriptional signatures compared to the test differential transcription signature. In some embodiments, a statistical test (e.g., t-test, ANOVA) is used to compare the plurality of compound-treated transcriptional signatures to the test differential transcription signature and each compound-treated transcriptional signature that has statistically significant similarity (e.g., P-value less than 0.15, less than 0.10, less than 0.05, less than 0.01) to the test differential transcription signature is considered to match the test differential transcription signature.
[0335] In accordance with block 2312, through the identification of the first compound, a compound that transitions the hepatic disease state to the healthy state is identified.
[0336] Examples
[0337] Example 1 - Manufacturing a human hepatic abnormality detection system. A method for manufacturing a human hepatic abnormality detection system is provided. At a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors there was obtained, in electronic form, first information comprising single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells. The first plurality of nucleiwas approximately 832 thousand nuclei. Each nucleus in the first plurality of nuclei was obtained from a flash frozen liver tissue sample in a set of 103 flash frozen liver tissue samples.
[0338] Each frozen liver tissue sample in the plurality of frozen liver tissue samples was from a different subject in a cohort of 103 subjects. This represents a cohort that is an order of magnitude larger than publicly available single cell RNAseq of sorted immune cells.
[0339] Nuclei isolation was carried out by lysing cells, which releases cytoplasmic RNA into suspension (ambient RNA). Ambient RNA is a common source of technical variability between samples in single-nuclei datasets and may confound differential expression & clustering analysis. CellBender was used to evaluate and remove ambient RNA signature with minimal loss of other information. See Fleming et al.. 2019, “CellBender removebackground: a deep generative model for unsupervised removal of background noise from scRNA-seq datasets,” bioRxiv 791699, doi: 10.1101 / 791699, which is hereby incorporated by reference. As illustrated in Figure 11, this ambient RNA correction reduced hepatocyte marker score in non-hepatocyte cell types without affecting cell type specific marker scores. In fact, the ambient RNA correction reduced ambient and mitochondrial RNA but did not affect other data quality metrics such as median number of counts per cell, median number of genes per cell, nuclear marker frequency, or cell type frequency. Ambient RNA correction removed technical variability and improved cell type distinguishability in 103 / 109 NASH atlas samples. Only the samples in which ambient RNA correction were used for the cohort of subjects. In Figure 11, “pre-CB” refers to prior to ambient RNA correction with CellBender whereas “CB” refers to after ambient RNA correction with CellBender.
[0340] The first plurality of nuclei included a different subset of nuclei from a liver tissue sample from each subject in the cohort of subjects. The first information further included first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state. At least a first subset of subjects in the cohort of subjects had the first hepatic state and a second subset of subjects in the cohort of subjects had the second hepatic state. Figure 9 illustrates summary statistics for certain first metadata for the cohort of subjects. Fibrosis states and NAS scores summarized in Figure 9 are described above in conjunction with block 220 of Figure 2B. Histopathological assessment for NAFLD Activity Score (NAS) and fibrosis score was performed on harvested liver tissue. Fibrosis score (F0- F4) captures severity of fibrosis observed. NAS score (0-8) captured features of NASH: steatosis, lobular inflammation and hepatocyte ballooning. The cohort in this example had 17healthy (fibrosis scale FO) subjects, 16 nonalcoholic fatty liver FO subjects, and 70 subjects that had fibrosis within the range Fl to F4. The first metadata also includes the following information for each of the 103 subjects: sex, body mass index, age, ethnicity, cause of death, serology (CMV, EBV status), final lab profile, alcohol use, tobacco use, illicit drug use, medical history, and medications.
[0341] The respective single-nucleus transcriptome data for the plurality of genes for each nucleus in the first plurality of nucleic was barcoded with the subject in the cohort of subjects originating the respective single-nucleus transcriptome data.
[0342] Analysis of the transcriptome data for each of the nuclei in the first plurality of nuclei revealed that the first plurality of nuclei originated from the various cell types listed in the left hand column of Figure 10. The middle column of Figure 10 illustrates how many cells of each of these cell types is represented in the first plurality of nuclei, and the righthand column of Figure 10 provides the average number per subject in the cohort of subject per cell type.
[0343] The first plurality of nuclei was clustered into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus transcriptome data for the plurality of genes for each unique pair of nuclei in the first plurality of nuclei and (ii) evaluating the plurality of distances with a criterion function. The plurality of distances includes a separate distance for each unique pair of nuclei in the first plurality of nuclei. Each respective distance in the plurality of distances represents a different pair of nuclei in the first plurality of nuclei and quantifies a distance between (i) a respective first vector formed by the single-nucleus transcriptome data for the plurality of genes for a respective first nucleus in the different pair of nuclei and (ii) a respective second vector formed by the single-nuclei transcriptome data for the plurality of genes for a respective second nucleus in the different pair of nuclei. Each respective cluster in the plurality of clusters represents a corresponding subset of nuclei of the first plurality of nuclei that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei within the corresponding subset of nuclei with the criterion function. While this clustering was performed using principal components of the single-nucleus transcriptome data in high dimensional space, Figure 3 illustrates a two-dimensional UMAP rendering of the 832 thousand nuclei present across these high dimensional clusters.
[0344] Relative to single cell RNA sequencing, the single nucleotide sequencing of the present example exhibits lower frequency of infiltrating immune cells, lower librarycomplexity and higher ambient RNA. As discussed above, the ambient RNA was corrected by identifying it and removing it from the transcription data for each nuclei.
[0345] One major use for the dataset is illustrated in Figures 3 and 6. Another major use is illustrated in Figure 4. Figures 3 and 6 illustrates how the transcription data has been clustered and the cell types of each cluster identified (based on abundance of transcriptional cell type signatures in each cluster) and then visualized in two-dimensions using UMAP. Figure 4 illustrates how transcriptional clusters also captured transitions from one hepatic state to another. Thus, the embodiment of Figures 3 and 6 allows for the study of specific cell types whereas the embodiment of Figure 4 allows for the determination of the molecular basis (e.g., differential gene expression) that distinguishes different hepatic states from each other.
[0346] Figure 6 illustrates a UMAP of the clustering of transcriptome data limited to those nuclei in the plurality of nuclei that have the transcriptional cell type signature of stellate cells. Figure 4 illustrates the difference within the clustering of Figure 6 between nuclei from stellate cells from healthy subjects (left panel) and nuclei from stellate cells from subjects with F3 fibrosis (right panel). Figure 4 illustrates how stellate cell transitioning from a quiescent to an activated state is significantly associated with worsening of fibrosis in NASH subjects.
[0347] Figure 12 illustrates a UMAP of the clustering of transcriptome data limited to those nuclei in the plurality of nuclei that have the transcriptional cell type signature of macrophages. A total of 22,000 nuclei in the first plurality of nuclei from the 103 subjects of the cohort were used in the clustering illustrated in Figure 12. In Figure 12, cluster 1202 is predominantly Kupffer cells characterized by expression of the gene signature comprising genes CD163, MARCO, TIMD4, CD5L, ADGRE1 (F4 / 80). In Figure 12, cluster 1204 is predominantly scar-associated macrophages (SAM) characterized by expression of the gene signature comprising genes TREM2, CD9, SPPl(OPN), IL IB, CCR2, and LGALS3. In Figure 12, cluster 1206 is predominantly monocyte-derived macrophages characterized by expression of the gene signature comprising genes CD14, FCGR3A (CD16), MNDA, ITGAM (CD1 IB), and CX3CR1. Figure 13 illustrates the difference within the clustering of Figure 12 between nuclei from macrophages from healthy subjects (left panel) and nuclei from macrophages from subjects with F3 fibrosis (right panel). Figure 13 illustrates how there is a clinical cell state transition between liver resident cells (left panel) in cells from healthy subjects and enrichment of infiltrating M3 macrophages in fibrotic subjects. Figure 21 illustrates how canonical polarization states can be assigned to the macrophage andKupfer cell clusters of Figure 12. Figures 22A and 22B illustrates how there is a significant shift towards fibrosis, steatosis, and inflammation towards the M4>M3 state. Figures 22A and 22B also illustrates that the KCM2lowstate is more associated with a healthy state. Analysis of this data suggest that markers for M3 macrophages include ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB. Thus, referring to Figure 13, evaluation of ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB would be instructive in looking for compounds that push cells from the fibrotic state on the right to the healthy state on the left.
[0348] Figure 14 illustrates a UMAP of the clustering of transcriptome data limited to those nuclei in the plurality of nuclei that have the transcriptional cell type signature of T cells.Figure 15 illustrates the difference within the clustering of Figure 14 between nuclei from T cells from healthy subjects (left panel) and nuclei from T cells from subjects with F3 fibrosis (right panel). Figure 15 illustrates that, unlike for macrophages and stellate cells, there is no clear transition within the clustering between T cells from healthy subjects (left panel) and T cells from fibrotic subjects (right panel).
[0349] REFERENCES CITED AND ALTERNATIVE EMBODIMENTS
[0350] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0351] The present invention can be implemented as a computer program product that includes a computer program mechanism embedded in a non-transitory computer readable storage medium. For instance, the computer program product could contain the program modules shown in any combination of Figures 1-2. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, or any other non-transitory computer readable data or program storage product.
[0352] The present description includes example systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative implementations. For purposes of explanation, numerous specific details are set forth in order to provide an understanding of various implementations of the inventive subject matter. It will be evident, however, to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures and techniques have not been shown in detail.
[0353] Many modifications and variations of this invention can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. The specific embodiments described herein are offered by way of example only. The embodiments were chosen and described in order to best explain the principles of the invention and its practical applications, to thereby enable others skilled in the art to best utilize the invention and various embodiments with various modifications as are suited to the particular use contemplated. The invention is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
WHAT IS CLAIMED IS:
1. A method for manufacturing a human hepatic abnormality detection system, the method comprising: at a computer system having one or more processors, and memory storing one or more programs for execution by the one or more processors: obtaining, in electronic form, first information comprising:(i) single nucleus or single cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples, each liver tissue sample in the plurality of liver tissue samples is from a different subject in a cohort of subjects, the first plurality of nuclei or cells includes a different subset of nuclei or cells from a liver tissue sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state, wherein at least a first subset of subjects in the cohort of subjects have the first hepatic state and a second subset of subjects in the cohort of subjects have the second hepatic state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells,each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or singlecell transcriptome data for the plurality of genes for a respective first nucleus or first cell in the different pair of nuclei and (ii) a respective second vector formed by the single nucleus or single cell transcriptome data for the plurality of genes for a respective second nucleus or second cell in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; and using the first metadata to identify a first cluster in the plurality of clusters with the first hepatic state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first hepatic state.
2. The method of claim 1, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first hepatic state or the second hepatic state comprises a histologically graded disease status for the respective subject.
3. The method of claim 1, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first hepatic state or the second hepatic state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each liver tissue in the plurality of liver tissue samples.
4. The method of any one of claims 1-3, wherein each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50 liver tissue samples, or 100 or more liver tissue samples.
5. The method of any one of claims 1-4, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
6. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which race is a covariate with respect to the first hepatic state.
7. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which sex is a covariate with respect to the first hepatic state.
8. The method of claim 5, wherein the method further comprises using the first cluster to determine an extent to which age is a covariate with respect to the first hepatic state.
9. The method of any one of claims 1-8, wherein sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
10. The method of any one of claims 1-9, the method further comprising filtering the singlenucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
11. The method of any one of claims 1-10, wherein the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
12. The method claim 11, wherein the method further comprises using the first cluster to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first hepatic state.
13. The method of any one of claims 1-12, wherein the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
14. The method claim 11, wherein the method further comprises using the first cluster to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first hepatic state.
15. The method of any one of claims 1-14, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and overlaying genotype data for each subject represented in the first cluster with a hepatic state of each subject in the first cluster.
16. The method of any one of claims 1-15, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and determining an extent to which a genotype is a covariate for the first hepatic state using the genotype data for each subject represented in the first cluster.
17. The method of any one of claims 1-16, the method further comprising using the hepatic abnormality detection system to associate a test subject with the first hepatic state by a procedure comprising: obtaining, in electronic form, second information comprising single-nucleus or singlecell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells, wherein each nucleus or cell in the second plurality of nuclei or cells is obtained from a liver tissue sample obtained from the test subject; co-clustering the first plurality of nuclei or cells and the second plurality of nuclei or cells into the plurality of clusters; and identifying the test subject as having the first hepatic state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
18. The method of any one of claims 1-17, wherein the first hepatic state is fibrosis and the second hepatic state is absence of fibrosis.
19. The method of any one of claims 1-17, wherein the first hepatic state is absence of fibrosis and the second hepatic state is fibrosis.
20. The method of any one of claims 1-17, wherein the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
21. The method of any one of claims 1-17, wherein the first hepatic state is presence of liver inflammation and the second hepatic state is absence of liver inflammation.
22. The method of any one of claims 1-17, wherein the first hepatic state is absence of liver inflammation and the second hepatic state is presence of liver inflammation.
23. The method of any one of claims 1-17, wherein the first hepatic state is a first stage of liver inflammation and the second hepatic state is a second stage of liver inflammation.
24. The method of any one of claims 1-17, wherein the first hepatic state is presence of liver steatosis and the second hepatic state is absence of liver steatosis.
25. The method of any one of claims 1-17, wherein the first hepatic state is absence of liver steatosis and the second hepatic state is presence of liver steatosis.
26. The method of any one of claims 1-17, wherein the first hepatic state is a first stage of liver steatosis and the second hepatic state is a second stage of liver steatosis.
27. The method of any one of claims 1-26, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, or sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, or sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
28. The method of any one of claims 1-27, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, and sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, and sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
29. The method of any one of claims 1-28, the method further comprising: using the first metadata to identify a second cluster in the plurality of clusters with the second hepatic state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second hepatic state.
30. The method of claim 29, wherein the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
31. The method of claim 29, the method further determining that the first cluster comprises quiescent hepatic stellate cells or nuclei of quiescent hepatic stellate cells and the second cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells.
32. The method of claim 29, the method further determining that the first cluster comprises a first type of activated hepatic stellate cells or nuclei of the first type of activated hepatic stellate cells and the second cluster comprises a second type of activated hepatic stellate cells or nuclei of the second type of activated hepatic stellate cells.
33. The method of claim 29, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises RGS5+ stellate cells or nuclei of RGS5+ stellate cells.
34. The method of claim 29, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises vascular smooth muscle cells (VSMCs) or nuclei of VSMCs.
35. The method of claim 29, wherein the first hepatic state is absence of fibrosis or inflammation and the second hepatic state is presence of fibrosis or inflammation.
36. The method of claim 35, the method further determining that the first cluster comprises quiescent stellate cells or nuclei of quiescent stellate cells and the second cluster comprises activated stellate cells or nuclei of activated stellate cells.
37. The method of claim 29, the method further determining that the first cluster comprises Ml or M2 macrophage cells or nuclei of Ml or M2 macrophage cells and the second cluster comprises M3 macrophage cells or nuclei of M3 macrophage cells.
38. The method of claim 29, the method further determining that the first cluster comprises M3 macrophages or nuclei of M3 macrophages and the second cluster comprises cells or nuclei of cells other than M3 macrophage cells.
39. The method of any one of claims 1-38, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
40. The method of any one of claims 1-39, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
41. The method of any one of claims 1-40, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
42. The method of any one of claims 1-41, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
43. The method of any one of claims 1-42, wherein the first hepatic state or the second hepatic state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis (NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
44. The method of any one of claims 1-43, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
45. The method of any one of claims 1-44, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
46. The method of claim 29, wherein the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
47. The method of any one of claims 29-46, wherein the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises: accessing, in electronic form, a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds; determining a test differential transcription signature by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
48. The method of claim 47, wherein the hepatic disease state is selected from an acute stage, a chronic stage, a clinical stage, a flare-up, a remission, a progressive stage, a refractory, a subclinical stage, and a terminal phase of a hepatic disease.
49. The method of claim 47 or 48, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
50. The method of any one of claims 47 to 49, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
51. The method of any one of claims 47 to 50, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
52. The method of any one of claims 47 to 51, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the singlenucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
53. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
54. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nuclei or single-cell RNA sequencing (scRNA-seq) data.
55. The method of any one of claims 47 to 52, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nuclei or single-cell RNA sequencing (scRNA-seq) data.
56. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises hepatic stellate cells.
57. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises a first type of activated hepatic stellate cells.
58. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises a second type of activated hepatic stellate cells.
59. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of hepatic stellate cells.
60. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of a first type of activated hepatic stellate cells.
61. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of a second type of activated hepatic stellate cells.
62. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises macrophages.
63. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises M3 macrophages.
64. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises Kupffer M2hlghcells.
65. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of macrophages.
66. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of M3 macrophages.
67. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells consists of Kupffer M2hlghcells.
68. The method of any one of claims 47 to 55, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
69. The method of any one of claims 47 to 68, wherein each corresponding compound- treated sample of one or more cells is a frozen sample.
70. The method of any one of claims 47 to 69, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
71. The method of any one of claims 47 to 70, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
72. A computer system for manufacturing a human hepatic abnormality detection system, the computer system comprising: one or more processors; and memory addressable by the one or more processors, the memory storing at least one program for execution by the one or more processors, the at least one program comprising instructions for: obtaining, in electronic form, first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus in a first plurality of nuclei or cells, wherein each nucleus in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples,each liver tissue sample in the plurality of liver tissue samples is from a different subject in a cohort of subjects, the first plurality of nuclei or cells includes a different subset of nuclei or cell from a liver tissue sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state, wherein at least a first subset of subjects in the cohort of subjects have the first hepatic state and a second subset of subjects in the cohort of subjects have the second hepatic state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells, each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or singlecell transcriptome data for the plurality of genes for a respective first nucleus or cell in the different pair of nuclei and (ii) a respective second vector formed by the singlenucleus or single-cell transcriptome data for the plurality of genes for a respective second nucleus in the different pair of nuclei, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; andusing the first metadata to identify a first cluster in the plurality of clusters with the first hepatic state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first hepatic state.
73. The computer system of claim 72, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first hepatic state or the second hepatic state comprises a histologically graded disease status for the respective subject.
74. The computer system of claim 72, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the first hepatic state or the second hepatic state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each liver tissue in the plurality of liver tissue samples.
75. The computer system of any one of claims 72-74, wherein each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50 liver tissue samples, or 100 or more liver tissue samples.
76. The computer system of any one of claims 72-75, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
77. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which race is a covariate with respect to the first hepatic state.
78. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which sex is a covariate with respect to the first hepatic state.
79. The computer system of claim 76, wherein the method further comprises using the first cluster to determine an extent to which age is a covariate with respect to the first hepatic state.
80. The computer system of any one of claims 72-79, wherein sex, race, age, and / or physical condition are each represented in a balanced manner in the cohort of subjects.
81. The computer system of any one of claims 72-80, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
82. The computer system of any one of claims 72-81, wherein the first metadata includes one or more first biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a first standardized set of biomarkers that each define a unique cell type.
83. The computer system claim 82, wherein the method further comprises using the first cluster to determine an extent to which a first biomarker in the first standardized set of biomarkers is a covariate with respect to the first hepatic state.
84. The computer system of any one of claims 72-83, wherein the first metadata includes one or more second biomarker annotations for each nucleus or cell in the plurality of nuclei or cells drawn from a second standardized set of biomarkers that each define a unique cell type at a unique stage.
85. The computer system claim 84, wherein the method further comprises using the first cluster to determine an extent to which a second biomarker in the second standardized set of biomarkers is a covariate with respect to the first hepatic state.
86. The computer system of any one of claims 72-85, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and overlaying genotype data for each subject represented in the first cluster with a hepatic state of each subject in the first cluster.
87. The computer system of any one of claims 72-86, the method further comprising: obtaining genotype data for each subject in the plurality of subjects; and determining an extent to which a genotype is a covariate for the first hepatic state using the genotype data for each subject represented in the first cluster.
88. The computer system of any one of claims 72-87, the method further comprising using the hepatic abnormality detection system to associate a test subject with the first hepatic state by a procedure comprising: obtaining, in electronic form, second information comprising single-nucleus or singlecell transcriptome data for the plurality of genes for each nucleus or cell in a second plurality of nuclei or cells, wherein each nucleus or cell in the second plurality of nuclei or cells is obtained from a liver tissue sample obtained from the test subject; co-clustering the first plurality of nuclei or cells and the second plurality of nuclei or cells into the plurality of clusters; and identifying the test subject as having the first hepatic state when nuclei or cells from the second plurality of nuclei or cells co-cluster into the first cluster.
89. The computer system of any one of claims 72-88, wherein the first hepatic state is fibrosis and the second hepatic state is absence of fibrosis.
90. The computer system of any one of claims 72-88, wherein the first hepatic state is absence of fibrosis and the second hepatic state is fibrosis.
91. The computer system of any one of claims 72-88, wherein the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
92. The computer system of any one of claims 72-88, wherein the first hepatic state is presence of liver inflammation and the second hepatic state is absence of liver inflammation.
93. The computer system of any one of claims 72-88, wherein the first hepatic state is absence of liver inflammation and the second hepatic state is presence of liver inflammation.
94. The computer system of any one of claims 72-88, wherein the first hepatic state is a first stage of liver inflammation and the second hepatic state is a second stage of liver inflammation.
95. The computer system of any one of claims 72-88, wherein the first hepatic state is presence of liver steatosis and the second hepatic state is absence of liver steatosis.
96. The computer system of any one of claims 72-88, wherein the first hepatic state is absence of liver steatosis and the second hepatic state is presence of liver steatosis.
97. The computer system of any one of claims 72-88, wherein the first hepatic state is a first stage of liver steatosis and the second hepatic state is a second stage of liver steatosis.
98. The computer system of any one of claims 72-97, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, or sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, or sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, or sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
99. The computer system of any one of claims 72-98, wherein the first information further comprises second metadata for each subject in the cohort of subjects comprising age, body mass index, and sex, and the method further comprises using the second information to prune the cohort of subjects based on age, body mass index, and sex thereby causing the cohort of subjects to be free of confounding for age, body mass index, and sex, wherein the pruning causes a subset of subjects to be removed from the cohort of subjects prior to the clustering.
100. The computer system of any one of claims 72-99, the method further comprising: using the first metadata to identify a second cluster in the plurality of clusters with the second hepatic state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the second hepatic state.
101. The computer system of claim 100, wherein the first hepatic state is a first stage of fibrosis and the second hepatic state is a second stage of fibrosis.
102. The computer system of claim 100, the method further determining that the first cluster comprises quiescent hepatic stellate cells or nuclei of quiescent hepatic stellate cells and the second cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells.
103. The computer system of claim 100, the method further determining that the first cluster comprises a first type of activated hepatic stellate cells or nuclei of the first type of activated hepatic stellate cells and the second cluster comprises a second type of activated hepatic stellate cells or nuclei of the second type of activated hepatic stellate cells.
104. The computer system of claim 100, the method further determining that the first cluster comprises Ml or M2 macrophage cells or nuclei of Ml or M2 macrophage cells and the second cluster comprises M3 macrophage cells or nuclei of M3 macrophage cells.
105. The computer system of claim 100, the method further determining that the first cluster comprises Kupffer M2hlghcells or nuclei of Kupffer M2hlghcells and the second cluster comprises Kupffer M2lowcells or nuclei of Kupffer M2lowcells.
106. The computer system of claim 100, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises RGS5+ stellate cells or nuclei of RGS5+ stellate cells.
107. The computer system of claim 100, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises vascular smooth muscle cells or nuclei of VSMCs.
108. The computer system of claim 100, wherein the first hepatic state is absence of fibrosis or inflammation and the second hepatic state is presence of fibrosis or inflammation.
109. The computer system of claim 108, the method further determining that the first cluster comprises quiescent stellate cells or nuclei of quiescent stellate cells and the second cluster comprises activated stellate cells or nuclei of activated stellate cells.
110. The computer system of claim 108, the method further determining that the first cluster comprises Ml or M2 macrophages or nuclei of Ml or M2 macrophages and the second cluster comprises M3 macrophages or nuclei of M3 macrophages.
111. The computer system of any one of claims 72-110, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
112. The computer system of any one of claims 72-111, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
113. The computer system of any one of claims 72-112, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
114. The computer system of any one of claims 72-113, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
115. The computer system of any one of claims 72-114, wherein the first hepatic state or the second hepatic state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non- Alcoholic Steatohepatitis (NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
116. The computer system of any one of claims 72-115, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
117. The computer system of any one of claims 71-116, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
118. The computer system of claim 100, wherein the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
119. The computer system of any one of claims 72-118, wherein the first cluster represents a hepatic disease state and the second cluster represents a hepatic healthy state, and the method further comprises: accessing, in electronic form, a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound- treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of at least 10 compounds; determining a test differential transcription signature by differential comparison of the transcriptional signature of the nuclei or cells of the first cluster and the second cluster; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
120. The method of claim 119, wherein the hepatic disease state is selected from an acute stage, a chronic stage, a clinical stage, a flare-up, a remission, a progressive stage, a refractory, a subclinical stage, and a terminal phase of a hepatic disease.
121. The computer system of claim 119 or 120, wherein the control sample and each corresponding compound-treated sample each comprises a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent is dimethyl sulfoxide (DMSO).
122. The computer system of claim 119 or 120, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
123. The computer system of any one of claims 119 or 120, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
124. The computer system of any one of claims 119 to 123, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the single-nucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
125. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
126. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
127. The computer system of any one of claims 119 to 124, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
128. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises hepatic stellate cells.
129. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises a first type of activated hepatic stellate cells.
130. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises a second type of activated hepatic stellate cells.
131. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of hepatic stellate cells.
132. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of a first type of activated hepatic stellate cells.
133. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of a second type of activated hepatic stellate cells.
134. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises macrophages.
135. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises M3 macrophages.
136. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises Kupffer M2hlghcells.
137. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of macrophages.
138. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells consists of M3 macrophages or Kupffer M2hlghcells.
139. The computer system of any one of claims 119 to 127, wherein each corresponding140. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
141. The computer system of any one of claims 119 to 127, wherein each corresponding compound-treated sample of one or more cells is a frozen sample.
142. The computer system of any one of claims 119 to 141, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
143. The computer system of any one of claims 119 to 142, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
144. A non-transitory computer readable storage medium, wherein the non-transitory computer readable storage medium stores instructions, which when executed by a computerI l lsystem, cause the computer system to perform a method for manufacturing a human hepatic abnormality detection system, the method comprising: obtaining, in electronic form, first information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus in a first plurality of nuclei or cells, wherein each nucleus in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples, each liver tissue sample in the plurality of liver tissue samples is from a different subject in a cohort of subjects, the first plurality of nuclei cells includes a different subset of nuclei or cells from a liver tissue sample from each subject in the cohort of subjects, and the first plurality of nuclei or cells comprises at least 1000 nuclei or cells, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has a first hepatic state or a second hepatic state, wherein at least a first subset of subjects in the cohort of subjects have the first hepatic state and a second subset of subjects in the cohort of subjects have the second hepatic state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters by (i) computing a plurality of distances using the single-nucleus or single-cell transcriptome data for the plurality of genes for each unique pair of nuclei or cells in the first plurality of nuclei or cells and (ii) evaluating the plurality of distances with a criterion function, wherein the plurality of distances includes a separate distance for each unique pair of nuclei or cells in the first plurality of nuclei or cells, each respective distance in the plurality of distances represents a different pair of nuclei or cells in the first plurality of nuclei or cells and quantifies a distance between (i) a respective first vector formed by the single-nucleus or singlecell transcriptome data for the plurality of genes for a respective first nucleus in the different pair of nuclei or cells and (ii) a respective second vector formed by thesingle nucleus or single cell transcriptome data for the plurality of genes for a respective second nucleus in the different pair of nuclei or cells, and each respective cluster in the plurality of clusters represents a corresponding subset of nuclei or cells of the first plurality of nuclei or cells that are clustered together based on evaluation of distances in the plurality of distances representing different pairs of nuclei or cells within the corresponding subset of nuclei or cells with the criterion function; and using the first metadata to identify a first cluster in the plurality of clusters with the first hepatic state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the first hepatic state.
145. A method for identifying a compound that transitions a hepatic disease state to a healthy state, the method comprising: obtaining information comprising:(i) single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a first plurality of nuclei or cells, wherein each nucleus or cell in the first plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples from a cohort of subjects, and(ii) first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the hepatic disease state or the healthy state, wherein at least a first subset of subjects in the cohort of subjects have the hepatic disease state and a second subset of subjects in the cohort of subjects have the healthy state, and wherein the respective single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells is barcoded with the subject in the cohort of subjects originating the respective single-nucleus or single-cell transcriptome data; clustering the first plurality of nuclei or cells into a plurality of clusters based on the single-nucleus or single-cell transcriptome data for the plurality of genes for each pair of nuclei or cells in the first plurality of nuclei or cells; using the information to identify a first cluster in the plurality of clusters with the hepatic disease state by determining that the first cluster includes nuclei or cells from subjects in the cohort of subjects that have the hepatic disease state;using the information to identify a second cluster in the plurality of clusters with the healthy state by determining that the second cluster includes nuclei or cells from subjects in the cohort that have the healthy state; accessing a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set, wherein, the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds; determining a test differential transcription signature by differential comparison of a transcriptional signature of the nuclei or cells of the first cluster and the second cluster; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures, thereby identifying a first compound associated with a compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature.
146. The method of claim 145, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the hepatic disease state or the healthy state comprises a histologically graded disease status for the respective subject.
147. The method of claim 145, wherein the first metadata for each respective subject in the cohort of subjects indicating at least for each respective subject in the cohort of subjects whether the respective subject has the hepatic disease state or the healthy state comprises a histologically graded disease status for the respective subject determined in accordance with a consistent, verified handling of each liver tissue in the plurality of liver tissue samples.
148. The method of any one of claims 145-147, wherein the first metadata for each respective subject in the cohort of subjects further indicates one or more features of the subject selected from the group consisting of sex, race, age, and physical condition.
149. The method of any one of claims 145-148, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
150. The method of any one of claims 145-149, wherein the hepatic disease state is fibrosis, a stage of fibrosis, presence of liver inflammation, a stage of liver inflammation, liver steatosis, or a stage of liver steatosis.
151. The method of any one of claims 145-150, the method further determining that the first cluster comprises nuclei or cells of a first type of activated hepatic stellate cells and the second cluster comprises nuclei or cells of a second type of activated hepatic stellate cells.
152. The method of any one of claims 145-150, the method further determining that the first cluster comprises M3 macrophages or nuclei of M3 macrophages and the second cluster comprises Ml or M2 macrophages or nuclei of Ml or M2 macrophages.
153. The method of any one of claims 145-150, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises RGS5+ stellate cells or nuclei of RGS5+ stellate cells.
154. The method of any one of claims 145-150, the method further determining that the first cluster comprises activated hepatic stellate cells or nuclei of activated hepatic stellate cells and the second cluster comprises vascular smooth muscle cells or nuclei of VSMCs.
155. The method of any one of claims 145-150, the method further determining that the first cluster comprises quiescent stellate cells or nuclei of quiescent stellate cells and the second cluster comprises activated stellate cells or nuclei of activated stellate cells.
156. The method of any one of claims 145-155, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
157. The method of any one of claims 145-156, wherein the first plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
158. The method of any one of claims 145-157, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
159. The method of any one of claims 145-158, wherein the hepatic disease state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis (NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
160. The method of any one of claims 145-159, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
161. The method of any one of claims 145-159, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
162. The method of any one of claims 145-161, wherein the first cluster represents the hepatic disease state and the second cluster represents the healthy state, and the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
163. The method of any one of claims 145-162, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide (DMSO).
164. The method of any one of claims 145-163, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
165. The method of any one of claims 145-153, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
166. The method of any one of claims 145-165, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the single — nucleus-assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
167. The method of any one of claims 145-165, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
168. The method of any one of claims 145-165, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
169. The method of any one of claims 145-165, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
170. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises hepatic stellate cells.
171. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises a first type of activated hepatic stellate cells.
172. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises a second type of activated hepatic stellate cells.
173. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of hepatic stellate cells.
174. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of a first type of activated hepatic stellate cells.
175. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of a second type of activated hepatic stellate cells.
176. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises macrophages.
177. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises M3 macrophages.
178. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells comprises Kupffer M2hlghcells.
179. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of macrophages.
180. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of M3 macrophages.
181. The method of any one of claims 145-169, wherein each corresponding compound- treated sample of one or more cells consists of Kupffer M2hlghcells.
182. The method of any one of claims 145-181, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
183. The method of any one of claims 145-182, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
184. The method of any one of claims 145-183, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
185. A method for identifying a compound that transitions a hepatic disease state to a healthy state, the method comprising: obtaining single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a plurality of nuclei or cells, wherein each nucleus or cell in the plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples from a cohort of subjects, wherein at least a first subset of subjects in the cohort of subjects have the hepatic disease state and a second subset of subjects in the cohort of subjects have the healthy state; selecting a second plurality of nuclei or cells from the first plurality of nuclei or cells on the basis that each nuclei or cell in the second plurality of nuclei or cell is from a stellate cell; clustering the second plurality of nuclei or cells into a first cluster representing a quiescent state and a second cluster representing an activated state using the single-cell or single-nucleus transcriptome data for a plurality of genes for each nucleus or cell in the second plurality of nuclei or cells; determining a test differential transcription signature by differential comparison of expression of a second plurality of genes between the first cluster and the second cluster; accessing a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature of the second plurality of genes ina plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set of the second plurality of genes, wherein, the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures to identify a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature, wherein the first compound is identified as the compound that transitions the hepatic disease state to the healthy state by the comparing.
186. The method of claim 185, wherein the second plurality of genes comprises at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 genes in the group consisting of LRAT, LHX2, NCAM1, NES, HGF, HAND2, RELN, ECM1, BAMBI, ETS1, NOTCH1, COLECI 1, ETS2 PLIN2, PPARG, RBP1, SPARCL1, TCF21, GATA4, GATA6, ACTA2, COL1A1, COL3A1, PDGFRB, PDGFRA, TIMP1, VCL, CCL2, MMP2, LAMC3, and LXN.
187. The method of claim 185, wherein the second plurality of genes consists of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or all 31 genes in the group consisting of LRAT, LHX2, NCAM1, NES, HGF, HAND2, RELN, ECM1, BAMBI, ETS1, NOTCH1, COLECI 1, ETS2 PLIN2, PPARG, RBP1, SPARCL1, TCF21, GATA4, GATA6, ACTA2, COL1A1, COL3A1, PDGFRB, PDGFRA, TIMP1, VCL, CCL2, MMP2, LAMC3, and LXN.
188. The method of any one of claims 185-187, wherein each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50 liver tissue samples, or 100 or more liver tissue samples.
189. The method of any one of claims 185-188, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
190. The method of any one of claims 185-189, wherein the hepatic disease state is fibrosis, a stage of fibrosis, presence of liver inflammation, a stage of liver inflammation, liver steatosis, or a stage of liver steatosis.
191. The method of any one of claims 185-190, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
192. The method of any one of claims 185-191, wherein the plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
193. The method of any one of claims 185-192, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
194. The method of any one of claims 185-193, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
195. The method of any one of claims 185-194, wherein the hepatic disease state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis (NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
196. The method of any one of claims 185-195, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
197. The method of any one of claims 185-195, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
198. The method of any one of claims 185-197, wherein the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
199. The method of any one of claims 185-198, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide.
200. The method of any one of claims 185-198, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
201. The method of any one of claims 185-198, wherein the control sample and each corresponding compound-treated sample comprises dimethyl sulfoxide.
202. The method of any one of claims 185-201, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the singlenucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
203. The method of any one of claims 185-201, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
204. The method of any one of claims 185-201 wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
205. The method of any one of claims 185-201, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
206. The method of any one of claims 185-205, wherein each corresponding compound- treated sample of one or more cells comprises hepatic stellate cells.
207. The method of any one of claims 185-205, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
208. The method of any one of claims 185-207, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
209. The method of any one of claims 185-208, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
210. A method for identifying a compound that transitions a hepatic disease state to a healthy state, the method comprising: obtaining single-nucleus or single-cell transcriptome data for a plurality of genes for each nucleus or cell in a plurality of nuclei or cells, wherein each nucleus or cell in the plurality of nuclei or cells is obtained from a liver tissue sample in a plurality of liver tissue samples from a cohort of subjects, wherein at least a first subset of subjects in the cohort of subjects have the hepatic disease state and a second subset of subjects in the cohort of subjects have the healthy state;selecting a second plurality of nuclei or cells from the first plurality of nuclei or cells on the basis that each nuclei or cell in the second plurality of nuclei or cells is from a macrophage cell; clustering the second plurality of nuclei or cells into a first cluster representing a Ml and / or M2 state and a second cluster representing an M3 state; determining a test differential transcription signature by differential comparison of expression of a second plurality of genes between the first cluster and the second cluster; accessing a plurality of compound-specific differential transcriptional signatures, wherein each respective compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures is a difference between (i) a respective compound-treated transcriptional signature of the second plurality of genes in a plurality of compound-treated transcriptional signatures and (ii) a baseline transcriptional signature data set of the second plurality of genes, wherein, the baseline transcriptional signature data set is from a control sample of one or more cells of a cell type, and each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures is from a corresponding compound-treated sample of one or more cells of the cell type separately treated with a different compound in a plurality of compounds; and comparing the test differential transcription signature to each respective compoundspecific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures to identify a first compound associated with a compound-specific differential transcriptional signature in the plurality of compound-specific differential transcriptional signatures that matches the test differential transcription signature, wherein the first compound is identified as the compound that transitions the hepatic disease state to the healthy state by the comparing.
211. The method of claim 210, wherein the second plurality of genes comprises at least 2, 3, 4, 5, 6, 7, or 8 genes in the group consisting of ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB.
212. The method of claim 210, wherein the second plurality of genes consists of 2, 3, 4, 5, 6, 7, or 8 genes in the group consisting of ITGAX, PPARG, CD83, SPP1, CD9, LPL, TREM2, and GPNMB.
213. The method of any one of claims 210-212, wherein each liver tissue sample in the plurality of liver tissue samples is from a different subject in the cohort of subjects and the plurality of liver tissue samples comprises 5 liver tissue samples, 20 liver tissue samples, 50 liver tissue samples, or 100 or more liver tissue samples.
214. The method of any one of claims 210-213, the method further comprising filtering the single-nucleus or single-cell transcriptome data for the plurality of genes for each nucleus or cell in the first plurality of nuclei or cells to remove counts of ambient RNA molecules, doublets and / or empty droplets.
215. The method of any one of claims 210-214, wherein the hepatic disease state is fibrosis, a stage of fibrosis, presence of liver inflammation, a stage of liver inflammation, liver steatosis, or a stage of liver steatosis.
216. The method of any one of claims 210-214, wherein the plurality of genes comprises 100 or more genes, 200 or more genes, 500 or more genes, 1000 or more genes, 2000 or more genes, 5000 or more genes, 10000 or more genes, or 15,000 or more genes.
217. The method of any one of claims 210-216, wherein the plurality of nuclei or cells comprises 10,000 or more nuclei or cells, 50,000 or more nuclei or cells, 100,000 or more nuclei or cells, 250,000 or more nuclei or cells, 500,000 or more nuclei or cells, 600,000 or more nuclei or cells, or 1 x 106or more nuclei or cells.
218. The method of any one of claims 210-217, wherein the cohort of subjects comprises 25 or more subjects, 50 or more subjects, 75 or more subjects, or 100 or more subjects.
219. The method of any one of claims 210-218, wherein the first subset of subjects comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects and the second subset of subjects is other than the first subset of subjects and comprises 5 or more subjects, 10 or more subjects, or 15 or more subjects.
220. The method of any one of claims 210-219, wherein the hepatic disease state is associated with Non-Alcoholic Fatty Liver Disease (NAFLD), Non-Alcoholic Steatohepatitis(NASH), NASH with fibrosis, fibrosis, simple fatty liver or steatosis, cirrhosis, hepatocellular carcinoma (HCC), liver failure, pre-diabetes, diabetes, vascular and cardiac diseases, atherosclerosis, hyperlipidemia, hyperglycemia, infarctions, ictus, or hypertension.
221. The method of any one of claims 210-220, wherein the method informs a response to a drug compound in a patient or in a plurality of patients.
222. The method of any one of claims 210-221, wherein the method informs a response to a dosing amount, duration, and / or frequency of a drug in a patient or in a plurality of patients.
223. The method of any one of claims 210-222, wherein the method further comprises identifying a metabolic pathway comprising a set of genes, wherein one or more genes in the set of genes are overexpressed or under-expressed in the first cluster relative to the second cluster.
224. The method of any one of claims 210-223, wherein the control sample and each corresponding compound-treated sample is exposed to a solvent, wherein the solvent is the same solvent for the control sample and each corresponding compound-treated sample, optionally wherein the solvent comprises dimethyl sulfoxide.
225. The method of any one of claims 210-223, wherein the control sample and each corresponding compound-treated sample each comprises a polar aprotic solvent, wherein the polar aprotic solvent is the same polar aprotic solvent for the control sample and each corresponding compound-treated sample.
226. The method of any one of claims 210-223, wherein the control sample and each corresponding compound-treated sample comprises DMSO.
227. The method of any one of claims 210-226, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus assay and / or single-cell assay data, optionally wherein the singlenucleus assay and / or single-cell assay data is selected from single-cell ribonucleic acid (RNA) sequencing (scRNA-seq) data, single-nucleus RNA sequencing (snRNA-seq) data, scTag-seq data, single-cell assay for transposase-accessible chromatin using sequencing(scATAC-seq) data, CyTOF / SCoP data, E-MS / Abseq data, miRNA-seq data, CITE-seq data, or any combinations thereof.
228. The method of any one of claims 210-226, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises or consists of single-cell RNA sequencing (scRNA-seq) data.
229. The method of any one of claims 210-226, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures comprises single-nucleus RNA sequencing (scRNA-seq) data.
230. The method of any one of claims 210-226, wherein each respective compound-treated transcriptional signature in the plurality of compound-treated transcriptional signatures consists of single-nucleus RNA sequencing (scRNA-seq) data.
231. The method of any one of claims 210-230, wherein each corresponding compound- treated sample of one or more cells comprises or consists of one or more cells from an organ, cells from a tissue, stem cells, human cells, cells from umbilical cord blood, cells from peripheral blood, bone marrow cells, cells from a solid tissue, differentiated cells and / or cells from a cell line.
232. The method of any one of claims 210-231, wherein each respective compound in the plurality of compounds has a molecular weight of less than 2000 Daltons.
233. The method of any one of claims 210-232, wherein each respective compound in the plurality of compounds satisfies at least three criteria of the Lipinski rule of five criteria, or optionally each of the Lipinski rule of five criteria.
Citation Information
Patent Citations
Machine learning implementation for multi-analyte assay development and testing
US20210174958A1
Systems and methods for associating compounds with physiological conditions using fingerprint analysis
US20220403335A1
Cited By
Method for automatically constructing pathological image data set and training cell nucleus detection and classification based on space transcriptome technology
CN121884006A