Systems and methods for methylation analysis of liver disease

EP4750919A1Pending Publication Date: 2026-06-03ACTIVE GENOMES EXPRESSED DIAGNOSTICS CORP

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
ACTIVE GENOMES EXPRESSED DIAGNOSTICS CORP
Filing Date
2024-07-25
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Current diagnostic methods for liver disease, particularly non-alcoholic steatohepatitis (NASH) and metabolic dysfunction-associated steatohepatitis (MASH), are invasive, expensive, and often detect the disease at advanced stages, leading to poor prognosis and high mortality rates.

Method used

A method involving the assay of bodily samples to determine quantitative measures of NASH/MASH-specific methylated biomarkers in nucleic acids, followed by computer processing using a trained machine learning algorithm to detect the presence or absence of NASH or MASH.

Benefits of technology

This method enables early detection of liver disease, distinguishes between different types of liver diseases, and provides a non-invasive, reliable diagnostic tool, potentially reducing mortality and healthcare costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024039559_30012025_PF_FP_ABST
    Figure US2024039559_30012025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for detecting a presence or an absence of a liver disease in a subject. An exemplary method includes assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in a nucleic acid from the bodily sample, computer processing the set of liver-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference, and detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR METHYLATION ANALYSIS OF LIVERDISEASECROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 528,755, filed July 25, 2023, which is incorporated by reference herein in its entirety.BACKGROUND

[0002] Liver disease is one of the most common chronic diseases in the United States, affecting about 100 million people in the U.S. Liver disease is often associated with diabetes, obesity, and metabolic syndrome among several other risk factors. While 80 million patients have simple steatosis (the benign form of liver disease), 20 million patients have nonalcoholic steatohepatitis (NASH, the advanced form of liver disease) that can lead to advanced stage liver diseases and liver cancer. The incidence of liver disease in the United States has significantly increased in the past two decades, with projections to continue rising. Blood tests for liver functions and imaging procedures (e.g., abdominal ultrasound) are commonly used for liver disease diagnosis. Liver biopsies are accurate, however, expensive, invasive, and thus, less ineffective to serve the population at large. It is a challenge to obtain a reliable, accurate, and painless diagnosis for liver disease at an early stage, especially when distinguishing among different types of liver diseases.SUMMARY

[0003] Liver disease is typically asymptomatic until patients have progressed to advanced stages. Many patients are diagnosed at late stages, during which the prognosis of the disease can be poor, mortality rates and healthcare spending high. Therefore, there are needs for a reliable and non-invasive diagnostic test assay to detect a presence or an absence of liver disease at an early stage. Moreover, there are unmet medical / clinical needs for a diagnostic test / assay capable of distinguishing between different types of liver diseases such that patients can be accurately diagnosed and appropriately treated.

[0004] In an aspect, the present disclosure provides a method for detecting a presence or an absence of non-alcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH) in a subject, the method comprising: (a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of NASH / MASH-specific methylated biomarkers in nucleic acids obtained or derived from the bodily sample; (b) computer processing the set of NASH / MASH-specific methylatedbiomarkers using a trained machine learning algorithm or compared against a reference (e.g., a set of reference DNA sequences); and (c) detecting the presence or the absence of NASH or MASH in the subject, based at least in part on the computer processing in (b).

[0005] In some embodiments, the set of NASH / MASH-specific methylated biomarkers comprises a member selected from the group listed in Table 2.

[0006] In some embodiments, the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a solid tissue sample, and a lymphatic fluid sample.

[0007] In some embodiments, the subject is diagnosed with NASH or MASH, or is suspected of having NASH or MASH.

[0008] In some embodiments, the subject has a risk factor for NASH or MASH. In some embodiments, the risk factor for NASH / MASH comprises type 2 diabetes, obesity, metabolic syndrome, family history of liver disease, genetic factors associated with NASH / MASH or liver fibrosis, and / or polycystic ovary syndrome.

[0009] In some embodiments, the subject is asymptomatic for NASH or MASH.

[0010] In some embodiments, the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

[0011] In some embodiments, the assaying in (a) further comprises amplifying the nucleic acid. In some embodiments, the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification (e.g., loop-mediated isothermal amplification (LAMP)).

[0012] In some embodiments, the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing, pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylation-specific PCR.

[0013] In some embodiments, the method further comprises using primers or probes configured to selectively enrich the nucleic acid for the set of liver-specific methylated biomarkers. In some embodiments, the primers or probes are nucleic acid primers, nucleic acid probes, or peptide nucleic acid (PNA) probes. In some embodiments, the nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of liverspecific methylated biomarkers.

[0014] In some embodiments, the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads for a sequence block and a total number of sequence reads for the sequence block, or (ii) a mean methylation level of the sequence block.

[0015] In some embodiments, the quantitative measures are determined as (i) a first combined beta value of at least two NASH / MASH-specific hypermethylated biomarkers, or (ii) a second combined beta value of at least two NASH / MASH-specific hypomethylated biomarkers. In some embodiments, the reference comprises a first reference beta value of the at least two NASH / MASH-specific hypermethylated biomarkers and a second reference beta value of the at least two NASH / MASH-specific hypomethylated biomarkers, and the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with NASH or MASH.

[0016] In some embodiments, the method further comprises administering a treatment to the subject based on the detected presence of NASH or MASH in the subject. In some embodiments, the treatment is selected from the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention. In some embodiments, the treatment is a pharmaceutical treatment selected from the group consisting of: Resmetiron / Rezdiffra (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione, ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication (naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 lU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

[0017] In some embodiments, the method further comprises assaying the bodily sample to determine quantitative measures of a set of NASH / MASH-specific hypermethylated or hypomethylated cytosine-phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample. In some embodiments, the set of NASH / MASH-specific hypermethylated or hypomethylated CpG sites comprises a member selected from the group listed in Table 2.

[0018] In some embodiments, the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

[0019] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0020] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0021] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0022] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a positive predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0023] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0024] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.99.

[0025] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of NASH or MASH and a second set of independent training samples associated with an absence or non-elevated susceptibility of NASH or MASH.

[0026] In another aspect, the present disclosure provides a method for detecting a presence or an absence of a liver disease in a subject, the method comprising: (a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in nucleic acids obtained or derived from the bodily sample, wherein the set of liver-specific methylated biomarkers comprises a member selected from the group listed in Table 1; (b) computer processing the set of liver-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference; and (c) detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing in (b).

[0027] In some embodiments, the liver disease is selected from the group consisting of steatosis, nonalcoholic steatohepatitis (NASH), metabolic dysfunction-associated steatohepatitis (MASH), non-alcoholic fatty liver disease (NAFLD), metabolic dysfunction- associated steatotic liver disease (MASLD), liver fibrosis, cirrhosis, and chronic liver disease.

[0028] In some embodiments, the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a mine sample, a solid tissue sample, and a lymphatic fluid sample.

[0029] In some embodiments, the subject is diagnosed with the liver disease or is suspected of having the liver disease.

[0030] In some embodiments, the subject is asymptomatic for the liver disease.

[0031] In some embodiments, the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

[0032] In some embodiments, the assaying in (a) further comprises amplifying the nucleic acid. In some embodiments, the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification (e.g., loop-mediated isothermal amplification (LAMP)).

[0033] In some embodiments, the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing,pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylation-specific PCR.

[0034] In some embodiments, the method further comprises using primers or probes configured to selectively enrich the nucleic acid for the set of liver-specific methylated biomarkers. In some embodiments, the primers or probes are nucleic acid primers or nucleic acid probes. In some embodiments, the nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of liver-specific methylated biomarkers.

[0035] In some embodiments, the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads in a sequence block and a total number of sequence reads in the sequence block, or (ii) a mean methylation level of the sequence block.

[0036] In some embodiments, the quantitative measures are determined as (i) a first combined beta value of at least two liver-specific hypermethylated biomarkers, or (ii) a second combined beta value of at least two liver-specific hypomethylated biomarkers. In some embodiments, the reference comprises a first reference beta value of the at least two liver-specific hypermethylated biomarkers and a second reference beta value of the at least two liver-specific hypomethylated biomarkers, and wherein the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with liver disease.

[0037] In some embodiments, the method further comprises administering a treatment to the subject based on the detected presence of the liver disease in the subject. In some embodiments, the treatment is selected from the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention. In some embodiments, the treatment is a pharmaceutical treatment selected from the group consisting of: Resmetiron / Rezdiffra (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione, ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication(naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 lU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

[0038] In some embodiments, the method further comprises assaying the bodily sample to determine quantitative measures of a set of liver-specific hypermethylated or hypomethylated cytosine-phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample. In some embodiments, the set of liver- specific hypermethylated or hypomethylated CpG sites comprises a member selected from the group listed in Table 1.

[0039] In some embodiments, the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

[0040] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0041] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0042] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0043] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a positive predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0044] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, atleast about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0045] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.99.

[0046] In some embodiments, the trained machine learning algorithm is configured to distinguish between a presence or an absence of each of a plurality of different liver diseases.

[0047] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of the liver disease and a second set of independent training samples associated with an absence or non-elevated susceptibility of the liver disease.

[0048] In another aspect, the present disclosure provides a method for detecting a presence or an absence of a liver disease in a subject, the method comprising: (a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in a nucleic acid from the bodily sample, wherein the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads in a sequence block and a total number of sequence reads in the sequence block, (ii) a mean methylation level of the sequence block, (iii) a first combined beta value of at least one liver-specific hypermethylated biomarker, or (iv) a second combined beta value of at least one liver-specific hypomethylated biomarkers; (b) computer processing the set of liverspecific methylated biomarkers using a trained machine learning algorithm or compared against a reference; and (c) detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing in (b).

[0049] In some embodiments, the set of liver- specific methylated biomarkers comprises a member selected from the group listed in Table 1.

[0050] In some embodiments, the liver disease is selected from the group consisting of steatosis, nonalcoholic steatohepatitis (NASH), metabolic dysfunction-associated steatohepatitis (MASH), non-alcoholic fatty liver disease (NAFLD), metabolic dysfunction- associated steatotic liver disease (MASLD), liver fibrosis, cirrhosis, and chronic liver disease.

[0051] In some embodiments, the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a tissue sample, and a lymphatic fluid sample.

[0052] In some embodiments, the subject is diagnosed with the liver disease or is suspected of having the liver disease.

[0053] In some embodiments, the subject is asymptomatic for the liver disease.

[0054] In some embodiments, the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

[0055] In some embodiments, the assaying in (a) further comprises amplifying the nucleic acid. In some embodiments, the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification (e.g., loop-mediated isothermal amplification (LAMP)).

[0056] In some embodiments, the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing, pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylation-specific PCR.

[0057] In some embodiments, the method further comprises using primers or probes configured to selectively enrich the nucleic acid for the set of liver-specific methylated biomarkers. In some embodiments, the primers or probes are nucleic acid primers or nucleic acid probes. In some embodiments, the nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of liver-specific methylated biomarkers.

[0058] In some embodiments, the reference comprises a first reference beta value of the at least one liver-specific hypermethylated biomarkers and a second reference beta value of the at least one liver-specific hypomethylated biomarkers, and wherein the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with liver disease.

[0059] In some embodiments, the method further comprises administering a treatment to the subject based on the detected presence of the liver disease in the subject. In some embodiments, the treatment is selected from the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention. In some embodiments, the treatment is a pharmaceutical treatment selected from the group consisting of: Resmetiron / Rezdiffra (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione,ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof. In some embodiments, the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication (naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 IU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

[0060] In some embodiments, the method further comprises assaying the bodily sample to determine quantitative measures of a set of liver-specific hypermethylated or hypomethylated cytosine-phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample.

[0061] In some embodiments, the set of liver- specific hypermethylated or hypomethylated CpG sites comprises a member selected from the group listed in Table 1.

[0062] In some embodiments, the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

[0063] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0064] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0065] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0066] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a positive predictive value of at leastabout 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0067] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

[0068] In some embodiments, the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.99.

[0069] In some embodiments, the trained machine learning algorithm is configured to distinguish between a presence or an absence of each of a plurality of different liver diseases.

[0070] In some embodiments, the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of the liver disease and a second set of independent training samples associated with an absence or non-elevated susceptibility of the liver disease.

[0071] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0072] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.

[0073] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE

[0074] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the present disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:

[0076] FIG. 1 illustrates an example workflow of determining the presence of liver disease and NASH / MASH in a subject based on the methylation status of liver-specific and NASH / MASH-specific methylated biomarkers.

[0077] FIG. 2 illustrates an example workflow of detecting the methylation status of methylated biomarkers and cytosine-phosphate-guanine (CpG) sites.

[0078] FIG. 3 illustrates a heatmap of selected liver-specific hypomethylated sequence blocks (N = 477).

[0079] FIG. 4 illustrates a heatmap of selected liver-specific hypomethylated sequence blocks (N = 2,642).

[0080] FIG. 5 illustrates a heatmap of selected liver-specific hypermethylated sequence blocks (N = 243).

[0081] FIG. 6 illustrates a histogram of NASH / MASH versus simple steatosis using whole genome bisulfite sequencing (WGBS) assessment.

[0082] FIG. 7A-7D illustrate box plots of beta values of hypermethylated biomarkers that distinguish NASH / MASH from simple steatosis.

[0083] FIG. 8A and 8B illustrate box plots of beta values of hypomethylated biomarkers that distinguish NASH / MASH from simple steatosis.

[0084] FIG. 9A illustrates box plots of combined hypermethylated biomarkers that distinguish NASH / MASH from simple steatosis.

[0085] FIG. 9B illustrates box plots of combined hypomethylated biomarkers that distinguish NASH / MASH from simple steatosis.

[0086] FIG. 10 is a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION

[0087] While various embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed.

[0088] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0089] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0090] While there have been significant improvements in diagnosing and treating liver disease, early detection offers patients the best opportunity for early diagnosis, treatment, and remission. Imaging procedures (e.g., abdominal ultrasound, computerized tomography (CT) scanning, magnetic resonance imaging (MRI), or elastography) are often used as initial tests for liver disease diagnosis. However, these techniques often lack the ability of distinguishing different types of liver diseases. Liver tissue examination (e.g., liver biopsy) are often performed after liver disease has progressed to later stages. Treatment at this stage can be challenging and offers poor prognosis. Moreover, the progression of liver disease, for example, from chronic liver damage and inflammation to cirrhosis and cancer is highly variable among patients. The course and rate of liver disease progression are often unpredictable using prognostic tools that are currently available. Therefore, early detection strategies that are easy to access with low cost are needed.

[0091] Epigenetic mechanisms governing gene expression and cellular phenotype have been brought to attention providing insights in disease diagnosis and therapeutics design. One mechanism of epigenetic modifications is through DNA methylation, in which DNA methyltransferase (DNMT) adds a methyl group to cytosine (C) bases to regulate gene expression. DNMT is involved in metabolic reprogramming, apoptosis, inflammation, among many other mechanisms. Methylation in non-coding regions of genomic DNA (e.g., upstream, or downstream regulatory regions) can turn on or turn off the expression of genes leading to increasing or decreasing levels of proteins encoded by the genes. Thus, the identification of methylation patterns in the regulatory regions is an indication of whether or not genes controlled by the regulatory regions are turned on or off. Further, the regulation of RNA spicing activity and specificity may lead to the production of alternatively spliced mRNA transcripts and downstream alternative protein variants / isoforms, which may have significantly altered functions and interactions with other proteins.

[0092] The present disclosure enables the identification of “on / off” state of genes or alternatively spliced mRNA transcripts implicated in different types of liver diseases and prediction as to whether a patient has an elevated risk of developing corresponding diseases. The identified methylated biomarkers enable the detection of liver disease and states of disease, and have the ability to distinguish among different types of liver disease. The identified methylation patterns and biomarkers exhibit both tissue-specificity and disease specificity in a superior manner than transcriptional expression, genetic mutations, or proteins. In some embodiments, the identified methylated biomarkers may be liver-specific and thus, can be used to determine whether a patient has an elevated risk of developing a liver disease. In other embodiments, the methylated biomarkers may be disease specific. These methylated biomarkers can be used to determine whether a patient has an elevated risk of developing a particular type of liver disease and to distinguish from other types of liver diseases.

[0093] One aspect described herein provides a method for detecting a presence or an absence of a liver disease in a subject. In some embodiments, the method may comprise (a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in a nucleic acid from the bodily sample, (b) computer processing the set of liver-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference, and (c) detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing.

[0094] The subject may have a high risk for a liver disease selected from the group consisting of viral hepatitis, steatosis, nonalcoholic steatohepatitis (NASH), metabolic dysfunction- associated steatohepatitis (MASH), non-alcoholic fatty liver disease (NAFLD), metabolic dysfunction-associated steatotic liver disease (MASLD), liver fibrosis, cirrhosis, liver cancer, and other forms of chronic liver disease. The terms “nonalcoholic steatohepatitis (NASH)” and “metabolic dysfunction-associated steatohepatitis (MASH)” may be used interchangeably herein. Alternatively or in addition to, the subject may have one or more of other diseases, physical conditions, and lifestyles that increase the risk of a liver disease, including heavy alcohol use, blood transfusion, obesity, diabetes, polycystic ovary syndrome, and metabolic syndrome.

[0095] The bodily sample may generally refer to, for example, a sample obtained or derived from the body of the subject through any suitable means of collection. The bodily sample may contain nucleic acid from which methylation patterns may be determined. The bodily sample may be obtained from an organ or system of interest (e.g., liver, spleen, kidney, bile duct, breast, prostate, lymph node). The bodily sample may be from a tissue or organ where components of which were originally formed or where components metastasized to after originating in a different tissue or organ (e.g., lung, pancreas, intestine, spleen, prostate, or cardiovascular). The bodily sample may comprise excreta (e.g., urine and stool) of the patient. In some embodiments, the bodily sample may be selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a tissue sample, and a lymphatic fluid sample. A solid or semi-solid tissue sample may be processed and liquefied.

[0096] The bodily sample may be processed prior to methylation analysis. In some embodiments, nucleic acid may be isolated from the bodily sample. In other embodiments, a circulating material may be extracted from the bodily sample. For example, when the bodily sample comprises blood, plasma may be isolated from the blood, and cell-free DNA (e.g., circulating cell-free DNA) may be extracted from the plasma.

[0097] The nucleic acid isolated from the bodily sample may be assayed for methylation. The methylation measurement may be implemented using one or more of bisulfite sequencing (e.g., post-bisulfate adapter-tagging, PBAT), targeted methylation sequencing (e.g., targeted bisulfite sequencing or methyl sequencing), pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, methylation arrays, or methylation specific polymerase chain reaction (PCR). In bisulfite sequencing, for example, bisulfite treatment of nucleic acid may leave methylated cytosine (e.g., 5mC) unaffected, while non-methylatedcytosine (C) may be converted into uracil (U). Subsequent polymerase chain reaction (PCR) amplification may convert uracil (U) into thymine (T). During the amplification process, primers for loci of interest may be barcoded for downstream analysis. The nucleic acid sample may further be sequenced to identify methylated cytosines (5mC) at single base resolution. The sequencing output may allow quantitative analysis, such as the methylation percent of target oligonucleotide sequences and copies per mL (e.g., of blood sample, plasma sample, or cfDNA sample) of target differentially methylated regions (DMRs), as well as methylation patterns (e.g., two or more methylated and / or unmethylated cytosine residues) over DNA sequence blocks.

[0098] In some embodiments, reads (e.g., sequences of base pairs) may be aligned to target sequences and filtered for quality scores. A quality score in a next-generation sequencing (NGS) process may refer to a probability that a given base is called incorrectly by the sequencer. Genomic regions may contain differentially methylated status. A quality score threshold may be imposed on the reads (e.g., within 10% tolerance of similarity to target sequence).

[0099] The methylation analysis may comprise determination of degree of methylation, for example, a percentage of cytosine (C) residues in a nucleic acid that have methyl groups. In some embodiments, the methylation status may be determined as the beta value of the nucleic acid sample. A beta value may generally refer to the percent methylation of the nucleic acid sample reported on a scale between 0-100%. Beta values greater than 50% may indicate a methylated state and thus, the genomic region is silenced, whereas beta values less than 50% may indicate a unmethylated state and thus, the genomic region is activated.

[0100] A plurality of bioinformatics pipelines may be applied to analyze methylation status and patterns. In some embodiments, Bismark may be used following bisulfite sequencing (as described by, for example, Felix Krueger and Simon R. Andrews, Bismark: a flexible aligner and methylation caller for Bisulfite-Seq applications, Bioinformatics, 27(11): 1571-1572 (2011), which is incorporated by reference herein). Reads from bisulfite sequencing may be converted into C-to-T and G-to-A (equivalent to C-to-T conversion on the reverse strand) versions. Each read may be aligned to equivalently pre-converted forms of a reference genome using four parallel instances. In particular, C-to-T read conversion may be aligned with forward strand C-to-T converted reference genome region and forward strand G-to-A converted reference genome region. G-to-A read conversion may be aligned with forward strand C-to-T converted reference genome region and forward strand G-to-A converted reference genome region. The methylation state of positions involving cytosines may bedetermined by comparing the read sequence with the corresponding reference genomic sequence.

[0101] In some embodiments, the quantitative measures of liver-specific methylated biomarkers may be determined as a ratio between a number of methylated sequence reads in a sequence block and a total number of sequence reads within the sequence block, referring to the count-based approach. The size of the sequence block may be between about 50 bp and 300 bp, 100 bp and 300 bp, 100 bp and 200 bp. The determination of methylated count total count may minimize potential bias toward the low coverage reads in the datasets.

[0102] In other embodiments, the quantitative measures of liver-specific methylated biomarkers may be determined as a mean methylation level of a sequence block, referring to the ratio-based approach. For example, the methylation level of each nucleotide in a sequence block may be determined. The mean methylation level may be determined using methylation levels of a plurality of nucleotide sites.

[0103] The quantitative measures of liver-specific methylated biomarkers may be determined as at least one of the count-based approach or ratio-based approach. In some embodiments, both the count-based and ratio-based approaches may be used to generate a methylation matrix. The rows of the matrix may indicate binned block sites, for example, chrl: 1,000, 000. The columns of the matrix may indicate bodily samples collected from a variety of patients. The bodily samples may comprise bodily fluid samples and solid tissue samples. The bodily fluid samples may comprise blood samples. The solid tissue samples may comprise liver tissue samples and non-liver tissue samples.

[0104] Liver-specific methylated biomarkers (e.g., developed or identified biomarkers) may be detected based on one or more detection criteria. The criteria may comprise a difference between a mean methylation level of a liver-specific methylated biomarker and / or sequence block detected in a liver tissue sample and a mean methylation level of the biomarker and / or sequence block detected in a blood sample, which may be at least 10%, 20%, 30%, 40%, or 50%. The blood sample may comprise peripheral blood mononuclear cells (PBMCs) from which cellular genomic DNA may be extracted and assayed for methylation. In some embodiments, a mean methylation level of a liver-specific hypermethylated biomarker detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50% higher than a mean methylation level of the biomarker detected in a blood sample. In other embodiments, a mean methylation level of a liver-specific hypomethylated biomarker detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50% lower than a mean methylation level of the biomarker detected in a blood sample.

[0105] The criteria may comprise a difference between a mean methylation level of a liverspecific biomarker detected in a liver tissue sample and a mean methylation level of the biomarker detected in a non-liver tissue sample may be at least 10%, 20%, 30%, 40%, 50%. In some embodiments, a mean methylation level of a liver-specific hypermethylated biomarker detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, 50% higher than a mean methylation level of the biomarker detected in a non-liver tissue sample. In other embodiments, a mean methylation level of a liver-specific hypomethylated biomarker detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, 50% lower than a mean methylation level of the biomarker detected in a non-liver tissue sample.

[0106] In some embodiments, the mean p-value and false discovery rate (FDR) of the methylated biomarkers may be calculated. Accordingly, the criteria of determining liverspecific methylated biomarkers may further comprise pre-determined thresholds for p-value and / or FDR. For example, the pre-determined threshold of p-value may be 0.05, 0.04, 0.03, 0.02, 0.01. Only those sequence blocks that have a mean p-value below the pre-determined threshold may be selected. Alternatively or in addition to, the pre-determined threshold for FDR may be 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04. Only those sequence blocks that have an FDR below the pre-determined threshold may be selected.

[0107] The criteria may comprise the number of overlapping samples between liver tissue samples and non-liver samples (e.g., blood samples and non-liver tissue samples). In some embodiments, biomarkers with less than or equal to 2 overlapping sites may be considered as liver specific methylated biomarkers. The term “overlapping,” as used herein, generally refers to individual samples whose percent methylation at a given site lies within the range of percent methylation seen for that site in another tissue.

[0108] In some embodiments, the liver-specific methylated biomarkers may comprise at least one of a hypermethylated biomarker or a hypomethylated biomarker. In other embodiments, the liver-specific methylated biomarkers may comprise a plurality of hypermethylated biomarkers and hypomethylated biomarkers. For example, the liver-specific methylated biomarkers may comprise a member selected from the group listed in Table 1.

[0109] One aspect described herein detects a liver-specific hypermethylated or hypomethylated cytosine-phosphate-guanine (CpG) sites. In some embodiments, at least one of the count-based approach or ratio-based approach may be used to determine quantitative measures of a set of liver-specific hypermethylated or hypomethylated CpG sites in the nucleic acid from a bodily sample. For example, both the count-based and ratio-based approaches may be used to generate a methylation matrix. The rows of the matrix mayindicate binned sequence block sites, for example, chrl : 1,000,000. The columns of the matrix may indicate bodily samples collected from a variety of patients, including liver tissue samples, non-liver tissue samples, and blood samples.

[0110] The liver- specific hypermethylated or hypomethylated CpG sites maybe detected following one or more detection criteria. The criteria may comprise a difference between a mean methylation level of a liver-specific hypermethylated or hypomethylated CpG site detected in a liver tissue sample and a mean methylation level of the CpG site detected in a blood sample may be at least 10%, 20%, 30%, 40%, or 50%. In some embodiments, a mean methylation level of a liver-specific hypermethylated CpG site detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, 50% higher than a mean methylation level of the CpG site detected in a blood sample. In other embodiments, a mean methylation level of a liver-specific hypomethylated CpG site detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50% lower than a mean methylation level of the CpG site detected in a blood sample.

[0111] The criteria may comprise a difference between a mean methylation level of a liverspecific hypermethylated or hypomethylated CpG site detected in a liver tissue sample and a mean methylation level of the CpG site detected in a non-liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50%. In some embodiments, a mean methylation level of a CpG site detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50% higher than a mean methylation level of the CpG site detected in a non-liver tissue sample. In other embodiments, a mean methylation level of a CpG site detected in a liver tissue sample may be at least 10%, 20%, 30%, 40%, or 50% lower than a mean methylation level of the CpG site detected in a non-liver tissue sample.

[0112] In some embodiments, the mean p- values and FDRs of the CpG sites may be calculated. Accordingly, the criteria of identifying liver- specific hypermethylated or hypomethylated CpG sites may further comprise pre-determined thresholds for p- value and / or FDR. For example, the pre-determined threshold of p-value may be 0.05, 0.04, 0.03, 0.02, 0.01. Only those CpG sites that have a mean p-value below the pre-determined threshold may be selected. Alternatively or in addition to, the pre-determined threshold for FDR may be 0.08, 0.07, 0.06, 0.05, 0.04. Only those CpG sites that have an FDR below the pre-determined threshold may be selected.

[0113] The criteria may further comprise pre-determined thresholds for deviations between patient samples. In some embodiments, the standard deviations thresholds for methylation levels of bodily samples may be about 20%, 18%, 15%, 12%, 10%, 8%, 5%. Only thosesamples that have a standard deviation lower than the pre-determined threshold may be selected.

[0114] In some embodiments, the liver-specific hypermethylated or hypomethylated CpG sites may comprise at least one of a hypermethylated CpG site or a hypomethylated CpG site. In other embodiments, the liver-specific hypermethylated or hypomethylated CpG sites may comprise a plurality of hyper-methylated CpG sites and hypo-methylated CpG sites. For example, the liver-specific methylated biomarkers may comprise a member selected from the group listed in Table 1.

[0115] The detection of presence of a liver disease in a subject based on detecting liverspecific methylated biomarkers and hypomethylated or hypomethylated CpG sites may allow the liver disease to be diagnosed at an early stage. Based on the presence of the liver disease, the patient may receive proper and effective treatment, including lifestyle intervention, pharmacologic therapy, and surgical intervention. For example, when diagnosed with steatosis, the patient may receive lifestyle intervention including weight control, alcohol control, and regular exercise. When diagnosed with some other liver disease or liver cancer, the patient may receive pharmacologic therapy and / or surgical intervention.

[0116] The detection of presence of a liver disease in a subject may also allow the determination of efficacy of the treatment administered to the patient. In some embodiments, the method may provide quantitative measures of a set of liver-specific methylated biomarkers and / or hypermethylated or hypomethylated CpG sites to determine whether a liver disease is present in the patient. The quantitative measures of liver-specific methylated biomarkers before and after the treatment may allow health care providers to determine the efficacy of the treatment and whether there is a resistance to the treatment.

[0117] Another aspect described herein provides a method for detecting a presence or an absence of nonalcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH) in a subject. In some embodiments, the method may comprise assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of NASH / MASH-specific methylated biomarkers in a nucleic acid from the bodily sample, computer processing the set of NASH / MASH-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference, and detecting the presence or the absence of the NASH / MASH in the subject, based at least in part on the computer processing.

[0118] In some embodiments, at least one of the count-based approaches or ratio-based approaches may be used to determine quantitative measures of NASH / MASH-specificmethylated biomarkers. For example, both the count-based and ratio-based approaches may be used to generate a methylation matrix. The rows of the matrix may indicate binned block sites, for example, chrl : 1,000,000. The columns of the matrix may indicate bodily samples collected from a variety of patients, including liver tissue samples, non-liver tissue samples, and blood samples. In some embodiments, one or more control samples may also be assayed, including bodily samples obtained or derived from patients who are diagnosed with, for example, simple steatosis.

[0119] NASH / MASH-specific methylated biomarkers may be detected following one or more detection criteria. The criteria may comprise a mean methylation status (e.g., average beta value) of a bodily sample obtained from a patient diagnosed with NASH / MASH that may be at least about 10%, 20%, 30%, 40%, 50% different from that of a control sample. In some embodiments, the mean methylation status of a NASH / MASH-specific hypermethylated biomarker in patient samples may be at least about 10%, 20%, 30%, 40%, 50% higher than that in control samples. In other embodiments, the mean methylation level of a NASH / MASH-specific hypomethylated biomarker in patient samples may be at least about 10%, 20%, 30%, 40%, 50% lower than that in control samples.

[0120] The criteria may further comprise a pre-determined threshold of the number of hypermethylated or hypomethylated CpG sites in a sequence block. In some embodiments, the pre-determined threshold for the number of hypermethylated or hypomethylated CpG sites in a sequence block per group (e.g., sample group, control group) may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17 ,18, 19, or 20. A sequence block with the number of hypermethylated or hypomethylated CpG sites larger than the pre-determined threshold may be selected. In other embodiments, the criteria may comprise a pre-determined threshold range of the number of hypermethylated or hypomethylated CpG sites in a sequence block. For example, the pre-determined threshold range of the number of hypermethylated or hypomethylated CpG sites in a sequence block may be 2-20, 2-18, 2-15, 2-12, 2-10, 2-7, 3- 17, 3-14, 3-10, 3-7, 4-13, 4-11, 4-9, 5-12, 5-10, 5-8, 6-11, 6-9, or 7-10. A sequence block with the number of hypermethylated or hypomethylated CpG sites within the pre-determined threshold range may be selected.

[0121] The criteria may further comprise pre-determined thresholds for deviations between patient samples. In some embodiments, the standard deviation thresholds for methylation levels of bodily samples may be about 20%, 18%, 15%, 12%, 10%, 8%, 5%. Only those samples that have a standard deviation lower than the pre-determined threshold may be selected. For example, the standard deviation of bodily samples obtained from patients withNASH / MASH may be lower than 10%, and the standard deviation of control samples may also be lower than 10%.

[0122] The NASH / MASH-specific methylated biomarkers may comprise at least one of a hypermethylated biomarker or a hypomethylated biomarker. In some embodiments, the NASH / MASH-specific methylated biomarkers may comprise NPPA-AS1, PRDM16-DT, HIF1AN, DNAH10, ALG10B, KIF7, SHISA9, BAIAP3, RBFOX3, LOC644669, TPRX1, A1BG, ADGRL1, CLEC4GP1, CREG2, GBX2, LOC101929452, TLX2, HSPA12B, LOC730668, SETD7, KCNN2, UNC5A, LINC02119, LINC01449, FIGNL1, ERICH1, OLFM1, or TMEM187.

[0123] In some embodiments, the quantitative measures of NASH / MASH-specific methylated biomarkers may be determined as at least one of a first combined beta value of at least two NASH / MASH-specific hypermethylated biomarkers or a second combined beta value of at least two NASH / MASH-specific hypomethylated biomarkers. For example, at least two NASH / MASH-specific hypermethylated biomarkers may be selected from PRDM16-DT, HIF1AN, LOC644669, TPRX1, A1BG, ADGRL1, CLEC4GP1, CREG2, LOC101929452, TLX2, LOC730668, KCNN2, UNC5A, LINC02119, FIGNL1, and TMEM187. Or at least two NASH / MASH-specific hypomethylated biomarkers may be selected from NPPA-AS1, DNAH10, ALG10B, KIF7, SHISA9, BA1AP3, RBFOX3, GBX2, HSPA12B, SETD7, LINC01449, ERICH1, andOLFMl.

[0124] Accordingly, the reference may comprise a first reference beta value of the at least two NASH / MASH-specific hypermethylated biomarkers and a second reference beta value of the at least two NASH / MASH-specific hypomethylated biomarkers. The first and second reference beta values may be determined by assaying control samples. The first combined beta value may be at least 10%, 12%, 15%, 18%, 20%, 22%, 25%, 28%, 30%, 32%, 35%, 38%, 40%, 42%, 45%, 48%, or 50% higher than the first reference beta value, thereby determining the presence of NASH / MASH in the patients from which the bodily samples are obtained. The second combined beta value may be at least 10%, 12%, 15%, 18%, 20%, 22%, 25%, 28%, 30%, 32%, 35%, 38%, 40%, 42%, 45%, 48%, or 50% lower than the second reference beta value, thereby determining the presence of NASH / MASH in the patient from which the bodily samples are obtained.

[0125] FIG. 1 illustrates an example workflow 100 of determining the presence of liver disease and NASH / MASH in a subject based on liver-specific and NASH / MASH-specific methylated biomarkers, in accordance with some embodiments of the present disclosure. A bodily sample is first obtained or derived from a patient to be diagnosed and nucleic acidisolated from the bodily sample (see operations 110 and 120). The isolated nucleic acid is assayed to determine the methylation status (see operation 130). For example, the quantitative measures of the methylation status of liver-specific methylated biomarkers and NASH / MASH-specific methylated biomarkers may be determined (see operations 142 and 144). The detected methylation status of the liver-specific methylated biomarkers are used to determine whether a liver disease is present in the patient, and the NASH / MASH-specific methylated biomarkers to determine whether NASH / MASH is present (see operations 152 and 154).

[0126] FIG. 2 illustrates an example workflow 200 of detecting methylated biomarkers and CpG sites, in accordance with some embodiments of the present disclosure. Bodily samples from patients to be diagnosed (e.g., test samples) and reference or control samples may be obtained, from which nucleic acid may be isolated and analyzed for methylation status (e.g., see operations 110-130 of FIG. 1). For example, the test sample may be a liver tissue sample, whereas a control sample may be a non-liver tissue sample or blood sample. Following the methylation assay (e.g., bisulfite sequencing, enzymatic methylation sequencing or nanopore sequencing), the generated sample and control datasets are normally distributed with a p- value smaller than or equal to, for example, 0.01 (operation 210). For each sequence block, a mean p-value and false discovery rate (FDR) are calculated (operation 220). In some embodiments, pre-determined thresholds for p-values and FDRs may be used to select sequence blocks for further processing. The pre-determined thresholds may vary depending on the types of methylated biomarkers and hypermethylated or hypomethylated CpG sites. For example, to detect liver-specific methylated biomarkers, the pre-determined thresholds for p-value and FDR of sequence blocks may be 0.01 and 0.1, respectively. To detect liverspecific hypermethylated or hypomethylated CpG sites, the pre-determined thresholds for p- value and FDR may be 0.01 and 0.05, respectively. The methylation status of each sequence block in the test samples may be compared with control samples (see operation 230). For example, the methylation status of each sequence block in the test liver tissue sample may be compared with control dataset derived from blood samples and non-liver tissue samples, respectively. For the detection of liver-specific methylated biomarkers, the number of overlapping samples between control samples (e.g., blood samples and non-liver tissue samples) and test liver tissue samples may be determined (operation 240). Only those biomarkers with less than or equal to 2 overlapping sites may be considered as liver specific methylated biomarkers.

[0127] Methylated biomarkers and methylated CpG sites can be detected by comparing the sample dataset and control dataset (operation 250). In some embodiments, a difference between a mean methylation level of a liver-specific biomarker detected in a liver tissue sample and a mean methylation level of the biomarker detected in a blood sample may be at least 40%. In other embodiments, a difference between a mean methylation level of a liverspecific biomarker detected in a liver tissue sample and a mean methylation level of the biomarker detected in a non-liver tissue sample may be at least 20%. Table 1 lists 186 liverspecific methylated biomarkers, including 36 hypermethylated biomarkers and 150 hypomethylated biomarkers.

[0128] The following liver-specific biomarkers are included in Table 1: ABCG8, ACACB, ACACB, ACADS, ADGRF3, AGBL2, AGXT, ALB, ALB, ALB, AMFR, ANKAR, AOC4P, AOC4P, APOA4, APOA4, AQP9, B4GALNT4, BDH1, BDH1, BDH1, C16orf70, CA5A, CA5A, CCL16, CD28, CLDN14, COQ8A, CSAD, CYP7A1, DERA, DHODH, DIPK1A, EP300-AS1, EP78, ERMARD, EXOC3L4, FAAP20, FAF2, FAM149A, FAM20C, FAM71D, FANCE, FASN, FASN, FGA, FNDC3B, FRZB, FSD1L, FTCD-AS1, FTCD- AS1, FTCD-AS1, FUOM, GALR3, GAP43, GGT1, GLB1, GLS2, GOLGA6L9, GPAM, GPAM, GPAM, GPAM, GPR21, GPR21, GPR88, GRIK2, GRK5, GSTZ1, HACD3, HAL,HAL, HPD, IRF7, ITIH1, ITIH4, ITIH4-AS1, JAK1, KCNS1, KDM2B, KLF6, KNSTRN, KYAT1, LARP4, LINC00167, LINC00933, LINC01213, LINC01700, LOC101928093, LOCI 02724265, LOCI 02724467, LOC343052, LRRC19, LRRC4C, MAP7D1, MATN2, MBL2, MIR1322, MIR3138, MIR4266, MIR4677, MIR4783, MIR576, MIR576, MIR576, MIR6815, MST1P2, MYEOV, NADK, NEK4, NID 1 , NOTCH2NLB, NPFFR1 , NR3C1, OIP5-AS1, PANXI, PCSK6, PCSK6-AS1, PCSK9, PDS5A, PECR, PEPD, PFN3, PFN3, PGBD4, PGBD5, PKLR, PLTP, PRKY, PRR5, RABGAP1L, RANBP3L, RBFOX2,RBFOX2, RBMS3, RET, RNASE4, RNF113B, RNF44, RPL31, RRP1B, RTN4, SECISBP2, SERPINA1, SERPINA10, SERPINA6, SERPINF2, SH2B2, SIM1, SLC10A1, SLC16A14, SLC25A33, SLC2A9, SLC2A9, SLC2A9, SLC6A12, SMIM14, SNORA50A, SNORA59B, SNORD168, SNORD168, SNORD98, SNRNP200, SPTAN1, STAT5B, SVIL-AS1, TBC1D14, TEF, TFR2, TFR2, TFR2, TFR2, TFR2, TFR2, TOR3A, TUT7, UBP1, UGDH, VASN, VSTM2L, WWC1, ZC3H6, ZFAND2A, ZFYVE1, ZHX2, and ZYGI IB.

[0129] Table 1. 186 liver-specific methylated biomarkers, including 36 hypermethylated biomarkers and 150 hypomethylated biomarkers.

[0130] Table 1 provides exemplary liver-specific methylated biomarkers. Differential methylation regions (DMRs) are reported in average beta value (3) which represents a log ratio of percent DMR. Annotations indicate chromosomal region, direction as hypermethylated (hyper) or hypomethylated (hypo), nearest associated gene, number of overlaps between liver tissue and other tissue or number of overlaps between liver tissue and peripheral blood mononucleates (PBMC). Mean methylation (average beta value p) in liver,Blood and Other Tissues of the same DMR are reported. P-value, false discovery rate (FDR), associated gene, genomic annotation and oligonucleotide sequence are also reported.Duplicated “associated genes,” for example BDH1, are listed multiple times, indicating that there have multiple distinct hypermethylated or hypomethylated regions associated with the same gene, but multiple distinct targets can be used. Candidate regions are derived from 100 base pair (bp) binning sequences; however, larger regions can be assessed computationally using by human genome coordinates reported.

[0131] In some embodiments, a difference between a mean methylation level of a CpG site detected in a liver tissue sample and a mean methylation level of the CpG site detected in a blood sample may be at least 40%. In other embodiments, a difference between a meanmethylation level of a CpG site detected in a liver tissue sample and a mean methylation level of the CpG site detected in a non-liver tissue sample may be at least 20%.

[0132] NASH / MASH-specific methylated biomarkers can also be detected by comparing the sample dataset and control dataset. In some embodiments, a difference between a mean methylation level of a NASH / MASH-specific biomarker detected in a liver tissue sample obtained from a patient with NASH / MASH and a mean methylation level of the biomarker in a control sample from a patient with simple steatosis may be at least 20%. A number of CpG sites in a sequence block may be more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 per group (e.g., sample group and control group). Table 2 lists 583 NASH / MASH-specific methylated biomarkers, including 322 hypermethylated biomarkers and 261 hypomethylated biomarkers. The last column of Table 2 lists genomic regions in which the methylation status (e.g., beta value) is associated with the presence of NASH / MASH. These regions are found in cell-free DNA and can be used to distinguish NASH / MASH from simple steatosis. Chromosome LD reflections are found in diverse areas including but not limited to introns, exons, promotors, transcription start sites and inter-genic regions. Average beta difference indicates the difference in the average methylation status between test samples and control samples. The differences change in either direction (e.g., hypermethylated or hypomethylated). For example, some of the NASH / MASH-specified methylated biomarkers are hypermethylated in disease state, where the methylation levels of these markers are higher in test samples. Other NASH / MASH-specified methylated biomarkers are hypomethylated in disease state, where the methylation levels of these markers are lower in test samples.

[0133] The following NASH / MASH-specific biomarkers are included in Table 2: MARCH9, A1BG, ABCA4, ABL2, ABLIM2, ABRAXAS 1, ACADL, ACTR3B, ACTR3BP5, ACTRT1, ADAMTS6, AD API, ADCY2, ADGRL1, AFF2, AJAP1, ALG10, ALG10B, ALK, AMZ1, ANKRD1 1, ANO7, AP3D1, APRT, ARAP2, ARHGAP45, ARHGEF18, ARMCX2, ASNSP1, ASPG, ATP11C, ATP4A, ATP5PF, BAIAP3, BARX1, BARX2, BEST2, BEX1, BICRAL, BMP8A, BRINP3, BRSK2, BYSL, C12orf45, C15orf40, Clorfl94, C5orf49, C9orf62, C9orf92, CACNA1C-IT3, CADPS2, CALML4, CALN1, CASK, CASP16P, CBX2, CCDC130, CCDC187, CDC123, CDC42EP4, CDH12, CDH4, CDH8, CEACAM8, CENPS- CORT, CHIT1, CHKA, CHMP1A, CHST7, CITED1, CLEC4GP1, CMKLR1, CMTM7, CNFN, COL6A2, CREB3L3, CREG2, CRMP1, CSNK1G1, CTAGE1, CTRC, CYP19A1, DBN1, DDX39A, DEFB122, DES, DHCR24, DIAPH3, DMBT1, DMD, DNAH10, DNPH1, DOCKS, DOK7, DPY19L2P4, DPY19L4, DSCAML1, DT, DTNA, DUOX2, DUSP21, EBF4, ECHS1, EFCAB6, EFNA2, EGFL7, EGFR-AS1, ELFN2, ELOVL6, EMBP1, EMX2,ENPP6, EPHA3, ERICH1, FA2H, FAM169A, FAM173A, FAM222A, FAM53A,FAM74A1, FAM86C1, FAM90A25P, FBXO42, FBXO7, FBXW7-AS1, FCGR2A, FFAR3,FGF12-AS2, FGFR3, FHIT, FIGNL1, FLJ33534, FLNA, FOXB1, FOXD1, FOXQ1, FRG2KP, FRMD4A, FRMD4B, FRMPD4, FSCB, FUCA1, FUNDCI, FUT9, GACAT3, GADL1, GALR1, GAMT, GBX2, GDF6, GDI1, GDNF-AS1, GMFG, GNA12, GNAS, GNPDA2, GOLGA6B, GPR4, GPRC5B, GPX4, GRIA3, GRIFIN, GRIN3B, GRM4,GTF2IRD2B, GUCY1B1, HCFC1, HCG2040054, HCK, HCN2, HDHD5, HERC2P4,HIF1AN, HIP1, HMG20B, HOXB13, HOXB6, HSPA12B, IBA57, IGLON5, IGSF10,IGSF21, IL2, IL27RA, IL2RG, INF2, INSL3, INSRR, IRF2BP1, IRX1, KANK3, KCNG2, KCNK9, KCNN2, KIAA1324, KIAA1522, KIF17, KIF4B, KIF7, KIRREL3-AS2, KLF2,KLF5, KLHL13, KLRG2, KRT16P3, L3MBTL4, LAMP1, LASPI, LAT, LDC1P, LDOC1,LHPP, LIG4, LINC00051, LINC00364, LINC00376, LINC00460, LINC00470, LINC00487, LINC00673, LINC00922, LINC00951, LINC01017, LINC01029, LINGO 1046, LINC01122, LINC01137, LINC01164, LINC01168, LINC01192, LINC01234, LINC01370, LINC01440, LINGO 1449, LINGO 1532, LINC01587, LINC01644, LINGO 1667, LINC01668, LINGO 1677, LINC01683, LINC01740, LINC01756, LINC01854, LINC01923, LINC01935, LINC01987,LINC02004, LINC02050, LINC02119, LINC02145, LINC02174, LINC02302, LINC02322,LINC02383, LINC02450, LINC02492, LINC02533, LINC02611, L1NC02645, LINC02702,LINC02737, LMLN, LOCI 00129697, LOCI 00130370, LOCI 00272217, LOC100287010, LOG 100287387, LOCI 00506178, LOCI 00506585, LOCI 00507373, LOCI 00996351,LOC100996654, LOC101927138, LOC101927394, LOC101927691, LOC101927815,LOC101927884, LOC101927948, LOC101928303, LOC101928358, LOC101928381,LOC101928932, LOC101929058, LOC101929452, LOC102723883, LOC102724152,LOC105371592, LOC105374313, LOC105377448, LOC105378047, LOC105378183,LOC105378311, LOC105379152, LOC107984282, LOC150935, LOC284933, LOC644669,LOO729968, LOC730668, LOXL3, LPP-AS1, LRRC38, LRRC47, LSR, MAN1C1,MAT2B, MBNL3, MEAK7, MELTF, MFAP2, MFAP3L, MGC12916, MGLL, MICALL2, MIDI, MINDY3, MIR1184-2, MIR196A1, MIR3667, MIR4251, MIR4283-1, MIR4454, MIR4486, MIR4501, MIR4531, MIR4648, MIR4666B, MIR5189, MIR548AQ, MIR548AU, MIR574, MIR592, MIR604, MIR6819, MIR759, MIR935, MIR9901, MLLT10P1, MOB3A, MORC4, MS4A6E, MTOR-AS1, MTUS2, MUC5AC, MYH8, NA, NAA10, NACC2, NANOS2, NCOA4, NEIL2, NEURL1-AS1, NFIA, NFIX, NKRF, NKX1-1, NKX2-5,NKX2-8, NME3, NMI, NOC2LP2, NOL4L, NPIPB9, NPPA-AS1, NPY5R, NRL, NTMT1, NUB1, NUPR2, NXT2, OCM, OLFM1, ONECUT3, OPCML, OPRM1, OR11H12, OR1F12,OR4C46, OR4X1, OR7D2, OTUB1, P3H2, PABPC1P2, PABPC5, PCDH15, PCDHA10,PEX14, PHF20L1, PHKA1, PHKA2, PID1, PIGBOS1, PKP3, PLAC9, PLCD1, PLEC,PLEKHA8P1, PLIN4, PMEPA 1, POTEA, PPP1R3B, PPP1R3D, PPP2R2A, PRDM16,PREP, PRF1, PROB1, PROCR, PROS1, PRR20D, PSD3, PTCHD1, PTGDS, PTGIR,PTPN4, PTPRD-AS1, QRFP, RAB39B, RABL2A, RALY-AS1, RALYL, RASA3,RBAKDN, RBFOX3, RBM6, REPS2, RIMS3, RIN2, RIPOR2, RNA18SN2, RNA28SN2,RNA5-8SN4, RND2, RNVU1-19, RNVU1-3, RPL19, RSPRY1, SARAF, SBK2, SCARF2,SCFD2, SENP5, SEPTIN9, SETD7, SETD9, SFTA1P, SH2B2, SH3GL1P2, SH3PXD2B,SHC2, SHISA9, SIM2, SIN3A, SLC22A11, SLC25A26, SLC25A43, SLC25A53, SLC29A3,SLC2A3, SLC2A4RG, SLC37A1, SLC6A11, SLC6A16, SLC6A1-AS1, SLF2, SLIT2,SMAD7, SMIM10, SNHG20, SNORA15B-1, SNORD148, SNTG1, SNX29, SP3, SPATA3-AS1, SPEN. SPHK2, SPIB, SPIN4, SPRED2. SPRY4-AS1, SPTBN2, SPTSSA, STS,SULT2B1, SULT4A1, SUMF1, SYTL4, TAC4, TACC2, TAFA5, TCEA3, TCF3, TCL1A,TEX28, TGM3, THEG, TIMP1, TKTL1, TLE3, TLNRD1, TLX2, TMEM102, TMEM161B-AS1, TMEM187, TMEM189-UBE2V1, TMEM250, TMEM38B, TMEM81, TMEM88,TNRC17, TOMM40, TOMM6, TOR4A, TP53TG3HP, TPM 1 -AS, TPO, TPRX1,TRAPPC10, TRH, TRIP10, TRPC5, TSC22D3, TTC34, UACA, UBASH3A, UHMK1,ULBP1, UNCI 19B, UNC5A, UNC93B1, UPK3A, USP17L2, USP6, USP9X, VAV2,VIPR2, VPS53, VWA2, WBP1L, WNT3A, WWC3, XAF1, ZBTB8OS. ZIM3, ZMYND12.ZNF185, ZNF436, ZNF454, ZNF469, ZNF587, ZNF608, ZNF628, ZNF692, ZNF716,ZNF730, ZNF750, ZNF839, ZNRF1, and ZP3.

[0134] Table 2. NASH / MASH-specific methylated biomarkers.

[0135] FIG. 3 illustrates a heatmap (generated from open-source R heatmap package) of the methylation level of selected liver-specific sequence blocks (N = 477). The x-axis lists the bodily samples that are processed, including liver tissue samples and blood samples. The y- axis lists selected DMR blocks (100 bp in size). The row z-value of these sequence blocks corresponding to liver tissue samples is -2, as opposed to the row z-value of close to 1 corresponding to blood samples. It indicates these sequence blocks represent liver-specific hypomethylated biomarkers. The methylation level of these biomarkers is significantly lower in liver tissue samples compared to blood samples. These selected liver- specific hypomethylation sequence blocks are commonly detected by the method as described herein and a reference method (N. Loyfer et aL, A DNA Methylation Atlas for Normal Human CellTypes, Nature, 613, 355-364 (2023)).

[0136] FIG. 4 illustrates a heatmap of the methylation level of selected liver-specific sequence blocks (N = 2,642). Similar to FIG. 3, the x-axis lists the bodily samples that are processed, including liver tissue samples and blood samples. The y-axis lists selected DMR blocks (100 bp in size). The row z- value of these blocks corresponding to liver tissue samples is -2, as opposed to the row z-value of close to 1 corresponding to blood samples. It indicates these blocks represent liver-specific hypomethylated biomarkers. The expression of these biomarkers is significantly lower in liver tissue samples compared to blood samples.

[0137] FIG. 5 illustrates a heatmap of the methylation level of selected liver-specific sequence blocks (N = 243). The x-axis lists the bodily samples that are processed, including liver tissue samples and blood samples. The y-axis lists selected DMR blocks (100 bp in size). The row z-value of these sequence blocks corresponding to liver tissue samples is close to 2, as opposed to the row z-value of close to -1 corresponding to blood samples. It indicates these sequence blocks represent liver-specific hypermethylated biomarkers. The expression of these biomarkers is significantly higher in liver tissue samples compared to blood samples.

[0138] FIG. 6 illustrates a histogram of NASH / MASH versus simple steatosis using whole genome bisulfite sequencing (WGBS) assessment. Blood samples are collected from 24 patients, of which 12 patients are diagnosed with simple steatosis (“control”) and 12 patients are diagnosed with NASH / MASH (“disease”)- The methylation analysis is performed via WGBS. The x-axis of FIG. 6 indicates a difference in methylation level between bodily samples from patients with NASH / MASH and from patients with simple steatosis. For example, a positive difference may indicate the methylation level of a corresponding biomarker in a NASH / MASH patient sample is higher than that in a simple steatosis patient sample, i.e., hypermethylated. A negative difference may indicate the methylation level of a corresponding biomarker in a NASH / MASH patient sample is lower than that in a simple steatosis patient sample, i.e., hypomethylated. The y-axis indicates a frequency of NASH / MASH-specific methylated biomarkers found to have different levels in NASH / MASH patients and in simple steatosis patients. As illustrated, a total of 583 NASH / MASH -specific methylated biomarkers are detected that can distinguish a disease state (NASH) from a control state (e.g., simple steatosis).

[0139] FIG. 7A-7D illustrate box plots of beta values of hypermethylated biomarkers that distinguish NASH / MASH from simple steatosis. The biomarkers are detected using cell-free DNA isolated from patient blood samples. Each box plot illustrates an average beta value comparison of a hypermethylated biomarker in bodily samples isolated from patients with NASH / MASH and simple steatosis (control) samples. As illustrated, the methylation level ofthese hypermethylation regions are significantly higher in NASH / MASH patient samples, indicating these regions are NASH / MASH-specific.

[0140] FIG. 8A and 8B illustrate box plots of beta values of hypomethylated biomarkers that distinguish patients diagnosed with NASH / MASH from patients diagnosed with simple steatosis. Each box plot illustrates an average beta value comparison of a hypomethylation region in patient blood samples and control samples. As illustrated, the methylation level of these hypomethylation regions are significantly lower in NASH / MASH patient samples, indicating these regions are NASH / MASH-specific.

[0141] In some embodiments, instead of using a single methylated biomarker or sequence block, a combination of methylated biomarkers or sequence blocks may be used to detect a liver disease and distinguish one liver disease from another. FIG. 9 A illustrates box plots of combined hypermethylated biomarkers that distinguish NASH / MASH from simple steatosis. FIG. 9B illustrates box plots of combined hypomethylated biomarkers that distinguish NASH / MASH from simple steatosis. Combining methylated biomarkers may yield a set of biomarkers that when combined, show zero overlap between disease state (e.g., NASH / MASH) and control state. The combined methylated biomarkers markers have between 90%-99% accuracy from initial pilot.EXAMPLES

[0142] Example 1 - Detection of NASH / MASH-Speeific Methylated biomarkers

[0143] Methods and systems of the present disclosure were used to detect NASH / MASH- specific methylated biomarkers. The bodily samples obtained from patients diagnosed with non-alcoholic fatty liver disease (NAFLD) or NASH / MASH were used as disease samples to detect NASH / MASH-specific methylated biomarkers. Samples obtained from patients diagnosed with simple steatosis were used as controls. In particular, blood was collected from each patient and plasma was isolated from each blood sample. Cell-free DNA was extracted from plasma using Qiagen QIAamp circulating nucleic acid kit (Qiagen Sciences, Germantown, Maryland). Subsequently, cell-free DNA was treated with bisulfite to convert unmethylated cytosines (C) to uracil (U). Methylated cytosine (5mC and 5hmC) residues were protected from bisulfite conversion and thus were not converted to uridine residues. Thus, cytosine residues remaining in the bisulfite-treated DNA molecules corresponded to methylated cytosine residues. The bisulfite-treated DNA was purified, followed by NGS library preparation, PCR amplification, and sequencing on a sequencing system (e.g., Illumina NextSeq 500, NextSeq 2000 or NovaSeq 6000 System). The NGS-compatible sequencing libraries were prepared using xGen methyl- sequencing DNA library prep kit(Integrated DNA Technologies, Coralville, Iowa). The sequencing adapters (e.g., Illumina TruSeq methyl capture EPIC library prep protocol kit) were used for sequencing.

[0144] The sequencing reads obtained after bisulfite sequencing were converted into C-to-T and G-to-A (equivalent to C-to-T conversion on the reverse strand) versions, and aligned to equivalently converted versions of a reference genome (e.g., hg38 human genome). Each of them was aligned to equivalently pre-converted forms of a reference genome using four parallel instances. The methylation state of each individual cytosine residue was determined by comparing the read sequence with the corresponding sites in the reference genomic sequence. Subsequently, the sequences associated with adapter dimers were trimmed using Trim Galore, a wrapper script that automates quality and adapter trimming as well as quality control. The sequences associated with adapter dimers were quality trimmed with a minimum quality score greater than 25. The trimmed raw sequencing reads were aligned against the reference genome using the Bismark align tool. After sequence alignment, the Bismark methylation extractor tool was used to summarize the level of methylation of CpG sites. For comparison with consortium-scale datasets, including epigenome roadmap, BLUEPRINT, and datasets used in a reference method (N. Loyfer et al., A DNA Methylation Atlas for Normal Human Cell Types, Nature, 613, 355-364 (2023), uninformedly processed downstream datasets (e.g., DNA methylation level in CpG sites) were used.

[0145] The quantitative measures of NASH / MASH-specific methylated biomarkers were determined using both the count-based approach and ratio-based approach. The count-based approach calculated a ratio between a number of methylated sequence reads that align with a sequence block and a total number of sequence reads that align with the same sequence block. The ratio-based approach calculated the methylation level for each cytosine in a sequence block, and generated a mean methylation level of the sequence block using multiple nucleotide sites. Based on the two approaches, a methylation matrix was formed, where columns listed a variety of tissue samples and rows listed binned block sites (e.g. chrl: 1,000,000). The NASH / MASH-specific methylated biomarkers were detected based on the following detection criteria:

[0146] (1) The sequence blocks were determined for both disease samples and control samples using the count-based and ratio-based approaches;

[0147] (2) The methylation status of the disease samples was compared with that of control samples using non-parametric test (e.g., Mann- Whitney U test, with p- value lower than 0.01 and FDR lower than 0.1);

[0148] (3) The number of CpG sites in the sequence blocks were more than 2-11 sites per group, determined by quantifying the beta value (methylation status) of adjacent CpG sites within a sequence block that contain similar methylation status.

[0149] (4) The methylation level for missing sequence blocks was imputed by determining a group-averaged methylation level across other CpG sites within the sequence block, and the imputed sequence blocks were at most three samples per group;

[0150] (5) A mean methylation status (e.g., average beta value) between disease samples and control samples that had at least a 20% difference in either direction were selected. In other words, a mean methylation status of hypermethylated biomarkers in the disease samples had to be at least 20% higher than that in control samples, and a mean methylation status of hypomethylated biomarkers in the disease samples had to be at least 20% lower than that in control samples to be selected;

[0151] (6) A standard deviation of both disease samples and control samples was lower than 10%.

[0152] A total of 583 NASH / MASH-specific methylated biomarkers were identified following the above detection criteria, listed in Table 2. For gene annotation, information on the nearest genes, exons, introns, intergenic regions, and UTRs was extracted using Homer annotation tools (annotatePeaks.pl DMR hg38).

[0153] Kits

[0154] The present disclosure provides kits for identifying or monitoring a liver disease of a subject. A kit may comprise probes for identifying a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences for each of a plurality of liver disease- associated biomarkers or methylated CpG sites in a bodily sample of the subject. A quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences at each of a plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample may be indicative of one or more liver diseases. The probes may be selective for the sequences at the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. A kit may comprise instructions for using the probes to process the bodily sample to generate datasets indicative of a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences at each of the plurality of liver disease- associated biomarkers or methylated CpG sites in a bodily sample of the subject.

[0155] The probes in the kit may be selective for the sequences for the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. The probes in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) moleculescorresponding to the plurality of liver disease-associated biomarkers or methylated CpG sites. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with nucleic acid sequences from one or more of the plurality of liver disease-associated biomarkers or methylated CpG sites or genomic regions. The plurality of liver disease-associated biomarkers or methylated CpG sites or genomic regions may comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more distinct liver disease-associated biomarkers or methylated CpG sites or genomic regions. The plurality of liver disease-associated biomarkers or methylated CpG sites or genomic regions may comprise one or more members selected from the group listed in Table 1 or Table 2.

[0156] The instructions in the kit may comprise instructions to assay the bodily sample using the probes that are selective for the sequences at the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. These probes may be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) from one or more of the plurality of liver disease-associated biomarkers or methylated CpG sites. These nucleic acid molecules may be primers or enrichment sequences. The instructions to assay the bodily sample may comprise introductions to perform array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the bodily sample to generate datasets indicative of a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences at each of the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. A quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences at each of a plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample may be indicative of one or more liver diseases.

[0157] The instructions in the kit may comprise instructions to measure and interpret assay readouts, which may be quantified at one or more of the plurality of liver disease-associated biomarkers or CpG sites to generate the datasets indicative of a quantitative measure (e.g., indicative of a presence, absence, or relative amount) of sequences at each of the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to the plurality of liver disease-associated biomarkers or methylated CpG sites may generate the datasets indicative of a quantitative measure (e.g., indicative of a presence,absence, or relative amount) of sequences at each of the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily sample. Assay readouts may comprise quantitative PCR (qPCR) values, digital PCR (dPCR) values, droplet digital PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof.

[0158] Machine learning algorithms

[0159] After using one or more assays to process one or more bodily samples derived from the subject to generate one or more datasets indicative of the liver disease, a trained algorithm may be used to process one or more of the datasets (e.g., at each of a plurality of liver disease-associated biomarkers or methylated CpG sites) to determine the liver disease. For example, the trained algorithm may be used to determine quantitative measures of sequences at each of the plurality of liver disease-associated biomarkers or methylated CpG sites in the bodily samples. The trained algorithm may be configured to identify the liver disease with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more than 99% for at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, or more than about 500 independent samples.

[0160] The trained algorithm may comprise a supervised machine learning algorithm. The trained algorithm may comprise a classification and regression tree (CART) algorithm. The supervised machine learning algorithm may comprise, for example, a Random Forest, a support vector machine (SVM), a neural network, or a deep learning algorithm. The trained algorithm may comprise an unsupervised machine learning algorithm.

[0161] The trained algorithm may be configured to accept a plurality of input variables and to produce one or more output values based on the plurality of input variables. The plurality of input variables may comprise one or more datasets indicative of a liver disease. For example, an input variable may comprise a number of sequences corresponding to or aligning to each of the plurality of liver disease-associated biomarkers or methylated CpG sites. The plurality of input variables may also include clinical health data of a subject.

[0162] The trained algorithm may comprise a classifier, such that each of the one or more output values comprises one of a fixed number of possible values (e.g., a linear classifier, a logistic regression classifier, etc.) indicating a classification of the bodily sample by the classifier. The trained algorithm may comprise a binary classifier, such that each of the one ormore output values comprises one of two values (e.g., {0, 1}, {positive, negative}, or {high- risk, low-risk}) indicating a classification of the bodily sample by the classifier. The trained algorithm may be another type of classifier, such that each of the one or more output values comprises one of more than two values (e.g., {0, 1, 2}, {positive, negative, or indeterminate}, or {high-risk, intermediate-risk, or low-risk}) indicating a classification of the bodily sample by the classifier. The output values may comprise descriptive labels, numerical values, or a combination thereof. Some of the output values may comprise descriptive labels. Such descriptive labels may provide an identification or indication of the disease or disorder state of the subject, and may comprise, for example, positive, negative, high-risk, intermediaterisk, low-risk, or indeterminate. Such descriptive labels may provide an identification of a treatment for the subject’s liver disease, and may comprise, for example, a therapeutic intervention, a duration of the therapeutic intervention, and / or a dosage of the therapeutic intervention suitable to treat a liver disease. Such descriptive labels may provide an identification of secondary clinical tests that may be appropriate to perform on the subject, and may comprise, for example, an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRJ) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a bodily cytology, or any combination thereof. For example, such descriptive labels may provide a prognosis of the liver disease of the subject. As another example, such descriptive labels may provide a relative assessment of the liver disease of the subject. Some descriptive labels may be mapped to numerical values, for example, by mapping “positive” to 1 and “negative” to 0.

[0163] Some of the output values may comprise numerical values, such as binary, integer, or continuous values. Such binary output values may comprise, for example, {0, 1 }, {positive, negative}, or {high-risk, low-risk}. Such integer output values may comprise, for example, {0, 1, 2}. Such continuous output values may comprise, for example, a probability value of at least 0 and no more than 1. Such continuous output values may comprise, for example, an unnormalized probability value of at least 0. Such continuous output values may indicate a prognosis of the liver disease of the subject. Some numerical values may be mapped to descriptive labels, for example, by mapping 1 to “positive” and 0 to “negative.”

[0164] Some of the output values may be assigned based on one or more cutoff values. For example, a binary classification of samples may assign an output value of “positive” or 1 if the sample indicates that the subject has at least a 50% probability of having a liver disease (e.g., liver disease). For example, a binary classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has less than a 50%probability of having a liver disease (e.g., liver disease). In this case, a single cutoff value of 50% is used to classify samples into one of the two possible binary output values. Examples of single cutoff values may include about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, and about 99%.

[0165] As another example, a classification of samples may assign an output value of “positive” or 1 if the sample indicates that the subject has a probability of having a liver disease (e.g., liver disease) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The classification of samples may assign an output value of “positive” or 1 if the sample indicates that the subject has a probability of having a liver disease (e.g., liver disease) of more than about 50%, more than about 55%, more than about 60%, more than about 65%, more than about 70%, more than about 75%, more than about 80%, more than about 85%, more than about 90%, more than about 91%, more than about 92%, more than about 93%, more than about 94%, more than about 95%, more than about 96%, more than about 97%, more than about 98%, or more than about 99%.

[0166] The classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has a probability of having a liver disease (e.g., liver disease) of less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. The classification of samples may assign an output value of “negative” or 0 if the sample indicates that the subject has a probability of having a liver disease (e.g., liver disease) of no more than about 50%, no more than about 45%, no more than about 40%, no more than about 35%, no more than about 30%, no more than about 25%, no more than about 20%, no more than about 15%, no more than about 10%, no more than about 9%, no more than about 8%, no more than about 7%, no more than about 6%, no more than about 5%, no more than about 4%, no more than about 3%, no more than about 2%, or no more than about1%.

[0167] The classification of samples may assign an output value of “indeterminate” or 2 if the sample is not classified as “positive”, “negative”, 1, or 0. In this case, a set of two cutoff values is used to classify samples into one of the three possible output values. Examples of sets of cutoff values may include { 1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, { 15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}. Similarly, sets of n cutoff values may be used to classify samples into one of n+1 possible output values, where n is any positive integer.

[0168] The trained algorithm may be trained with a plurality of independent training samples. Each of the independent training samples may comprise a bodily sample from a subject, associated datasets obtained by assaying the bodily sample (as described elsewhere herein), and one or more known output values corresponding to the bodily sample (e.g., a clinical diagnosis, prognosis, absence, or treatment efficacy of a liver disease of the subject).Independent training samples may comprise bodily samples and associated datasets and outputs obtained or derived from a plurality of different subjects. Independent training samples may comprise bodily samples and associated datasets and outputs obtained at a plurality of different time points from the same subject (e.g., on a regular basis such as weekly, biweekly, or monthly). Independent training samples may be associated with presence of the liver disease (e.g., training samples comprising bodily samples and associated datasets and outputs obtained or derived from a plurality of subjects known to have the liver disease). Independent training samples may be associated with absence of the liver disease (e.g., training samples comprising bodily samples and associated datasets and outputs obtained or derived from a plurality of subjects who are known to not have a previous diagnosis of the liver disease or who have received a negative test result for the liver disease).

[0169] The trained algorithm may be trained with at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent training samples. The independent training samples may comprise bodily samples associated with presence of the liver disease and / or bodily samples associated with absence of the liver disease. The trained algorithm may be trained with no more than about 500, no more than about 450, no more than about 400, no more than about 350, no more than about 300, no more than about 250, no more than about 200, no more than about 150, no more than about 100, or no more than about 50 independent trainingsamples associated with presence of the liver disease. In some embodiments, the bodily sample is independent of samples used to train the trained algorithm.

[0170] The trained algorithm may be trained with a first number of independent training samples associated with presence of the liver disease and a second number of independent training samples associated with absence of the liver disease. The first number of independent training samples associated with presence of the liver disease may be no more than the second number of independent training samples associated with absence of the liver disease. The first number of independent training samples associated with presence of the liver disease may be equal to the second number of independent training samples associated with absence of the liver disease. The first number of independent training samples associated with presence of the liver disease may be greater than the second number of independent training samples associated with absence of the liver disease.

[0171] The trained algorithm may be configured to identify the liver disease at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more; for at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent training samples. The accuracy of identifying the liver disease by the trained algorithm may be calculated as the percentage of independent test samples (e.g., subjects known to have the liver disease or subjects with negative clinical test results for the liver disease) that are correctly identified or classified as having or not having the liver disease.

[0172] The trained algorithm may be configured to identify the liver disease with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The PPV of identifying the liver disease using the trained algorithm may be calculated as the percentage of bodily samples identified or classified as having the liver disease that correspond to subjects that truly have the liver disease.

[0173] The trained algorithm may be configured to identify the liver disease with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The NPV of identifying the liver disease using the trained algorithm may be calculated as the percentage of bodily samples identified or classified as not having the liver disease that correspond to subjects that truly do not have the liver disease.

[0174] The trained algorithm may be configured to identify the liver disease with a clinical sensitivity at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical sensitivity of identifying the liver disease using the trained algorithm may be calculated as the percentage of independent test samples associated with presence of the liver disease (e.g., subjects known to have the liver disease) that are correctly identified or classified as having the liver disease.

[0175] The trained algorithm may be configured to identify the liver disease with a clinical specificity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, atleast about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical specificity of identifying the liver disease using the trained algorithm may be calculated as the percentage of independent test samples associated with absence of the liver disease (e.g., subjects with negative clinical test results for the liver disease) that are correctly identified or classified as not having the liver disease.

[0176] The trained algorithm may be configured to identify the liver disease with an Area- Under-Curve (AUG) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, at least about 0.99, or more. The AUG may be calculated as an integral of the Receiver Operating Characteristic (ROC) curve (e.g., the area under the ROC curve) associated with the trained algorithm in classifying bodily samples as having or not having the liver disease.

[0177] The trained algorithm may be adjusted or tuned to improve one or more of the performance, accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or AUG of identifying the liver disease. The trained algorithm may be adjusted or tuned by adjusting parameters of the trained algorithm (e.g., a set of cutoff values used to classify a bodily sample as described elsewhere herein, or weights of a neural network). The trained algorithm may be adjusted or tuned continuously during the training process or after the training process has been completed.

[0178] After the trained algorithm is initially trained, a subset of the inputs may be identified as most influential or most important to be included for making high-quality classifications. For example, a subset of the plurality of liver disease-associated biomarkers or methylated CpG sites may be identified as most influential or most important to be included for makinghigh-quality classifications or identifications of liver diseases (or sub-types of liver diseases). The plurality of liver disease-associated biomarkers or methylated CpG sites or a subset thereof may be ranked based on classification metrics indicative of each genomic locus’s influence or importance toward making high-quality classifications or identifications of liver diseases (or sub-types of liver diseases). Such metrics may be used to reduce, in some cases significantly, the number of input variables (e.g., predictor variables) that may be used to train the trained algorithm to a desired performance level (e.g., based on a desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, AUG, or a combination thereof). For example, if training the trained algorithm with a plurality comprising several dozen or hundreds of input variables in the trained algorithm results in an accuracy of classification of more than 99%, then training the trained algorithm instead with only a selected subset of no more than about 5, no more than about 10, no more than about 15, no more than about 20, no more than about 25, no more than about 30, no more than about 35, no more than about 40, no more than about 45, no more than about 50, or no more than about 100 such most influential or most important input variables among the plurality can yield decreased but still acceptable accuracy of classification (e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%). The subset may be selected by rank-ordering the entire plurality of input variables and selecting a predetermined number (e.g., no more than about 5, no more than about 10, no more than about 15, no more than about 20, no more than about 25, no more than about 30, no more than about 35, no more than about 40, no more than about 45, no more than about 50, or no more than about 100) of input variables with the best classification metrics.

[0179] Identifying or monitoring a liver disease

[0180] After using a trained algorithm to process the dataset, the liver disease may be identified or monitored in the subject. The identification may be based at least in part on quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of DNA or RNA methylation at the liver disease-associated biomarkers or methylated CpG sites).

[0181] The liver disease may be identified in the subject at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The accuracy of identifying the liver disease by the trained algorithm may be calculated as the percentage of independent test samples (e.g., subjects known to have the liver disease or subjects with negative clinical test results for the liver disease) that are correctly identified or classified as having or not having the liver disease.

[0182] The liver disease may be identified in the subject with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The PPV of identifying the liver disease using the trained algorithm may be calculated as the percentage of bodily samples identified or classified as having the liver disease that correspond to subjects that truly have the liver disease.

[0183] The liver disease may be identified in the subject with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The NPV of identifying the liver disease using the trained algorithm may be calculated as the percentage of bodily samples identified or classified as not having the liver disease that correspond to subjects that truly do not have the liver disease.

[0184] The liver disease may be identified in the subject with a clinical sensitivity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical sensitivity of identifying the liver disease using the trained algorithm may be calculated as the percentage of independent test samples associated with presence of the liver disease (e.g., subjects known to have the liver disease) that are correctly identified or classified as having the liver disease.

[0185] The liver disease may be identified in the subject with a clinical specificity of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more. The clinical specificity of identifying the liver disease using the trained algorithm may be calculated as the percentage of independent test samples associated with absence of the liver disease (e.g., subjects with negative clinical test results for the liver disease) that are correctly identified or classified as not having the liver disease.

[0186] In an aspect, the present disclosure provides a method for determining that a subject is at risk of liver disease, comprising assaying a bodily sample derived from the subject to generate a dataset that is indicative of said liver disease risk at a specificity of at least 80%, and using a trained algorithm that is trained on samples independent of the bodily sample todetermine that the subject is at risk of liver disease at an accuracy of at least about 50%, at least about 55%. at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.

[0187] After the liver disease is identified in a subject, a sub-type of the liver disease (e.g., selected from among a plurality of sub-types of the liver disease) may further be identified. The sub-type of the liver disease may be determined based at least in part on the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites). For example, the subject may be identified as being at risk of a sub-type of liver disease (e.g., selected from among a plurality of sub-types of liver disease). After identifying the subject as being at risk of a sub-type of liver disease, a clinical intervention for the subject may be selected based at least in part on the sub-type of liver disease for which the subject is identified as being at risk. In some embodiments, the clinical intervention is selected from a plurality of clinical interventions (e.g., clinically indicated for different sub-types of liver disease).

[0188] In some embodiments, the trained algorithm may determine that the subject is at risk of liver disease of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more.

[0189] The trained algorithm may determine that the subject is at risk of liver disease at an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or more.

[0190] Upon identifying the subject as having the liver disease, the subject may be optionally provided with a therapeutic intervention (e.g., prescribing an appropriate course of treatment to treat the liver disease of the subject). The therapeutic intervention may comprise a prescription of an effective dose of a drug, a further testing or evaluation of the liver disease, a further monitoring of the liver disease, or a combination thereof. If the subject is currently being treated for the liver disease with a course of treatment, the therapeutic intervention may comprise a subsequent different course of treatment (e.g., to increase treatment efficacy due to non-efficacy of the current course of treatment).

[0191] The therapeutic intervention may comprise recommending the subject for a secondary clinical test to confirm a diagnosis of the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell-free biological cytology, or any combination thereof.

[0192] The quantitative measures of sequence reads of the dataset at the panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites). In such cases, the quantitative measures of the dataset of the patient may change during the course of treatment. For example, the quantitative measures of the dataset of a patient with decreasing risk of the liver disease due to an effective treatment may shift toward the profile or distribution of a healthy subject (e.g., a subject without a liver disease). Conversely, for example, the quantitative measures of the dataset of a patient with increasing risk of the liver disease due to an ineffective treatment may shift toward the profile or distribution of a subject with higher risk of the liver disease or a more advanced liver disease.

[0193] The liver disease of the subject may be monitored by monitoring a course of treatment for treating the liver disease of the subject. The monitoring may comprise assessing the liver disease of the subject at two or more time points. The assessing may be based at least on the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNAat the liver disease-associated biomarkers or methylated CpG sites) determined at each of the two or more time points.

[0194] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of one or more clinical indications, such as (i) a diagnosis of the liver disease of the subject, (ii) a prognosis of the liver disease of the subject, (iii) an increased risk of the liver disease of the subject, (iv) a decreased risk of the liver disease of the subject, (v) an efficacy of the course of treatment for treating the liver disease of the subject, and (vi) a non-efficacy of the course of treatment for treating the liver disease of the subject.

[0195] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of a diagnosis of the liver disease of the subject. For example, if the liver disease was not detected in the subject at an earlier time point but was detected in the subject at a later time point, then the difference is indicative of a diagnosis of the liver disease of the subject. A clinical action or decision may be made based on this indication of diagnosis of the liver disease of the subject, such as, for example, prescribing a new therapeutic intervention for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the diagnosis of the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell- free biological cytology, or any combination thereof.

[0196] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of a prognosis of the liver disease of the subject.

[0197] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkersor methylated CpG sites) determined between the two or more time points may be indicative of the subject having an increased risk of the liver disease. For example, if the liver disease was detected in the subject both at an earlier time point and at a later time point, and if the difference is a positive difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) increased from the earlier time point to the later time point), then the difference may be indicative of the subject having an increased risk of the liver disease. A clinical action or decision may be made based on this indication of the increased risk of the liver disease, e.g., prescribing a new therapeutic intervention or switching therapeutic interventions (e.g., ending a current treatment and prescribing a new treatment) for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the increased risk of the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell- free biological cytology, or any combination thereof.

[0198] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of the subject having a decreased risk of the liver disease. For example, if the liver disease was detected in the subject both at an earlier time point and at a later time point, and if the difference is a negative difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) decreased from the earlier time point to the later time point), then the difference may be indicative of the subject having a decreased risk of the liver disease. A clinical action or decision may be made based on this indication of the decreased risk of the liver disease (e.g., continuing or ending a current therapeutic intervention) for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the decreased risk of the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emissiontomography (PET) scan, a PET-CT scan, a cell-free biological cytology, or any combination thereof.

[0199] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of an efficacy of the course of treatment for treating the liver disease of the subject. For example, if the liver disease was detected in the subject at an earlier time point but was not detected in the subject at a later time point, then the difference may be indicative of an efficacy of the course of treatment for treating the liver disease of the subject. A clinical action or decision may be made based on this indication of the efficacy of the course of treatment for treating the liver disease of the subject, e.g., continuing or ending a current therapeutic intervention for the subject. The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the efficacy of the course of treatment for treating the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell-free biological cytology, or any combination thereof.

[0200] In some embodiments, a difference in the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) determined between the two or more time points may be indicative of a non-efficacy of the course of treatment for treating the liver disease of the subject. For example, if the liver disease was detected in the subject both at an earlier time point and at a later time point, and if the difference is a positive or zero difference (e.g., the quantitative measures of sequence reads of the dataset at a panel of liver disease-associated biomarkers or methylated CpG sites (e.g., quantitative measures of RNA transcripts or DNA at the liver disease-associated biomarkers or methylated CpG sites) increased or remained at a constant level from the earlier time point to the later time point), and if an efficacious treatment was indicated at an earlier time point, then the difference may be indicative of a non-efficacy of the course of treatment for treating the liver disease of the subject. A clinical action or decision may be made based on this indication of the non-efficacy of the course of treatment for treating the liver disease of the subject, e.g., ending a current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention for the subject.The clinical action or decision may comprise recommending the subject for a secondary clinical test to confirm the non-efficacy of the course of treatment for treating the liver disease. This secondary clinical test may comprise an imaging test, a blood test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest X-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell- free biological cytology, or any combination thereof.

[0201] In some embodiments, the methods and systems of the present disclosure may receive clinical health data of the subject, wherein the clinical health data comprises a plurality of quantitative or categorical measures of said subject. A trained algorithm may be used to process the clinical health data of the subject to determine a risk score indicative of the risk of liver disease of the subject.

[0202] In some embodiments, for example, the clinical health data comprises one or more quantitative measures of the subject, such as age, weight, height, body mass index (BMI), blood pressure, heart rate, glucose levels, etc. As another example, the clinical health data can comprise one or more categorical measures, such as race, ethnicity, history of medication or other clinical treatment, history of tobacco use, history of alcohol consumption, daily activity or fitness level, genetic test results, blood test results, imaging results, etc.

[0203] In some embodiments, the computer-implemented method for predicting a risk of liver disease of a subject is performed using a computer or mobile device application. For example, a subject can use a computer or mobile device application to input her own clinical health data, including quantitative and / or categorical measures. The computer or mobile device application can then use a trained algorithm to process the clinical health data to determine a risk score indicative of the risk of liver disease of the subject. The computer or mobile device application can then display a report indicative of the risk score indicative of the risk of liver disease of the subject.

[0204] In some embodiments, the risk score indicative of the risk of liver disease of the subject can be refined by performing one or more subsequent clinical tests for the subject. For example, the subject can be referred by a physician for one or more subsequent clinical tests (e.g., an ultrasound imaging or a blood test) based on the initial risk score. Next, the computer or mobile device application may process results from the one or more subsequent clinical tests using a trained algorithm to determine an updated risk score indicative of the risk of liver disease of the subject.

[0205] In some embodiments, the risk score comprises a likelihood of the subject having a liver disease within a pre-determined duration of time. For example, the pre-determinedduration of time may be about 1 hour, about 2 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 1.5 days, about 2 days, about 2.5 days, about 3 days, about 3.5 days, about 4 days, about 4.5 days, about 5 days, about 5.5 days, about 6 days, about 6.5 days, about 7 days, about 8 days, about 9 days, about 10 days, about 12 days, about 14 days, about 3 weeks, about 4 weeks, about 5 weeks, about 6 weeks, about 7 weeks, about 8 weeks, about 9 weeks, about 10 weeks, about 11 weeks, about 12 weeks, about 13 weeks, or more than about 13 weeks.

[0206] Outputting a report of the liver disease

[0207] After the liver disease is identified or an increased risk of the liver disease is monitored in the subject, a report may be electronically outputted that is indicative of (e.g., identifies or provides an indication of) the liver disease of the subject. The subject may not display a liver disease (e.g., is asymptomatic of the liver disease such as a liver disease). The report may be presented on a graphical user interface (GUI) of an electronic device of a user. The user may be the subject, a caretaker, a physician, a nurse, or another health care worker.

[0208] The report may include one or more clinical indications such as (i) a diagnosis of the liver disease of the subject, (ii) a prognosis of the liver disease of the subject, (iii) an increased risk of the liver disease of the subject, (iv) a decreased risk of the liver disease of the subject, (v) an efficacy of the course of treatment for treating the liver disease of the subject, and (vi) a non-efficacy of the course of treatment for treating the liver disease of the subject. The report may include one or more clinical actions or decisions made based on these one or more clinical indications. Such clinical actions or decisions may be directed to therapeutic interventions, induction or inhibition of labor, or further clinical assessment or testing of the liver disease of the subject.

[0209] For example, a clinical indication of a diagnosis of the liver disease of the subject may be accompanied with a clinical action of prescribing a new therapeutic intervention for the subject. As another example, a clinical indication of an increased risk of the liver disease of the subject may be accompanied with a clinical action of prescribing a new therapeutic intervention or switching therapeutic interventions (e.g., ending a current treatment and prescribing a new treatment) for the subject. As another example, a clinical indication of a decreased risk of the liver disease of the subject may be accompanied with a clinical action of continuing or ending a current therapeutic intervention for the subject. As another example, a clinical indication of an efficacy of the course of treatment for treating the liver disease of the subject may be accompanied with a clinical action of continuing or ending a currenttherapeutic intervention for the subject. As another example, a clinical indication of a nonefficacy of the course of treatment for treating the liver disease of the subject may be accompanied with a clinical action of ending a current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention for the subject.

[0210] Computer systems

[0211] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 10 shows a computer system 1001 that is programmed or otherwise configured to detecting a presence or an absence of a liver disease in a subject, in accordance with some embodiments. The computer system 1001 can regulate various aspects of the present disclosure, for example, implementing algorithms. The computer system 1001 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device.

[0212] The computer system 1001 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 1005, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1001 also includes memory or memory location 1010 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1015 (e.g., hard disk), communication interface 1020 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1025, such as cache, other memory, data storage and / or electronic display adapters. The memory 1010, storage unit 1015, interface 1020 and peripheral devices 1025 are in communication with the CPU 1005 through a communication bus (solid lines), such as a motherboard. The storage unit 1015 can be a data storage unit (or data repository) for storing data. The computer system 1001 can be operatively coupled to a computer network (“network”) 1030 with the aid of the communication interface 1020. The network 1030 can be the Internet, an intranet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 1030 in some cases is a telecommunication and / or data network. The network 1030 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1030, in some cases with the aid of the computer system 1001, can implement a peer-to-peer network, which may enable devices coupled to the computer system 1001 to behave as a client or a server.

[0213] The CPU 1005 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1010. The instructions can be directed to the CPU 1005, which can subsequently program or otherwise configure the CPU 1005 to implement methods of thepresent disclosure. Examples of operations performed by the CPU 1005 can include fetch, decode, execute, and writeback.

[0214] The CPU 1005 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1001 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0215] The storage unit 1015 can store files, such as drivers, libraries and saved programs. The storage unit 1015 can store user data, e.g., user preferences and user programs. The computer system 1001 in some cases can include one or more additional data storage units that are external to the computer system 1001, such as located on a remote server that is in communication with the computer system 1001 through an intranet or the Internet.

[0216] The computer system 1001 can communicate with one or more remote computer systems through the network 1030. For instance, the computer system 1001 can communicate with a remote computer system of a user (e.g., a mobile device). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1001 via the network 1030.

[0217] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1001, such as, for example, on the memory 1010 or electronic storage unit 1015. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 1005. In some cases, the code can be retrieved from the storage unit 1015 and stored on the memory 1010 for ready access by the processor 1005. In some situations, the electronic storage unit 1015 can be precluded, and machine-executable instructions are stored on memory 1010.

[0218] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.

[0219] Aspects of the systems and methods provided herein, such as the computer system 1001, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, suchas memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.

[0220] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may beinvolved in carrying one or more sequences of one or more instructions to a processor for execution.

[0221] The computer system 1001 can include or be in communication with an electronic display 1035 that comprises a user interface (UI) 1040 for providing, for example, a dashboard. Examples of UI include, without limitation, a graphical user interface (GUI) and web-based user interface.

[0222] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1005.

[0223] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A method for detecting a presence or an absence of a non-alcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH) in a subject, the method comprising:(a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of NASH / MASH-specific methylated biomarkers in nucleic acids obtained or derived from the bodily sample;(b) computer processing the set of NASH / MASH-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference sequence; and(c) detecting the presence or the absence of NASH or MASH in the subject, based at least in part on the computer processing in (b).

2. The method of claim 1, wherein the set of NASH / MASH-specific methylated biomarkers comprises a member selected from the group listed in Table 2.

3. The method of claim 1, wherein the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a solid tissue sample, and a lymphatic fluid sample.

4. The method of claim 1, wherein the subject is diagnosed with NASH or MASH, or is suspected of having NASH or MASH.

5. The method of claim 1, wherein the subject has a risk factor for NASH or MASH.

6. The method of claim 5, wherein the risk factor for NASH or MASH comprises type 2 diabetes, obesity, metabolic syndrome, family history of liver disease, genetic factors associated with NASH / MASH or liver fibrosis, and / or polycystic ovary syndrome.

7. The method of claim 1, wherein the subject is asymptomatic for NASH or MASH.

8. The method of claim 1, wherein the nucleic acid comprises deoxyribonucleic acid(DNA) or ribonucleic acid (RNA).

9. The method of claim 8, wherein the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

10. The method of claim 1, wherein the assaying in (a) further comprises amplifying the nucleic acid.

11. The method of claim 10, wherein the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification.

12. The method of claim 1, wherein the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing, pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylationspecific PCR.

13. The method of claim 1, further comprising using primers or probes configured to selectively enrich the nucleic acid for the set of NASH / MASH-specific methylated biomarkers.

14. The method of claim 13, wherein the primers or probes are nucleic acid primers, nucleic acid probes, or peptide nucleic acid (PNA) probes.

15. The method of claim 14, wherein the nucleic acid primers, nucleic acid probes, or peptide nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of NASH / MASH-specific methylated biomarkers.

16. The method of claim 1, wherein the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads that align with a sequence block and a total number of sequence reads that align with the same sequence block, or (ii) a mean methylation level of the sequence block.

17. The method of claim 1, wherein the quantitative measures are determined as (i) a first combined beta value of at least two NASH / MASH-specific hypermethylated biomarkers, or (ii) a second combined beta value of at least two NASH / MASH-specific hypomethylated biomarkers.

18. The method of claim 17, wherein the reference comprises a first reference beta value of the at least two NASH / MASH-specific hypermethylated biomarkers and a second reference beta value of the at least two NASH / MASH-specific hypomethylated biomarkers, and wherein the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with NASH or MASH.

19. The method of claim 1, further comprising administering a treatment to the subject based on the detected presence of NASH or MASH in the subject.

20. The method of claim 19, wherein the treatment is selected f rom the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention.

21. The method of claim 19, wherein the treatment is a pharmaceutical treatment selected from the group consisting of: Resmetiron / Rezdiffra (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof.

22. The method of claim 19, wherein the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione, ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof.

23. The method of claim 19, wherein the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication (naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 lU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

24. The method of claim 1, further comprising assaying the bodily sample to determine quantitative measures of a set of NASH / MASH-specific hypermethylated or hypomethylated cytosine-phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample.

25. The method of claim 24, wherein the set of NASH / MASH-specific hypermethylated or hypomethylated CpG sites comprises a member selected from the group listed in Table 2.

26. The method of claim 1, wherein the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

27. The method of claim 1, wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

28. The method of claim 1, wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

29. The method of claim 1 , wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at leastabout 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

30. The method of claim 1, wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a positive predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

31. The method of claim 1, wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

32. The method of claim 1 , wherein the trained machine learning algorithm is configured to detect the presence or the absence of NASH or MASH with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.99.

33. The method of claim 1, wherein the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of NASH or MASH and a second set of independent training samples associated with an absence or non-elevated susceptibility of NASH or MASH.

34. A method for detecting a presence or an absence of a liver disease in a subject, the method comprising:(a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in nucleic acids obtained or derived from the bodily sample, wherein the set of liver-specific methylated biomarkers comprises a member selected from the group listed in Table 1;(b) computer processing the set of liver-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference; and(c) detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing in (b).

35. The method of claim 34, wherein the liver disease is selected from the group consisting of steatosis, nonalcoholic steatohepatitis (NASH), metabolic dysfunction-associated steatohepatitis (MASH), non-alcoholic fatty liver disease (NAFLD), metabolic dysfunction-associated steatotic liver disease (MASLD), liver fibrosis, cirrhosis, and chronic liver disease.

36. The method of claim 34, wherein the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a solid tissue sample, and a lymphatic fluid sample.

37. The method of claim 34, wherein the subject is diagnosed with the liver disease or is suspected of having the liver disease.

38. The method of claim 34, wherein the subject is asymptomatic for the liver disease.

39. The method of claim 34, wherein the nucleic acid comprises deoxyribonucleic acid(DNA) or ribonucleic acid (RNA).

40. The method of claim 39, wherein the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

41. The method of claim 34, wherein the assaying in (a) further comprises amplifying the nucleic acid.

42. The method of claim 41, wherein the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification).

43. The method of claim 34, wherein the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing, pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylationspecific PCR.

44. The method of claim 34, further comprising using primers or probes configured to selectively enrich the nucleic acid for the set of liver-specific methylated biomarkers.

45. The method of claim 44, wherein the primers or probes are nucleic acid primers, nucleic acid probes, or peptide nucleic acid (PNA) probes.

46. The method of claim 45, wherein the nucleic acid primers, nucleic acid probes, or peptide nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of liver- specific methylated biomarkers.

47. The method of claim 34, wherein the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads that align with a sequence block and a total number of sequence reads that align with the same sequence block, or (ii) a mean methylation level of the sequence block.

48. The method of claim 34, wherein the quantitative measures are determined as (i) a first combined beta value of at least two liver-specific hypermethylated biomarkers, or (ii) a second combined beta value of at least two liver-specific hypomethylated biomarkers.

49. The method of claim 48, wherein the reference comprises a first reference beta value of the at least two liver-specific hypermethylated biomarkers and a second reference beta value of the at least two liver-specific hypomethylated biomarkers, and wherein the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with liver disease.

50. The method of claim 34, further comprising administering a treatment to the subject based on the detected presence of the liver disease in the subject.

51. The method of claim 50, wherein the treatment is selected from the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention.

52. The method of claim 50, wherein the treatment is a pharmaceutical treatment selected from the group consisting of: Resmetiron / Rezdifffa (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof.

53. The method of claim 50, wherein the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione, ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof.

54. The method of claim 50, wherein the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication (naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 lU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

55. The method of claim 34, further comprising assaying the bodily sample to determine quantitative measures of a set of liver-specific hypermethylated or hypomethylated cytosine- phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample.

56. The method of claim 55, wherein the set of liver-specific hypermethylated or hypomethylatiedCpG sites comprises a member selected from the group listed in Table 1.

57. The method of claim 34, wherein the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

58. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

59. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

60. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

61. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a positive predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

62. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

63. The method of claim 34, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.99.

64. The method of claim 34, wherein the trained machine learning algorithm is configured to distinguish between a presence or an absence of each of a plurality of different liver diseases.

65. The method of claim 34, wherein the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of the liver disease and a second set of independent training samples associated with an absence or non-elevated susceptibility of the liver disease.

66. A method for detecting a presence or an absence of a liver disease in a subject, the method comprising:(a) assaying a bodily sample obtained or derived from the subject to determine quantitative measures of a set of liver-specific methylated biomarkers in a nucleic acid from the bodily sample, wherein the quantitative measures are determined as (i) a ratio between a number of methylated sequence reads that align with a sequence block and a total number of reads that align with the same sequence block, (ii) a mean methylation level of the sequence block, (iii) a first combined beta value of at least one liver-specific hypermethylated biomarkers, or (iv) a second combined beta value of at least one liver-specific hypomethylated biomarkers;(b) computer processing the set of liver-specific methylated biomarkers using a trained machine learning algorithm or compared against a reference; and(c) detecting the presence or the absence of the liver disease in the subject, based at least in part on the computer processing in (b).

67. The method of claim 66, wherein the set of liver-specific methylated biomarkers comprises a member selected from the group listed in Table 1.

68. The method of claim 66, wherein the liver disease is selected from the group consisting of steatosis, nonalcoholic steatohepatitis (NASH), metabolic dysfunction- associated steatohepatitis (MASH), non-alcoholic fatty liver disease (NAFLD), metabolic dysfunction-associated steatotic liver disease (MASLD), liver fibrosis, cirrhosis, and chronic liver disease.

69. The method of claim 66, wherein the bodily sample is selected from the group consisting of a whole blood sample, a plasma sample, a serum sample, a saliva sample, a stool sample, a urine sample, a solid tissue sample, and a lymphatic fluid sample.

70. The method of claim 66, wherein the subject is diagnosed with the liver disease or is suspected of having the liver disease.

71. The method of claim 66, wherein the subject is asymptomatic for the liver disease.

72. The method of claim 66, wherein the nucleic acid comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).

73. The method of claim 72, wherein the assaying in (a) further comprises (i) subjecting the bodily sample to conditions that are sufficient to isolate, enrich, or extract DNA molecules or RNA molecules, and (ii) analyzing the DNA molecules or RNA molecules.

74. The method of claim 73, wherein the assaying in (a) further comprises amplifying the nucleic acid.

75. The method of claim 74, wherein the amplifying comprises polymerase chain reaction (PCR), rolling circle amplification, or isothermal amplification.

76. The method of claim 66, wherein the assaying in (a) further comprises use of DNA sequencing, RNA sequencing, bisulfite sequencing (BS-seq), targeted methylation sequencing, pyrosequencing, enzymatic methylation sequencing (EM-Seq), nanopore sequencing, enzymatic treatment, a methylation array, droplet digital PCR or methylationspecific PCR.The method of claim 66, further comprising using primers or probes configured to selectively enrich the nucleic acid for the set of liver-specific methylated biomarkers.

78. The method of claim 77, wherein the primers or probes are nucleic acid primers, nucleic acid probes, or peptide nucleic acid probes.

79. The method of claim 78, wherein the nucleic acid primers, nucleic acid probes, or peptide nucleic acid probes have sequence complementarity with nucleic acid sequences of the set of liver- specific methylated biomarkers.

80. The method of claim 66, wherein the reference comprises a first reference beta value of the at least one liver-specific hypermethylated biomarkers and a second reference beta value of the at least one liver-specific hypomethylated biomarkers, and wherein the first and second reference beta values are determined by assaying a control bodily sample obtained or derived from a subject who is not diagnosed with liver disease.

81. The method of claim 66, further comprising administering a treatment to the subject based on the detected presence of the liver disease in the subject.

82. The method of claim 81, wherein the treatment is selected from the group consisting of lifestyle intervention, pharmacologic therapy, and surgical intervention.

83. The method of claim 81, wherein the treatment is selected from the group consisting of: Resmetiron / Rezdiffra (THRb agonist), Lanifibranor (Pan-PPAR agonist), Semaglutide (GLP1-RA), and a combination thereof.

84. The method of claim 81, wherein the treatment is selected from the group consisting of: Rencofilstat (Cyclophilin inhibitor), PXL065 (Deuterium-modified thiazolidinedione, ION224 (DGAT2 inhibitor), Denifanstat (FASN inhibitor), Efruxifermin (FGF21 agonist), Pegozafermin (FGF21 agonist), Tirzepatide (GLP1-RA / GIP / GR), BI456906 (GLP1- RA / GIP / GR), Pemvidutide (GLP1-RA / GIP / GR), Cotadutide (GLP1-RA / GIP / GR), HSD17B13 (GSK4532990), Saroglitazar (PP AR agonists), Icosabutate (structurally engineered fatty acids), VK2809 (THRb agonist), and a combination thereof.

85. The method of claim 81, wherein the treatment is selected from the group consisting of: lifestyle modifications (diet and exercise), losing 5-10% of body weight, bariatric surgery, weight loss medication (naltrexone, phentermine, topiramate, bupropion, orlistat), vitamin E (>800 lU / d), GLP-1 receptor agonist, Dual GLP-1 and GIP agonist, Pioglitazone, and a combination thereof.

86. The method of claim 66, further comprising assaying the bodily sample to determine quantitative measures of a set of liver-specific hypermethylated or hypomethylated cytosine- phosphate-guanine (CpG) sites in the nucleic acid from the bodily sample.

87. The method of claim 86, wherein the set of liver-specific hypermethylated or hypomethylated CpG sites comprises a member selected from the group listed in Table 1.

88. The method of claim 66, wherein the trained machine learning algorithm comprises a deep learning algorithm, a linear or logistic regression, a support vector machine (SVM), a neural network, a Random Forest, a gradient boosted algorithm.

89. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

90. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

91. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about70%, at least about 75%, at least about 80%, at Least about 85%, at least about 90%, at least about 95%, or at least about 99%.

92. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a positive predictive value of at least about 50%, at least about 55%, at least about 60%, at Least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

93. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with a negative predictive value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at Least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%.

94. The method of claim 66, wherein the trained machine learning algorithm is configured to detect the presence or the absence of the liver disease with an area under receiver operating characteristic curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, or at least about 0.

99.

95. The method of claim 66, wherein the trained machine learning algorithm is configured to distinguish between a presence or an absence of each of a plurality of different liver diseases.

96. The method of claim 66, wherein the trained machine learning algorithm is trained using a first set of independent training samples associated with a presence or elevated susceptibility of the liver disease and a second set of independent training samples associated with an absence or non-elevated susceptibility of the liver disease.