Method and system for detecting and evaluating liver conditions

By analyzing DNA methylation patterns in cell-free biological samples and using machine learning algorithms, the invasiveness of liver biopsy has been addressed, enabling highly sensitive and specific detection of liver diseases and supporting personalized treatment and management.

CN121152886APending Publication Date: 2025-12-16HEPTA BIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480020162.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-18
Filing Date
2024-01-17
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing liver biopsy methods are invasive and risky, limiting their widespread use in the detection of liver diseases, especially in the assessment of liver fibrosis.

Method used

By analyzing cell-free biological samples such as plasma cfDNA samples, the methylation pattern or methylation level of DNA molecules is determined, and these data are processed using trained machine learning algorithms to generate electronic reports indicating the presence or risk of liver disease.

Benefits of technology

It provides a non-invasive method that can detect and monitor liver diseases with high sensitivity and specificity, including early and late-stage liver diseases, supporting personalized treatment decisions and the management of liver diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121152886A_ABST
    Figure CN121152886A_ABST
Patent Text Reader

Abstract

In one aspect, the disclosure provides a method for identifying whether a subject has or is at increased risk of developing a liver disease (e.g., non-alcoholic fatty liver disease, non-alcoholic steatohepatitis, hepatitis, cirrhosis, or cancer), comprising providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject, determining the cfDNA sample or a derivative thereof to determine a methylation pattern or methylation level of a DNA molecule of the cfDNA sample, processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for liver disease, and generating, based at least in part on the output, an electronic report indicating that the subject has a liver disease or an increased risk of developing a liver disease.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 439,716, filed January 18, 2023, which is incorporated herein by reference in its entirety. Background Technology

[0003] Liver disease can be caused by a variety of pathologies, such as infection, genetic conditions, obesity, and alcohol abuse. Blood tests can be used to measure the levels of enzyme biomarkers in the blood. Liver function tests, such as the international normalized ratio (INR), can be used to assess the extent of coagulopathy, an indicator of liver dysfunction. Imaging tools such as ultrasound, magnetic resonance imaging (MRI), or computed tomography (CT) can be used to visualize signs of liver damage, scarring, or tumors.

[0004] Incorporation

[0005] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, or patent application is expressly and individually indicated to be incorporated by reference. If any publication, patent, or patent application incorporated by reference conflicts with the disclosure in this specification, this specification is intended to supersede and / or give precedence to any such conflicting material. Summary of the Invention

[0006] Liver biopsy is currently considered the gold standard for evaluating liver fibrosis in patients with fatty liver disease. However, the inherent risks and invasiveness of biopsy evaluation may limit its widespread use. Improved diagnostic tools for detecting liver disease may be crucial for effective disease management and treatment.

[0007] In view of the need for improved diagnostic tools for detecting liver diseases, this disclosure provides methods, systems, and kits for identifying or monitoring liver diseases by processing cell-free biological samples obtained from or derived from a subject. Cell-free biological samples (e.g., plasma samples) obtained from a subject can be analyzed to identify liver diseases, which may include, for example, measuring the presence, absence, or relative assessment of liver disease. Such subjects may include subjects with one or more liver diseases and subjects without one or more liver diseases. Liver diseases may include, for example, alcoholic fatty liver disease (AFLD), alcohol-related liver disease (ALD), metabolic and alcohol-related / associative liver disease (MetALD), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), steatotic liver disease (SLD), metabolic dysfunction-associated fatty liver disease (MAFLD), metabolic dysfunction-associated steatotic liver disease (MASLD), metabolic dysfunction-associated steatohepatitis (MASH), cryptogenic steatotic liver disease (cryptogenic SLD), hepatitis, cancer (e.g., hepatocellular carcinoma or hepatobiliary carcinoma), and cirrhosis.

[0008] In one aspect, this disclosure provides a method for identifying whether a subject has liver disease or is at an increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) determining the cfDNA sample or a derivative thereof to determine the methylation pattern or methylation level of the DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating that the subject has liver disease or is at an increased risk of developing liver disease, based at least in part on the output.

[0009] In another aspect, this disclosure provides a method for monitoring liver disease in a subject, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) measuring the cfDNA sample or a derivative thereof to determine the methylation pattern or methylation level of the DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating the progression of the subject's liver disease, at least in part based on the output.

[0010] In another aspect, this disclosure provides a method for identifying the prognosis of liver disease in subjects who have liver disease or are at an increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) measuring the cfDNA sample or a derivative thereof to determine the methylation pattern or methylation level of the DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating the prognosis of subjects who have liver disease or are at an increased risk of developing liver disease, based at least in part on the output.

[0011] In another aspect, this disclosure provides a method for identifying a treatment for a subject who has liver disease or is at increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) measuring the cfDNA sample or a derivative thereof to determine the methylation pattern or methylation level of the DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating a treatment for a subject who has liver disease or is at increased risk of developing liver disease, based at least in part on the output.

[0012] In another aspect, this disclosure provides a method for determining a treatment response in a subject with liver disease or at an increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) measuring the cfDNA sample or a derivative thereof to determine the methylation pattern or methylation level of the DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating a treatment response in a subject with liver disease or at an increased risk of developing liver disease, based at least in part on the output.

[0013] In some implementations, the determination includes identifying the methylation pattern and methylation level of DNA molecules in a cfDNA sample, wherein a trained ML algorithm is used to process the methylation pattern and methylation level.

[0014] In some implementations, the assay includes sequencing.

[0015] In some embodiments, the method further includes treating the DNA molecules of the cfDNA sample with a reaction mixture containing an enzyme for methylation-sensing sequencing prior to sequencing.

[0016] In some embodiments, the method further includes treating the DNA molecules of the cfDNA sample with a reaction mixture containing bisulfite prior to sequencing.

[0017] In some implementations, the assay includes amplification.

[0018] In some implementations, amplification includes polymerase chain reaction (PCR).

[0019] In some implementations, cfDNA samples are obtained from or derived from plasma samples, serum samples, urine samples, saliva samples, or liver tissue samples.

[0020] In some implementations, the method also includes grading whole blood samples derived from the object to provide cfDNA samples.

[0021] In some embodiments, (a) includes subjecting the cfDNA sample to conditions sufficient to separate, enrich, or extract a set of DNA molecules, and wherein (b) includes measuring the DNA molecules.

[0022] In some implementations, (b) includes using nucleic acid primers or probes to selectively enrich sets of DNA molecules corresponding to sets of one or more genomic regions.

[0023] In some implementations, one or more genomic regions are selected from the genes listed in Table 1.

[0024] In some implementations, the nucleic acid primers or probes are sequence complementary to the set of nucleic acid sequences of one or more genomic regions.

[0025] In some implementations, cfDNA samples are measured without nucleic acid isolation, enrichment, or extraction.

[0026] In some implementations, the subjects are asymptomatic for liver disease.

[0027] In some implementations, the output indicates with at least 50% accuracy whether a cfDNA sample is positive for liver disease.

[0028] In some implementations, accuracy is determined by calculating the percentage of independent samples that are correctly identified as having or not having liver disease.

[0029] In some implementations, the output indicates whether a cfDNA sample is positive for liver disease with at least 50% clinical sensitivity.

[0030] In some implementation schemes, the clinical sensitivity is at least 50%.

[0031] In some implementations, the output indicates with at least 50% clinical specificity whether a cfDNA sample is positive for liver disease.

[0032] In some implementation schemes, the clinical specificity is at least 50%.

[0033] In some implementations, the output indicates whether a cfDNA sample is positive for liver disease with a positive predictive value of at least 50%.

[0034] In some implementations, the output indicates whether a cfDNA sample is positive for liver disease with a negative predictive value of at least 50%.

[0035] In some implementations, the output indicates whether the cfDNA sample is positive for liver disease with an area under the receiver operating characteristics (AUROC) of at least 0.50.

[0036] In some implementations, the output indicates whether a cfDNA sample is positive for liver disease with a positive likelihood ratio of at least about 1.3.

[0037] In some implementations, the output indicates whether the cfDNA sample is negative for liver disease with a negative likelihood ratio of up to about 0.75.

[0038] In some implementations, liver disease is defined as early-stage liver disease.

[0039] In some implementations, liver disease is defined as advanced liver disease.

[0040] In some implementations, the liver disease is non-alcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH).

[0041] In some implementations, liver disease is fibrosis.

[0042] In some implementations, liver disease is cirrhosis.

[0043] In some implementations, liver disease is hepatocellular carcinoma (HCC).

[0044] In some implementations, the liver disease is hepatobiliary cancer, including, for example, bile duct cancer, angiosarcoma, gallbladder cancer, or undifferentiated embryonic sarcoma of the liver (UESL).

[0045] In some implementations, liver disease is defined as viral hepatitis.

[0046] In some implementations, the liver disease is non-alcoholic fatty liver disease (NAFLD) or metabolic dysfunction-associated fatty liver disease (MASLD).

[0047] In some implementations, the liver disease is non-alcoholic fatty liver disease (NAFL) or steatosis.

[0048] In some implementations, the liver disease is metabolic dysfunction-associated fatty liver disease (MAFLD).

[0049] In some implementations, liver disease is alcohol-related liver disease (ALD).

[0050] In some implementations, liver disease is defined as metabolic and alcohol-related liver disease (MetALD).

[0051] In some implementations, the method also includes providing the subject with a therapeutic intervention for liver disease, at least in part based on the output.

[0052] In some implementation schemes, the liver disease is NASH, and the therapeutic interventions include vitamin E supplementation, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise programs, diet programs, or bariatric surgery.

[0053] In some implementations, the liver disease is NASH, and the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, an FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, an SCD-1 (stearoyl-CoA desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, an FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted gene modifier (including, for example, PNPLA3 or HSD17B13), a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupling agent or uncoupling agent, a structurally engineered fatty acid, or a combination thereof.

[0054] In some implementations, the liver disease is NAFLD, and the therapeutic interventions are vitamin E supplementation, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise programs, diet programs, bariatric surgery, or combinations thereof.

[0055] In some implementations, the liver disease is NAFLD, and the therapeutic intervention is a GLP1 receptor agonist, an FGF analog, a THR agonist, an SCD-1 inhibitor, a FAS inhibitor, an FXR agonist, an ACC inhibitor, a PPAR agonist, a gene-targeting modifier (including, for example, PNPLA3 or HSD17B13), a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclic protein inhibitor, a pan-cysteine ​​inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactolectin-3 inhibitor, a mitochondrial uncoupling agent or uncoupling agent, a structurally engineered fatty acid, or a combination thereof.

[0056] In some implementations, the method also includes monitoring the subject’s liver disease at two or more time points, at least in part based on the output.

[0057] In some implementations, the method also includes determining the likelihood or risk score that the subject has liver disease or is at an increased risk of having liver disease.

[0058] In some implementations, the method also includes determining the molecular subtype, grade, stage, or severity of liver disease.

[0059] In some implementations, the method also includes determining the prognosis of liver disease.

[0060] In some implementations, the method also includes determining the eligibility of the subject as a liver transplant donor or recipient.

[0061] In some implementations, a subject is deemed eligible as a liver transplant donor if they are not identified as having liver disease or at an increased risk of developing liver disease.

[0062] In some implementations, a subject is deemed eligible as a liver transplant recipient if they are identified as having liver disease or at an increased risk of developing liver disease.

[0063] In some implementations, the trained ML algorithm is trained using an independent set of samples associated with the presence or increased risk of liver disease.

[0064] In some implementations, the trained ML algorithm is trained using a first independent set of samples associated with the presence or increased risk of liver disease and a second independent set of samples associated with the absence or no increased risk of liver disease.

[0065] In some implementations, (c) also includes using a trained ML algorithm or another trained algorithm to process the subject’s clinical health data set.

[0066] In some implementations, clinical health data include one or more quantitative measurements selected from age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) level, alanine aminotransferase (ALT) level, gamma-glutamyl transferase (GGT) level, platelet count, triglyceride level, glycated hemoglobin (HbA1c) level, creatinine level, insulin level, prothrombin time, haptoglobin level, and glucose level.

[0067] In some implementations, clinical health data include one or more categorical measurements selected from race, ethnicity, drug history or other clinical treatment history, alcohol use history, daily activities or fitness levels, genetic test results, blood test results, and imaging results.

[0068] In some implementations, the trained ML algorithm includes a supervised ML algorithm.

[0069] In some implementations, supervised ML algorithms include classifiers or regressions.

[0070] In some implementations, supervised ML algorithms include deep learning algorithms, support vector machines (SVM), neural networks, random forests, linear regression, or logistic regression.

[0071] In some implementations, the methylation pattern or methylation level is represented by parameters of distribution, sufficient statistical data, or near-sufficient statistical data.

[0072] In another aspect, this disclosure provides a method for determining whether a subject has liver disease or is at an increased risk of developing liver disease, comprising: (a) providing a cell-free nucleic acid sample derived from the subject; (b) measuring the cell-free nucleic acid sample or a derivative thereof to determine the methylome of the cell-free nucleic acid sample; and (c) processing the methylome using a trained machine learning (ML) algorithm to determine whether the subject has liver disease or is at an increased risk of developing liver disease, wherein the determination has at least about 70% sensitivity and at least about 70% specificity. Attached Figure Description

[0073] The novel features of the invention are set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description and accompanying drawings (also referred to herein as “Figures”), which illustrate illustrative embodiments utilizing the principles of the invention, in which:

[0074] Figure 1 An example workflow for a method used to identify or monitor the liver disease status of a subject is shown.

[0075] Figure 2A computer system is shown that is programmed or otherwise configured to implement the methods provided herein.

[0076] Figure 3 A schematic diagram of example training data is shown.

[0077] Figure 4 The score distribution of cfDNA methylation data is shown to distinguish between samples with non-alcoholic steatohepatitis (NASH) and those without (healthy) NASH.

[0078] Figure 5 The score distribution of cfDNA methylation data distinguishes between risky and non-risky NASH samples, where risky NASH is defined as an individual with NASH and stage 2 or higher fibrosis.

[0079] Figure 6 The score distribution of cfDNA methylation data is shown to distinguish NASH samples with cirrhosis from those without.

[0080] Figure 7 The score distribution of cfDNA methylation data is shown to distinguish between early NASH samples, late NASH samples, and non-NASH (healthy) samples. Detailed Implementation

[0081] While various embodiments of the invention have been shown and described herein, those skilled in the art will understand that such embodiments are provided by way of example only. Various modifications, alterations, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0082] Differential patterns of nucleic acid molecules may aid in the detection or stratification of liver diseases. This article provides methods and systems for determining nucleic acids for the detection or stratification of liver diseases. For example, methylation patterns of circulating deoxyribonucleic acid (DNA) in human plasma can be detected and used to stratify the severity of liver fibrosis in patients with NAFLD.

[0083] Liver disease refers to several conditions that affect and damage the liver. There are four main stages of liver disease: 1) inflammation; 2) fibrosis; 3) cirrhosis; and 4) liver failure or liver cancer. Early liver disease can be characterized by inflammation, enlargement, or fibrosis of the liver. Over time, liver disease leads to cirrhosis (scarring). As more and more scar tissue replaces healthy liver tissue, the liver cannot function properly. If left untreated, liver disease can lead to more serious conditions such as impaired liver function and cancer. Late-stage liver disease, also known as end-stage liver disease or late-stage liver disease, can be characterized by irreversible cirrhosis, liver failure, and stage 4 hepatitis C. Fatty liver disease (SLD) encompasses all the various causes of steatosis.

[0084] Nonalcoholic fatty liver disease (NAFLD) is a common chronic pathology associated with progressive histological changes in the liver parenchyma. These NAFLD-related changes range from simple fat accumulation within hepatocytes (also known as hepatic steatosis or fatty liver) to more severe histological changes characterized by hepatocyte damage, fibrosis, and inflammation, which are hallmarks of nonalcoholic steatohepatitis (NASH). NASH is also known as metabolic dysfunction-associated steatohepatitis (MASH).

[0085] Nonalcoholic fatty liver disease (NAFLD) is a common cause of chronic liver pathology worldwide. The prevalence of NAFLD is closely associated with rising rates of diabetes, obesity, and metabolic syndrome in the general population. Simple steatosis is the earliest stage of NAFLD, usually non-progressive and asymptomatic. Appropriate lifestyle and dietary adjustments at this early stage can restore the affected liver to a healthy state. Because simple steatosis can progress to severe fibrotic stages and promote carcinogenesis, timely detection and risk stratification of NAFLD are necessary.

[0086] NAFLD is also known as metabolic dysfunction-associated steatotic liver disease (MASLD). MASLD encompasses patients with hepatic steatosis and at least one of the five cardiometabolic risk factors. In addition to MASLD alone, there is another category called metabolic and alcohol-related / associative liver disease (MetALD), which refers to MASLD patients with high weekly alcohol consumption (e.g., 140 g / week for women and 210 g / week for men). Patients with liver disease without metabolic parameters and of unknown etiology may be referred to as cryptogenic steatotic liver disease (cryptogenic SLD). The methods described in this article can be used to identify, stratify, or differentiate any type or subtype of liver disease, as described in this article and in Rinella et al. Hepatology 78(6):p 1966-1986, December 2023 DOI:10.1097 / HEP.0000000000000520, which is incorporated herein by reference in its entirety.

[0087] Extracellular circulating nucleic acids (cfDNAs) found in bodily fluids, including blood, can serve as promising non-invasive biomarkers for liver diseases. For example, epigenetic features of circulating cfDNA (such as methylation patterns) can help detect the presence of disease and monitor its progression. Intracellular miRNAs are typically involved in the regulation of gene expression, but after being released by apoptotic cells, they can remain highly stable in the extracellular environment for extended periods. Therefore, circulating nucleic acid profiles can reflect pathogenic processes in body tissues and organs, enabling highly sensitive, non-invasive detection of liver diseases.

[0088] definition

[0089] As used herein, the term "nucleic acid" generally refers to a polymer of nucleotides of any length, including deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nucleic acids can have any three-dimensional structure and can perform any known or unknown function. Non-limiting examples of nucleic acids include DNA, ribonucleic acid (RNA), coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If modifications are present, the modification of the nucleotide structure can be performed before or after nucleic acid assembly. The nucleotide sequence of a nucleic acid can be broken down by non-nucleotide components. Nucleic acids can be further modified after polymerization, such as by conjugation or binding with a reporter reagent.

[0090] As used herein, the terms “nucleic acid molecule,” “nucleic acid sequence,” “nucleic acid fragment,” “oligonucleotide,” and “polynucleotide” generally refer to polynucleotides, such as deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs and / or combinations thereof (e.g., mixtures of DNA and RNA). Nucleic acid molecules can have various lengths. The length of a nucleic acid molecule can be at least about 5 bases, 10 bases, 20 bases, 30 bases, 40 bases, 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 110 bases, 120 bases, 130 bases, 140 bases, 150 bases, 160 bases, 170 bases, 180 bases, 190 bases, 200 bases, 300 bases, 400 bases, 500 bases, 1 kilobits (kb), 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, or 50 kb, or it can have any number of bases between any two of the above values. Oligonucleotides typically consist of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (when the polynucleotide is RNA, thymine (T) corresponds to uracil (U)). Therefore, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" are intended at least in part as a letter representation of a polynucleotide molecule. Alternatively, these terms can also be used for the polynucleotide molecule itself. This letter representation can be entered into a database on a computer with a central processing unit and / or used for bioinformatics applications such as functional genomics and homology retrieval. Oligonucleotides may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.

[0091] As used herein, the terms “nucleic acid molecule,” “nucleic acid sequence,” “nucleic acid fragment,” “oligonucleotide,” and “polynucleotide” generally refer to polynucleotides, such as deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs and / or combinations thereof (e.g., mixtures of DNA and RNA). Nucleic acid molecules can have various lengths. The length of a nucleic acid molecule can be at least 5 bases, at least 10 bases, at least 20 bases, at least 30 bases, at least 40 bases, at least 50 bases, at least 60 bases, at least 70 bases, at least 80 bases, at least 90 bases, at least 100 bases, at least 110 bases, at least 120 bases, at least 130 bases, at least 140 bases, at least 150 bases, at least 160 bases, at least 170 bases, at least 180 bases, at least 190 bases, at least 200 bases, at least 300 bases, at least 400 bases, at least 500 bases, at least 1,000 bases (kb), at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 10 kb, at least 50 kb, or any number of bases between any two of the above values. Oligonucleotides typically consist of a specific sequence of four nucleotide bases: adenine (A), cytosine (C), guanine (G), and thymine (T) (when the polynucleotide is RNA, thymine (T) corresponds to uracil (U)). Therefore, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" are intended at least in part as a letter representation of a polynucleotide molecule. Alternatively, these terms can also be applied to the polynucleotide molecule itself. This letter representation can be entered into a database on a computer with a central processing unit and / or used for bioinformatics applications such as functional genomics and homology retrieval. Oligonucleotides may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.

[0092] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule with a nucleotide sequence in a starting population of nucleic acid molecules, for which the determination of the presence, quantity, and / or sequence of these nucleic acid molecules, or variations of one or more of these nucleic acid molecules, is desired. Target nucleic acids can be any type of nucleic acid, including DNA, RNA, and their analogues. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.

[0093] As used herein, the term “target” generally refers to a biomarker gene or a genomic region within a biomarker region. As used herein, the term “reference” generally refers to a sample obtained or derived from a subject diagnosed with liver disease or a subject who has received a negative clinical indication of liver disease (e.g., a healthy or control subject without liver disease).

[0094] As used herein, the terms “locus” or “region” are generally used interchangeably and refer to a specific genomic region on the genome, indicated by chromosome number, start position, and end position.

[0095] As used herein, the term "object" generally refers to an entity or medium that has testable or detectable genetic information. An object can be a person or individual, such as a patient. An object can be a vertebrate, such as a mammal. Non-limiting examples of mammals include rats, apes, humans, farm animals, sporting animals, and pets.

[0096] As used herein, the term "sample" generally refers to a biological sample, such as a sample obtained from or derived from an object. A sample can be obtained from tissues and / or cells, or from the environment of tissues and / or cells. A sample can be a cell-free biological sample or a substantially cell-free biological sample, or it can be processed or fractionated to produce a cell-free biological sample. For example, cell-free biological samples can include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and derivatives thereof. Cell-free biological samples can be obtained from or derived from an object using EDTA collection tubes, cell-free RNA collection tubes, or cell-free DNA collection tubes. Cell-free biological samples can be derived from whole blood samples through fractionation. In some embodiments, the biological sample or its derivatives may contain cells. For example, the biological sample can be a blood sample or its derivatives (e.g., blood collected via a collection tube or blood drop), a liver tissue sample, a vaginal sample (e.g., a vaginal swab), or a cervical sample (e.g., a cervical swab). In some examples, samples may comprise, be obtained from, or be derived from tissue biopsies (e.g., liver biopsies), cell biopsies, blood (e.g., whole blood), plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, feces, urine, extracellular fluid, dried blood spots, cultured cells, culture media, waste tissue, plant matter, synthetic proteins, bacterial and / or viral samples, fungal tissue, archaea, or protozoa. Samples may be isolated from their source prior to collection. Non-limiting examples include fingerprints, saliva, urine, blood, feces, semen, or other bodily fluids isolated from their original source prior to collection. In some examples, samples are isolated from their original source (cells, tissues, bodily fluids such as blood, environmental samples, etc.) during sample preparation. Samples may or may not be purified or otherwise enriched from their original source. In some embodiments, the original source is homogenized prior to further processing. Samples may be filtered or centrifuged to remove the erythrocyte sedimentation rate (ESR) layer, lipids, or particulate matter. Samples can also be purified or enriched for nucleic acids, or treated with RNase or DNase. Samples may contain intact, fragmented, or partially degraded tissues and / or cells.

[0097] Samples may be obtained from subjects who have or are suspected of having a disease or condition, and subjects may or may not have a confirmed diagnosis of the disease or condition. A second opinion may be required from the subject. A disease or condition may be an infectious disease, an immune disorder or disease, cancer, a genetic disease, a degenerative disease, a lifestyle disease, or an injury. Infectious diseases may be caused by bacteria, viruses, fungi, and / or parasites. Cancer may be hepatocellular carcinoma (HCC) or hepatobiliary cancer, such as cholangiocarcinoma, angiosarcoma, gallbladder cancer, or undifferentiated embryonal sarcoma of the liver (UESL).

[0098] Sample components (including nucleic acids) can be labeled, for example using identifiable markers, to achieve sample multiplexing. Some non-limiting examples of identifiable markers include fluorophores, magnetic nanoparticles, and nucleic acid barcodes. Fluorophores can include fluorescent proteins such as GFP, YFP, RFP, eGFP, mCherry, tdtomato, FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 680, Alexa Fluor 750, Pacific Blue, coumarin, BODIPY FL, Pacific Green, Oregon Green, Cy3, Cy5, PacificOrange, TRITC, Texas Red, phycoerythrin, allophycocyanin, or other fluorophores. Prior to sequencing, one or more barcode tags (e.g., by coupling or linking) can be attached to cell-free nucleic acids (e.g., cfDNA) in a sample. The barcode can uniquely tag the cfDNA molecules in the sample. Alternatively, the barcode can non-uniquely tag the cfDNA molecules in the sample. The non-unique tagging of the barcode allows additional information obtained from the cfDNA molecules (e.g., at least a portion of the endogenous sequence of the cfDNA molecule) to be used, in conjunction with the non-unique tag, as a unique identifier for the cfDNA molecules in the sample (e.g., for unique identification from other molecules). For example, a uniquely identified cfDNA sequence read (e.g., from a given template molecule) can be detected at least in part based on sequence information containing one or more consecutive base regions at one or both ends of the read, the length of the read, and / or the sequence of the barcode attached to one or both ends of the read. DNA molecules can be uniquely identified without labeling by dividing a DNA (e.g., cfDNA) sample into many (e.g., at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, or at least about 100,000) distinct discrete subunits (e.g., partitions, pores, or droplets) before amplification, such that the amplified DNA molecules can be uniquely distinguished and identified as originating from their respective individual input DNA molecules.

[0099] Multiplexing can be performed on any number of samples. For example, multiplexing analysis can include at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 or more samples. Identifiable markers can provide a method to trace the origin of each sample or can guide different samples to different regions or solid supports.

[0100] Any number of samples can be mixed before analysis without labeling or multiplexing. For example, multiplexing analysis can contain at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 or more samples. Sample multiplexing without labeling can be performed using a combined cell design, where samples are mixed into cells in a manner that allows the signals from individual samples to be resolved from the analyzed cells using computational solution multiplexing.

[0101] Samples can be enriched prior to sequencing. For example, cfDNA molecules can selectively or non-selectively enrich one or more regions of the genome or transcriptome of a target genome. For example, cfDNA molecules can selectively enrich one or more regions of the genome or transcriptome of a target genome through targeted sequence capture (e.g., using kits), selective amplification, or targeted amplification. As another example, cfDNA molecules can non-selectively enrich one or more regions of the target genome or transcriptome through universal amplification. In some embodiments, amplification includes universal amplification, whole-genome amplification, or non-selective amplification. cfDNA molecules can be size-selected, selecting fragments within a predetermined length range. For example, DNA fragments can be size-selected prior to ligation, with lengths ranging from approximately 40 base pairs (bp) to approximately 250 bp. As another example, DNA fragments can be size-selected after ligation, with lengths ranging from approximately 160 bp to approximately 400 bp.

[0102] As used herein, “amplifying” and “amplification” are used interchangeably and generally refer to the generation of one or more copies of nucleic acid or “amplification products.” The term “DNA amplification” generally refers to the generation of one or more copies of a DNA molecule or “amplified DNA products.” The term “reverse transcription amplification” generally refers to the generation of deoxyribonucleic acid (DNA) from a glyconucleotide (RNA) template via the action of reverse transcriptase. Amplification can be performed using polymerase chain reaction (PCR), a reaction based on the synthesis of a new DNA strand complementary to the initial template strand using DNA polymerase.

[0103] As used herein, the term "polymerase chain reaction" or "PCR" generally refers to a method for increasing the concentration of a fragment of a target sequence in a mixture of genomic DNA without cloning or purification. This process for amplifying the target sequence may involve introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target sequence, followed by a series of precise thermal cycles in the presence of a DNA polymerase. The two primers can be complementary to the other strand of the double-stranded target sequence. For amplification, the mixture can be denatured, and the primers annealed to their complementary sequences within the target molecule. After annealing, the primers can be extended with a polymerase to form a new pair of complementary strands. Denaturation, primer annealing, and polymerase extension can be repeated multiple times (e.g., denaturation, annealing, and extension constitute a "cycle"; multiple "cycles" may be possible) to obtain a high concentration of amplified fragments of the desired target sequence. The length of the amplified fragment of the desired target sequence is determined by the relative positions of the primers, and thus this length is a controllable parameter. Due to the reproducible aspect of this process, the method is called "polymerase chain reaction" or "PCR". Since the amplified fragment of the desired target sequence becomes the dominant sequence in the mixture (in terms of concentration), the amplified fragment can be referred to as a "PCR amplite," "PCR product," or "amplifier."

[0104] As used herein, the term "methylation" refers to 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC), which includes cytosine residues that are part of the sequence CG (also known as CpG dinucleotides). Some CG dinucleotides in the human genome are methylated, while others are not. Furthermore, methylation can be cell-specific and tissue-specific, allowing a particular CG dinucleotide to be methylated in one cell while remaining unmethylated in another; or methylated in one tissue while remaining unmethylated in another. DNA methylation is a crucial regulator of gene transcription. Abnormal DNA methylation patterns (hypermethylation and hypomethylation) compared to normal tissues may be associated with a large number of human malignancies. In some embodiments, the 5hmC residues of the sequence can be glycosylated, followed by bisulfite treatment, bisulfite-free treatment, or digestion with a methylation-sensitive restriction enzyme. For example, glycosylation can be performed using glucosyltransferases.

[0105] As used herein, the terms “methylation state,” “methylation status,” and “methylation profile” generally refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine (e.g., a DNA molecule) is considered methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered unmethylated.

[0106] As used in this article, "DNA template" usually refers to sample DNA containing the target sequence. At the start of the reaction, high temperature is applied to the original double-stranded DNA molecules to cause the strands to separate.

[0107] As used in this article, the term "primer" generally refers to a short segment of single-stranded DNA that is complementary to the DNA template. Polymerase synthesizes new DNA starting from the end of the primer.

[0108] As used herein, the term "sensitivity" or "clinical sensitivity" generally refers to the percentage of positive diagnostic results obtained from a collection of diseased samples. For example, such diseased samples may be analyzed to detect DNA methylation values ​​above a threshold that distinguishes between disease (e.g., liver disease) and non-disease (e.g., healthy or control) samples. In some embodiments, a positive result is defined as a histologically confirmed disease with a reported DNA methylation value above the threshold (e.g., a disease-related range); and a false negative is defined as a histologically confirmed disease with a reported DNA methylation value below the threshold (e.g., a non-disease-related range). Sensitivity values ​​can reflect the probability that a DNA methylation measurement of a given biomarker obtained from a diseased sample falls within the disease-related measurement range. The clinical relevance of the calculated sensitivity value can represent an estimate of the probability that a given biomarker can detect or predict the presence of a clinical condition when applied to subjects with a clinical condition.

[0109] As used herein, the term "specificity" or "clinical specificity" generally refers to the percentage of negative diagnostic results obtained from a collection of non-disease samples. For example, such non-disease samples can be analyzed to detect DNA methylation values ​​below a threshold that distinguishes between diseased (e.g., liver disease) and non-diseased (e.g., non-liver disease) samples. In some embodiments, a negative result is defined as a histologically confirmed non-disease sample that reports a DNA methylation value below the threshold (e.g., within the range associated with no disease), and a false positive result is defined as a histologically confirmed non-disease sample that reports a DNA methylation value above the threshold (e.g., within the range associated with disease). A specificity value can reflect the probability that a DNA methylation measurement of a given biomarker obtained from a non-liver disease (e.g., healthy or control) sample falls within the range of non-disease-related measurements. The clinical relevance of a calculated specificity value can represent an estimate of the probability that a given biomarker can detect or predict the absence of a clinical condition when applied to subjects without a clinical condition.

[0110] As used herein, the term “AUC” or “AUROC” generally refers to the area under the receiver operating characteristic (ROC) curve. An ROC curve can be a graph of the true positive rate (TPR) versus the false positive rate (FPR) at multiple different thresholds or cutoffs for a diagnostic test, illustrating the trade-off between sensitivity and specificity depending on the chosen cutoff (e.g., any increase in sensitivity is accompanied by a decrease in specificity). The area under the ROC curve (AUC) can be a measure of the accuracy of a diagnostic test (e.g., the larger the area, the more accurate the diagnosis), with an optimal value of 1. In contrast, the ROC curve for a randomized test may lie diagonally, with an AUC of 0.5 (e.g., indicating a randomized or worthless test).

[0111] The method disclosed herein

[0112] Current diagnostic tools for liver disease may be difficult to obtain and incomplete. Blood tests can be used to measure the levels of enzyme biomarkers in the blood. Liver function tests, such as the international normalized ratio (INR), can be used to assess the extent of coagulopathy, an indicator of liver dysfunction. Imaging tools such as ultrasound, MRI, or CT can be used to visualize signs of liver damage, scarring, or tumors. Liver biopsy is currently the gold standard for evaluating liver fibrosis in patients with fatty liver. However, the inherent risks and invasiveness of biopsy evaluation limit its widespread application. Therefore, there is an urgent clinical need for accurate, cost-effective, and non-invasive diagnostic methods to detect and monitor liver disease, thereby enabling effective disease management and treatment.

[0113] This disclosure provides methods, systems, and kits for identifying or monitoring liver diseases by processing cell-free biological samples obtained from or derived from a subject. Cell-free biological samples (e.g., plasma samples) obtained from a subject can be analyzed to identify liver diseases, which may include, for example, measuring the presence, absence, or relative assessment of liver disease. Such subjects may include subjects with one or more liver diseases and subjects without one or more liver diseases. Liver diseases may include, for example, alcoholic or non-alcoholic fatty liver disease, non-alcoholic steatohepatitis, hepatitis, cancer (e.g., hepatocellular carcinoma), and cirrhosis.

[0114] Figure 1An example workflow of a method for identifying or monitoring the liver disease state of a subject according to embodiments disclosed herein is illustrated. In one aspect, this disclosure provides a method 100 for identifying or monitoring the liver disease state of a subject. Method 100 may include measuring a first cell-free biological sample derived from the subject by a first assay to generate a first dataset (operation 101). Next, based at least in part on the generated first dataset, method 100 may optionally include measuring a second cell-free biological sample derived from the subject by a second assay (e.g., an assay different from the first assay) to generate a second dataset indicating the liver disease state with higher specificity than the first dataset (operation 102). For example, DNA molecules extracted from the second cell-free plasma sample may be sequenced to generate a set of sequence reads indicating the liver disease state of the subject. In some embodiments, a first cell-free biological sample is obtained from the subject at a first time point for treatment with the first assay. Then, optionally, a second cell-free biological sample is obtained from the same subject at a second time point for treatment with the second assay. In some implementations, cell-free biological samples can be obtained from the subject and then aliquoted to produce a first cell-free biological sample and a second cell-free biological sample, which are then processed with a first assay and a second assay, respectively. Next, a trained machine learning algorithm can be used to process the first and / or second datasets to determine the subject's liver disease status (operation 103). The trained machine learning algorithm can be configured to identify liver disease with at least approximately 80% accuracy in 50 independent samples. A report can then be generated electronically indicating (e.g., identifying or providing an indication of) the presence or susceptibility to liver disease in the subject (operation 104).

[0115] Cell-free biological samples can be obtained from subjects with a liver disease state (e.g., liver disease or condition), subjects suspected of having a liver disease state, or subjects that do not have and are not suspected of having a liver disease state. The disease or condition can be a disease or condition affecting the liver. Non-limiting examples of such diseases or conditions include fatty liver disease, alcoholic fatty liver disease, non-alcoholic fatty liver disease, steatohepatitis, non-alcoholic steatohepatitis, hepatitis (e.g., hepatitis A, hepatitis B, or hepatitis C), liver cancer (e.g., hepatocellular carcinoma), cholangiocarcinoma (including, for example, bile duct cancer, angiosarcoma, gallbladder cancer, or undifferentiated embryonic sarcoma of the liver (UESL)), cirrhosis, hemochromatosis, Wilson's disease, obesity, diabetes, hypertension, and other liver conditions disclosed herein.

[0116] Samples may be obtained before and / or after treatment of a subject with a disease or condition. Samples may be obtained during treatment or throughout the treatment regimen. Multiple samples may be obtained from a subject to monitor the effects of treatment over time, including samples taken before treatment begins. Samples may also be obtained from a subject to monitor abnormal tissue-specific cell death or organ transplantation.

[0117] Samples may be obtained from subjects suspected of having a disease or condition. Samples may be obtained from subjects exhibiting unexplained symptoms such as fatigue, nausea or vomiting, yellowing of the skin or eyes (jaundice), swelling of the legs or ankles, abdominal swelling (ascites), abdominal pain, itchy skin, weight gain, weight loss, pain, aches, tremors, weakness, lethargy, disorientation, or confusion. Samples may be obtained from subjects exhibiting explained symptoms. Samples may be obtained from subjects at risk of developing a disease or condition due to one or more factors such as family and / or personal history, age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) level, alanine aminotransferase (ALT) level, gamma-glutamyl transferase (GGT), platelet count, triglyceride level, haptoglobin level, glucose level, environmental exposure, lifestyle risk factors, the presence of other risk factors, or a combination thereof.

[0118] Samples can be obtained from healthy subjects or individuals. In some embodiments, samples can be obtained longitudinally from the same subject or individual. In some embodiments, longitudinally acquired samples can be analyzed for the purpose of monitoring an individual's health status and detecting health problems early (e.g., early diagnosis of liver disease). In some embodiments, samples can be collected at home or a point of care and then transported by mail, courier, or other means of transport for analysis. For example, a home user can collect a bloodstain sample by pricking their finger. The bloodstain sample can be dried, then transported by mail for analysis. In some embodiments, longitudinally acquired samples can be used to monitor responses to stimuli expected to affect health, athletic performance, or cognitive performance. Non-limiting examples include responses to medications, diets, and / or exercise programs. In some embodiments, individual samples have multiple uses, allowing for methylation profiling analysis to obtain clinically relevant information, but also for obtaining information about an individual's personal or family ancestry.

[0119] In some implementations, a biological sample is a nucleic acid sample containing one or more nucleic acid molecules. The nucleic acid molecules can be cell-free or substantially cell-free nucleic acid molecules, such as cell-free DNA (cfDNA) or cell-free RNA (cfRNA) or mixtures thereof. Nucleic acid molecules can be derived from a variety of sources, including human, mammalian, non-human mammalian, ape, monkey, chimpanzee, reptile, amphibian, or avian sources. Furthermore, samples can be extracted from a variety of animal fluids containing cell-free sequences, including but not limited to blood, serum, plasma, bone marrow, vitreous humor, sputum, feces, urine, tears, sweat, saliva, semen, mucosal secretions, mucus, cerebrospinal fluid, pleural fluid, amniotic fluid, and lymph.

[0120] Cell-free biological samples may contain one or more analytes that can be measured, such as cfRNA molecules suitable for measurement to generate transcriptome data, cfDNA molecules suitable for measurement to generate genomic data, proteins suitable for measurement to generate proteome data, metabolites suitable for measurement to generate metabolome data, or mixtures or combinations thereof. One or more such analytes (e.g., cfRNA molecules, cfDNA molecules, proteins, or metabolites) can be isolated or extracted from one or more cell-free biological samples of the subject for use in downstream assays using one or more suitable assays.

[0121] After obtaining cell-free biological samples from a subject, the samples can be processed to generate a dataset indicative of the subject's liver disease status. For example, proteomics data (e.g., quantitative measurements of DNA or RNA transcripts at liver disease-related genomic loci), proteomic data (quantitative measurements of proteins in a dataset containing liver disease-related proteins), and / or metabolomics data (quantitative measurements of metabolites in a dataset containing liver disease-related metabolites) can indicate liver disease status. Processing cell-free biological samples obtained from a subject may include: (i) subjecting the sample to conditions sufficient to isolate, enrich, or extract multiple nucleic acid molecules, proteins, and / or metabolites; and (ii) measuring multiple nucleic acid molecules, proteins, and / or metabolites to generate a dataset. In some embodiments, the quantitative measurement of DNA may include the presence, absence, or extent of methylation, hypermethylation, and / or hypomethylation. Alternatively, or in combination, the quantitative measurement of DNA may include the presence, absence, or extent of variant patterns. Variant patterns may include gene mutations, single nucleotide polymorphisms (SNPs), or copy number variations. Alternatively, or in combination, quantitative measurements of DNA may include the presence, absence, or extent of viral genomic patterns.

[0122] In some implementations, multiple nucleic acid molecules are extracted from cell-free biological samples and sequenced to generate multiple sequencing reads. The nucleic acid molecules may contain RNA or DNA. Nucleic acid molecules (e.g., RNA or DNA) can be extracted from cell-free biological samples using various methods, such as nucleic acid extraction kits. Extraction methods can extract all RNA or DNA molecules from a sample. Alternatively, extraction methods can selectively extract a portion of RNA or DNA molecules from a sample. RNA molecules extracted from the sample can be converted into DNA molecules via reverse transcription (RT).

[0123] Nucleic acid molecules can be sequenced using any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), ligation sequencing, hybridization sequencing, and RNA-Seq (Illumina).

[0124] Sequencing can include nucleic acid amplification (e.g., amplification of RNA or DNA molecules). In some implementations, nucleic acid amplification is polymerase chain reaction (PCR). Appropriate rounds of PCR (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to adequately amplify the initial amount of nucleic acid (e.g., RNA or DNA) to the input required for subsequent sequencing. In some cases, PCR can be used for global amplification of target nucleic acids. This amplification can include using a ligand sequence that can be ligated to different molecules before PCR amplification using universal primers. PCR can be performed using a variety of commercial kits, such as those offered by Life Technologies, Affymetrix, Promega, Qiagen, etc. In other cases, only certain target nucleic acids in a population of nucleic acids can be amplified. Specific primers (which may bind to ligands) can be used to selectively amplify certain targets for downstream sequencing. PCR can include targeted amplification of one or more genomic loci, such as genomic loci associated with liver diseases. Sequencing can include the simultaneous use of RT and PCR, such as the OneStep RT-PCR kit protocols from Qiagen, NEB, Thermo Fisher Scientific, or Bio-Rad.

[0125] RNA or DNA molecules isolated or extracted from cell-free biological samples can be labeled, for example, using identifiable markers, to allow multiplexing of multiple samples. Any number of RNA or DNA samples can be multiplexed. For example, a multiplexing reaction can involve RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more initial cell-free biological samples. For example, multiple cell-free biological samples can be labeled with sample barcodes, allowing each DNA molecule to be traced back to the sample (and object) from which it originated. Such markers can be attached to RNA or DNA molecules by ligation or by PCR amplification using primers.

[0126] After sequencing nucleic acid molecules, appropriate bioinformatics processing can be performed on the sequence reads to generate data indicating the presence, absence, or relative assessment of liver disease. For example, the sequence reads can be aligned with one or more reference genomes (e.g., genomes of one or more species, such as the human genome). The aligned sequence reads can then be quantified at one or more genomic loci to generate a dataset indicating liver disease. For instance, quantifying sequences corresponding to multiple genomic loci associated with liver disease can generate a dataset indicating liver disease.

[0127] In some cases, cell-free biological samples can be processed without any nucleic acid extraction. For example, liver disease in a subject can be identified or monitored using probes configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple liver disease-associated genomic loci. The probes can be nucleic acid primers. The probes can have sequence complementarity with one or more nucleic acid sequences from multiple liver disease-associated genomic loci or genomic regions. Multiple liver disease-related genomic loci or regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100 or more different liver disease-related genomic loci or regions. Multiple liver disease-associated genomic loci or regions may contain one or more members selected from the genes listed in Table 1 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, approximately 25, approximately 30, approximately 35, approximately 40, approximately 45, approximately 50, approximately 55, approximately 60, approximately 65, approximately 70, approximately 75, approximately 80, approximately 85, approximately 90, approximately 95, approximately 100, approximately 200, approximately 300, approximately 400, approximately 500, approximately 600, approximately 700, approximately 800, approximately 900, approximately 1000, or more). Liver disease-associated genomic loci or regions may be associated with age, race, ethnicity, BMI, blood glucose levels, or other liver disease states or complications.

[0128] Table 1

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157] Probes can be nucleic acid molecules (e.g., RNA or DNA) that are sequence complementary to nucleic acid sequences (e.g., RNA or DNA) of one or more genomic loci (e.g., liver disease-related genomic loci). These nucleic acid molecules can be primers or enriched sequences. Assaying cell-free biological samples using probes selective for one or more genomic loci (e.g., liver disease-related genomic loci) can include using array hybridization (e.g., microarray-based hybridization), PCR, or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some implementations, DNA or RNA can be measured by one or more of the following methods: isothermal DNA / RNA amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassay, electrochemical assay, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assay, molecular inverted probe, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (e.g., CRISPR genotyping PCR (ctPCR), specific high-sensitivity enzyme reporter gene unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR trans reporter gene (DETECTR), and CRISPR-mediated simulated multiple event recording device (CAMERA)), and laser transmission spectroscopy (LTS).

[0158] Assay readouts can be performed to quantify one or more genomic loci (e.g., liver disease-associated genomic loci) to generate data indicating liver disease status. For example, quantification of array hybridization or PCR corresponding to multiple genomic loci (e.g., liver disease-associated genomic loci) can generate data indicating liver disease status. Assay readouts can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values ​​thereof. The assay can be a home test configured for home use.

[0159] In some implementations, multiple assays are used to process cell-free biological samples of the subject. For example, a first assay may be used to process a first cell-free biological sample obtained or derived from the subject to generate a first dataset; and based at least in part on the first dataset, a second assay, different from the first assay, may be used to process a second cell-free biological sample obtained or derived from the subject to generate a second dataset indicative of liver disease status. The first assay may be used to screen or process cell-free biological samples of a set of subjects, while a second or subsequent assay may be used to screen or process cell-free biological samples of a smaller subset of the set of subjects. The first assay may be low-cost and / or highly sensitive in detecting one or more liver disease states (e.g., liver disease or condition), and is suitable for screening or processing cell-free biological samples of a relatively large set of subjects. The second assay may be higher-cost and / or more specific in detecting one or more liver disease states, and is suitable for screening or processing cell-free biological samples of a relatively small set of subjects (e.g., a subset of subjects screened using the first assay). The second assay may generate a second dataset with higher specificity (e.g., for one or more liver disease states) than the first dataset generated using the first assay. As an example, one or more cell-free biological samples can be treated on a large set of objects using cfDNA assays, and then one or more cell-free biological samples can be treated on a smaller subset of objects using metabolomics assays, or vice versa. The smaller subset of objects can be selected at least in part based on the results of the first assay.

[0160] Alternatively, multiple assays can be used to process cell-free biological samples from the subject simultaneously. For example, a first assay can be used to process a first cell-free biological sample obtained or derived from the subject to generate a first dataset indicative of liver disease status; and a second assay, different from the first assay, can be used to process a second cell-free biological sample obtained or derived from the subject to generate a second dataset indicative of liver disease status. Either or both of the first and second datasets can then be analyzed to assess the subject's liver disease status. For example, a single diagnostic index or score can be generated based on a combination of the first and second datasets. As another example, separate diagnostic indices or scores can be generated based on the first and second datasets.

[0161] Metabolomics assays can be used to process cell-free biological samples. For example, metabolomics analysis can be used to identify quantitative measurements (e.g., indicating presence, absence, or relative amount) of each of a variety of liver disease-related metabolites in a target cell-free biological sample. Metabolomics analysis can be configured to process a target cell-free biological sample, such as a blood sample (or a derivative thereof). Quantitative measurements (e.g., indicating presence, absence, or relative amount) of liver disease-related metabolites in a cell-free biological sample can indicate one or more liver diseases. Metabolites in a cell-free biological sample can be generated as a result of one or more metabolic pathways corresponding to liver disease-related genes (e.g., as a final product or byproduct). Determining one or more metabolites in a cell-free biological sample can include isolating or extracting metabolites from the cell-free biological sample. Metabolomics analysis can be used to generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of each of a variety of liver disease-related metabolites in a target cell-free biological sample.

[0162] Metabolomics assays can analyze a wide range of metabolites in cell-free biological samples, such as small molecules, lipids, amino acids, peptides, nucleotides, hormones and other signaling molecules, cytokines, minerals and elements, polyphenols, fatty acids, dicarboxylic acids, alcohols and polyols, alkanes and alkenes, keto acids, glycolipids, carbohydrates, hydroxy acids, purines, prostaglandins, catecholamines, acyl phosphates, phospholipids, cyclic amines, amino ketones, nucleosides, glycerides, aromatic acids, retinoids, amino alcohols, pterin, steroids, carnitine, leukotrienes, indoles, porphyrins, glycophosphates, coenzyme A derivatives, glucuronic acids, ketones, glycophosphates, inorganic ions and gases, sphingolipids, bile acids, alcohol phosphates, amino acid phosphates, aldehydes, quinones, pyrimidines, pyridoxal, tricarboxylic acids, acylglycine, cobalamin derivatives, lipoamide, biotin, and polyamines.

[0163] Metabolomics analysis may include one or more of the following: mass spectrometry (MS), targeted MS, gas chromatography (GC), high performance liquid chromatography (HPLC), capillary electrophoresis (CE), nuclear magnetic resonance (NMR) spectroscopy, ion mobility spectrometry, Raman spectroscopy, electrochemical assays, or immunoassays.

[0164] Methylation-specific assays can be used to process cell-free biological samples. For example, methylation-specific assays can be used to identify quantitative measurements (e.g., indicating presence, absence, or relative amount) of methylation at each of multiple liver disease-related genomic loci in a cell-free biological sample of a subject. Additionally, or alternatively, methylation-specific assays can be used to identify qualitative measurements (e.g., based on relative amount methylation patterns) of methylation at multiple liver disease-related genomic loci in a cell-free biological sample of a subject. Methylation-specific assays can be used to process cell-free biological samples of a subject, such as blood samples (or derivatives thereof). Quantitative measurements of methylation at liver disease-related genomic loci in a cell-free biological sample (e.g., indicating presence, absence, or relative amount) can indicate one or more liver disease states. Qualitative measurements of methylation at liver disease-related genomic loci in a cell-free biological sample (e.g., based on relative amount methylation patterns) can indicate one or more liver disease states. Methylation-specific assays can be used to generate datasets indicating quantitative and / or qualitative measurements of methylation at each of multiple liver disease-related genomic loci in a cell-free biological sample of a subject.

[0165] Methylation-specific assays may include one or more of the following: methylation-sensing sequencing (e.g., with or without bisulfite treatment), enzyme methylation sequencing, methylation-specific PCR (MSP), methylation-sensitive restriction enzyme (MSRE) digestion, pyrosequencing, methylation-sensitive single-strand conformation analysis (MS-SSCA), high-resolution melting analysis (HRM), methylation-sensitive single nucleotide primer extension (MS-SnuPE), base-specific cleavage / MALDI-TOF, microarray-based methylation assays, methylation-specific PCR, targeted bisulfite sequencing, oxidized bisulfite sequencing, mass spectrometry-based bisulfite sequencing, or reduced representative bisulfite sequencing (RRBS).

[0166] Bisulfite sequencing or processing involves treating DNA with a bisulfite (such as sodium bisulfite) to convert cytosine residues into uracil residues, while 5-methylcytosine residues remain unaffected. Therefore, bisulfite-treated DNA retains only the methylated cytosine residues.

[0167] Targeted bisulfite sequencing involves hybridization, where pre-designed oligonucleotides can be used to probe or target specific genomic regions of interest, such as CpG islands, gene promoters, and other important methylation regions (e.g., liver disease-related genomic loci). Targeted bisulfite sequencing can include amplification to amplify multiple bisulfite-converted DNA regions in a single reaction. Specific primers can be designed to capture regions of interest and evaluate site-specific DNA methylation patterns.

[0168] Pyrosequencing is a synthetic sequencing method that quantitatively monitors the real-time incorporation of nucleotides by converting released pyrophosphatase into corresponding optical signals. Analyzing DNA methylation patterns using pyrosequencing combines a simple reaction protocol with reproducible and accurate measurements of the degree of methylation at multiple adjacent CpGs, offering high quantitative resolution. After bisulfite treatment and PCR amplification, the degree of methylation at each CpG position in the sequence can be determined based on the T / C ratio. The purification and sequencing process can be repeated on the same template to analyze other CpGs in the same amplification product.

[0169] RRBS is a high-efficiency, high-throughput technique for analyzing whole-genome methylation profiles at the single-nucleotide level. RRBS can be combined with restriction enzyme and bisulfite sequencing to enrich regions of the genome with high CpG content. RRBS can reduce the amount of nucleotides required for sequencing to 1% of the genome. The fragment containing the reduced genome can still contain most promoters, as well as regions that are difficult to analyze using conventional bisulfite sequencing methods (such as repetitive sequences).

[0170] In some cases, bisulfite conversion methods can damage sample DNA, leading to fragmentation, loss, and bias, thus limiting their practicality. Bisulfite-free methylation sequencing methods can convert methylated cytosine while minimizing these drawbacks. For example, bisulfite-free methylation sequencing of cfDNA may be advantageous because the concentration of cfDNA in plasma can be very low and it can be a limiting resource in liquid biopsy applications.

[0171] Enzymatic methylation sequencing provides a bisulfite-free method that minimizes damage to sample DNA for methylation detection. Such enzymatic methods can offer higher mapping efficiency, more uniform GC coverage, detection of more CpG with fewer sequence reads, and a more uniform dinucleotide distribution. Enzymatic methylation sequencing methods may include treatment with methylcytosine dioxygenases (such as deca-undeca transloses (TET)); glucosyltransferases (such as β-glucosyltransferase (BGT)); and / or cytidine deaminases (such as activation-induced (cytidine) deaminase (AID) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC)). Methylcytosine dioxygenases can be used to convert 5mC and 5hmC residues to 5caC to protect these methylated residues from deamination in downstream processing operations. Non-limiting examples of methylcytosine dioxygenases include TET1, TET2, TET3, and their catalytically active variants or fusion proteins. Glucosyltransferases can be used to add glucose residues to 5hmC while protecting these methylated residues from downstream deamination. Cytidine deaminases can be used to deaminate 5mC residues to uracil and 5hmC residues to thymine. Non-restrictive examples of cytidine deaminases include APOBEC3A and its catalytically active variants or fusion proteins. Combinations of one or more enzymes can be used for bisulfite-free methylation sequencing.

[0172] TET-assisted pyridineborane sequencing (TAPS) uses the TET enzyme to oxidize the 5mC and 5hmC residues to 5caC. Then, pyridineborane is used to reduce 5caC to dihydrouracil, which is then converted to thymine after amplification. TAPS can be performed in two other ways: TAPSβ and chemically assisted pyridineborane sequencing (CAPS). In TAPSβ, β-glucosyltransferase is used to tagged 5hmC with glucose to protect it from oxidation and reduction reactions, thus enabling specific detection of 5mC. In CAPS, potassium perruthenate is used as a chemical substitute for TET and specifically oxidizes 5hmC, allowing for direct detection of 5hmC.

[0173] Methylation-specific PCR (MSP) is a qualitative analysis of DNA methylation. MSP offers advantages such as ease of design and execution, high sensitivity for detecting small amounts of methylated DNA, and rapid screening of large numbers of samples without the need for expensive laboratory equipment. This assay may require modification of genomic DNA with sodium bisulfite, followed by PCR amplification using two separate primer sets: one pair designed to recognize the methylated version of the bisulfite-modified sequence, and the other pair designed to recognize the unmethylated version. The amplicons are visualized after agarose gel electrophoresis using ethidium bromide staining. An amplicon of the expected size produced by either primer pair indicates the presence of DNA with the corresponding methylation state in the original sample.

[0174] In some implementations, methylation-sensitive restriction enzyme (MSRE) digestion can be used to analyze the methylation status of cytosine residues in CpG sequences. These enzymes may be unable to cleave methylated cytosine residues, thus leaving the methylated DNA fragment intact. Sample DNA obtained from or derived from a subject can be digested with one or more MSREs. For example, the liver disease-related genomic loci described herein may contain at least one specific MSRE recognition sequence (recognition site). Sample DNA can be cut (digested) according to its methylation level, where higher methylation levels result in lower levels of enzyme digestion. For example, if a DNA sample from a healthy subject has a lower level of CpG methylation at the recognition sequence than another DNA sample from a patient with liver disease, that DNA can be cut to a greater extent.

[0175] For example, DNA molecules can be extracted from a biological sample. A first portion of the extracted DNA molecule may be fragmented at CpG sites, such as by MSRE digestion, while a second portion may not undergo such fragmentation. Next, qPCR amplification (e.g., using qPCR primers) can be performed on at least one biomarker locus (i.e., an internal reference locus). A cycle threshold (Ct) value for each amplified region in a set of genomic regions (e.g., liver disease-related biomarkers) can be obtained and normalized based on the internal reference locus. The qPCR signal intensity of the biomarker locus can be calculated, where signal intensity = 2^[Ct, biomarker restriction locus - Ct, internal reference locus]. A probability score can then be calculated reflecting the correlation between the subject's biomarker signal intensity and a "disease" reference value and / or the correlation between the subject's biomarker signal intensity and a "healthy" reference value.

[0176] In some implementations, control loci can be designed to exclude MSRE restriction sites. In some implementations, a fixed proportion of control DNA is added to the sample DNA of all subjects. In some implementations, at least one pair of qPCR primers is designed for each target genomic region of the biomarker. For each patient, two qPCR reactions are run independently on the same qPCR target: the first qPCR reaction is run on a first portion of the sample DNA containing the MSRE-digested DNA template, and the second qPCR reaction is run on a second portion of the sample DNA containing the undigested DNA template. The undigested template can be used to represent fully methylated DNA. After MSRE digestion and purification, the same amount of DNA can be used for both the digested and undigested templates. The signal intensity of the qPCR reaction can be generated based on a cycle threshold (Ct) value. The Ct value refers to the number of cycles required for the fluorescence signal to cross a given cycle threshold (e.g., the signal exceeds the background level). The Ct level can be inversely proportional to the amount of target nucleic acid in the sample (e.g., the lower the Ct level of a given sample, the greater the amount of target nucleic acid in the sample). For each locus in a given sample, the difference in Ct values ​​(ΔCt) between the first qPCR reaction (run on a digested DNA template) and the second qPCR reaction (run on an undigested DNA template) can be calculated and used to indicate the DNA methylation level of the sample. Therefore, the ΔCt value can represent the DNA methylation level of the target region of the subject. For example, undigested DNA may have a lower Ct value, while digested DNA from a normal individual may have a higher Ct value, resulting in a larger absolute ΔCt value. Conversely, the ΔCt value in a subject with liver disease may be smaller (e.g., close to 0).

[0177] Proteomics assays can be used to process cell-free biological samples. For example, proteomics assays can be used to identify quantitative measurements (e.g., indicating presence, absence, or relative amount) of each of a variety of liver disease-related proteins or peptides in a cell-free biological sample of a subject. Proteomics assays can be configured to process cell-free biological samples of a subject, such as blood samples (or derivatives thereof). Quantitative measurements (e.g., indicating presence, absence, or relative amount) of liver disease-related proteins or peptides in a cell-free biological sample can indicate one or more liver disease states. Proteins or peptides in a cell-free biological sample can be produced as a result of one or more biochemical pathways corresponding to liver disease-related genes (e.g., as a final product or byproduct). Measuring one or more proteins or peptides in a cell-free biological sample can include isolating or extracting proteins or peptides from the cell-free biological sample. Proteomics assays can be used to generate datasets indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of each of a variety of liver disease-related proteins or peptides in a cell-free biological sample of a subject.

[0178] Proteomics assays can analyze a variety of proteins or peptides in cell-free biological samples, such as proteins produced under different cellular conditions (e.g., development, cell differentiation, or cell cycle). Proteomics assays can include one or more of the following: antibody-based immunoassays, Edman degradation assays, mass spectrometry-based assays (e.g., matrix-assisted laser desorption / ionization (MALDI) and electrospray ionization (ESI)), top-down proteomics assays, bottom-up proteomics assays, mass spectrometry immunoassays (MSIA), stable isotope standard and antipeptide antibody capture (SISCAPA) assays, fluorescence two-dimensional differential gel electrophoresis (2-DDIGE) assays, quantitative proteomics assays, protein microarray assays, or reversed-phase protein microarray assays. Proteomics assays can detect post-translational modifications of proteins or peptides (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, and nitrosation). Proteomics assays can identify or quantify one or more proteins or peptides from databases (e.g., Human Protein Atlas, Peptide Atlas, and UniProt).

[0179] Reagent test kit

[0180] This disclosure provides a kit for identifying or monitoring liver disease states in a subject. The kit may contain probes for quantitative measurements (e.g., indicating presence, absence, or relative amount) of sequences at each of multiple liver disease-related genomic loci in a cell-free biological sample of the subject. Quantitative measurements (e.g., indicating presence, absence, or relative amount) of sequences at each of multiple liver disease-related genomic loci in the cell-free biological sample may indicate one or more liver disease states. The probes may be selective for sequences at multiple liver disease-related genomic loci in the cell-free biological sample. The kit may include instructions for processing the cell-free biological sample using the probes to generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of sequences at each of multiple liver disease-related genomic loci in the cell-free biological sample of the subject.

[0181] The probes in the kit can be selective for sequences at multiple liver disease-associated genomic loci in cell-free biological samples. The probes in the kit can be configured to selectively enrich nucleic acid molecules (e.g., RNA or DNA) corresponding to multiple liver disease-associated genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can be sequence complementary to one or more nucleic acid sequences from multiple liver disease-associated genomic loci or genomic regions. Multiple liver disease-associated genomic loci or genomic regions may contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, or at least 45. At least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000 or more different liver disease-related genomic loci or regions. Multiple liver disease-related genomic loci or regions may contain one or more members selected from the genes listed in Table 1.

[0182] The kit instructions may include instructions for measuring cell-free biological samples using probes selective for sequences at multiple liver disease-related genomic loci. These probes may be nucleic acid molecules (e.g., RNA or DNA) that are sequence-complementary to one or more nucleic acid sequences (e.g., RNA or DNA) from multiple liver disease-related genomic loci. These nucleic acid molecules may be primers or enriched sequences. The instructions for measuring cell-free biological samples may include instructions to process the cell-free biological samples by performing array hybridization, PCR, or nucleic acid sequencing to generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of sequences at each of the multiple liver disease-related genomic loci in the cell-free biological sample. Quantitative measurements (e.g., indicating presence, absence, or relative amount) of sequences at each of the multiple liver disease-related genomic loci in the cell-free biological sample may indicate one or more liver disease states.

[0183] The kit instructions may include instructions for measuring and interpreting assay readouts that can be quantified at one or more of multiple liver disease-related genomic loci to generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple liver disease-related genomic loci in a cell-free biological sample. For example, quantification by array hybridization or polymerase chain reaction (PCR) corresponding to multiple liver disease-related genomic loci can generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple liver disease-related genomic loci in a cell-free biological sample. Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values ​​thereof.

[0184] The kit may include a metabolomics assay for the quantitative measurement (e.g., indicating presence, absence, or relative amount) of each of several liver disease-related metabolites in a cell-free biological sample of a subject. The quantitative measurement (e.g., indicating presence, absence, or relative amount) of liver disease-related metabolites in a cell-free biological sample may indicate one or more liver disease states. Metabolites in a cell-free biological sample may be generated as a result of one or more metabolic pathways corresponding to liver disease-related genes (e.g., as a final product or byproduct). The kit may include instructions for isolating or extracting metabolites from a cell-free biological sample, and / or instructions for using a metabolomics assay to generate a dataset of quantitative measurements (e.g., indicating presence, absence, or relative amount) of each of several liver disease-related metabolites in a cell-free biological sample of a subject.

[0185] Machine learning models

[0186] After processing one or more cell-free biological samples derived from an object using one or more assays to generate one or more datasets indicative of liver disease or condition, a trained algorithm can be used to process the one or more datasets (e.g., at each of multiple liver disease-related genomic loci) to determine the liver disease state. For example, a trained algorithm can be used to determine sequence quantification measurements at each of multiple liver disease-related genomic loci in a cell-free biological sample. The trained algorithm can be configured to identify liver disease states with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more than 99% for at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, or more than about 500 independent samples.

[0187] Trained algorithms can include supervised machine learning algorithms. Trained algorithms can include Classification and Regression Tree (CART) algorithms. Supervised machine learning algorithms can include classifiers or regression algorithms. Supervised machine learning algorithms can include, for example, deep learning algorithms, support vector machines (SVM), neural networks, random forests, linear regression, or logistic regression. Trained algorithms can also include unsupervised machine learning algorithms.

[0188] The trained algorithm can be configured to receive multiple input variables and produce one or more output values ​​based on these input variables. The multiple input variables can contain one or more datasets indicating liver disease status. For example, the input variables can contain multiple sequences corresponding to or aligned to each of multiple liver disease-related genomic loci. The multiple input variables can also include the subject's clinical health data.

[0189] The trained algorithm may include a classifier such that each of one or more output values ​​contains one of a fixed number of possible values ​​(e.g., a linear classifier, logistic regression classifier, etc.), indicating the classifier's classification of cell-free biological samples. The trained algorithm may include a binary classifier such that each of one or more output values ​​contains one of two values ​​(e.g., {0, 1}, {positive, negative}, or {high risk, low risk}), indicating the classifier's classification of cell-free biological samples. The trained algorithm may be another type of classifier such that each of one or more output values ​​contains one of more than two values ​​(e.g., {0, 1, 2}, {positive, negative, or uncertain}, or {high risk, intermediate risk, or low risk}), indicating the classifier's classification of cell-free biological samples. Output values ​​may contain descriptive labels, numerical values, or combinations thereof. Some output values ​​may contain descriptive labels. Such descriptive labels can provide identification or indication of the liver disease or condition status of the object. Such descriptive labels may include, for example, positive, negative, high risk, intermediate risk, low risk, or uncertain. Such descriptive labels can provide identification of treatments for the liver disease state of the subject, and such treatments may include, for example, therapeutic interventions (e.g., vitamin E supplements, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise programs, diet programs, bariatric surgery, GLP1 (glucagon-like peptide-1) receptor agonists, FGF (fibroblast growth factor) analogs, THR (thyroid hormone receptor) agonists, SCD-1 (stearoyl-CoA desaturase 1) inhibitors, FAS (fatty acid synthase) inhibitors, FXR (farnesol X receptor) agonists, A... The descriptive label may include: CC (acetyl-CoA carboxylase) inhibitors, PPAR (peroxisome proliferator-activated receptor) agonists, targeted gene modifiers (including, for example, PNPLA3 or HSD17B13), LOXL2 (lysyl oxidase-like 2) inhibitors, pan-cyclic protein inhibitors, pan-cysteine ​​inhibitors, chemokine receptor (e.g., CCR2 / CCR5) inhibitors, galactolectin-3 inhibitors, mitochondrial uncoupling agents or uncoupling agents, structurally engineered fatty acids, or any combination thereof; the duration of the therapeutic intervention; and / or the dose of the therapeutic intervention appropriate for the liver disease condition. Such descriptive labels can provide identification of a possible secondary clinical test for the subject, and such secondary clinical test may include, for example, blood tests, liver biopsy, imaging tests, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound scans, chest X-rays, positron emission tomography (PET), PET-CT scans, cell-free biological cytology, or any combination thereof. For example, such descriptive labels can provide a prediction of the subject's liver disease condition. As another example, such descriptive labels can provide a relative assessment of an object’s liver disease status (e.g., presence, stage, or subtype).Some descriptive labels can be mapped to numerical values; for example, "positive" can be mapped to 1 and "negative" to 0.

[0190] Some output values ​​can contain numerical values, such as binary, integer, or continuous values. Such binary output values ​​can contain, for example, {0, 1}, {positive, negative}, or {high risk, low risk}. Such integer output values ​​can contain, for example, {0, 1, 2}. Such continuous output values ​​can contain, for example, probability values ​​that are at least 0 and no greater than 1. Such continuous output values ​​can contain, for example, non-normalized probability values ​​that are at least 0. Such continuous output values ​​can indicate the prognosis of an object's liver disease state. Some numerical values ​​can be mapped to descriptive labels, for example, mapping 1 to "positive" and 0 to "negative".

[0191] Some output values ​​can be assigned based on one or more cutoff values. For example, if the probability of a sample indicating that the subject has liver disease is at least 50%, the binary classification of the sample can be assigned an output value of "positive" or 1. Conversely, if the probability of a sample indicating that the subject has liver disease is less than 50%, the binary classification of the sample can be assigned an output value of "negative" or 0. In this case, a single cutoff value of 50% is used to classify the sample into one of two possible binary output values. Examples of single cutoff values ​​can include approximately 1%, approximately 2%, approximately 5%, approximately 10%, approximately 15%, approximately 20%, approximately 25%, approximately 30%, approximately 35%, approximately 40%, approximately 45%, approximately 50%, approximately 55%, approximately 60%, approximately 65%, approximately 70%, approximately 75%, approximately 80%, approximately 85%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, and approximately 99%.

[0192] As another example, if the probability that the sample indicates the subject has liver disease is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or higher, then the sample can be classified as “positive” or an output value of 1. If the probability that the sample indicates the subject has a liver disease state exceeds approximately 50%, approximately 55%, approximately 60%, approximately 65%, approximately 70%, approximately 75%, approximately 80%, approximately 85%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, or approximately 99%, the sample can be classified as "positive" or an output value of 1.

[0193] If the probability of the sample indicator having liver disease is less than approximately 50%, less than approximately 45%, less than approximately 40%, less than approximately 35%, less than approximately 30%, less than approximately 25%, less than approximately 20%, less than approximately 15%, less than approximately 10%, less than approximately 9%, less than approximately 8%, less than approximately 7%, less than approximately 6%, less than approximately 5%, less than approximately 4%, less than approximately 3%, less than approximately 2%, or less than approximately 1%, the sample can be classified as "negative" or an output value of 0. If the probability of the sample indicator having liver disease is not greater than approximately 50%, not greater than approximately 45%, not greater than approximately 40%, not greater than approximately 35%, not greater than approximately 30%, not greater than approximately 25%, not greater than approximately 20%, not greater than approximately 15%, not greater than approximately 10%, not greater than approximately 9%, not greater than approximately 8%, not greater than approximately 7%, not greater than approximately 6%, not greater than approximately 5%, not greater than approximately 4%, not greater than approximately 3%, not greater than approximately 2%, or not greater than approximately 1%, the sample can be classified as "negative" or an output value of 0.

[0194] If a sample is not classified as “positive,” “negative,” 1, or 0, its classification can be assigned an output value of “uncertain” or 2. In this case, a set of two cutoff values ​​is used to classify the sample into one of three possible output values. Examples of sets of cutoff values ​​can include {1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, {15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}. Similarly, a set of n cutoff values ​​can be used to classify a sample into one of n+1 possible output values, where n is any positive integer.

[0195] The trained algorithm can be trained using multiple independent samples. Each independent sample may contain a cell-free biological sample from the object, a relevant dataset obtained by measuring that cell-free biological sample (as described herein), and one or more known output values ​​corresponding to that cell-free biological sample (e.g., clinical diagnosis, prognosis, absence, or therapeutic efficacy of the object's liver disease state). Independent samples may contain cell-free biological samples obtained or derived from multiple different objects, along with relevant datasets and outputs. Independent samples may contain cell-free biological samples obtained from the same object at multiple different time points (e.g., regularly, such as weekly, bi-weekly, or monthly), along with relevant datasets and outputs. Independent samples may be associated with the presence of a liver disease state (e.g., training samples may contain cell-free biological samples obtained or derived from multiple objects known to have a liver disease state, along with relevant datasets and outputs). Independent samples may be associated with the absence of a liver disease state (e.g., training samples may contain cell-free biological samples obtained or derived from multiple objects known not to have a liver disease state, or who have received negative test results for a liver disease state, along with relevant datasets and outputs).

[0196] The trained algorithm can be trained using at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent samples. Independent samples may include cell-free biological samples associated with the presence of a liver disease state and / or cell-free biological samples associated with the absence of a liver disease state. The trained algorithm can be trained using no more than about 500, no more than about 450, no more than about 400, no more than about 350, no more than about 300, no more than about 250, no more than about 200, no more than about 150, no more than about 100, or no more than about 50 independent samples associated with the presence of a liver disease. In some implementations, cell-free biological samples are independent of the samples used to train the trained algorithm.

[0197] The trained algorithm can be trained using a first number of independent samples associated with the presence of liver disease and a second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease may not be greater than the second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease may be equal to the second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease may be greater than the second number of independent samples associated with the absence of liver disease.

[0198] The trained algorithm can be configured to target at least approximately 5, at least approximately 10, at least approximately 15, at least approximately 20, at least approximately 25, at least approximately 30, at least approximately 35, at least approximately 40, at least approximately 45, at least approximately 50, at least approximately 100, at least approximately 150, at least approximately 200, at least approximately 250, at least approximately 300, at least approximately 350, at least approximately 400, at least approximately 450, or at least approximately 500 independent samples with at least approximately 50%, at least approximately 55%, at least approximately 6 The accuracy of the trained algorithm in identifying liver disease states can be calculated as the percentage of independent samples (e.g., subjects known to have a liver disease state or subjects with negative clinical test results for a liver disease state) that are correctly identified or classified as having or not having a liver disease state.

[0199] The trained algorithm can be configured to identify liver disease states with a positive predictive value (PPV) of at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher. The PPV for identifying liver disease states using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as having liver disease states that correspond to objects that actually have liver disease states.

[0200] The trained algorithm can be configured to identify liver disease states with at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher negative predictive values ​​(NPV). The NPV for identifying liver disease states using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as not having liver disease states, corresponding to objects that truly do not have liver disease states.

[0201] The trained algorithm can be configured to have at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, and at least approximately 9 Clinical sensitivity of 0%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or higher, for identifying liver disease states. The clinical sensitivity for identifying liver disease states using a trained algorithm can be calculated as the percentage of independent samples (e.g., subjects known to have liver disease states) that are correctly identified or classified as having a liver disease state in relation to its presence.

[0202] The trained algorithm can be configured to have at least approximately 5%, at least approximately 10%, at least approximately 15%, at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, and at least approximately 9 The clinical specificity for identifying liver disease states is 0%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or higher. The clinical specificity for identifying liver disease states using a trained algorithm can be calculated as the percentage of independent samples (e.g., subjects with negative clinical test results for liver disease states) correctly identified or classified as not having a liver disease state, in relation to the absence of such a state.

[0203] The trained algorithm can be configured to identify liver disease states with an area under the curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, at least about 0.99 or higher. AUC can be calculated as the integral of the receiver operating characteristic (ROC) curve, such as the area under the ROC curve (AUROC) associated with a trained algorithm that classifies cell-free biological samples as having or not having a liver disease state.

[0204] The trained algorithm can be tuned or adjusted to improve one or more of the following: performance, accuracy, PPV, NPV, clinical sensitivity, clinical specificity, or AUC in identifying liver disease states. The trained algorithm can be tuned or adjusted by modifying its parameters (e.g., the set of cutoff values ​​for classifying cell-free biological samples as described elsewhere in this document, or the weights of the neural network). The trained algorithm can be continuously tuned or adjusted during or after training.

[0205] After initial training of the trained algorithm, a subset of the input can be identified as the most influential or most important subset for high-quality classification. For example, a subset of multiple liver disease-related genomic loci can be identified as the most influential or most important subset for high-quality classification or identification of liver diseases (or subtypes). Multiple liver disease-related genomic loci or subsets thereof can be ranked based on classification metrics that indicate the influence or importance of each genomic locus for high-quality classification or identification of liver diseases (or subtypes). Such metrics can be used to reduce (in some cases significantly reduce) the number of input variables (e.g., predictor variables) required to train the trained algorithm to the desired performance level (e.g., based on the required minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, AUC, positive likelihood ratio, negative likelihood ratio, or combinations thereof). For example, if training a trained algorithm with multiple input variables containing dozens or hundreds of input variables results in a classification accuracy exceeding 99%, then alternatively training the trained algorithm with only a selected subset of the most influential or important input variables—no more than approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 100—can produce a lower but still acceptable classification accuracy. Accuracy (e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%). A subset can be selected by sorting all multiple input variables and choosing a predetermined number of input variables with the best classification index (e.g., no more than about 5, no more than about 10, no more than about 15, no more than about 20, no more than about 25, no more than about 30, no more than about 35, no more than about 40, no more than about 45, no more than about 50, or no more than about 100).

[0206] The accuracy of a trained algorithm can be context-dependent. In some cases, accuracy can be based on training samples from a general population. In other cases, accuracy can be based on training samples from a high-risk group (e.g., a group suspected of having liver disease). Several factors can be considered when interpreting the test performance of a trained algorithm, including: 1) the prevalence of the disease or condition, such as how many people in the target population have the disease; and 2) whether the test is used to diagnose the disease (i.e., a positive (confirmatory) test) or to confirm that the subject does not have the disease (i.e., a negative (exclusion) test).

[0207] On the other hand, metrics such as pre-test / post-test probabilities, Bayes factor, likelihood ratio, or information gain can be context-independent. These metrics measure the amount of new information provided by the test. For example, the pre-test / post-test probability ratio can be calculated by dividing the probability that "a subject in the target population has the condition" by the probability that "a subject in the target population with a given test result has the condition." As an example, approximately 5% of the US population has NASH; therefore, the pre-test probability of NASH in the US population is 5%. If the test detects that 50% of the subjects do indeed have NASH, the post-test probability is 50%, and the pre-test / post-test ratio is 10. As another example, if approximately 40% of the subjects in a high-risk group have NASH, and a hypothesis test is performed on this high-risk group, the test detects that 50% of the people do indeed have NASH, and the pre-test / post-test ratio is 1.25.

[0208] Identifying or monitoring liver disease status

[0209] After processing a dataset using a trained algorithm, the liver disease status of an individual can be identified or monitored. This identification can be based, at least in part, on quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites.

[0210] The liver disease status of subjects can be identified with an accuracy of at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, or higher. The accuracy of identifying liver disease status using a trained algorithm can be calculated as the percentage of independent samples (e.g., subjects known to have a liver disease status or subjects with negative clinical test results indicating a liver disease status) correctly identified or classified as having or not having a liver disease status.

[0211] The liver disease status of a subject can be identified with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or higher. The PPV for identifying liver disease states using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as having liver disease states that correspond to objects that actually have liver disease states.

[0212] The liver disease status of a subject can be identified with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or higher. The NPV for identifying liver disease states using a trained algorithm can be calculated as the percentage of cell-free biological samples identified or classified as not having liver disease states, corresponding to objects that truly do not have liver disease states.

[0213] It can be at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 90%. Clinical sensitivity of 1%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or higher, is used to identify the liver disease status of subjects. The clinical sensitivity for identifying liver disease status using a trained algorithm can be calculated as the percentage of independent samples (e.g., subjects known to have a liver disease status) that are correctly identified or classified as having a liver disease status in relation to the presence of that status.

[0214] It can be at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 90%. The clinical specificity for identifying the liver disease status of subjects is 1%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or higher. The clinical specificity for identifying the liver disease status using a trained algorithm can be calculated as the percentage of independent samples (e.g., subjects with negative clinical test results for liver disease status) correctly identified or classified as not having a liver disease status, in relation to the absence of such a status.

[0215] Likelihood ratios can be used to evaluate the performance of diagnostic tests. Liver disease states in subjects can be identified or excluded based on likelihood ratios (e.g., positive or negative likelihood ratios). Likelihood ratios are independent of disease prevalence in the training population and are therefore more representative of disease prevalence in the target population. Because likelihood ratios are independent of disease prevalence, they can be more directly correlated with the performance of a given diagnostic test.

[0216] The positive likelihood ratio can be calculated as sensitivity / (1-specificity). It can be expressed as at least about 1, at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16. Positive likelihood ratios of at least approximately 17, at least approximately 18, at least approximately 19, at least approximately 20, at least approximately 30, at least approximately 40, at least approximately 50, at least approximately 60, at least approximately 70, at least approximately 80, at least approximately 90, at least approximately 100, at least approximately 200, at least approximately 300, at least approximately 400, at least approximately 500, at least approximately 600, at least approximately 700, at least approximately 800, at least approximately 900, or at least approximately 1000 were used to identify the liver disease status of the subjects.

[0217] The negative likelihood ratio can be calculated as (1 - sensitivity) / specificity. It can be expressed as approximately 1, approximately 0.99, approximately 0.95, approximately 0.9, approximately 0.8, approximately 0.7, approximately 0.75, approximately 0.6, approximately 0.5, approximately 0.4, approximately 0.3, approximately 0.25, approximately 0.2, approximately 0.1, approximately 0.09, approximately 0.08, approximately 0.07, approximately 0.06, and so on. The negative likelihood ratios of approximately 0.05, 0.04, 0.03, 0.02, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003, 0.002, or 0.001 exclude the liver disease status of the subjects.

[0218] In one aspect, this disclosure provides a method for determining the risk of developing liver disease in a subject, comprising measuring a cell-free biological sample derived from the subject to generate a dataset indicating the risk of developing liver disease with at least 80% specificity, and using a trained algorithm trained on samples independent of the cell-free biological sample to determine the risk of developing liver disease in the subject with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or higher.

[0219] After a subject's liver disease is identified, a subtype of the liver disease can be further identified (e.g., selected from multiple subtypes of liver disease). The subtype of liver disease can be determined at least in part based on: quantitative measurements of sequence reads from a dataset of sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from a dataset of sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites. For example, a subject can be identified as being at risk for a subtype of liver disease (e.g., selected from multiple subtypes of liver disease). After a subject is identified as being at risk for a subtype of liver disease, a clinical intervention can be selected for that subject at least in part based on the subtype of liver disease to which the subject is identified as being at risk. In some embodiments, the clinical intervention is selected from a variety of clinical interventions (e.g., clinically indicated for different subtypes of liver disease).

[0220] In some implementations, the trained algorithm can determine the risk of liver disease in a subject to be at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or higher.

[0221] The trained algorithm can achieve a success rate of at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 81%, at least approximately 82%, at least approximately 83%, at least approximately 84%, at least approximately 85%, at least approximately 86%, at least approximately 87%, at least approximately 88%, at least approximately 89%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, at least approximately 99%, at least approximately 99.1%, at least approximately 99.2%, at least approximately 99.3%, at least approximately 99.4%, and at least approximately 99.5%. The accuracy rate is at least approximately 99.6%, at least approximately 99.7%, at least approximately 99.8%, at least approximately 99.9%, at least approximately 99.99%, at least approximately 99.999%, or higher to identify subjects at risk of liver disease.

[0222] Once a subject is identified as having a liver disease state, therapeutic interventions may be optionally provided (e.g., prescribing an appropriate treatment procedure to treat the subject's liver disease state). Therapeutic interventions may include prescribing an effective dose of medication, further testing or evaluation of the liver disease state, further monitoring of the liver disease state, exercise programs, dietary programs, weight-loss surgery, or a combination thereof. Therapeutic interventions may include vitamin E supplements, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise programs, diet programs, bariatric surgery, GLP1 (glucagon-like peptide-1) receptor agonists, FGF (fibroblast growth factor) analogs, THR (thyroid hormone receptor) agonists, SCD-1 (stearoyl-CoA desaturase 1) inhibitors, FAS (fatty acid synthase) inhibitors, FXR (farnesyl X receptor) agonists, ACC (acetyl-CoA carboxylase) inhibitors, PPAR (peroxisome proliferator-activated receptor) agonists, targeted gene modifiers (including, for example, PNPLA3 or HSD17B13), LOXL2 (lysyl oxidase-like 2) inhibitors, pan-cyclic protein inhibitors, pan-cysteine ​​inhibitors, chemokine receptors (e.g., CCR2 / CCR5) inhibitors, galactolectin-3 inhibitors, mitochondrial uncouplers or uncoupling agents, structurally engineered fatty acids, or combinations thereof. If the subject is currently receiving treatment for liver disease in a therapeutic process, the therapeutic intervention may include a different subsequent treatment process (e.g., to improve the efficacy of the current treatment process since it is ineffective).

[0223] Therapeutic interventions may include recommending secondary clinical testing to confirm the diagnosis of liver disease. This secondary clinical testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0224] After identifying a candidate as having a liver disease condition, the patient may optionally be determined to be ineligible for liver transplantation. Conversely, after identifying a candidate as not having liver disease, the patient may optionally be determined to be eligible for liver transplantation. If a candidate is not identified as having liver disease or is at an increased risk of developing liver disease, the patient may be determined to be eligible as a liver transplant donor. If a candidate is identified as having liver disease or is at an increased risk of developing liver disease, the patient may be determined to be eligible as a liver transplant recipient.

[0225] Various therapeutic interventions and clinical tests for liver disease can be used in conjunction with the methods described herein. For example, a therapeutic intervention can be administered to a subject after the presence of liver disease has been determined. As another example, a preventative intervention can be administered to a subject after the presence of an elevated risk of liver disease has been determined. Examples of liver disease interventions and clinical tests can be found in Vittal et al., Clin Liver Dis. 2019 Aug; 23(3):417–432; Marroni et al., World J Gastroenterol. 2018 Jul 14; 24(26):2785–2805; Leoni et al., World J Gastroenterol. 2018 Aug 14; 24(30):3361–3373; and Sumida et al., J Gastroenterol. 2018 Mar; 53(3):362–376, each of which is incorporated herein by reference in its entirety.

[0226] Quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of datasets containing sets of liver disease-related metabolites can be evaluated over a period of time to monitor patients (e.g., subjects with liver disease or undergoing treatment for liver disease). In this context, the quantitative measurements of a patient's dataset can change during the treatment process. For example, the quantitative measurements of a dataset of patients whose risk of liver disease is reduced due to effective treatment may shift towards a profile or distribution of healthy subjects (e.g., subjects without liver disease or condition). Conversely, the quantitative measurements of a dataset of patients whose risk of liver disease is increased due to ineffective treatment may shift towards a profile or distribution of subjects with a higher risk of liver disease or more severe liver disease.

[0227] Liver disease in a subject can be monitored by monitoring the treatment process for that subject. Monitoring may include assessing the subject's liver disease status at two or more time points. The assessment may be based, at least in part, on quantitative measurements of sequence reads from datasets of identified sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts at liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets of datasets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of datasets of datasets of liver disease-related metabolites at each of the two or more time points.

[0228] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate one or more clinical indicators, such as (i) a diagnosis of liver disease in the subject; (ii) a prognosis of liver disease in the subject; (iii) an increased risk of liver disease in the subject; (iv) a reduced risk of liver disease in the subject; (v) the efficacy of a treatment regimen for treating liver disease in the subject; and (vi) the inefficiency of a treatment regimen for treating liver disease in the subject.

[0229] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate a diagnosis of liver disease in a subject. For example, if liver disease was not detected in a subject at an earlier time point but was detected at a later time point, this difference indicates a diagnosis of liver disease in the subject. Clinical actions or decisions can be made based on this indication of a diagnosis of liver disease in a subject, such as prescribing a new therapeutic intervention for the subject. Clinical actions or decisions may include recommending that the subject undergo secondary clinical testing to confirm the diagnosis of the liver disease state. Such secondary clinical testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0230] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate the prognosis of a subject's liver disease status.

[0231] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate an increased risk of a subject having a liver disease state. For example, if a subject's liver disease state is detected at both earlier and later time points, and if the difference is positive (e.g., an increase in quantitative measurements of sequence reads or RNA transcripts from datasets containing sets of liver disease-related genomic loci, proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites from an earlier time point to a later time point), then that difference can indicate an increased risk of a subject having a liver disease state. Clinical actions or decisions can be made based on this indication of an increased risk to the liver disease state, such as prescribing a new therapeutic intervention or changing a therapeutic intervention (e.g., ending the current treatment and prescribing a new one). Clinical actions or decisions may include recommending that the subject undergo secondary clinical testing to confirm the increased risk to the liver disease state. This secondary clinical testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0232] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate a reduced risk of a liver disease state in the subject. For example, if liver disease is detected in the subject at both earlier and later time points, and if the difference is negative (e.g., a decrease in quantitative measurements of sequence reads or RNA transcripts from datasets containing sets of liver disease-related genomic loci, proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites from an earlier time point to a later time point), this difference can indicate a reduced risk of a liver disease state in the subject. Clinical actions or decisions (e.g., continuing or ending current therapeutic interventions) can be made for the subject based on this indication of a reduced risk of a liver disease state. Clinical actions or decisions may include recommending that the subject undergo secondary clinical testing to confirm a reduced risk of liver disease status. This secondary clinical testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0233] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate the efficacy of a treatment process for a liver disease state in a subject. For example, if liver disease was detected in a subject at an earlier time point but not at a later time point, this difference can indicate the efficacy of a treatment process for the subject's liver disease. Clinical actions or decisions can be made based on this indication of the efficacy of the treatment process for the subject's liver disease, such as continuing or discontinuing the subject's current therapeutic intervention. Clinical actions or decisions may include recommending that the subject undergo secondary clinical testing to confirm the efficacy of the treatment process for the liver disease state. The secondary clinical test may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0234] In some implementations, differences in quantitative measurements of sequence reads from datasets containing sets of liver disease-related genomic loci (e.g., quantitative measurements of DNA or RNA transcripts from liver disease-related genomic loci), proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, determined between two or more time points, can indicate the ineffectiveness of a treatment process for the liver disease state of the subject. For example, if the liver disease state of the subject is detected at both earlier and later time points, and if the difference is positive or zero (e.g., quantitative measurements of sequence reads or RNA transcripts from datasets containing sets of liver disease-related genomic loci, proteomics data containing quantitative measurements of proteins from datasets containing sets of liver disease-related proteins, and / or metabolomics data containing quantitative measurements of sets of liver disease-related metabolites, increasing or remaining constant from the earlier time point to the later time point), and if an effective treatment was indicated at the earlier time point, then the difference can indicate the ineffectiveness of a treatment process for the liver disease of the subject. Clinical actions or decisions can be made based on the indication of ineffectiveness of the treatment process for the patient's liver disease. For example, this could involve ending the current therapeutic intervention and / or switching to (e.g., prescribing) a different new therapeutic intervention. Clinical actions or decisions may include recommending that the patient undergo secondary clinical testing to confirm the ineffectiveness of the treatment process for the liver disease. This secondary clinical testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.

[0235] In another aspect, this disclosure provides a computer-implemented method for predicting the risk of liver disease in a subject, comprising: (a) receiving clinical health data of the subject, wherein the clinical health data includes multiple quantitative or categorical measurements of the subject; (b) processing the clinical health data of the subject using a trained algorithm to determine a risk score indicating the risk of liver disease in the subject; and (c) electronically outputting a report of the risk score indicating the risk of liver disease in the subject.

[0236] For example, in some implementations, clinical health data includes one or more quantitative measurements of the subject, such as age, weight, height, body mass index (BMI), blood pressure, heart rate, and blood glucose levels. As another example, clinical health data may include one or more categorical indicators, such as race, ethnicity, medical history, medication history or other clinical treatment history, tobacco use history, alcohol consumption history, daily activity or fitness level, genetic testing results, blood test results, and imaging results.

[0237] In some implementations, the computer-implemented method for predicting the risk of liver disease in a subject is performed using a computer or mobile device application. For example, the subject can input their own clinical health data, including quantitative and / or categorical indicators, using the computer or mobile device application. The computer or mobile device application can then process the clinical health data using a trained algorithm to determine a risk score indicating the subject's risk of liver disease. The computer or mobile device application can then display a report of the risk score indicating the subject's risk of liver disease.

[0238] In some implementations, a risk score indicating the risk of liver disease in a subject can be refined by performing one or more follow-up clinical tests. For example, a physician may recommend one or more follow-up clinical tests (e.g., imaging tests or blood tests) based on an initial risk score. A computer or mobile device application can then use a trained algorithm to process the results from one or more follow-up clinical tests to determine an updated risk score indicating the risk of liver disease in the subject.

[0239] In some implementations, the risk score includes the likelihood that the subject has liver disease over a predetermined period of time. For example, the scheduled duration can be approximately 1 hour, approximately 2 hours, approximately 4 hours, approximately 6 hours, approximately 8 hours, approximately 10 hours, approximately 12 hours, approximately 14 hours, approximately 16 hours, approximately 18 hours, approximately 20 hours, approximately 22 hours, approximately 24 hours, approximately 1.5 days, approximately 2 days, approximately 2.5 days, approximately 3 days, approximately 3.5 days, approximately 4 days, approximately 4.5 days, approximately 5 days, approximately 5.5 days, approximately 6 days, approximately 6.5 days, approximately 7 days, approximately 8 days, approximately 9 days, approximately 10 days, approximately 12 days, approximately 14 days, approximately 3 weeks, approximately 4 weeks, approximately 5 weeks, approximately 6 weeks, approximately 7 weeks, approximately 8 weeks, approximately 9 weeks, approximately 10 weeks, approximately 11 weeks, approximately 12 weeks, approximately 5 months, approximately 6 months, approximately 7 months, approximately 8 months, approximately 9 months, approximately 10 months, approximately 11 months, approximately 1 year, approximately 2 years, approximately 3 years, approximately 4 years, approximately 5 years, or more than approximately 5 years.

[0240] After identifying a subject's liver disease status or detecting an increased risk of liver disease, a report indicating the subject's liver disease (e.g., identifying or providing an indication of liver disease) can be generated electronically. The subject may not exhibit liver disease (e.g., be asymptomatic for liver disease). The report may be presented on a user's electronic device's graphical user interface (GUI). The user can be the subject, caregiver, physician, nurse, or other healthcare professional.

[0241] The report may include one or more clinical indications, such as (i) a diagnosis of the subject's liver disease; (ii) the prognosis of the subject's liver disease; (iii) an increased risk of the subject's liver disease; (iv) a reduced risk of the subject's liver disease; (v) the efficacy of the treatment regimen used to treat the subject's liver disease; and (vi) the ineffectiveness of the treatment regimen used to treat the subject's liver disease. The report may include one or more clinical actions or decisions made based on these clinical indications. Such clinical actions or decisions may involve therapeutic interventions, induction or suppression of labor, or further clinical evaluation or testing of the subject's liver disease.

[0242] For example, a clinical indication of a diagnosis of liver disease in a subject may be accompanied by clinical action of prescribing a new therapeutic intervention. As another example, a clinical indication of an increased risk of liver disease in a subject may be accompanied by clinical action of prescribing a new therapeutic intervention or changing the therapeutic intervention (e.g., ending the current treatment and prescribing a new one). As another example, a clinical indication of a decreased risk of liver disease in a subject may be accompanied by clinical action of continuing or ending the current therapeutic intervention. As another example, a clinical indication of the efficacy of a treatment process for a subject's liver disease may be accompanied by clinical action of continuing or ending the current therapeutic intervention. As yet another example, a clinical indication of the ineffectiveness of a treatment process for a subject's liver disease may be accompanied by clinical action of ending the current therapeutic intervention and / or changing (e.g., prescribing) a different new therapeutic intervention.

[0243] Computer System

[0244] This disclosure provides a computer system that is programmed to implement the methods of this disclosure. Figure 2 A computer system 201 is shown, which is programmed or otherwise configured for, for example: (i) training and testing trained algorithms; (ii) processing data using trained algorithms to determine the liver disease status of a subject; (iii) determining a quantitative measurement indicating the liver disease status of a subject; (iv) identifying or monitoring the liver disease status of a subject; and (v) electronically outputting a report indicating the liver disease status of a subject.

[0245] Computer system 201 can control various aspects of the analysis, calculation, and generation of this disclosure, such as: (i) training and testing trained algorithms; (ii) processing data using trained algorithms to determine the liver disease status of a subject; (iii) determining quantitative measurements indicative of the liver disease status of a subject; (iv) identifying or monitoring the liver disease status of a subject; and (v) electronically outputting a report indicative of the liver disease status of a subject. Computer system 201 can be a user's electronic device or a computer system remotely located relative to an electronic device. The electronic device can be a mobile electronic device.

[0246] Computer system 201 includes a central processing unit (CPU, also referred to herein as a “processor” or “computer processor”) 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 201 also includes memory or storage units 210 (e.g., random access memory, read-only memory, flash memory), electronic storage units 215 (e.g., hard disk), a communication interface 220 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 225, such as cache, other memory, data storage, and / or electronic display adapters. Memory 210, storage units 215, interface 220, and peripheral devices 225 communicate with CPU 205 via a communication bus (solid line) (such as a motherboard). Storage unit 215 may be a data storage unit (or data repository) for storing data. Computer system 201 may be operatively coupled to computer network (“network”) 230 via communication interface 220. Network 230 may be the Internet, the Internet of Things, and / or an extranet, or an intranet and / or extranet communicating with the Internet.

[0247] In some cases, network 230 is a telecommunications and / or data network. Network 230 may contain one or more computer servers that can support distributed computing, such as cloud computing. For example, one or more computer servers may enable cloud computing through network 230 (“the cloud”) to perform various aspects of the analysis, computation, and generation disclosed herein, such as: (i) training and testing trained algorithms; (ii) processing data using trained algorithms to determine the liver disease status of an object; (iii) determining a quantitative measurement indicative of the liver disease status of an object; (iv) identifying or monitoring the liver disease status of an object; and (v) electronically outputting a report indicative of the liver disease status of an object. Such cloud computing may be provided by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. In some cases, network 230 may enable a peer-to-peer network via computer system 201, which may allow devices coupled to computer system 201 to act as clients or servers.

[0248] CPU 205 may include one or more computer processors and / or one or more graphics processing units (GPUs). CPU 205 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as memory 210. The instructions may be directed to CPU 205, which may subsequently be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by CPU 205 may include fetching, decoding, executing, and writing back.

[0249] CPU 205 may be part of a circuit, such as an integrated circuit. One or more other components of system 201 may be contained in this circuit. In some cases, this circuit is an application-specific integrated circuit (ASIC).

[0250] Storage unit 215 may store files, such as drivers, libraries, and saved programs. Storage unit 215 may store user data, such as user preferences and user programs. In some cases, computer system 201 may include one or more additional data storage units located outside computer system 201, such as on a remote server communicating with computer system 201 via an intranet or the Internet.

[0251] Computer system 201 can communicate with one or more remote computer systems via network 230. For example, computer system 201 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets, or tablet PCs (e.g., tablet PCs). iPad Galaxy Tab), telephone, smartphone (e.g.) iPhone, Android-compatible devices (or personal digital assistant). Users can access computer system 201 via network 230.

[0252] The methods described herein can be implemented using machine-executable code (e.g., a computer processor) stored in an electronic storage location (e.g., memory 210 or electronic storage unit 215) of computer system 201. The machine-executable code or machine-readable code can be provided in software form. During use, the code can be executed by processor 205. In some cases, the code can be retrieved from storage unit 215 and stored in memory 210 for easy access by processor 205. In some cases, electronic storage unit 215 can be excluded, and machine-executable instructions can be stored in memory 210.

[0253] The code can be pre-compiled and configured for use on machines with processors suitable for executing it, or it can be compiled at runtime. The code can be provided in the form of a programming language, which can be selected to enable the code to be executed in a pre-compiled or just-in-time compiled manner.

[0254] Various aspects of the systems and methods provided herein, such as computer system 201, can be embodied in a programmatic form. These aspects of the technology can be considered "products" or "manufactured goods," typically carried or contained in some type of machine-readable medium in the form of machine (or processor) executable code and / or associated data. Machine-executable code can be stored in electronic storage units, such as memory (e.g., read-only memory, random access memory, flash memory) or hard disks. "Storage" type media can include any or all tangible memory of computers, processors, etc., or related modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software content can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable the loading of software from one computer or processor to another, such as from a management server or host to a computer platform for an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used through physical interfaces between local devices, wired and optical route networks, and various air links. Physical elements that carry such waves (such as wired or wireless links, optical links, etc.) can also be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as "computer or machine-readable medium" refer to any medium involved in providing instructions to a processor for execution.

[0255] Therefore, machine-readable media (such as computer-executable code) can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, any storage device such as any computer or similar device, such as a database that can be used to implement the figures shown. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, copper wires, and optical fibers, which include wires that form the bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as sound or light waves generated during radio frequency (RF) and infrared (IR) data communication. Therefore, common forms of computer-readable media include: floppy disks, floppy hard disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cardstock, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or cards, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may involve carrying one or more sequences of one or more instructions to a processor for execution.

[0256] Computer system 201 may include or communicate with an electronic display 235, which includes a user interface (UI) 240 for providing, for example, (i) a visual display indicating the training and testing of a trained algorithm; (ii) a visual display indicating data indicating the liver disease status of an object; (iii) a quantitative measurement of the liver disease status of an object; (iv) an identification of the object having a liver disease status; or (v) an electronic report indicating the liver disease status of an object. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0257] The methods and systems disclosed herein can be implemented by one or more algorithms. The algorithms can be implemented in software when executed by the central processing unit 205. The algorithms can, for example, (i) train and test the trained algorithm; (ii) process data using the trained algorithm to determine the liver disease status of a subject; (iii) determine a quantitative measurement indicating the liver disease status of the subject; (iv) identify or monitor the liver disease status of the subject; and (v) electronically output a report indicating the liver disease status of the subject.

[0258] cfDNA methylation

[0259] In some implementations, cfDNA methylation data (observations) obtained from biological samples comprises a collection of sequenced DNA fragments that have undergone transformation conditions, converting unmethylated cytosine sites to thymine, to provide the methylation status of cytosine sites in the DNA fragments. Each DNA fragment may consist of multiple base pair reads, some of which indicate whether methylated sites are methylated or unmethylated. This paper provides machine learning models and systems for inferring relevant outcomes from such cfDNA methylation data. Non-limiting examples of such outcomes include: (i) the presence or absence of a disease; (ii) the type or subtype of the disease; (iii) the type, dosage, or combination of treatments used to treat the disease; (iv) the predicted response of a subject to treatment for the disease; (v) the risk of a subject developing a late form of the disease; and (vi) the outcome (prognosis) of the subject.

[0260] A dataset can contain cfDNA methylation data from one or more objects, at least some of which have one or more tags as described herein. A challenge in training ML models is generating models from datasets that can infer outcomes from new, previously untrained cfDNA methylation data. In cfDNA methylation data, each fragment can be assigned a location in the genome. ML models can represent the data through data representation, characterization, or feature engineering. For large datasets (e.g., with millions of data points), deep neural networks can be used to generate data representations in a purely data-driven manner. Such networks can be designed to build complex underlying datasets without requiring strong assumptions derived purely from the data. However, in cases with small to medium sample sizes, such as cfDNA methylation data, inferring outcomes using purely data-driven representations (without any assumptions) can be challenging. This paper presents a method for representing data with high dimensionality and small sample sizes in ML models for inferring outcomes with high accuracy and sensitivity. The method described herein includes providing a compact probability distribution of multiple fragments; using the compact probability distribution during intermediate training to provide a training model; and using the training model to characterize the multiple fragments.

[0261] cfDNA methylation data consists of numerous fragments; however, these fragments can originate from any location in the genome, and samples can have varying numbers of fragments. This non-uniform sparse data can also pose challenges to ML training methods. Furthermore, the human genome contains approximately 28 million methylation sites, several orders of magnitude larger than the largest feasible clinical study using cfDNA methylation data. Training with input data that is several orders of magnitude larger than the number of training data points can be challenging for ML training methods.

[0262] DNA data (including cfDNA methylation data) can be generated by sequencers, which can be an expensive and time-consuming process. ML training methods can circumvent these drawbacks by utilizing data from diverse studies, regardless of the data acquisition method and data source. Such data sources can include, but are not limited to:

[0263] —cfDNA and noncfDNA, for example, combining data obtained from cfDNA with data obtained from tissue samples;

[0264] —Different methylation assays, for example, combining data obtained from bisulfite conversion with data obtained from enzyme conversion assays; and

[0265] —Different sequencing methods, for example, combining data obtained from microarrays with data obtained from next-generation sequencing.

[0266] Such flexibility allows the use of pre-acquired data, such as publicly available data.

[0267] With approximately 3.2 billion genomic locations, of which about 28 million are potentially methylated, cfDNA methylation data can be enormous. Each fragment in cfDNA can have an average of about 150 base pairs. Therefore, for example, for a given sample, a cfDNA methylation dataset at 30x sequencing depth might require at least 48GB and 250MB of storage space, respectively, to store base pairs and methylation states. A training process involving 500 samples might require multiple rounds. Therefore, a machine learning training method capable of handling such large datasets is needed.

[0268] Distributing training across computer clusters can help overcome these challenges. However, this approach can have several drawbacks. Training can be extremely slow and time-consuming because multiple computers need to communicate with each other during training. For example, training a model with all segments based on 1,000 observations might require approximately 1,500 core hours and thousands of computers. Alternatively, data can be partitioned into different regions of the genome and processed independently. However, this approach may hinder ML models from learning subtle interactions between different genomic regions.

[0269] This paper provides a machine learning (ML) approach that mitigates the challenges described herein. The method disclosed herein includes: providing a probability distribution based on cfDNA methylation data from a set of fragments from a biological sample; and training this probability distribution on an ML model. The method is not trained on the set of fragments, but rather on a probability distribution of the set of fragments. This probability distribution can represent the state of the sample; the list of observed fragments can be derived from a probability distribution mediated by blood sampling and sequencing of the set of DNA fragments. Specifically, the method disclosed herein includes transforming the set of input fragments into a probability distribution most likely to generate that input fragment.

[0270] Representing data using probability distributions offers several advantages. Probability distributions can be non-sparse and have a predefined, fixed complexity. They can represent the likelihood of observing different methylation patterns. Representing the state of methylation patterns, they are therefore less susceptible to variation due to assays, sequencing methods, and other factors. Such characteristics are desirable given the availability of sequencing data in the public domain (e.g., from the National Institutes of Health and other research institutions). Furthermore, probability distributions are much smaller in size, making them easier to use in training or distributed systems. This, in turn, makes building complex models more feasible. Building complex models can be very expensive and time-consuming if computation is high. Therefore, probability distributions provide a simpler way to make training complex models feasible. Additionally, the probability distribution representing a given sample can be computed without needing or knowing other samples (e.g., training other samples). Therefore, this process can be easily distributed across computer clusters. The procedure does not leak information between samples and can be freely performed without cross-validation or training and testing datasets. Such representations are also suitable for building models capable of producing high-quality inference.

[0271] Cell-free DNA methylation data can be derived from a large number of cells throughout the body. Assuming each cell possesses multiple features (Z), a cell can be represented as a mixture of these features. A sample can be represented as the proportion of different cells, and therefore, as the proportion of such hidden features. Therefore, the primary task is to determine the optimal Z features from a dataset of probabilistic distributions that can be estimated for a set of fragments from the cfDNA methylation dataset.

[0272] Figure 3 The illustration shows a schematic of the example training dataset. Mathematically, the underlying dataset is a D-dimensional random variable (i.e., the number of methylation sites), which is partially observed. Each observation (i.e., a participant) consists of multiple segments. Each segment corresponds to a set of values ​​that correspond to a portion of the data in D-dimensional space.

[0273] As described in this article, observations can be represented by a distribution in D-dimensional space, which is expressed by φ. s (Each observation) is a feature, not a set of fragments. The distribution parameter φ s It is a statistical data set of fragments. For a large class of distributions, such as the exponential family, the distribution parameter (φ) s The distribution φ can be explicitly expressed as its sufficient statistical data. For others, in general, the parameters of the distribution can be expressed as approximately sufficient statistical data. For these general cases, φ can be calculated by maximizing the likelihood of a class of distributions using the following equation. s :

[0274]

[0275] Such probability distributions can be characterized in various ways. For example, the probability distribution of a sample can be represented using a Markov model, where the probability of observing a methylation state depends on its genomic location and the state of previous methylation sites. Such a model can be established by quantifying the number of observed states and the number of k-mers at each genomic location or methylation site, which can be determined using the following equation:

[0276] f = {s i ;i∈(a,b)a <b<D}, Where s is the state of k-mer at a specific position.

[0277] Assume all data are represented as parameters of a probability distribution (i.e., all estimates of all observations). Several methods can be used to estimate the aforementioned hidden Z features. One method involves maximizing the likelihood using the following formula:

[0278]

[0279] Where θz is a distribution on D, similar to φ used to describe features. This likelihood can be maximized using the following expectation-maximization equation:

[0280] expect:

[0281] In the expected step, based on the current estimate of θ, the most likely q can be determined. i,z .

[0282] maximize:

[0283] In the maximization step, the most likely θ can be determined based on the current estimate of q.

[0284] The output of the above method is a set of Z parameters (θ) that describe the hidden features of the dataset.

[0285] These estimates can be made without relying on distribution assumptions, such as Gaussian or Bernoulli distributions.

[0286] Because this specific data representation greatly reduces the data size, most computations can be processed on a general-purpose computer or easily distributed across multiple computers for faster runtime.

[0287] The result of the first operation is a representative distribution corresponding to the Z unknown features. These features do not need to be known in advance or specified by experts.

[0288] Since this first operation can be used to estimate a set of biological characteristics, data from various sources can be merged and / or aggregated, including cfDNA data, data from different assays (e.g., RNA data, proteomics data, metabolomics data, etc.), data with different sequencing depths, and / or data generated by different sequencing methods.

[0289] Given a set of Z features (representative distributions), the set of segments can be transformed into a fixed set of features in several ways. For example, observations can be represented as histograms of locations and the aforementioned features. A Z×D zero matrix can be used as a starting point. For each segment, a Z×1 vector can be incremented at the segment location in D using the following equation:

[0290]

[0291] For each Z component.

[0292] Alternatively, or additionally, a fragment can be represented by its informativeness relative to a feature. For example, the probability of observing a fragment in an observation can be determined using the following equation:

[0293] FragFreq=p(f;φ i )

[0294] The ratio of the expected number of features to the total number of features can be determined using the following equation:

[0295]

[0296] Then, each segment can be represented as corresponding to the observation φ i The number of Z+1 FragFreq × InverseSampleFreq with Z features.

[0297] By representing the observations as a fixed size, these representations can be additive. Therefore:

[0298] • The representation of two sets from a fragment is equal to the sum of the representations of each set.

[0299] • By adding the D÷A columns of the matrix together, the representation can be simplified from Z×D to Z×A.

[0300] In the above representation, the set of fragments is used only once and can be computed solely based on known θ parameters. Therefore, this method overcomes the challenges described in this paper. Since the probability distribution can provide a smaller and biologically more accurate representation of the sample, the above ML method achieves computational feasibility without dividing the genome into small regions.

[0301] Example

[0302] Example 1: Classification of liver diseases using methylation data from patient plasma samples.

[0303] Plasma samples were collected from individual patients previously diagnosed with various liver diseases, including non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), and cirrhosis. The DNA methylation pattern across the entire genome was determined using the methods described above. First, cell-free DNA (cfDNA) was extracted from biological samples (e.g., plasma isolated from blood). The extracted DNA was then treated with sodium bisulfite to convert unmethylated cytosine to uracil, while methylated cytosine remained unchanged. Library preparation was then performed on the sodium bisulfite-treated DNA, including end repair and A-tailing, where the DNA ends were passivated and an adenine nucleotide was added to the 3' end of each strand. Subsequently, a specific linker was ligated to the DNA ends to enable the DNA to bind to the sequencing platform and provide primer binding sites during amplification. The linker-ligated DNA was then subjected to PCR amplification. The amplified DNA was sequenced using high-throughput DNA sequencing technology to determine the methylation pattern of DNA molecules in the cfDNA samples, resulting in approximately 500 million cfDNA reads with information about approximately 28 million CpGs.

[0304] Additionally, independent data derived from methylation microarrays are used to generate the features described using the methods described above. These microarrays contain data from various cell types in both healthy and diseased states, such as liver cells, brain cells, and heart cells. This method does not use labels indicating cell type or condition, but relies solely on methylation microarray data.

[0305] The methylation data is processed by computer to generate a set of three features (Z=3) distributed in the index family.

[0306] For each plasma sample (each containing approximately 500 million cfDNA reads), the generated features are used to convert the cfDNA into a fixed set of features. The cfDNA fragments are mapped to specific genomic locations, and then the fragments are converted into Z=3 features, one for each feature. The fragment frequency and desampling frequency for each fragment are then calculated, and another feature is calculated as the fragment frequency multiplied by the desampling frequency, resulting in 3+1=4 features per fragment.

[0307] The features of each fragment are added to the CpG position of its first CpG, ultimately transforming the entire sample into approximately 4 times 28 million features.

[0308] As described above, the process is further enhanced by additive features to further reduce the dimension of the sample representation from 4 multiplied by approximately 28 million to 4 multiplied by 100, for a total of 4 * 100 = 400 features.

[0309] While various machine learning training methods may be applicable to these representations, a simplified approach using a 1-nearest neighbor classifier is employed to demonstrate the effectiveness of the method disclosed in this paper. Using independent microarray data, the average representation of liver disease was computed, and a score was calculated for each sample indicating the distance between the sample and the average liver disease representation.

[0310] This method has been repeatedly used in a variety of applications, including for distinguishing between NASH and non-NASH (healthy) samples. Figure 4 ), distinguishing between risky NASH and non-risky NASH samples, where risky NASH is defined as individuals with NASH and stage 2 or higher fibrosis ( Figure 5 ), distinguishing between NASH samples with and without cirrhosis ( Figure 6 ), and to differentiate between early NASH, late NASH, and non-NASH (healthy) samples. Figure 7 ).

[0311] Figure 4 and Figure 6 The results shown indicate that the disclosed method can be used to identify subjects with liver disease. Figure 4 The identification of NASH is shown, and Figure 6 The identification of cirrhosis is shown. Figure 5 The disclosed method is shown to also be used for stratifying subjects with liver disease based on prognosis. Figure 7 The disclosed method is shown to be useful in differentiating between early and late-stage liver disease.

Claims

1. A method for identifying whether a subject has a liver disease or is at increased risk of developing a liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or the methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for the liver disease; and (d) based at least in part on the output, generating an electronic report indicative of the subject having the liver disease or being at the increased risk of developing the liver disease.

2. The method of claim 1, wherein the assaying comprises identifying the methylation pattern and the methylation level of the DNA molecules of the cfDNA sample, wherein the methylation pattern and the methylation level are processed using the trained ML algorithm.

3. The method of claim 1 or 2, wherein the assaying comprises sequencing.

4. The method of claim 3, further comprising, prior to the sequencing, treating the DNA molecules of the cfDNA sample with a reaction mixture comprising an enzyme for methylation-aware sequencing.

5. The method of claim 3, further comprising, prior to the sequencing, treating the DNA molecules of the cfDNA sample with a reaction mixture comprising bisulfite.

6. The method of any one of claims 1-5, wherein the assaying comprises amplification.

7. The method of claim 6, wherein the amplification comprises polymerase chain reaction (PCR).

8. The method of any one of claims 1-7, wherein the cfDNA sample is obtained or derived from a plasma sample, a serum sample, a urine sample, a saliva sample, or a liver tissue sample.

9. The method of any one of claims 1-8, further comprising fractionating a whole blood sample derived from the subject to provide the cfDNA sample.

10. The method of any one of claims 1-9, wherein (a) comprises subjecting the cfDNA sample to conditions sufficient to isolate, enrich, or extract a collection of DNA molecules, and wherein (b) comprises assaying the DNA molecules.

11. The method of claim 10, wherein (b) comprises selectively enriching the collection of DNA molecules corresponding to a panel of one or more genomic regions using nucleic acid primers or probes.

12. The method of claim 11, wherein the one or more genomic regions are selected from the genes listed in Table 1.

13. The method of claim 10 or 11, wherein the nucleic acid primers or probes have sequence complementarity to nucleic acid sequences of the panel of one or more genomic regions.

14. The method of any one of claims 1-9, wherein the cfDNA sample is assayed without nucleic acid isolation, enrichment, or extraction.

15. The method of any one of claims 1-14, wherein the subject is asymptomatic for the liver disease.

16. The method of any one of claims 1-15, wherein the output indicates whether the cfDNA sample is positive for the liver disease with at least 50% accuracy.

17. The method of claim 16, wherein the accuracy is determined by calculating the percentage of independent samples that are correctly identified as having or not having the liver disease.

18. The method of any one of claims 1-17, wherein the output indicates whether the cfDNA sample is positive for the liver disease with at least 50% clinical sensitivity.

19. The method of claim 18, wherein the clinical sensitivity is at least 50%.

20. The method of any one of claims 1-19, wherein the output indicates whether the cfDNA sample is positive for the liver disease with at least 50% clinical specificity.

21. The method of claim 20, wherein the clinical specificity is at least 50%.

22. The method of any one of claims 1-21, wherein the output indicates whether the cfDNA sample is positive for the liver disease with at least 50% positive predictive value.

23. The method of any one of claims 1-22, wherein the output indicates whether the cfDNA sample is positive for the liver disease with at least 50% negative predictive value.

24. The method of any one of claims 1-23, wherein the output indicates whether the cfDNA sample is positive for the liver disease with an area under the receiver operating characteristic (AUROC) of at least 0.

50.

25. The method of any one of claims 1-24, wherein the output indicates whether the cfDNA sample is positive for the liver disease with a positive likelihood ratio of at least about 1.

3.

26. The method of any one of claims 1-25, wherein the output indicates whether the cfDNA sample is negative for the liver disease with a negative likelihood ratio of at most about 0.

75.

27. The method of any one of claims 1-26, wherein the liver disease is early stage liver disease.

28. The method of any one of claims 1-26, wherein the liver disease is late stage liver disease.

29. The method of any one of claims 1-28, wherein the liver disease is nonalcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH).

30. The method of any one of claims 1-28, wherein the liver disease is fibrosis.

31. The method of any one of claims 1-28, wherein the liver disease is cirrhosis.

32. The method of any one of claims 1-28, wherein the liver disease is hepatocellular carcinoma (HCC).

33. The method of any one of claims 1-28, wherein the liver disease is hepatobiliary cancer.

34. The method of any one of claims 1-28, wherein the liver disease is viral hepatitis.

35. The method of any one of claims 1-28, wherein the liver disease is nonalcoholic fatty liver disease (NAFLD) or metabolic dysfunction-associated fatty liver disease (MASLD).

36. The method of any one of claims 1-28, wherein the liver disease is nonalcoholic fatty liver (NAFL) or steatosis.

37. The method of any one of claims 1-28, wherein the liver disease is metabolic dysfunction-associated fatty liver disease (MAFLD).

38. The method of any one of claims 1-28, wherein the liver disease is alcohol-related liver disease (ALD).

39. The method of any one of claims 1-28, wherein the liver disease is metabolic and alcohol-related / associated liver disease (MetALD).

40. The method of any one of claims 1-39, further comprising providing, to the subject, a therapeutic intervention for the liver disease based at least in part on the output.

41. The method of claim 40, wherein the liver disease is NASH, and wherein the therapeutic intervention is a vitamin E supplement, a weight loss agent, an anti-hypertensive agent, an anti-diabetic agent, a cholesterol-lowering agent, an exercise regimen, a diet regimen, a weight loss surgery, or a combination thereof.

42. The method of claim 40, wherein the liver disease is NASH, and wherein the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, a FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, a SCD-1 (stearoyl-CoA desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, a FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted genetic modifier, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pannexin inhibitor, a caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galectin-3 inhibitor, a mitochondrial uncoupler or decoupler, a structurally engineered fatty acid, or a combination thereof.

43. The method of claim 40, wherein the liver disease is NAFLD, and wherein the therapeutic intervention is a vitamin E supplement, a weight loss agent, an anti-hypertensive agent, an anti-diabetic agent, a cholesterol-lowering agent, an exercise regimen, a diet regimen, a weight loss surgery, or a combination thereof.

44. The method of claim 40, wherein the liver disease is NAFLD, and wherein the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, a FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, a SCD-1 (stearoyl-CoA desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, a FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted genetic modifier, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pannexin inhibitor, a caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galectin-3 inhibitor, a mitochondrial uncoupler or decoupler, a structurally engineered fatty acid, or a combination thereof.

45. The method of any one of claims 1-44, further comprising monitoring the liver disease in the subject at two or more time points based at least in part on the output.

46. The method of any one of claims 1-45, further comprising determining a likelihood or risk score that the subject has the liver disease or is at an increased risk of developing the liver disease.

47. The method of any one of claims 1-46, further comprising determining a molecular subtype, grade, stage, or severity of the liver disease.

48. The method of any one of claims 1-47, further comprising determining a prognosis of the liver disease.

49. The method of any one of claims 1-48, further comprising determining an outcome of the liver disease.

50. The method of any one of claims 1-49, further comprising determining the eligibility of the subject as a liver transplant donor or a liver transplant recipient.

51. The method of claim 50, wherein the subject is determined to be eligible as the liver transplant donor if the subject is not identified as having the liver disease or at an increased risk of developing the liver disease.

52. The method of claim 50, wherein the subject is determined to be eligible as the liver transplant recipient if the subject is identified as having the liver disease or at an increased risk of developing the liver disease.

53. The method of any one of claims 1-52, wherein the trained ML algorithm is trained using a set of independent samples related to the presence or increased risk of the liver disease.

54. The method of any one of claims 1-53, wherein (c) further comprises processing a set of clinical health data of the subject using the trained ML algorithm or another trained algorithm.

55. The method of claim 54, wherein the clinical health data comprises one or more quantitative measures selected from age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) level, alanine aminotransferase (ALT) level, gamma-glutamyl transferase (GGT), platelet count, triglyceride level, glycosylated hemoglobin (HbAlc) level, creatinine level, insulin level, prothrombin time, haptoglobin level, and glucose level.

56. The method of claim 54 or 55, wherein the clinical health data comprises one or more categorical measures selected from race, ethnicity, medication history or other clinical treatment history, alcohol use history, level of daily activity or physical fitness, genetic test results, blood test results, and imaging results.

57. The method of any one of claims 1-56, wherein the trained ML algorithm comprises a supervised ML algorithm.

58. The method of claim 57, wherein the supervised ML algorithm comprises a classifier or a regression.

59. The method of claim 57 or 58, wherein the supervised ML algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, a random forest, a linear regression, or a logistic regression.

60. The method of any one of claims 1-59, wherein the methylation pattern or the methylation level is represented by a profile, a sufficient statistic, or a parameter that is close to a sufficient statistic.

61. A method for monitoring a liver disease in a subject, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or a methylation level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or the methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for the liver disease; and (d) generating, based at least in part on the output, an electronic report indicative of a progression of the liver disease in the subject.

62. A method for identifying a liver disease prognosis of a subject having a liver disease or at an increased risk of developing a liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or a methylation level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or the methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for the liver disease; and (d) generating, based at least in part on the output, an electronic report indicative of the prognosis of the subject having the liver disease or at the increased risk of developing the liver disease.

63. A method for identifying a treatment for a subject having a liver disease or at increased risk of developing a liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or the methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for the liver disease; and (d) generating, based at least in part on the output, an electronic report indicative of the treatment for the subject having the liver disease or at the increased risk of developing the liver disease.

64. A method for determining a treatment response of a subject having a liver disease or at increased risk of developing a liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample derived from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or the methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for the liver disease; and (d) generating, based at least in part on the output, an electronic report indicative of the treatment response of the subject having the liver disease or at the increased risk of developing the liver disease.

65. A method for determining whether a subject has a liver disease or is at increased risk of developing a liver disease, the method comprising: (a) providing a cell-free nucleic acid sample derived from the subject; (b) assaying the cell-free nucleic acid sample, or a derivative thereof, to determine a methylation profile of the cell-free nucleic acid sample; and (c) processing the methylation profile using a trained machine learning (ML) algorithm to determine whether the subject has the liver disease or is at the increased risk of developing the liver disease, wherein the determination has a sensitivity of at least about 70% and a specificity of at least about 70%. ​