Method and system for detecting and assessing liver conditions - Patents.com
A cfDNA-based method using a machine learning algorithm for liver disease detection and monitoring addresses the invasiveness of biopsies, offering accurate and timely diagnostic solutions for liver conditions.
Patent Information
- Application Number
- JP2025542273
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-18
- Filing Date
- 2024-01-17
- Publication Date
- 2026-02-10
AI Technical Summary
Liver biopsy, the current gold standard for assessing liver fibrosis, is invasive and limited in widespread use, necessitating improved diagnostic tools for detecting and managing liver diseases such as fatty liver disease.
A method and system utilizing cell-free deoxyribonucleic acid (cfDNA) samples to determine methylation patterns through a trained machine learning algorithm, generating reports on liver disease presence, risk, progression, and potential treatments.
Provides accurate, non-invasive detection and monitoring of liver diseases with at least 50% accuracy, sensitivity, and specificity, enabling timely intervention and management.
Smart Images

Figure 2026504962000001_ABST
Abstract
Description
[Technical Field]
[0001] cross reference
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 439,716, filed January 18, 2023, which is incorporated herein by reference in its entirety. [Background technology]
[0002]
[0002] Liver disease can be caused by a variety of conditions, including infections, genetic disorders, obesity, and alcohol abuse. Blood tests can be used to measure the levels of enzyme biomarkers in the blood. Liver function tests, such as the international normalized ratio (INR), can be used to assess the degree of coagulation disorder, which is an indicator of liver dysfunction. Diagnostic imaging tools, such as ultrasound, magnetic resonance imaging (MRI), or computed tomography (CT), can be used to visualize signs of damage, scarring, or tumors in the liver. Incorporation by Reference
[0003] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, it is intended that the present specification supersede and / or take precedence over such conflicting material. Summary of the Invention [Means for solving the problem]
[0003]
[0004] Liver biopsy is the current gold standard for assessing liver fibrosis in patients with fatty liver disease. However, the inherent risks and invasiveness of biopsy evaluation may limit its widespread use. Improved diagnostic tools for detecting liver disease are considered essential for effective disease management.
[0004]
[0005] Recognizing the need for improved diagnostic tools for detecting liver disease, the present disclosure provides a method, system and kit for identifying or monitoring liver disease by processing acellular biological samples obtained from or derived from subject.The acellular biological samples (such as plasma samples) obtained from subject can be analyzed to identify liver disease, which can include, for example, determining the presence, absence or relative evaluation of liver disease.Such subject can include the subject with one or more liver diseases and the subject without one or more liver diseases. Liver diseases include, for example, alcoholic fatty liver disease (AFLD), alcohol-related liver disease (ALD), metabolic dysfunction alcohol-related liver disease (MetALD), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), fatty liver disease (SLD), metabolic dysfunction-associated fatty liver disease (MAFLD), metabolic dysfunction-associated fatty liver disease (MASLD), metabolic dysfunction-associated steatohepatitis (MASH), fatty liver disease of unknown cause (SLD of unknown cause), hepatitis, cancer (e.g., hepatocellular carcinoma or hepatobiliary carcinoma), and cirrhosis.
[0005]
[0006] In one aspect, the present disclosure provides a method for identifying whether a subject has or is at increased risk of developing liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicative of whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicative of the subject having or at increased risk of developing liver disease based at least in part on the output.
[0006]
[0007] In another aspect, the present disclosure provides a method for monitoring liver disease in a subject, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating progression of liver disease in the subject based at least in part on the output.
[0007]
[0008] In another aspect, the present disclosure provides a method for identifying a liver disease prognosis for a subject having or at increased risk of developing liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating, based at least in part on the output, an electronic report indicating a prognosis for the subject having or at increased risk of developing liver disease.
[0008]
[0009] In another aspect, the present disclosure provides a method for identifying a treatment for a subject having or at increased risk of developing liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating an electronic report indicating a treatment for the subject having or at increased risk of developing liver disease based at least in part on the output.
[0009]
[0010] In another aspect, the present disclosure provides a method for determining treatment response for a subject having or at increased risk of developing liver disease, the method comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from the subject; (b) assaying the cfDNA sample, or a derivative thereof, to determine a methylation pattern or methylation level of DNA molecules in the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained ML algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; and (d) generating, based at least in part on the output, an electronic report indicating treatment response for the subject having or at increased risk of developing liver disease.
[0010]
[0011] In some embodiments, the assay comprises identifying methylation patterns and methylation levels of DNA molecules of the cfDNA sample, wherein the methylation patterns and methylation levels are processed using a trained ML algorithm.
[0011]
[0012] In some embodiments, the assay comprises performing sequencing.
[0013] In some embodiments, the method further comprises treating the DNA molecules of the cfDNA sample with a reaction mixture comprising an enzyme for methylation-aware sequencing prior to sequencing.
[0012]
[0014] In some embodiments, the method further comprises treating the DNA molecules of the cfDNA sample with a reaction mixture comprising bisulfite prior to sequencing.
[0015] In some embodiments, the assay comprises amplification.
[0013]
[0016] In some embodiments, the amplification comprises polymerase chain reaction (PCR).
[0017] In some embodiments, the cfDNA sample is obtained or derived from a plasma sample, a serum sample, a urine sample, a saliva sample, or a liver tissue sample.
[0014]
[0018] In some embodiments, the method further comprises fractionating a whole blood sample from the subject to provide a cfDNA sample.
[0019] In some embodiments, (a) comprises subjecting the cfDNA sample to conditions sufficient to isolate, enrich, or extract a set of DNA molecules, and (b) comprises assaying the DNA molecules.
[0015]
[0020] In some embodiments, (b) comprises using nucleic acid primers or probes to selectively enrich a set of DNA molecules corresponding to a panel of one or more genomic regions.
[0016]
[0021] In some embodiments, the one or more genomic regions are selected from the group consisting of the genes listed in Table 1.
[0022] In some embodiments, the nucleic acid primers or probes have sequence complementarity to nucleic acid sequences of a panel of one or more genomic regions.
[0017]
[0023] In some embodiments, the cfDNA sample is assayed without nucleic acid isolation, enrichment, or extraction.
[0024] In some embodiments, the subject is asymptomatic for the liver disease.
[0018]
[0025] In some embodiments, the output indicates with at least 50% accuracy whether the cfDNA sample is positive for liver disease.
[0026] In some embodiments, accuracy is determined by calculating the percentage of independent samples that are correctly identified as having or not having liver disease.
[0019]
[0027] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with a clinical sensitivity of at least 50%.
[0028] In some embodiments, the clinical sensitivity is at least 50%.
[0020]
[0029] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with at least 50% clinical specificity.
[0030] In some embodiments, the clinical specificity is at least 50%.
[0021]
[0031] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with a positive predictive value of at least 50%.
[0032] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with a negative predictive value of at least 50%.
[0022]
[0033] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with an area under the receiver operating characteristic curve (AUROC) of at least 0.50.
[0034] In some embodiments, the output indicates whether the cfDNA sample is positive for liver disease with a positive likelihood ratio of at least about 1.3.
[0023]
[0035] In some embodiments, the output indicates whether the cfDNA sample is negative for liver disease with a negative likelihood ratio of up to about 0.75.
[0036] In some embodiments, the liver disease is early stage liver disease.
[0024]
[0037] In some embodiments, the liver disease is advanced liver disease.
[0038] In some embodiments, the liver disease is non-alcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH).
[0025]
[0039] In some embodiments, the liver disease is fibrosis.
[0040] In some embodiments, the liver disease is cirrhosis.
[0041] In some embodiments, the liver disease is hepatocellular carcinoma (HCC).
[0026]
[0042] In some embodiments, the liver disease is hepatobiliary cancer, including, for example, cholangiocarcinoma, angiosarcoma, gallbladder cancer, or undifferentiated embryonal sarcoma of the liver (UESL).
[0043] In some embodiments, the liver disease is viral hepatitis.
[0027]
[0044] In some embodiments, the liver disease is non-alcoholic fatty liver disease (NAFLD) or metabolic dysfunction-associated fatty liver disease (MASLD).
[0045] In some embodiments, the liver disease is non-alcoholic fatty liver disease (NAFL) or steatosis.
[0028]
[0046] In some embodiments, the liver disease is metabolic dysfunction-associated fatty liver disease (MAFLD).
[0047] In some embodiments, the liver disease is alcohol-related liver disease (ALD).
[0029]
[0048] In some embodiments, the liver disease is metabolic dysfunction alcohol-related liver disease (MetALD).
[0049] In some embodiments, the method further includes providing a therapeutic intervention for the liver disease to the subject based at least in part on the output.
[0030]
[0050] In some embodiments, the liver disease is NASH and the therapeutic intervention is vitamin E supplementation, a weight loss agent, an antihypertensive agent, an antidiabetic agent, a cholesterol-lowering agent, exercise therapy, diet therapy, or bariatric surgery.
[0031]
[0051] In some embodiments, the liver disease is NASH, and the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, an FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, an SCD-1 (stearoyl-coenzyme A desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, an FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted gene modifier including, for example, PNPLA3 or HSD17B13, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or a combination thereof.
[0032]
[0052] In some embodiments, the liver disease is NAFLD and the therapeutic intervention is vitamin E supplementation, a weight loss agent, an antihypertensive agent, an antidiabetic agent, a cholesterol-lowering agent, exercise therapy, diet therapy, bariatric surgery, or a combination thereof.
[0033]
[0053] In some embodiments, the liver disease is NAFLD and the therapeutic intervention is a GLP1 receptor agonist, an FGF analog, a THR agonist, an SCD-1 inhibitor, a FAS inhibitor, an FXR agonist, an ACC inhibitor, a PPAR agonist, e.g., a targeted gene modifier including PNPLA3 or HSD17B13, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or a combination thereof.
[0034]
[0054] In some embodiments, the method further includes monitoring the subject for liver disease at two or more time points based at least in part on the output.
[0055] In some embodiments, the method further comprises determining a likelihood or risk score that the subject has or is at increased risk of having liver disease.
[0035]
[0056] In some embodiments, the method further comprises determining the molecular subtype, grade, stage, or severity of the liver disease.
[0057] In some embodiments, the method further comprises determining a prognosis for the liver disease.
[0036]
[0058] In some embodiments, the method further comprises determining the subject's eligibility as a liver transplant donor or liver transplant recipient.
[0059] In some embodiments, a subject is determined to be eligible as a liver transplant donor if the subject is not identified as having or being at increased risk of developing liver disease.
[0037]
[0060] In some embodiments, a subject is determined to be eligible as a liver transplant recipient if the subject is identified as having or being at increased risk of developing liver disease.
[0038]
[0061] In some embodiments, the trained ML algorithm is trained using an independent set of samples that are associated with the presence or increased risk of liver disease.
[0062] In some embodiments, the trained ML algorithm is trained using a first set of independent samples associated with the presence of or increased risk of liver disease, and a second set of independent samples associated with the absence of or no increased risk of liver disease.
[0039]
[0063] In some embodiments, (c) further includes processing the set of clinical health data of the subject using the trained ML algorithm or another trained algorithm.
[0040]
[0064] In some embodiments, the clinical health data includes one or more quantitative measures selected from the group consisting of age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) level, alanine transaminase (ALT) level, gamma-glutamyltransferase (GGT), platelet count, triglyceride level, glycated hemoglobin (HbA1c) level, creatinine level, insulin level, prothrombin time, haptoglobin level, and glucose level.
[0041]
[0065] In some embodiments, the clinical health data includes one or more categorical measures selected from the group consisting of race, ethnicity, history of medication or other clinical treatment, alcohol consumption history, level of daily activity or fitness, genetic test results, blood test results, and imaging results.
[0042]
[0066] In some embodiments, the trained ML algorithm comprises a supervised ML algorithm.
[0067] In some embodiments, the supervised ML algorithm comprises a classifier or a regression.
[0043]
[0068] In some embodiments, the supervised ML algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, a random forest, a linear regression, or a logistic regression.
[0044]
[0069] In some embodiments, the methylation pattern or level is represented by a parameter of a distribution, a sufficient statistic, or an approximately sufficient statistic.
[0070] In another aspect, the disclosure provides a method for determining whether a subject has or is at increased risk of developing liver disease, the method comprising: (a) providing a cell-free nucleic acid sample from the subject; (b) assaying the cell-free nucleic acid sample, or a derivative thereof, to determine a methylome of the cell-free nucleic acid sample; and (c) processing the methylome using a trained machine learning (ML) algorithm to determine whether the subject has or is at increased risk of developing liver disease, wherein the determination has a sensitivity of at least about 70% and a specificity of at least about 70%.
[0045]
[0071] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG") [Brief explanation of the drawings]
[0046] [Figure 1]
[0072] FIG. 1 shows an exemplary workflow of a method for identifying or monitoring a liver disease state in a subject. [Figure 2]
[0073] FIG. 2 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein. [Figure 3]
[0074] FIG. 3 shows a schematic diagram of exemplary training data. [Figure 4]
[0075] Figure 4 shows the score distribution of cfDNA methylation data that distinguish non-alcoholic steatohepatitis (NASH) samples from non-NASH (healthy) samples. [Figure 5]
[0076] Figure 5 shows the score distribution of cfDNA methylation data distinguishing at-risk from non-risk NASH samples, where at-risk NASH is defined as individuals with stage 2 or above NASH and fibrosis. [Figure 6]
[0077] Figure 6 shows the score distribution of cfDNA methylation data that distinguishes NASH samples with cirrhosis from NASH samples without cirrhosis. [Figure 7]
[0078] Figure 7 shows the score distribution of cfDNA methylation data distinguishing early NASH samples, late NASH samples, and non-NASH (healthy) samples. DETAILED DESCRIPTION OF THE INVENTION
[0047]
[0079] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0048]
[0080] The differential pattern in nucleic acid molecules can be useful for detecting or stratifying liver disease.Provided herein is the method and system for detecting or stratifying liver disease using nucleic acid assay.For example, the methylation pattern of circulating deoxyribonucleic acid (DNA) can be detected in human plasma and used to stratify the severity of liver fibrosis in patients with NAFLD.
[0049]
[0081] Liver disease refers to several conditions that affect and damage the liver. There are four main stages of liver disease: 1) inflammation; 2) fibrosis; 3) cirrhosis; and 4) liver failure or liver cancer. Early-stage liver disease may be characterized by liver inflammation or enlargement or fibrosis. Over time, liver disease can lead to cirrhosis (scarring). As healthy liver tissue is gradually replaced by scar tissue, the liver can no longer function properly. If left untreated, liver disease can lead to more serious conditions, such as liver failure and cancer. Advanced liver disease, also known as end-stage liver disease or late liver disease, can be characterized by irreversible cirrhosis, liver failure, and stage 4 hepatitis C. Fatty liver disease (SLD) encompasses all of the various etiologies of steatosis.
[0050]
[0082] Nonalcoholic fatty liver disease (NAFLD) is a common chronic condition associated with progressive histological changes in the liver parenchyma. These NAFLD-related changes range from simple fat accumulation in hepatocytes, also known as hepatic steatosis or fatty liver, to more severe histological changes characterized by hepatocellular injury, fibrosis, and inflammation, which are characteristic of nonalcoholic steatohepatitis (NASH). NASH is also known as metabolic dysfunction-associated steatohepatitis (MASH).
[0051]
[0083] Nonalcoholic fatty liver disease (NAFLD) is a common cause of chronic liver pathology worldwide. The prevalence of NAFLD strongly correlates with the increasing incidence of diabetes, obesity, and metabolic syndrome in the general population. The earliest stage of NAFLD, simple steatosis, is often nonprogressive and remains asymptomatic. Appropriate lifestyle and dietary modifications at this early stage can reverse the diseased liver to a healthy state. Timely detection and risk stratification of NAFLD are essential because simple steatosis can progress to a severe fibrotic stage, potentially promoting carcinogenesis.
[0052]
[0084] NAFLD is also called metabolic dysfunction-associated fatty liver disease (MASLD).MASLD includes patients with hepatic steatosis and at least one of five cardiometabolic risk factors.Another category other than pure MASLD, called metabolic dysfunction alcohol-related liver disease (MetALD), refers to patients with MASLD who consume more alcohol per week (for example, 140g / week and 210g / week for women and men, respectively).Patients with liver disease whose metabolic parameters are not obtained and whose cause is unknown can be called unknown cause fatty liver disease (unknown cause SLD). The methods described herein can be used, for example, to identify, stratify, or differentiate between any liver disease type or subtype described herein and in Rinella et al., Hepatology 78(6):1966-1986, December 2023, DOI:10.1097 / HEP.0000000000000520, which is incorporated herein by reference in its entirety.
[0053]
[0085] Extracellular circulating nucleic acids found in biological fluids, including blood, may serve as promising non-invasive biomarkers for liver disease. For example, the epigenetic signature of circulating cfDNA, such as methylation patterns, may be useful for detecting the presence of disease and monitoring disease progression. Intracellular miRNAs are usually involved in the regulation of gene expression, but after being released by apoptotic cells, miRNAs can remain highly stable in the extracellular environment for a long period of time. Therefore, the profile of circulating nucleic acids may reflect pathogenic processes in the body's tissues and organs, enabling highly sensitive non-invasive detection of liver disease.
[0054] definition
[0086] As used herein, the term "nucleic acid" generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or their analogs. Nucleic acids may have any three-dimensional structure and may perform any function, known or unknown. Non-limiting examples of nucleic acids include DNA, ribonucleic acid (RNA), coding or non-coding regions of a gene or gene fragment, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be made before or after assembly of the nucleic acid. The sequence of nucleotides in a nucleic acid may be interrupted by non-nucleotide components. Nucleic acids can be further modified after polymerization, such as by conjugation or conjugation with a reporter agent.
[0055]
[0087] As used herein, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" generally refer to polynucleotides, such as deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs and / or combinations thereof (e.g., mixtures of DNA and RNA). Nucleic acid molecules can have a variety of lengths. Nucleic acid molecules can have a length of at least about 5 bases, 10 bases, 20 bases, 30 bases, 40 bases, 50 bases, 60 bases, 70 bases, 80 bases, 90 bases, 100 bases, 110 bases, 120 bases, 130 bases, 140 bases, 150 bases, 160 bases, 170 bases, 180 bases, 190 bases, 200 bases, 300 bases, 400 bases, 500 bases, 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, or 50 kb, or any number of bases between any two of the above values. Oligonucleotides typically consist of specific sequences of the four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (though if the polynucleotide is RNA, uracil (U) is substituted for thymine (T)). Thus, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" are intended, at least in part, to be alphabetical representations of polynucleotide molecules. Alternatively, these terms can be applied to the polynucleotide molecules themselves. This alphabetical representation can be input into a database on a computer having a central processing unit and / or used in bioinformatics applications, such as functional genomics and homology searching. Oligonucleotides may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0056]
[0088] As used herein, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" generally refer to polynucleotides such as deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs and / or combinations thereof (e.g., mixtures of DNA and RNA). Nucleic acid molecules can be of various lengths. Nucleic acid molecules may have a length of at least 5 bases, at least 10 bases, at least 20 bases, at least 30 bases, at least 40 bases, at least 50 bases, at least 60 bases, at least 70 bases, at least 80 bases, at least 90 bases, at least 100 bases, at least 110 bases, at least 120 bases, at least 130 bases, at least 140 bases, at least 150 bases, at least 160 bases, at least 170 bases, at least 180 bases, at least 190 bases, at least 200 bases, at least 300 bases, at least 400 bases, at least 500 bases, at least 1 kilobase (kb), at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 10 kb, at least 50 kb, or any number of bases between any two of the above values. Oligonucleotides typically consist of specific sequences of the four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (though if the polynucleotide is RNA, uracil (U) is substituted for thymine (T)). Thus, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" are intended, at least in part, to be alphabetical representations of polynucleotide molecules. Alternatively, these terms can be applied to the polynucleotide molecules themselves. This alphabetical representation can be input into a database on a computer having a central processing unit and / or used in bioinformatics applications, such as functional genomics and homology searching. Oligonucleotides may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0057]
[0089] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule in a starting population of nucleic acid molecules having a nucleotide sequence whose presence, amount, and / or sequence, or changes in one or more of these, are desired to be determined. The target nucleic acid can be any type of nucleic acid, including DNA, RNA, and analogs thereof. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.
[0058]
[0090] As used herein, the term "target" generally refers to a genomic region within a marker gene or marker region.As used herein, the term "reference" generally refers to a sample obtained or derived from a subject diagnosed with liver disease or a subject who has been diagnosed with liver disease and has been found to be negative for clinical indications of liver disease (e.g., a healthy subject or control subject who does not have liver disease).
[0059]
[0091] As used herein, the terms "locus" or "region" are generally interchangeable and refer to a specific genomic region on the genome represented by a chromosome number, a start position, and an end position.
[0060]
[0092] As used herein, the term "subject" generally refers to an entity or medium that has testable or detectable genetic information. A subject may be a human or an individual, such as a patient. A subject may be a vertebrate, e.g., a mammal. Non-limiting examples of mammals include mice, monkeys, humans, farm animals, sport animals, and pets.
[0061]
[0093] As used herein, the term "sample" generally refers to a biological sample obtained or derived from a subject, for example. A sample can be obtained from tissues and / or cells, or from the environment of tissues and / or cells. A sample can be a cell-free or substantially cell-free biological sample, or can be processed or fractionated to produce a cell-free biological sample. For example, a cell-free biological sample includes cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and derivatives thereof. A cell-free biological sample can be obtained from a subject or derived from a subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free DNA collection tube. A cell-free biological sample can be obtained from a whole blood sample by fractionation. In some embodiments, a biological sample or its derivatives can contain cells. For example, the biological sample may be a blood sample or a derivative thereof (e.g., blood collected by a collection tube or blood drop), a liver tissue sample, a vaginal sample (e.g., a vaginal swab), or a cervical sample (e.g., a cervical swab). In some examples, the sample may comprise, be obtained from, or be derived from a tissue biopsy sample (e.g., a liver tissue biopsy sample), a cell biopsy sample, blood (e.g., whole blood), plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, urine, extracellular fluid, dried blood spot, cultured cells, culture medium, discarded tissue, plant material, synthetic protein, bacterial and / or viral sample, fungal tissue, archaea, or protozoa. The sample may be isolated from the source prior to collection. Non-limiting examples include fingerprints, saliva, urine, blood, stool, semen, or other bodily fluids isolated from their primary source prior to collection. In some examples, the sample is isolated from its primary source (cells, tissues, bodily fluids such as blood, environmental samples, etc.) during sample preparation. The sample may be purified or otherwise concentrated from its primary source, or may be unpurified or unconcentrated. In some embodiments, the primary source is homogenized before further processing. The sample may also be filtered or centrifuged to remove buffy coat, lipids, or particulate matter.The sample may also be purified or enriched for nucleic acids or treated with RNase or DNase. The sample may contain intact, fragmented, or partially degraded tissues and / or cells.
[0062]
[0094] The sample can be obtained from a subject suspected of having or having a disease or disorder, and the subject may or may not have been diagnosed with the disease or disorder. The subject may need a second opinion. The disease or disorder can be an infectious disease, an immune disorder or disease, cancer, a genetic disease, a degenerative disease, a lifestyle-related disease, or an injury. The infectious disease can be caused by bacteria, viruses, fungi, and / or parasites. The cancer can be hepatocellular carcinoma (HCC) or hepatobiliary carcinoma, including, for example, cholangiocarcinoma, angiosarcoma, gallbladder cancer, or hepatic undifferentiated embryonal sarcoma (UESL) of the liver.
[0063]
[0095] Sample components (including nucleic acids) can be tagged with, for example, distinguishable tags to enable sample multiplexing. Some non-limiting examples of distinguishable tags include fluorophores, magnetic nanoparticles, and nucleic acid barcodes. Fluorophores include fluorescent proteins such as GFP, YFP, RFP, eGFP, mCherry, tdtomato, FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 680, Alexa Fluor 750, Pacific Blue, coumarin, BODIPY FL, Pacific Green, Oregon Green, Cy3, Cy5, Pacific Orange, TRITC, Texas Red, phycoerythrin, allophycocyanin, or other fluorophores. One or more barcode tags can be attached (e.g., by coupling or ligation) to cell-free nucleic acids (e.g., cfDNA) in a sample before sequencing. The barcodes can uniquely tag cfDNA molecules in a sample. Alternatively, the barcodes can non-uniquely tag cfDNA molecules in a sample. The barcodes can non-uniquely tag cfDNA molecules in a sample, so that additional information obtained from the cfDNA molecules (e.g., at least a portion of the endogenous sequence of the cfDNA molecules) combined with the non-unique tags can serve as a unique identifier (e.g., for uniquely identifying the cfDNA molecules in the sample from other molecules). For example, cfDNA sequence reads with unique identity (e.g., from a given template molecule) can be detected at least in part based on sequence information including one or more consecutive base regions at one or both ends of the sequence read, the length of the sequence read, and / or the sequence of the attached barcode at one or both ends of the sequence read.DNA molecules can be uniquely identified without tagging by dividing the DNA (e.g., cfDNA) sample into a large number (e.g., at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, or at least about 100,000) of different, distinct subunits (e.g., compartments, wells, or droplets) prior to amplification, such that the amplified DNA molecules can be uniquely separated and identified as derived from each individual input molecule of DNA.
[0064]
[0096] Any number of samples can be multiplexed.For example, multiplex analysis can comprise at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 or more samples.Identifiable tags can provide a way to check each sample for its origin, or can lead different samples to be divided into different areas or solid supports.
[0065]
[0097] Any number of samples can be mixed before analysis without tagging or multiplexing.For example, multiplex analysis can include at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 or more samples. Samples can be multiplexed without tagging using a combinatorial pooling design, where samples are mixed into pools in a manner that allows signals from individual samples to be separated from the analyzed pool using computational demultiplexing.
[0066]
[0098] The sample can be enriched before sequencing. For example, cfDNA molecules can be selectively enriched for one or more regions from a subject's genome or transcriptome, or non-selectively enriched. For example, cfDNA molecules can be selectively enriched for one or more regions from a subject's genome or transcriptome by targeted sequence capture (e.g., using a panel), selective amplification, or targeted amplification. As another example, cfDNA molecules can be non-selectively enriched for one or more regions from a subject's genome or transcriptome by global amplification. In some embodiments, amplification includes global amplification, whole genome amplification, or non-selective amplification. cfDNA molecules can be size-selected for fragments having a length within a predetermined range. For example, size selection can be performed on DNA fragments with lengths ranging from about 40 base pairs (bp) to about 250 bp before adaptor ligation. As another example, size selection can be performed on DNA fragments with lengths ranging from about 160 bp to about 400 bp after adaptor ligation.
[0067]
[0099] As used herein, the terms "amplifying" and "amplification" are used interchangeably and generally refer to producing one or more copies of a nucleic acid or "amplification product." The term "DNA amplification" generally refers to producing one or more copies of a DNA molecule or "amplified DNA product." The term "reverse transcription amplification" generally refers to the production of deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template by the action of reverse transcriptase. Amplification can be performed by the polymerase chain reaction (PCR), which relies on the synthesis of a new strand of DNA complementary to the original template strand using a DNA polymerase.
[0068]
[0100] As used herein, the term "polymerase chain reaction" or "PCR" generally refers to a method for increasing the concentration of a segment of a target sequence in a mixture of genomic DNA without cloning or purification. This process for amplifying a target sequence can involve introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target sequence, followed by thermal cycling in a precise order in the presence of a DNA polymerase. The two primers may be complementary to each strand of a double-stranded target sequence. To achieve amplification, the mixture is denatured, allowing the primers to anneal to their complementary sequences within the target molecule. Following annealing, the primers can be extended by a polymerase to form a new pair of complementary strands. Denaturation, primer annealing, and polymerase extension can be repeated multiple times (e.g., denaturation, annealing, and extension constitute one "cycle," and there can be multiple "cycles") to obtain a highly concentrated amplified segment of the desired target sequence. The length of the amplified segment of the desired target sequence is determined by the relative positions of the primers with respect to each other, and therefore this length is a controllable parameter. Because of the iterative aspect of the process, this method is referred to as the "polymerase chain reaction," or "PCR." Because the desired amplified segment of the target sequence becomes the predominant sequence (in terms of concentration) in the mixture, the amplified segment can be referred to as "PCR amplified," "PCR product," or "amplicon."
[0069]
[0101] As used herein, the term "methylation" refers to 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC), which includes cytosine residues that are part of the sequence CG, also referred to as CpG dinucleotides. Some CG dinucleotides in the human genome are methylated, while others are unmethylated. Additionally, methylation can be cell- and tissue-specific, such that a particular CG dinucleotide can be methylated in a particular cell but not in another cell, or in a particular tissue but not in another tissue. DNA methylation is an important regulator of gene transcription. Abnormal DNA methylation patterns, both hypermethylated and hypomethylated, compared to normal tissues can be associated with numerous human malignancies. In some embodiments, the 5hmC residues in the sequence can be subjected to glucosylation, followed by bisulfite treatment, bisulfite-free enzyme treatment, or methylation-sensitive restriction enzyme digestion. For example, glucosylation can be performed using glucosyltransferase.
[0070]
[0102] As used herein, the terms "methylation state," "methylation situation," and "methylation profile" generally refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule (e.g., a DNA molecule) containing a methylated cytosine is considered to be methylated (e.g., a nucleic acid molecule in a methylation state is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0071]
[0103] As used herein, the term "DNA template" generally refers to a sample DNA containing a target sequence. At the start of the reaction, high temperature is applied to the original double-stranded DNA molecule, causing the strands to separate from each other.
[0072]
[0104] As used herein, the term "primer" generally refers to a short piece of single-stranded DNA that is complementary to a DNA template. Polymerase begins synthesizing new DNA from the end of the primer.
[0073]
[0105] As used herein, the term "sensitivity" or "clinical sensitivity" generally refers to the percentage of a set of disease samples that yield a positive diagnostic result. For example, such disease samples can be analyzed to detect DNA methylation values above a threshold that distinguishes disease (e.g., liver disease) samples from non-disease (e.g., healthy or control) samples. In some embodiments, a positive result is defined as a histologically confirmed disease that reports a DNA methylation value above a threshold (e.g., a disease-associated range), and a false negative result is defined as a histologically confirmed disease that reports a DNA methylation value below a threshold (e.g., a disease-unassociated range). The sensitivity value may reflect the probability that a DNA methylation measurement value for a given marker obtained from a diseased sample falls within the range of disease-associated measurements. The clinical relevance of a calculated sensitivity value may represent an estimate of the probability that a given marker, when applied to a subject with a clinical condition, can detect or predict the presence of that clinical condition.
[0074]
[0106] As used herein, the term "specificity" or "clinical specificity" generally refers to the percentage of a set of non-disease samples that yield a negative diagnostic result. For example, such non-disease samples can be analyzed to detect DNA methylation values below a threshold that distinguishes disease (e.g., liver disease) samples from non-disease (e.g., non-liver disease) samples. In some embodiments, a negative is defined as a histologically confirmed non-disease sample reporting a DNA methylation value below a threshold (e.g., a non-disease range), and a false positive is defined as a histologically confirmed non-disease sample reporting a DNA methylation value above a threshold (e.g., a disease-associated range). The specificity value may reflect the probability that a DNA methylation measurement value for a given marker obtained from a non-liver disease (e.g., healthy or control) sample falls within the range of non-disease-associated measurements. The clinical relevance of a calculated specificity value may represent an estimate of the probability that a given marker can detect or predict the absence of a clinical condition when applied to subjects without that clinical condition.
[0075]
[0107] As used herein, the term "AUC" or "AUROC" generally refers to the area under the receiver operating characteristic (ROC) curve. The ROC curve may be a plot of the true positive rate (TPR) against the false positive rate (FPR) for several different possible thresholds or cut points for a diagnostic test, thereby showing the trade-off between sensitivity and specificity depending on the selected cut point (e.g., every increase in sensitivity is accompanied by a decrease in specificity). The area under the ROC curve (AUC) can be a measure of the accuracy of a diagnostic test (e.g., the larger the area, the higher the diagnostic accuracy), with an optimal value of 1. In comparison, a random test may have a ROC curve that is diagonal and an AUC of 0.5 (e.g., representing a random or worthless test).
[0076] Methods of the present disclosure
[0108] Current diagnostic tools for liver disease can be inaccessible and imperfect. Blood tests can be used to measure the levels of enzyme biomarkers in the blood. Liver function tests, such as the international normalized ratio (INR), can be used to assess the degree of coagulation disorder, an indicator of liver dysfunction. Imaging tools, such as ultrasound, MRI, or CT, can be used to visualize signs of liver damage, scarring, or tumors. Liver biopsy is the current gold standard for assessing liver fibrosis in patients with fatty liver disease. However, the inherent risks and invasiveness of biopsy evaluation limit its widespread use. Therefore, there is an urgent clinical need for accurate, affordable, and noninvasive diagnostic methods for detecting and monitoring liver disease toward effective disease management treatment.
[0077]
[0109] The present disclosure provides a method, system, and kit for identifying or monitoring liver disease by processing acellular biological samples obtained from or derived from a subject.Acellular biological samples (e.g., plasma samples) obtained from a subject can be analyzed to identify liver disease, which can include, for example, determining the presence, absence, or relative value of liver disease.Such subjects can include subjects with one or more liver diseases and subjects without one or more liver diseases.Liver diseases include, for example, alcoholic or non-alcoholic fatty liver disease, non-alcoholic steatohepatitis, hepatitis, cancer (e.g., hepatocellular carcinoma), and cirrhosis.
[0078]
[0110] 1 illustrates an exemplary workflow of a method for identifying or monitoring a subject's liver disease state according to embodiments disclosed herein. In one aspect, the present disclosure provides a method 100 for identifying or monitoring a subject's liver disease state. Method 100 may include assaying a first acellular biological sample from the subject with a first assay to generate a first dataset (operation 101). Next, based at least in part on the generated first dataset, method 100 may optionally include assaying a second acellular biological sample from the subject with a second assay (e.g., a different assay from the first assay) to generate a second dataset that indicates the liver disease state with greater specificity than the first dataset (operation 102). For example, DNA molecules extracted from the second acellular plasma sample can be sequenced to generate a set of sequence reads indicative of the subject's liver disease state. In some embodiments, the first acellular biological sample is obtained from the subject at a first time point for processing in the first assay. Optionally, a second acellular biological sample is then obtained from the same subject at a second time point for processing in a second assay. In some embodiments, the acellular biological sample can be obtained from the subject and then aliquoted to generate a first acellular biological sample and a second acellular biological sample, which can then be processed in a first assay and a second assay, respectively. The first dataset and / or the second dataset can then be processed using a trained machine learning algorithm to determine the subject's liver disease status (operation 103). The trained machine learning algorithm may be configured to identify liver disease with at least about 80% accuracy for 50 independent samples. A report indicating (e.g., identifying or providing an indication of) the presence or susceptibility of the subject to liver disease is then electronically generated (operation 104).
[0079]
[0111] Acellular biological samples can be obtained from subjects with liver disease conditions (e.g., liver disease or liver pathology), from subjects suspected of having liver disease conditions, or from subjects without or not suspected of having liver disease conditions.Diseases or disorders can be diseases or disorders that affect the liver.Non-limiting examples of such diseases or disorders include fatty liver disease, alcoholic fatty liver disease, non-alcoholic fatty liver disease, steatohepatitis, non-alcoholic steatohepatitis, hepatitis (e.g., hepatitis A, hepatitis B, or hepatitis C), liver cancer (e.g., hepatocellular carcinoma), hepatobiliary cancer, including cholangiocarcinoma, angiosarcoma, gallbladder cancer, or hepatic undifferentiated embryonal sarcoma (UESL), liver cirrhosis, hemochromatosis, Wilson's disease, obesity, diabetes, hypertension, and other liver pathologies disclosed herein.
[0080]
[0112] The sample can be obtained before and / or after the subject is treated with a disease or disorder.The sample can be obtained before and / or after the subject is treated with a disease or disorder.The sample can be obtained during treatment or treatment regimen.Multiple samples can be obtained from the subject to monitor the effect of treatment over time, including starting before the start of treatment.The sample can be obtained from the subject to monitor abnormal tissue-specific cell death or organ transplantation.
[0081]
[0113] Sample can be obtained from the subject who is suspected of having disease or disorder.Sample can be obtained from the subject who experiences unexplained symptoms, such as fatigue, nausea or vomiting, yellowing of skin or eyes (jaundice), swelling of legs or ankles, abdominal swelling (ascites), abdominal pain, itchy skin, weight gain, weight loss, aching pain, pain, tremors, weakness, drowsiness, or disorientation or confusion.Sample can be obtained from the subject who has explained symptoms. The sample can be obtained from a subject who is at risk of developing a disease or disorder due to one or more factors, e.g., family and / or personal history, age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) levels, alanine transaminase (ALT) levels, gamma-glutamyltransferase (GGT), platelet count, triglyceride levels, haptoglobin levels, glucose levels, environmental exposures, lifestyle risk factors, the presence of other risk factors, or a combination thereof.
[0082]
[0114] Samples can be obtained from healthy subjects or individuals. In some embodiments, samples can be obtained longitudinally from the same subject or individual. In some embodiments, longitudinally obtained samples can be analyzed for the purposes of monitoring an individual's health and early detection of health problems (e.g., early diagnosis of liver disease). In some embodiments, samples can be collected in a home environment or at a hospital and then transported by mail, courier delivery, or other transportation method before analysis. For example, a home user can collect a blood spot sample by pricking their finger. The blood spot sample can be dried and then transported by mail before analysis. In some embodiments, longitudinally obtained samples can be used to monitor responses to stimuli expected to affect health, athletic performance, or cognitive performance. Non-limiting examples include response to drug therapy, dietary therapy, and / or exercise therapy. In some embodiments, individual samples are multipurpose, allowing for methylation profiling to obtain clinically relevant information, but can also be used to obtain information about an individual's personal or family ancestry.
[0083]
[0115] In some embodiments, biological sample is a nucleic acid sample that contains one or more nucleic acid molecules.Nucleic acid molecule can be acellular or substantially acellular nucleic acid molecule, such as cell-free DNA (cfDNA) or cell-free RNA (cfRNA) or a mixture thereof.Nucleic acid molecule can be derived from various sources, including human, mammal, non-human mammal, ape, monkey, chimpanzee, reptile, amphibian or avian sources.In addition, sample can be extracted from various animal body fluids that contain acellular sequences, including but not limited to blood, serum, plasma, bone marrow, vitreous, sputum, feces, urine, tears, sweat, saliva, semen, mucous secretions, mucus, cerebrospinal fluid, pleural fluid, amniotic fluid and lymphatic fluid.
[0084]
[0116] The acellular biological sample may contain one or more analytes that can be assayed, such as cfRNA molecules suitable for assaying to generate transcriptomics data, cfDNA molecules suitable for assaying to generate genomic data, proteins suitable for assaying to generate proteomics data, metabolites suitable for assaying to generate metabolomics data, or a mixture or combination thereof. One or more such analytes (e.g., cfRNA molecules, cfDNA molecules, proteins, or metabolites) can be isolated or extracted from one or more acellular biological samples of a subject for downstream assay using one or more suitable assays.
[0085]
[0117] After obtaining an acellular biological sample from a subject, the sample can be processed to generate a dataset indicative of the subject's liver disease status. For example, the presence, absence, or quantitative assessment of the sample's nucleic acid molecules in a panel of liver disease-associated genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-associated genomic loci), proteomic data including quantitative measures of proteins in a dataset of a panel of liver disease-associated proteins, and / or metabolomic data including quantitative measures of a panel of liver disease-associated metabolites can indicate the liver disease status. Processing the acellular biological sample obtained from a subject can include (i) subjecting the sample to conditions sufficient to isolate, enrich, or extract multiple nucleic acid molecules, proteins, and / or metabolites, and (ii) assaying the multiple nucleic acid molecules, proteins, and / or metabolites to generate a dataset. In some embodiments, the quantitative measure of DNA can include the presence, absence, or degree of methylation, hypermethylation, and / or hypomethylation. Alternatively, or in combination, the quantitative measure of DNA can include the presence, absence, or degree of a variant pattern. The variant pattern may include genetic mutations, single nucleotide polymorphisms (SNPs), or copy number variations. Alternatively, or in combination, the quantitative measure of DNA may include the presence, absence, or degree of viral genome patterns.
[0086]
[0118] In some embodiments, multiple nucleic acid molecules are extracted from acellular biological samples and subjected to sequencing to generate multiple sequencing reads. The nucleic acid molecules may include RNA or DNA. Nucleic acid molecules (e.g., RNA or DNA) can be extracted from acellular biological samples by various methods, such as using a nucleic acid extraction kit. The extraction method can extract all RNA molecules or DNA molecules from the sample. Alternatively, the extraction method can selectively extract a portion of RNA molecules or DNA molecules from the sample. The RNA molecules extracted from the sample can be converted into DNA molecules by reverse transcription (RT).
[0087]
[0119] Sequencing of nucleic acid molecules can be performed by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing-by-synthesis (SBS), sequencing-by-ligation, sequencing-by-hybridization, and RNA-Seq (Illumina).
[0088]
[0120] Sequencing can involve nucleic acid amplification (e.g., of RNA or DNA molecules). In some embodiments, the nucleic acid amplification is polymerase chain reaction (PCR). A suitable number of PCR rounds (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to sufficiently amplify an initial amount of nucleic acid (e.g., RNA or DNA) to a desired input amount for subsequent sequencing. In some cases, PCR can be used for global amplification of target nucleic acids. This amplification can involve first using adapter sequences that can be ligated to different molecules, followed by PCR amplification using universal primers. PCR can be performed using any of several commercially available kits, such as those provided by Life Technologies, Affymetrix, Promega, Qiagen, etc. In other cases, only certain target nucleic acids within a population of nucleic acids can be amplified. Specific primers, optionally with adapter ligation, can be used to selectively amplify certain targets for downstream sequencing. PCR can include targeted amplification of one or more genomic loci, such as genomic loci associated with liver disease. Sequencing can involve simultaneous RT and PCR, for example, using the OneStep RT-PCR kit protocol from Qiagen, NEB, Thermo Fisher Scientific, or Bio-Rad.
[0089]
[0121] The RNA molecules or DNA molecules that are isolated or extracted from acellular biological samples can be tagged with, for example, distinguishable tags, so as to enable multiple samples to be multiplexed.Any number of RNA samples or DNA samples can be multiplexed.For example, multiplexing reaction can comprise at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more than 100 RNA or DNA from the initial acellular biological samples. For example, multiple cell-free biological samples can be tagged with sample barcodes to trace each DNA molecule back to the sample (and subject) from which it originated. Such tags can be attached to RNA or DNA molecules by ligation or by PCR amplification using primers.
[0090]
[0122] After nucleic acid molecule is subjected to sequencing, suitable bioinformatics process can be carried out on sequence read to generate data indicating the presence, absence or relative evaluation of liver disease.For example, sequence read can be aligned with one or more reference genomes (for example, one or more species genomes, for example, human genome).The aligned sequence read can be quantified at one or more genome loci to generate a data set indicating liver disease.For example, the quantification of the sequences corresponding to multiple genome loci associated with liver disease can generate a data set indicating liver disease.
[0091]
[0123] In some cases, acellular biological samples can be processed without any nucleic acid extraction.For example, by using a probe that is configured to selectively enrich the nucleic acid (for example, RNA or DNA) molecules corresponding to multiple liver disease-related genome loci, liver disease in subjects can be identified or monitored.The probe can be a nucleic acid primer.The probe can have sequence complementarity with the nucleic acid sequence from one or more of multiple liver disease-related genome loci or genome regions. The plurality of liver disease associated genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, or more distinct liver disease associated genomic loci or genomic regions. The plurality of liver disease-associated genomic loci or genomic regions can include one or more members (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, or more) selected from the group consisting of the genes listed in Table 1. A liver disease-associated genomic locus or region may be associated with age, race, ethnicity, BMI, blood glucose level, or other liver disease condition or comorbidity.
[0092] Table 1-1
[0093] Table 1-2
[0094] Table 1-3
[0095] Table 1-4
[0096] Table 1-5
[0097] Table 1-6
[0098] Table 1-7
[0099] Table 1-8
[0100] Table 1-9
[0101] Table 1-10
[0102] Table 1-11
[0103] Table 1-12
[0104] Table 1-13
[0105] Table 1-14
[0106] Table 1-15
[0107] Table 1-16
[0108] Table 1-17
[0109] Table 1-18
[0110] Table 1-19
[0111] Table 1-20
[0112] Table 1-21
[0113] Table 1-22
[0114] Table 1-23
[0115] Table 1-24
[0116] Table 1-25
[0117] Table 1-26
[0118] Table 1-27
[0119] Table 1-28
[0124] The probes can be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity with the nucleic acid sequences (e.g., RNA or DNA) of one or more genomic loci (e.g., liver disease-associated genomic loci). These nucleic acid molecules can be primers or enrichment sequences. Assays of cell-free biological samples using probes selective for one or more genomic loci (e.g., liver disease-associated genomic loci) can include the use of array hybridization (e.g., microarray-based), PCR, or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some embodiments, DNA or RNA may be assayed by one or more of isothermal DNA / RNA amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (e.g., CRISPR typing PCR (ctPCR), specific highly sensitive enzyme reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR transreporter (DETECTR), and CRISPR-mediated analog multi-event recording device (CAMERA)), and laser transmission spectroscopy (LTS).
[0120]
[0125] The assay readout can be quantified at one or more genomic loci (e.g., liver disease-associated genomic loci) to generate data indicative of a liver disease state. For example, array hybridization or PCR quantification corresponding to multiple genomic loci (e.g., liver disease-associated genomic loci) can generate data indicative of a liver disease state. The assay readout can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof. The assay can be a home test configured to be performed in a home environment.
[0121]
[0126] In some embodiments, a plurality of assays are used to process a cell-free biological sample from a subject.For example, a first assay can be used to process a first cell-free biological sample obtained from or derived from a subject to generate a first data set, and a second assay different from the first assay can be used to process a second cell-free biological sample obtained from or derived from a subject based at least in part on this first data set to generate a second data set that indicates a liver disease state.The first assay can be used to screen or process a cell-free biological sample from a set of subjects, while a second or subsequent assay can be used to screen or process a cell-free biological sample from a smaller subset of the set of subjects.The first assay can have low cost and / or high sensitivity for detecting one or more liver disease states (e.g., liver disease or liver pathology), which is suitable for screening or processing a cell-free biological sample from a relatively large set of subjects. The second assay may be more cost-effective and / or more specific for detecting one or more liver disease conditions, making it suitable for screening or processing acellular biological samples from a relatively small set of subjects (e.g., a subset of subjects screened using the first assay). The second assay can generate a second dataset with higher specificity (e.g., for one or more liver disease conditions) than the first dataset generated using the first assay. As an example, one or more acellular biological samples can be processed using a cfDNA assay for a large set of subjects, followed by a metabolomics assay for a smaller subset of subjects, or vice versa. The smaller subset of subjects can be selected at least in part based on the results of the first assay.
[0122]
[0127] Alternatively, multiple assays can be used to simultaneously process a cell-free biological sample of a subject.For example, a first assay can be used to process a first cell-free biological sample obtained from or derived from a subject to generate a first data set indicating a liver disease state, and a second assay different from the first assay can be used to process a second cell-free biological sample obtained from or derived from a subject to generate a second data set indicating a liver disease state.Then, any or all of the first data set and the second data set can be analyzed to evaluate the liver disease state of the subject.For example, a single diagnostic index or diagnostic score can be generated based on the combination of the first data set and the second data set.As another example, separate diagnostic indexes or diagnostic scores can be generated based on the first data set and the second data set.
[0123]
[0128] A cell-free biological sample can be processed using a metabolomics assay. For example, a metabolomics assay can be used to identify quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of multiple liver disease-related metabolites in a cell-free biological sample of a subject. A metabolomics assay can be configured to process a cell-free biological sample of a subject, such as a blood sample (or a derivative thereof). Quantitative measures (e.g., indicating the presence, absence, or relative amount) of liver disease-related metabolites in the cell-free biological sample can indicate one or more liver diseases. Metabolites in the cell-free biological sample can be produced as a result of one or more metabolic pathways corresponding to liver disease-related genes (e.g., as end products or by-products). Assaying one or more metabolites of the cell-free biological sample can include isolating or extracting the metabolites from the cell-free biological sample. A metabolomics assay can be used to generate a dataset indicating quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of multiple liver disease-related metabolites in the cell-free biological sample of a subject.
[0124]
[0129] Metabolomic assays can analyze a variety of metabolites in cell-free biological samples, such as small molecules, lipids, amino acids, peptides, nucleotides, hormones and other signaling molecules, cytokines, minerals and elements, polyphenols, fatty acids, dicarboxylic acids, alcohols and polyols, alkanes and alkenes, ketoacids, glycolipids, carbohydrates, hydroxyacids, purines, prostanoids, catecholamines, acylphosphates, phospholipids, cyclic amines, aminoketones, nucleosides, glycerolipids, aromatic acids, retinoids, aminoalcohols, pterins, steroids, carnitine, leukotrienes, indoles, porphyrins, sugar phosphates, Coenzyme A derivatives, glucuronides, ketones, sugar phosphates, inorganic ions and gases, sphingolipids, bile acids, alcohol phosphates, amino acid phosphates, aldehydes, quinones, pyrimidines, pyridoxal, tricarboxylic acids, acylglycines, cobalamin derivatives, lipoamides, biotin, and polyamines.
[0125]
[0130] The metabolomic assay may include, for example, one or more of mass spectrometry (MS), targeted MS, gas chromatography (GC), high performance liquid chromatography (HPLC), capillary electrophoresis (CE), nuclear magnetic resonance (NMR) spectroscopy, ion mobility spectroscopy, Raman spectroscopy, electrochemical assays, or immunoassays.
[0126]
[0131] Acellular biological samples can be processed using a methylation-specific assay. For example, a methylation-specific assay can be used to identify a quantitative measure of methylation (e.g., indicating presence, absence, or relative amount) of each of a plurality of liver disease-related genomic loci in acellular biological samples of a subject. Additionally or alternatively, a methylation-specific assay can also be used to identify a qualitative measure of methylation (e.g., a methylation pattern based on relative amount) of a plurality of liver disease-related genomic loci in acellular biological samples of a subject. A methylation-specific assay can be configured to process acellular biological samples of a subject, such as a blood sample (or a derivative thereof). A quantitative measure of methylation (e.g., indicating presence, absence, or relative amount) of liver disease-related genomic loci in acellular biological samples can indicate one or more liver disease states. A qualitative measure of methylation (e.g., a methylation pattern based on relative amount) of liver disease-related genomic loci in acellular biological samples can indicate one or more liver disease states. Methylation-specific assays can be used to generate datasets that indicate quantitative and / or qualitative measures of methylation at each of a plurality of liver disease-associated genomic loci in a cell-free biological sample of a subject.
[0127]
[0132] Methylation-specific assays may include, for example, one or more of methylation-aware sequencing (e.g., using bisulfite treatment or bisulfite-free treatment), enzymatic methylation sequencing, methylation-specific PCR (MSP), methylation-sensitive restriction enzyme (MSRE) digestion, pyrosequencing, methylation-sensitive single-strand conformation analysis (MS-SSCA), high-resolution melting analysis (HRM), methylation-sensitive single-nucleotide primer extension (MS-SnuPE), base-specific cleavage / MALDI-TOF, microarray-based methylation assays, methylation-specific PCR, targeted bisulfite sequencing, oxidative bisulfite sequencing, mass spectrometry-based bisulfite sequencing, or reduced representation bisulfite sequencing (RRBS).
[0128]
[0133] Bisulfite sequencing or processing involves treating DNA with bisulfite (e.g., sodium bisulfite), which converts cytosine residues to uracil residues, while leaving 5-methylcytosine residues unaffected. As a result, DNA treated with bisulfite may retain only methylated cytosines.
[0129]
[0134] Targeted bisulfite sequencing involves hybridization, which uses predesigned oligonucleotides to probe or target specific genomic regions of interest, such as CpG islands, gene promoters, and other important methylated regions (e.g., liver disease-related genomic loci). Targeted bisulfite sequencing can also involve amplification, which amplifies multiple bisulfite-converted DNA regions in a single reaction. Specific primers can be designed to capture the regions of interest and evaluate site-specific DNA methylation patterns.
[0130]
[0135] Pyrosequencing is a sequencing-by-synthesis method that quantitatively monitors real-time nucleotide incorporation via the enzymatic conversion of released pyrophosphate into a proportional light signal. Analysis of DNA methylation patterns by pyrosequencing combines a simple reaction protocol with reproducible and accurate measurement of the degree of methylation at several closely spaced CpGs with high quantitative resolution. After bisulfite treatment and PCR amplification, the degree of methylation at each CpG position in the sequence can be determined from the ratio of T to C. The purification and sequencing process can be repeated on the same template to analyze other CpGs in the same amplification product.
[0131]
[0136] RRBS is an efficient, high-throughput method for analyzing genome-wide methylation profiles at the single-nucleotide level. RRBS combines restriction enzyme and bisulfite sequencing to enrich for regions of the genome with high CpG content. RRBS can reduce the amount of nucleotides required for sequencing to as little as 1% of the genome. Fragments containing the reduced genome may still contain most of the promoters and regions such as repetitive sequences that are difficult to profile using conventional bisulfite sequencing approaches.
[0132]
[0137] In some cases, bisulfite conversion methods can lead to DNA damage, resulting in fragmentation, loss, and bias, thereby limiting their usefulness. Bisulfite-free methylation sequencing methods can minimize these drawbacks while enabling the conversion of methylated cytosines. For example, bisulfite-free methylation sequencing of cfDNA can be advantageous because cfDNA may exist at very low concentrations in plasma, which can be a limited resource in liquid biopsy applications.
[0133]
[0138] Enzymatic methylation sequencing offers a bisulfite-free approach that minimizes damage to sample DNA for methylation detection. Such enzymatic approaches can provide higher mapping efficiency, more uniform GC coverage, detection of more CpGs with fewer sequence reads, and more uniform dinucleotide distribution. Enzymatic methylation sequencing methods can involve treatment with methylcytosine dioxygenases, such as ten-eleven translocation (TET) enzymes; glucosyltransferases, such as β-glucosyltransferase (BGT); and / or cytidine deaminases, such as activation-induced (cytidine) deaminase (AID) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC). Methylcytosine dioxygenases can be used to convert 5mC and 5hmC residues to 5caC, protecting these methylated residues from deamination in downstream processing steps. Non-limiting examples of methylcytosine dioxygenases include TET1, TET2, TET3, and their catalytically active variants or fusion proteins. Glucosyltransferases can be used to add glucosyl groups to 5hmC and protect these methylated residues from downstream deamination. Cytidine deaminases can be used to deaminate 5hmC residues to uracil and 5hmC residues to thymine. Non-limiting examples of cytidine deaminases include AP0BEC3A and its catalytically active variants or fusion proteins. For bisulfite-free methylation sequencing, a combination of one or more enzymes can be used.
[0134]
[0139] TET-assisted pyridine borane sequencing (TAPS) uses the TET enzyme to oxidize 5mC and 5hmC residues to 5caC. Pyridine borane is then used to reduce 5caC to dihydrouracil, which is subsequently converted to thymine after amplification. TAPS can be performed in two other ways: TAPSβ and chemically assisted pyridine borane sequencing (CAPS). TAPSβ uses β-glucosyltransferase to label 5hmC with glucose, protecting it from oxidation and reduction reactions and allowing for specific detection of 5mC. In CAPS, potassium perruthenate acts as a chemical surrogate for TET to specifically oxidize 5hmC, thus allowing for direct detection of 5hmC.
[0135]
[0140] Methylation-specific PCR (MSP) is a qualitative DNA methylation assay. MSP has advantages such as ease of design and implementation, sensitivity to detect small amounts of methylated DNA, and the ability to rapidly screen a large number of samples without expensive laboratory equipment. This assay requires modification of genomic DNA with sodium bisulfite and two independent primer sets for PCR amplification (one pair designed to recognize the methylated version of the bisulfite-modified sequence and the other pair designed to recognize the unmethylated version of the bisulfite-modified sequence). Amplicons can be visualized using ethidium bromide staining after agarose gel electrophoresis. Amplicons of the expected size generated from either primer pair can indicate the presence of DNA with the respective methylation status in the original sample.
[0136]
[0141] In some embodiments, methylation-sensitive restriction enzyme (MSRE) digestion can be used to analyze the methylation status of cytosine residues in CpG sequences. These enzymes cannot cleave methylated cytosine residues, leaving methylated DNA fragments intact. Sample DNA obtained from or derived from a subject can be digested with one or more MSREs. For example, the liver disease-related genomic loci described herein can contain at least one specific MSRE recognition sequence (recognition site). The sample DNA is cleaved (digested) based on its methylation level; the more highly methylated the sample, the less digested by the enzyme. For example, if a DNA sample from a healthy subject has less methylation for a CpG in its recognition sequence than another DNA sample from a liver disease patient, the DNA can be cleaved more highly.
[0137]
[0142] For example, DNA molecules can be extracted from a biological sample. A first portion of the extracted DNA molecules can be subjected to CpG site fragmentation conditions, such as MSRE digestion, while a second portion of the extracted DNA molecules can be not subjected to such fragmentation conditions. Next, qPCR amplification of at least one biomarker locus, which is an internal control locus, can be performed (e.g., using qPCR primers). A cycle threshold (Ct) value can be obtained for each amplified region of a set of genomic regions (e.g., liver disease-related biomarkers) and normalized based on the internal control locus. The qPCR signal intensity can be calculated for the biomarker locus, where signal intensity = 2^[Ct, biomarker restriction locus - Ct, internal control locus]. Subsequently, a probability score can be calculated that reflects the correlation between the biomarker signal intensity in the subject and the "disease" reference and / or the correlation between the biomarker signal intensity in the subject and the "healthy" reference.
[0138]
[0143] In some embodiments, the control locus can be designed to exclude the MSRE restriction site. In some embodiments, a fixed proportion of control DNA is added to the sample DNA for all test subjects. In some embodiments, at least one pair of qPCR primers is designed for each target genomic region of the biomarker. For each patient, two qPCR reactions are performed independently on the same qPCR target: the first qPCR reaction is performed on a first portion of the sample DNA containing the MSRE-digested DNA template, and the second qPCR reaction is performed on a second portion of the sample DNA containing the undigested DNA template. The undigested template can be used to represent fully methylated DNA. After purification of the MSRE digest, the same amount of DNA can be used for the digested and undigested templates. The signal intensity of the qPCR reaction can be generated from the cycle threshold (Ct) value. The Ct value refers to the number of cycles required for the fluorescent signal to exceed a predetermined cycle threshold (e.g., the signal exceeds the background level). The Ct level can be inversely proportional to the amount of target nucleic acid in the sample (e.g., the lower the Ct level of a given sample, the greater the amount of target nucleic acid in the sample). For each locus in a given sample, the Ct difference (delta Ct) between the first qPCR reaction (performed on the digested DNA template) and the second qPCR reaction (performed on the undigested DNA template) can be calculated and used to indicate the DNA methylation level of the sample. Thus, the delta Ct value can represent the subject's DNA methylation level for the target region. For example, undigested DNA may have a low Ct value, while digested DNA from a normal individual may have a high Ct value, resulting in a large absolute delta Ct value. Otherwise, the delta Ct value from a subject with liver disease may be small (e.g., close to 0).
[0139]
[0144] Acellular biological samples can be processed using proteomic assays. For example, a proteomic assay can be used to identify quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of multiple liver disease-related proteins or polypeptides in a subject's acellular biological sample. A proteomic assay can be configured to process a subject's acellular biological sample, such as a blood sample (or a derivative thereof). Quantitative measures (e.g., indicating the presence, absence, or relative amount) of liver disease-related proteins or polypeptides in the acellular biological sample can indicate one or more liver disease conditions. Proteins or polypeptides in the acellular biological sample can be produced as a result of one or more biochemical pathways corresponding to liver disease-related genes (e.g., as end products or by-products). Assaying one or more proteins or polypeptides in the acellular biological sample can include isolating or extracting the proteins or polypeptides from the acellular biological sample. A proteomic assay can be used to generate a dataset indicating quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of multiple liver disease-related proteins or polypeptides in the subject's acellular biological sample.
[0140]
[0145] Proteomic assays can analyze various proteins or polypeptides in acellular biological samples, such as proteins produced under different cellular conditions (e.g., development, cell differentiation, or cell cycle). Proteomic assays can include, for example, one or more of antibody-based immunoassays, Edman degradation assays, mass spectrometry-based assays (e.g., matrix-assisted laser desorption / ionization (MALDI) and electrospray ionization (ESI)), top-down proteomic assays, bottom-up proteomic assays, mass spectrometry immunoassays (MSIA), stable isotope standard capture with anti-peptide antibodies (SISCAPA) assays, fluorescent two-dimensional differential gel electrophoresis (2-D DIGE) assays, quantitative proteomic assays, protein microarray assays, or reverse-phase protein microarray assays. Proteomic assays can detect post-translational modifications of proteins or polypeptides (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, and nitrosylation). Proteomic assays can identify or quantify one or more proteins or polypeptides from databases (e.g., Human Protein Atlas, PeptideAtlas, and UniProt).
[0141] kit
[0146] The present disclosure provides a kit for identifying or monitoring a subject's liver disease state.The kit can include a probe for identifying a quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of a plurality of liver disease-related genomic loci in a cell-free biological sample of the subject.The quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of a plurality of liver disease-related genomic loci in a cell-free biological sample can indicate one or more liver disease states.The probe can be selective for the sequence at each of a plurality of liver disease-related genomic loci in a cell-free biological sample.The kit can include instructions for using the probe to process the cell-free biological sample to generate a data set that indicates a quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of a plurality of liver disease-related genomic loci in a cell-free biological sample of the subject.
[0142]
[0147] The probes in the kit may be selective for sequences at multiple liver disease-associated genomic loci in a cell-free biological sample. The probes in the kit may be configured to selectively enrich nucleic acid molecules (e.g., RNA or DNA) corresponding to multiple liver disease-associated genomic loci. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with nucleic acid sequences from one or more of the multiple liver disease-associated genomic loci or genomic regions. The multiple liver disease-associated genomic loci or genomic regions may be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, or at least 45. The plurality of liver disease-associated genomic loci or genomic regions may comprise one or more members selected from the group consisting of the genes listed in Table 1.
[0143]
[0148] The instructions in the kit may include instructions for assaying the acellular biological sample using probes selective for sequences at multiple liver disease-related genomic loci in the acellular biological sample. These probes may be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity with nucleic acid sequences (e.g., RNA or DNA) derived from one or more of the multiple liver disease-related genomic loci. These nucleic acid molecules may be primers or enrichment sequences. The instructions for assaying the acellular biological sample may include instructions for performing array hybridization, PCR, or nucleic acid sequencing to process the acellular biological sample and generate a data set that indicates a quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple liver disease-related genomic loci in the acellular biological sample. The quantitative measure (e.g., indicating presence, absence, or relative amount) of the sequence at each of the multiple liver disease-related genomic loci in the acellular biological sample may indicate one or more liver disease states.
[0144]
[0149] The instructions in the kit can include instructions for measuring and interpreting assay readings, which can be quantified at one or more of a plurality of liver disease-related genomic loci to generate a data set that shows a quantitative measure (e.g., indicating the presence, absence, or relative amount) of the sequence at each of a plurality of liver disease-related genomic loci in acellular biological samples.For example, by quantification of array hybridization or polymerase chain reaction (PCR) corresponding to a plurality of liver disease-related genomic loci, a data set can be generated that shows a quantitative measure (e.g., indicating the presence, absence, or relative amount) of the sequence at each of a plurality of liver disease-related genomic loci in acellular biological samples.The assay readings can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or their normalized values.
[0145]
[0150] The kit may include a metabolomics assay for identifying quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of a plurality of liver disease-related metabolites in a cell-free biological sample of a subject. The quantitative measures (e.g., indicating the presence, absence, or relative amount) of liver disease-related metabolites in the cell-free biological sample may indicate one or more liver disease states. The metabolites in the cell-free biological sample may be produced as a result of one or more metabolic pathways corresponding to liver disease-related genes (e.g., as end products or by-products). The kit may include instructions for isolating or extracting metabolites from the cell-free biological sample and / or for using a metabolomics assay to generate a dataset indicating quantitative measures (e.g., indicating the presence, absence, or relative amount) of each of a plurality of liver disease-related metabolites in the cell-free biological sample of the subject.
[0146] Machine learning models
[0151] After using one or more assays to process one or more acellular biological samples from a subject to generate one or more data sets that indicate liver disease or liver pathology, a trained algorithm can be used to process one or more of the data sets (for example, at each of multiple liver disease-related genomic loci) to determine liver disease status.For example, a trained algorithm can be used to determine a quantitative measure of the sequence at each of multiple liver disease-related genomic loci in the acellular biological sample. The trained algorithm may be configured to identify liver disease states with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or greater than 99% for at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least 500, or greater than about 500 independent samples.
[0147]
[0152] The trained algorithm may include a supervised machine learning algorithm. The trained algorithm may include a classification and regression tree (CART) algorithm. The supervised machine learning algorithm may include a classifier or a regression. The supervised machine learning algorithm may include, for example, a deep learning algorithm, a support vector machine (SVM), a neural network, a random forest, a linear regression, or a logistic regression. The trained algorithm may include an unsupervised machine learning algorithm.
[0148]
[0153] The trained algorithm may be configured to receive multiple input variables and generate one or more output values based on the multiple input variables.The multiple input variables may include one or more data sets that indicate liver disease status.For example, the input variables may include several sequences that correspond to or align with each of multiple liver disease-related genomic loci.The multiple input variables may also include clinical health data of the subject.
[0149]
[0154] The trained algorithm may include a classifier in which one or more output values each include one of a fixed number of possible values indicating the classification of the acellular biological sample by the classifier (e.g., a linear classifier, a logistic regression classifier, etc.). The trained algorithm may include a binary classifier in which one or more output values each include one of two values indicating the classification of the acellular biological sample by the classifier (e.g., {0, 1}, {positive, negative}, or {high risk, low risk}). The trained algorithm may also be another type of classifier in which one or more output values each include one of more than two values indicating the classification of the acellular biological sample by the classifier (e.g., {0, 1, 2}, {positive, negative, or indeterminate}, or {high risk, medium risk, or low risk}). The output values may include descriptive labels, numerical values, or a combination thereof. Some of the output values may include descriptive labels. Such descriptive labels can provide an identification or indication of the subject's liver disease or disorder status. Such descriptive labels may include, for example, positive, negative, high risk, medium risk, low risk, or indeterminate.Such descriptive labeling can provide identification of a treatment for a subject's liver disease state, which can include, for example, a therapeutic intervention suitable for treating the liver disease state (e.g., vitamin E supplementation, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise regimens, dietary regimens, bariatric surgery, GLP1 (glucagon-like peptide-1) receptor agonists, FGF (fibroblast growth factor) analogs, THR (thyroid hormone receptor) agonists, SCD-1 (stearoyl-coenzyme A desaturase 1) inhibitors, FAS (fatty acid synthase) inhibitors, FXR (farnesoid X receptor) agonists, The therapeutic intervention may include a therapeutic agent, such as an agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted gene modifier (e.g., including PNPLA3 or HSD17B13), a L0XL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or any combination thereof), a duration of the therapeutic intervention, and / or a dosage of the therapeutic intervention. Such descriptive labeling may provide identification of secondary clinical tests that may be appropriate to perform on the subject, which may include, for example, blood tests, liver biopsies, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, acellular biological cytology, or any combination thereof. For example, such descriptive labels can provide a prognosis of the subject's liver disease state. As another example, such descriptive labels can provide a relative assessment of the subject's liver disease state (e.g., presence or absence, stage, or subtype). Some descriptive labels can be mapped to numerical values, for example, by mapping "positive" to 1 and "negative" to 0.
[0150]
[0155] Some of the output values may include numerical values, such as binary, integer, or continuous values. Such binary output values may include, for example, {0, 1}, {positive, negative}, or {high risk, low risk}. Such integer output values may include, for example, {0, 1, 2}. Such continuous output values may include, for example, probability values of at least 0 and less than or equal to 1. Such continuous output values may include, for example, an unnormalized probability value of at least 0. Such continuous output values may indicate a prognosis for the subject's liver disease status. Some numerical values may be mapped to descriptive labels, for example, by mapping 1 to "positive" and 0 to "negative."
[0151]
[0156] Some of the output values can be assigned based on one or more cutoff values. For example, a binary classification of a sample can assign an output value of "positive" or 1 if the sample indicates that the subject has at least a 50% probability of having a liver disease condition. For example, a binary classification of a sample can assign an output value of "negative" or 0 if the sample indicates that the subject has less than a 50% probability of having a liver disease condition. In this case, a single cutoff value of 50% is used to classify the sample into one of two possible binary output values. Examples of single cutoff values may include about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, and about 99%.
[0152]
[0157] As another example, a sample classification can be assigned an output value of "positive" or 1 if the sample indicates that the probability that the subject has liver disease is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more. The classification of the sample can be assigned an output value of "positive" or 1 if the sample indicates that the probability that the subject has a liver disease condition is greater than about 50%, greater than about 55%, greater than about 60%, greater than about 65%, greater than about 70%, greater than about 75%, greater than about 80%, greater than about 85%, greater than about 90%, greater than about 91%, greater than about 92%, greater than about 93%, greater than about 94%, greater than about 95%, greater than about 96%, greater than about 97%, greater than about 98%, or greater than about 99%.
[0153]
[0158] The classification of the sample can be assigned an output value of "negative" or 0 if the sample indicates that the probability that the subject has liver disease is less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1%. The classification of the sample can be assigned an output value of "negative" or 0 if the sample indicates that the probability that the subject has a liver disease condition is about 50% or less, about 45% or less, about 40% or less, about 35% or less, about 30% or less, about 25% or less, about 20% or less, about 15% or less, about 10% or less, about 9% or less, about 8% or less, about 7% or less, about 6% or less, about 5% or less, about 4% or less, about 3% or less, about 2% or less, or about 1% or less.
[0154]
[0159] Sample classification can assign an output value of "indeterminate" or 2 if the sample cannot be classified as "positive," "negative," 1, or 0. In this case, a set of two cutoff values is used to classify the sample into one of three possible output values. Example sets of cutoff values include {1%, 99%}, {2%, 98%}, {5%, 95%}, {10%, 90%}, {15%, 85%}, {20%, 80%}, {25%, 75%}, {30%, 70%}, {35%, 65%}, {40%, 60%}, and {45%, 55%}. Similarly, a set of n cutoff values can be used to classify the sample into one of n+1 possible output values, where n is any positive integer.
[0155]
[0160] The trained algorithm may be trained using multiple independent samples. Each of the independent samples may include an acellular biological sample from a subject, an associated dataset (as described herein) obtained by assaying the acellular biological sample, and one or more known output values corresponding to the acellular biological sample (e.g., clinical diagnosis, prognosis, absence, or treatment effectiveness of the subject's liver disease state). The independent samples may include acellular biological samples and associated datasets and outputs obtained or derived from multiple different subjects. The independent samples may include acellular biological samples and associated datasets and outputs obtained from the same subject at multiple different time points (e.g., periodically, for example, weekly, biweekly, or monthly). The independent samples may be associated with the presence of a liver disease state (e.g., a training sample including acellular biological samples and associated datasets and outputs obtained or derived from multiple subjects known to have a liver disease state). The independent sample can be associated with the absence of a liver disease state (e.g., a training sample including a cell-free biological sample and associated datasets and output obtained or derived from multiple subjects who are known to have no previous diagnosis of a liver disease state or who have tested negative for a liver disease state).
[0156]
[0161] The trained algorithm can be trained using at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent samples. The independent samples can include acellular biological samples associated with the presence of a liver disease state and / or acellular biological samples associated with the absence of a liver disease state. The trained algorithm can be trained using about 500 or less, about 450 or less, about 400 or less, about 350 or less, about 300 or less, about 250 or less, about 200 or less, about 150 or less, about 100 or less, or about 50 or less independent samples associated with the presence of liver disease. In some embodiments, the acellular biological sample is independent of the sample used to train the trained algorithm.
[0157]
[0162] The trained algorithm can be trained using a first number of independent samples associated with the presence of liver disease and a second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease can be less than or equal to the second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease can be equal to the second number of independent samples associated with the absence of liver disease. The first number of independent samples associated with the presence of liver disease can be greater than the second number of independent samples associated with the absence of liver disease.
[0158]
[0163] The trained algorithm may identify liver disease by at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138%, at least about 139%, The algorithm may be configured to identify at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, or at least about 500 independent samples with an accuracy of at least about 98%, at least about 99%, or better. The accuracy of identifying a liver disease state by the trained algorithm can be calculated as the percentage of independent samples (e.g., subjects known to have a liver disease state or subjects with negative laboratory test results for a liver disease state) that are correctly identified or classified as having or not having a liver disease state.
[0159]
[0164] The trained algorithm may be configured to identify a liver disease state with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or greater. The PPV of identifying a liver disease state using a trained algorithm can be calculated as the percentage of acellular biological samples that are identified or classified as having a liver disease state that correspond to subjects that truly have the liver disease state.
[0160]
[0165] The trained algorithm may be configured to identify a liver disease state with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or greater. The NPV of identifying a liver disease state using the trained algorithm can be calculated as the percentage of cell-free biological samples that are identified or classified as not having a liver disease state, corresponding to subjects that truly do not have the liver disease state.
[0161]
[0166] The trained algorithm may identify liver disease status as at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 100%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138%, at least about 139%, at least about 140%, at least about 141%, at least about 142%, at least about 143%, at least about 144%, at least about 145%, at least about 146%, at least about 147%, at least about 148%, at least about 149%, at least about 15 The clinical sensitivity of identifying a liver disease state using a trained algorithm may be calculated as the percentage of independent samples associated with the presence of the liver disease state (e.g., subjects known to have the liver disease state) that are correctly identified or classified as having the liver disease state.
[0162]
[0167] The trained algorithm may assess liver disease status as at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, The algorithm may be configured to identify a liver disease condition with a clinical specificity of at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater. The clinical specificity of identifying a liver disease condition using the trained algorithm can be calculated as the percentage of independent samples associated with the absence of a liver disease condition (e.g., subjects with negative laboratory test results for a liver disease condition) that are correctly identified or classified as not having the liver disease condition.
[0163]
[0168] The trained algorithm may be configured to identify a liver disease state with an area under the curve (AUC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.81, at least about 0.82, at least about 0.83, at least about 0.84, at least about 0.85, at least about 0.86, at least about 0.87, at least about 0.88, at least about 0.89, at least about 0.90, at least about 0.91, at least about 0.92, at least about 0.93, at least about 0.94, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, at least about 0.99, or more. The AUC can be calculated as the integral of the receiver operating characteristic (ROC) curve, e.g., the area under the ROC curve (AUROC), associated with the trained algorithm in classifying acellular biological samples as having or not having a liver disease state.
[0164]
[0169] The trained algorithm can be adjusted or fine-tuned to improve one or more of the performance, accuracy, PPV, NPV, clinical sensitivity, clinical specificity or AUC of identifying liver disease state.The trained algorithm can be adjusted or fine-tuned by adjusting the parameters of the trained algorithm (for example, the set of cutoff values used to classify acellular biological samples, as described elsewhere herein, or the weights of neural network).The trained algorithm can be continuously adjusted or fine-tuned during the training process or after the training process is completed.
[0165]
[0170] After the trained algorithm is initially trained, a subset of inputs can be identified as the most influential or most important to be included in order to achieve high-quality classification.For example, a subset of multiple liver disease-related genomic loci can be identified as the most influential or most important to be included in order to achieve high-quality classification or identification of liver disease (or liver disease subtypes).To achieve high-quality classification or identification of liver disease (or liver disease subtypes), multiple liver disease-related genomic loci or a subset thereof can be ranked based on classification metrics that indicate the influence or importance of each genomic locus.The use of such metrics can, in some cases, significantly reduce the number of input variables (e.g., predictor variables) that can be used to train the trained algorithm to a desired performance level (e.g., based on a desired minimum accuracy, PPV, NPV, clinical sensitivity, clinical specificity, AUC, positive likelihood ratio, negative likelihood ratio, or a combination thereof). For example, if training a trained algorithm using a plurality of input variables, including tens or hundreds, in the trained algorithm results in greater than 99% classification accuracy, instead, training the trained algorithm using only a selected subset of the most influential or most significant input variables of the plurality, such as about 5 or less, about 10 or less, about 15 or less, about 20 or less, about 25 or less, about 30 or less, about 35 or less, about 40 or less, about 45 or less, about 50 or less, or about 100 or less, may result in a reduced, but still acceptable, classification accuracy (e.g., at least 99%). At least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% can be obtained.The subset can be selected by rank-ordering the plurality of input variables overall and selecting a predetermined number (e.g., about 5 or less, about 10 or less, about 15 or less, about 20 or less, about 25 or less, about 30 or less, about 35 or less, about 40 or less, about 45 or less, about 50 or less, or about 100 or less) of input variables having the best classification metrics.
[0166]
[0171] The accuracy of the trained algorithm may be context-dependent. In some cases, the accuracy may be based on training samples from the general population. In other cases, the accuracy may be based on training samples from a high-risk population, such as a population suspected of having liver disease. To interpret the test performance of a trained algorithm, several factors may be considered: 1) the prevalence of the disease or condition, for example, how many people in the target population have the disease; and 2) whether the test is for diagnosing the disease, i.e., whether it is a positive (rule-in) test, or whether the test is for confirming that the subject does not have the disease, i.e., whether it is a negative (rule-out) test.
[0167]
[0172] On the other hand, metrics such as pre-test / post-test probability, Bayes factor, likelihood ratio, or information gain may be context-independent. These metrics measure the amount of new information provided by a test. For example, the pre-test / post-test probability ratio can be calculated by dividing the probability of a subject in a target population having a certain condition by the probability of a subject in the target population having a given test result having that condition. As an example, approximately 5% of the US population has NASH; therefore, the pre-test probability of NASH in the US population is 5%. If 50% of subjects detected by the test actually have NASH, the post-test probability is 50% and the pre-test / post-test ratio is 10. As another example, if approximately 40% of subjects in a high-risk population have NASH and a hypothetical test is performed on this high-risk population, 50% of those detected by the test truly have NASH, and the pre-test / post-test ratio is 1.25.
[0168] Identifying or monitoring liver disease status
[0173] After using trained algorithm to process data set, can identify or monitor the liver disease state in subject.Identification can be based at least in part on the quantitative measure of the sequence reads of data set in the panel of liver disease-related genomic loci (for example, the quantitative measure of DNA or RNA transcripts in liver disease-related genomic loci), the proteomic data comprise the quantitative measure of the protein of data set in the panel of liver disease-related proteins, and / or the metabolomic data comprise the quantitative measure of the panel of liver disease-related metabolites.
[0169]
[0174] The liver disease state can be identified in the subject with an accuracy of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more.The accuracy of identifying the liver disease state by the trained algorithm can be calculated as the percentage of independent samples (for example, the subjects who are known to have liver disease state or the subjects whose clinical test results for liver disease state are negative) that are correctly identified or classified as having or not having liver disease state.
[0170]
[0175] A liver disease state can be identified in a subject with a positive predictive value (PPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more. The PPV of identifying a liver disease state using a trained algorithm can be calculated as the percentage of acellular biological samples that are identified or classified as having a liver disease state that correspond to subjects that truly have the liver disease state.
[0171]
[0176] A liver disease state can be identified in a subject with a negative predictive value (NPV) of at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more. The NPV of identifying a liver disease state using the trained algorithm can be calculated as the percentage of cell-free biological samples that are identified or classified as not having a liver disease state, corresponding to subjects that truly do not have the liver disease state.
[0172]
[0177] The liver disease state is assessed in at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about The clinical sensitivity of identifying a liver disease state using a trained algorithm can be calculated as the percentage of independent samples associated with the presence of a liver disease state (e.g., subjects known to have a liver disease state) that are correctly identified or classified as having the liver disease state.
[0173]
[0178] The liver disease state is assessed in at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about The liver disease state can be identified with a clinical specificity of about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999%, or greater. The clinical specificity of identifying a liver disease state using a trained algorithm can be calculated as the percentage of independent samples associated with the absence of a liver disease state (e.g., subjects with negative laboratory test results for a liver disease state) that are correctly identified or classified as not having a liver disease state.
[0174]
[0179] Likelihood ratio can be used to evaluate the performance of diagnostic test.Liver disease state can be identified or excluded in subject based on likelihood ratio, for example, positive likelihood ratio or negative likelihood ratio.Likelihood ratio can be independent of the disease prevalence in training population, and therefore can be more representative of the disease prevalence in target population.Because likelihood ratio is independent of disease prevalence, likelihood ratio can be more directly related to the performance of given diagnostic test.
[0175]
[0180] The positive likelihood ratio can be calculated as sensitivity / (1-specificity). The liver disease state can be assessed as at least about 1, at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 26, at least about 27, at least about 28, at least about 29, at least about 30, at least about 31, at least about 32, at least about 33, at least about 34, at least about 35, at least about 36, at least about 37, at least about 38, at least about 39, at least about 40, at least about 41, at least about 42, at least about 43, at least about 44, at least about 45, at least about 46, at least about 47, at least about 48, at least about 49, at least about 50, at least about 51, at least about 52, at least about 53, at least about 54, at least about 55, at least about 56, at least about 57, at least about 58, at least about 59, at least about 60, at least about In some embodiments, the HIV-1 positive gene can be identified in a subject with a positive likelihood ratio of at least about 6, at least about 17, at least about 18, at least about 19, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, or at least about 1000.
[0176]
[0181] The negative likelihood ratio can be calculated as (1-sensitivity) / specificity. Liver disease status is categorized as up to about 1, up to about 0.99, up to about 0.95, up to about 0.9, up to about 0.8, up to about 0.7, up to about 0.75, up to about 0.6, up to about 0.5, up to about 0.4, up to about 0.3, up to about 0.25, up to about 0.2, up to about 0.1, up to about 0.09, up to about 0.08, up to about 0.07, up to about 0.06, up to about It can be ruled out in subjects having a negative likelihood ratio of about 0.05, at most about 0.04, at most about 0.03, at most about 0.02, at most about 0.01, at most about 0.009, at most about 0.008, at most about 0.007, at most about 0.006, at most about 0.005, at most about 0.004, at most about 0.003, at most about 0.002, or at most about 0.001.
[0177]
[0182] In one aspect, the disclosure provides a method for determining that a subject is at risk for developing liver disease, comprising assaying an acellular biological sample from the subject to generate a dataset indicative of risk for developing liver disease with at least 80% specificity, and using a trained algorithm trained on a sample independent of the acellular biological sample to determine that the subject is at risk for developing liver disease with at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75% specificity. , at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more accurate.
[0178]
[0183] After liver disease is identified in a subject, liver disease subtypes (e.g., selected from multiple liver disease subtypes) can be further identified. Liver disease subtypes can be determined at least in part based on quantitative measures of sequence reads in a dataset of a panel of liver disease-related genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-related genomic loci), proteomic data including quantitative measures of proteins in a dataset of a panel of liver disease-related proteins, and / or metabolomic data including quantitative measures of a panel of liver disease-related metabolites. For example, a subject can be identified as being at risk for liver disease subtypes (e.g., selected from multiple liver disease subtypes). After identifying a subject at risk for liver disease subtypes, a clinical intervention for the subject can be selected at least in part based on the liver disease subtype that the subject is identified as being at risk for. In some embodiments, the clinical intervention is selected from multiple clinical interventions (e.g., clinically indicated for different liver disease subtypes).
[0179]
[0184] In some embodiments, the trained algorithm can determine that a subject has at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more risk of liver disease.
[0180]
[0185] The trained algorithm determines that the subject is at risk for liver disease by at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138%, at least can be determined with an accuracy of about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, at least about 99.99%, at least about 99.999% or greater.
[0181]
[0186] Upon identifying a subject as having a liver disease state, the subject can optionally be administered a therapeutic intervention (e.g., prescribed an appropriate course of treatment to treat the subject's liver disease state). Therapeutic intervention can include prescribing an effective dose of a drug, further testing or evaluation of the liver disease state, further monitoring of the liver disease state, exercise therapy, diet therapy, bariatric surgery, or a combination thereof. Therapeutic interventions include vitamin E supplementation, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise therapy, diet therapy, bariatric surgery, GLP1 (glucagon-like peptide-1) receptor agonists, FGF (fibroblast growth factor) analogs, THR (thyroid hormone receptor) agonists, SCD-1 (stearoyl-coenzyme A desaturase 1) inhibitors, FAS (fatty acid synthase) inhibitors, FXR (farnesoid X receptor) agonists, ACC (acetyl-CoA capsaicin) inhibitors, and combinations thereof. The therapeutic intervention may include a cyclophilin inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or a combination thereof. If the subject is currently undergoing treatment for a liver disease state with a treatment course, the therapeutic intervention may include a subsequent different treatment course (e.g., to enhance the effectiveness of the treatment because the current treatment course is ineffective).
[0182]
[0187] Therapeutic intervention can include recommending a secondary laboratory test to the subject to confirm the diagnosis of the liver disease state, which can include a blood test, a liver biopsy, an imaging test, a computed tomography (CT) scan, a magnetic resonance imaging (MRI) scan, an ultrasound scan, a chest x-ray, a positron emission tomography (PET) scan, a PET-CT scan, a cell-free biological cytology, or any combination thereof.
[0183]
[0188] After identifying that a subject has liver disease state, the subject can be determined to be ineligible for liver disease transplantation in some cases.After identifying that a subject does not have liver disease state, the subject can be determined to be eligible for liver disease transplantation in some cases.If a subject is not identified as having liver disease or as having an increased risk of developing liver disease, the subject can be determined to be eligible as a liver transplant donor.If a subject is identified as having liver disease or as having an increased risk of developing liver disease, the subject can be determined to be eligible as a liver transplant recipient.
[0184]
[0189] Various therapeutic interventions and clinical tests for liver disease can be used in combination with the method described herein.For example, after determining that a subject has liver disease, therapeutic interventions can be administered to the subject.As another example, if it is determined that a subject has an increased risk of liver disease, preventive interventions can be administered to the subject.Examples of interventions and clinical tests for liver disease are described in Vittal et al., Clin Liver Dis., August 2019;23(3):417-432; Marroni et al., World J Gastroenterol., July 14, 2018;24(26):2785-2805; Leoni et al., World J Gastroenterol., August 14, 2018;24(30):3361-3373; and Sumida et al., J Gastroenterol., March 2018;53(3):362-376, each of which is incorporated herein by reference in its entirety.
[0185]
[0190] Quantitative measures of sequence reads of a dataset of a panel of liver disease-related genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-related genomic loci), proteomic data including quantitative measures of proteins of a dataset of a panel of liver disease-related proteins, and / or metabolomic data including quantitative measures of a panel of liver disease-related metabolites can be evaluated over a period of time to monitor patients (e.g., subjects with liver disease or subjects receiving treatment for liver disease).In such cases, the quantitative measures of the patient's dataset may change during the course of treatment.For example, the quantitative measures of a dataset of a patient whose risk of liver disease is reduced due to effective treatment may shift toward the profile or distribution of a healthy subject (e.g., a subject who does not have liver disease or liver pathology).Conversely, the quantitative measures of a dataset of a patient whose risk of liver disease is increased due to ineffective treatment may shift toward the profile or distribution of subjects who are at a higher risk of liver disease or have more advanced liver disease.
[0186]
[0191] The liver disease of object can be monitored by monitoring the treatment course for treating the liver disease of object.Monitoring can include evaluating the liver disease state of object at two or more time points.Evaluation can be based on at least the quantitative measure of the sequence reads of the data set of the panel of liver disease-related genomic loci (for example, the quantitative measure of the DNA or RNA transcripts of the liver disease-related genomic loci), the proteomic data, comprising the quantitative measure of the protein of the data set of the panel of liver disease-related proteins, and / or the metabolome data, comprising the quantitative measure of the panel of liver disease-related metabolites, determined at each of two or more time points.
[0187]
[0192] In some embodiments, differences in quantitative measures of sequence reads of a dataset at a panel of liver disease-associated genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-associated genomic loci), proteomic data comprising quantitative measures of proteins of a dataset at a panel of liver disease-associated proteins, and / or metabolomic data comprising quantitative measures of a panel of liver disease-associated metabolites determined between two or more time points may indicate one or more clinical indications, for example, (i) a diagnosis of liver disease in a subject, (ii) a prognosis of liver disease in a subject, (iii) an increased risk of liver disease in a subject, (iv) a decreased risk of liver disease in a subject, (v) the effectiveness of a treatment course for treating liver disease in a subject, and (vi) the ineffectiveness of a treatment course for treating liver disease in a subject.
[0188]
[0193] In some embodiments, the difference between the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci (e.g., a quantitative measure of DNA or RNA transcripts at liver disease-related genomic loci), the proteomic data including a quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or the metabolomic data including a quantitative measure of a panel of liver disease-related metabolites determined between two or more time points can indicate a diagnosis of liver disease in the subject.For example, if liver disease is not detected in the subject at an earlier time point, but is detected in the subject at a later time point, the difference indicates a diagnosis of liver disease in the subject.Based on this indication of the diagnosis of liver disease in the subject, a clinical action or decision can be made, such as prescribing a new therapeutic intervention for the subject.The clinical action or decision can include recommending a secondary laboratory test for the subject to confirm the diagnosis of liver disease state. This secondary laboratory testing may include blood tests, liver biopsy, imaging tests, computed tomography (CT) scan, magnetic resonance imaging (MRI) scan, ultrasound scan, chest x-ray, positron emission tomography (PET) scan, PET-CT scan, acellular biological cytology, or any combination thereof.
[0189]
[0194] In some embodiments, differences in quantitative measures of sequence reads of a dataset at a panel of liver disease-associated genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-associated genomic loci), proteomic data including quantitative measures of proteins of a dataset at a panel of liver disease-associated proteins, and / or metabolomic data including quantitative measures of a panel of liver disease-associated metabolites determined between two or more time points may indicate prognosis of a subject's liver disease status.
[0190]
[0195] In some embodiments, the difference between two or more time points determined in the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci (e.g., quantitative measure of DNA or RNA transcripts at liver disease-related genomic loci), proteomic data comprising quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or metabolomic data comprising quantitative measure of a panel of liver disease-related metabolites can indicate a subject with an increased risk of liver disease state.For example, if a liver disease state is detected in a subject at both an earlier time point and a later time point, and the difference is positive (e.g., the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci or RNA transcripts, proteomic data comprising quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or metabolomic data comprising quantitative measure of a panel of liver disease-related metabolites increases from an earlier time point to a later time point), the difference can indicate a subject with an increased risk of liver disease state. Based on this indication of increased risk of liver disease state, clinical action or decision can be made, for example, prescribe new therapeutic intervention to the subject or switch therapeutic intervention (for example, terminate current treatment and prescribe new treatment).Clinical action or decision can include recommending secondary laboratory tests to the subject to confirm increased risk of liver disease state.This secondary laboratory test can include blood test, liver biopsy, imaging test, computed tomography (CT) scan, magnetic resonance imaging (MRI) scan, ultrasound scan, chest X-ray, positron emission tomography (PET) scan, PET-CT scan, acellular biological cytology, or any combination thereof.
[0191]
[0196] In some embodiments, the difference between the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci (for example, the quantitative measure of DNA or RNA transcripts at liver disease-related genomic loci), the proteomic data comprising the quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or the metabolomic data comprising the quantitative measure of a panel of liver disease-related metabolites determined between two or more time points can indicate a subject with a reduced risk of liver disease state.For example, if liver disease is detected in a subject at both an earlier time point and a later time point, and the difference is negative (for example, the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci or RNA transcripts, the proteomic data comprising the quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or the metabolomic data comprising the quantitative measure of a panel of liver disease-related metabolites decreases from an earlier time point to a later time point), the difference can indicate a subject with a reduced risk of liver disease state. Based on this indication that the risk of liver disease state is reduced, a clinical action or decision can be taken on the subject (e.g., continuing or terminating current therapeutic intervention).Clinical action or decision can include recommending a secondary laboratory test to the subject to confirm the reduced risk of liver disease state.This secondary laboratory test can include blood tests, liver biopsy, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest X-rays, positron emission tomography (PET) scans, PET-CT scans, acellular biological cytology, or any combination thereof.
[0192]
[0197] In some embodiments, the difference between the quantitative measure of sequence reads of a dataset of a panel of liver disease-related genomic loci (for example, the quantitative measure of DNA or RNA transcripts at liver disease-related genomic loci), the proteomic data including the quantitative measure of proteins of a dataset of a panel of liver disease-related proteins, and / or the metabolomic data including the quantitative measure of a panel of liver disease-related metabolites determined between two or more time points can indicate the effectiveness of a treatment course for treating a subject's liver disease condition.For example, if liver disease is detected in a subject at an earlier time point but not in the subject at a later time point, this difference can indicate the effectiveness of a treatment course for treating the subject's liver disease.Based on this indication of the effectiveness of a treatment course for treating a subject's liver disease, clinical action or decision can be made, for example, to continue or terminate the current therapeutic intervention for the subject.Clinical action or decision can include recommending a secondary clinical test for the subject to confirm the effectiveness of a treatment course for treating a liver disease condition. The secondary laboratory tests may include blood tests, liver biopsies, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest x-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.
[0193]
[0198] In some embodiments, differences in quantitative measures of sequence reads of a dataset at a panel of liver disease-associated genomic loci (e.g., quantitative measures of DNA or RNA transcripts at liver disease-associated genomic loci), proteomic data comprising quantitative measures of proteins of a dataset at a panel of liver disease-associated proteins, and / or metabolomic data comprising quantitative measures of a panel of liver disease-associated metabolites determined between two or more time points may indicate that a treatment course for treating the liver disease condition of the subject is ineffective. For example, if liver disease status is detected in a subject at both an earlier time point and a later time point, and the difference is negative or zero (for example, if the quantitative measure of sequence reads in a dataset of a panel of liver disease-related genomic loci or RNA transcripts, if the proteomic data comprising the quantitative measure of proteins in a dataset of a panel of liver disease-related proteins, and / or the metabolomic data comprising the quantitative measure of a panel of liver disease-related metabolites increases from an earlier time point to a later time point or remains at a constant level), and if effective treatment is shown at an earlier time point, this difference can indicate that the treatment course for treating the subject's liver disease is ineffective.Based on this indication that the treatment course for treating the subject's liver disease is ineffective, clinical action or decision can be made, for example, the subject can be terminated from the current therapeutic intervention and / or switched to a different new therapeutic intervention (e.g., prescribed).Clinical action or decision can include recommending a secondary clinical test for the subject to confirm that the treatment course for treating liver disease is ineffective. The secondary laboratory tests may include blood tests, liver biopsies, imaging tests, computed tomography (CT) scans, magnetic resonance imaging (MRI) scans, ultrasound scans, chest x-rays, positron emission tomography (PET) scans, PET-CT scans, cell-free biological cytology, or any combination thereof.
[0194]
[0199] In another aspect, the present disclosure provides a computer-implemented method for predicting a subject's risk of liver disease, the method including: (a) receiving clinical health data of the subject, wherein the clinical health data includes a plurality of quantitative or categorical measures of the subject; (b) processing the subject's clinical health data using a trained algorithm to determine a risk score indicative of the subject's risk of liver disease; and (c) electronically outputting a report showing the risk score indicative of the subject's risk of liver disease.
[0195]
[0200] In some embodiments, for example, the clinical health data includes one or more quantitative measures of a subject, such as age, weight, height, body mass index (BMI), blood pressure, heart rate, and glucose level. As another example, the clinical health data may include one or more categorical measures, such as race, ethnicity, medical history, history of medication or other clinical treatment, smoking history, alcohol consumption history, level of daily activity or fitness, genetic test results, blood test results, and imaging results.
[0196]
[0201] In some embodiments, the computer-implemented method for predicting the risk of liver disease of a subject is carried out using a computer or mobile device application.For example, a subject can use a computer or mobile device application to input their own clinical health data, including quantitative and / or categorical measures.The computer or mobile device application can then use a trained algorithm to process the clinical health data and determine a risk score that indicates the subject's risk of liver disease.The computer or mobile device application can then display a report that shows the risk score that indicates the subject's risk of liver disease.
[0197]
[0202] In some embodiments, the risk score that indicates the risk of liver disease of the subject can be refined by carrying out one or more subsequent clinical tests on the subject.For example, the subject can be referred by a doctor for one or more subsequent clinical tests (for example, imaging test or blood test) based on the initial risk score.Then, the application of computer or mobile device can use the trained algorithm to process the results of one or more subsequent clinical tests to determine an updated risk score that indicates the risk of liver disease of the subject.
[0198]
[0203] In some embodiments, the risk score comprises the possibility that the subject has liver disease within a predetermined period of time.For example, the predetermined period of time can be about 1 hour, about 2 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 1.5 days, about 2 days, about 2.5 days, about 3 days, about 3.5 days, about 4 days, about 4.5 days, about 5 days, about 5.5 days, about 6 days, about 6.5 days, about 7 days, about It can be 8 days, about 9 days, about 10 days, about 12 days, about 14 days, about 3 weeks, about 4 weeks, about 5 weeks, about 6 weeks, about 7 weeks, about 8 weeks, about 9 weeks, about 10 weeks, about 11 weeks, about 12 weeks, about 5 months, about 6 months, about 7 months, about 8 months, about 9 months, about 10 months, about 11 months, about 1 year, about 2 years, about 3 years, about 4 years, about 5 years, or more than about 5 years.
[0199]
[0204] After the liver disease state is identified in the subject or the increased risk of liver disease is monitored, a report can be electronically outputted, which indicates (for example, identifies or gives an indication of) the liver disease of the subject.The subject may not have liver disease (for example, is asymptomatic for liver disease).The report can be displayed on the graphical user interface (GUI) of the user's electronic device.The user can be the subject, a caregiver, a doctor, a nurse, or another medical professional.
[0200]
[0205] The report may include one or more clinical indicators: (i) a diagnosis of liver disease in the subject; (ii) a prognosis of liver disease in the subject; (iii) an increased risk of liver disease in the subject; (iv) a decreased risk of liver disease in the subject; (v) the effectiveness of a treatment course for treating liver disease in the subject; and (vi) ineffectiveness of a treatment course for treating liver disease in the subject. The report may include one or more clinical actions or decisions made based on these one or more clinical indicators. Such clinical actions or decisions may be directed to therapeutic intervention, induction or inhibition of labor, or further clinical evaluation or testing of the subject's liver disease.
[0201]
[0206] For example, the clinical indication that the subject is diagnosed with liver disease can be accompanied by the clinical action of prescribing a new therapeutic intervention for the subject.As another example, the clinical indication that the subject's risk of liver disease is increasing can be accompanied by the clinical action of prescribing a new therapeutic intervention for the subject or switching therapeutic intervention (for example, terminating current treatment and prescribing a new treatment).As another example, the clinical indication that the subject's risk of liver disease is decreasing can be accompanied by the clinical action of continuing or terminating the current therapeutic intervention for the subject.As another example, the clinical indication that the treatment course for treating the subject's liver disease is effective can be accompanied by the clinical action of continuing or terminating the current therapeutic intervention for the subject.As another example, the clinical indication that the treatment course for treating the subject's liver disease is ineffective can be accompanied by the clinical action of terminating the current therapeutic intervention for the subject and / or switching to a different new therapeutic intervention (for example, prescribing).
[0202] Computer Systems
[0207] The present disclosure provides a computer system that is programmed to implement the method of the present disclosure. Figure 2 shows a computer system 201 that is programmed or otherwise configured to, for example, (i) train and test a trained algorithm, (ii) use the trained algorithm to process data to determine a subject's liver disease state, (iii) determine a quantitative measure that indicates the subject's liver disease state, (iv) identify or monitor the subject's liver disease state, and (v) electronically output a report that indicates the subject's liver disease state.
[0203]
[0208] The computer system 201 can control various aspects of the analysis, calculation, and generation of the present disclosure, such as (i) training and testing the trained algorithm, (ii) processing data using the trained algorithm to determine the subject's liver disease status, (iii) determining a quantitative measure indicative of the subject's liver disease status, (iv) identifying or monitoring the subject's liver disease status, and (v) electronically outputting a report indicative of the subject's liver disease status. The computer system 201 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0204]
[0209] Computer system 201 includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") 205, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 201 also includes memory or memory locations 210 (e.g., random access memory, read-only memory, flash memory), electronic storage 215 (e.g., a hard disk), a communication interface 220 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 225, such as cache, other memory, data storage, and / or an electronic display adapter. Memory 210, storage 215, interface 220, and peripheral devices 225 communicate with CPU 205 via a communication bus (solid lines), such as a motherboard. Storage 215 may be a data storage device (or data repository) for storing data. Computer system 201 may be operably coupled to a computer network ("network") 230 with the aid of communication interface 220. Network 230 may be the Internet, an internet and / or extranet, or an intranet and / or extranet in communication with the Internet.
[0205]
[0210] Network 230 may, in some cases, be a telecommunications and / or data network. Network 230 may include one or more computer servers that enable distributed computing, such as cloud computing. For example, one or more computer servers may enable cloud computing to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as (i) training and testing trained algorithms, (ii) processing data using the trained algorithms to determine a subject's liver disease status, (iii) determining quantitative measures indicative of a subject's liver disease status, (iv) identifying or monitoring a subject's liver disease status, and (v) electronically outputting a report indicative of a subject's liver disease status, via network 230 (the "cloud"). Such cloud computing may be provided by cloud computing platforms such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. Network 230 may, in some cases, implement a peer-to-peer network with the aid of computer system 201, which may enable devices coupled to computer system 201 to act as clients or servers.
[0206]
[0211] CPU 205 may include one or more computer processors and / or one or more graphics processing units (GPUs). CPU 205 may execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, for example, memory 210. The instructions may be directed to CPU 205, which may then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by CPU 205 may include fetch, decode, execute, and writeback.
[0207]
[0212] The CPU 205 may be part of a circuit, for example, an integrated circuit. One or more other components of the system 201 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0208]
[0213] The storage device 215 can store files, e.g., drivers, libraries, and saved programs. The storage device 215 can store user data, e.g., user preferences and user programs. The computer system 201 can optionally include one or more additional data storage devices external to the computer system 201, for example, located on a remote server that communicates with the computer system 201 via an intranet or the Internet.
[0209]
[0214] Computer system 201 can communicate with one or more remote computer systems via network 230. For example, computer system 201 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 201 via network 230.
[0210]
[0215] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored in electronic storage locations of computer system 201, such as memory 210 or electronic storage 215. The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by processor 205. In some cases, the code may be retrieved from storage 215 and stored in memory 210 for easy access by processor 205. In some circumstances, electronic storage 215 may be omitted, and machine-executable instructions are stored in memory 210.
[0211]
[0216] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled at run time. The code may be supplied in a programming language that may be selected to allow the code to be executed in a pre-compiled or compiled manner.
[0212]
[0217] Aspects of the systems and methods provided herein, e.g., computer system 201, can be embodied in programming. Various aspects of the present technology can be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, e.g., memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage"-type media can include any or all of the tangible memory of a computer, processor, etc., or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. The software, in whole or in part, can be communicated from time to time via the Internet or various other telecommunications networks. Such communication can, for example, enable loading of the software from one computer or processor to another, e.g., from a management server or host computer to an application server's computer platform. Thus, other types of media onto which software elements may be stored include optical, electrical, and electromagnetic waves, such as those used over physical interfaces between local devices, over wired and optical landline networks, and via various air links. Physical elements that carry such waves, e.g., wired or wireless links, optical links, etc., may also be considered media onto which software may be stored. As used herein, unless limited to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0213]
[0218] Thus, machine-readable media, e.g., computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, some of the storage devices in any computer, such as those used to implement the databases shown in the figures, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0214]
[0219] The computer system 201 can include or be in communication with an electronic display 235 that includes a user interface (UI) 240 for providing, for example, (i) a visual display showing the training and testing of a trained algorithm, (ii) a visual display of data indicative of a subject's liver disease status, (iii) a quantitative measure of a subject's liver disease status, (iv) an identification of a subject having a liver disease status, or (v) an electronic report indicative of a subject's liver disease status. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0215]
[0220] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 205. The algorithms can, for example, (i) train and test a trained algorithm, (ii) process data using the trained algorithm to determine the liver disease status of the subject, (iii) determine a quantitative measure indicative of the liver disease status of the subject, (iv) identify or monitor the liver disease status of the subject, and (v) electronically output a report indicative of the liver disease status of the subject.
[0216] cfDNA methylation
[0221] In some embodiments, the cfDNA methylation data (observation) obtained from biological samples comprises a set of sequenced DNA fragments that are subjected to conversion conditions, such that unmethylated cytosine sites are converted to thymine, to provide the methylation status of cytosine sites in the DNA fragments.Each DNA fragment may consist of several base pair reads, indicating whether the methylation site is methylated or unmethylated.Provided herein are machine learning models and systems that are useful for predicting relevant outcomes from such cfDNA methylation data.Non-limiting examples of such outcomes include: (i) the presence or absence of disease; (ii) the type or subtype of disease; (iii) the type, dosage, or combination of treatments for disease treatment; (iv) the predicted response of the subject to disease treatment; (v) the risk that the subject will develop an advanced form of disease; and (vi) the outcome (prognosis) of the subject.
[0217]
[0222] The dataset can include cfDNA methylation data from one or more subjects, at least some of which have one or more markers described herein. The challenge in training an ML model is to generate a model that can use the dataset to predict outcomes from new cfDNA methylation data that has not been previously trained. In cfDNA methylation data, each fragment can be assigned to a location in the genome. ML models can represent data through data representation, featurization, or feature engineering. For example, for large datasets with millions of data points, data representations can be generated in a purely data-driven manner using deep neural networks. Such networks can be designed to build complex underlying datasets purely from data without strong assumptions. However, when the sample size is small to moderate, such as in the case of cfDNA methylation data, it can be difficult to predict outcomes using a purely data-driven representation without any assumptions. Provided herein is a method for representing data with high dimensionality and small sample sizes in an ML model to predict outcomes with high accuracy and sensitivity. The methods described herein include providing a compact probability distribution of a plurality of fragments, using the compact probability distribution in intermediate training to provide a trained model, and using the trained model to characterize the plurality of fragments.
[0218]
[0223] cfDNA methylation data consists of a large number of fragments; however, these fragments can originate anywhere in the genome, and samples can have different numbers of fragments. These heterogeneous and sparse data can also pose challenges to ML training methods. Furthermore, there are approximately 28 million methylation sites in the human genome, which is several orders of magnitude more than the largest clinical study that can be achieved using cfDNA methylation data. Training using data with input dimensions several orders of magnitude larger than the number of training data can pose challenges to ML training methods.
[0219]
[0224] DNA data, including cfDNA methylation data, can be generated using a sequencer, which can be an expensive and time-consuming process. ML training methods can be used to circumvent these drawbacks by utilizing data from various studies, regardless of acquisition method and data source. Such data sources include, but are not limited to: - combining data obtained from cfDNA and non-cfDNA, e.g., cfDNA, with data obtained from tissue samples; - combining data obtained from a different methylation assay, for example, bisulfite conversion, with data obtained from an enzymatic conversion assay; and - Combining data obtained from different sequencing methods, for example microarrays, with data obtained from next generation sequencing.
[0220]
[0225] Such flexibility may allow the use of pre-obtained data, for example publicly available data.
[0226] Considering that there are about 3.2 billion genome positions, of which about 28 million can be methylated, cfDNA methylation data can be very large.Each fragment in cfDNA can have about 150 base pairs on average.Therefore, for example, for a given sample, the cfDNA methylation data set with 30 times sequencing depth can require at least 48 gigabytes and 250 megabytes of storage capacity for base pairs and methylation status, respectively.A training procedure involving 500 samples can require several times of processing.There is a need for an ML training method that can handle such large data sets.
[0221]
[0227] Distributing training across a cluster of computers can help overcome these challenges. However, this approach can have various drawbacks. Because multiple computers need to communicate with each other during training, training can be extremely slow and time-consuming. For example, training a model using all fragments with 1,000 observations can require approximately 1,500 core hours and thousands of computers. Alternatively, the data can be split into separate regions, for example, across different genomes, and processed independently. However, this approach can hinder the ML model from learning the subtle interactions between different genomic regions.
[0222]
[0228] The present disclosure provides an ML method that alleviates the problems described herein. The disclosed method includes: providing a probability distribution based on cfDNA methylation data of a set of fragments from a biological sample; and training the probability distribution based on an ML model. Instead of training on a set of fragments, the disclosed method includes training on a probability distribution of the set of fragments. The probability distribution can represent the state of the sample, and the list of observed fragments can be extracted from such a probability distribution mediated by blood sampling and sequencing of the set of DNA fragments. Specifically, the disclosed method includes converting a set of input fragments into a probability distribution that is most likely to generate the input fragments.
[0223]
[0229] Representing data using a probability distribution has various advantages. The probability distribution may not be sparse but may have a predetermined complexity. The probability distribution may represent the likelihood of observing different methylation patterns. This probability distribution represents the state of the methylation pattern and is therefore less susceptible to variations in assays, sequencing methods, and other factors. Such a feature may be desirable because data in the public domain, such as sequencing data from the National Institutes of Health and other research institutions, is available. Furthermore, probability distributions are much smaller in size, which may make them much easier to use in training or distributed systems. As a result, building complex models may be more feasible. Building complex models can be very expensive and time-consuming when computation is expensive. Therefore, probability distributions can provide a simpler approach that makes training complex models feasible. In addition, a probability distribution representing a given sample can be calculated without the need or knowledge of other samples (e.g., training other samples). Therefore, this procedure can be easily distributed across computer clusters. This procedure allows for the freedom of no information leakage between samples and therefore no cross-validation or training and testing datasets are required. Such representations may also be suitable for building models that generate high-quality inferences.
[0224]
[0230] Cell-free DNA methylation data can be derived from a large number of cells throughout the body. Assuming that each cell has several features (Z), a cell can be represented by a mixture of those features. A sample can be represented as a proportion of multiple distinct cells, and therefore as a proportion of such hidden features. Therefore, the first task is to determine the best Z features that can estimate those features for a set of fragments from the cfDNA methylation dataset from a dataset of models from a set of probability distributions.
[0225]
[0231] Figure 3 shows a schematic diagram of an exemplary training dataset. From a mathematical perspective, the underlying dataset is a D-dimensional random variable (i.e., the number of methylation sites) that is observed in parts. Each observation (i.e., participant) consists of several fragments. A fragment corresponds to a set of values that corresponds to a portion of the D-dimensional space.
[0226]
[0232] As described herein, observations are made not as a set of fragments but as φ s It can be formulated as a distribution in D-dimensional space characterized by the parameters φ of the distribution (one for each observation). s is a statistic of the set of fragments. For large distribution classes, e.g., the exponential family, the parameters of the distribution (φ s ) can be expressed explicitly as their sufficient statistics. In other cases, in the general case, the parameters of the distribution can be expressed by statistics that are close enough. For these general cases, φ s can be calculated by maximizing the likelihood for a class of distributions using the following formula:
[0227]
number
[0233] Such probability distributions can be characterized in several ways. For example, a probability distribution representation of a sample can be expressed using a Markov model, in which the probability of observing a methylation state depends on its genomic location as well as the state of the previous methylation site. Such a model can be created by quantifying the number of states observed and the number of k-mers at each genomic location or methylation site, which can be determined using the following formula:
[0228]
number
[0229]
[0234] All data are expressed as parameters of a probability distribution (i.e., all φ for all observations). i Assuming that (Z is estimated), several techniques can be used to estimate the mentioned hidden Z features. One technique involves maximizing the likelihood using the following formula:
[0230]
number
[0231]
number
[0235] In the expectation step, we find the most likely q based on our current estimate of θ. i,z can be determined.
[0232]
number
[0236] In the maximization step, the most likely θ can be determined based on the current estimate of q.
[0233]
[0237] The output of the above method is a set of Z parameters (θ) that describe the hidden features of the dataset.
[0238] These estimates may not rely on distribution assumptions, eg, Gaussian or Bernoulli distributions.
[0234]
[0239] Because of the greatly reduced data size due to this particular representation of the data, most calculations can be processed on a general-purpose computer or easily distributed across multiple computers for faster execution times.
[0235]
[0240] The result of the first operation is a representative distribution corresponding to the unknown Z features. These features do not need to be known a priori or assigned by an expert.
[0241] Because this first operation can be used to infer a set of biological features, data can be incorporated and / or aggregated from a variety of sources, including cfDNA data, data from different assays (e.g., RNA data, proteomics data, metabolomics data, etc.), data with different sequencing depths, and / or data generated from different sequencing methods.
[0236]
[0242] Given a set of Z features (representative distributions), the set of fragments can be transformed into a fixed set of features in several ways. For example, observations can be represented as a histogram over location and the above features. A Z x D zero matrix can be used as a starting point. For each fragment, a Z x 1 vector can be incremented with the fragment's location within D using the following formula:
[0237]
number
[0243] For each Z component
[0244] Alternatively, or additionally, fragments can be represented by how informative they are about the feature. For example, the probability of observing a fragment in a given observation can be determined using the following formula: FragFreq = p(f;φ i )
[0245] The fraction of features expected to give rise to fragment f relative to the total number of features can be determined using the following formula:
[0238]
number
[0246] Next, each fragment is observed by iand for Z features, it can be expressed as Z+1 numbers corresponding to FragFreq × InverseSampleFreq.
[0239]
[0247] If the observations are represented as fixed sizes, then these representations can be additive. Thus: · The representation from two sets of fragments is equal to the sum of the representations from each set.
[0240] ·By adding together the D÷A columns of the matrix, we can reduce the representation from Z×D to Z×A.
[0248] The set of fragments is used only once in the above representation and can be calculated based only on the known θ parameter. Therefore, this method overcomes the problems described herein. Because the probability distribution can provide a smaller and more biologically accurate representation of the sample, the above ML method does not require fragmentation of the genome into small regions in order for the method to be computationally feasible. [Example]
[0241] Example 1: Classification of liver disease using methylation data from patient plasma samples.
[0249] Plasma samples were collected from individual patients previously diagnosed with various liver diseases, including nonalcoholic fatty liver disease (NAFLD), nonalcoholic steatohepatitis (NASH), and cirrhosis. Using the method described above, genome-wide DNA methylation patterns were determined. First, cell-free DNA (cfDNA) was extracted from biological samples, e.g., plasma isolated from blood. The extracted DNA was then treated with sodium bisulfite to convert unmethylated cytosines to uracil, while methylated cytosines remained unchanged. The bisulfite-treated DNA was then subjected to library preparation, including end repair and A-tailing, to blunt the DNA ends and add adenine nucleotides to the 3' ends of each strand. Following this, specific adapters were ligated to the DNA ends to enable binding to the sequencing platform and to provide sites for primer binding during amplification. The adapter-ligated DNA was then subjected to PCR amplification. Using high-throughput DNA sequencing technology, the amplified DNA was sequenced to determine the methylation patterns of DNA molecules in the cfDNA samples, resulting in the generation of approximately 500 million cfDNA reads with information on approximately 28 million CpGs.
[0242]
[0250] Furthermore, independent data from methylation microarrays were utilized to generate the described signatures using the above method. These microarrays contained data from multiple cell types, such as liver, brain, and heart cells, in both healthy and disease states. This approach avoided the use of labels indicating cell type or state and relied solely on methylation microarray data.
[0243]
[0251] The methylation data were computationally processed to generate a set of three features (Z=3) distributed within an exponential family.
[0252] For each plasma sample (each containing approximately 500 million cfDNA reads), the generated features were used to convert the cfDNA into a fixed feature set. cfDNA fragments were mapped to specific genomic locations, and then the fragments were converted into Z=3 features, one for each feature. The fragment frequency and inverse sample frequency were then calculated for each fragment, and another feature was calculated as fragment frequency × inverse sample frequency, resulting in 3 + 1 = 4 features per fragment.
[0244]
[0253] The feature of each fragment was added to the CpG position of its first CpG, ultimately converting the entire sample into 4 x approximately 28 million features.
[0254] This process was further enhanced with additive features as described above to further reduce the dimensionality of the sample representation from 4 x approximately 28 million features to 4 x 100, for a total of 4*100 = 400 features.
[0245]
[0255] While various machine learning training methods can be applied to these representations, a simplified approach using a 1-nearest neighbor classifier was used to demonstrate the effectiveness of the disclosed method. Using independent microarray data, the mean representation of liver disease was calculated, and a score indicating the distance between the sample and the mean liver disease representation was calculated for each sample.
[0246]
[0256] These methods were repeated for several applications, including distinguishing NASH from non-NASH (healthy) samples (Figure 4), distinguishing at-risk NASH from non-risk NASH samples (at-risk NASH is defined as individuals with NASH and stage 2 or greater fibrosis) (Figure 5), distinguishing NASH samples with and without cirrhosis (Figure 6), and distinguishing early NASH, late NASH, and non-NASH (healthy) samples (Figure 7).
[0247]
[0257] The results shown in Figure 4 and Figure 6 demonstrate that the disclosed method can be used to identify subjects with liver pathology. Figure 4 shows the identification of NASH, and Figure 6 shows the identification of cirrhosis. Figure 5 shows that the disclosed method can also be used to stratify subjects with liver pathology based on prognosis. Figure 7 shows that the disclosed method can be used to distinguish between early and late liver disease.
Claims
1. 1. A method for identifying whether a subject has or is at increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from a subject; (b) assaying the cfDNA sample or a derivative thereof to determine the methylation pattern or level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; (d) generating an electronic report indicative of the subject having or at increased risk of developing liver disease based at least in part on the output; and A method comprising:
2. 2. The method of claim 1, wherein the assay comprises identifying methylation patterns and methylation levels of DNA molecules of the cfDNA sample, and the methylation patterns and methylation levels are processed using a trained ML algorithm.
3. The method of claim 1 or 2, wherein the assay comprises performing sequencing.
4. 4. The method of claim 3, further comprising treating the DNA molecules of the cfDNA sample with a reaction mixture containing an enzyme for methylation-aware sequencing prior to sequencing.
5. 4. The method of claim 3, further comprising treating the DNA molecules of the cfDNA sample with a reaction mixture containing bisulfite prior to sequencing.
6. The method of any one of claims 1 to 5, wherein the assay comprises amplification.
7. The method of claim 6 , wherein the amplification comprises polymerase chain reaction (PCR).
8. 8. The method of any one of claims 1 to 7, wherein the cfDNA sample is obtained or derived from a plasma sample, a serum sample, a urine sample, a saliva sample or a liver tissue sample.
9. 9. The method of any one of claims 1 to 8, further comprising fractionating a whole blood sample from the subject to provide a cfDNA sample.
10. 10. The method of any one of claims 1-9, wherein (a) comprises subjecting a cfDNA sample to conditions sufficient to isolate, enrich or extract a set of DNA molecules, and (b) comprises assaying the DNA molecules.
11. 11. The method of claim 10, wherein (b) comprises using nucleic acid primers or probes to selectively enrich a set of DNA molecules corresponding to a panel of one or more genomic regions.
12. 12. The method of claim 11, wherein the one or more genomic regions are selected from the group consisting of the genes listed in Table 1.
13. 12. The method of claim 10 or 11, wherein the nucleic acid primers or probes have sequence complementarity to nucleic acid sequences of a panel of one or more genomic regions.
14. 10. The method of any one of claims 1 to 9, wherein the cfDNA sample is assayed without nucleic acid isolation, enrichment or extraction.
15. The method of any one of claims 1 to 14, wherein the subject is asymptomatic for the liver disease.
16. 16. The method of any one of claims 1-15, wherein the output indicates with at least 50% accuracy whether the cfDNA sample is positive for liver disease.
17. 17. The method of claim 16, wherein accuracy is determined by calculating the percentage of independent samples that are correctly identified as having or not having liver disease.
18. 18. The method of any one of claims 1-17, wherein the output indicates whether the cfDNA sample is positive for liver disease with a clinical sensitivity of at least 50%.
19. 19. The method of claim 18, wherein the clinical sensitivity is at least 50%.
20. 20. The method of any one of claims 1-19, wherein the output indicates whether the cfDNA sample is positive for liver disease with a clinical specificity of at least 50%.
21. 21. The method of claim 20, wherein the clinical specificity is at least 50%.
22. 22. The method of any one of claims 1-21, wherein the output indicates whether the cfDNA sample is positive for liver disease with a positive predictive value of at least 50%.
23. 23. The method of any one of claims 1-22, wherein the output indicates whether the cfDNA sample is positive for liver disease with a negative predictive value of at least 50%.
24. 24. The method of any one of claims 1-23, wherein the output indicates whether the cfDNA sample is positive for liver disease with an area under the receiver operating characteristic curve (AUROC) of at least 0.
50.
25. 25. The method of any one of claims 1-24, wherein the output indicates whether the cfDNA sample is positive for liver disease with a positive likelihood ratio of at least about 1.
3.
26. 26. The method of any one of claims 1-25, wherein the output indicates whether the cfDNA sample is negative for liver disease with a negative likelihood ratio of at most about 0.
75.
27. The method according to any one of claims 1 to 26, wherein the liver disease is early stage liver disease.
28. The method according to any one of claims 1 to 26, wherein the liver disease is an advanced stage liver disease.
29. 29. The method of any one of claims 1 to 28, wherein the liver disease is non-alcoholic steatohepatitis (NASH) or metabolic dysfunction-associated steatohepatitis (MASH).
30. The method of any one of claims 1 to 28, wherein the liver disease is fibrosis.
31. The method according to any one of claims 1 to 28, wherein the liver disease is cirrhosis.
32. The method of any one of claims 1 to 28, wherein the liver disease is hepatocellular carcinoma (HCC).
33. The method according to any one of claims 1 to 28, wherein the liver disease is hepatobiliary cancer.
34. The method of any one of claims 1 to 28, wherein the liver disease is viral hepatitis.
35. The method according to any one of claims 1 to 28, wherein the liver disease is non-alcoholic fatty liver disease (NAFLD) or metabolic dysfunction-associated fatty liver disease (MASLD).
36. The method of any one of claims 1 to 28, wherein the liver disease is non-alcoholic fatty liver (NAFL) or steatosis.
37. The method according to any one of claims 1 to 28, wherein the liver disease is metabolic dysfunction-associated fatty liver disease (MAFLD).
38. The method according to any one of claims 1 to 28, wherein the liver disease is alcohol-related liver disease (ALD).
39. 29. The method of any one of claims 1 to 28, wherein the liver disease is metabolic dysfunction alcohol-related liver disease (MetALD).
40. 40. The method of any one of claims 1-39, further comprising providing a therapeutic intervention for liver disease to a subject based at least in part on the output.
41. 41. The method of claim 40, wherein the liver disease is NASH and the therapeutic intervention is vitamin E supplementation, weight loss agents, antihypertensive agents, antidiabetic agents, cholesterol-lowering agents, exercise therapy, diet therapy, bariatric surgery, or a combination thereof.
42. 41. The method of claim 40, wherein the liver disease is NASH and the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, an FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, an SCD-1 (stearoyl-coenzyme A desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, an FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted gene modifier, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or a combination thereof.
43. 41. The method of claim 40, wherein the liver disease is NAFLD and the therapeutic intervention is vitamin E supplementation, a weight loss agent, an antihypertensive agent, an antidiabetic agent, a cholesterol-lowering agent, exercise therapy, diet therapy, bariatric surgery, or a combination thereof.
44. 41. The method of claim 40, wherein the liver disease is NAFLD and the therapeutic intervention is a GLP1 (glucagon-like peptide-1) receptor agonist, an FGF (fibroblast growth factor) analog, a THR (thyroid hormone receptor) agonist, an SCD-1 (stearoyl-coenzyme A desaturase 1) inhibitor, a FAS (fatty acid synthase) inhibitor, an FXR (farnesoid X receptor) agonist, an ACC (acetyl-CoA carboxylase) inhibitor, a PPAR (peroxisome proliferator-activated receptor) agonist, a targeted gene modifier, a LOXL2 (lysyl oxidase-like 2) inhibitor, a pan-cyclophilin inhibitor, a pan-caspase inhibitor, a chemokine receptor (e.g., CCR2 / CCR5) inhibitor, a galactin-3 inhibitor, a mitochondrial uncoupler or uncoupler, a structurally engineered fatty acid, or a combination thereof.
45. 45. The method of any one of claims 1-44, further comprising monitoring the subject for liver disease at two or more time points based at least in part on the output.
46. 46. The method of any one of claims 1 to 45, further comprising determining a probability or risk score that the subject has or is at increased risk of having liver disease.
47. 47. The method of any one of claims 1 to 46, further comprising determining the molecular subtype, grade, stage, or severity of the liver disease.
48. 48. The method of any one of claims 1 to 47, further comprising determining the prognosis of the liver disease.
49. 49. The method of any one of claims 1 to 48, further comprising determining the outcome of the liver disease.
50. 50. The method of any one of claims 1 to 49, further comprising determining the subject's eligibility as a liver transplant donor or liver transplant recipient.
51. 51. The method of claim 50, wherein the subject is determined to be eligible as a liver transplant donor if the subject is not identified as having liver disease or as being at increased risk for developing liver disease.
52. 51. The method of claim 50, wherein the subject is determined to be eligible as a liver transplant recipient if the subject is identified as having liver disease or as being at increased risk of developing liver disease.
53. 53. The method of any one of claims 1 to 52, wherein the trained ML algorithm is trained using an independent set of samples associated with the presence of or increased risk of liver disease.
54. 54. The method of any one of claims 1-53, wherein (c) further comprises processing the subject's set of clinical health data using a trained ML algorithm or another trained algorithm.
55. 55. The method of claim 54, wherein the clinical health data comprises one or more quantitative measures selected from the group consisting of age, weight, height, body mass index (BMI), blood pressure, heart rate, aspartate aminotransferase (AST) levels, alanine transaminase (ALT) levels, gamma-glutamyltransferase (GGT), platelet count, triglyceride levels, glycated hemoglobin (HbA1c) levels, creatinine levels, insulin levels, prothrombin time, haptoglobin levels, and glucose levels.
56. 56. The method of claim 54 or 55, wherein the clinical health data includes one or more categorical measures selected from the group consisting of race, ethnicity, history of medication or other clinical treatment, alcohol consumption history, level of daily activity or fitness, genetic test results, blood test results, and imaging results.
57. The method of any one of claims 1 to 56, wherein the trained ML algorithm comprises a supervised ML algorithm.
58. 58. The method of claim 57, wherein the supervised ML algorithm comprises a classifier or a regression.
59. 59. The method of claim 57 or 58, wherein the supervised ML algorithm comprises a deep learning algorithm, a support vector machine (SVM), a neural network, a random forest, linear regression, or logistic regression.
60. 60. The method of any one of claims 1 to 59, wherein the methylation pattern or level is represented by a parameter of a distribution, a sufficient statistic, or an approximately sufficient statistic.
61. 1. A method for monitoring liver disease in a subject, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from a subject; (b) assaying the cfDNA sample or a derivative thereof to determine the methylation pattern or level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; (d) generating an electronic report indicating progression of liver disease in the subject based at least in part on the output; A method comprising:
62. 1. A method for identifying a liver disease prognosis in a subject having or at increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from a subject; (b) assaying the cfDNA sample or a derivative thereof to determine the methylation pattern or level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; (d) generating an electronic report indicating a prognosis for the subject having or at increased risk of developing liver disease based at least in part on the output; and A method comprising:
63. 1. A method for identifying a treatment for a subject having or at increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from a subject; (b) assaying the cfDNA sample or a derivative thereof to determine the methylation pattern or level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; (d) generating an electronic report indicating treatment for a subject having or at increased risk of developing liver disease based at least in part on the output; and A method comprising:
64. 1. A method for determining treatment response in a subject having or at increased risk of developing liver disease, comprising: (a) providing a cell-free deoxyribonucleic acid (cfDNA) sample from a subject; (b) assaying the cfDNA sample or a derivative thereof to determine the methylation pattern or level of DNA molecules of the cfDNA sample; (c) processing the methylation pattern or methylation level using a trained machine learning (ML) algorithm to generate an output indicating whether the cfDNA sample is positive for liver disease; (d) generating an electronic report indicating treatment response for the subject having or at increased risk of developing liver disease based at least in part on the output; A method comprising:
65. 1. A method for determining whether a subject has or is at increased risk of developing liver disease, comprising: (a) providing a cell-free nucleic acid sample from a subject; (b) assaying the cell-free nucleic acid sample or a derivative thereof to determine the methylome of the cell-free nucleic acid sample; (c) processing the methylome using a trained machine learning (ML) algorithm to determine whether the subject has or is at increased risk for developing liver disease, wherein the determination has a sensitivity of at least about 70% and a specificity of at least about 70%. A method comprising: