Liver cancer methylation markers and machine models
A machine learning-based approach using methylation fraction and variance scores, combined with demographic and protein information, addresses the limitations of current liver cancer detection methods, achieving improved sensitivity and specificity for early-stage HCC detection.
Patent Information
- Application Number
- PCT/US2025/031786
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-04
- Filing Date
- 2025-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Current methods for detecting liver cancer, particularly hepatocellular carcinoma (HCC), suffer from low sensitivity and specificity, especially in early stages, with existing biomarkers like alpha-fetoprotein (AFP) assays and ultrasonography providing inadequate performance, leading to late diagnoses and poor prognosis.
A method utilizing methylation fraction and methylation variance scores, combined with demographic and protein information, is determined using machine learning models to assess liver cancer status, employing sequence read data from genomic regions and integrating these scores into ensemble models for improved detection.
The method achieves high sensitivity and specificity in detecting liver cancer, enabling early-stage detection and improving treatment outcomes by overcoming low signal-to-noise issues in circulating tumor DNA, reducing sequencing costs, and enhancing model generalizability across diverse patient populations.
Smart Images

Figure IMGF000058_0001 
Figure IMGF000059_0001 
Figure IMGF000060_0001
Abstract
Description
LIVER CANCER METHYLATION MARKERS AND MACHINE MODELSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority benefit of U.S. Provisional Application No. 63 / 654,819, filed on May 31, 2024, and U.S. Provisional Application No. 63 / 690,523, filed on September 4, 2024, the contents of each of which are hereby incorporated by reference in their entirety.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (333412001340SEQLIST.xml; Size: 10,090 bytes; and Date of Creation: May 22, 2025) is herein incorporated by reference in its entirety.TECHNICAL FIELD OF THE DISCLOSURE
[0003] The present invention relates generally to methods for the determination of a methylation cancer score and / or a demographic -protein cancer score in an individual suspected of having a liver cancer using machine learning models. More specifically, the present invention relates to methods for the determination of a liver cancer status in an individual based on a methylation cancer score and / or a demographic -protein cancer score using machine learning models.BACKGROUND OF THE DISCLOSURE
[0004] Liver cancer, in particular the hepatocellular carcinoma (HCC), is the fifth most common neoplasm in the world and the leading cause of death among cirrhotic patients. Any focal liver lesion in a patient with cirrhosis is suggestive of HCC, and early detection may permit curative treatment in 30%-40% of patients (Bruix J, et al. J Hepatol. 2001;35:421-430). a- fetoprotein (AFP) assay is the most frequent biologic screening test, but the diagnostic performance is poor. The radiologic modality most widely used for screening is ultrasonography, with a sensitivity around 45% for early detection of HCC (Tzartzeva K, et al. Gastroenterology 2018;154:1706-18).
[0005] The threat of HCC is expected to continue to grow in the coming years (Llovet J M, et al., Liver Transpl. 2004 Feb. 10(2 Suppl 1):S115-20). HCC prognosis remains dismal (5-year survival <15%) due to frequent diagnoses at late, non-curable stages (Ferlay J, et al. Int J Cancer. 2015;136:E359-86). Accordingly, there is a great need for early detection of HCC to improve the survival rate of these patients.
[0006] In an effort to improve HCC detection, the GALAD score was developed by combining age, sex, AFP, Lens culinaris agglutinin-reactive AFP (AFP-L3%), and des-gamma- carboxy prothrombin (DCP) (Johnson P et al. Cancer Epidemiol Biomarkers Prev. 2014;23:144- 53). The GALAD score has shown superior HCC detection performance compared with AFP, but there is room for improvement (Berhane S et al. Clin Gastroenterol Hepatol.2016;14:e876.14; Best J et al. Clin Gastroenterol Hepatol. 2020;18:e724), indicating the still unmet need for HCC surveillance tests with substantially improved performance characteristics.SUMMARY OF THE INVENTION
[0007] In some aspects, provided herein is a method for determining a methylation cancer score of an individual suspected of having a liver cancer, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; and inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual.
[0008] In any of the preceding embodiments, for the methylation fraction value for the genomic region of one or more first genomic regions, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with any of the one or more first genomic regions. In any of the preceding embodiments, for the methylation fraction value for the genomic region of the one or more first genomic regions, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having: (i) at least one base overlapping with the genomic region of the one or more first genomic regions, (ii) at least one methylated CpG site, and (iii) at least 50% of total CpG sites having a methylation.
[0009] In any of the preceding embodiments, for the methylation variance value for the genomic region of the one or more second genomic regions, reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having: (i) at least one base overlapping with the genomic region of the one or more second genomic regions, and (ii) at least one CpG site.
[0010] In any of the preceding embodiments, the method further comprises determining a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In any of the preceding embodiments, the liver cancer status is a liver cancer-positive status when the methylation cancer score is greater than the predetermined methylation cancer score threshold.
[0011] In any of the preceding embodiments, the method further comprises processing, at the one or more processors, extracted the methylation fraction value for each of the one or more first genomic regions before inputting into the first trained machine learning model. In any of the preceding embodiments, the method further comprises processing, at the one or more processors, extracted the methylation variance value for each of the one or more second genomic regions before inputting into the second trained machine learning model. In any of the preceding embodiments, the method further comprises processing, at the one or more processors, extracted the methylation fraction value for each of the one or more first genomic regions before inputting into the first trained machine learning model and processing, at the one or more processors, extracted the methylation variance value for each of the one or more second genomic regions before inputting into the second trained machine learning model. In any of the preceding embodiments, the processing of the extracted methylation value and the extracted methylationvariance value comprises inputting missing data, feature selection, data transformation, and / or data scaling.
[0012] In any of the preceding embodiments, the method further comprises receiving, at the one or more processors, demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual. In any of the preceding embodiments, the AFP level is a concentration of AFP. In any of the preceding embodiments, the AFP-L3 level is a percentage of AFP-L3 relative to total AFP. In any of the preceding embodiments, the DCP level is a concentration of DCP.
[0013] In any of the preceding embodiments, the method further comprises determining, using the one or more processors, a demographic-protein cancer score based on the demographic and protein information. In any of the preceding embodiments, the method further comprises determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic -protein cancer score and a predetermined demographic -protein cancer score threshold. In any of the preceding embodiments, a liver cancer-positive status is determined when the methylation cancer score is greater than the predetermined methylation cancer score threshold. In any of the preceding embodiments, a liver cancer-positive status is determined when the demographic -protein cancer score is greater than the predetermined demographic -protein cancer score threshold. In any of the preceding embodiments, a liver cancer-positive status is determined when (i) the methylation cancer score is greater than the predetermined methylation cancer score threshold, and / or (ii) the demographic-protein cancer score is greater than the predetermined demographic-protein cancer score threshold.
[0014] In some aspects, provided herein is a method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence readdata, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining the liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold.
[0015] In some aspects, provided herein is a method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining a liver cancer status of theindividual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and a predetermined demographic-protein cancer score threshold.
[0016] In any of the preceding embodiments, the method further comprises performing a sequencing assay to generate the sequence read data. In any of the preceding embodiments, the sequencing assaying is a pair-end sequencing assay. In any of the preceding embodiments, performing the sequencing assay comprises: enzymatically converting unmethylated cytosines to uracils in the sample; amplifying nucleic acids in the sample using polymerase chain reaction (PCR); and sequencing the amplified nucleic acids using a next generation sequencing (NGS) technique. In any of the preceding embodiments, the enzymatic conversion comprises: treating the sample with an oxidizing agent to oxidize methylated cytosines; and / or treating the sample with a glucosyltransferase to glucosylate methylated cytosines; and treating the sample with a deaminating agent to deaminate unmethylated cytosines to uracils. In any of the preceding embodiments, the oxidizing agent is tet methylcytosine dioxygenase 2 (TET2). In any of the preceding embodiments, the glucosyltransferase is T4 phage beta-glucosyltransferase (T4-BGT). In any of the preceding embodiments, the deaminating agent is apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC).
[0017] In any of the preceding embodiments, the first trained machine learning model comprises a neural network model or trained deep learning model. In any of the preceding embodiments, the second trained machine learning model comprises a neural network model or trained deep learning model. In any of the preceding embodiments, the first and second trained machine learning models comprise a neural network model or trained deep learning model. In any of the preceding embodiments, the first trained machine learning model comprises a support vector machine model, a random forest machine model, or a logistic regression machine model. In any of the preceding embodiments, the second trained machine learning model comprises a support vector machine model, a random forest machine model, or a logistic regression machine model. In any of the preceding embodiments, the first and second trained machine learning models comprise a support vector machine model, a random forest machine model, or a logistic regression machine model. In any of the preceding embodiments, the method further comprises a cross-validation procedure.
[0018] In any of the preceding embodiments, the first trained machine learning model is trained using one or more training data sets comprising paired methylation fraction values from a plurality of individuals. In any of the preceding embodiments, the second trained machine learning model is trained using one or more training data sets comprising paired methylation variance values from a plurality of individuals. In any of the preceding embodiments, the first trained machine learning model is trained using one or more training data sets comprising paired methylation fraction values from a plurality of individuals; and, the second trained machine learning model is trained using one or more training data sets comprising paired methylation variance values from a plurality of individuals.
[0019] In any of the preceding embodiments, the trained ensemble machine learning model comprises a bagging machine learning model, a boosting machine learning model, a stacking machine learning model and a random forest machine learning model. In any of the preceding embodiments, the trained ensemble machine learning model is trained using one or more training data sets comprising paired methylation fraction values and methylation variance values from a plurality of individuals.
[0020] In any of the preceding embodiments, the method further comprises obtaining the sample from the individual.
[0021] In any of the preceding embodiments, the sample comprises a tissue biopsy sample or a liquid biopsy sample. In any of the preceding embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In any of the preceding embodiments, the sample is a liquid biopsy sample and comprises circulating tumor cells. In any of the preceding embodiments, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA).
[0022] In any of the preceding embodiments, the sample is a first sample, and wherein the demographic and protein information is obtained from a second sample from the individual. In any of the preceding embodiments, the method further comprises obtaining the second sample from the individual. In any of the preceding embodiments, the first sample and the second sample are the same sample. In any of the preceding embodiments, the first sample and the second sample are different samples. In any of the preceding embodiments, the second sample comprises a tissue biopsy sample or a liquid biopsy sample. In any of the precedingembodiments, the second sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In any of the preceding embodiments, the second sample is a liquid biopsy sample and comprises circulating tumor cells. In any of the preceding embodiments, the second sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA).
[0023] In any of the preceding embodiments, the one or more first genomic regions and the one or more second genomic regions are each individually between about 50 base pairs and about 500 base pairs in length. In any of the preceding embodiments, the one or more first genomic regions and the one or more second genomic regions are each individually 100 or 200 base pairs in length. In any of the preceding embodiments, the one or more first genomic regions are each individually 100 base pairs in length, and wherein the one or more second genomic regions are each individually 200 base pairs in length.
[0024] In any of the preceding embodiments, the one or more first genomic regions overlap with the one or more second genomic regions. In any of the preceding embodiments, the one or more first genomic regions do not overlap with the one or more second genomic regions.
[0025] In any of the preceding embodiments, the one or more first genomic regions comprise at least 900 regions. In any of the preceding embodiments, the one or more second genomic regions comprise at least 900 regions.
[0026] In any of the preceding embodiments, the one or more first genomic regions or second genomic regions individually comprise any of the genomic regions listed in Table 1.
[0027] In any of the preceding embodiments, the individual is a human.
[0028] In any of the preceding embodiments, the individual is suspected of having a liver cancer. In any of the preceding embodiments, the liver cancer is hepatocellular carcinoma (HCC).
[0029] In some aspects, provided herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions andtotal reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
[0030] In any of the preceding embodiments, the system further comprises instructions to determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In any of the preceding embodiments, the system further comprises instructions to receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP- L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; determine a demographic -protein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic -protein cancer score threshold.
[0031] In some aspects, provided herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive sequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one ormore first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
[0032] In any of the preceding embodiments, the non-transitory computer-readable storage medium further comprises instructions to determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In any of the preceding embodiments, the non-transitory computer-readable storage medium further comprises instructions to receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; determine a demographic-protein cancer score based on the demographic and protein information; and determine the liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic-protein cancer score threshold.BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Various aspects of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0034] FIG. 1 shows a non-limiting exemplary method for determining a liver cancer status of an individual, in accordance with some embodiments provided herein.
[0035] FIGS. 2A and 2B show the receiver operating characteristic (ROC) curves from the cross-validation data set for individuals having stage 1 hepatocellular carcinoma (HCC). Each curve represents a single cross-validation run. TPR (True Positive Rate) is calculated as the sensitivity, and the FPR (False Positive Rate) is calculated as 1- specificity. The vertical solidline represents a specificity of 0.9, and the diagonal dashed line represents chance level (identity line). FIG. 2A shows the ROC curves for the HCC case-control individuals, from Cohort 1 and 2, where only stage 1 HCC was included in the test fold. FIG. 2B shows the ROC curves for the HCC case-control individuals, from Cohort 1 and 2, and the prospective HCC individuals from Cohort 3, where only stage 1 HCC was included in the test fold.
[0036] FIGS. 3A and 3B show the ROC curves from the cross-validation data set for individuals having stage 2 or higher stages of HCC. Each curve represents a single cross- validation run. TPR is calculated as the sensitivity, and the FPR is calculated as 1- specificity. FIG. 3A shows the ROC curves for the HCC case-control individuals, from Cohort 1 and 2, where only stage 2 or higher stages of HCC were included in the test fold. FIG. 3B shows the ROC curves for the HCC case-control individuals, from Cohort 1 and 2, and the prospective HCC individuals from Cohort 3, where only stage 2 or higher stages of HCC were included in the test fold.
[0037] FIG. 4 shows the ROC curves from the cross-validation data set for individuals having stage 2 or higher stages for the HCC case-control individuals, from Cohort 1 and 2, and the prospective HCC individuals from Cohort 3, where only Cohort 3 was included in the test fold. Each curve represents a single cross-validation run. TPR is calculated as the sensitivity, and the FPR is calculated as 1- specificity.
[0038] FIG. 5 shows an exemplary computing device or system, in accordance with some embodiments provided herein.
[0039] FIG. 6 shows an exemplary computer system or computer network, in accordance with some embodiments of the systems provided herein.
[0040] FIG. 7 shows area under the curve (AUC) AUC plots for Stage 1 HCC samples taken from patients in “Non-progression” (n=19) versus “Progression” (n=22) categories.
[0041] FIG. 8 shows the median difference in methylation scores at diagnosis / pre-treatment and at an endpoint for patients in “Non-Progression” versus “Progression” categories.
[0042] FIG. 9 shows a stratification of patients using methylation scores at the time of patient diagnosis / pre-treatment based on a methylation score threshold of 0.75.
[0043] FIG. 10A shows cumulative sensitivities a 0, 6, 12, and 18 months for all HCC using either HelioLiver (HL) or Ultrasound (US) testing.
[0044] FIG. 10B shows cumulative sensitivities a 0, 6, 12, and 18 months for early-stage HCC using either HL or US testing.
[0045] FIG. 10C shows the size distribution for HCC diagnosis using either HL or US testing.DETAILED DESCRIPTION OF THE DISCLOSURE
[0046] Cancer is characterized by an abnormal growth of a cell caused by one or more mutations or modifications of a gene leading to dysregulated balance of cell proliferation and cell death. DNA methylation silences expression of tumor suppression genes, and presents itself as one of the first neoplastic changes. Methylation patterns found in neoplastic tissue and plasma demonstrate homogeneity, and in some instances are utilized as a sensitive diagnostic marker. For example, cMethDNA assay has been shown in one study to be about 91% sensitive and about 96% specific when used to diagnose metastatic breast cancer. In another study, circulating tumor DNA (ctDNA) was about 87.2% sensitive and about 99.2% specific when it was used to identify KRAS gene mutation in a large cohort of patients with metastatic colon cancer (Bettegowda et al., Detection of Circulating Tumor DNA in Early- and Late-Stage Human Malignancies. Sci. Transl. Med, 6(224) :ra24. 2014). The same study further demonstrated that ctDNA is detectable in >75% of patients with advanced pancreatic, ovarian, colorectal, bladder, gastroesophageal, breast, melanoma, hepatocellular, and head and neck cancers (Bettegowda et al.).
[0047] “Liquid biopsies” assaying the methylation of circulating cell-free DNA (cfDNA) released from cancer cells have been actively explored as a promising noninvasive biomarker to sensitively detect various cancer types, including HCC, at early stages (Tran N et al., JHEP Rep. 2021 ;3: 100304). Comprehensive methylome profiling of HCC tissue / plasma samples combined with machine-learning analysis has previous enabled the detection of cfDNA methylation markers associated with the presence of HCC in patients with chronic liver diseases as a potential HCC detection biomarkers (Hao X, et al., Proc Natl Acad Sci U S A. 2017; 114:7414- 9.; Xu R-H, et al. Nat Mater. 2017;16:1155-61).
[0048] The present invention is based, at least in part, on the inventors’ discovery of a method for determining a methylation cancer score using methylation fraction and methylation variance, and in certain embodiments demographic and protein information, that enables robust and accurate detection of liver cancer, such as hepatocellular carcinoma (HCC). The methods described herein have demonstrated high sensitivity and specificity for classifying individuals. The implementation of the subject methods will enable easy, flexible, noninvasive, and accurate liver cancer detection at early stages, and significantly improve treatment outcomes for a transformative reduction of liver cancer mortality. As discussed herein, these findings represent a significant advancement in the field of liver cancer testing.
[0049] In certain aspects of the disclosure provided herein, it is clinically pertinent that patients with liver cirrhosis are at a high risk of developing HCC, and it is recommended they undergo routine surveillance via biannual abdominal ultrasound according to current AASLD guidelines. However, ultrasound often fails to detect early stage and small lesions. Moreover, ultrasound suffers from low adherence rates (9%-20%) and the majority of high-risk patients end up not being regularly surveilled. Socioeconomic factors, ease of access, and scheduling difficulties can also contribute to the low adherence rates observed. Other factors due to which ultrasound is a sub-optimal screening / surveillance modality for HCC include: (a) operator / technician skill and experience, (b) ultrasound results are subjective and are impacted by the skill / experience of the operator, (c) BMI of patients, e.g., many high-risk patients also have a high BMI which impacts the quality of ultrasound results. Studies have shown that about 20% of patients who need to be surveilled have moderately to severely limited visualization for ultrasound due to obesity and / or MAFLD or alcohol related liver cirrhosis. Due to such factors, often by the time a patient is diagnosed, they already have advanced disease for which treatments and therapies are limited. Currently no viable and validated alternative to ultrasound exists. The blood-based methods provided herein were developed to address these challenges posed by ultrasound to improve patient outcomes.
[0050] During the development of the inventions provided herein, it was noted that patients with early-stage HCC have low levels of circulating tumor DNA. Thus, cfDNA methylation sequencing-based methods face the challenge of overcoming this low signal-to-noise issue. Assays based on mutations alone have poor predictive power for detecting HCC especially for smaller / early- stage lesions. Unless sequenced to sufficiently high coverage, there isn’t sufficientsignal to reliably detect cancer. Whole genome cfDNA methylation sequencing based methods remain costly given the high sequencing coverage requirements. The disclosure provided herein overcomes these issues and more by, e.g., reducing dimensionality of the feature space to enrich for cancer signal while keeping sequencing coverage high and sequencing costs reasonable. Nevertheless, this still results in millions of reads per patient which cannot be processed or analyzed by hand. The disclosure provided herein is based, at least in part, on extensive feature engineering work to featurize and quantify cfDNA methylation patterns to optimize the performance of custom downstream ML / AI model / algorithms. s.
[0051] The final cfDNA methylation model is a complex ensemble model which has two base models as input. Each base model uses a different cfDNA methylation feature as input discovered to provide improve specificity and sensitivity of results. The architecture of the methods, models, and features is the results of extensive development and training on hundreds of patients from multiple sies across the country over a period of several years.
[0052] More specifically, two new cfDNA methylation-based features, methylation fraction and methylation variance, were developed for inputs to downstream AI / ML models. Such features were found to amply cancer signal in the blood. Use of these feature required the development of novel AI / ML models and training on proprietary data. For example, the models were trained using a diverse set of liver cirrhosis etiologies to ensure model generalizability and high predictive power in new patients with varying causes liver cirrhosis. Furthermore, the models described herein were shaped based on iterative improvements via validation and cross validation.
[0053] Thus, in some aspects, disclosed herein are methods and systems of determining a liver cancer status an individual based on various DNA methylation metrics of specific genomic regions, including methylation fraction and methylation variance, and / or demographic and protein information of the individual. Use of the term genomic regions is used, in certain embodiments, to describe the origin of a sequence found in cfDNA, and is not meant to imply that the analyzed DNA is genomic DNA.
[0054] In some embodiments, the invention relates generally to non-invasive methods for the determination of a methylation cancer score and / or a demographic-protein cancer score in an individual using machine learning models. More specifically, the present invention relates tomethods for the determination of a liver cancer status in an individual based on a methylation cancer score and / or a demographic-protein cancer score using machine learning models.
[0055] All publications, including patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.
[0056] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.I. Definitions
[0057] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0058] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For instance, where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictate otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in thedisclosure. In some embodiments, two opposing and open ended ranges are provided for a feature, and in such description it is envisioned that combinations of those two ranges are provided herein. For example, in some embodiments, it is described that a feature is greater than about 10 units, and it is described (such as in another sentence) that the feature is less than about 20 units, and thus, the range of about 10 units to about 20 units is described herein.
[0059] The term “about” as used herein refers to the usual error range for the respective value readily known in this technical field. Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X.”
[0060] As used herein, including in the appended claims, the singular forms “a,” “or,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.” It is understood that aspects and variations described herein include embodiments “consisting” and / or “consisting essentially of’ such aspects and variations.
[0061] As used herein, a “subject” or an “individual,” which are terms that are used interchangeably, is a mammal. In some embodiments, a “mammal” includes humans, nonhuman primates, domestic and farm animals, and zoo, sports, or pet animals, such as dogs, horses, rabbits, cattle, pigs, hamsters, gerbils, mice, ferrets, rats, cats, monkeys, etc. In some embodiments, the subject or individual is human.
[0062] In some embodiments, the term “diagnosis” is used herein to refer to the identification or classification of a molecular or pathological state, disease or condition. For example, “diagnosis” may refer to identification of a particular type of cancer, e.g., a liver cancer. “Diagnosis” may also refer to the classification of a particular type of cancer, e.g., by histology (e.g., a hepatocellular carcinoma), by DNA methylation status in a particular gene or genes and / or proteins, or combination of both.
[0063] Those skilled in the art will recognize that several embodiments are possible within the scope and spirit of the present disclosure. The following description illustrates the disclosure and, of course, should not be construed in any way as limiting the scope of the inventions described herein.II. Methods
[0064] Disclosed herein, in certain embodiments, are methods of determining a methylation cancer score of an individual suspected of having a liver cancer. The methylation cancer score of the individual can be used, optionally in combination with demographic and protein information from the individual (e.g., a demographic -protein cancer score), to determine a liver cancer status of the individual. In some embodiments, the liver cancer status may enable the selection of individuals suspected of having liver cancer for treatment. In some instances, the methods comprise utilizing a methylation cancer score (e.g., a combined score based on a methylation fraction score and a methylation variance score) and / or a demographic-protein cancer score described herein, to determine a liver cancer status of an individual.
[0065] FIG. 1 shows a non-limiting exemplary method for determining a liver cancer status of an individual, in accordance with some embodiments provided herein. At 102, the method comprises receiving, at one or more processors, sequence read data from a sample from the individual (e.g., an individual suspected of having a liver cancer). Although not shown, the method may further comprise obtaining the sequence read data from the individual, such as by performing a sequencing assay. The sequence read data may contain information from multiple individuals, and the sequence read data from these individual can be uniquely indexed. Sequence read data that contains information from multiple individuals may be subsequently demultiplexed to generate independent sequence read data for each individual, and the demultiplexed sequence read data can be used in the methods provided herein. Demultiplexing may include processing through a custom bioinformatics pipeline. For example, the sequence read data may be trimmed to remove any adapter or primer sequences prior to alignment to a reference genome (e.g., a human reference genome, such as hg38). The trimmed and aligned read sequence data can be filtered to remove non-primary alignments, supplementary alignments, and alignments that fail a platform or vendor quality control check. Moreover, the filtered read sequence data can be flagged to allow for the removal of duplicate reads. The filtered and duplicate-flagged read sequence data can be used to calculate a next generation sequencing (NGS) quality control (QC) metric for each sample, and / or to calculate CpG site level and read level methylation metrics (e.g., methylation fraction value and / or methylation variance value). For generating cfDNA methylation features for downstream machine learning models, duplicate reads (and their alignment) and reads with a quality metric that fall below apredetermined threshold can be excluded. The sequence read data may be subsequently extracted to determine certain features of the sequence read data. For example, the sequence read data can contain information regarding the methylation status, such as methylation fraction and methylation variance, of genomic region.
[0066] Specifically, at 104, the method comprises extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions. At 106, the method comprises extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions.
[0067] The extracted methylation fraction values for each of the one or more first genomic regions and the extracted methylation variance values for each of the one or more second genomic regions may be used to determine a methylation fraction score and a methylation variance score, respectively. The extracted methylation fraction value for each of the one or more first genomic regions are inputted into a first trained machine learning model to output a methylation fraction score, and the extracted methylation variance value for each of the one or more second genomic regions are inputted into a second trained machine learning model to output a methylation variance score. Accordingly, at 108 the method comprises inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score. At 110, the method comprises inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score.
[0068] The methylation fraction score and the methylation variance score may be used in combination to determine a methylation cancer score. For example, the methylation fraction score and the methylation variance score can be used as the input for a trained ensemble machine learning model to output the methylation cancer score of the individual. Thus, the method comprises inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output themethylation cancer score of the individual, at 112. A liver cancer status of the individual can be subsequently determined by comparing the methylation cancer score with a predetermined methylation cancer score threshold, such as at 118.
[0069] However, a method of determining a liver cancer status of an individual may also comprise determining the liver cancer status based on a demographic-protein cancer score. In such embodiments, at 114, the method comprises receiving, at one or more processors, demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alphafetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma- carboxy prothrombin (DCP) level in the individual. Next, the method comprises determining, using the one or more processors, a demographic-protein cancer score based on the demographic and protein information at 116. Finally, at 118, the method comprises determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and a predetermined demographic -protein cancer score threshold.
[0070] In some aspects, provided herein is a method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual suspected of having a liver cancer; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output amethylation cancer score; and determining the liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments, the method further comprises: receiving, at the one or more processors, demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual. In some embodiments, the method further comprises determining, using the one or more processors, a demographic-protein cancer score based on the demographic and protein information. In some embodiments, the method further comprises determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic -protein cancer score and a predetermined demographicprotein cancer score threshold. In some embodiments, the method further comprises: performing a sequencing reaction on the sample from the individual, wherein the sequencing reaction generates the sequence read data. In some embodiments, the sequencing reaction comprises NGS. In some embodiments, the NGS comprises paired-end sequencing. In some embodiments, the NGS comprises a step of demultiplexing raw sequence read data using barcodes. In some embodiments, the method further comprises processing the sequence read data. In some embodiments, the processing comprises removing low-quality sequence read data and / or removing duplicate sequence read data. In some embodiments, the processing occurs before the sequence read data is received at the one or more processors.
[0071] In some aspects, provided herein is a method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites perread across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic -protein cancer score and a predetermined demographic -protein cancer score threshold. In some embodiments, the method further comprises: performing a sequencing reaction on the sample from the individual, wherein the sequencing reaction generates the sequence read data. In some embodiments, the sequencing reaction comprises NGS. In some embodiments, the NGS comprises paired-end sequencing. In some embodiments, the NGS comprises a step of demultiplexing raw sequence read data using barcodes. In some embodiments, the method further comprises processing the sequence read data. In some embodiments, the processing comprises removing low-quality sequence read data and / or removing duplicate sequence read data. In some embodiments, the processing occurs before the sequence read data is received at the one or more processors.
[0072] In some aspects, provided herein is a method for determining a liver cancer status of an individual, the method comprising: performing a sequencing reaction on a sample from the individual, wherein the sequencing reaction generates sequence read data; processing the sequence read data; receiving, at one or more processors, the sequence read data and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more secondgenomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and a predetermined demographic -protein cancer score threshold.
[0073] Aspects of the methods disclosed herein are described in more detail below in a modular fashion. Such presentation is not to be construed as limiting the scope of combinations of the various aspects encompassed by the disclosure of the present application to form a method for processing components, or products thereof, of a sample.A. Samples
[0074] The disclosed methods may be used with a variety of samples. For example, in some instances, the sample is a biological sample isolated from an individual. Examples of a sample include, but are not limited to, a tumor sample, a tissue sample, a biopsy sample (e.g., a tissue biopsy, a liquid biopsy, or both), a blood sample (e.g., a peripheral whole blood sample), a blood plasma sample, a blood serum sample, a lymph sample, a saliva sample, a sputum sample, a urine sample, a gynecological fluid sample, a circulating tumor cell (CTC) sample, a cerebral spinal fluid (CSF) sample, a pericardial fluid sample, a pleural fluid sample, an ascites (peritoneal fluid) sample, a feces (or stool) sample, or other body fluid, secretion, and / or excretion sample (or cell sample derived therefrom). In certain instances, the sample may be frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.
[0075] In some embodiments, the sample comprises a biopsy sample. In some embodiments, the sample comprises a tissue biopsy sample or a liquid biopsy sample. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). In some embodiments, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof. In some embodiments, the sample is a cfDNA sample and comprises ctDNA.
[0076] In some embodiments, the sample is a liquid sample (e.g., a liquid biopsy sample). In some embodiments, the liquid sample comprises blood and other liquid samples of biological origin (including, but not limited to, peripheral blood, sera, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, cerumen, breast milk, broncheoalveolar lavage fluid, semen, prostatic fluid, cowper’s fluid or pre-ejaculatory fluid, female ejaculate, sweat, tears, cyst fluid, pleural and peritoneal fluid, pericardial fluid, ascites, lymph, chyme, chyle, bile, interstitial fluid, menses, pus, sebum, vomit, vaginal secretions / flushing, synovial fluid, mucosal secretion, stool water, pancreatic juice, lavage fluids from sinus cavities, bronchopulmonary aspirates, blastocyl cavity fluid, or umbilical cord blood. In some embodiments, the biological fluid is blood, a blood derivative or a blood fraction, e.g., serum or plasma. In a specific embodiment, a sample comprises a blood sample. In another embodiment, a serum sample is used. In another embodiment, a sample comprises urine. In some embodiments, the liquid sample also encompasses a sample that has been manipulated in any way after their procurement, such as by centrifugation, filtration, precipitation, dialysis, chromatography, treatment with reagents, washed, or enriched for certain cell populations.
[0077] In some embodiments, the sample is a tissue sample (e.g., a tissue biopsy sample). In some instances, a tissue corresponds to any cell(s). Different types of tissue correspond to different types of cells (e.g., liver, lung, blood, connective tissue, and the like), but also healthy cells vs. tumor cells or to tumor cells at various stages of neoplasia, or to displaced malignant tumor cells. In some embodiments, a tissue sample further encompasses a clinical sample, and also includes cells in culture, cell supernatants, organs, and the like. Samples also comprise fresh-frozen and / or formalin- fixed, paraffin-embedded tissue blocks, such as blocks preparedfrom clinical or pathological biopsies, prepared for pathological analysis or study by immunohistochemistry. Examples of tissues include, but are not limited to, connective tissue, muscle tissue, nervous tissue, epithelial tissue, and blood. Tissue samples may be collected from any of the organs within an animal or human body. Examples of human organs include, but are not limited to, the brain, heart, lungs, liver, kidneys, pancreas, spleen, thyroid, mammary glands, uterus, prostate, large intestine, small intestine, bladder, bone, skin, etc.
[0078] In some embodiments, the sample is a normal sample (e.g., normal or control tissue without disease, or normal or control body fluid, stool, blood, serum, amniotic fluid), most importantly in healthy stool, blood, serum, amniotic fluid or other body fluid. In other embodiments, the sample is from an individual having or at risk of a disease, such as an individual having or at risk of liver cancer. In some embodiments, the sample is hypermethylated or hypomethylated compared to a normal sample. For example, the hypermethylated or hypomethylated sample may have a decreased or increased (respectively) methylation frequency of at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100% in comparison to a normal sample. In one embodiment, the sample is also hypomethylated or hypermethylated in comparison to a previously obtained sample analysis of the same patient having or at risk of a disease (e.g., liver cancer), particularly to compare progression of a disease.
[0079] In some aspects, the methods provided herein comprises obtaining a sample from an individual. In some embodiments, the sample may be obtained from the individual by tissue resection (e.g., surgical resection), needle biopsy, bone marrow biopsy, bone marrow aspiration, skin biopsy, endoscopic biopsy, fine needle aspiration, oral swab, nasal swab, vaginal swab or a cytology smear, scrapings, washings or lavages (such as a ductal lavage or bronchoalveolar lavage), etc.
[0080] In some embodiments, a single sample is used to determine a methylation cancer score and a demographic-protein cancer score of an individual. In some embodiments, the single sample is obtained from an individual, e.g., the same individual. In some embodiments, the method comprises obtaining the sample from the individual to determine the methylation cancer score and the demographic -protein cancer score of the individual. In some embodiments, the sample is a cfDNA sample.
[0081] In some embodiments, a first sample is used to determine a methylation cancer score a demographic-protein cancer score of an individual. In some embodiments, the method comprises obtaining the first sample from the individual to determine the methylation cancer score. In some embodiments, a second sample is used to determine a demographic-protein cancer score of the individual. In some embodiments, the method comprises obtaining the second sample from the individual to determine the demographic -protein cancer score. In some embodiments, the first sample and the second sample are from the same individual. In some embodiments, the first sample and the second sample are the same type of sample, such as any of the sample types described herein. In some embodiments, the first sample and the second sample are a different type of sample, such as any of the sample types described herein. In some embodiments, the first sample is a cfDNA sample. In some embodiments, the second sample is a cfDNA sample.B. Individuals
[0082] The disclosed methods may be used to determine a methylation cancer score, a demographic-protein score, and / or a liver cancer status of a sample obtained from a variety of individuals, such as individuals suspected of having a liver cancer.
[0083] In some embodiments, the individual is a mammal. In some embodiments, the mammal is a human. In some embodiments, the mammal is a non-human. In some embodiments, the human is a male. In some embodiments, the human is a female. In some embodiments, the individual is suspected of having a liver cancer. None of the terms require or are limited to situations characterized by the supervision (e.g. constant or intermittent) of a health care worker (e.g. a doctor, a registered nurse, a nurse practitioner, a physician’s assistant, an orderly or a hospice worker).
[0084] In some embodiments, the sample is obtained (e.g., collected) from an individual having a liver cancer or at risk of having a liver cancer. For example, in some embodiments, the individual has a genetic predisposition to a liver cancer (e.g., having a genetic mutation that increases his or her baseline risk for developing a liver cancer). In some embodiments, the individual has been exposed to an environmental perturbation (e.g., radiation or a chemical) that increases his or her risk for developing a liver cancer. In some embodiments, the individual isbeing monitored for development of a liver cancer. In some embodiments, the individual is being monitored for liver cancer progression or regression, e.g., after being treated with an anticancer therapy (or anti-cancer treatment). In some embodiments, the individual is being monitored for relapse of liver cancer. In some embodiments, the individual has been, or is being treated, for liver cancer. In some embodiments, the liver cancer is HCC. In some embodiments, the liver cancer is malignant. In some embodiments, the individual is healthy e.g., does not have the liver cancer and / or another disease). In some embodiments, the individual has been diagnosed with benign disease. In some embodiments, the benign disease is a benign tumor, diabetes, liver cirrhosis, chronic hepatitis B or hepatitis C virus infection, chronic obstructive pulmonary disease (COPD), or other benign diseases. In some embodiments, the individual has an early stage liver cancer, such as stage I or II.C. Sequencing and methylation status determination
[0085] According to the provided methods, sequence read data of a sample from an individual suspected of having a liver cancer may be obtained to perform a methylation analysis, such as to determine a methylation fraction value and a methylation fraction score of an individual. Sequence read data can be in various forms and remain compatible with the present invention. For example, in some embodiments, a sample from an individual may be sequenced (e.g., using any of the provided sequencing methods) to produce a plurality of sequence fragments, with each sequence fragment comprising one or more sequence reads. The sample from the individual may analyzed, according to the provided methods, at the sequence fragment level or at the sequence read level. As used herein, the term “sequence read data” encompasses data on a single sequence read or reads of a sequence fragment (whether or not the single sequence read or reads of a sequence fragment are connected or linked in some manner).
[0086] In some embodiments, the sequence read data is obtained from sequencing, such as a sequencing assay. In some embodiments, the provided method (e.g., a method of determining a methylation cancer score in an individual and / or a method of determining a liver cancer status in an individual) comprises performing a sequencing assay to generate the sequence read data. In some embodiments, the sequencing assay is a paired-end sequencing assay. Paired-end sequencing comprises the sequencing of both ends of a sequence read (e.g., sequence fragment)and generating high-quality, alignable sequence data. In some embodiments, the sequencing assay is a non-disruptive methylation sequencing assay. In some embodiments, the sequencing assay comprises performing a NGS assay.
[0087] In some embodiments, a genomic region of DNA comprises a cytosine methylation site. In some instances, cytosine methylation comprises 5-methylcytosine (5-mCyt) and 5- hydroxymethylcytosine. In some cases, a cytosine methylation site occurs in a CpG dinucleotide motif. In other cases, a cytosine methylation site occurs in a CHG or CHH motif, in which ‘H’ is adenine, cytosine or thymine. In some instances, one or more CpG dinucleotide motif or CpG site forms a CpG island, a short DNA sequence rich in CpG dinucleotide. In some instances, CpG islands are typically, but not always, between about 0.2 to about 1 kb in length. In some instances, a genomic region comprises a CpG island.
[0088] In some embodiments, DNA (e.g., cfDNA) for sequencing is isolated from a sample of an individual by any means standard in the art, including the use of commercially available kits. Briefly, wherein the DNA of interest is encapsulated in by a cellular membrane the biological sample is disrupted and lysed by enzymatic, chemical or mechanical means. In some cases, the DNA solution is then cleared of proteins and other contaminants e.g. by digestion with proteinase K. The DNA is then recovered from the solution. In such cases, this is carried out by means of a variety of methods including salting out, organic extraction or binding of the DNA to a solid phase support. In some instances, the choice of method is affected by several factors including time, expense and required quantity of DNA.
[0089] Wherein the sample DNA is not enclosed in a membrane (e.g., circulating DNA from a cell free sample such as blood or urine) methods standard in the art for the isolation and / or purification of DNA are optionally employed (See, for example, Bettegowda et al. Detection of Circulating Tumor DNA in Early- and Late-Stage Human Malignancies. Sci. Transl. Med, 6(224): ra24. 2014). Such methods include the use of a protein degenerating reagent e.g. chaotropic salt e.g. guanidine hydrochloride or urea; or a detergent e.g. sodium dodecyl sulphate (SDS), cyanogen bromide. Alternative methods include but are not limited to ethanol precipitation or propanol precipitation, vacuum concentration amongst others by means of a centrifuge. In some cases, the person skilled in the art also make use of devices such as filter devices e.g. ultrafiltration, silica surfaces or membranes, magnetic particles, polystyrol particles,polystyrol surfaces, positively charged surfaces, and positively charged membranes, charged membranes, charged surfaces, charged switch membranes, charged switched surfaces.
[0090] In some instances, once the nucleic acids have been extracted, methylation analysis (e.g., sequencing and determination of methylation status) is carried out by any means known in the art. A variety of methylation analysis procedures are known in the art and may be used to practice the methods disclosed herein. These assays allow for determination of the methylation state of one or a plurality of CpG sites within a tissue sample. In addition, these methods may be used for absolute or relative quantification of methylated nucleic acids. Such methylation assays involve, among other techniques, two major steps. The first step is a methylation specific reaction or separation, such as (i) bisulfite treatment, (ii) methylation specific binding, or (iii) methylation specific restriction enzymes, or (iv) enzymatic conversion.-The second major step involves (i) amplification and detection, or (ii) direct detection, by a variety of methods such as (a) PCR (sequence- specific amplification) such as Taqman(R), (b) DNA sequencing of untreated and bisulfite-treated DNA, (c) sequencing by ligation of dye-modified probes (including cyclic ligation and cleavage), (d) pyrosequencing, (e) single-molecule sequencing, (f) mass spectroscopy, or (g) Southern blot analysis.
[0091] Additionally, restriction enzyme digestion of PCR products amplified from bisulfite- converted DNA may be used, e.g., the method described by Sadri and Hornsby (1996, Nucl. Acids Res. 24:5058- 5059), or COBRA (Combined Bisulfite Restriction Analysis) (Xiong and Laird, 1997, Nucleic Acids Res. 25:2532- 2534). COBRA analysis is a quantitative methylation assay useful for determining DNA methylation levels at specific gene loci in small amounts of DNA. Briefly, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite- treated DNA. Methylation-dependent sequence differences are first introduced into the DNA by standard bisulfite treatment according to the procedure described by Frommer et al. (Frommer et al, 1992, Proc. Nat. Acad. Sci. USA, 89, 1827-1831). PCR amplification of the bisulfite converted DNA is then performed using primers specific for the CpG sites of interest, followed by restriction endonuclease digestion, gel electrophoresis, and detection using specific, labeled hybridization probes. Methylation levels in the original DNA sample are represented by the relative amounts of digested and undigested PCR product in a linearly quantitative fashion across a wide spectrum of DNA methylation levels. In addition, this technique can be reliably applied to DNA obtained from micro-dissectedparaffin- embedded tissue samples. Typical reagents (e.g., as might be found in a typical COBRA- based kit) for COBRA analysis may include, but are not limited to: PCR primers for specific gene (or methylation-altered DNA sequence or CpG island); restriction enzyme and appropriate buffer; gene-hybridization oligo; control hybridization oligo; kinase labeling kit for oligo probe; and radioactive nucleotides. Additionally, bisulfite conversion reagents may include: DNA denaturation buffer; sulfo nation buffer; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity column); desulfonation buffer; and DNA recovery components.
[0092] In an embodiment, the methylation status of selected CpG sites is determined using methylation-Specific PCR (MSP). MSP allows for assessing the methylation status of virtually any group of CpG sites within a CpG island, independent of the use of methylation- sensitive restriction enzymes (Herman et al, 1996, Proc. Nat. Acad. Sci. USA, 93, 9821- 9826; U.S. Pat. Nos. 5,786,146, 6,017,704, 6,200,756, 6,265,171 (Herman and Baylin); U.S. Pat. Pub. No.2010 / 0144836 (Van Engeland et al)). Briefly, DNA is modified by a deaminating agent such as sodium bisulfite to convert unmethylated, but not methylated cytosines to uracil, and subsequently amplified with primers specific for methylated versus unmethylated DNA. In some instances, typical reagents (e.g., as might be found in a typical MSP- based kit) for MSP analysis include, but are not limited to: methylated and unmethylated PCR primers for specific gene (or methylation- altered DNA sequence or CpG island), optimized PCR buffers and deoxynucleotides, and specific probes. One may use quantitative multiplexed methylation specific PCR (QM-PCR), as described by Fackler et al. Fackler et al, 2004, Cancer Res. 64(13) 4442-4452; or Fackler et al, 2006, Clin. Cancer Res. 12(11 Pt 1) 3306-3310.
[0093] In an embodiment, the methylation profile of selected CpG sites is determined using MethyLight and / or Heavy Methyl Methods. The MethyLight and Heavy Methyl assays are a high-throughput quantitative methylation assay that utilizes fluorescence- based real-time PCR (Taq Man(R)) technology that requires no further manipulations after the PCR step (Eads, C.A. et al, 2000, Nucleic Acid Res. 28, e 32; Cottrell et al, 2007, J. Urology 177, 1753, U.S. Pat. Nos. 6,331,393 (Laird et al)). Briefly, the MethyLight process begins with a mixed sample of DNA that is converted, in a sodium bisulfite reaction, to a mixed pool of methylation-dependent sequence differences according to standard procedures (the bisulfite process converts unmethylated cytosine residues to uracil). Fluorescence-based PCR is then performed either inan “unbiased” (with primers that do not overlap known CpG methylation sites) PCR reaction, or in a “biased” (with PCR primers that overlap known CpG dinucleotides) reaction. In some cases, sequence discrimination occurs either at the level of the amplification process or at the level of the fluorescence detection process, or both. In some cases, the MethyLight assay is used as a quantitative test for methylation patterns in the DNA sample, wherein sequence discrimination occurs at the level of probe hybridization. In this quantitative version, the PCR reaction provides for unbiased amplification in the presence of a fluorescent probe that overlaps a particular putative methylation site. An unbiased control for the amount of input DNA is provided by a reaction in which neither the primers, nor the probe overlie any CpG dinucleotides. Alternatively, a qualitative test for genomic methylation is achieved by probing of the biased PCR pool with either control oligonucleotides that do not “cover” known methylation sites (a fluorescence- based version of the “MSP” technique), or with oligonucleotides covering potential methylation sites. Typical reagents (e.g., as might be found in a typical MethyLight- based kit) for MethyLight analysis may include, but are not limited to: PCR primers for specific gene (or methylation- altered DNA sequence or CpG island); TaqMan(R) probes; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0094] Quantitative MethyLight uses bisulfite to convert DNA and the methylated sites are amplified using PCR with methylation independent primers. Detection probes specific for the methylated and unmethylated sites with two different fluorophores provides simultaneous quantitative measurement of the methylation. The Heavy Methyl technique begins with bisulfate conversion of DNA. Next specific blockers prevent the amplification of unmethylated DNA. Methylated DNA does not bind the blockers and their sequences will be amplified. The amplified sequences are detected with a methylation specific probe. (Cottrell et al, 2004, Nuc. Acids Res. 32:el0, the contents of which is hereby incorporated by reference in its entirety).
[0095] The Ms-SNuPE technique is a quantitative method for assessing methylation differences at specific CpG sites based on bisulfite treatment of DNA, followed by singlenucleotide primer extension (Gonzalgo and Jones, 1997, Nucleic Acids Res. 25, 2529-2531). Briefly, DNA is reacted with sodium bisulfite to convert unmethylated cytosine to uracil while leaving 5-methylcytosine unchanged. Amplification of the desired target sequence is then performed using PCR primers specific for bisulfite-converted DNA, and the resulting product is isolated and used as a template for methylation analysis at the CpG site(s) of interest. In somecases, small amounts of DNA are analyzed (e.g., micro-dissected pathology sections), and the method avoids utilization of restriction enzymes for determining the methylation status at CpG sites. Typical reagents (e.g., as is found in a typical Ms-SNuPE-based kit) for Ms- SNuPE analysis include, but are not limited to: PCR primers for specific gene (or methylation- altered DNA sequence or CpG island); optimized PCR buffers and deoxynucleotides; gel extraction kit; positive control primers; Ms-SNuPE primers for specific gene; reaction buffer (for the Ms- SNuPE reaction); and radioactive nucleotides. Additionally, bisulfite conversion reagents may include: DNA denaturation buffer; sulfonation buffer; DNA recovery regents or kit (e.g., precipitation, ultrafiltration, affinity column); desulfonation buffer; and DNA recovery components.
[0096] In some embodiments, next generation sequencing data is generated for the whole genome. In some embodiments, next generation sequencing data is generated for a targeted set of genomic regions, e.g., using a hybridization-based target enrichment protocol.
[0097] In another embodiment, the methylation status of selected CpG sites is determined using differential Binding-based Methylation Detection Methods. For identification of differentially methylated regions, one approach is to capture methylated DNA. This approach uses a protein, in which the methyl binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC) (Gebhard et al, 2006, Cancer Res. 66:6118-6128; and PCT Pub. No. WO 2006 / 056480 A2 (Relhi)). This fusion protein has several advantages over conventional methylation specific antibodies. The MBD FC has a higher affinity to methylated DNA and it binds double stranded DNA. Most importantly the two proteins differ in the way they bind DNA. Methylation specific antibodies bind DNA stochastically, which means that only a binary answer can be obtained. The methyl binding domain of MBD-FC, on the other hand, binds DNA molecules regardless of their methylation status. The strength of this protein - DNA interaction is defined by the level of DNA methylation. After binding DNA, eluate solutions of increasing salt concentrations can be used to fractionate non- methylated and methylated DNA allowing for a more controlled separation (Gebhard et al, 2006, Nucleic Acids Res. 34: e82). Consequently this method, called Methyl-CpG immunoprecipitation (MCIP), not only enriches, but also fractionates DNA according to methylation level, which is particularly helpful when the unmethylated DNA fraction should be investigated as well.
[0098] In an alternative embodiment, a 5 -methyl cytidine antibody to bind and precipitate methylated DNA. Antibodies are available from Abeam (Cambridge, MA), Diagenode (Sparta, NJ) or Eurogentec (c / o AnaSpec, Fremont, CA). Once the methylated fragments have been separated they may be sequenced using microarray based techniques such as methylated CpG- island recovery assay (MIRA) or methylated DNA immunoprecipitation (MeDIP) (Pelizzola et al, 2008, Genome Res. 18, 1652-1659; O’Geen et al, 2006, BioTechniques 41(5), 577-580, Weber et al, 2005, Nat. Genet. 37, 853-862; Horak and Snyder, 2002, Methods Enzymol, 350, 469-83; Lieb, 2003, Methods Mol Biol, 224, 99-109). Another technique is methyl-CpG binding domain column / segregation of partly melted molecules (MBD / SPM, Shiraishi et al, 1999, Proc. Natl. Acad. Sci. USA 96(6):2913-2918).
[0099] In some embodiments, methods for detecting methylation include randomly shearing or randomly fragmenting the DNA, cutting the DNA with a methylation-dependent or methylation-sensitive restriction enzyme and subsequently selectively identifying and / or analyzing the cut or uncut DNA. Selective identification can include, for example, separating cut and uncut DNA (e.g., by size) and quantifying a sequence of interest that was cut or, alternatively, that was not cut. See, e.g., U.S. Patent No. 7,186,512. Alternatively, the method can encompass amplifying intact DNA after restriction enzyme digestion, thereby only amplifying DNA that was not cleaved by the restriction enzyme in the area amplified. See, e.g., U.S. Patents No. 7,910,296; No. 7,901,880; and No. 7,459,274. In some embodiments, amplification can be performed using primers that are gene specific.
[0100] For example, there are methyl-sensitive enzymes that preferentially or substantially cleave or digest at their DNA recognition sequence if it is non-methylated. Thus, an unmethylated DNA sample is cut into smaller fragments than a methylated DNA sample. Similarly, a hypermethylated DNA sample is not cleaved. In contrast, there are methylsensitive enzymes that cleave at their DNA recognition sequence only if it is methylated. Methyl- sensitive enzymes that digest unmethylated DNA suitable for use in methods of the technology include, but are not limited to, Hpall, Hhal, Maell, BstUI and Acil. In some instances, an enzyme that is used is Hpall that cuts only the unmethylated sequence CCGG. In other instances, another enzyme that is used is Hhal that cuts only the unmethylated sequence GCGC. Both enzymes are available from New England BioLabs(R), Inc. Combinations of two or more methyl-sensitive enzymes that digest only unmethylated DNA are also used. Suitableenzymes that digest only methylated DNA include, but are not limited to, Dpnl, which only cuts at fully methylated 5’-GATC sequences, and McrBC, an endonuclease, which cuts DNA containing modified cytosines (5-methylcytosine or 5-hydroxymethylcytosine or N4- methylcytosine) and cuts at recognition site 5’... PumC(N4o-3ooo) PumC... 3’ (New England BioLabs, Inc., Beverly, MA). Cleavage methods and procedures for selected restriction enzymes for cutting DNA at specific sites are well known to the skilled artisan. For example, many suppliers of restriction enzymes provide information on conditions and types of DNA sequences cut by specific restriction enzymes, including New England BioLabs, Pro-Mega Biochems, Boehringer-Mannheim, and the like. Sambrook et al. (See Sambrook et al. Molecular Biology: A Laboratory Approach, Cold Spring Harbor, N.Y. 1989) provide a general description of methods for using restriction enzymes and other enzymes.
[0101] In some instances, a methylation-dependent restriction enzyme is a restriction enzyme that cleaves or digests DNA at or in proximity to a methylated recognition sequence, but does not cleave DNA at or near the same sequence when the recognition sequence is not methylated. Methylation-dependent restriction enzymes include those that cut at a methylated recognition sequence (e.g., Dpnl) and enzymes that cut at a sequence near but not at the recognition sequence (e.g., McrBC). For example, McrBC’ s recognition sequence is 5’ RmC (N40-3000) RmC 3 ‘ where “R” is a purine and “mC” is a methylated cytosine and “N40-3000” indicates the distance between the two RmC half sites for which a restriction event has been observed. McrBC generally cuts close to one half-site or the other, but cleavage positions are typically distributed over several base pairs, approximately 30 base pairs from the methylated base. McrBC sometimes cuts 3’ of both half sites, sometimes 5’ of both half sites, and sometimes between the two sites. Exemplary methylation-dependent restriction enzymes include, e.g., McrBC, McrA, MrrA, Bisl, Glal and Dpnl. One of skill in the art will appreciate that any methylation-dependent restriction enzyme, including homologs and orthologs of the restriction enzymes described herein, is also suitable for use with one or more methods described herein.
[0102] In some cases, a methylation- sensitive restriction enzyme is a restriction enzyme that cleaves DNA at or in proximity to an unmethylated recognition sequence but does not cleave at or in proximity to the same sequence when the recognition sequence is methylated. Exemplary methylation-sensitive restriction enzymes are described in, e.g., McClelland et al, 22(17)NUCLEIC ACIDS RES. 3640-59 (1994). Suitable methylation- sensitive restriction enzymes that do not cleave DNA at or near their recognition sequence when a cytosine within the recognition sequence is methylated at position C5 include, e.g., Aat II, Aci I, Acd I, Age I, Alu I, Asc I, Ase I, AsiS I, Bbe I, BsaA I, BsaH I, BsiE I, BsiW I, BsrF I, BssH II, BssK I, BstB I, BstN I, BstU I, Cla I, Eae I, Eag I, Fau I, Fse I, Hha I, HinPl I, HinC II, Hpa II, Hpy99 I, HpyCH4 IV, Kas I, Mbo I, Mlu I, MapAl I, Msp I, Nae I, Nar I, Not I, Pml I, Pst I, Pvu I, Rsr II, Sac II, Sap I, Sau3A I, Sfl I, Sfo I, SgrA I, Sma I, SnaB I, Tsc I, Xma I, and Zra I. Suitable methylation-sensitive restriction enzymes that do not cleave DNA at or near their recognition sequence when an adenosine within the recognition sequence is methylated at position N6 include, e.g., Mbo I. One of skill in the art will appreciate that any methylation- sensitive restriction enzyme, including homologs and orthologs of the restriction enzymes described herein, is also suitable for use with one or more of the methods described herein. One of skill in the art will further appreciate that a methylation- sensitive restriction enzyme that fails to cut in the presence of methylation of a cytosine at or near its recognition sequence may be insensitive to the presence of methylation of an adenosine at or near its recognition sequence. Likewise, a methylation-sensitive restriction enzyme that fails to cut in the presence of methylation of an adenosine at or near its recognition sequence may be insensitive to the presence of methylation of a cytosine at or near its recognition sequence. For example, Sau3AI is sensitive (i.e., fails to cut) to the presence of a methylated cytosine at or near its recognition sequence, but is insensitive (i.e., cuts) to the presence of a methylated adenosine at or near its recognition sequence. One of skill in the art will also appreciate that some methylation-sensitive restriction enzymes are blocked by methylation of bases on one or both strands of DNA encompassing of their recognition sequence, while other methylation-sensitive restriction enzymes are blocked only by methylation on both strands, but can cut if a recognition site is hemi-methylated.
[0103] In alternative embodiments, adaptors are optionally added to the ends of the randomly fragmented DNA, the DNA is then digested with a methylation-dependent or methylation-sensitive restriction enzyme, and intact DNA is subsequently amplified using primers that hybridize to the adaptor sequences. In this case, a second step is performed to determine the presence, absence or quantity of a particular gene in an amplified pool of DNA. In some embodiments, the DNA is amplified using real-time, quantitative PCR.
[0104] In other embodiments, the methods comprise quantifying the average methylation density in a target sequence within a population of DNA. In some embodiments, the method comprises contacting DNA with a methylation-dependent restriction enzyme or methylationsensitive restriction enzyme under conditions that allow for at least some copies of potential restriction enzyme cleavage sites in the locus to remain uncleaved; quantifying intact copies of the locus; and comparing the quantity of amplified product to a control value representing the quantity of methylation of control DNA, thereby quantifying the average methylation density in the locus compared to the methylation density of the control DNA.
[0105] In some instances, the quantity of methylation of a locus of DNA is determined by providing a sample of DNA comprising the locus, cleaving the DNA with a restriction enzyme that is either methylation- sensitive or methylation-dependent, and then quantifying the amount of intact DNA or quantifying the amount of cut DNA at the DNA locus of interest. The amount of intact or cut DNA will depend on the initial amount of DNA containing the locus, the amount of methylation in the locus, and the number (i.e., the fraction) of nucleotides in the locus that are methylated in the DNA. The amount of methylation in a DNA locus can be determined by comparing the quantity of intact DNA or cut DNA to a control value representing the quantity of intact DNA or cut DNA in a similarly-treated DNA sample. The control value can represent a known or predicted number of methylated nucleotides. Alternatively, the control value can represent the quantity of intact or cut DNA from the same locus in another (e.g., normal, nondiseased) cell or a second locus.
[0106] By using at least one methylation-sensitive or methylation-dependent restriction enzyme under conditions that allow for at least some copies of potential restriction enzyme cleavage sites in the locus to remain uncleaved and subsequently quantifying the remaining intact copies and comparing the quantity to a control, average methylation density of a locus can be determined. If the methylation- sensitive restriction enzyme is contacted to copies of a DNA locus under conditions that allow for at least some copies of potential restriction enzyme cleavage sites in the locus to remain uncleaved, then the remaining intact DNA will be directly proportional to the methylation density, and thus may be compared to a control to determine the relative methylation density of the locus in the sample. Similarly, if a methylation-dependent restriction enzyme is contacted to copies of a DNA locus under conditions that allow for at least some copies of potential restriction enzyme cleavage sites in the locus to remain uncleaved, thenthe remaining intact DNA will be inversely proportional to the methylation density, and thus may be compared to a control to determine the relative methylation density of the locus in the sample. Such assays are disclosed in, e.g., U.S. Patent No. 7,910,296.
[0107] The methylated CpG island amplification (MCA) technique is a method that can be used to screen for altered methylation patterns in DNA, and to isolate specific sequences associated with these changes (Toyota et al, 1999, Cancer Res. 59, 2307-2312, U.S. Pat. No. 7,700,324 (Issa et al)). Briefly, restriction enzymes with different sensitivities to cytosine methylation in their recognition sites are used to digest DNAs from primary tumors, cell lines, and normal tissues prior to arbitrarily primed PCR amplification. Fragments that show differential methylation are cloned and sequenced after resolving the PCR products on high- resolution polyacrylamide gels. The cloned fragments are then used as probes for Southern analysis to confirm differential methylation of these regions. Typical reagents (e.g., as might be found in a typical MCA-based kit) for MCA analysis may include, but are not limited to: PCR primers for arbitrary priming DNA; PCR buffers and nucleotides, restriction enzymes and appropriate buffers; gene-hybridization oligos or probes; control hybridization oligos or probes.
[0108] In some embodiments, the methods provided herein further comprise performing the sequencing assay. In some embodiments, the methods provided herein further comprise performing a non-disruptive methylation sequencing technique. In some embodiments, the non- disruptive methylation sequencing technique is an enzymatic methyl-seq (EM-seq) technique. In some embodiments, the non-disruptive methylation sequencing technique comprises: (a) enzymatically modifying methylated cytosines (such as 5-methylcytosine (5 me) and 5- hydroxymethylcytosine (5 hmC)) to prevent deamination in further enzymatic steps; (b) enzymatically converting unmethylated cytosines to uracils; (c) performing PCR amplification (thereby converting uracils to thymines; and (d) sequencing using a NGS technique. Various techniques for performing a non-disruptive methylation sequencing technique have been described in the art. See, e.g., Vaisvila et al., Genome Res, 31, 2021, which is incorporated herein in its entirety. In some embodiments, enzymatically modifying methylated cytosines is performed using TET2 and / or T4-BGT. In some embodiments, the non-disruptive methylation sequencing technique comprises enzymatically converting unmethylated cytosines to uracil using APOBEC3A. In some embodiments, the non-disruptive methylation sequencing techniquecomprises subjecting a sample comprising DNA, such as a cfDNA sample, to a NGS library preparation technique.
[0109] In some embodiments, the methods provided herein comprise indirect methylation sequencing methods which detects unmethylated cytosines by C-to-T transition through processing a cfDNA sample extracted from the biological sample with an agent or agents to convert unmethylated cytosines to uracil which are then read out as thymine at the sequencing level. In some embodiments, the method provided herein comprise direct methylation sequencing methods which detects methylated cytosines by C-to-T transition through processing a cfDNA sample extracted from the biological sample with an agent or agents to convert methylated cytosines to thymine, e..g, Wang, T., Fowler, J.M., Liu, L. et al. Direct enzymatic sequencing of 5-methylcytosine at single-base resolution. Nat Chem Biol 19, 1004-1012 (2023). http s : / / doi . org / 10.1038 / s41589-023-01318 - 1 , which is hereby incorporated herein by reference in its entirety.
[0110] In some embodiments, the NGS library preparation technique comprises shearing the DNA, such as to obtain a DNA size of less than about 500 base pairs, such as less than about any of 450 base pairs, 400 base pairs, 350 base pairs, or 300 base pairs. In some embodiments, the NGS library preparation technique comprises a step of end prep of sheared DNA. In some embodiments, the NGS library preparation technique comprises a step of adaptor ligation. In some embodiments, the NGS library preparation technique comprises a step of cleaning up adaptor ligated DNA. In some embodiments, the cleaned and ligated DNA is subjected to oxidative enzymes, such as TET2, and / or glucosyltransferase, such as T4-BGT, to modify methylated cytosines (5-methylcytosines and 5-hydroxymethylcytosines). In some embodiments, the NGS library preparation technique comprises a step of cleaning enzyme oxidized and / or glucosylated DNA. In some embodiments, the oxidized and / or glucosylated DNA is further subjected to enzymatic cytosine deamination (such as using APOBEC3A). In some embodiments, the NGS library preparation technique comprises a step of PCR amplification of the deaminated DNA. In some embodiments, the PCR amplification comprises the use of dualindexed primer pairs. In some embodiments, the NGS library preparation technique comprises a step of pooling two or more DNA samples that have been amplified with orthogonal dualindexed primer pairs. In some embodiments, the pooling comprises isolating target DNAs. Insome embodiments, the isolating comprises a hybridization-based target enrichment protocol. In some embodiments, the NGS library preparation technique comprises a step of sequencing and quantification. In some embodiments, the method comprises adding a control to the sample comprising DNA, e.g., prior to performing any enzymatic conversion steps.
[0111] In some embodiments, the step of sequencing comprises NGS. Thus, in some embodiments, the step of sequencing comprises loading an enzymatically converted library into an NGS platform (e.g., an NGS performing hardware) to generate a loaded library. In some embodiments, the step of sequencing comprises binding the loaded libraries to a flow cell, such as a microfluidic chamber equipped for both enzymatic reactions and imaging wherein oligonucleotides are bound to a solid support. In some embodiments, the binding comprises annealing an adapter region of a single- stranded DNA from the library to a complementary oligonucleotide that is bound to the flow cell. In some embodiments, the step of sequencing comprises amplifying the bound, single- stranded DNA (e.g., bridge amplification). In some embodiments, the amplified, bound DNA comprises a first strand bound to the solid support of the flow cell wherein all of the first strand are bound to the flow cell from the same adapter sequence. In some embodiments, the amplified, bound DNA comprises a second strand bound to the solid support of the flow cell wherein the second strand is the reverse complement of the first strand. In some embodiments, the step of sequencing comprises: (i) cleaving the second strands of the amplified, bound DNA from the solid support of the flow cell; (ii) washing away the cleaved second strands; and, (iii) sequencing the first strands of the amplified, bound DNA. In some embodiments, the step of sequencing comprises (i) cleaving the first strands of the amplified, bound, DNA from the solid support of the flow cell; (ii) washing away the cleaved first strands; and, (iii) sequencing the second strands.
[0112] Additional methylation detection methods include those methods described in, e.g., U.S. Patents No. 7,553,627; No. 6,331,393; U.S. Patent Serial No. 12 / 476,981; U.S. Patent Publication No. 2005 / 0069879; Rein, et al, 26(10) NUCLEIC ACIDS RES. 2255-64 (1998); and Olek et al, 17(3) NAT. GENET. 275-6 (1997).
[0113] In another embodiment, the methylation status of selected CpG sites is determined using Methylation-Sensitive High Resolution Melting (HRM). Recently, Wojdacz et al. reported methylation- sensitive high resolution melting as a technique to assess methylation. (Wojdacz and Dobrovic, 2007, Nuc. Acids Res. 35(6) e41; Wojdacz et al. 2008, Nat. Prot. 3(12)1903-1908; Balic et al, 2009 J. Mol. Diagn. 11 102- 108; and US Pat. Pub. No. 2009 / 0155791 (Wojdacz et al)). A variety of commercially available real time PCR machines have HRM systems including the Roche LightCycler480, Corbett Research RotorGene6000, and the Applied Biosystems 7500. HRM may also be combined with other amplification techniques such as pyrosequencing as described by Candiloro et al. (Candiloro et al, 2011, Epigenetics 6(4) 500-507).
[0114] In another embodiment, the methylation status of selected CpG locus is determined using a primer extension assay, including an optimized PCR amplification reaction that produces amplified targets for analysis using mass spectrometry. The assay can also be done in multiplex. Mass spectrometry is a particularly effective method for the detection of polynucleotides associated with the differentially methylated regulatory elements. The presence of the polynucleotide sequence is verified by comparing the mass of the detected signal with the expected mass of the polynucleotide of interest. The relative signal strength, e.g., mass peak on a spectra, for a particular polynucleotide sequence indicates the relative population of a specific allele, thus enabling calculation of the allele ratio directly from the data. This method is described in detail in PCT Pub. No. WO 2005 / 012578A1 (Beaulieu et al), which is hereby incorporated by reference in its entirety. For methylation analysis, the assay can be adopted to detect bisulfite introduced methylation dependent C to T sequence changes. These methods are particularly useful for performing multiplexed amplification reactions and multiplexed primer extension reactions (e.g., multiplexed homogeneous primer mass extension (hME) assays) in a single well to further increase the throughput and reduce the cost per reaction for primer extension reactions.
[0115] Other methods for DNA methylation analysis include restriction landmark genomic scanning (RLGS, Costello et al, 2002, Meth. Mol Biol, 200, 53-70), methylation- sensitive- representational difference analysis (MS-RDA, Ushijima and Yamashita, 2009, Methods Mol Biol 507, 1 17-130). Comprehensive high-throughput arrays for relative methylation (CHARM) techniques are described in WO 2009 / 021141 (Feinberg and Irizarry). The Roche(R) NimbleGen(R) microarrays including the Chromatin Immunoprecipitation-on- chip (ChlP-chip) or methylated DNA immunoprecipitation-on-chip (MeDIP-chip). These tools have been used for a variety of cancer applications including melanoma, liver cancer and lung cancer (Koga et al, 2009, Genome Res., 19, 1462-1470; Acevedo et al, 2008, Cancer Res., 68, 2641-2651; Rauchet al, 2008, Proc. Nat. Acad. Sci. USA, 105, 252-257). Others have reported bisulfate conversion, padlock probe hybridization, circularization, amplification and next generation or multiplexed sequencing for high throughput detection of methylation (Deng et al, 2009, Nat. Biotechnol 27, 353-360; Ball et al, 2009, Nat. Biotechnol 27, 361-368; U.S. Pat. No. 7,611,869 (Fan)). As an alternative to bisulfate oxidation, Bayeyt et al. have reported selective oxidants that oxidize 5-methylcytosine, without reacting with thymidine, which are followed by PCR or pyro sequencing (WO 2009 / 049916 (Bayeyt et al).
[0116] In some instances, quantitative amplification methods (e.g., quantitative PCR or quantitative linear amplification) are used to quantify the amount of intact DNA within a locus flanked by amplification primers following restriction digestion. Methods of quantitative amplification are disclosed in, e.g., U.S. Patents No. 6, 180,349; No. 6,033,854; and No. 5,972,602, as well as in, e.g., DeGraves, et al, 34(1) BIOTECHNIQUES 106-15 (2003); Deiman B, et al., 20(2) MOL. BIOTECHNOL. 163-79 (2002); and Gibson et al, 6 GENOME RESEARCH 995-1001 (1996).
[0117] Following reaction or separation of nucleic acid in a methylation specific manner, the nucleic acid in some cases are subjected to sequence-based analysis. For example, once it is determined that one particular genomic sequence from a sample is hypermethylated or hypomethylated compared to its counterpart, the amount of this genomic sequence can be determined. Subsequently, this amount can be compared to a standard control value and used to determine the present of liver cancer in the sample. In many instances, it is desirable to amplify a nucleic acid sequence using any of several nucleic acid amplification procedures which are well known in the art. Specifically, nucleic acid amplification is the chemical or enzymatic synthesis of nucleic acid copies which contain a sequence that is complementary to a nucleic acid sequence being amplified (template). The methods and kits may use any nucleic acid amplification or detection methods known to one skilled in the art, such as those described in U.S. Pat. Nos. 5,525,462 (Takarada et al); 6,114,117 (Hepp et al); 6,127,120 (Graham et al); 6,344,317 (Umovitz); 6,448,001 (Oku); 6,528,632 (Catanzariti et al); and PCT Pub. No. WO 2005 / 111209 (Nakajima et al).
[0118] In some embodiments, the nucleic acids are amplified by PCR amplification using methodologies known to one skilled in the art. One skilled in the art will recognize, however, that amplification can be accomplished by any known method, such as ligase chain reaction(LCR), Q -replicas amplification, rolling circle amplification, transcription amplification, selfsustained sequence replication, nucleic acid sequence-based amplification (NASBA), each of which provides sufficient amplification. Branched-DNA technology is also optionally used to qualitatively demonstrate the presence of a sequence of the technology, which represents a particular methylation pattern, or to quantitatively determine the amount of this particular genomic sequence in a sample. Nolte reviews branched-DNA signal amplification for direct quantitation of nucleic acid sequences in clinical samples (Nolte, 1998, Adv. Clin. Chem. 33:201-235).
[0119] The PCR process is well known in the art and include, for example, reverse transcription PCR, ligation mediated PCR, digital PCR (dPCR), or droplet digital PCR (ddPCR). For a review of PCR methods and protocols, see, e.g., Innis et al, eds., PCR Protocols, A Guide to Methods and Application, Academic Press, Inc., San Diego, Calif. 1990; U.S. Pat. No. 4,683,202 (Mullis). PCR reagents and protocols are also available from commercial vendors, such as Roche Molecular Systems. In some instances, PCR is carried out as an automated process with a thermostable enzyme. In this process, the temperature of the reaction mixture is cycled through a denaturing region, a primer annealing region, and an extension reaction region automatically. Machines specifically adapted for this purpose are commercially available.
[0120] In some embodiments, amplified sequences are also measured using invasive cleavage reactions such as the Invader(R) technology (Zou et al, 2010, Association of Clinical Chemistry (AACC) poster presentation on July 28, 2010, “Sensitive Quantification of Methylated Markers with a Novel Methylation Specific Technology; and U.S. Pat. No. 7,011,944 (Prudent et al)).
[0121] Suitable NGS technologies are widely available. Examples include the 454 Life Sciences platform (Roche, Branford, CT) (Margulies et al. 2005 Nature, 437, 376-380);Illumina’s Genome Analyzer, GoldenGate Methylation Assay, or Infinium Methylation Assays, i.e., Infinium HumanMethylation 27K BeadArray or VeraCode GoldenGate methylation array (Illumina, San Diego, CA; Bibkova et al, 2006, Genome Res. 16, 383-393; U.S. Pat. Nos. 6,306,597 and 7,598,035 (Macevicz); 7,232,656 (Balasubramanian et al.)); QX200™ Droplet Digital™ PCR System from Bio-Rad; or DNA Sequencing by Ligation, SOLiD System (Applied Biosystems / Life Technologies; U.S. Pat. Nos. 6,797,470, 7,083,917, 7,166,434, 7,320,865, 7,332,285, 7,364,858, and 7,429,453 (Barany et al); the Helicos True SingleMolecule DNA sequencing technology (Harris et al, 2008 Science, 320, 106-109; U.S. Pat. Nos. 7,037,687 and 7,645,596 (Williams et al); 7, 169,560 (Lapidus et al); 7,769,400 (Harris)), the single molecule, real-time (SMRT™) technology of Pacific Biosciences, and sequencing (Soni and Meller, 2007, Clin. Chem. 53, 1996-2001); semiconductor sequencing (Ion Torrent; Personal Genome Machine); DNA nanoball sequencing; sequencing using technology from Dover Systems (Polonator), and technologies that do not require amplification or otherwise transform native DNA prior to sequencing (e.g., Pacific Biosciences and Helicos), such as nanopore-based strategies (e.g., Oxford Nanopore, Genia Technologies, and Nabsys). These systems allow the sequencing of many nucleic acid molecules isolated from a specimen at high orders of multiplexing in a parallel fashion. Each of these platforms allow sequencing of clonally expanded or non-amplified single molecules of nucleic acid fragments. Certain platforms involve, for example, (i) sequencing by ligation of dye- modified probes (including cyclic ligation and cleavage), (ii) pyrosequencing, and (iii) single-molecule sequencing.
[0122] Pyrosequencing is a nucleic acid sequencing method based on sequencing by synthesis, which relies on detection of a pyrophosphate released on nucleotide incorporation. Generally, sequencing by synthesis involves synthesizing, one nucleotide at a time, a DNA strand complimentary to the strand whose sequence is being sought. Study nucleic acids may be immobilized to a solid support, hybridized with a sequencing primer, incubated with DNA polymerase, ATP sulfurylase, luciferase, apyrase, adenosine 5’ phosphsulfate and luciferin. Nucleotide solutions are sequentially added and removed. Correct incorporation of a nucleotide releases a pyrophosphate, which interacts with ATP sulfurylase and produces ATP in the presence of adenosine 5’ phosphsulfate, fueling the luciferin reaction, which produces a chemiluminescent signal allowing sequence determination. Machines for pyrosequencing and methylation specific reagents are available from Qiagen, Inc. (Valencia, CA). See also Tost and Gut, 2007, Nat. Prot. 2 2265-2275. An example of a system that can be used by a person of ordinary skill based on pyro sequencing generally involves the following steps: ligating an adaptor nucleic acid to a study nucleic acid and hybridizing the study nucleic acid to a bead; amplifying a nucleotide sequence in the study nucleic acid in an emulsion; sorting beads using a picoliter multiwell solid support; and sequencing amplified nucleotide sequences by pyrosequencing methodology (e.g., Nakano et al, 2003, J. Biotech. 102, 117-124). Such a system can be used to exponentially amplify amplification products generated by a processdescribed herein, e.g., by ligating a heterologous nucleic acid to the first amplification product generated by a process described herein.
[0123] In some embodiments, the sequence read data contains information from a single individual. In some embodiments, the sequence read data from a single individual does not require demultiplexing.
[0124] In some embodiments, the sequence read data contains information from a plurality of individuals that are uniquely indexed. In some embodiments, the sequence read data from a plurality of individuals is demultiplexed. In some embodiments, the demultiplexing of the sequence read data occurs prior to receiving the sequence read data at one or more processors. In some embodiments, the demultiplexing generates independent sequence read data for each individual of the plurality of individuals.
[0125] In some embodiments, the sequence read data is processed. In some embodiments, the processing comprises applying one or more bioinformatics processes. In some embodiments, the processing comprises removing duplicate sequence read data. In some embodiments, the processing comprises removing sequence read data that does not meet a predetermined quality threshold. In some embodiments, the processing occurs before the sequence read data is received at the one or more processors.
[0126] In some embodiments, the processing comprises trimming the sequence read data to remove adapter or primer sequences prior to alignment to a reference genome. In some embodiments, the reference genome is a human reference genome. In some embodiments, the human reference genome is hg38.
[0127] In some embodiments, the processing comprises filtering the sequence read data. In some embodiments, the filtering occurs after the sequence read is trimmed and processed. In some embodiments, the filtering comprises removing non-primary alignments, supplementary alignments, and alignments that fail a platform or a vendor quality control check. In some embodiments, the filtering comprises flagging duplicate sequence read data. In some embodiments, the flagged, duplicate sequence read data may be removed.
[0128] In some embodiments, the processing comprises using filtered, and optionally duplicate-flagged, sequence read data to calculate a NGS QC metric for the sequence read data. In some embodiments, the processing comprises using filtered, and optionally duplicate-flagged,sequence read data to calculate CpG site level and read level methylation metrics (e.g., a methylation fraction value and / or methylation variance value). In some embodiments, the duplicate sequence read data is removed from the sequence read data before it is received at the one or more processors. In some embodiments, the sequence read data that falls below a predetermined threshold for the NGS QC metric is removed from the sequence read data before it is received at the one or more processors.D. Methylation cancer score
[0129] In some aspects, the provided methods may be used to determine a methylation cancer score of an individual suspected of having a liver cancer. A methylation score may be an output of a machine learning model (e.g., an ensemble machine learning model) whose input was a methylation fraction score and a methylation variance score. The methylation cancer score can be subsequently used to determine a liver cancer status of the individual, which is based on the methylation cancer score and a predetermined methylation cancer score threshold. The individual may be determined to have a liver-cancer positive status if the methylation cancer score is greater than the predetermined methylation cancer score threshold.
[0130] In some embodiments, provided herein is a method for determining a methylation score of an individual suspected of having a liver cancer, the method comprising: receiving, at one or more processors, sequence read data; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; and inputting, using the one or more processors, the methylation fraction score and themethylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual.
[0131] In some embodiments, the method for determining a methylation score of an individual suspected of having a liver cancer comprises: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; processing, at the one or more processors, extracted the methylation fraction value for each of the one or more first genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; processing, at the one or more processors, extracted the methylation variance value for each of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; and inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual. In some embodiments, the processing comprises inputting missing data, feature selection, data transformation, and / or data scaling.
[0132] In some embodiments, the method for determining a methylation score of an individual suspected of having a liver cancer comprises: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance of inthe number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual; and determining a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments, the liver cancer status is a liver cancer-positive status when the methylation cancer score is greater than the predetermined methylation cancer score threshold.
[0133] When performing the subject methods, although a sequence read may be associated with a genomic region based the sequence read having a single overlapping base with the genomic region, when assessing methylation fraction values and methylation variance values, the entire sequence read is evaluated. Thus, in some embodiments a methylation fraction value and / or a methylation variance value, may include data, such as of a CpG site, from outside the genomic region. z. Methylation fraction
[0134] In some aspects, a methylation fraction value is based on a ratio of methylated sequence reads associated with a genomic region of one or more first genomic regions and total reads associated with the one or more first genomic regions. An extracted methylation fraction value may be inputted into a first machine learning model, using one or more processors, to output a methylation fraction score. The methylation fraction score can then be used as part of an input for an ensemble machine learning model which outputs a methylation cancer score (e.g., based on the methylation fraction score and a methylation variance score).
[0135] In some embodiments, the methylation fraction value is extracted from each of one or more first genomic regions, using one or more processors and sequence read data obtained from a sample from an individual. In some embodiments, the extracted methylation fraction value isbased on a ratio of methylated sequence reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions. In some embodiments, the extracted methylation fraction value for each of the one or more first genomic regions is processed, using one or more processors, before inputting into the first trained machine learning model. In some embodiments, the processing comprises inputting missing data, feature selection, data transformation, and / or data scaling. In some embodiments, the extracted methylation fraction value for each of the one or more first genomic regions is inputted, using one or more processors, into a first trained machine learning model to output a methylation fraction score.
[0136] In some embodiments, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with any of the one or more first genomic regions. In some embodiments, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with any of the one or more first genomic regions, such as any of at least two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more, bases overlapping with any of the one or more first genomic regions. In some embodiments, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having less than 20 bases overlapping with any of the one or more first genomic regions, such as less than any of 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, nine, eight, seven, six, five, four, three, or two bases overlapping with any of the one or more first genomic regions. In some embodiments, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having about any of one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 bases overlapping with any of the one or more first genomic regions.
[0137] In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with the genomic region of the one or more first genomic regions. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with the genomic region of the one or more first genomic regions, such as any of at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20,25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, or more, bases overlapping with the genomic region of the one or more first genomic regions. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having less than 200 bases overlapping with the genomic region of the one or more first genomic regions, such as less than any of 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, nine, eight, seven, six, five, four, three, or two bases overlapping with the genomic region of the one or more first genomic regions. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having about any of one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases overlapping with the genomic region of the one or more first genomic regions.
[0138] In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having at least one methylated CpG site. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having at least one methylated CpG site, such as at least any of two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more, methylated CpG sites. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having less than 20 methylated CpG sites, such as less than any of 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, nine, eight, seven, six, five, four, three, or two methylated CpG sites. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having any of about one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 methylated CpG sites.
[0139] In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having at least 50% of total CpG sites having a methylation. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of readsfrom the sequence read data having at least 50% of total CpG sites having a methylation, such as any of at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more of total CpG sites having a methylation. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having less than 99% of total CpG sites having a methylation, such as less than any of 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, or less, of total CpG sites having a methylation. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having any of about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% of total CpG sites having a methylation.
[0140] In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having: (i) between about one and about 200 bases overlapping with the genomic region of the one or more first genomic regions, (ii) between about one and about 20 methylated CpG sites, and (iii) between at least 50% and at least 99% of total CpG sites having a methylation. In some embodiments, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from the sequence read data having: (i) at least one base overlapping with the genomic region of the one or more first genomic regions, (ii) at least one methylated CpG site, and (iii) at least 50% of total CpG sites having a methylation. ii. Methylation variance
[0141] In some aspects, a methylation variance value is based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of one or more second genomic regions. An extracted methylation variance value may be inputted into a second machine learning model, using one or more processors, to output a methylation variance score. The methylation variance score can then be used as part of an input for an ensemble machine learning model which outputs a methylation cancer score (e.g., based on the methylation variance score and a methylation fraction score).
[0142] In some embodiments, the methylation variance value is extracted from each of one or more second genomic regions, using one or more processors and sequence read data obtainedfrom a sample from an individual. In some embodiments, the extracted methylation variance value is based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions. In some embodiments, the extracted methylation variance value for each of the two or more first genomic regions is processed, using one or more processors, before inputting into the second trained machine learning model. In some embodiments, the processing comprises inputting missing data, feature selection, data transformation, and / or data scaling. In some embodiments, the extracted methylation variance value for each of the one or more second genomic regions is inputted, using one or more processors, into a second trained machine learning model to output a methylation variance score.
[0143] In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having at least one base overlapping with the genomic region of the one or more second genomic regions. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having at least one base overlapping with the genomic region of the one or more second genomic regions, such at least any of two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, or more, bases overlapping with the genomic region of the one or more second genomic regions. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having less than 200 bases overlapping with the genomic region of the one or more second genomic regions, such as less than any of 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, nine, eight, seven, six, five, four, three, or two bases overlapping with the genomic region of the one or more second genomic regions. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bases overlapping with the genomic region of the one or more second genomic regions.
[0144] In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having at least one CpG site. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having at least one CpG site, such as any of at least two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 CpG sites. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having less than 20 CpG sites, such as less than any of 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, nine, eight, seven, six, five, four, three, or two CpG sites. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 CpG sites.
[0145] In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having: (i) between at least one base and at least 200 bases overlapping with the genomic region of the one or more second genomic regions, and (ii) between at least one and at least 20 CpG sites. In some embodiments, the reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having: (i) at least one base overlapping with the genomic region of the one or more second genomic regions, and (ii) at least one CpG site.E. Genomic regions
[0146] In any of the provided methods, the sequence read data may comprise genomic regions, such as one or more first genomic regions and one or more second genomic regions. These genomic regions correspond to particular start and end sites within a chromosome of an individual from which the sample is obtained. Each genomic region may contain no CpG sites, a single CpG site, or a plurality of CpG sites. In some examples, the genomic regions are evaluated to determine the methylation fraction value and / or the methylation variance value.
[0147] In some embodiments, the sequence read data comprises sequence data of one or more genomic regions. In some embodiments, the one or more genomic regions do not overlap. In some embodiments, the one or more genomic regions overlap.
[0148] In some embodiments, the sequence read data comprises data of more than one genomic region. In some embodiments, the sequence read data comprises data of between about 100 and about 5,000 genomic regions, such as between about 100 and about 2,000 genomic regions, between about 1,000 and about 3,000 genomic regions, between about 2,000 and about 4,000 genomic regions, or between about 3,000 and about 5,000 genomic regions. In some embodiments, the sequence read data comprises data of at least about 100 genomic regions, such as data of at least about any of 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, or more genomic regions. In some embodiments, the sequence read data comprises data of less than about 5,000 genomic regions, such as data of less than about any of 4,500, 4,000, 3,500, 3,000, 2,500, 2,000, 1,500, 1,000, 950, 900, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, 250, 200, 150, 100, or fewer genomic regions. In some embodiments, the sequence read data comprises data of 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, or 5,000 genomic regions.
[0149] In some embodiments, the one or more genomic regions are each individually the same length. In some embodiments, the one or more genomic regions comprise genomic regions of different lengths. In some embodiments, the one or more genomic regions are each individually between about 50 and about 500 base pairs in length, such as between about 50 and about 150 base pairs in length, between about 100 and about 200 base pairs in length, between about 150 and about 250 base pairs in length, between about 200 and about 3000 base pairs in length, between about 250 and about 350 base pairs in length, between about 300 and about 400 base pairs in length, between about 350 and about 450 base pairs in length, or between about 400 and about 500 base pairs in length. In some embodiments, the one or more genomic regions are each individually at least about 50 base pairs in length, such at least about any of 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, or more base pairs in length. In some embodiments, the one or more genomic regions are each individually less than about 500 base pairs in length, such as less than about any of 490, 480, 470, 460, 450, 440, 430, 420, 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, or fewer base pairs in length. In some embodiments,the one or more genomic regions are each individually 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs in length. In some embodiments, the one or more genomic regions are each individually 100 or 200 base pairs in length. In some embodiments, the one or more genomic regions are each individually 100 base pairs in length. In some embodiments, the one or more genomic regions are each individually 200 base pairs in length.
[0150] In some embodiments, the one or more genomic regions (e.g., one or more first genomic regions and / or one or more second genomic regions) are used to extract a feature. For example, in an embodiment, the one or more genomic regions may be used to extract a methylation fraction value. In some embodiments, one or more first genomic regions are used to extract a methylation fraction value. In another embodiments, the one or more genomic regions can be used to extract a methylation variance value. In some embodiments, one or more second genomic regions are used to extract a methylation variance value.
[0151] In some embodiments, the one or more genomic regions are one or more first genomic regions. In some embodiments, a method provided herein comprises extracting for each of one or more first genomic regions, using one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions.
[0152] In some embodiments, the sequence read data comprises data of more than one first genomic region. In some embodiments, the sequence read data comprises data of between about 50 and about 2,500 first genomic regions, such as between about 50 and about 500 first genomic regions, between about 250 and about 1,000 first genomic regions, between about 500 and about 2,000 first genomic regions, or between about 1,000 and about 2,500 first genomic regions. In some embodiments, the sequence read data comprises data of at least about 50 first genomic regions, such as data of at least about any of 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, 2,500, or more first genomic regions. In some embodiments, the sequence read data comprises data of less than about 2,500 first genomic regions, such as data of less than about any of 2,000, 1,500, 1,000, 950, 900, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, 250, 200, 150, 100, or fewer firstgenomic regions. In some embodiments, the sequence read data comprises data of 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, or 2,500 first genomic regions.
[0153] In some embodiments, the one or more first genomic regions are each individually the same length. In some embodiments, the one or more first genomic regions comprise genomic regions of different lengths. In some embodiments, the one or more first genomic regions are each individually between about 50 and about 500 base pairs in length, such as between about 50 and about 150 base pairs in length, between about 100 and about 200 base pairs in length, between about 150 and about 250 base pairs in length, between about 200 and about 3000 base pairs in length, between about 250 and about 350 base pairs in length, between about 300 and about 400 base pairs in length, between about 350 and about 450 base pairs in length, or between about 400 and about 500 base pairs in length. In some embodiments, the one or more first genomic regions are each individually at least about 50 base pairs in length, such at least about any of 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, or more base pairs in length. In some embodiments, the one or more first genomic regions are each individually less than about 500 base pairs in length, such as less than about any of 490, 480, 470, 460, 450, 440, 430, 420, 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, or fewer base pairs in length. In some embodiments, the one or more first genomic regions are each individually 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs in length. In some embodiments, the one or more first genomic regions are each individually 100 or 200 base pairs in length. In some embodiments, the one or more first genomic regions are each individually 100 base pairs in length.
[0154] In some embodiments, the one or more genomic regions are one or more second genomic regions. In some embodiments, a method provided herein comprises extracting for each of one or more second genomic regions, using one or more processors and the sequence read data, a methylation variance value based on the variance of in the number of methylated CpGsites per read across reads associated with a genomic region of the one or more second genomic regions.
[0155] In some embodiments, the sequence read data comprises data of more than one second genomic region. In some embodiments, the sequence read data comprises data of between about 50 and about 2,500 second genomic regions, such as between about 50 and about 500 second genomic regions, between about 250 and about 1,000 second genomic regions, between about 500 and about 2,000 second genomic regions, or between about 1,000 and about 2,500 second genomic regions. In some embodiments, the sequence read data comprises data of at least about 50 second genomic regions, such as data of at least about any of 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, 2,500, or more second genomic regions. In some embodiments, the sequence read data comprises data of less than about 2,500 second genomic regions, such as data of less than about any of 2,000, 1,500, 1,000, 950, 900, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, 250, 200, 150, 100, or fewer second genomic regions. In some embodiments, the sequence read data comprises data of 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, or 2,500 second genomic regions.
[0156] In some embodiments, the one or more second genomic regions are each individually the same length. In some embodiments, the one or more second genomic regions comprise genomic regions of different lengths. In some embodiments, the one or more second genomic regions are each individually between about 50 and about 500 base pairs in length, such as between about 50 and about 150 base pairs in length, between about 100 and about 200 base pairs in length, between about 150 and about 250 base pairs in length, between about 200 and about 3000 base pairs in length, between about 250 and about 350 base pairs in length, between about 300 and about 400 base pairs in length, between about 350 and about 450 base pairs in length, or between about 400 and about 500 base pairs in length. In some embodiments, the one or more second genomic regions are each individually at least about 50 base pairs in length, such at least about any of 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, or more base pairs in length. In some embodiments, the one or more second genomic regions are each individually less than about 500 base pairs in length, such as less than about any of 490, 480, 470, 460, 450, 440, 430, 420, 410,400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, or fewer base pairs in length. In some embodiments, the one or more second genomic regions are each individually 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 215, 220, 230, 240, 250, 260, 270, 280, 290, 300. 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 base pairs in length. In some embodiments, the one or more second genomic regions are each individually 100 or 200 base pairs in length. In some embodiments, the one or more second genomic regions are each individually 200 base pairs in length.
[0157] In some embodiments, the sequence read data comprises data of one or more first genomic regions and one or more second genomic regions. In some embodiments, the one or more first genomic regions do not overlap with the one or more second genomic regions. In some embodiments, the one or more first genomic regions do overlap with the one or more second genomic regions. In some embodiments, the sequence read data comprises data of between about 500 and about 2,500 first genomic regions and data of between about 500 and about 2,500 second genomic regions. In some embodiments, the sequence read data comprises data of at least about 900 first genomic regions and data of at least about 900 second genomic regions. In some embodiments, the sequence read data comprises data of about 900 first genomic regions and data of about 900 second genomic regions. In some embodiments, the one or more fist genomic regions and the one or more second genomic regions are each individually between about 50 base pairs and about 500 base pairs in length. In some embodiments, the one or more first genomic regions and the one or more second genomic regions are each individually 100 or 200 base pairs in length. In some embodiments, the one or more first genomic regions are each individually 100 base pairs in length, and wherein the one or more second genomic regions are each individually 200 base pairs in length.
[0158] A variety of first genomic regions and second genomic regions are compatible for use in the methods provided herein. Exemplary genomic regions (e.g., exemplary first genomic regions and / or second genomic regions) are provided in Table 1 below. In some embodiments, the one or more first genomic regions or second genomic regions individually comprise any of the genomic regions listed in Table 1. In Table 1, the position numbers refer to a chromosome start and end site that corresponds to a genomic region.Table 1: Genomic regions useful for the provided invention.F. Demographic-protein cancer score
[0159] A method for determining a methylation cancer score of an individual suspected of having a liver cancer, as provided herein, may further comprise determining a demographicprotein cancer score of the individual. The demographic-protein cancer score may be used alone, or in tandem with the methylation cancer score, to determine a liver cancer status of the individual. Thus, in some aspects, provided herein is a method of determining a liver cancer status of an individual based on a methylation cancer score and demographic-protein cancer score of the individual. The demographic-protein cancer score may comprise methods as provided in Johnson P et al. Cancer Epidemiol Biomarkers Prev. 2014;23:144-53.
[0160] In some embodiments, the method comprises receiving demographic information of the individual suspected of having a liver cancer. Demographics are statistics that describe populations and their characteristics. Demographic information refers to socioeconomic information expressed statistically, including sex, age, race, geographic location, employment status, level of education, income level, marriage status, and more. Demographic information of the individual may therefore come in many forms. For example, in some embodiments, the demographic information of the individual comprises information regarding the sex of the individual, the age of the individual, the race of the individual, the geographic location of the individual, the employment status of the individual, the level of education of the individual, the income level of the individual, and / or the marriage status of the individual. In someembodiments, the demographic information comprises one or more pieces of demographic information. In some embodiments, the demographic information comprises information regarding the sex of the individual. In some embodiments, the demographic information comprises information regarding the age of the individual. In some embodiments, the demographic information comprises information regarding the sex of the individual and the age of the individual.
[0161] In some embodiments, the method comprises receiving protein information from the individual suspected of having a liver cancer. While a wide range of proteins may be employed as liver cancer protein markers, the liver protein markers employed in many embodiments of the instant methods include proteins selected from the group consisting of: AFP, lens culinaris agglutinin-reactive AFP (AFP-L3), des-gamma carboxy prothrombin (DCP), osteopontin, midkine (MDK), dikkopf-1 (DKK1), glypican-3 (GPC-3), alpha- 1 fucosidase (AFU), and golgi protein-73 (GP-73).
[0162] The liver protein markers may be expressed as a concentration of the protein in the individual and / or a percentage of the protein in the individual. In some embodiments, the AFP level is a concentration of AFP. In some embodiments, the concentration of AFP is a concentration in the blood of the individual, such as reported in ng / mL. In some embodiments, the AFP-L3 level is a percentage (AFP-L3%) based on the ratio of AFP-L3 to total AFP. In some embodiments, the DCP level is a concentration of DCP. In some embodiments, the concentration of DCP is a concentration in the blood of the individual, such as reported in ng / mL.
[0163] Data regarding protein information can be obtained from a variety of techniques. In some embodiments, the protein information is based on respective serum concentrations measured from the individual. In some embodiments, the serum concentrations are derived from the sample obtained from the individual. In some embodiments, serum concentrations of AFP, AFP-L3%, and DCP are measured by using commercially available assays.
[0164] In some embodiments, the method comprises receiving protein information regarding AFP level in the individual, AFP-L3% in the individual, DCP level in the individual, MDK level in the individual, DKK1 level in the individual, GPC-3 level in the individual, AFU level in the individual, and GP-73 level in the individual. In some embodiments, the method comprisesreceiving protein information regarding AFP level in the individual. In some embodiments, the method comprises receiving protein information regarding AFP-L3% in the individual. In some embodiments, the method comprises receiving protein information regarding DCP level in the individual. In some embodiments, the method comprises receiving protein information regarding AFP concentration, AFP-L3%, and DCP concentration.
[0165] In some embodiments, the method comprises obtaining the protein information regarding AFP level in the individual, AFP-L3% in the individual, DCP level in the individual, MDK level in the individual, DKK1 level in the individual, GPC-3 level in the individual, AFU level in the individual, and GP-73 level in the individual. In some embodiments, the method comprises obtaining protein information regarding AFP level in the individual. In some embodiments, the method comprises obtaining protein information regarding AFP-L3% in the individual. In some embodiments, the method comprises obtaining protein information regarding DCP level in the individual. In some embodiments, the method comprises obtaining protein information regarding AFP concentration, AFP-L3%, and DCP concentration.
[0166] In some embodiments, the method comprises receiving, at one or more processors, demographic and protein information from the individual suspected of having a liver cancer. In some embodiments, the demographic information comprises information regarding the sex of the individual and the age of the individual. In some embodiments, the protein information comprises information regarding AFP concentration, AFP-L3%, and DCP concentration. In some embodiments, the AFP concentration, AFP-L3%, and DCP concentration are based on respective serum concentrations measured from the individual. In some embodiments, the serum concentrations are derived from the sample obtained from the individual.
[0167] In some embodiments, the method comprises determining, using the one or more processors, a demographic-protein cancer score based on the demographic and protein information. In some embodiments, the method comprises determining a liver cancer status of the individual based on the demographic -protein cancer score and a predetermined demographicprotein cancer score threshold. In some embodiments, the method comprises determining a liver cancer status of the individual based on the demographic-protein cancer score and a predetermined demographic -protein cancer score threshold and a methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments, a liver cancer-positive status is determined when the demographic -protein cancer score is greater than thepredetermined demographic -protein cancer score threshold. In some embodiments, a liver cancer-positive status is determined when the demographic-protein cancer score is greater than the predetermined demographic-protein cancer score threshold and the methylation cancer score is greater than the predetermined methylation cancer score threshold.G. Machine learning models and training thereof
[0168] The provided embodiments concern the use of one or more machine learning models. In some embodiments, the subject methods comprises use of one or more trained machine learning models, such as a first trained machine learning model, a second trained machine learning model, and a trained ensemble machine learning model. In some aspects, the provided methods further comprise training one or more of the trained machine learning models. For example, the methods may comprise training i) a first trained machine learning model, ii) a second trained machine learning model, and / or iii) a trained ensemble machine learning model.
[0169] In some aspects, provided herein is a method for determining a methylation cancer score of an individual suspected of having a liver cancer using trained machine learning models (e.g., a first trained machine learning model, a second trained machine learning model, and a trained ensemble machine learning model). In some embodiments, the method comprises: inputting, using one or more processors, an extracted methylation fraction value for each of one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, an extracted methylation variance value for each of one or more second genomic regions into a second trained machine learning model to output a methylation variance score; and inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual. In some embodiments, the method further comprises training i) the first trained machine learning model, ii) the second trained machine learning model, and / or iii) the trained ensemble machine learning model.
[0170] In other aspects, provided herein is a method for determining a liver cancer status of an individual using trained machine learning models (e.g., a first trained machine learning model, a second trained machine learning model, and a trained ensemble machine learningmodel). In some embodiments, the method comprises: inputting, using one or more processors, an extracted methylation fraction value for each of one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, an extracted methylation variance value for each of one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining the liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments, the method comprises: inputting, using one or more processors, an extracted methylation fraction value for each of one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, an extracted methylation variance value for each of one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) a demographic -protein cancer score and a predetermined demographic-protein cancer score threshold. In some embodiments, the method further comprises training i) the first trained machine learning model, ii) the second trained machine learning model, and / or iii) the trained ensemble machine learning model.
[0171] In certain embodiments, the methylation fraction value for each genomic region of one or more first genomic regions are mathematically combined and the combined value (e.g., the methylation fraction score) is correlated to a methylation cancer score of the individual, which can be further correlated to the underlying diagnostic question. In certain embodiments, the variance fraction value for each genomic region of one or more second genomic regions are mathematically combined and the combined value (e.g., the methylation variance score) is correlated to a methylation cancer score of the individual, which can be further correlated to the underlying diagnostic question. In some embodiments, these correlations are performed by a first trained machine learning model and a second trained machine learning model, respectively.In certain other embodiments, the methylation fraction score and the methylation variance score are mathematically combined and the combined value (e.g., the methylation cancer score) is correlated to the underlying diagnostic question.
[0172] In some instances, the methylation fraction values and / or the methylation variance values, respectively, may be combined by any appropriate state of the art machine learning model (e.g., a first trained machine learning model and a second trained machines learning model, respectively), to output a methylation fraction score and a methylation variance score, respectively. Well-known mathematical methods, including machine learning models, for correlating values employ methods like discriminant analysis (DA) (e.g., linear-, quadratic-, regularized-DA), Discriminant Functional Analysis (DFA), Kernel Methods (e.g., SVM), Multidimensional Scaling (MDS), Nonparametric Methods (e.g., k-Nearest-Neighbor Classifiers), PLS (Partial Least Squares), Tree-Based Methods (e.g., Logic Regression, CART, Random Forest Methods, Boosting / Bagging Methods), Generalized Linear Models (e.g., Logistic Regression), Principal Components based Methods (e.g., SIMCA), Generalized Additive Models, Fuzzy Logic based Methods, Neural Networks and Genetic Algorithms based Methods. The skilled artisan will have no problem in selecting an appropriate method to evaluate a methylation fraction value combination and a methylation variance value combination as described herein. In one embodiment, the method used in a correlating methylation fraction value of each genomic region of one or more first genomic regions and the method used in correlating a methylation variance value of each genomic region of one or more second genomic regions, e.g., to determine a methylation fraction score and a methylation variance score, respectively, is selected from DA (e.g., Linear-, Quadratic-, Regularized Discriminant Analysis), DFA, Kernel Methods (e.g., SVM), MDS, Nonparametric Methods (e.g., k-Nearest-Neighbor Classifiers), PLS (Partial Least Squares), Tree-Based Methods (e.g., Logic Regression, CART, Random Forest Methods, Boosting Methods), or Generalized Linear Models (e.g., Logistic Regression), and Principal Components Analysis. Details relating to these statistical methods are found in the following references: Ruczinski et al., 12 J. OF COMPUTATIONAL AND GRAPHICAL STATISTICS 475-511 (2003); Friedman, J. H„ 84 J. OF THE AMERICAN STATISTICAL ASSOCIATION 165-75 (1989); Hastie, Trevor, Tibshirani, Robert, Friedman, Jerome, The Elements of Statistical Learning, Springer Series in Statistics (2001); Breiman, L., Friedman, J. H., Olshen, R. A., Stone, C. J. Classification and regression trees, California:Wadsworth (1984); Breiman, L„ 45 MACHINE LEARNING 5-32 (2001); Pepe, M. S., The Statistical Evaluation of Medical Tests for Classification and Prediction, Oxford Statistical Science Series, 28 (2003); and Duda, R. O., Hart, P. E., Stork, D. O., Pattern Classification, Wiley Interscience, 2nd Edition (2001).
[0173] Once the methylation fraction score and the methylation variance score are determined, such as by using a first machine learning model and a second machine learning model respectively, these scores may be combined in order to determine a methylation cancer score of the individual suspected of having a liver cancer. In some instances, the methylation fraction score and the methylation variance score may be combined by any appropriate state of the art ensemble machine learning model, to output a methylation cancer score. In some embodiments, the ensemble machine learning model is a stacking machine learning model, a bagging machine learning model, a boosting machine learning model, and a random forest machine learning model.
[0174] Models for utilizing a methylation cancer score and / or a demographic-protein cancer score to predict the liver cancer status for future samples can also be used. These models may be based on the Compound Covariate Predictor (Radmacher et al. Journal of Computational Biology 9:505-511, 2002), Diagonal Linear Discriminant Analysis (Dudoit et al. Journal of the American Statistical Association 97:77-87, 2002), Nearest Neighbor Classification (also Dudoit et al.), and Support Vector Machines with linear kernel (Ramaswamy et al. PNAS USA 98:15149-54, 2001). The models incorporated markers that were differentially methylated at a given significance level (e.g. 0.01, 0.05 or 0.1) as assessed by the random variance t-test (Wright G. W. and Simon R. Bioinformatics 19:2448-2455, 2003). The prediction error of each model using cross validation, preferably leave-one-out cross-validation (Simon et al. Journal of the National Cancer Institute 95:14-18, 2003can be estimated. For each leave-one-out cross- validation training set, the entire model building process is repeated. In some instances, it is also evaluated in whether the cross-validated error rate estimate for a model is significantly less than one would expect from random prediction. In some cases, the class labels are randomly permuted and the entire leave-one-out cross-validation process is then repeated. The significance level is the proportion of the random permutations that gives a cross-validated error rate no greater than the cross-validated error rate obtained with the real sequence read data.
[0175] Another classification method is the greedy-pairs method described by Bo and Jonassen (Genome Biology 3(4):research0017.1-0017.11, 2002). The greedy-pairs approach starts with ranking all markers based on their individual t-scores on the training set. This method attempts to select pairs of markers that work well together to discriminate the classes.
[0176] Furthermore, a binary tree classifier for utilizing methylation cancer score and / or demographic-protein score is optionally used to predict the class of future individual samples. The first node of the tree incorporated a binary classifier that distinguished two subsets of the total set of classes. The individual binary classifiers are based on the “Support Vector Machines” incorporating markers that were differentially expressed among markers at the significance level (e.g. 0.01, 0.05 or 0.1) as assessed by the random variance t-test (Wright G. W. and Simon R. Bioinformatics 19:2448-2455, 2003). Classifiers for all possible binary partitions are evaluated and the partition selected is that for which the cross-validated prediction error is minimum. The process is then repeated successively for the two subsets of classes determined by the previous binary split. The prediction error of the binary tree classifier can be estimated by cross-validating the entire tree building process. This overall cross-validation includes re-selection of the optimal partitions at each node and re-selection of the markers used for each cross-validated training set as described by Simon et al. (Simon et al. Journal of the National Cancer Institute 95:14-18, 2003). Several-fold cross validation in which a fraction of the samples is withheld, a binary tree developed on the remaining samples, and then class membership is predicted for the samples withheld. This is repeated several times, each time withholding a different percentage of the samples. The samples are randomly partitioned into fractional test sets (Simon R and Lam A. BRB-ArrayTools User Guide, version 3.2. Biometric Research Branch, National Cancer Institute).
[0177] Thus, in one embodiment, the methylation cancer score and the demographic -protein cancer score are rated by their correct correlation to the disease (e.g., liver cancer status in an individual), preferably by p-value test.
[0178] In additional embodiments, factors such as the value, level, feature, characteristic, property, etc. of a transcription rate, mRNA level, translation rate, protein level, biological activity, cellular characteristic or property, genotype, phenotype, etc. can be utilized in addition prior to, during, or after administering a therapy to an individual to enable further analysis of the individual’s liver cancer status.
[0179] In some embodiments, a diagnostic test to correctly predict liver cancer status is measured as the sensitivity of the assay, the specificity of the assay or the area under a receiver operated characteristic (“ROC”) curve. In some instances, sensitivity is the percentage of true positives that are predicted by a test to be positive, while specificity is the percentage of true negatives that are predicted by a test to be negative. In some cases, an ROC curve provides the sensitivity of a test as a function of 1- specificity. The greater the area under the ROC curve, for example, the more accurate or powerful the predictive value of the test. Other useful measures of the utility of a test include positive predictive value and negative predictive value. Positive predictive value is the percentage of people who test positive that are actually positive. Negative predictive value is the percentage of people who test negative that are actually negative.
[0180] In some embodiments, the methylation cancer score (e.g., a combination of the methylation fraction score and the methylation variance score) and / or the demographic-protein cancer score show a statistical difference in different samples of at least p<0.05, p<10'2, p<10'3, p< 10'4or p< 10'5. Diagnostic tests that use these score may show an ROC of at least 0.6, at least about 0.7, at least about 0.8, or at least about 0.9. In some instances, the methylation cancer score and / or the demographic -protein cancer score have different values in different subjects with or without liver cancer. In additional instances, the methylation cancer score and / or the demographic-protein cancer score for different subtypes of liver cancer have different values. In certain embodiments, the methylation cancer score and / or the demographic-protein cancer score are measured in an individual sample using the methods described herein and compared, for example, to a predefined methylation cancer score threshold or a predefined demographicprotein cancer score threshold, respectively, and are used to determine whether the individual has liver cancer, which liver cancer subtype does the individual have, and / or what is the prognosis of the individual having liver cancer. In other embodiments, the methylation cancer score and / or the demographic -protein cancer score in an individual sample are compared, for example, to a predefined methylation cancer score or a predefined demographic-protein cancer score, respectively. In some embodiments, the measurement(s) is then compared with a relevant diagnostic amount(s), cut-off(s), or multivariate model scores that distinguish between the presence or absence of liver cancer, between liver cancer subtypes, and between a “good” or a “poor” prognosis. As is well understood in the art, by adjusting the particular diagnostic cut- off(s) used in an assay, one can increase sensitivity or specificity of the diagnostic assaydepending on the preference of the diagnostician. In some embodiments, the particular diagnostic cut-off is determined, for example, by measuring the methylation cancer score and / or the demographic-protein cancer score in a statistically significant number of samples from individuals with or without liver cancer and from patients with different liver cancer subtypes and drawing the cut-off to suit the desired levels of specificity and sensitivity.
[0181] In some aspects, the subject methods further comprise training i) a first trained machine learning model, ii) a second trained machine learning model, and / or iii) a trained ensemble machine learning model. In some embodiments, the machine learning models are trained on data from one or more individuals, e.g., a population of individuals. In some embodiments, the population of individuals comprises individuals that have a liver cancer. In some embodiments, the population of individual comprises individuals that do not have a liver cancer. In some embodiments, the population of individuals comprises individuals that have various stages of liver cancer, including liver cancer in remission. In some embodiments, the first trained machine learning model, the second trained machine learning model, and the trained ensemble machine learning model are trained on data from the same one or more individuals. In some embodiments, the first trained machine learning model, the second trained machine learning model, and the trained ensemble machine learning model are trained on data from different one or more individuals. In some embodiments, the training data provided to any machine learning model is labeled, such that data associated with an individual having a liver cancer is labeled as such and data associated with an individual not having a liver cancer is labeled as such.
[0182] In some embodiments, the machine learning models may be trained using particular one or more first genomic regions and particular one or more second genomic regions. In some embodiments, as described herein, the one or more first genomic regions and the one or more second genomic regions may be different lengths and may or may not overlap. Once a machine learning model of the subject methods is trained, it is deployed (e.g., used in the subject methods) using the same specifications / features on which it was trained. For example, a first trained machine learning model will use the same one or more first genomic regions as the one or more first genomic regions on which it was trained.H. Methods of diagnosis and treatment
[0183] The subject methods (e.g., a method of determining a methylation cancer score of an individual suspected of having a liver cancer or a method of determining a liver cancer status of an individual) may be employed to diagnose cancer, for example. In particular embodiments, the subject methods may be employed to diagnose an individual with a liver cancer, such as hepatocellular carcinoma (HCC). Thus, disclosed herein, in certain embodiments, are methods of diagnosing liver cancer and selecting subjects suspected of having liver cancer for treatment. In some embodiments, disclosed herein is a method of selecting a subject suspected of having liver cancer for treatment. In further embodiments, the subject methods may be used to diagnose and individual with a liver cancer and subsequently treat the individual for the liver cancer. In some embodiments, the subject diagnosed of having liver cancer is further treated with a therapeutic agent. In some embodiments, the subject diagnosed of having liver cancer is administered a therapeutic agent.
[0184] Thus, in some embodiments, provided herein are methods of treating a liver cancer in an individual, comprising performing the subject methods (e.g., a method of determining a methylation cancer score of an individual suspected of having a liver cancer or a method of determining a liver cancer status of an individual) to diagnose the liver cancer in the individual and administering the individual a therapeutic agent to treat the liver cancer.
[0185] In some embodiments, provided is a method of treating a liver cancer in an individual, the method comprising diagnosing the individual as having a liver cancer according to a method described herein, administering to the individual a therapeutic agent to treat the liver cancer. In some embodiments, the therapeutic agent to treat the liver cancer in the individual is one or more of sorafenib tosylate, doxorubicin, fluorouracil, cisplatin, bevacizumab, atezolizumab, tremelimumab, durvalumab, duvalumab, lenvatinib mesylate, cabozantinib-s- malate, regorafenib, ramucirumab, nivolumab, ipilimumab, ramucirumab, nivolumab, pembrolizumab, fuutibatinib, or pemigatinib. In some embodiments, the diagnosing the individual is based on an obtained cfDNA samples, such as obtained from a blood sample from the individual.
[0186] In some embodiments, the individual is diagnosed of having a liver cancer according to the subject methods. In some embodiments, as described herein, the liver cancer comprises arelapsed liver cancer, a refractory liver cancer, and / or a metastatic liver cancer. In some embodiments, the individual is diagnosed of having a relapsed liver cancer, a refractory liver cancer, and / or a metastatic liver cancer. In some embodiments, the liver cancer is any type of liver cancer, such as the liver cancers described herein. In some embodiments, the liver cancer comprises hepatocellular carcinoma (HCC), fibrolamellar HCC, cholangiocarcinoma, angiosarcoma, or hepatoblastoma.
[0187] In some embodiments, the methylation cancer score and / or the liver cancer status is used to determine the prognosis of an individual having a liver cancer. As such “making a diagnosis” or “diagnosing”, as used herein, is further inclusive of making determining a risk of developing cancer or determining a prognosis, which can provide for predicting a clinical outcome (with or without medical treatment), selecting an appropriate treatment (or whether treatment would be effective), or monitoring a current treatment and potentially changing the treatment, based on the methylation cancer score and / or the liver cancer status determined according to use of the subject methods. Further, in some embodiments of the presently disclosed subject matter, multiple determinations of the methylation cancer score and / or the liver cancer status over time can be made to facilitate diagnosis and / or prognosis. A temporal change in the methylation cancer score, for instance, can be used to predict a clinical outcome, monitor the progression of liver cancer, and / or monitor the efficacy of appropriate therapies directed against the liver cancer.
[0188] Thus, in some embodiments, provided here is a method of determining a prognosis of an individual. In some embodiments, the prognosis if performed after a treatment and may assist with making further medical decisions, such as whether to maintain or switch treatment. In some embodiments, the prognosis is practiced to determine how aggressive a cancer may be.
[0189] In some embodiments, the individual diagnosed of having a liver cancer is further treated with a therapeutic agent. Exemplary therapeutic agents include, but are not limited to, sorafenib tosylate, doxorubicin, fluorouracil, cisplatin, bevacizumab, atezolizumab, tremelimumab, durvalumab, duvalumab, lenvatinib mesylate, cabozantinib-s-malate, regorafenib, ramucirumab, nivolumab, ipilimumab, ramucirumab, nivolumab, pembrolizumab, fuutibatinib, pemigatinib, or a combination thereof.
[0190] In some embodiments, the methods of treatment provided herein comprise subjecting the individual determined to have a liver cancer to an ablation treatment (such as using radiofrequency ablation, microwave ablation, one or more ethanol injections, and / or external-beam radiotherapy), resection treatment, or liver transplantation. In some embodiments, the methods of treatment provided herein comprise subjecting the individual determined to have a liver cancer to a transarterial therapy (TACE) comprising intraarterial infusion of a cytotoxic agent followed by embolization of the one or more vessels that feed a tumor. In some embodiments, the methods of treatment provided herein comprise subjecting the individual determined to have a liver cancer to a selective internal radiation therapy (SIRT) comprising intraarterial infusion of yttrium-90 resin microspheres. In some embodiments, the methods of treatment provided herein comprise administering to the individual determined to have a liver cancer one or more of sorafenib, lenvatinib, regorafenib, cabozantinib, ramucirumab, tremelumumab, nivolumab, pembrolizumab, lenvatinib plus pembrolizumab, atezolizumab plus bevacizumab, erlotinib, brivanib, sunitinib, linifanib, everolimus, pegylated arginine deiminase (ADI-PEG20), doxorubicin FOLFOX (fluorouracil, leucovorin [folinic acid], and oxaliplatin), or tiv antinib.II. Systems
[0191] Also provided herein are systems designed to implement any of the disclosed methods for determining a methylation cancer score of an individual suspected of having a liver cancer or a liver cancer status of an individual. The provided systems may comprise one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that may be executed by the one or more processors.
[0192] In some embodiments, the disclosed systems may be used for determining a methylation cancer score and / or a liver cancer status of an individual in any of a variety of samples as described herein (e.g., a tissue sample, biopsy sample, hematological sample, or liquid biopsy sample derived from the individual).
[0193] In some aspects, provided herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receivesequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
[0194] In some embodiments, the system further comprises instructions to determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments the system further comprises instructions to receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual; determine a demographic -protein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic-protein cancer score threshold.
[0195] In some aspects, provided herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data a sample from an individual and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with agenomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual; determine a demographicprotein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic -protein cancer score threshold.
[0196] In some instances, the embodiments the described systems may further comprise sample processing and library preparation workstations, microplate-handling robotics, fluid dispensing systems, temperature control modules, environmental control chambers, additional data storage modules, data communication modules (e.g., Bluetooth®, WiFi, intranet, or internet communication hardware and associated software), display modules, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, sequencing data analysis software packages), etc., or any combination thereof. In some embodiments, the systems may comprise, or be part of, a computer system or computer network as described elsewhere herein.III. Computer Systems and Networks
[0197] Further provided herein are computer systems and networks (e.g., non-transitory computer-readable storage mediums) designed to implement any of the disclosed methods for determining a methylation cancer score of an individual suspected of having a liver cancer or a liver cancer status of an individual. The non-transitory computer-readable storage medium maystore one or more programs, the one or more programs comprising instructions that may be executed by one or more processors of a system.
[0198] In some aspects, provided herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive sequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
[0199] In some embodiments, the non-transitory computer-readable storage medium further comprises instructions to determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold. In some embodiments the non-transitory computer-readable storage medium further comprises instructions to receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; determine a demographic-protein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic-protein cancer score threshold.
[0200] In some aspects, provided herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, whichwhen executed by one or more processors of a system, cause the system to: receive sequence read data a sample from an individual and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, AFP level in the individual, AFP-L3 level in the individual, and DCP level in the individual; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual; determine a demographic-protein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic -protein cancer score threshold.
[0201] FIG. 5 illustrates an example of a computing device or system in accordance with certain embodiments of the provided methods. Device 500 can be a host computer connected to a network. Device 500 can be a client computer or a server. As shown in FIG. 5, device 500 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server or handheld computing device (portable electronic device) such as a phone or tablet. The device can include, for example, one or more processor(s) 510, input devices 520, output devices 530, memory or storage devices 540, communication devices 560, and nucleic acid sequencers 570. Software module 550 residing in memory or storage device 540 may comprise, e.g., an operating system as well as software for executing the methods described herein. Input device 520 and output device 530 can generally correspond to those described herein, and can either be connectable or integrated with the computer.
[0202] Input device 520 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 530 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.
[0203] Storage device 540 can be any suitable device that provides storage (e.g., an electrical, magnetic or optical memory including a RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 560 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a wired media (e.g., a physical system bus 580, Ethernet connection, or any other wire transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).
[0204] Software module 550, which can be stored as executable instructions in storage device 940 and executed by processor(s) 510, can include, for example, an operating system and / or the processes that embody the functionality of the methods of the present disclosure (e.g., as embodied in the devices as described herein).
[0205] Software module 550 can also be stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described herein, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium can be any medium, such as storage device 540, that can contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer- readable storage media may include memory units like hard drives, flash drives and distribute modules that operate as a single functional unit. Also, various processes described herein may be embodied as modules configured to operate in accordance with the embodiments and techniques described above. Further, while processes may be shown and / or described separately, those skilled in the art will appreciate that the above processes may be routines or modules within other processes.
[0206] Software module 550 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instructionexecution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic or infrared wired or wireless propagation medium.
[0207] Device 500 may be connected to a network (e.g., network 604, as shown in FIG. 6 and / or described below), which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0208] Device 500 can be implemented using any operating system, e.g., an operating system suitable for operating on the network. Software module 550 can be written in any suitable programming language, such as C, C++, Java, Python, and / or R. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or through a Web browser as a Web-based application or Web service, for example. In some embodiments, the operating system is executed by one or more processors, e.g., processor(s) 510.
[0209] Device 500 can further include a nucleic acid sequencer 570, which can be any suitable nucleic acid sequencing instrument.
[0210] FIG. 6 illustrates an example of a computing process in accordance with one embodiment. In system 600, device 500 (e.g., as described above and illustrated in FIG. 5) is connected to network 604, which is also connected to device 606. In some embodiments, device 1206 is a sequencer. Exemplary sequencers can include, without limitation, Roche / 454’s Genome Sequencer (GS) FLX System, Illumina / Solexa’s Genome Analyzer (GA), Illumina’s HiSeq® 2500, HiSeq® 3000, HiSeq® 4000 and NovaSeq® 6000 Sequencing Systems, Life / APG’s Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator’s G.007 system, Helicos BioSciences’ HeliScope Gene Sequencing system, or Pacific Biosciences’ PacBio® RS system.
[0211] Devices 500 and 606 may communicate, e.g., using suitable communication interfaces via network 604, such as a Local Area Network (LAN), Virtual Private Network (VPN), or the Internet. In some embodiments, network 1204 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 500 and 606 may communicate, in part or in whole, via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. Additionally, devices 500 and 606 may communicate, e.g., using suitable communication interfaces, via a second network, such as a mobile / cellular network. Communication between devices 500 and 606 may further include or communicate with various servers such as a mail server, mobile server, media server, telephone server, and the like. In some embodiments, Devices 500 and 606 can communicate directly (instead of, or in addition to, communicating via network 604), e.g., via wireless or hardwired communications, such as Ethernet, IEEE 802.11b wireless, or the like. In some embodiments, devices 500 and 606 communicate via communications 608, which can be a direct connection or can occur via a network (e.g., network 604).
[0212] One or all of devices 500 and 606 generally include logic (e.g., http web server logic) or are programmed to format data, accessed from local or remote databases or other sources of data and content, for providing and / or receiving information via network 604 according to various examples described herein.IV. Examples
[0213] The following examples are included for illustrative purposes only and are not intended to limit the scope of the invention.Example 1: Biomarker discovery and targeted panel designDNA isolation and. Quantification
[0214] Extraction of cfDNA from plasma samples obtained from individuals was performed using a QIAsymphony DSP Circulating DNA Kit (QIAGEN) according to the manufacturer’s recommendations. Isolated cfDNA was size selected by using SPRIselect beads (Beckman Coulter). DNA concentration was measured using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher) according to the manufacturer’s recommendations.
[0215] Tumor and corresponding far site samples of the same tissue were obtained from individuals. Samples were formalin-fixed paraffin-embedded (FFPE) to preserve the tissue. Isolation of genomic DNA from tissue samples was performed using a QIAsymphony DSP DNA Mini kit (QIAGEN) according to the manufacturer’s recommendations. Genomic DNA was sheared to approximately 170 bp and size selected by using SPRIselect beads (Beckman Coulter). The concentration was measured using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher) according to the manufacturer’s recommendations.Library preparation and. enzymatic conversion of cfDN A
[0216] Two reactions were performed according to the manufacturer’s recommendations. The first reaction used ten-eleven translocation dioxygenase 2 (TET2) and T4 phage betaglucosyltransferase (T4-BGT). TET2 is a Fe(II) / alpha- ketoglutarate-dependent dioxygenase that catalyzes the oxidization of 5-methylcytosine to 5- hydroxy methylcytosine (5hmC), 5- formylcytosine (5fC), and 5-carboxycytosine (5caC) in three consecutive steps with the concomitant formation of CO2 and succinate. T4-BGT catalyzes the glucosylation of the formed 5hmC as well as pre-existing genomic 5hmCs to 5-(P-glucosyloxymethyl)cytosine (5gmC). These reactions protect 5mC and 5hmC against deamination by APOBEC3A. This ensures that only cytosines are deaminated to uracils, thus enabling the discrimination of cytosine from its methylated and hydroxy methylated forms. The following sections characterize the catalytic actions of TET2, T4-phage beta-glucosyltransferase (T4-BGT) and apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like 3 A (APOBEC3A).
[0217] Samples of cfDNA or sheared genomic DNA were transferred to individual wells of a 96 well plate to begin NGS library construction. NEBNext Ultra II Reagents (NEB) were then used according to the manufacturer’s instructions for DNA end repair, A-tailing and adaptor ligation of EM-seq adaptor (A5mCA5mCT5mCTTT5mC5mC5mCTA5mCA5mCGA5mCG5mCT5mCTT5mC5mCGAT5 mC*T (SEQ ID NO: 1) and [Phos]GAT5mCGGAAGAG5mCA5mCA5mCGT5mCTGAA5mCT5mC5mCAGT5mCA (SEQ ID NO: 2)). The ligated DNA samples were purified using magnetic beads according to the manufacturer’s protocol. The purified adapter ligated DNA samples were eluted from the magnetic beads and transferred to a new 96 well plate. Adapter ligated DNA samples were then oxidized and glucosylated in reactions containing TET2 and T4-BGT (NEB). The oxidationreaction was initiated by adding an Fe (II) solution and oxidation and gluco sylation reactions were then incubated for 1 h at 37°C. Following this, proteinase K (NEB) was added and the reactions were incubated for 30 min at 37°C to stop the oxidation and glucosylation reaction. The oxidized and glucosylated DNA samples were then purified using magnetic beads according to the manufacturer’s instructions. The purified oxidized and glucosylated DNA samples were eluted in water and transferred to a new 96 well plate. The oxidized and glucosylated DNA samples were then denatured by the addition of formamide (Sigma- Aldrich) and incubation at 85°C for 10 min. The denatured DNA samples were then deaminated in reactions containing APOBEC3A. The deamination reactions were incubated at 37°C for 3 h. The deaminated DNA samples were purified using magnetic beads according to the manufacturer’s protocol. The purified deaminated DNA samples were eluted from magnetic beads and transferred to a new 96 well plate. NEBNext Unique Dual Index Primers and NEBNext Q5U Master Mix (NEB) were then added to the denatured DNA samples and the DNA samples were PCR amplified to generate indexed NGS libraries. The indexed NGS libraries were then purified using magnetic beads, and transferred to a new 96 well plate.Hybridizing the amplified enzymatic-converted library
[0218] A biomarker discovery panel (~42 Mb) was designed using multiple public datasets to explore regions of the genome that could be useful in distinguishing between cancer and noncancer individuals. This panel was designed to allow a multiplexed consideration of cfDNA methylation sequencing based features such as: fragmentation, cfDNA methylation, and nucleosome occupancy. In addition, the TWIST methylome panel (~ 120Mb), developed by TWIST Bioscience was used for targeted cfDNA sequencing. Targeted cfDNA methylation sequencing was performed using these two panels to identify the most informative biomarkers for diagnosing hepatocellular carcinoma (HCC) in patients with liver cirrhosis.
[0219] An optimized target enrichment protocol comprising the following steps was used: pool indexed NGS libraries for hybridization, hybridize capture probes with pooled libraries, bind hybridized targets to streptavidin beads, wash beads, post-capture PCR amplification step, amplified library purification, and a quality control step that includes amplified library quantification. Libraries were sequenced using a NGS platform.Designing the final panel
[0220] Multiple pairs of cancer and non-cancer samples were sequenced. Within each dataset, all samples were balanced for age and sex between the cancer and non-cancer groups. Illumina 450K Methylation array data consisting of tumor tissue samples and solid tissue normal samples from TCGA (The Cancer Genome Atlas) were also analyzed.
[0221] To achieve the desired sequencing coverage of approximately 200x while capturing as many informative regions as possible, the target size of genomic regions for use in the panel was determined to be between 400-500 Kb. Univariate feature selection, as well as cross validated feature selection stratagems were employed to identify the combination of biomarkers to be included in the panel. Several metrics and methods were used to quantify cancer signal and rank the most discriminative markers between the pairs of cancer and non-cancer samples, such as: Mann- Whitney U (MWU) test statistic, Area Under the Receiver Operating Characteristic (AUC), fold change, sensitivity at 90%, 95% and 100% specificity, F-statistic, and Boruta feature selection algorithm.
[0222] Features that were concordant between the plasma datasets and FFPE tissue and / or TCGA datasets were given a higher priority. In the cross validated regime, features that were most often selected by the feature selection metric / method / algorithm among different folds were given a higher ranking. Sample labels (cancer / non-cancer) were randomly permuted to show that the features selected indeed carried true signal and were not observed due to random chance. Perturbations of the data, such as randomly removing samples or repeating samples while leaving the correct sample labels intact, were performed to determine that the features selected were stable.
[0223] The final panel size was -474 Kb, which was within the target panel size range of 400-500 Kb.Example 2: Algorithm development and machine learning model training
[0224] The development of the algorithm for determining a liver cancer status of an individual involved a combination of two components: a methylation cancer score and a demographic-protein cancer score. The methylation cancer score was internally developed and the demographic-protein cancer score was based on the GALAD (Gender, Age, AFP-L3, AFP,Des-gamma-carboxy prothrombin (DCP)) algorithm which is a known metric in the literature used for HCC detection (Johnson et. al., 2014). The methylation cancer score and the demographic-protein cancer score were independently compared to their respective predetermined thresholds and used to determine a liver cancer status of an individual.Methylation cancer score
[0225] The methylation cancer score was calculated using an ensemble machine learning model (e.g., a stacking machine learning model), where the outputs of two base machine learning models are inputted into a meta-classifier to output the methylation cancer score.
[0226] The first base machine learning model used the methylation fraction values as an input, which is based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions. The methylation fraction value undergoes several preprocessing procedures designed to enhance the model predictive ability before employing a Random Forest classifier to produce the methylation fraction score.
[0227] The second base machine learning model used the methylation variance values as an input, which is based on the variance of in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions. The methylation variance value undergoes several preprocessing steps before the transformed features are fed into a Logistic Regression classifier to produce the methylation variance score.
[0228] The methylation fraction score and the methylation variance score are then used as the input for a stacking machine learning model which outputs the methylation cancer score.Demographic-protein cancer score
[0229] The demographic-protein cancer score (e.g., the GALAD algorithm) was determined based on the following information from the individual: concentration of AFP, concentration of AFP-L3, concentration of DCP, age, and gender (sex). The GALAD algorithm is a logistic regression that was fitted on data from HCC positive and HCC negative individuals from a single medical center in the United Kingdom Johnson et. al., 2014). The model was validated on thousands of samples from Germany and Japan (Best et. al., 2020). The output of the GALAD algorithm is the demographic-protein cancer score.
[0230] The threshold value for the demographic -protein cancer score was chosen as 0.5, which corresponds to a GALAD score of 0. This choice was motivated by interpreting the GALAD score as a probability where a GALAD score of 0 would correspond to a probability of 0.5. A natural prior for such a threshold would be a probability of 0.5.Determining a liver cancer status in an individual
[0231] A liver cancer status in an individual is determined by assessing whether the methylation cancer score or the demographic-protein cancer score passes a predetermined threshold. Model analysis and choice of features were guided by repeated cross-validation (number of repetitions = 15) using random seeds. The threshold for the methylation cancer score was chosen such that after the ‘OR’ operation with the demographic -protein cancer score, the target specificity was ~ 90%. From the results of the repeated cross-validation described in the following Examples, the threshold value for the methylation cancer score was chosen as 0.57.
[0232] Test results with a methylation cancer score greater than 0.57 or a demographicprotein cancer score greater than 0.5 are reported as “Positive” for liver cancer. Test results with both a methylation cancer scoreless than or equal to 0.57 and a demographic -protein cancer score less than or equal to 0.5 are reported as “Negative” for liver cancer.Example 3: Cross-validation for stage 1 liver cancer
[0233] The ability to differentiate stage 1 liver cancer from non-cancer was evaluated using the final panel described in Example 2. Training was conducted using two HCC case-control cohorts - “Cohort 1” and “Cohort 2” - as well as a prospective HCC cohort - “Cohort 3”. All specimens from Cohort 1 , Cohort 2 and Cohort 3 were collected from participants of Institutional Review Board (IRB) approved studies and all participants provided informed consent. Cohort 1 consisted of participants enrolled in the ELITE sample collection study (clinicaltrials.gov ID: NCT05181826). Cohort 2 consisted of participants enrolled in the LIVER- 1 sample collection protocol (clinicaltrials.gov ID: NCT05199259). Cohort 3 consisted of participants who both enrolled in the CLiMB clinical study (clinicaltrials.gov ID: NCT03694600) and provided informed consent prior to June 20, 2020. Serum and whole blood samples were collected from study participants at clinical sites. Serum samples were collected by using BD vacutainer SST tubes (Becton Dickinson), processed to serum according to themanufacturer’s recommendations and shipped to a central laboratory. Upon receipt at the central laboratory, serum samples were stored frozen at approximately -75°C until analysis. Whole blood samples were collected by using PAXgene ccfDNA tubes (PreAnalytiX) and shipped to a central laboratory. Whole blood samples were then processed to plasma by centrifugation at the central laboratory and stored frozen at approximately -75°C until analysis.
[0234] In Batch 1, the case-control cohorts, Cohort 1 and Cohort 2, were used to optimize selection of features and model types by employing a fifteen-fold cross validation. Only stage 1 cancer was included in the test fold for each iteration. The AUC utilizing only the methylation cancer score (without the demographic -protein cancer score) was used to evaluate diagnostic performance (FIG. 2A).
[0235] In Batch 1, the case-control cohorts, Cohort 1 and Cohort 2, as well as the prospective cohort, Cohort 3, were used to optimize selection of features and model types by employing a fifteen-fold cross validation. Only stage 1 cancer was included in the test fold for each iteration. The AUROC utilizing only the methylation cancer score (without the demographic-protein cancer score) was used to evaluate diagnostic performance (FIG. 2B).
[0236] Batch 1 (without Cohort 3) achieved a sensitivity of 0.62 at a specificity of 0.904 with an AUC of 0.826. Batch 2 (with Cohort 3) achieved a sensitivity of 0.67 at a specificity of 0.903 with an AUC of 0.845.
[0237] Table 2 shows AUC information, sensitivity, and specificity of Batch 1 and Batch 2. AUC was calculated using only the methylation cancer score. Sensitivity and specificity were calculated using the methylation cancer score and the demographic -protein cancer score.Table 2. Cross-validation results of stage 1 liver cancer in test fold.Example 4: Cross-validation for stage 2 or higher liver cancer
[0238] In Batch 3, the case-control cohorts, Cohort 1 and Cohort 2, were used to optimize selection of features and model types by employing a fifteen-fold cross validation. Only stage 2 and above cancer was included in the test fold for each iteration. The AUROC utilizing only the methylation cancer score (without the demographic-protein cancer score) was used to evaluate diagnostic performance (FIG. 3A).
[0239] In Batch 4, the case-control cohorts, Cohort 1 and Cohort 2, as well as the prospective cohort, Cohort 3, were used to optimize selection of features and model types by employing a fifteen-fold cross validation. Only stage 2 and above cancer was included in the test fold for each iteration. The AUROC utilizing only the methylation cancer score (without the demographic-protein cancer score) was used to evaluate diagnostic performance (FIG. 3B).
[0240] Batch 3 (without Cohort 3) achieved a sensitivity of 0.794 at a specificity of 0.91 with an AUC of 0.87. Batch 4 (with Cohort 3) achieved a sensitivity of 0.79 at a specificity of 0.904 with an AUC of 0.88.
[0241] Table 3 shows AUC information, sensitivity, and specificity of Batch 3 and Batch 4. AUC was calculated using only the methylation cancer score. Sensitivity and specificity were calculated using the methylation cancer score and the demographic -protein cancer score.Table 3. Cross-validation results of stage 2 or higher liver cancer in test fold.Example 5: Cross-validation of Cohort 3
[0242] In Batch 5, the case-control cohorts, Cohort 1 and Cohort 2, as well as the prospective cohort, Cohort 3, were used to optimize selection of features and model types by employing a fifteen-fold cross validation. Only Cohort 3 was included in the test fold for each iteration. The AUROC utilizing only the methylation cancer score (without the demographicprotein cancer score) was used to evaluate diagnostic performance (FIG. 4).
[0243] Batch 5 achieved a sensitivity of 0.5578 at a specificity of 0.90 with an AUC of 0.82.
[0244] Table 4 shows AUC information, sensitivity, and specificity of Batch 5. AUC, sensitivity, and specificity were calculated using only the methylation cancer score.Table 4. Cross-validation results of Cohort 3 in test fold.Example 6: Detection of Early Stage HCC
[0245] This example demonstrates that the HelioLiver test (HL), as described in Examples 1-5, can detect more tumors and at an earlier stage compared to an Ultrasound test (US), when used in a surveillance setting. Monte Carlo simulations were used to assess longitudinal test performance.
[0246] First, various parameters were evaluated for stage 1 HCC patients in non-progression following treatment (n=19) and progression following treatment (n=22) categories. As illustrated in FIG. 7, methylation score using the HelioLiver test is the best predictor of patient outcomes, with an AUC of 0.897, as compared to other widely used indices. Moreover, FIG. 8 shows methylation score difference for these patients, which suggests that a decrease in methylation score may be indicative of an improvement in the patient’s condition. Furthermore, patients with early-intermediate stage HCC in this cohort who have a methylation score of >0.75 at diagnosis have a shorter time to progression compared to those who have a methyl score <0.75 (FIG. 9). The data demonstrates the use of the claimed methods to determine whether or not a treatment is working, thus assisting in important treatment decisions to improve patient outcomes.
[0247] Cumulative sensitivities across HCCs (FIG. 10A) and for early-stage HCCs (FIG. 10B) were evaluated at various timepoints using either the HE or US tests. For each timepoint, sensitivity was calculated as the number of tumors detected divided by the total number of tumors detected at the first timepoint. These results indicate that the HE test is more sensitive at HCC detection compared to the US test.
[0248] Additionally, the HL test was able to detect tumors of smaller size compared to the US test across all timepoints (FIG. 10C). These results indicate that the HL test is capable of earlier HCC detection compared to the US test.
[0249] The present invention is not intended to be limited in scope to the particular disclosed embodiments, which are provided, for example, to illustrate various aspects of the invention.Various modifications to the compositions and methods described will become apparent from the description and teachings herein. Such variations may be practiced without departing from the true scope and spirit of the disclosure and are intended to fall within the scope of the present disclosure.
Claims
CLAIMSWhat is claimed is:
1. A method for determining a methylation cancer score of an individual suspected of having a liver cancer, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; and inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output the methylation cancer score of the individual.
2. The method of claim 1, wherein, for the methylation fraction value for the genomic region of one or more first genomic regions, the total reads associated with the one or more first genomic regions is the sum of reads from the sequence read data having at least one base overlapping with any of the one or more first genomic regions.
3. The method of claim 1 or 2, wherein, for the methylation fraction value for the genomic region of the one or more first genomic regions, the methylated reads associated with the genomic region of the one or more first genomic regions is the sum of reads from thesequence read data having: (i) at least one base overlapping with the genomic region of the one or more first genomic regions, (ii) at least one methylated CpG site, and (iii) at least 50% of total CpG sites having a methylation.
4. The method of any one of claims 1-3, wherein, for the methylation variance value for the genomic region of the one or more second genomic regions, reads associated with the genomic region of the one or more second genomic regions are reads from the sequence read data having: (i) at least one base overlapping with the genomic region of the one or more second genomic regions, and (ii) at least one CpG site.
5. The method of any one of claims 1-4, further comprising: determining a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold.
6. The method of claim 5, wherein the liver cancer status is a liver cancer-positive status when the methylation cancer score is greater than the predetermined methylation cancer score threshold.
7. The method of any one of claims 1-6, further comprising: processing, at the one or more processors, extracted the methylation fraction value for each of the one or more first genomic regions before inputting into the first trained machine learning model; and / or processing, at the one or more processors, extracted the methylation variance value for each of the one or more second genomic regions before inputting into the second trained machine learning model.
8. The method of claim 7, wherein the processing of the extracted methylation value and the extracted methylation variance value comprises inputting missing data, feature selection, data transformation, and / or data scaling.
9. The method of any one of claims 1-8, further comprising: receiving, at the one or more processors, demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alphafetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des- gamma-carboxy prothrombin (DCP) level in the individual.
10. The method of claim 9, wherein the AFP level is a concentration of AFP.
11. The method of claim 9 or 10, wherein the AFP-L3 level is a percentage of AFP-L3 relative to total AFP.
12. The method of any one of claims 9-11, wherein the DCP level is a concentration of DCP.
13. The method of any one of claims 9-12, further comprising: determining, using the one or more processors, a demographic-protein cancer score based on the demographic and protein information.
14. The method of claim 13, further comprising: determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and a predetermined demographic-protein cancer score threshold.
15. The method of claim 14, wherein a liver cancer-positive status is determined when (i) the methylation cancer score is greater than the predetermined methylation cancer score threshold, and / or (ii) the demographic -protein cancer score is greater than the predetermined demographic -protein cancer score threshold.
16. A method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score;inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score; inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining the liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold.
17. A method for determining a liver cancer status of an individual, the method comprising: receiving, at one or more processors, sequence read data from a sample from the individual and demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; extracting for each of one or more first genomic regions, using the one or more processors and the sequence read data, a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extracting for each of one or more second genomic regions, using the one or more processors and the sequence read data, a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; inputting, using the one or more processors, the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning model to output a methylation fraction score; inputting, using the one or more processors, the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning model to output a methylation variance score;inputting, using the one or more processors, the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score; and determining a liver cancer status of the individual based on (a) the methylation cancer score and a predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and a predetermined demographic-protein cancer score threshold.
18. The method of any one of claims 1-17, further comprising: performing a sequencing assay to generate the sequence read data.
19. The method of claim 18, wherein the sequencing assaying is a pair-end sequencing assay.
20. The method of claim 18 or 19, wherein performing the sequencing assay comprises: enzymatically converting unmethylated cytosines to uracils in the sample; amplifying nucleic acids in the sample using polymerase chain reaction (PCR); and sequencing the amplified nucleic acids using a next generation sequencing (NGS) technique.
21. The method of claim 20, wherein the enzymatic conversion comprises: treating the sample with an oxidizing agent to oxidize methylated cytosines; and / or treating the sample with a glucosyltransferase to glucosylate methylated cytosines; and treating the sample with a deaminating agent to deaminate unmethylated cytosines to uracils.
22. The method of claim 21, wherein the oxidizing agent is tet methylcytosine dioxygenase 2 (TET2) and the glucosyltransferase is T4 phage beta-glucosyltransferase (T4-BGT).
23. The method of claim 21 or 22, wherein the deaminating agent is apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC).
24. The method of any one of claims 1-23, wherein the first and / or second trained machine learning models comprise a neural network model or trained deep learning model.
25. The method of claim 24, wherein the first and / or second trained machine learning models comprise a support vector machine model, a random forest machine model, or a logistic regression machine model.
26. The method of claim 24 or 25, further comprising a cross-validation procedure.
27. The method of any one of claims 1-26, wherein:the first trained machine learning model is trained using one or more training data sets comprising paired methylation fraction values from a plurality of individuals; and / or the second trained machine learning model is trained using one or more training data sets comprising paired methylation variance values from a plurality of individuals.
28. The method of any one of claims 1-27, wherein the trained ensemble machine learning model comprises a bagging machine learning model, a boosting machine learning model, a stacking machine learning model and a random forest machine learning model.
29. The method of any one of claims 1-28, wherein the trained ensemble machine learning model is trained using one or more training data sets comprising paired methylation fraction values and methylation variance values from a plurality of individuals.
30. The method of any one of claims 1-29, further comprising: obtaining the sample from the individual.
31. The method of any one of claims 1-30, wherein the sample comprises a tissue biopsy sample or a liquid biopsy sample.
32. The method of any one of claims 1-31, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.
33. The method of any one of claims 1-32, wherein the sample is a liquid biopsy sample and comprises circulating tumor cells.
34. The method of any one of claims 1-33, wherein the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA).
35. The method of any one of claims 9-15 and 17-34, wherein the sample is a first sample, and wherein the demographic and protein information is obtained from a second sample from the individual.
36. The method of any one of claims 9-15 and 17-35, further comprising: obtaining the second sample from the individual.
37. The method of claim 35 or 36, wherein the first sample and the second sample are the same sample.
38. The method of claim 35 or 36, wherein the first sample and the second sample are different samples.
39. The method of any one of claims 35-38, wherein the second sample comprises a tissue biopsy sample or a liquid biopsy sample.
40. The method of any one of claims 35-39, wherein the second sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.
41. The method of any one of claims 35-40, wherein the second sample is a liquid biopsy sample and comprises circulating tumor cells.
42. The method of any one of 35-41, wherein the second sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA).
43. The method of any one of claims 1-42, wherein the one or more first genomic regions and the one or more second genomic regions are each individually between about 50 base pairs and about 500 base pairs in length.
44. The method of claim 43, wherein the one or more first genomic regions and the one or more second genomic regions are each individually 100 or 200 base pairs in length.
45. The method of claim 43 or 44, wherein the one or more first genomic regions are each individually 100 base pairs in length, and wherein the one or more second genomic regions are each individually 200 base pairs in length.
46. The method of any one of claims 1-45, wherein the one or more first genomic regions overlap with the one or more second genomic regions.
47. The method of any one of claims 1-46, wherein the one or more first genomic regions do not overlap with the one or more second genomic regions.
48. The method of any one of claims 1-47, wherein the one or more first genomic regions comprise at least 900 regions.
49. The method of any one of claims 1-48, wherein the one or more second genomic regions comprise at least 900 regions.
50. The method of any one of claims 1-49, wherein the one or more first genomic regions or second genomic regions individually comprise any of the genomic regions listed in Table 1.
51. The method of any of claims 1-50, wherein the individual is a human.
52. The method of any one of claims 1-51, wherein the individual is suspected of having a liver cancer.
53. The method of claim 52, wherein the liver cancer is hepatocellular carcinoma (HCC).
54. A system comprising:one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
55. The system of claim 54, further comprising instructions to: determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold.
56. The system of claim 54, further comprising instructions to: receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual;determine a demographic -protein cancer score based on the demographic and protein information; and determine a liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic -protein cancer score threshold.
57. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a system, cause the system to: receive sequence read data from a sample from an individual suspected of having a liver cancer; extract for each of one or more first genomic regions a methylation fraction value based on a ratio of methylated reads associated with a genomic region of the one or more first genomic regions and total reads associated with the one or more first genomic regions; extract for each of one or more second genomic regions a methylation variance value based on the variance in the number of methylated CpG sites per read across reads associated with a genomic region of the one or more second genomic regions; input the extracted methylation fraction value for each of the one or more first genomic regions into a first trained machine learning models to output a methylation fraction score; input the extracted methylation variance value for each of the one or more second genomic regions into a second trained machine learning models to output a methylation variance score; and input the methylation fraction score and the methylation variance score into a trained ensemble machine learning model to output a methylation cancer score of the individual.
58. The non-transitory computer-readable storage medium of claim 57, further comprising instructions to: determine a liver cancer status of the individual based on the methylation cancer score and a predetermined methylation cancer score threshold.I l l9. The non-transitory computer-readable storage medium of claim 57, further comprising instructions to: receive demographic and protein information from the individual, wherein the demographic and protein information comprises information regarding: sex of the individual, age of the individual, alpha-fetoprotein (AFP) level in the individual, AFP-L3 level in the individual, and des-gamma-carboxy prothrombin (DCP) level in the individual; determine a demographic -protein cancer score based on the demographic and protein information; and determine the liver cancer status of the individual based on (a) the methylation cancer score and the predetermined methylation cancer score threshold, and / or (b) the demographic-protein cancer score and the predetermined demographic -protein cancer score threshold.
Citation Information
Patent Citations
Methods for multimodal epigenetic sequencing assays
US20230323473A1
Methods and systems for detecting cancer via nucleic acid methylation analysis
US20240084397A1
Optimization of model-based featurization and classification
US20240161867A1
Systems and methods for predicting hematological conditions using methylation data
WO2023172772A1