Compositions and methods for TET-assisted pyridine borane sequencing for cell-free DNA
Patent Information
- Application Number
- JP2024505327
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-27
- Filing Date
- 2022-07-26
- Publication Date
- 2025-07-17
AI Technical Summary
Current methods for detecting cancer through cell-free DNA methylation are limited by low depth, targeted, or low resolution, qualitative enrichment-based sequencing, which fail to comprehensively capture the cfDNA methylome, and lack tissue of origin information necessary for accurate cancer localization.
A TET-assisted pyridine borane sequencing (TAPS) method, specifically optimized for cell-free DNA (cfDNA), provides high-quality, genome-wide methylation signatures by converting 5mC, 5hmC, 5caC, and 5fC to DHU, allowing for quantitative analysis of methylation biomarkers, tissue of origin, and DNA fragmentation profiles.
TAPS achieves unique mapping rates of at least 80% and deduplicated mapping rates of at least 70%, enabling accurate identification of cancer types like HCC and PDAC, and provides multimodal information for early cancer detection and tissue localization.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Detailed Description of the Invention
[0001] [Technical field] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 203,565, filed July 27, 2021, the contents of which are incorporated by reference herein in their entirety.
[0002] The contents of the electronic sequence listing (sequencelisting.xml, size: 8,000 bytes, and creation date: July 26, 2022) are incorporated by reference in their entirety into this specification.
[0003] The present disclosure provides compositions and methods related to TET-assisted pyridine borane sequencing (TAPS). Specifically, the present disclosure provides cfDNA-optimized TAPS (cfTAPS) that provides high quality and deep whole genome cell-free methylomes. The compositions and methods provided herein facilitate obtaining multimodal information on cfDNA features, including DNA methylation, tissue of origin, and DNA fragmentation for disease diagnosis and treatment. [Background technology] Recent advances in cancer research provide new ways to treat cancer, but early detection still represents the best chance to cure cancer. Early treatment not only significantly improves patient survival but also significantly reduces costs. Circulating cell-free DNA (cfDNA), which is free-floating DNA in plasma derived from cell death in various healthy and diseased tissues, has great potential for developing early cancer detection assays. Genetic information in cfDNA, such as mutations and copy number variations (CNVs), shows potential utility for monitoring cancer progression and treatment. However, given the low percentage of tumor DNA in early disease, it is difficult to detect genetic alterations. Moreover, genetic alterations are weakly informative about the tissue of origin, which is required to determine the location of malignant tumors.
[0004] In contrast, widespread epigenetic changes such as DNA methylation in both cancer cells and the tumor microenvironment occur early in tumor development. Recent studies have shown that cfDNA methylation is one of the most promising biomarkers for early cancer detection by providing thousands of methylation changes that can be combined to overcome detection limits and information on the tissue of origin that allows cancer localization with high confidence. DNA methylation is best determined by whole-genome, base-resolution, and quantitative sequencing methods such as bisulfite sequencing. However, bisulfite sequencing is DNA damaging and expensive. Thus, current cfDNA methylation sequencing is limited by being low-depth, targeted, or low-resolution, and qualitative enrichment-based sequencing, thus incompletely capturing the cfDNA methylome. Summary of the Invention The embodiments of the present disclosure include a method for obtaining a methylation signature. According to these embodiments, the method includes isolating cell-free DNA (cfDNA) from a sample, preparing a sequencing library comprising cfDNA, and carrying out TET-assisted pyridine borane sequencing (TAPS) on the sequencing library to obtain a methylation signature of cfDNA. In some embodiments, the methylation signature is a whole genome methylation signature.
[0005] In some embodiments, the unique mapping rate obtained from TAPS to cfDNA is at least 80% and / or the unique de-duplicated mapping rate is at least 70%.
[0006] In some embodiments, preparing the sequencing library comprises ligating sequencing adaptors to the isolated cfDNA.
[0007] In some embodiments, carrier DNA is added to the sequencing library prior to performing TAPS.
[0008] In some embodiments, the method further includes identifying at least one methylation biomarker from the cfDNA whole-genome methylation signature and determining whether the methylation biomarker is indicative of cancer.
[0009] In some embodiments, the methylation biomarkers comprise differentially methylated regions (DMRs).
[0010] In some embodiments, the method further comprises classifying the sample based on the DMR as compared to a reference DMR.
[0011] In some embodiments, the reference DMR corresponds to a non-cancerous control or a cancerous control.
[0012] In some embodiments, the method further includes identifying at least one methylation biomarker from the cfDNA whole-genome methylation signature and determining a tissue of origin corresponding to the methylation biomarker.
[0013] In some embodiments, the method further comprises classifying the sample based on tissue of origin biomarkers.
[0014] In some embodiments, the method further comprises identifying a DNA fragmentation profile and determining whether the fragmentation profile is indicative of cancer.
[0015] In some embodiments, the method further includes identifying at least one sequence variant from the cfDNA and determining whether the sequence variant is indicative of cancer.
[0016] In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5mC modifications in cfDNA and providing a quantitative measure of the frequency of the 5mC modification.
[0017] In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5hmC modifications in cfDNA and providing a quantitative measure of the frequency of the 5hmC modification.
[0018] In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5caC modifications in cfDNA and providing a quantitative measure of the frequency of the 5caC modifications.
[0019] In some embodiments, performing TAPS on a sequencing library to obtain a genome-wide methylation signature includes identifying 5fC modifications in cfDNA and providing a quantitative measure of the frequency of the 5fC modification.
[0020] The present disclosure also includes a method of determining whether a subject has cancer using any of the methods described herein. In some embodiments, the cancer comprises hepatocellular carcinoma (HCC) or pancreatic ductal adenocarcinoma (PDAC).
[0021] The present disclosure also includes a method of determining whether a subject has early stage cancer using any of the methods described herein. In some embodiments, the cancer comprises early stage hepatocellular carcinoma (HCC) or early stage pancreatic ductal adenocarcinoma (PDAC).
[0022] In yet another preferred embodiment, the present invention provides a multimodal method of analyzing cfDNA in a patient sample, comprising isolating cfDNA from the patient sample; converting 5mC and / or 5hmC residues in the sample to DHU residues to provide a modified cfDNA sample; sequencing the modified cfDNA sample to identify methylated regions in the sample, where a cytosine (C) to thymine (T) transition or a cytosine (C) to DHU transition in the modified cfDNA sample compared to an unmodified reference cfDNA provides a location of either 5mC or 5hmC in the cfDNA; and performing one or more additional analytical steps on the modified cfDNA selected from the group consisting of: a) determining copy number variation of one or more targets in the modified cfDNA sample, b) determining a tissue of origin or one or more targets in the modified cfDNA sample, c) determining a fragmentation profile of the modified cfDNA sample, and d) identifying one or more single nucleotide mutations in the modified cfDNA sample.
[0023] In some embodiments, sequencing the modified cfDNA sample to identify methylated regions in the sample, including identifying at least one differentially methylated region (DMR).
[0024] In some embodiments, the multimodal method further comprises classifying the sample based on the DMR compared to a reference DMR.
[0025] In some embodiments, the reference DMR corresponds to a non-cancerous control or a cancerous control.
[0026] In some embodiments, determining copy number variation (CNV) of one or more targets in the modified cfDNA sample comprises determining observed read counts for target sequences across the genome by dividing the reference genome into bins and counting the number of reads in each bin.
[0027] In some embodiments, the presence of a copy number abnormality greater than 500 kb is indicative of a CNV in the patient.
[0028] In some embodiments, determining the tissue of origin or one or more targets in the modified cfDNA sample comprises tissue deconvolution of data obtained from sequencing the modified cfDNA sample.
[0029] In some embodiments, tissue deconvolution involves comparing the DNA methylation values determined in the modified cfDNA sample to reference DMRs from two or more different tissues.
[0030] In some embodiments, determining the fragmentation profile of the modified cfDNA sample comprises classifying fragment lengths and fragment periodicity in the modified cfDNA sample.
[0031] In some embodiments, classifying fragment lengths and fragment periodicity in the modified cfDNA sample further comprises calculating the percentage of cfDNA fragments between 300 and 500 bp in 10 bp length range bins.
[0032] In some embodiments, the step of identifying one or more single nucleotide mutations in the modified cfDNA sample further includes distinguishing C to T SNPs from 5mC or 5hmC at specific positions in the cfDNA by comparing the sequencing results after TAPS, where the presence of a T read at a specific position in the complement to the original bottom strand of the cfDNA is indicative of a C to T SNP and the presence of a C read at a specific position in the complement to the original bottom strand of the cfDNA is indicative of a 5mC or 5hmC.
[0033] In some embodiments, two or more of steps a, b, c, and d are performed on modified cfDNA.
[0034] In some embodiments, three or more of steps a, b, c, and d are performed on modified cfDNA.
[0035] In some embodiments, all of steps a, b, c, and d are performed on modified cfDNA.
[0036] In some embodiments, the unique mapping rate resulting from the sequencing step is at least 80% and / or the unique de-duplicated mapping rate is at least 70%.
[0037] In some embodiments, the sequencing step further comprises preparing a sequencing library comprising cfDNA by ligating sequencing adaptors to the isolated cfDNA.
[0038] In some embodiments, carrier DNA is added to the cfDNA.
[0039] In some embodiments, the multimodal method provides a cfDNA whole-genome methylation signature, and the method further comprises identifying at least one methylation biomarker from the cfDNA whole-genome methylation signature and determining whether the methylation biomarker is indicative of cancer.
[0040] In some embodiments, the multimodal method further includes identifying 5mC modifications in cfDNA and providing a quantitative measure of the frequency of the 5mC modifications.
[0041] In some embodiments, the multimodal method further comprises identifying 5hmC modifications in cfDNA and providing a quantitative measure of the frequency of the 5hmC modifications.
[0042] In some embodiments, the multimodal method further includes identifying a 5caC modification in cfDNA and providing a quantitative measure of the frequency of the 5caC modification.
[0043] In some embodiments, the multimodal method further comprises detecting a 5fC modification in cfDNA and providing a quantitative measure of the frequency of the 5fC modification.
[0044] In some embodiments, the step of converting 5mC and / or 5hmC residues in the sample to DHU residues to provide a modified cfDNA sample includes oxidizing the 5mC and / or 5hmC residues to provide 5caC and / or 5fC residues, and reducing the 5caC and / or 5fC residues to DHU residues.
[0045] In some embodiments, the step of oxidizing 5mC and / or 5hmC residues to provide 5caC and / or 5fC residues comprises treatment of the sample with a Tet enzyme.
[0046] In some embodiments, the step of oxidizing 5mC and / or 5hmC residues to provide 5caC and / or 5fC residues comprises treatment of the sample with a chemical oxidizing agent such that one or more 5fC residues are generated.
[0047] In some embodiments, the step of reducing 5caC and / or 5fC residues to DHU residues comprises treatment of the sample with a borane reducing agent.
[0048] Embodiments of the present disclosure also include methods of determining whether a subject has early stage cancer using any of the multimodal methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1: cfDNA analysis by TAPS. (A) Schematic of the TAPS approach for cfDNA analysis. cfDNA is isolated from 1–3 mL of plasma. 10 ng of cfDNA is ligated to Illumina sequencing adaptors and topped up with 100 ng of carrier DNA. 5mC and 5hmC in DNA are then oxidized to 5caC by mTet1CD enzyme, reduced to DHU by PyBr, amplified, and detected as T in the final sequencing run. Computational analysis of TAPS data allows for the simultaneous characterization of multiple cfDNA features, including DNA methylation, tissue of origin, fragmentation patterns, and CNVs. (B) Number of total reads, uniquely mapped reads, and uniquely mapped PCR deduplication reads in 87 cfDNA TAPS libraries. The total number of reads, as well as the average percentage of uniquely mapped reads and deduplication reads compared to total reads, are shown above the bar graphs. Error bars represent standard error. (C) Conversion and false positive rates of 5mC in 85 cfDNA TAPS libraries based on spike-in controls with modified or unmodified cytosines at known positions. Each dot represents an individual sample.
[0049] [Figure 2A] cfDNA methylation in clinical samples. Cancer stage distribution of 21 HCC and 23 PDAC patients included in this study.
[0050] FIG. 2B: cfDNA methylation in clinical samples. Mean per CpG genomic modification levels in non-cancer control, HCC and PDAC cfDNA. Each dot represents an individual sample.
[0051] [Figure 2C] cfDNA methylation in clinical samples. PCA plot of cfDNA methylation in 1 kb genomic windows in non-cancer controls and HCC.
[0052] [Figure 2D] cfDNA methylation in clinical samples. PCA plot of cfDNA methylation in 1 kb genomic windows in non-cancer controls and PDAC.
[0053] [Figure 2E] cfDNA methylation in clinical samples. Regional overexpression analysis showed that regulatory regions were most correlated with PC2 for HCC and PC1 for PDAC.
[0054] [Figure 2F] cfDNA methylation in clinical samples. Receiver operating characteristic (ROC) curves of model classification performance based on differentially methylated enhancers in HCC and non-cancer controls (n=51, HCC=21, non-cancer controls=30).
[0055] FIG. 2G: cfDNA methylation in clinical samples. LOO cancer prediction scores for HCC and non-cancer controls. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC.
[0056] [Figure 2H] cfDNA methylation in clinical samples. ROC curves of model classification performance based on differentially methylated enhancers between PDAC and non-cancer controls (n=53, PDAC=23, non-cancer controls=30).
[0057] FIG. 2I: cfDNA methylation in clinical samples. LOO cancer prediction scores for PDAC and non-cancer controls. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as PDAC.
[0058] Figure 3: cfTAPS allows the analysis of tissue of origin and fragmentation patterns in cfDNA. (A) Mean tissue contribution in non-cancer individuals estimated by NNLS. Tissue contributions less than 1.5% are summarized as "other". (B) Box plots showing estimated liver cancer contributions within non-cancer, HCC, and PDAC groups. Statistical significance was assessed by paired t-test. ns-not significant. (C) Length distribution of cfDNA fragments in the three groups. For each sample, the proportion (P) of long cfDNA fragments (300-500bp) in 10 base pair intervals was used as a fragmentation feature for PCA analysis and machine learning. (D) Box plots showing the proportion of short (70-150bp) and long (300-500bp) fragments in non-cancer controls, PDAC, and HCC. Kruskal-Wallis tests were performed to test for differences in fragment size distribution between groups. Statistically significant differences are marked with asterisks (*P value < 0.05, **P value < 0.01, ***P value < 0.001, ****P value < 0.0001). (E) PCA plot of the proportion of cfDNA 10 bp fragments in non-cancer controls and HCC (left panel), and non-cancer controls and PDAC (right panel).
[0059] Figure 4. Integrating multimodal features from cfTAPS enhances multi-cancer detection. (A) Heatmap showing the performance of individual models for multi-cancer prediction and the predicted probability for each patient. Each vertical column is a patient. Detection yes / no means the patient is correctly or incorrectly classified based on a particular feature. Prediction score means the probability of classifying a patient into a particular group based on a particular feature. (B) Schematic detailing how we integrate multiple features (DNA methylation, tissue contribution, and fragmentation percentage) extracted from cfTAPS data for multi-cancer prediction. (C) Actual and predicted patient status calculated with LOO cross-validation.
[0060] Figure 5: cfDNA TAPS. (A) Agarose gel of 10 representative cfDNA TAPS libraries after post-amplification cleanup. All cfDNA TAPS libraries were prepared from 10 ng of cfDNA and amplified with 7 PCR cycles. (B) Number of mapped read pairs for hg38, spike-in, and carrier DNA in 87 cfDNA TAPS libraries. The average percentage of mapped read pairs compared to total read pairs is shown above the bar graph. Error bars represent standard error. (C) Number of total reads, uniquely mapped reads, and uniquely mapped PCR-deduped reads in cfDNA WGBS (EGAD00001004317) (24). The total number of reads, and the average percentage of uniquely mapped reads and deduped reads compared to total reads are shown above the bar graph. Error bars represent standard error. (D) Correlation between technical replicates of cfDNA TAPS libraries prepared from the same cfDNA samples sequenced to low depth 2.6×. Methylation was calculated in 100 kb windows.
[0061] [Figure 6A] Global cfDNA methylation patterns in cancer and controls. Age and sex distribution of pancreatitis, cirrhosis, PDAC, HCC, and non-cancer control patients in the cfTAPS cohort.
[0062] FIG. 6B: Global cfDNA methylation patterns in cancer and controls. Genome-wide distribution of CpG modifications in cfDNA in non-cancer controls, HCC, and PDAC. Bar plots show the distribution of average CpG modifications for each group. Overlay line plots show the CpG methylation distribution in each patient.
[0063] FIG. 6C: Global cfDNA methylation patterns in cancer and controls. Correlation plot of mean cfDNA CpG modification levels and tumor size (mm) in HCC patients.
[0064] [Figure 6D] Global cfDNA methylation patterns in cancer and controls. Correlation plot of mean cfDNA CpG modification levels and tumor stage in HCC patients.
[0065] FIG. 6E: Global cfDNA methylation patterns in cancer and controls. Correlation plot of tumor size (mm) for PDAC patients. Each point represents an individual patient. The dashed line represents the linear trend fitted with linear regression. The shaded area represents the 95% confidence interval of the fitted model. Pearson correlation coefficients (cor) and P values are indicated on the plot.
[0066] FIG. 6F: Global cfDNA methylation patterns in cancer and controls. Correlation plot of tumor stage for PDAC patients. Each point represents an individual patient. The dashed line represents the linear trend fitted with linear regression. The shaded area represents the 95% confidence interval of the fitted model. Pearson correlation coefficients (cor) and P values are indicated on the plot.
[0067] FIG. 6G: Global cfDNA methylation patterns in cancer and controls. Distribution of CpG modification levels on chromosome 4 in cfDNA of non-cancer controls, HCC and PDAC. Each line represents an individual patient. The mean CpG modification value was calculated for each 1 Mb window along chromosome 4 and Gaussian smoothed (smoothing window size 10).
[0068] [Figure 6H] Global cfDNA methylation patterns in cancer and controls. Methylation variance in 1 Mb genomic windows in non-cancer controls, HCC and PDAC.
[0069] FIG. 6I: Global cfDNA methylation patterns in cancer and controls. PCA plots of cfDNA methylation in 1 kb genomic windows in non-cancer controls and HCC, non-cancer subjects and PDAC (Crohn's disease and colitis are colored green and yellow, respectively).
[0070] FIG. 7A: HCC and PDAC prediction based on cfDNA DMRs. Overview of the training and validation approach of the LOO model. The total number of samples is labeled as n. In each iteration, the model training set consisted of n-1 samples. Differentially methylated enhancers (for HCC) or promoters (for PDAC) were selected for model building. Prediction models were evaluated on hold-out test samples in each fold. Cirrhosis and pancreatitis samples were not included in DMR identification and model building.
[0071] FIG. 7B: HCC and PDAC prediction based on cfDNA DMR. HCC cancer prediction score of cirrhosis samples. Each blue dot represents the prediction score of an individual LOO model. Black dots indicate the average probability score of a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC.
[0072] FIG. 7C: HCC and PDAC prediction based on cfDNA DMR. Gene Ontology analysis of genes associated with differentially methylated enhancers based on HCC cfDNA using Enrichr for NCI-Nature Pathway Interaction (P-value<0.002). The top 10 categories selected based on P-value are shown in the graph. Gene-enhancer interactions were assigned using the GeneHancer reference database.
[0073] FIG. 7D: HCC and PDAC prediction based on cfDNA DMRs. Methylation of representative differentially methylated enhancers in HCC cfDNA for DLC1 gene (two-tailed t-test P value=8.765e-06).
[0074] FIG. 7E: HCC and PDAC prediction based on cfDNA DMR. PDAC cancer prediction score for pancreatitis samples. Each yellow dot represents the prediction score of an individual LOO model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as PDAC.
[0075] FIG. 7F: HCC and PDAC prediction based on cfDNA DMR. Gene Ontology analysis of the closest genes to differentially methylated promoters based on PDAC cfDNA using Enrichr for NCI-Nature Pathway Interaction (P-value < 0.002). The top 10 categories selected based on P-value are shown on the graph.
[0076] FIG. 7G: HCC and PDAC prediction based on cfDNA DMR. Representative differentially methylated promoter methylation in PDAC cfDNA for the RB1 gene (two-tailed t-test P value=0.0017).
[0077] FIG. 7H: HCC and PDAC prediction based on cfDNA DMR. HCC cancer prediction scores of an independent cfDNA WGBS dataset (EGAD00001004317). Each point represents the prediction score of an individual LOO model. Grey points belong to non-cancer controls, and red points belong to HCC. Black points indicate the average probability score of a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC.
[0078] FIG. 7I: HCC and PDAC prediction based on cfDNA DMRs. Proportion of ref DMRs that could be detected in downsampled reads. DMRs identified in the original LOO model training were treated as ref DMRs.
[0079] Figure 8A: Tissue of origin of cfDNA. t-SNE plot of the reference tissue methylation atlas.
[0080] FIG. 8B: Tissue of origin of cfDNA. Average tissue contribution in HCC and PDAC individuals.
[0081] Figure 8C: Tissue of origin of cfDNA. Box plots showing estimated T cell contribution in non-cancer, HCC and PDAC cfDNA samples.
[0082] [Figure 8D] Tissue of origin of cfDNA. ROC curve of model performance using tissue contribution to classify HCC versus non-cancer.
[0083] FIG. 8E: LOO cancer prediction scores for HCC and non-cancer controls using classifiers trained on tissue contributions from tissues of origin of cfDNA. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC.
[0084] FIG. 8F: Tissue of origin of cfDNA. Cancer score of liver cirrhosis samples using HCC vs. non-cancer classifier. Each blue dot represents the prediction score of an individual model. Black dots show the average probability score of a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC.
[0085] FIG. 8G: Tissue of origin of cfDNA. ROC curve of model performance using tissue contribution to classify PDAC versus control.
[0086] FIG. 8H: Tissue of origin of cfDNA. LOO cancer prediction scores for PDAC and non-cancer controls using classifiers built on tissue contribution. The dashed line represents a probability score threshold. Samples with a probability score above this threshold were predicted as PDAC.
[0087] FIG. 8I: Tissue of origin of cfDNA. PDAC cancer scores for pancreatitis samples using PDAC vs. non-cancer classifiers. Each yellow dot represents the prediction score of an individual model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as PDAC.
[0088] FIG. 9A: CNV analysis in cfDNA. Heatmap of CNV estimation from cfDNA in 100 kb bins.
[0089] FIG. 9B: CNV analysis in cfDNA. cfDNA samples with CNVs greater than 500k.
[0090] FIG. 10: cfDNA fragmentation patterns for cancer prediction. (A) Fragment size distribution of cfDNA in public whole genome bisulfite sequencing data. Frequencies were calculated by dividing the number of fragments of a particular length by the total number of fragments. (B) ROC curves of HCC and non-cancer control prediction scores from a generalized linear model using the percentage of long cfDNA fragments (300-500 bp) in 10 bp bins as features. (C) Cancer prediction scores of HCC and non-cancer controls in classifiers trained using LOO cross-validation. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC. (D) HCC cancer prediction scores of cirrhosis samples in these classifiers. Each blue dot represents the prediction score of an individual model. The black dot indicates the average prediction score. The dashed line represents the probability score threshold, and samples with an average probability score above this threshold were predicted as HCC. (E) ROC curves of PDAC and non-cancer control prediction scores from a generalized linear model using the proportion of long cfDNA fragments (300-500 bp) in 10 bp bins as features. (F) LOO cancer prediction scores for PDAC and non-cancer controls for a classifier built on cfDNA fragment frequencies in the 10 bp length range. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as PDAC. (G) PDAC cancer prediction scores for pancreatitis samples for a classifier built on cfDNA fragment frequencies in the 10 bp length range. Each yellow dot represents the prediction score of an individual model. The black dot indicates the average prediction score. The dashed line represents the probability score threshold. Samples with a mean probability score above this threshold were predicted as PDAC.
[0091] FIG. 11A: cfTAPS multi-cancer detection. Performance of methylation, tissue contribution, and fragmentation fraction models in three-class classification. The top panel shows the accuracy of each classifier, and the bottom panel shows actual and predicted patient status in the LOO cross-validation analysis.
[0092] [Figure 11B] Multiple cancer detection by cfTAPS. Heatmap showing the methylation status of selected genomic regions used for cancer type prediction.
[0093] [Figure 11C] Multiple cancer detection by cfTAPS. Gene ontology analysis using Enrichr for NCI-Nature Pathway Interaction for the closest genes of selected DMRs for three-class classification.
[0094] FIG. 12 is a schematic diagram of the different patterns resulting from C to T SNPs and methylated cytosines in the target sequence before and after TAPS. In the diagram, OT means original top, OB means original bottom, CTOT means complementary to original top, and CTOB means complementary to original bottom. [Mode for carrying out the invention] Recently, a bisulfite-free DNA methylation sequencing method, TET-assisted pyridine borane sequencing (TAPS), has been developed, as described in International PCT Application No. PCT / US2019 / 012627, filed January 8, 2019, which claims priority to U.S. Provisional Patent Application No. 62 / 614,798, filed January 8, 2018, U.S. Provisional Patent Application No. 62 / 660,523, filed April 20, 2018, and U.S. Provisional Patent Application No. 62 / 771,409, filed November 26, 2018, each of which is incorporated herein by reference in its entirety. TAPS is based on the use of a mild chemical reaction to directly detect DNA methylation and has shown improved sequence quality, mapping rate, and coverage compared to bisulfite sequencing, while halving the sequencing cost. The combination of direct methylation detection and the non-destructive nature of TAPS is useful not only for DNA methylation analysis, but also for simultaneous genetic analysis in cfDNA, as further described herein, and can enhance non-invasive cancer detection by liquid biopsy.Embodiments of the present disclosure include cfDNA-optimized TAPS (cfTAPS) to deliver high quality and deep whole genome methylome from as low as 10 ng of cfDNA.
[0095] As further described herein, cfTAPS is applied to cfDNA of hepatocellular carcinoma (HCC) and pancreatic ductal adenocarcinoma (PDAC), two types of cancer that have particularly poor prognosis, mainly due to detection at advanced disease stages. Non-invasive methods for early detection of PDAC and HCC are not available, and this method contributes to their late diagnosis. For decades, detection of HCC has relied on liver ultrasound combined with serum alpha-fetoprotein (AFP) measurement. However, these methods have low specificity and sensitivity. There is no blood test to detect or diagnose PDAC. Carbohydrate antigen 19-9 (CA19-9) is used to monitor the treatment and development of PDAC, but its sensitivity and specificity are too low to diagnose or screen PDAC. Therefore, novel approaches for PDAC and HCC detection are urgently needed.
[0096] The results provided herein show that the wealth of information from cfTAPS enables integrated multimodal epigenetic and heritable analysis of differential methylation, tissue of origin, and fragmentation profiles to accurately distinguish cfDNA samples from patients with HCC and PDAC from controls and patients with precancerous inflammatory conditions. In addition, the results provided herein show the successful optimization and application of cfTAPS to characterize genome-wide base-resolution methylomes in cfDNA from HCC, PDAC, and non-cancer controls. Using as little as 10 ng of cfDNA, cfTAPS libraries demonstrated significantly improved sequencing quality and depth compared to previous cfDNA WGBS. Indeed, using less cfDNA input than previous studies, cfDNA TAPS generated the most comprehensive cell-free methylation to date. The much higher yield of informative reads allows cfTAPS to extract more information from a given amount of cfDNA, making it a viable option for large-scale cfDNA methylation studies. Use of TAPS has resulted in superior unique mapping rates and deduplicated unique mapping rates compared to other methods. In some embodiments, the unique mapping rate is at least 65% and / or the unique deduplicated mapping rate is at least 55%. In some embodiments, the unique mapping rate is at least 70% and / or the unique deduplicated mapping rate is at least 60%. In some embodiments, the unique mapping rate is at least 75% and / or the unique deduplicated mapping rate is at least 65%. In some embodiments, the unique mapping rate is at least 80% and / or the unique deduplicated mapping rate is at least 70%. In some embodiments, the unique mapping rate is at least 85% and / or the unique deduplicated mapping rate is at least 72%. In some embodiments, the unique mapping rate is at least 90% and / or the unique deduplicated mapping rate is at least 75%.
[0097] Deep sequencing achieved by cfTAPS allows detailed analysis of cell-free methylome and genome-wide discovery of methylation biomarkers for early cancer detection. Although no significant global hypomethylation was observed, suggesting that the proportion of cfDNA derived from tumor cells is low (supported by the lack of CNVs in most cancer patients included herein), the results show that local methylation signals in regulatory regions such as enhancers and promoters contain cancer-specific information that can accurately distinguish HCC and PDAC from controls. This is particularly important considering the inflammation-enriched real-world control group used in the patient cohort, and the HCC model disclosed herein can correctly identify all HCC and control patients from the cfDNA WGBS dataset as an independent validation.
[0098] Another important advantage of cfDNA methylation for early cancer detection is its ability to determine tissue of origin information. Using currently available public WGBS tissue database, we performed genome-wide tissue deconvolution of cfTAPS data, and the results showed increased liver tumor contribution in HCC cfDNA and distinct immune signatures in cancer cfDNA. Tissue deconvolution itself can be used for cancer detection. Finally, TAPS directly converts modified cytosines, so it retains the most underlying genetic information compared to other approaches that convert unmodified cytosines. In this disclosure, CNV and fragmentation information are extracted from cfTAPS, the latter of which is lost in cfDNA WGBS. The results further showed that an integrated approach that combines differential methylation, tissue of origin, and fragmentation profile can improve model performance for multi-cancer detection.
[0099] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting. 1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. In case of conflict, the present document including definitions shall control. Although preferred methods and materials are described below, methods and materials similar or equivalent to those described herein can also be used in the practice or testing of this disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.
[0100] As used herein, the terms "comprise," "include," "having," "has," "can," "containing," and variations thereof are intended to be open-ended transitional phrases, terms, or words that do not exclude the possibility of additional acts or structures. The singular forms "a," "and," and "the" include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprising," "consisting of," and "consisting essentially of" the embodiments or elements presented herein, whether or not expressly stated.
[0101] For the recitation of numerical ranges herein, each intervening numerical value therebetween to the same degree of precision is expressly contemplated. For example, for the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0102] For the recitation of numerical ranges herein, each intervening numerical value therebetween to the same degree of precision is expressly contemplated. For example, for the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0103] As used herein, "correlates to" refers to compared to.
[0104] As used herein, "methylation" refers to cytosine methylation at the C5 or N4 position of cytosine, the N6 position of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is usually not methylated, since typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplified template. However, "unmethylated DNA" or "methylated DNA" can also refer to amplified DNA whose original template was not methylated or was methylated, respectively.
[0105] As a result, as used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, which is not present in recognized typical nucleotide bases.For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring.Therefore, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide.
[0106] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.
[0107] As used herein, the "methylation state," "methylation profile," "methylation status," and "methylation signature" of a nucleic acid molecule refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule that contains a methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0108] As used herein, "methylation frequency" or "percent (%) methylation" refers to the number of instances where a molecule or locus is methylated compared to the number of instances where the molecule or locus is unmethylated. Methylation state frequencies can be used to describe a population of individuals or a sample from a single individual. For example, a nucleotide locus with a methylation state frequency of 50% is methylated 50% of the time and unmethylated 50% of the time. Such frequencies can be used to describe, for example, the degree to which a nucleotide locus or nucleic acid region is methylated in a population of individuals or a collection of nucleic acids. Thus, if the methylation in a first population or pool of nucleic acid molecules is different from the methylation in a second population or pool of nucleic acid molecules, the frequency of the methylation state of the first population or pool will be different from the frequency of the methylation state of the second population or pool. Such frequencies can also be used to describe, for example, the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, such frequencies can be used to describe the degree to which a group of cells from a tissue sample is methylated or unmethylated at a nucleotide locus or nucleic acid region.
[0109] As used herein, the term "genome-wide cfDNA methylation signature" refers to a signature obtained by any method that interrogates candidate methylation markers across the entire breadth of the genome, rather than a narrow set of candidate sites (as with array-based techniques).
[0110] As used herein, the term "copy number variation" (abbreviated as CNV) refers to a situation in which the copy number of a particular segment of DNA varies between the genomes of different individuals.
[0111] As used herein, the term "unique mapping rate" refers to the criteria used to validate sequencing data, specifically the percentage of sequencing reads that map to exactly one location in the reference genome. In some embodiments, the unique mapping rate can be calculated as the percentage of reads with a defined parameter (e.g., 500, 120, 1000, 20) (e.g., MAPQ≧1 using bwa alignment) compared to the total number of sequenced reads.
[0112] As used herein, the term "unique de-duplicated mapping rate" refers to the proportion of de-duplicated sequencing reads (after removing duplicates) that map to exactly one location in a reference genome. In some preferred embodiments, the unique de-duplicated mapping rate can be determined by calculating the proportion of properly mapped reads after removing PCR duplicates (e.g., using MarkDuplicates (Picard)) compared to the total number of sequenced reads.
[0113] As used herein, the term "tissue deconvolution" refers to sorting the sequenced cfDNA in a sample into its tissue of origin and determining the relative contribution from the tissue.In some preferred embodiments, cfDNA methylation is compared with the methylation value (e.g., at DMR) in a reference atlas.These methods preferably use a regression method in which the origin percentage of cfDNA is the regression coefficient.
[0114] As used herein, the term "patient" or "subject" refers to an organism that is the subject of the various tests provided by the present technology. The term "subject" includes animals, preferably mammals, including humans. In a preferred embodiment, the subject is a primate. In an even more preferred embodiment, the subject is a human. Further with respect to the diagnostic method, the preferred subject is a vertebrate subject. The preferred vertebrates are warm-blooded animals, and the preferred warm-blooded vertebrates are mammals. The preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Thus, veterinary therapeutic uses are provided herein. As such, the present technology is provided for the diagnosis of mammals, such as humans, and those mammals that are economically important, such as animals raised on farms for human consumption, important because they are endangered, such as the Amur tiger, and / or animals that are socially important to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores such as cats and dogs; swine, including pigs, hogs, and wild boars; ruminants and / or ungulates, such as cattle, oxen, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses. 2. TET-assisted pyridine borane sequencing (TAPS)
[0005] Embodiments of the present disclosure provide a bisulfite-free base resolution method (TAPS) for detecting 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) in sequences, including for use with circulating cell-free DNA. As disclosed in International PCT Application No. PCT / US2019 / 012627 (which claims priority to U.S. Provisional Patent Applications Nos. 62 / 614,798, filed January 8, 2018, 62 / 660,523, filed April 20, 2018, and 62 / 771,409, filed November 26, 2018, each of which is incorporated herein by reference in its entirety), TAPS involves the use of mild enzymatic and chemical reactions to directly and quantitatively detect 5mC and 5hmC at base resolution without affecting unmodified cytosines. The present disclosure also provides methods for detecting 5-formylcytosine (5fC) and 5-carboxylcytosine (5caC) at base resolution without affecting unmodified cytosines. Thus, the methods provided herein provide for mapping of 5mC, 5hmC, 5fC, and 5caC, overcoming the shortcomings of conventional methods such as bisulfite sequencing.
[0115] Methods for identifying 5mC. In some embodiments, the methods of the present disclosure include identifying 5mC in a DNA sample (target DNA or whole genome) and providing a quantitative measure of the frequency of 5mC modification at each position where the modification is identified in the DNA. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5mC at each position in the DNA. According to these embodiments, the methods for identifying 5mC may include the use of a protecting group. In other embodiments, the methods for identifying 5mC do not require the use of a protecting group (e.g., cfTAPS, further described below).
[0116] When using a protecting group to identify 5mC in DNA (e.g., cfDNA) without 5hmC, the 5hmC in the sample is protected from conversion to 5caC and / or 5fC. In some embodiments, the 5hmC in the sample DNA is rendered unreactive to subsequent steps by adding a protecting group to the 5hmC. In one embodiment, the protecting group is a sugar, including modified sugars, such as glucose or 6-azido-glucose (6-azido-6-deoxy-D-glucose). The sugar protecting group can be added to the hydroxymethyl group of 5hmC by contacting the DNA sample with a uridine diphosphate (UDP) sugar in the presence of one or more glucosyltransferase enzymes. In some embodiments, the glucosyltransferase is T4 bacteriophage β-glucosyltransferase (βGT), T4 bacteriophage α-glucosyltransferase (αGT), and derivatives and analogs thereof. βGT is an enzyme that catalyzes the chemical reaction in which a beta-D-glucosyl (glucose) residue is transferred from UDP-glucose to 5-hydroxymethylcytosine residues in nucleic acids.
[0117] Methods for identifying 5hmC. In some embodiments, the methods of the present disclosure include identifying 5mC or 5hmC in a DNA sample (target DNA or whole genome). In some embodiments, the methods provide a quantitative measure of the frequency of 5mC or 5hmC modification at each position where the modification is identified in the DNA. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5mC or 5hmC at each position in the DNA. According to these embodiments, the methods for identifying 5mC or 5hmC provide the positions of 5mC and 5hmC, but do not distinguish between the two cytosine modifications. Rather, both 5mC and 5hmC are converted to DHU. The presence of DHU can be detected directly, or the modified DNA can be replicated by known methods, where DHU is converted to T. In some embodiments, the methods for identifying 5hmC include the use of a protecting group. In other embodiments, the methods for identifying 5hmC do not require the use of a protecting group (e.g., cfTAPS, further described below).
[0118] Method for identifying 5mC and / or method for identifying 5hmC. The present disclosure provides a method for identifying 5mC and identifying 5hmC in DNA (e.g., cfDNA) by carrying out a method for identifying 5mC on a first DNA sample and carrying out a method for identifying 5mC or 5hmC on a second DNA sample. In some embodiments, the first and second DNA samples are derived from the same DNA sample. For example, the first and second samples can be separate aliquots taken from a sample that contains DNA (e.g., cfDNA) to be analyzed.
[0119] 5mC and 5hmC (unprotected) are converted to 5fC and 5caC before conversion to DHU, so 5fC and 5caC present in DNA samples are detected as 5mC and / or 5hmC. However, considering that 5fC and 5caC are at very low levels in genomic DNA under normal conditions, this is often acceptable when analyzing methylation and hydroxymethylation in DNA samples. 5fC and 5caC signals can be eliminated by protecting 5fC and 5caC from conversion to DHU, for example, by hydroxylamine conjugation and EDC coupling, respectively. According to these embodiments, the method identifies the location and percentage of 5hmC in DNA by comparing the location and percentage of 5mC with the location and percentage of 5mC or 5hmC (combined). Alternatively, the location and frequency of 5hmC modification in DNA can be measured directly.
[0120] In some embodiments, the step of converting 5hmC to 5fC includes oxidizing 5hmC to 5fC by contacting the DNA with, for example, potassium perruthenate (KRuO4) (as described in Science. 2012, 33, 934-937 and WO2013017853, which are incorporated herein by reference) or Cu(II) / TEMPO (copper(II) perchlorate and 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO)) (as described in Chem. Commun., 2017, 53, 5756-5759 and WO2017039002, which are incorporated herein by reference). The 5fC in the DNA sample is then converted to DHU by methods disclosed herein (e.g., by borane reaction).
[0121] In some embodiments, identifying 5fC and / or 5caC provides the location of 5fC and / or 5caC but does not distinguish between these two cytosine modifications. Rather, both 5fC and 5caC are converted to DHU, which is detected by the methods described herein.
[0122] Methods for identifying 5caC. In some embodiments, the methods include identifying 5caC in a DNA sample (target DNA or whole genome) and provide a quantitative measure of the frequency of 5caC modification at each position where the modification is identified in the DNA. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5caC at each position in the DNA. According to these embodiments, the methods for identifying 5caC may include the use of a protecting group. In other embodiments, the methods for identifying 5caC do not require the use of a protecting group (e.g., cfTAPS, which is further described below).
[0123] In some embodiments, identification of 5caC in DNA can be performed when 5fC is protected (and 5mC and 5hmC are not converted to DHU). In some embodiments, adding a protecting group to 5fC in a DNA sample includes contacting the DNA with an aldehyde-reactive compound, including, for example, hydroxylamine derivatives, hydrazine derivatives, and hydrazide derivatives. Hydroxylamine derivatives include ashydroxylamine, hydroxylamine hydrochloride, hydroxylammonium sulfate acid, hydroxylamine phosphate, O-methylhydroxylamine, O-hexylhydroxylamine, O-pentylhydroxylamine, O-benzylhydroxylamine, particularly O-ethylhydroxylamine (EtONH2), O-alkylated or O-arylated hydroxylamines, acids or salts thereof. Hydrazine derivatives include N-alkylhydrazine, N-arylhydrazine, N-benzylhydrazine, N,N-dialkylhydrazine, N,N-diarylhydrazine, N,N-dibenzylhydrazine, N,N-alkylbenzylhydrazine, N,N-arylbenzylhydrazine, and N,N-alkylarylhydrazine. Hydrazide derivatives include -toluenesulfonylhydrazide, N-acylhydrazide, N,N-alkylacylhydrazide, N,N-benzylacylhydrazide, N,N-arylacylhydrazide, N-sulfonylhydrazide, N,N-alkylsulfonylhydrazide, N,N-benzylsulfonylhydrazide, and N,N-arylsulfonylhydrazide.
[0124] Method for identifying 5fC. In some embodiments, the method comprises identifying 5fC in a DNA sample (target DNA or whole genome) and provides a quantitative measure of the frequency of 5fC modification at each position where modification is identified in DNA. In some embodiments, the proportion of T at each transition position provides a quantitative level of 5fC at each position in DNA. According to these embodiments, the method for identifying 5fC can include the use of a protecting group. In other embodiments, the method for identifying 5fC does not require the use of a protecting group (e.g., cfTAPS, which is further described below).
[0125] In some embodiments, adding a blocking group to 5caC in a DNA sample can be accomplished by (i) contacting the DNA sample with a coupling agent, e.g., a carboxylic acid derivatizing agent, e.g., a carbodiimide derivative such as l-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) or N,N'-dicyclohexylcarbodiimide (DCC), and (ii) contacting the DNA sample with an amine, hydrazine, or hydroxylamine compound. Thus, for example, 5caC can be protected by treating the DNA sample with EDC, followed by benzylamine, ethylamine, or another amine to form an amide that protects 5caC from conversion to DHU (e.g., by pic-BH3). 3. TAPS for cfDNA (cfTAPS) The present disclosure provides a cfDNA-optimized TAPS (cfTAPS) method to provide a high-quality and high-depth whole genome cell-free methylome. As described further below, in one embodiment of the present disclosure, cfTAPS was applied to 85 cfDNA samples from patients with hepatocellular carcinoma (HCC) or pancreatic ductal adenocarcinoma (PDAC), and non-cancer controls. The most comprehensive cfDNA methylome to date was generated from as little as 10 ng of cfDNA (1-3 mL of plasma). The results provided herein showed that cfTAPS provides multimodal information on cfDNA characteristics, including DNA methylation, tissue of origin, and DNA fragmentation. The integrated analysis of these epigenetic and genetic features allows for accurate identification of early stage HCC and PDAC. Because the methods of the present disclosure utilize mild enzymatic and chemical reactions that avoid the substantial degradation of nucleic acids associated with methods such as bisulfite sequencing, the methods of the present disclosure are useful for analysis of low input samples, such as circulating cell-free DNA and single cell analysis.
[0126] According to these embodiments, the present disclosure provides a method for obtaining a methylation signature.In some embodiments, the method includes isolating cell-free DNA (cfDNA) from a sample, preparing a sequencing library comprising cfDNA, and performing TET-assisted pyridine borane sequencing (TAPS) on the sequencing library to obtain a methylation signature of cfDNA.In some embodiments, the methylation signature is a whole genome methylation signature.
[0127] In some embodiments, the use of carrier DNA results in higher library yields. As will be recognized by those of skill in the art based on the present disclosure, carrier DNA can be obtained by any means known to those of skill in the art, including, but not limited to, PCR amplification from a vector or plasmid template using one or more primers. In some embodiments, at least 1 ng of carrier DNA can be used. In some embodiments, at least 10 ng of carrier DNA can be used. In some embodiments, at least 25 ng of carrier DNA can be used. In some embodiments, at least 50 ng of carrier DNA can be used. In some embodiments, at least 100 ng of carrier DNA can be used. In some embodiments, at least 150 ng of carrier DNA can be used. In some embodiments, at least 200 ng of carrier DNA can be used. In some embodiments, at least 250 ng of carrier DNA can be used. In some embodiments, at least 500 ng of carrier DNA can be used. In some embodiments, about 1 ng to about 500 ng of carrier DNA can be used. In some embodiments, about 1 ng to about 500 ng of carrier DNA can be used. In some embodiments, about 50 ng to about 250 ng of carrier DNA can be used. In some embodiments, about 75 ng to about 150 ng of carrier DNA can be used. In some embodiments, about 50 ng to about 150 ng of carrier DNA can be used. In some embodiments, about 75 ng to about 125 ng of carrier DNA can be used.
[0128] In some embodiments, and as described herein, the method further comprises identifying at least one methylation biomarker from the cfDNA whole genome methylation signature and determining whether the methylation biomarker is indicative of cancer. In some embodiments, the methylation biomarker comprises a differentially methylated region (DMR). In some embodiments, the method further comprises classifying the sample based on the DMR compared to a reference DMR. In some embodiments, the reference DMR corresponds to a non-cancerous control or a cancerous control.
[0129] In some embodiments, and as described herein, the method further comprises identifying at least one methylation biomarker from the cfDNA whole genome methylation signature and determining a tissue of origin corresponding to the methylation biomarker. In some embodiments, the method further comprises classifying the sample based on the tissue of origin biomarker.
[0130] In some embodiments, and as described herein, the method further comprises identifying a DNA fragmentation profile and determining whether the fragmentation profile is indicative of cancer. According to these embodiments, the DNA fragmentation profile can be determined from cfTAPS whole genome sequencing data (e.g., alignment position of read pairs). In some preferred embodiments, the sequenced reads from cfTAPS are first aligned to a reference genome. Then, the length of cfDNA fragments is extracted from the alignment file generated from the sequencing data. The proportion of cfDNA fragments in 10bp intervals is used as the fragmentation profile of cell-free DNA.
[0131] In some embodiments, the method further includes identifying at least one sequence variant from the cfDNA and determining whether the sequence variant is indicative of cancer. For example, in some embodiments, cfTAPS can also distinguish methylation from C to T genetic variants or single nucleotide polymorphisms (SNPs), and thus can be used to detect genetic variants. In some embodiments, methylation and C to T SNPs can result in different patterns in cfTAPS. For example, methylation can result in T / G reads in the original top strand / original bottom strand and A / C reads in the complementary strand. In some embodiments, C to T SNPs can result in T / A reads in the original top strand / original bottom strand and the complementary strand. These different patterns are shown in FIG. 12. This further increases the utility of cfTAPS in providing both methylation information and genetic variants, and therefore mutations, in one experiment and sequencing run. This capability of the cfTAPS method disclosed herein provides for the integration of genomic and epigenetic analysis and a substantial reduction in sequencing costs by eliminating the need to perform standard whole genome sequencing (WGS).
[0132] According to the above embodiment, the method of the present disclosure includes the use of cfTAPS to generate information on methylation signature, methylation biomarker, DNA fragment profile, DNA sequence information (e.g., variant), and tissue of origin information in a single experiment to diagnose / detect cancer in a subject. As will be understood by those skilled in the art based on this disclosure, cfTAPS disclosed herein can be used to generate any combination of methylation signature, methylation biomarker, DNA fragment profile, DNA sequence information (e.g., variant), and tissue of origin information to diagnose / detect cancer in a subject. In some embodiments, a methylation signature can be obtained, and one or more of a methylation biomarker, a DNA fragment profile, DNA sequence information (e.g., variant), and tissue of origin information can also be obtained, which can be used to diagnose / detect cancer in a subject. In some embodiments, a methylation status of a biomarker can be obtained, and one or more of a methylation signature, a DNA fragment profile, DNA sequence information (e.g., variant), and tissue of origin information can also be obtained, which can be used to diagnose / detect cancer in a subject. In some embodiments, a DNA fragment profile can be obtained, and one or more of a methylation signature, a methylation biomarker, a DNA sequence information (e.g., variants), and information of tissue of origin can also be obtained, which can be used to diagnose / detect cancer in a subject. In some embodiments, a DNA sequence variant can be identified, and one or more of a methylation signature, a methylation biomarker, a DNA fragment profile, and information of tissue of origin can also be obtained, which can be used to diagnose / detect cancer in a subject. In some embodiments, information of tissue of origin can be obtained (e.g., from a whole genome cfDNA methylation signature), and one or more of a methylation signature, a methylation biomarker, a DNA fragment profile, and information of DNA sequence (e.g., variants), which can also be used to diagnose / detect cancer in a subject.
[0133] Thus, in some preferred embodiments, the invention provides a multimodal method of analyzing cfDNA in a patient sample, comprising: isolating cfDNA from the patient sample; converting 5mC and / or 5hmC residues in the sample to DHU residues to provide a modified cfDNA sample; and sequencing the modified cfDNA sample to identify methylated regions in the sample, where a cytosine (C) to thymine (T) transition or a cytosine (C) to DHU transition in the modified cfDNA sample provides a position of either 5mC or 5hmC in cfDNA, as compared to an unmodified reference cfDNA. a) determining copy number variation of one or more targets in a modified cfDNA sample; b) determining the tissue of origin or one or more targets in the modified cfDNA sample; c) determining the fragmentation profile of the modified cfDNA sample; and and d) performing one or more additional analytical steps on the modified cfDNA selected from the group consisting of: (a) identifying one or more single nucleotide mutations in the modified cfDNA sample.
[0134] In some preferred embodiments, the one or more additional steps is step a. In some preferred embodiments, the one or more additional steps is step b. In some preferred embodiments, the one or more additional steps is step c. In some preferred embodiments, the one or more additional steps is step d.
[0135] In some preferred embodiments, the one or more additional steps are steps a and b. In some preferred embodiments, the one or more additional steps are steps a and c. In some preferred embodiments, the one or more additional steps are steps a and d. In some preferred embodiments, the one or more additional steps are steps b and c. In some preferred embodiments, the one or more additional steps are steps b and d. In some preferred embodiments, the one or more additional steps are steps c and d.
[0136] In some preferred embodiments, the one or more additional steps are steps a, b, and c. In some preferred embodiments, the one or more additional steps are steps a, b, and d. In some preferred embodiments, the one or more additional steps are steps b, c, and d.
[0137] In some preferred embodiments, the one or more additional steps are all of steps a, b, c, and d.
[0138] In some embodiments, the unmodified reference cfDNA to which the modified cfDNA sample is compared can include any unmodified reference cfDNA, including, for example, publicly available reference cfDNA or an unmodified control sample from a patient.
[0139] In some embodiments, performing TAPS on the sequencing library to obtain a whole genome methylation signature comprises identifying a 5mC modification in cfDNA and providing a quantitative measure of the frequency of the 5mC modification. In some embodiments, performing TAPS on the sequencing library to obtain a whole genome methylation signature comprises identifying a 5hmC modification in cfDNA and providing a quantitative measure of the frequency of the 5hmC modification. In some embodiments, performing TAPS on the sequencing library to obtain a whole genome methylation signature comprises identifying a 5caC modification in cfDNA and providing a quantitative measure of the frequency of the 5caC modification. In some embodiments, performing TAPS on the sequencing library to obtain a whole genome methylation signature comprises identifying a 5fC modification in cfDNA and providing a quantitative measure of the frequency of the 5fC modification.
[0140] As will be recognized by those skilled in the art based on the present disclosure, the methods described herein (e.g., cfTAPS) can be used to diagnose / detect any type of cancer. The types of cancer that can be detected / diagnosed using the methods of the present disclosure include, but are not limited to, lung cancer, melanoma, colon cancer, colorectal cancer, neuroblastoma, breast cancer, prostate cancer, renal cell carcinoma, transitional cell carcinoma, bile duct cancer, brain cancer, non-small cell lung cancer, pancreatic cancer, liver cancer, gastric cancer, bladder cancer, esophageal cancer, mesothelioma, thyroid cancer, head and neck cancer, osteosarcoma, hepatocellular carcinoma, carcinoma of unknown primary, ovarian cancer, endometrial cancer, glioblastoma, Hodgkin's lymphoma, and non-Hodgkin's lymphoma. In some embodiments, the types of cancer or metastatic forms of cancer that can be detected / diagnosed by the methods of the present disclosure include, but are not limited to, carcinoma, sarcoma, lymphoma, germ cell tumor, and blastoma. In some embodiments, the cancer is an invasive and / or metastatic cancer (e.g., stage II cancer, stage III cancer, or stage IV cancer). In some embodiments, the cancer is an early stage cancer (e.g., stage 0 cancer, stage I cancer) and / or is not an invasive and / or metastatic cancer.
[0141] In some embodiments, the method of the present disclosure (e.g., cfTAPS) can be used to determine whether a subject has hepatocellular carcinoma (HCC) or pancreatic ductal adenocarcinoma (PDAC).In some embodiments, the method includes determining whether a subject has early stage hepatocellular carcinoma (HCC) or early stage pancreatic ductal adenocarcinoma (PDAC).
[0142] According to these embodiments, the present disclosure provides a method for quantitatively identifying one or more positions of 5mC, 5hmC, 5caC and / or 5fC in a nucleic acid with base resolution without affecting unmodified cytosines. In some embodiments, the nucleic acid is DNA. In some embodiments, the DNA is cfDNA (e.g., circulating cfDNA). In some embodiments, the nucleic acid is RNA. In some embodiments, the nucleic acid sample includes a target nucleic acid that is DNA or a target nucleic acid that is RNA. In some embodiments, the method is applied to the whole genome and is not limited to a specific target nucleic acid.
[0143] The nucleic acid can be any nucleic acid with cytosine modifications (i.e., 5mC, 5hmC, 5fC, and / or 5caC). The nucleic acid can be a single nucleic acid molecule in a sample, or the entire population of nucleic acid molecules in a sample (the entire genome or a subset thereof). The nucleic acid can be a natural nucleic acid from a source (e.g., a cell, tissue sample, etc.), or can be pre-converted into a form compatible with high-throughput sequencing, for example, by fragmentation, repair, and ligation with adapters for sequencing. Thus, the nucleic acid can include multiple nucleic acid sequences, such that the methods described herein can be used to generate libraries of target nucleic acid sequences that can be analyzed individually (e.g., by sequencing individual targets) or in groups (e.g., by high-throughput or next-generation sequencing methods).
[0144] Nucleic acid samples can be obtained from organisms from Monera (Bacteria), Protista, Fungi, Plantae, and Animalia Kingdoms. Nucleic acid samples may be obtained from a patient or subject, from an environmental sample, or from an organism of interest. In some embodiments, the sample is obtained from a human subject / patient, including but not limited to a human with cancer or suspected of having cancer. In some embodiments, the sample is obtained from tissues or cells from a human (e.g., obtained from a biopsy), including tissues or cells that are cancerous or suspected of being cancerous. In some embodiments, the nucleic acid sample is extracted from or derived from a cell or collection of cells, bodily fluids, tissue samples, organs, and organelles. In some embodiments, the nucleic acid sample is obtained from bodily fluids including, but not limited to, blood (plasma, serum, whole blood), urine, feces / fecal fluid, semen (male reproductive fluid), vaginal secretions, cerebrospinal fluid (CSF), ascites, synovial fluid, pleural fluid (pleural lavage), pericardial fluid, peritoneal fluid, amniotic fluid, saliva, nasal fluid, ear fluid, gastric fluid, breast milk, and any other bodily fluid that contains cfDNA, as well as cell culture supernatant. In some embodiments, the sample is obtained from a bodily fluid that is cancerous or suspected to be cancerous. The disclosed method utilizes mild enzymatic and chemical reactions that avoid substantial degradation of nucleic acids associated with methods such as bisulfite sequencing, making the disclosed method useful for analysis of low input samples, such as circulating cell-free DNA and single cell analysis.
[0145] In some embodiments, the DNA sample contains picogram amounts of DNA. In some embodiments, the DNA sample contains about 1 pg to about 900 pg of DNA, about 1 pg to about 500 pg of DNA, about 1 pg to about 100 pg of DNA, about 1 pg to about 50 pg of DNA, or about 1 to about 10 pg of DNA. In some embodiments, the DNA sample contains less than about 200 pg, less than about 100 pg of DNA, less than about 50 pg of DNA, less than about 20 pg of DNA, less than about 15 pg of DNA, less than about 10 pg of DNA, or less than about 5 pg of DNA.
[0146] In some embodiments, the DNA sample comprises nanogram amounts of DNA. Sample DNA for use in the disclosed methods can be any amount, including but not limited to DNA from a single cell or bulk DNA samples. In some embodiments, the methods can be performed on DNA samples comprising about 1 to about 500 ng of DNA, about 1 to about 200 ng of DNA, about 1 to about 100 ng of DNA, about 1 to about 50 ng of DNA, about 1 to about 10 ng of DNA, about 2 to about 5 ng of DNA. In some embodiments, the DNA sample comprises less than about 100 ng of DNA, less than about 50 ng of DNA, less than 40 ng of DNA, less than 30 ng of DNA, less than 20 ng of DNA, less than 15 ng of DNA, less than 5 ng of DNA, and less than 2 ng of DNA. In some embodiments, the DNA sample comprises microgram amounts of DNA.
[0147] The DNA sample used in the methods described herein may be from any source, including, for example, bodily fluids, tissue samples, organs, organelles, cells or collections of cells. In some embodiments, the DNA sample is obtained from a human subject / patient, including, but not limited to, a human with cancer or suspected of having cancer. In some embodiments, the DNA sample is obtained from tissues or cells from a human, including tissues or cells that are cancerous or suspected of being cancerous (e.g., obtained from a biopsy). In some embodiments, the DNA sample is extracted from or derived from a cell or collection of cells, bodily fluids, tissue samples, organs, and organelles. In some embodiments, the DNA sample is obtained from bodily fluids, including, but not limited to, blood (plasma, serum, whole blood), urine, feces / fecal fluid, semen (male reproductive fluid), vaginal secretions, cerebrospinal fluid (CSF), peritoneal fluid, synovial fluid, pleural fluid (pleural lavage), pericardial fluid, peritoneal fluid, amniotic fluid, saliva, nasal fluid, ear fluid, gastric fluid, breast milk, and any other bodily fluid that contains cfDNA, as well as cell culture supernatants. In some embodiments, the DNA sample is obtained from a body fluid that is cancerous or suspected to be cancerous. In some embodiments, the DNA sample is circulating cell-free DNA (cell-free DNA or cfDNA), which is DNA found in blood and not in cells. As will be understood by those skilled in the art based on this disclosure, cfDNA can be isolated from body fluids using methods known in the art. Commercially available kits are available for isolating cfDNA, including, for example, Circulating Nucleic Acid Kit (Qiagen). DNA samples can be obtained from enrichment steps, including but not limited to antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0148] The DNA can be any DNA with cytosine modifications (i.e., 5mC, 5hmC, 5fC, and / or 5caC), including, but not limited to, DNA fragments and / or genomic DNA. The DNA can be a single DNA molecule in a sample, or the entire population of DNA molecules in a sample (whole genome or a subset thereof). The DNA can be native DNA from a source, or can be pre-converted into a form compatible with high-throughput sequencing, for example, by fragmentation, repair, and ligation with adapters for sequencing. Thus, the DNA can include multiple DNA sequences, such that the methods described herein can be used to generate libraries of target DNA sequences that can be analyzed individually (e.g., by sequencing individual targets) or in groups (e.g., by high-throughput or next-generation sequencing methods).
[0149] According to these embodiments, the methods of the disclosure include a step of converting 5mC and 5hmC (or only 5mC if 5hmC is protected) to 5caC and / or 5fC. In some embodiments, this step includes contacting the DNA or RNA sample with a ten-eleven translocation (TET) enzyme. TET enzymes are a family of enzymes that catalyze the transfer of an oxygen molecule to the C5 methyl group at 5mC, resulting in the formation of 5-hydroxymethylcytosine (5hmC). TET further catalyzes the oxidation of 5hmC to 5fC and the oxidation of 5fC to form 5caC. TET enzymes useful in the methods of the disclosure include one or more of human TET1, TET2, and TET3; mouse TET1, TET2, and TET3; Naegleria TET (NgTET); Coprinopsis cinerea (CcTET); the catalytic domain of mouse TET1 (mTET1CD), and derivatives or analogs thereof. In some embodiments, the TET enzyme is NgTET. In some embodiments, the TET enzyme is human TET1 (hTET1). In some embodiments, the TET enzyme is mTET1CD.
[0150] The disclosed methods may also include a step of converting 5caC and / or 5fC in the nucleic acid sample to DHU. In some embodiments, this step includes contacting the DNA or RNA sample with a reducing agent, including borane reducing agents such as, for example, pyridine borane, 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride. In some embodiments, the reducing agent is pyridine borane and / or pic-BH3.
[0151] The method of the present disclosure may also include amplifying the copy number of the modified nucleic acid by methods known in the art. When the modified nucleic acid is DNA, the copy number can be increased by, for example, PCR, cloning, and primer extension. The copy number of each target DNA can be amplified by PCR using primers specific to a particular target DNA sequence. Alternatively, multiple different modified target DNA sequences can be amplified by cloning into a DNA vector by standard techniques. In some embodiments, the copy number of multiple different modified target DNA sequences is increased by PCR, for example, to generate a library for next-generation sequencing, in which a double-stranded adapter DNA is pre-ligated to the sample DNA (or modified sample DNA) and PCR is performed using a primer complementary to the adapter DNA.
[0152] In some embodiments, the method comprises detecting the sequence of the modified nucleic acid. The modified target DNA or RNA contains DHU at the position where one or more of 5mC, 5hmC, 5fC, and 5caC were present in the unmodified target DNA or RNA. DHU acts as T in DNA replication and sequencing methods. Thus, cytosine modification can be detected by any direct or indirect method known in the art that identifies C to T transitions. Such methods include sequencing methods such as Sanger sequencing, microarrays, and next-generation sequencing methods. C to T transitions can also be detected by restriction enzyme analysis, where the C to T transitions eliminate or introduce restriction endonuclease recognition sequences.
[0153] Embodiments of the present disclosure also provide kits for identifying 5mC and 5hmC in DNA. Such kits include reagents for identifying 5mC and 5hmC by the methods described herein. The kits may include reagents for identifying 5caC and reagents for identifying 5fC by the methods described herein. In some embodiments, the kits include a TET enzyme, a borane reducing agent, and instructions for carrying out the method. In some embodiments, the TET enzyme is TET1, and the borane reducing agent is selected from one or more of the group consisting of pyridine borane, 2-picoline borane (pic-BH3), borane, sodium borohydride, sodium cyanoborohydride, and sodium triacetoxyborohydride. In some embodiments, the TET1 enzyme is NgTet1 or mouse Tet1 (e.g., mTet1CD), and the borane reducing agent is pyridine borane and / or pic-BH3.
[0154] In some embodiments, the kit further comprises a 5hmC protecting group and a glycosyltransferase enzyme. In some embodiments, the protecting group added to the 5hmC is a sugar. In some embodiments, the sugar is a naturally occurring sugar or a modified sugar, e.g., glucose or a modified glucose. In some embodiments, the protecting group is added to the 5hmC by contacting the nucleic acid sample with a UDP linked to a sugar, e.g., UDP-glucose, or a UDP linked to a modified glucose, in the presence of a glucosyltransferase enzyme, e.g., T4 bacteriophage β-glucosyltransferase (βGT) and T4 bacteriophage α-glucosyltransferase (αGT), and derivatives and analogs thereof.
[0155] In some embodiments, the kit further comprises an oxidizing agent selected from potassium perruthenate (KRuO4) and / or Cu(II) / TEMPO (copper(II) perchlorate and 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO)). In some embodiments, the kit comprises a reagent for protecting 5fC in a nucleic acid sample. In some embodiments, the kit comprises an aldehyde reactive compound, including, for example, hydroxylamine derivatives, hydrazine derivatives, and hydrazide derivatives, as described herein. In some embodiments, the kit comprises a reagent for protecting 5caC, as described herein. In some embodiments, the kit comprises a reagent for isolating DNA or RNA. In some embodiments, the kit comprises a reagent for isolating low input DNA from a sample, for example, cfDNA from blood, plasma, or serum.
[0156] In some embodiments, the methods of the present disclosure include treating a patient (e.g., a patient with cancer, a patient with early stage cancer, or a patient suspected of having cancer). In some embodiments, the methods include determining a methylation signature provided herein and administering a treatment to the patient based on the results of determining the methylation signature. Treatment may include administering a pharmaceutical compound, a vaccine, performing surgery, imaging the patient, and / or performing another test. In some embodiments, the methods of the present disclosure can be used as part of clinical screening, methods of prognostic evaluation, methods of monitoring the outcome of a therapy, methods for identifying patients most likely to respond to a particular therapeutic treatment, methods of imaging a patient or subject, and methods for drug screening and development.
[0157] In some embodiments, the method of the present disclosure includes diagnosing cancer in a subject. The terms "diagnose" and "diagnosis" as used herein refer to a method by which a person skilled in the art can estimate and even determine whether a subject suffers from a given disease or condition, or is likely to develop a given disease or condition in the future. A person skilled in the art often makes a diagnosis based on one or more diagnostic indicators, such as methylation biomarkers and / or methylation signatures, which are indicative of the presence, severity, or absence of a condition (e.g., cancer).
[0158] In addition to diagnosis, clinical cancer prognosis involves determining the aggressiveness of cancer and the likelihood of tumor recurrence, and planning the most effective therapy. If it is possible to make a more accurate prognosis, or even if it is possible to evaluate the potential risk of suffering from cancer, it is possible to select an appropriate treatment, and in some cases a less harsh treatment, for the patient. Evaluation of a subject based on methylation signatures can be useful to separate subjects who have a good prognosis and / or are at low risk of developing cancer and do not require therapy or limited therapy, from subjects who are likely to develop cancer or suffer from cancer recurrence and may benefit from more intensive treatment. Thus, "making a diagnosis" or "diagnosing", as used herein, further includes making a determination of the risk of developing cancer or determining a prognosis (which can provide a prediction of clinical outcome (with or without medical treatment)), selecting an appropriate treatment (or whether a treatment is effective), or monitoring and potentially changing current treatment, based on the identification and evaluation of methylation signatures, as disclosed herein.
[0159] In some embodiments, the method of the disclosure includes determining whether to initiate or continue prevention or treatment of cancer in a subject. In some embodiments, the method includes providing a series of biological samples from a subject over a period of time, analyzing the series of biological samples to determine a methylation signature disclosed herein in each biological sample, and comparing any measurable changes in the methylation signature in each biological sample. Any changes in the methylation signature over a period of time can be used to predict the risk of developing cancer, predict clinical outcomes, determine whether to initiate or continue cancer prevention or therapy, and determine whether a current therapy is effectively treating cancer. For example, a first time point can be selected before the start of treatment and a second time point can be selected at a time point after the start of treatment. The methylation signature can be measured in each of the samples taken from the different time points, and qualitative and / or quantitative differences are shown. The changes in the methylation signature from the different samples can be correlated with the risk of developing cancer, prognosis, determining the effectiveness of treatment, and / or progression of cancer in the subject. In some embodiments, the methods and compositions of the invention are for the treatment or diagnosis of disease at an early stage, e.g., before symptoms of the disease are manifest, hi some embodiments, the methods and compositions of the invention are for the treatment or diagnosis of disease at a clinical stage.
[0160] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by those skilled in the art. For example, any nomenclature and techniques used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are well known and commonly used in the art. The meaning and scope of the terms must be clear. However, in the event of potential ambiguity, the definitions provided herein take precedence over any dictionary or extrinsic definitions. Furthermore, unless otherwise required by context, singular terms shall include the plural and plural terms shall include the singular. 4. Materials and Methods Experimental design. Whole blood samples from 30 non-cancer controls were obtained from John Radcliffe Hospital (ethical approval IDs 16 / YH / 0247 and 18 / WM / 0237). Pancreatitis blood samples from 8 patients were obtained from John Radcliffe Hospital. The study was approved by Oxfordshire REC-A (10 / H0604 / 51) and registered in the UK NIHR portfolio under study number 10776. PDAC patients consented for the study by Oxford Radcliffe Biobank (09 / H0606 / 5+5, project: 19 / A177) and whole blood samples were collected from 24 patients. Collection of plasma samples from 21 HCC patients and 4 cirrhosis patients was REC approved (ethical approval 2 / NE / 0395, IRAS project ID: 116370). No sample size calculation was performed. Sample size was determined based on availability. PDAC, HCC, pancreatitis, and cirrhosis samples were collected from subjects with clinically diagnosed disease. Non-cancer control samples were collected from individuals with no cancer diagnosis or history of cancer at the time of sample collection.
[0161] The primary goal of this study was the comprehensive multidimensional characterization of cfDNA in cancers and controls by whole-genome methylation sequencing using TAPS. cfDNA TAPS libraries were constructed and paired-end 150 bp sequenced on a NovaSeq 6000 sequencer (Illumina). Technical details are described in the following sections. Samples with <90% 5mC conversion calculated based on methylated lambda spike-in controls were excluded from downstream analyses.
[0162] Collection and preparation of cfDNA samples. Blood was collected in EDTA-coated Vacutainers. Plasma was separated from collected blood samples within 4 hours of collection. Plasma was collected by centrifugation at 1600xg for 10 min at 4°C and 16000xg for 10 min at 4°C and stored at -80°C for cfDNA purification. cfDNA from plasma was extracted using Qiamp Circulating Nuclei Acid Kit (Qiagen). cfDNA was quantified by Qubit Fluorometer (Life Technologies).
[0163] Preparation of carrier DNA and spike-in controls. Carrier DNA was prepared by PCR amplification of pNIC28-Bsa4 plasmid (Addgene, Cat. No. 26103) in a reaction containing 1 ng DNA template, 0.5 μM primers (forward: 5'-AGGCAACTTTATGCCCATGCAA-3' (SEQ ID NO: 2), reverse: 5'-CCAAGGGGTTATGCTAGTTATTGC-3' (SEQ ID NO: 3)), and 1x Phusion High-Fidelity PCR Master Mix (Thermo Scientific) with HF buffer. CpG-methylated lambda DNA and 2 kb unmodified spike-in control DNA were prepared as described above. CpG-methylated lambda DNA, carrier DNA and 2 kb of unmodified control were fragmented by Covaris M220 (peak incident power-50 W, duty factor-20%, cycles per burst (cpb)-200, time-150 s) and size-selected on 0.9–1.2× AMPure XP beads to select fragments of 150–250 bp.
[0164] Preparation of sequencing adapters. Adapter oligos (5'-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3' (SEQ ID NO: 4); 5'- / 5Phos / GATCGGAAGAGCACACGTCT-3' (SEQ ID NO: 5)) were obtained from IDT by HPLC purification. The adapter oligos were annealed together in a 50 μL reaction containing 15 μM of each oligo, 10 mM Tris-Cl (pH=8.0), 0.1 mM EDTA (pH=8.0) and 50 mM NaCl with a program of 95°C for 2 minutes, 140 cycles of 95°C for 20 seconds (with 0.5°C temperature reduction per cycle), and a hold at 4°C. 15 μM of the annealed Illumina multiplexed adapters were then aliquoted into small single-use vials and stored at -80°C.
[0165] mTet1CD oxidation. mTet1CD was prepared as previously described. DNA was incubated for 80 min at 37°C in a 50 μl reaction containing 50 mM HEPES buffer (pH 8.0), 100 μM ammonium iron(II) sulfate, 1 mM α-ketoglutarate, 2 mM ascorbic acid, 2 mM dithiothreitol, 100 mM NaCl, 1.2 mM ATP, and 4 μM mTet1CD. 0.8 U of proteinase K (New England Biolabs) was then added to the reaction mixture and incubated at 50°C for 1 h. The product was cleaned up with a Bio-Spin P-30 Gel Column (Bio-Rad) and 1.8× AMPure XP beads according to the manufacturer's instructions.
[0166] Pyridine borane reduction. The oxidized DNA in 35 μl of water was reduced in a 50 μl reaction containing 600 mM sodium acetate solution (pH 4.3) and 1 M pyridine borane (Alfa Aesar) in an Eppendorf ThermoMixer at 37° C., 850 rpm for 16 h. The product was purified using a Zymo-Spin column.
[0167] cfDNA TAPS. 10 ng of cfDNA was spiked in with 0.15% CpG methylated lambda DNA and 0.015% unmodified 2 kb control, used for end repair and A-tailing reactions, and ligated to Illumina Multiplexing adapters using the KAPA HyperPrep kit according to the manufacturer's protocol. 100 ng of carrier DNA was then added to the ligated library, and the samples were double oxidized with mTet1CD and reduced with pyridine borane as described above. The converted library was amplified for 7 cycles using KAPA Hifi Uracil Plus Polymerase and NEBNext® Multiplex Oligos for Illumina® (96 Unique Dual Index Primer Pairs), and cleaned up with 1× AMPure XP beads. cfDNA TAPS libraries were paired-end 150 bp sequenced on a NovaSeq 6000 sequencer (Illumina).
[0168] TAPS mapping and preprocessing. Raw sequencing reads were processed by trim_galore (version 0.6.2 www.bioinformatics.babraham.ac.uk / projects / trim_galore / ) to trim adapters and low quality bases with the following parameters --pairs--length 35--gzip--core 2. Clean reads were aligned to the human reference genome (GRCh38 ftp.ncbi.nlm.nih.gov / genomes / all / GCA / 000 / 001 / 405 / GCA_000001405.15_GRCh38 / seqs_for_alignment_pipelines.ucsc_ids / GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz.) combined with spike-in sequences using bwa mem (version 0.7.17-r1188) using the following parameters -I 500, 120, 1000, 20. Reads with MAPQ < 1 were excluded from further analysis. Duplicate reads were identified using Picard MarkDuplicates (version 2.18.29-SNAPSHOT). MethylDackel extracts (version 0.5.0 https: / / github.com / dpryan79 / MethylDackel) were used for methylation calling with the following parameters -q10-p13-t4--mergeContext--OT 10,140,75,75--OB 10,140,75,75. For further analysis, we excluded common SNPs (dbSNP153), blacklisted regions, centromeres, and CpG sites overlapping with sex chromosomes.
[0169] cfDNA WGBS analysis. cfDNA WGBS data were downloaded from EGAD00001004317. Raw sequencing reads were processed by trim_galore (version 0.6.2 www.bioinformatics.babraham.ac.uk / projects / trim_galore) and adapters and low quality bases were trimmed with the following parameters --pairs--length 35--gzip--core 2. Clean reads were aligned to the human reference genome (GRCh38) using bismark (Bismark Version: v0.22.0) with default parameters. deduplicate_bismark was used for duplicate removal. Samtools was used to filter fragments with -q 10 and only reads that mapped to the appropriate pairs were used for fragmentation analysis. Methylation was extracted from the duplicate-removed bam files with default parameters using bismark_methylation_extractor.
[0170] PCA for DNA methylation and feature over-representation analysis. The genome was binned into 1 kb windows. The number of methylated CpGs divided by the total number of sequenced CpGs was used to calculate the methylation level. Windows with a mean CpG coverage (total number of sequenced CpGs / total number of CpG positions) less than 2 were excluded for further analysis. Dimdesc with parameter proba=0.01 was used to determine the most contributing regions (maximum eigenvalue of each eigenvector) to each principal component obtained by the PCA function. Bedtools fisher was used to test the number of overlaps between the top 200 contributing regions (sorted by absolute correlation value) and selected genomic features. Selected genomic features included regulatory elements from Ensemble (ftp.ensembl.org / pub / release-97 / regulation / homo_sapiens / homo_sapiens.GRCh38.Regulatory_Build.regulatory_features.20190329.gff.gz) and CpG islands from UCSC (hgdownload.soe.ucsc.edu / goldenPath / hg38 / database / cpgIslandExt.txt.gz).
[0171] Two-class prediction using DNA methylation signatures. Two-class prediction models were trained and evaluated based on the LOO approach. Briefly, one sample was kept as the test set and the remaining samples were used for model training. DMRs (promoters for PDAC and enhancers for HCC) were identified in the training set by t-test (P-value < 0.002, methylation difference > 0.05). In each leave-one-out fold, 443-775 differentially methylated enhancers and 160-318 differentially methylated promoters were identified in the feature selection step for HCC versus non-cancer controls and PDAC versus non-cancer controls, respectively. In total, 1,521 enhancers and 531 promoters were selected during the cross-validation process. Prediction models were built for the selected DMRs and validated against the test samples using cv.Glmnet. This procedure was repeated N times, where N is the number of samples. ROC curves were prepared in R based on the prediction scores of held-out test samples from the cvglm model. Cirrhosis patients and cfDNA WGBS data were used as an independent validation set to evaluate the performance of the HCC model. Pancreatitis patients were used as an independent validation set to evaluate the performance of the PDAC model. Aligned BAM files were downsampled from 100M to 200M read pairs using samtools. For each downsampled set, DMRs were detected using the method described above. ref DMRs were defined as the sum of unique DMRs in the LOO cross-validation. The proportion of ref DMRs was calculated by dividing the overlapping DMRs between the downsampled set and the ref DMRs and total ref DMRs.
[0172] GO analysis of DMRs. Genes regulated by differentially methylated enhancers in HCC cfDNA were identified using the GeneHancer database. Genes closest to differentially methylated promoters in PDAC were identified as related using the following R packages: AnnotationHub (version 2.18.0), TxDb.Hsapiens.UCSC.hg38.knownGene (version 3.10.0) and org.Hs.eg.db (version 3.10.0). GO analysis was performed on these identified genes using the Enrichr tool against the NCI-Nature Pathway Interaction database.
[0173] Tissue Reference Map. CpG-level tissue methylation data was collated from six public sources (public methylation WGBS sources for the generation of tissue maps are not included in this disclosure but can be made available upon request). After filtering disease gender-specific and low coverage samples, 144 healthy adult tissue samples were retained and grouped into 32 physiologically distinct tissue groups (raw data related to cfDNA tissue contribution for each patient in the cfTAPS cohort are not included in this disclosure but can be made available upon request). 133 of the 144 samples were already aligned to hg38, and the remaining 11 samples were converted from hg19 to hg38 using the UCSC hgLiftOver tool.
[0174] A tissue-specific DMR discovery algorithm similar to that of Moss et al. was used to filter approximately 79,000 enhancers from the Ensembl Regulatory Build. Specifically, the algorithm performs pairwise one-vs-all comparisons for each tissue group in the reference atlas and selects regions that show the greatest median methylation difference and concordant methylation across the tissue group in question. As in Moss et al., pairwise tissue group correlations were also calculated and the DMRs that best separate each tissue group from the first and second most correlated tissues were included.
[0175] Tissue deconvolution by non-negative linear least squares regression. Tissue deconvolution was performed using non-negative linear least squares regression, implemented using the optimize function in Scipy in Python 3.8. Given a tissue reference matrix A and a vector y of methylation ratios observed in sample s, s Given, the tissue contribution x was estimated by solving the following minimization problem:
[0176] min||Ax-y s || 2 Let x≧0.
[0177] Fragmentation analysis. DNA fragment lengths were obtained from alignment files using Samtools. Fragmentation profiles were calculated as the percentage of cfDNA fragments in 10 bp length range bins. PCA analyses and plots were generated in R.
[0178] For fragmentation-based prediction, the proportion of cfDNA fragments (300-500 bp) in 10 bp length range bins was calculated. Models were built and trained by a leave-one-out approach using the cv.glmnet method. ROC curves were prepared in R based on the prediction scores from validation.
[0179] CNV analysis. The alignment files for each sample were downsampled to 225M read pairs by samtools view. The QDNAseq package was used for copy number variation analysis. Bin annotation was downloaded from QDNAseq.hg38 (github.com / asntech / QDNAseq.hg38) and a bin size of 100kb was used. Blacklisted regions or regions with less than 80 mappability were excluded for further analysis. Cutoffs of 0.8 and 1.2 were used to define copy number loss and gain, respectively, in the callBins function. Patients with copy number abnormalities with a length range of >500kb were classified as patients with CNV.
[0180] Three-class predictive models. Three-class predictive models were trained and evaluated based on the LOO approach. For DNA methylation, candidate features were first narrowed down to 824,320 1 kb windows encompassing mapping to regulatory regions as previously described. The methylation models aim to capture cancer type-specific methylation changes by selecting DMRs based on pairwise comparisons using t-tests. DMRs were then ranked by P-value and the top 5 DMRs in each pairwise comparison were selected for model training. Predictive models were built for selected DMRs among the training set using an SVM model implemented in the caret package (training method = "svmLinear2") and validated against test samples. This procedure was repeated N times, where N is the number of samples. For tissue contribution and fraction of fragmentation, models were built following the same method as for DMRs, using the raw matrix. These three models were integrated by taking the averaged (mean) predictions across the three modalities and the prediction selected for each case was the one with the largest averaged prediction score.
[0181] It should be understood that the foregoing detailed description and the accompanying examples are merely illustrative and should not be construed as limiting the scope of the present disclosure, which is defined solely by the appended claims and their equivalents.
[0182] Various changes and modifications to the disclosed embodiments will be apparent to those skilled in the art. Such changes and modifications, including but not limited to those related to the chemical structures, substituents, derivatives, intermediates, compounds, compositions, formulations, or methods of use of the present disclosure, can be made without departing from the spirit and scope thereof. EXAMPLES
[0183] 5. Working Example It will be readily apparent to those skilled in the art that other suitable modifications and adaptations of the disclosed methods described herein are readily applicable and recognizable, and may be made using suitable equivalents without departing from the scope of the disclosure or the aspects and embodiments disclosed herein. Having described the disclosure in detail, the disclosure will be more clearly understood by reference to the following examples, which are intended merely to illustrate certain aspects and embodiments of the disclosure, and should not be considered as limiting the scope of the disclosure. The disclosures of all journal references, U.S. patents, and publications referenced herein are incorporated herein by reference in their entirety.
[0184] The present disclosure has multiple aspects, which are illustrated by the following non-limiting examples. Example 1 Adaptation of TAPS for cfDNA sequencing. Experiments were performed to optimize the TAPS protocol to work with low-input cfDNA (10 ng, purified from 1–3 mL of plasma). Briefly, 10 ng of cfDNA was first ligated to Illumina adapters, and then 100 ng of carrier DNA was added to the sample prior to the TET oxidation and pyridine borane (PyBr) reduction steps (Figure 1A). We found that the addition of carrier DNA improved the recovery of cfDNA during the workflow, resulting in higher library yields compared to the standard TAPS protocol (Figure 5A). 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) in cfDNA were then oxidized by the mTet1CD enzyme to 5-carboxylcytosine (5caC) and reduced to dihydrouracil (DHU), which was amplified as T in the final PCR step (Figure 1A).
[0185] cfTAPS was applied to 87 cfDNA samples. Libraries were sequenced to an average of 360M read pairs (average depth of 11.6×, range 8.2–22×), yielding high unique mapping rates and unique de-duplicated mapping rates of 94.8% and 77.1%, respectively (Figure 1B, raw data related to sequencing statistics are not included in this disclosure but can be made available upon request). Of the mapped reads, 99.95% mapped to the human genome (Figure 5B). In comparison, recent cfDNA whole-genome bisulfite sequencing (WGBS) sequenced to a similar depth (average 371M read pairs) and obtained significantly lower unique mapping rates (63.6%) and unique de-duplicated mapping rates (53.9%), despite using a higher cfDNA input (from 5 mL of plasma) (Figure 5C). This highlights the advantage of cfTAPS in generating higher quality, more complex data than cfDNA WGBS while requiring less cfDNA input.
[0186] The accuracy of cfTAPS to detect 5mC was then evaluated based on spike-in controls with modified and unmodified cytosines at known positions. CpG methylated lambda DNA was used to estimate 5mC conversion. Two samples had low conversion rates below 85% and were excluded from downstream analysis (raw data related to sequencing statistics are not included in this disclosure but can be made available upon request). The remaining 85 samples had an average 5mC conversion rate of 97.0%, or a false negative rate (unconverted rate of 5mC) of 3.0% (Figure 1C). The false positive rate (unconverted rate of unmodified C) estimated based on unmodified amplicon spike-in was 0.28%, confirming that cfTAPS enables sensitive and specific detection of 5mC in cfDNA (Figure 1C). The high reproducibility of cfTAPS between technical replicates was further confirmed (Figure 5D). Example 2 Genome-wide DNA methylation from cfTAPS. We next performed experiments to characterize the cfDNA methylome in 85 cfDNA samples that passed initial quality control. The cohort included samples from 21 patients with HCC, 23 patients with PDAC, 30 non-cancer controls, 4 patients with cirrhosis, and 7 patients with pancreatitis (Figure 6A). Cirrhosis and pancreatitis are precancerous conditions that affect the liver and pancreas, respectively. Most PDAC and HCC patients in the cohort were in the non-metastatic stage, with 52% of PDAC patients and 67% of HCC patients in stages I and II (Figure 2A, clinical data related to the cfTAPS study cohort are not included in this disclosure but can be made available upon request). Of the 21 HCC patients, only 4 (19%) had elevated APF levels (>20 ng / mL). Of the 18 PDAC patients who had CA19-9 measurements, 16 (89%) had elevated levels of CA19-9 (greater than 37 U / mL). However, it is important to note that CA19-9 levels are often elevated in non-malignant conditions, including inflammatory diseases. Of note, the non-cancer controls were collected from endoscopy clinics and were overrepresented in gastrointestinal inflammatory conditions, such as Crohn's disease and colitis (clinical data related to the cfTAPS study cohort are not included in this disclosure, but can be made available upon request). These non-cancer controls are typically more difficult to distinguish from cancer patients than healthy control groups, but this may provide a more realistic comparison of diagnostic tests in the elderly population.
[0187] We analyzed the global methylation levels of cfDNA in cancer and control samples. cfDNA methylation showed a typical bimodal distribution in all groups with most CpG sites either fully methylated or unmethylated (Figure 5B). The average CpG methylation level in control samples was 75.5%, similar in cancer cfDNA (HCC: 74.9%, PDAC: 75.1%). The previously reported global cfDNA hypomethylation in HCC was only observed in a few samples with late stage or large tumor size (Figure 2B and Figure 6C-6F). In contrast, a higher variance of methylation in 1 Mb genomic windows was observed among cancer patients compared to controls (Figure 6G-6H).
[0188] Experiments were then performed to examine whether the genome-wide cfDNA methylation signature has the potential to discriminate between cancer patients and non-cancer controls. Principal component analysis (PCA) of cfDNA methylation in 1 kb genomic windows was first performed. Both HCC (Figure 2C) and PDAC samples (Figure 2D) showed partial separation from controls in principal component 2 (PC2) and PC1, respectively. Note that inflammatory patients (Crohn's disease and colitis) do not separate from other non-cancer controls (Figure 6I). Experiments were then performed to examine where the windows that contributed most to cancer / control separation were enriched in the genome. Results showed that the top 200 windows most highly correlated with PC2 in HCC cases were enriched in enhancers (Figure 2E). Conversely, the 200 windows most highly correlated with PC1 in PDAC cases were highly enriched in promoters (Figure 2E), suggesting that different cancer types have different cfDNA methylation signals. Example 3 Differential DNA methylation from cfTAPS. Because methylation patterns in regulatory regions contributed significantly to the distinction between cancer and controls in unsupervised analysis, experiments were performed to investigate the predictability of cfDNA methylation in enhancer and promoter regions for HCC and PDAC prediction, respectively, using a supervised machine learning approach with leave-one-out (LOO) cross-validation. Briefly, in each round of LOO cross-validation, one sample was used as a validation set, and the remaining samples were used for model training. Within each fold, differentially methylated enhancers and promoters for HCC and PDAC, respectively, were identified and used to train a regularized generalized linear model classifier (glmnet) to distinguish each cancer type from control samples. The model was then evaluated on hold-out test samples in each fold (Figure 7A). Cirrhosis and pancreatitis samples were not included in the model building, but were used as independent validation sets to evaluate the performance of the classifier for discriminating between cancer and pre-malignant conditions.
[0189] Significant prediction of HCC (AUC=0.99) was achieved based on the differentially methylated enhancers (Figures 2F-2G, raw data for the differentially methylated enhancers used for HCC prediction versus controls are not included in this disclosure but can be made available upon request). Furthermore, based on the prediction score, three out of four cirrhosis samples could be distinguished from HCC, suggesting that the model is capable of detecting cancer-specific features (Figure 7B). Gene ontology analysis was then performed on the differentially methylated enhancers and found significant enrichment in signaling pathways commonly affected in liver cancer, including regulation of RAC1 activity and IL8- and CXCR1-mediated signaling (Figure 7C). For example, in cfDNA of HCC patients, significant hypermethylation of enhancers regulating the expression of DLC1 gene, a tumor suppressor for human liver cancer involved in RAC1 and Rho signaling pathways, was observed (Figure 7D).
[0190] Accurate prediction of PDAC (AUC=0.98) was achieved based on differentially methylated promoters (Figures 2H-2I, raw data for differentially methylated promoters used for PDAC prediction versus controls are not included in this disclosure but can be made available upon request). Similarly, the classifier was able to predict 6 out of 7 pancreatitis samples as non-cancer despite not being trained on any pancreatitis samples (Figure 7E). Differentially methylated promoters in PDAC cfDNA were enriched in signaling pathways affected by PDAC, including RB1 regulation and p38 signaling pathways (Figure 7F). For example, results showed significant hypermethylation in the RB1 gene promoter, a well-studied tumor suppressor gene (Figure 7G). Hypermethylation of the RB1 promoter has been previously found in human cancers and downregulation of RB1 has been reported in pancreatic cancer.
[0191] Finally, the HCC model was validated on an independent dataset from a recent cfDNA WGBS study including four HCC patients and four non-cancer controls. The results showed that the model built on differentially methylated enhancers identified from cfTAPS data was able to correctly classify all HCC and non-cancer controls from this external dataset (Figure 7H). The high sequencing depth of cfTAPS is essential for de novo differential methylation analysis from cfDNA, and it is important to note that the identified differentially methylated regions (DMRs) were significantly reduced when the data were downsampled to 100-200M read pairs (Figure 7I). Taken together, cfTAPS enables genome-wide discovery of DMRs in cfDNA, and differential methylation patterns in regulatory regions allow accurate prediction of HCC and PDAC. Example 4 cfTAPS informs tissue of origin. cfDNA methylation has been shown to provide information of tissue of origin. Most approaches infer tissue contribution from cfDNA methylation using 450K methylation array tissue data, which covers less than 1% of CpGs in the human genome. To further utilize genome-wide information from cfTAPS for cfDNA deconvolution, CpG-level methylation data was collated from 144 publicly available tissue and blood cell WGBS and stratified into 32 physiologically distinct tissues and blood cell types, including liver tumor tissue (sources of public methylation WGBS data for the generation of tissue maps are not included in this disclosure, but can be made available upon request). Considering tissue-specific DNA methylation abundance in enhancer regions, an enhancer-aggregated reference map of tissue methylation was constructed. The resulting methylation reference map shows good clustering of blood and immune cell types, as well as physiologically relevant solid tissues (Figure 8A).
[0192] Tissue contributions in cfTAPS samples were calculated by performing non-negative linear least squares regression (NNLS). cfDNA tissue contributions were broadly similar between cancer and control groups, with a predominance of blood and immune cells and a low proportion of solid tissues, consistent with previous reports (Figure 3A, Figure 8B, raw data related to cfDNA tissue contributions for each patient in the cfTAPS cohort are not included in this disclosure but can be made available upon request). Importantly, significantly increased liver tumor contributions were observed in HCC alone (Figure 3B, paired t-test, P value 0.0016), and significantly increased memory T cell contributions were observed in PDAC samples (paired t-test, P value 0.028) (Figure 8C). A regularized generalized linear model was trained based on tissue contributions and evaluated across all samples using LOO cross-validation and shown to correctly separate the majority of samples of both cancer types (HCC vs. non-cancer controls: AUC = 0.77, PDAC vs. non-cancer controls: AUC = 0.81). However, these models perform poorly in distinguishing between pancreatitis and cirrhosis compared to methylation-based models (Figures 9D-8I). Tissue deconvolution is currently limited by the availability of public WGBS data. Nevertheless, these results indicate that cfTAPS provides valuable tissue-of-origin information for early cancer detection. Example 5 Fragmentation patterns from cfTAPS. Although the main purpose of cfTAPS is DNA methylation sequencing, it only induces base changes at modified cytosines, thus keeping the majority of DNA intact. Therefore, additional genetic information can be extracted from cfTAPS data to further improve the sensitivity of early cancer detection. First, we performed experiments to investigate CNVs from cfTAPS data. As expected in a non-progressive cancer cohort, CNVs were predicted only in four HCC patients and three PDAC patients (Figures 9A-9B). Next, we performed experiments to investigate whether cfTAPS can retain reliable cfDNA fragmentation information, which has been shown in recent years to change significantly during cancer development and has therefore been adopted for cancer detection assays.
[0193] It was first confirmed that the cfDNA fragmentation pattern detected by cfTAPS was consistent with that generated by whole genome sequencing (WGS), with a predominant peak at 167 bp, a second peak at approximately 320 bp, and smaller peaks at 10 bp intervals below 167 bp, reflecting the nucleosome fragmentation pattern (Figure 3C, raw data related to fragment length distribution in each individual are not included in this disclosure but can be made available upon request). In contrast, the fragmentation pattern was clearly different in the previously published cfDNA WGBS, likely due to the loss of the 10 bp oscillation in the cfDNA fragmentation profile due to DNA damage (Figure 10A). Consistent with previous cfDNA WGS, results showed that cancer patients had a higher frequency of cfDNA fragments <150 bp compared to non-cancer controls (Kruskal-Wallis test, HCC: P-value 6.871e-06, PDAC: P-value 0.006731) and a lower percentage of longer fragments between 310 and 500 bp (Kruskal-Wallis test, HCC: P-value 2.627e-07, PDAC: P-value 1.263e-06) (Figure 3D), further confirming the faithful preservation of cfDNA fragmentation information in cfTAPS.
[0194] We then developed a new approach for characterization of cfDNA fragmentation profiles using cfTAPS. Briefly, we divided the cfDNA fragmentation distribution into 10 bp bins and calculated the fraction of fragments in each 10 bp bin (Figure 3C). The fraction of cfDNA long fragments (300-500 bp) in the 10 bp bins separated PDAC and HCC from controls in an unsupervised analysis by PCA (Figure 3E). The results further showed that this cfDNA fragmentation signature could be used to distinguish HCC and PDAC from non-cancer controls with high accuracy (HCC AUC=0.92, PDAC AUC=0.84) (Figures 10B, 10C, 10E, and 10F). However, this approach was less accurate in distinguishing cancer from cirrhosis and pancreatitis compared to methylation-based classifiers (Figures 10D and 10G), suggesting that the fragmentation information is not cancer-specific. Example 6 Multi-cancer detection by cfTAPS. Experiments were then conducted to investigate the utility of cfTAPS for multi-cancer detection. The top 5 DMRs of each pairwise comparison (non-cancer controls to HCC, non-cancer controls to PDAC, HCC to PDAC) were selected as features for the multi-cancer differential methylation model. A support vector machine (SVM) model was trained to estimate the respective probability that a blood sample came from each group. Similar models were constructed using tissue contribution and fragmentation profile. Using LOO cross-validation, the results showed that the methylation model could achieve an overall accuracy of 0.77, which outperformed the tissue contribution and fragmentation profile models (respective accuracy of 0.62 and 0.46, Fig. 4A, Fig. 11A).
[0195] To further strengthen the multi-cancer prediction model, we constructed a multi-modal classifier that combined differential methylation, tissue contribution, and fragmentation profiles (Figure 4B). This integrated model took the averaged scores across the three modalities and used the most reliable prediction for each sample. The overall accuracy of the combined model was 0.86 (64 out of 74 correctly classified), and the accuracy for distinguishing controls from any cancer type was 0.92 (Figure 4C), highlighting the advantage of incorporating multi-modal information for cancer type prediction. Finally, we explored the DMRs used for multi-cancer prediction (Figure 11B, data related to the methylation features used for HCC, PDAC, and control prediction are not included in this disclosure but can be made available upon request). Interestingly, the results showed that genes near these regions were enriched in Notch and Wnt signaling, as well as EGFR (ErbB) signaling, providing biological support for these potential multi-cancer biomarkers (Figure 11C). [Brief description of the drawings]
[0196] [Figure 1]cfDNA analysis by TAPS. (A) Schematic of the TAPS approach for cfDNA analysis. cfDNA is isolated from 1–3 mL of plasma. 10 ng of cfDNA is ligated to Illumina sequencing adaptors and topped up with 100 ng of carrier DNA. 5mC and 5hmC in DNA are then oxidized to 5caC by mTet1CD enzyme, reduced to DHU by PyBr, amplified and detected as T in the final sequencing. Computational analysis of TAPS data allows for the simultaneous characterization of multiple cfDNA features including DNA methylation, tissue of origin, fragmentation patterns and CNVs. (B) Number of total reads, uniquely mapped reads and uniquely mapped PCR de-duplicated reads in 87 cfDNA TAPS libraries. The total number of reads and the average percentage of uniquely mapped reads and de-duplicated reads compared to total reads are shown above the bar graphs. Error bars represent standard error. (C) Conversion and false positive rates of 5mC in 85 cfDNA TAPS libraries based on spike-in controls with modified or unmodified cytosines at known positions. Each dot represents an individual sample. [Figure 2A] cfDNA methylation in clinical samples. Cancer stage distribution of 21 HCC and 23 PDAC patients included in this study. [Figure 2B] cfDNA methylation in clinical samples. Mean per CpG genomic modification levels in non-cancer controls, HCC and PDAC cfDNA. Each dot represents an individual sample. [Figure 2C] cfDNA methylation in clinical samples. PCA plot of cfDNA methylation in 1 kb genomic windows in non-cancer controls and HCC. [Figure 2D] cfDNA methylation in clinical samples. PCA plot of cfDNA methylation in 1 kb genomic windows in non-cancer controls and PDAC. [Figure 2E] cfDNA methylation in clinical samples.Regional overexpression analysis showed that regulatory regions were most correlated with PC2 for HCC and PC1 for PDAC. [Figure 2F] cfDNA methylation in clinical samples. Receiver operating characteristic (ROC) curves of model classification performance based on differentially methylated enhancers in HCC and non-cancer controls (n=51, HCC=21, non-cancer controls=30). [Figure 2G] cfDNA methylation in clinical samples. LOO cancer prediction scores for HCC and non-cancer controls. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC. [Figure 2H] cfDNA methylation in clinical samples. ROC curves of model classification performance based on differentially methylated enhancers between PDAC and non-cancer controls (n=53, PDAC=23, non-cancer controls=30). [Figure 2I] cfDNA methylation in clinical samples. LOO cancer prediction scores for PDAC and non-cancer controls. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as PDAC. [Diagram 3]cfTAPS allows the analysis of tissue of origin and fragmentation patterns in cfDNA. (A) Mean tissue contribution in non-cancer individuals estimated by NNLS. Tissue contributions less than 1.5% are summarized as "other". (B) Box plots showing estimated liver cancer contributions within non-cancer, HCC, and PDAC groups. Statistical significance was assessed by paired t-test. ns-not significant. (C) Length distribution of cfDNA fragments in the three groups. For each sample, the proportion (P) of long cfDNA fragments (300-500bp) in 10 base pair intervals was used as a fragmentation feature for PCA analysis and machine learning. (D) Box plots showing the proportion of short (70-150bp) and long (300-500bp) fragments in non-cancer controls, PDAC, and HCC. Kruskal-Wallis test was performed to test for differences in fragment size distribution between groups. Statistically significant differences are marked with asterisks (*P value < 0.05, **P value < 0.01, ***P value < 0.001, ****P value < 0.0001). (E) PCA plot of the proportion of cfDNA 10 bp fragments in non-cancer controls and HCC (left panel), and non-cancer controls and PDAC (right panel). [Figure 4] Integrating multimodal features from cfTAPS enhances multi-cancer detection. (A) Heatmap showing the performance of individual models for multi-cancer prediction and the predicted probability for each patient. Each vertical column is a patient. Detection Yes / No means the patient is correctly or incorrectly classified based on a particular feature. Prediction score means the probability of classifying a patient into a particular group based on a particular feature. (B) Schematic detailing how we integrate multiple features (DNA methylation, tissue contribution, and fragmentation percentage) extracted from cfTAPS data for multi-cancer prediction. (C) Actual and predicted patient status calculated with LOO cross-validation. [Diagram 5]cfDNA TAPS. (A) Agarose gel of 10 representative cfDNA TAPS libraries after post-amplification cleanup. All cfDNA TAPS libraries were prepared from 10 ng of cfDNA and amplified with 7 PCR cycles. (B) Number of mapped read pairs for hg38, spike-in, and carrier DNA in 87 cfDNA TAPS libraries. The average percentage of mapped read pairs compared to total read pairs is shown above the bar graph. Error bars represent standard error. (C) Number of total reads, uniquely mapped reads, and uniquely mapped PCR deduplication reads in cfDNA WGBS (EGAD00001004317) (24). The total number of reads and the average percentage of uniquely mapped reads and deduplication reads compared to total reads are shown above the bar graph. Error bars represent standard error. (D) Correlation between technical replicates of cfDNA TAPS libraries prepared from the same cfDNA samples sequenced to low depth 2.6×. Methylation was calculated in 100 kb windows. [Figure 6A] Global cfDNA methylation patterns in cancer and controls. Age and sex distribution of pancreatitis, cirrhosis, PDAC, HCC, and non-cancer control patients included in the cfTAPS cohort. [Figure 6B] Global cfDNA methylation patterns in cancer and controls. Genome-wide distribution of CpG modifications in cfDNA in non-cancer controls, HCC and PDAC. Bar plots show the distribution of average CpG modifications for each group. Overlay line plots show the CpG methylation distribution in each patient. [Figure 6C] Global cfDNA methylation patterns in cancer and controls. Correlation plot of mean cfDNA CpG modification levels and tumor size (mm) in HCC patients. [Figure 6D] Global cfDNA methylation patterns in cancer and controls. Correlation plot of mean cfDNA CpG modification levels and tumor stage in HCC patients. [Figure 6E]Global cfDNA methylation patterns in cancer and controls. Correlation plot of tumor size (mm) for PDAC patients. Each point represents an individual patient. The dashed line represents the linear trend fitted with linear regression. The shaded areas represent the 95% confidence intervals of the fitted model. Pearson correlation coefficients (cor) and P values are indicated on the plot. [Figure 6F] Global cfDNA methylation patterns in cancer and controls. Correlation plot of tumor stage for PDAC patients. Each point represents an individual patient. The dashed line represents the linear trend fitted with linear regression. The shaded regions represent the 95% confidence intervals of the fitted model. Pearson correlation coefficients (cor) and P values are indicated on the plot. [Figure 6G] Global cfDNA methylation patterns in cancer and controls. Distribution of CpG modification levels on chromosome 4 in cfDNA of non-cancer controls, HCC and PDAC. Each line represents an individual patient. Mean CpG modification values were calculated for each 1 Mb window along chromosome 4 and Gaussian smoothed (smoothing window size 10). [Figure 6H] Global cfDNA methylation patterns in cancer and controls. Methylation variance in 1 Mb genomic windows in non-cancer controls, HCC and PDAC. [Figure 6I] Global cfDNA methylation patterns in cancer and controls. PCA plots of cfDNA methylation in 1 kb genomic windows in non-cancer controls and HCC, non-cancer subjects and PDAC (Crohn's disease and colitis are colored green and yellow, respectively). [Figure 7A] HCC and PDAC prediction based on cfDNA DMRs. Overview of the training and validation approach of the LOO model. The total number of samples is labeled as n. In each iteration, the model training set consisted of n-1 samples. Differentially methylated enhancers (for HCC) or promoters (for PDAC) were selected for model building. Prediction models were evaluated on hold-out test samples in each fold. Cirrhosis and pancreatitis samples were not included in DMR identification and model building. [Figure 7B] HCC and PDAC prediction based on cfDNA DMR. HCC cancer prediction scores for cirrhosis samples. Each blue dot represents the prediction score of an individual LOO model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC. [Figure 7C] HCC and PDAC prediction based on cfDNA DMRs. Gene Ontology analysis of genes associated with differentially methylated enhancers based on HCC cfDNA using Enrichr for NCI-Nature Pathway Interaction (P-value < 0.002). The top 10 categories selected based on P-value are shown in the graph. Gene-enhancer interactions were assigned using the GeneHancer reference database. [Figure 7D] HCC and PDAC prediction based on cfDNA DMRs. Methylation of representative differentially methylated enhancers in HCC cfDNA for DLC1 gene (two-tailed t-test P value=8.765e-06). [Figure 7E] HCC and PDAC prediction based on cfDNA DMR. PDAC cancer prediction score for pancreatitis samples. Each yellow dot represents the prediction score of an individual LOO model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as PDAC. [Figure 7F] HCC and PDAC prediction based on cfDNA DMRs. Gene Ontology analysis of closest genes to differentially methylated promoters based on PDAC cfDNA using Enrichr for NCI-Nature Pathway Interaction (P-value < 0.002). The top 10 categories selected based on P-value are shown on the graph. [Figure 7G] HCC and PDAC prediction based on cfDNA DMR. Representative differentially methylated promoter methylation in PDAC cfDNA for the RB1 gene (two-tailed t-test P value=0.0017). [Figure 7H] HCC and PDAC prediction based on cfDNA DMR. HCC cancer prediction scores for an independent cfDNA WGBS dataset (EGAD00001004317). Each point represents the prediction score of an individual LOO model. Grey points belong to non-cancer controls and red points belong to HCC. Black points indicate the mean probability score of a particular sample. The dashed line represents the probability score threshold. Samples with a mean probability score above this threshold were predicted as HCC. [Figure 7I] HCC and PDAC prediction based on cfDNA DMRs. Proportion of ref DMRs that could be detected in downsampled reads. DMRs identified in the original LOO model training were treated as ref DMRs. [Figure 8A] Tissue of origin of cfDNA. t-SNE plot of the reference tissue methylation atlas. [Figure 8B] Tissue of origin of cfDNA. Average tissue contribution in HCC and PDAC individuals. [Figure 8C] Tissue of origin of cfDNA. Boxplots showing estimated T cell contribution in non-cancer, HCC and PDAC cfDNA samples. [Figure 8D] Tissue of origin of cfDNA. ROC curve of model performance using tissue contribution to classify HCC versus non-cancer. [Figure 8E] Tissue of origin of cfDNA. LOO cancer prediction scores for HCC and non-cancer controls using classifiers trained on tissue contribution. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC. [Figure 8F] Tissue of origin of cfDNA. Cancer scores of cirrhosis samples using HCC vs. non-cancer classifiers. Each blue dot represents the prediction score of an individual model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC. [Figure 8G] Tissue of origin of cfDNA. ROC curve of model performance using tissue contribution to classify PDAC versus controls. [Figure 8H] Tissue of origin of cfDNA. LOO cancer prediction scores for PDAC and non-cancer controls using classifiers built on tissue contribution. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as PDAC. [Figure 8I] Tissue of origin of cfDNA. PDAC cancer scores for pancreatitis samples using PDAC vs. non-cancer classifiers. Each yellow dot represents the prediction score of an individual model. Black dots indicate the average probability score for a particular sample. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as PDAC. [Figure 9A] CNV analysis in cfDNA. Heatmap of CNV estimates from cfDNA in 100kb bins. [Figure 9B] CNV analysis in cfDNA. cfDNA samples with >500k CNVs. [Figure 10]cfDNA fragmentation patterns for cancer prediction. (A) Fragment size distribution of cfDNA in public whole genome bisulfite sequencing data. Frequencies were calculated by dividing the number of fragments of a particular length by the total number of fragments. (B) ROC curves of HCC and non-cancer control prediction scores from a generalized linear model using the percentage of long cfDNA fragments (300-500 bp) in 10 bp bins as features. (C) Cancer prediction scores of HCC and non-cancer controls in classifiers trained using LOO cross-validation. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as HCC. (D) HCC cancer prediction scores of cirrhosis samples in these classifiers. Each blue dot represents the prediction score of an individual model. The black dot indicates the average prediction score. The dashed line represents the probability score threshold. Samples with an average probability score above this threshold were predicted as HCC. (E) ROC curves of PDAC and non-cancer control prediction scores from a generalized linear model using the proportion of long cfDNA fragments (300-500 bp) in 10 bp bins as features. (F) LOO cancer prediction scores for PDAC and non-cancer controls for a classifier built on cfDNA fragment frequencies in the 10 bp length range. The dashed line represents the probability score threshold. Samples with a probability score above this threshold were predicted as PDAC. (G) PDAC cancer prediction scores for pancreatitis samples for a classifier built on cfDNA fragment frequencies in the 10 bp length range. Each yellow dot represents the prediction score of an individual model. The black dot indicates the average prediction score. The dashed line represents the probability score threshold. Samples with a mean probability score above this threshold were predicted as PDAC. [Figure 11A] Multi-cancer detection with cfTAPS. Performance of methylation, tissue contribution and fragmentation fraction models in 3-class classification. The top panel shows the accuracy of each classifier, and the bottom panel shows actual and predicted patient status in a LOO cross-validation analysis. [Figure 11B] Multi-cancer detection by cfTAPS. Heatmap showing the methylation status of selected genomic regions used for cancer type prediction. [Figure 11C]Multiple cancer detection by cfTAPS. Gene ontology analysis using Enrichr for NCI-Nature Pathway Interaction for the closest genes of selected DMRs for 3-class classification. [Figure 12] Schematic diagram of the different patterns resulting from C to T SNPs and methylated cytosines in the target sequence before and after TAPS. In the diagram, OT means original top, OB means original bottom, CTOT means complementary to original top, and CTOB means complementary to original bottom.
Claims
1. A method for obtaining a methylation signature, comprising: isolating cell-free DNA (cfDNA) from a sample; preparing a sequencing library containing the cfDNA; performing Tet-assisted pyridine borane sequencing (TAPS) on the sequencing library to obtain a genome-wide methylation signature of the cfDNA.
2. The method according to claim 1, wherein the unique mapping rate obtained from the TAPS is at least 80%, and / or the unique deduplicated mapping rate is at least 70%.
3. The method according to claim 1, wherein preparing the sequencing library comprises ligating a sequencing adapter to the isolated cfDNA.
4. The method according to claim 1, wherein carrier DNA is added to the sequencing library before performing the TAPS.
5. The method according to claim 1, further comprising identifying at least one methylation biomarker from the genome-wide methylation signature of the cfDNA, and determining whether the methylation biomarker is an indicator of cancer.
6. The method according to claim 5, wherein the methylation biomarker comprises a differential methylation region (DMR).
7. The method according to claim 6, further comprising classifying the sample based on the DMR as compared to a reference DMR.
8. The method according to claim 7, wherein the reference DMR corresponds to a non-cancerous control or a cancerous control.
9. The method according to claim 1, further comprising identifying at least one methylation biomarker from the genome-wide methylation signature of the cfDNA, and determining the tissue of origin corresponding to the methylation biomarker.
10. The method according to claim 9, further comprising classifying the sample based on the biomarker of the tissue of origin.
11. The method according to claim 1, further comprising identifying a DNA fragmentation profile, and determining whether the fragmentation profile is an indicator of cancer.
12. The method according to claim 1, further comprising identifying at least one sequence variant from the cfDNA, and determining whether the sequence variant is an indicator of cancer.
13. Performing the TAPS on the array determination library to obtain the whole-genome methylation signature includes identifying one or more of 5mC modification, 5hmC modification, 5caC modification, or 5fC modification in the cfDNA, and providing a quantitative measure of the frequency of one or more of the 5mC modification, 5hmC modification, 5caC modification, or 5fC modification. The method according to claim 1.
14. A method for determining whether a subject has cancer using the method according to any one of claims 1 to 13.
15. The method according to claim 14, wherein the cancer includes hepatocellular carcinoma (HCC) or pancreatic ductal adenocarcinoma (PDAC).
16. A method for determining whether a subject has early-stage cancer using the method according to any one of claims 1 to 13.
17. The method according to claim 16, wherein the early-stage cancer includes early-stage hepatocellular carcinoma (HCC) or early-stage pancreatic ductal adenocarcinoma (PDAC).
18. A method for obtaining a methylation signature, comprising: Isolating cell-free DNA (cfDNA) from a sample; Preparing an array determination library containing the cfDNA; Adding carrier DNA to the array determination library; Performing TET-assisted pyridine borane sequencing (TAPS) on the array determination library to obtain the whole-genome methylation signature of the cfDNA; Identifying at least one methylation biomarker from the whole-genome methylation signature of the cfDNA and determining whether the methylation biomarker is an indicator of cancer; Identifying at least one methylation biomarker from the whole-genome methylation signature of the cfDNA and determining the tissue of origin corresponding to the methylation biomarker; Identifying a DNA fragmentation profile and determining whether the fragmentation profile is an indicator of cancer; Identifying at least one sequence variant from the cfDNA and determining whether the sequence variant is an indicator of cancer. The method.
19. The method according to claim 18, wherein the unique mapping rate obtained from the TAPS is at least 80%, and / or the unique deduplicated mapping rate is at least 70%.
20. The method according to claim 18, wherein preparing the sequencing library comprises ligating a sequencing adapter to the isolated cfDNA.
21. The method according to claim 18, wherein the methylation biomarker comprises a differentially methylated region (DMR).
22. The method according to claim 21, further comprising classifying the sample based on the DMR as compared to a reference DMR.
23. The method according to claim 22, wherein the reference DMR corresponds to a non-cancerous control or a cancerous control.
24. The method according to claim 18, further comprising classifying the sample based on a biomarker of the tissue of origin.
25. Performing the TAPS on the sequencing library to obtain the genome-wide methylation signature comprises identifying one or more of 5mC modification, 5hmC modification, 5caC modification, or 5fC modification in the cfDNA and providing a quantitative measure of the frequency of one or more of the 5mC modification, 5hmC modification, 5caC modification, or 5fC modification. The method according to claim 18.
26. A method of determining whether a subject has cancer using the method according to any one of claims 18 to 25.
27. The method according to claim 26, wherein the cancer comprises hepatocellular carcinoma (HCC) or pancreatic ductal adenocarcinoma (PDAC).
28. A method of determining whether a subject has early-stage cancer using the method according to any one of claims 18 to 25.
29. The method according to claim 28, wherein the early-stage cancer comprises early-stage hepatocellular carcinoma (HCC) or early-stage pancreatic ductal adenocarcinoma (PDAC).
30. A multimodal method for analyzing cfDNA in a patient sample, comprising: isolating the cfDNA from the patient sample; converting 5mC and / or 5hmC residues in the sample to DHU residues to provide a modified cfDNA sample; sequencing the modified cfDNA sample to identify methylated regions in the sample, wherein a shift from cytosine (C) to thymine (T) or from cytosine (C) to DHU in the modified cfDNA sample as compared to an unmodified reference cfDNA provides the position of either 5mC or 5hmC in the cfDNA. Identifying the methylated region. a) determining the copy number variation of one or more targets in the modified cfDNA sample; b) determining the tissue of origin or one or more targets in the modified cfDNA sample; c) determining the fragmentation profile of the modified cfDNA sample; and d) identifying one or more single nucleotide variants in the modified cfDNA sample, performing one or more additional analysis steps on the modified cfDNA selected from the group consisting of, and, said multimodal method.