A method for detecting cancer using extraembryonic methylated CpG islands.

By detecting extra germ layer-specific methylated CpG islands and combining them with computational analysis, a highly sensitive non-invasive cancer diagnostic method has been developed, which solves the problem of insufficient sensitivity in early cancer detection in existing technologies and achieves highly sensitive and specific diagnosis of a variety of cancers.

JP2026083227APending Publication Date: 2026-05-19PRESIDENT & FELLOWS OF HARVARD COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PRESIDENT & FELLOWS OF HARVARD COLLEGE
Filing Date
2026-03-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing cancer detection methods lack sufficient sensitivity and specificity in the early stages, especially liquid biopsy-based cfDNA methods, which are limited by tumor heterogeneity and have difficulty effectively detecting early-stage cancer.

Method used

By detecting extra germ layer-specific methylated CpG islands that are specific to most human cancer types and combining this with computational analysis of methylated alleles, a highly sensitive non-invasive cancer diagnostic method was developed. The MBD2 protein-based enrichment method was used to enrich methylated sequences of cfDNA.

Benefits of technology

It achieves highly sensitive, non-invasive early diagnosis of a variety of cancers, with 100% sensitivity and 95% specificity, and can detect cancers at various stages, including multiple cancer types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083227000044
    Figure 2026083227000044
  • Figure 2026083227000045
    Figure 2026083227000045
  • Figure 2026083227000046
    Figure 2026083227000046
Patent Text Reader

Abstract

Providing a method for cancer detection using extraembryonic methylated CpG islands. [Solution] The present invention relates to a method for characterizing cell-free DNA (cfDNA), detecting cancer, detecting cancer eradication, and determining the probability distribution of haplotypes. The method determines the proportion of fully methylated haplotypes using data from genomic sequences derived from methylated CpG islands (CGI) in the extraembryonic ectoderm (ExE) genome to characterize cfDNA samples and detect specific cancers. In one embodiment, the method described herein relates to characterizing cell-free DNA (cfDNA) samples derived from a target.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 126,863, filed on 17 December 2020, and U.S. Provisional Patent Application No. 63 / 246,306, filed on 20 September 2021, the entire teachings of which are incorporated herein by reference. [Background technology]

[0002] Background of the present invention The vast majority of cancer-related deaths are due to complications of metastatic disease. Modern cancer treatments generally fail against metastatic disease due to tumor evolution[1], allowing heterologous cancer cell populations to evade treatment, colonize new sites, and acquire novel traits that enable them to become more aggressive over time. Early diagnosis of the disease results in a significantly improved prognosis compared to advanced-stage disease and can be based on imaging-based or blood-based tests[2]. Serum-based protein biomarkers such as cancer antigen-125 (CA-125)[3], carcinoembryonic antigen (CEA)[4], and prostate-specific antigen (PSA)[5] have been used to track the progression of certain cancer types, but they lack the sensitivity and specificity required to detect early-stage disease.

[0003] Liquid biopsies based on cell-free DNA (cfDNA) analysis have attracted considerable interest due to their potential to identify cancer-causing mutations in the plasma of patients with early-stage disease. However, intertumor and intratumor heterogeneity limits the sensitivity of these methods, as recurrent clonal mutations are rare. More recent advances have been based on cfDNA methylation profiling to detect and classify reads originating from specific tumor types. While these approaches are promising, they need to be optimized for each tumor type. Therefore, due to tumor heterogeneity, there is a need for innovative methods for cancer detection with higher sensitivity. [Overview of the project] [Means for solving the problem]

[0004] Summary of the present invention The cancer screening method was discovered by detecting specific pan-oncological methylation signatures in cfDNA. Specifically, pan-oncological methylation signatures are based on loci that are preferentially methylated in the extraembryonic ectoderm, which are present across most human cancer types, unlike epiblasts.

[0005] Based on these findings, a highly sensitive identification of tumor-derived cfDNA was developed to enable non-invasive early diagnosis of human cancer. Computational analysis of methylation haplotypes identified from individual bisulfite-converted reads reduced background signals originating from normal cell types. The results provide the ability to detect extraembryonic methylation signatures in plasma samples from patients with cancerous disease at various stages. This invention improves upon previous screening methods by providing a highly sensitive, non-invasive pan-cancer diagnosis of disease based on cell-free methylation patterns in plasma.

[0006] In one embodiment, the present invention relates to a method for characterizing a cell-free DNA (cfDNA) sample derived from a subject, comprising the steps of: receiving sequencing data from the cfDNA sample, which includes methylation sequence reads for a genomic sequence, wherein the genomic sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not methylated in the corresponding epiblast or adult tissue; determining the proportion of haplotypes of the fully methylated genomic sequence; and, if the proportion of haplotypes is greater than a significance threshold, characterizing the cfDNA sample as containing fully methylated cfCDNA.

[0007] In certain embodiments, each haplotype includes five CGIs that are methylated in the ExE genome and not methylated in the corresponding epiblast or adult tissue. In certain embodiments, the cfDNA sample includes 0.01% to 0.1% tumor DNA. In certain embodiments, the sequencing data includes sequence information for less than 0.3% of the genome in question. In certain embodiments, the sequencing data includes sequence information substantially limited to one or more regions of the genome in question that have multiple CGIs that are methylated in the ExE genome and not methylated in the corresponding epiblast or adult tissue. In certain embodiments, fully methylated haplotypes are compared to one or more pre-established fully methylated haplotype signatures, and the cfDNA sample is further characterized as corresponding to or not corresponding to a pre-established fully methylated haplotype signature. In certain embodiments, the pre-established fully methylated haplotype signatures are identified by methods including random forest, support vector machine, or deep learning analysis. In certain embodiments, sequencing data including methylated sequence reads for genomic sequences from a cfDNA sample are enriched for the methylated sequences. In certain embodiments, the enrichment includes an MBD2 protein-based enrichment method. In certain embodiments, the cfDNA sample is obtained from plasma, urine, feces, menstrual fluid, or lymph. In some embodiments, the method further includes a step of determining the tissue of origin from the sequencing data.

[0008] In one embodiment, the present invention relates to a method for detecting cancer in a subject, comprising the steps of: receiving sequencing data including methylation sequence reads for a genomic sequence from a cfDNA sample derived from the subject, wherein the genomic sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not in the corresponding epiblast or adult tissue; determining the proportion of fully methylated haplotypes in the genomic sequence; and detecting cancer in the subject if the proportion of fully methylated haplotypes is greater than a significance threshold.

[0009] In certain embodiments, each haplotype includes five CGIs that are methylated in the ExE genome but not methylated in the corresponding epiblast or adult tissue. In certain embodiments, the cfDNA sample contains 0.01% to 0.1% tumor DNA. In certain embodiments, the sequencing data includes sequence information for less than 0.3% of the genome of interest. In certain embodiments, the sequencing data includes sequence information substantially limited to one or more regions of the genome of interest having multiple CGIs that are methylated in the ExE genome but not methylated in the corresponding epiblast or adult tissue. In certain embodiments, fully methylated haplotypes are compared to one or more pre-established fully methylated haplotype signatures corresponding to one or more tumor types to detect the presence or absence of one or more tumor types in the interest.

[0010] In certain embodiments, one or more tumor types include one or more of acute myeloid leukemia, bladder cancer, breast cancer, colon cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, or gastric cancer. In certain embodiments, pre-established fully methylated haplotype signatures corresponding to one or more tumor types are identified by methods including random forest, support vector machine, or deep learning analysis. In certain embodiments, sequencing data including methylated sequence reads for genomic sequences from cfDNA samples are enriched for sequences containing methylation. In certain embodiments, enrichment includes an MBD2 protein-based enrichment method. In certain embodiments, cfDNA samples are obtained from plasma, urine, feces, menstrual fluid, or lymph. In certain embodiments, the presence of cancer is detected in the sample with 100% sensitivity and 95% specificity. In certain embodiments, the cancer is stage I or stage III. In certain embodiments, cancer is selected from the group including adenocarcinoma, acute myeloid leukemia, bladder cancer, breast cancer, colon cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, stomach cancer, and uterine cancer. In certain embodiments, the method further includes the step of treating the subject for cancer if cancer is detected in the subject. In certain embodiments, the method further includes the step of determining the tissue of origin from sequencing data.

[0011] In one embodiment, the present invention relates to a method for detecting the eradication of cancer from a subject, comprising the steps of: receiving sequencing data including methylation sequence reads for a genomic sequence from a cfDNA sample derived from the subject after cancer treatment, wherein the genomic sequence includes a plurality of CGIs that are methylated in the ExE genome but not in the corresponding epiblast or adult tissue; determining the proportion of haplotypes of fully methylated genomic sequences; and detecting cancer in the subject if the proportion of fully methylated haplotypes is greater than a significance threshold, wherein if no cancer is detected in the subject, the cancer is eradicated from the subject.

[0012] In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple CGIs methylated in the extraembryonic ectoderm (ExE) genome. In certain embodiments, the genome sequence includes 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes one or more sequences provided in Table 3.

[0013] In one embodiment, the present invention relates to a method for determining the probability distribution of haplotypes, comprising the steps of: receiving sequencing data including methylated sequence reads for a genomic sequence from a cfDNA sample, wherein the genomic sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome and not methylated in the corresponding epiblast or adult tissue; assigning a training or validation set based on the methylated ExE CGI data; applying a machine learning method to estimate the probability distribution of all haplotypes across the ExE sites; and determining one or more classifications of tumor samples versus normal samples based on the predictive scores obtained from the machine learning method.

[0014] In certain embodiments, the machine learning method is random forest. In certain embodiments, the machine learning method is support vector machine. In certain embodiments, the machine learning method is deep learning. In certain embodiments, the method further includes a method step of evaluating the performance of the prediction, including performing an in-silico simulation by comparing sequencing reads randomly sampled from an epiblast or adult tissue with ExE reads. In certain embodiments, the method further includes a step of determining the tissue of origin from the sequencing data.

[0015] Some aspects of the present disclosure are directed to a method for determining tissue origin, comprising: a) receiving targeted bisulfite sequencing data comprising methylation sequence reads for genomic sequences from a cfDNA sample, wherein the genomic sequences comprise a plurality of CpG islands (CGIs) that are methylated in the genome of the extraembryonic ectoderm (ExE) and unmethylated in the corresponding epiblast or adult; and b) determining the tissue of origin by calculating the relative abundance of haplotypes from the methylated genomic regions by defining a tissue-specific index (TSI) for each haplotype. In certain embodiments, the TSI is calculated by the following formula:

Number

number

[0016] [Figure 1A] Figure 1A shows the mouse E6.5 conception product used to characterize the DNA methylation landscape of embryonic and extraembryonic tissues by comparing epiblast and ExE (exoembryonic ectoderm).

[0017] [Figure 1B]Figure 1B shows genetically conserved ExE hyperCGIs. The mean conservation score (phyloP30-way) is plotted as a function of the distance to the center of the CGI. Only CGIs close to the TSS (+ / -2000bp) are included.

[0018] [Figure 1C] Figure 1C shows mouse ExE hyperCGI lifted over to orthologous CGI in humans.

[0019] [Figure 1D] Figure 1D shows ExE HyperCGI accurately distinguishing cancer from normal samples. The performance of ExE HyperCGI in cancer prediction was tested using 13 TCGA cancer types, including matched normal tissue. Half samples were randomly selected to be trained by an SVM with a Gaussian kernel, and the resulting model was used to predict the remaining half samples as either tumor or normal. The results are shown as ROC curves, with the area under the curve (AUC) indicated.

[0020] [Figure 1E] Figure 1E shows that cancers are genetically heterogeneous and epigenetically homogeneous. Further summarizing the results from Figure 1D, we show the fraction of samples in each cancer type as accurately predicted by ExE Hyper-CGI. In parallel, we also show the fraction of samples containing the TP53 mutation.

[0021] [Figure 2A] Figure 2A shows an example of DNA methylation haplotypes. The methylation pattern of CpGs on each sequencing fragment represents a distinct DNA methylation haplotype that can be classified as unmethylated reads, mismatched reads, or fully methylated reads. The percentage of fully methylated reads (PMR) is defined as the fraction of fully methylated reads.

[0022] [Figure 2B]Figure 2B shows that using a percentage of fully methylated reads (PMRs) significantly reduces background noise in normal cells. Sequencing reads from publicly available WGBS data at the OTX2 locus were aggregated to increase coverage for tumor and normal samples, respectively.

[0023] [Figure 2C] Figure 2C shows an in silico simulation. Sequencing reads from ExE (tumor-like) cells were spiked onto reads from epiblast (normal-like) cells. The fraction of ExE-derived reads corresponds to 1%, 0.1%, or 0.01% in the three sets, respectively. In the negative control, all reads were randomly sampled from epiblast cells. Predictive results are shown using PMR, MHL, and mean methylation-based methods.

[0024] [Figure 3A] Figure 3A shows a typical workflow for the targeted bisulfite sequencing used. MBD enrichment is optional but can be used to specifically enrich methylated reads.

[0025] [Figure 3B] Figure 3B shows the uniformity of hybrid capture. On-target coverage was normalized by the average coverage in the designed region. This curve represents the fraction of loci with coverage higher than a given threshold.

[0026] [Figure 3C] Figure 3C shows the efficiency of targeted sequencing. To evaluate the efficiency of targeted sequencing, the same biological samples were profiled using WGBS and targeted BS. Normalized coverage is shown as a function of the distance to the center of the designed CGI.

[0027] [Figure 3D]Figure 3D shows the enrichment of methylated haplotypes by proteins containing a methyl-CpG binding domain (MBD). The enrichment efficiency is measured by the percentage of methylated reads.

[0028] [Figure 4A] Figure 4A shows the correlation of normalized counts between two assays (targeted BS with and without MBD enrichment). Targeted BS was performed on four samples (HuES64, HCT116, normal uterine, and uterine cancer) under two conditions, with and without MBD enrichment. For each DNA methylation haplotype, the correlation of normalized counts between the two assays was evaluated. All 32 DNA methylation haplotypes were classified into six classes based on the length of the fully methylated k-mer.

[0029] [Figure 4B] Figure 4B shows the normalized coverage of fully methylated reads compared between two assays of targeted BS with and without MBD enrichment for uterine cancer and normal uterine cancer. The Pearson correlation coefficient is also shown in the figure.

[0030] [Figure 4C] Figure 4C shows a comparison of normalized coverage of fully methylated reads between two assays, targeted BS and WGBS, for uterine cancer and normal uterus. The Pearson correlation coefficient is also shown in the figure.

[0031] [Figure 5A] Figure 5A shows the ultra-high sensitivity detection of cancer in diluted samples of HuES64 DNA mixed with HCT116.

[0032] [Figure 5B] Figure 5B shows ultra-high sensitivity detection of cancer in diluted samples of HuES64 DNA mixed with colon cancer DNA spikein.

[0033] [Figure 5C]Figure 5C shows ultra-sensitive detection of cancer in diluted samples of normal uterine DNA mixed with uterine cancer DNA spike-ins. The fractions of spike-ins in all three experiments included 1%, 0.1%, and 0.01%. Using NMR-based methods, the presence of spike-ins was predicted using an increasing number of top markers.

[0034] [Figure 6] Figure 6 shows that ExE HyperCGI accurately distinguishes cancer from normal samples. The performance of ExE HyperCGI in cancer prediction was tested using 13 TCGA cancer types, including matched normal tissue. The pan-cancer cohort consisted of 685 tumor samples and 710 normal samples, which were subdivided into training and validation sets of equal sample size. A random forest (RF) was implemented using the `randomForest` function of the `randomForest` R package with default parameter settings. False and true positive rates were calculated using the `roc` function of the `pROC` R package based on `out-of-bag` votes for the training data. RF was able to classify tumor samples with high specificity and sensitivity (AUC=0.98).

[0035] [Figure 7] Figure 7 shows the percentage of fully methylated reads (PMRs) and a comparison with three other metrics used in the literature. Five patterns of methylation haplotype combinations (schematic diagram) are used to illustrate the differences in methylation frequency, haplotype number, methylation haplotype load (MHL), and PMR.

[0036] [Figure 8]Figure 8 shows a schematic diagram of the method for quantifying DNA methylation using PMR. It is shown that 16 DNA methylation haplotypes represent schematic sequencing reads aligned to a gene locus. For each DNA methylation haplotype, fully methylated k-mers and the total number of k-mers were counted for a given width of k-mers. PMR is then defined as the proportion of fully methylated k-mers across all reads aligned to a given gene locus.

[0037] [Figure 9A] This paper demonstrates cancer prediction using mean methylation for simulated data. To evaluate the performance of mean methylation in cancer prediction, in silico simulations were performed by randomly sampling sequencing reads from normal-like tissue epiblasts and tumor-like tissue ExE as spike-ins. The spike-in fractions ranged from 0.01% to 1%, which is consistent with the fraction of ctDNA in cell-free DNA. Compared to epiblasts, ExE was identified as having higher mean methylation in ExE, as shown in red.

[0038] [Figure 9B] Figure 9B shows a simulated sample compared to an epiblast using the CGI defined in the previous step, with the resulting mean methylation difference represented as a box plot for each spike-in group.

[0039] [Figure 9C] Figure 9C shows the number of CGIs with increased or decreased mean methylation counted, and the significance p-values ​​estimated by a one-sided binomial test to predict the presence of ExE DNA.

[0040] [Figure 10A]Figure 10A shows cancer prediction using MHL on simulated data within a silico simulation, by randomly sampling sequencing reads from normal-like tissue epiblasts and tumor-like tissue ExE as spike-ins to evaluate MHL's performance in cancer prediction. The fraction of spike-ins ranges from 0.01% to 1%, which is consistent with the fraction of ctDNA in cell-free DNA. Comparing ExE to epiblasts, MHL identified higher CGI in ExE, as shown in red.

[0041] [Figure 10B] Figure 10B shows the simulated sample compared to the epiblast using the CGI defined in the previous step, and the resulting MHL difference is represented as a box plot for each spike-in group.

[0042] [Figure 10C] Figure 10C shows the number of CGIs with increased or decreased MHL counts, respectively, and the significance p-values ​​estimated by a one-sided binomial test to predict the presence of ExE DNA.

[0043] [Figure 11A] Figure 11A shows that in silico simulations were performed by randomly sampling sequencing reads from normal-like tissue epiblasts and tumor-like tissue ExE as spike-ins to evaluate the performance of PMR in cancer prediction. The spike-in fractions ranged from 0.01% to 1%, which is consistent with the fraction of ctDNA in cell-free DNA. Comparing ExE to epiblasts, PMR identified higher CGIs in ExE, as shown in red.

[0044] [Figure 11B] Figure 11B shows the simulated sample compared to the epiblast using the CGI defined in the previous step, and the resulting PMR difference is represented as a box plot for each spike-in group.

[0045] [Figure 11C] Figure 11C shows the number of CGIs with increased or decreased PMRs, respectively, and the significance p-value was estimated by a one-sided binomial test to predict the presence of ExE DNA.

[0046] [Figure 12] Figure 12 shows the identification of the optimal k-mer length for PMR. PMR is a function of k-mer length. To identify the optimal k-mer for cancer prediction, simulated data using a 0.01% ExE spike-in (method) with the PMR method were tested. Maximum sensitivity was achieved when the k-mer length was set to 5.

[0047] [Figure 13] Figure 13 shows that MHL is a biased metric for measuring DNA methylation across the entire assay. Targeted BS was performed on four samples (HuES64, HCT116, uterine cancer, and normal uterine cells) under two conditions: with and without MBD enrichment. MHL was compared for each of the four samples between the two assays (targeted BS with and without MBD enrichment).

[0048] [Figure 14] Figure 14 shows that PMR is a biased metric for measuring DNA methylation across the entire assay. Targeted BS was performed on four samples (HuES64, HCT116, uterine cancer, and normal uterine cells) under two conditions: with and without MBD enrichment. PMR was compared for each of the four samples between the two assays (targeted BS with and without MBD enrichment).

[0049] [Figure 15]Figure 15 shows NMR as an unbiased metric for measuring DNA methylation across the entire assay. Performance-targeted BS was performed on four samples (HuES64, HCT116, uterine cancer, and normal uterine cells) under two conditions: with and without MBD enrichment. NMR was compared for each of the four samples between the two assays (targeted BS with and without MBD enrichment). A Pearson correlation coefficient of 0.99 was observed for all four samples.

[0050] [Figure 16A] Figure 16A shows the detection of cancer in diluted samples using targeted BS with MBD enrichment. HuES64 DNA was mixed with HCT116 or colon cancer DNA spike-in, and normal uterine DNA was mixed with uterine cancer DNA spike-in. The spike-in fractions in all three experiments included 1%, 0.1%, and 0.01%. The experiments in Figure 16A were performed in parallel with 1 μg of input DNA.

[0051] [Figure 16B] Figure 16B shows parallel experiments using 50 ng of DNA. Using NMR basing, the presence of spike-ins was predicted using an increasing number of top markers.

[0052] [Figure 17A] Figure 17A shows an example of how an NMR-based cancer prediction pipeline works with HCT116 dilution data. HCT116 was compared to human ES cells (HuES64) to identify CGIs with higher NMR in HCT116, using a cutoff of 0.1. These CGIs were then derivedly ranked based on the difference in NMR between HCT116 and HuES64. The top 200 CGIs were selected as markers. A scatter plot of the NMRs is shown, with the selected markers highlighted in red. The NMR of the test samples was compared to that of HuES64.

[0053] [Figure 17B]Figure 17B shows the ΔNMR box plots for 1%, 0.1%, and 0.1% spike-in.

[0054] [Figure 17C] Figure 17C shows the number of markers counted for increases in NMR (ΔNMR>0) and decreases in NMR (ΔNMR<0) to test whether ΔNMR is statistically greater than 0. The P-value was calculated by a one-sided binomial test.

[0055] [Figure 18A] Figure 18A shows an example of how an NMR-based cancer prediction pipeline works for colon cancer dilution data. Colon cancer was compared to normal colon to identify CGIs with higher NMR in colon cancer, with a cutoff value of 0.1. These CGIs were then derivedly ranked based on the difference in NMR between tumor samples and Hu64ES. The top 200 CGIs were selected as markers. A scatter plot of NMR (normal) is shown.

[0056] [Figure 18B] Figure 18B shows a scatter plot of NMR (ES).

[0057] [Figure 18C] Figure 18C shows the ΔNMR box plots for 1%, 0.1%, and 0.1% spike-in.

[0058] [Figure 18D] Figure 18D shows the number of markers counted for increases in NMR (ΔNMR>0) and decreases in NMR (ΔNMR<0) to test whether ΔNMR is statistically greater than 0. The P-value was calculated by a one-sided binomial test.

[0059] [Figure 19]Figure 19 shows the identification of the optimal k-mer length for NMR. NMR is a function of k-mer length. To identify the optimal k-mer for cancer prediction, colon cancer spike-in data were tested with 0.01% colon cancer DNA. Maximum sensitivity was achieved when the k-mer length was set to 5.

[0060] [Figure 20A] Figure 20A shows cancer detection in diluted samples using mean methylation. HuES64 DNA was mixed with HCT116 or colon cancer DNA spike-in, and normal uterine DNA was mixed with uterine cancer DNA spike-in. The spike-in fractions in all three experiments included 1%, 0.1%, and 0.01%.

[0061] [Figure 20B] Figure 20B shows an MHL-based method for predicting the presence of spike-ins using an increasing number of top markers.

[0062] [Figure 21-1] Figure 21 shows the predicted fraction of tumor DNA in the colon cancer cohort. The predicted results for each sample are shown by the vertical dashed lines. [Figure 21-2] Figure 21 shows the predicted fraction of tumor DNA in the colon cancer cohort. The predicted results for each sample are shown by the vertical dashed lines.

[0063] [Figure 22-1] Figure 22 shows the predicted fraction of tumor DNA in the breast cancer cohort. The predicted results for each sample are shown in the figure, indicated by the vertical dashed lines. [Figure 22-2] Figure 22 shows the predicted fraction of tumor DNA in the breast cancer cohort. The predicted results for each sample are shown in the figure, indicated by the vertical dashed lines.

[0064] [Figure 23] Figure 23 shows diagrams of different CGI areas analyzed for cancer screening methods. [Modes for carrying out the invention]

[0065] Detailed description of the present invention Methods for characterizing cell-free DNA (cfDNA) samples

[0066] In one embodiment, the method described herein is for characterizing a cell-free DNA (cfDNA) sample derived from a subject, and includes the steps of: receiving sequencing data from the cfDNA sample, including methylation sequence reads for a genomic sequence, wherein the genomic sequence includes multiple CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not methylated in the corresponding epiblast or adult tissue; determining the proportion of haplotypes of fully methylated genomic sequences; and, if the proportion of haplotypes is greater than a significance threshold, characterizing the cfDNA sample as containing fully methylated cfCDNA.

[0067] In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple methylated CGIs in the ExE genome. In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple methylated CGIs in the ExE genome and chr14(human) bases 57,258,577~57,282,377. In certain embodiments, the genome sequence includes a continuous sequence of up to 8 megabases of the human genome containing multiple methylated CGIs in the extraembryonic ectoderm (ExE) genome. In certain embodiments, the genome sequence includes a continuous sequence of 6.1 megabases of the human genome containing multiple methylated CGIs in the extraembryonic ectoderm (ExE) genome. In certain embodiments, the genome sequence includes one or more sequences provided in Table 3.

[0068] In certain embodiments, the genome sequence includes 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes up to 100 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes up to 500 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes up to 1000 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence includes up to 1500 CGIs methylated in the ExE genome. In more specific embodiments, the genome sequence includes approximately 1,265 CGIs hypermethylated in the ExE tissue. In even more specific embodiments, the genome sequence includes approximately 473 CGIs hypermethylated in the ExE tissue.

[0069] As used herein, the significance threshold refers to the observed significance value known as the significance predictor (p-value) estimated by a one-sided binomial test to predict the presence of ExE DNA. In certain embodiments, for a 5% fraction of ctDNA in cell-free DNA, the p-value (i.e., the minimum p-value indicating significance) is 5.3 x 10⁻¹⁴. -145 In a specific embodiment, the P-value for 1% fraction of ctDNA in cell-free DNA is 3.9 x 10⁻¹⁰. -78 In a specific embodiment, the P-value for 0.1% fraction of ctDNA in cell-free DNA is 6.5 x 10⁻¹⁰. -19 In a specific embodiment, the P-value for 0.01% fraction of ctDNA in cell-free DNA is 6.3 x 10⁻¹⁰. -4 In a specific embodiment, the P-value for 5% fraction of ctDNA in cell-free DNA is 1.9 x 10⁻¹⁰. -78 In a specific embodiment, the P-value for 1% fraction of ctDNA in cell-free DNA is 7.4 x 10⁻¹⁰. -34It is. In certain embodiments, for a ctDNA fraction of 0.1% in cell-free DNA, the P-value is 4.2x10 -10 It is. In certain embodiments, for a ctDNA fraction of 0.01% in cell-free DNA, the P-value is 3.1x10 -2 It is. In certain embodiments, for a ctDNA fraction of 5% in cell-free DNA, the P-value is 4.5x10 -26 It is. In certain embodiments, for a ctDNA fraction of 1% in cell-free DNA, the P-value is 3.4x10 -15 It is. In certain embodiments, for a ctDNA fraction of 0.1% in cell-free DNA, the P-value is 1.1x10 -8 It is. In certain embodiments, for a ctDNA fraction of 0.01% in cell-free DNA, the P-value is 4.5x10 -6 It is. In certain embodiments, at a fraction of 1%, the P-value is 1.3x10 -58 It is. In certain embodiments, at a fraction of 0.1%, the P-value is 2.0x10 -37 It is. In certain embodiments, at a fraction of 0.01%, the P-value is 3.9x10 -9 It is. In certain embodiments, at a fraction of 1%, the P-value is 1.6x10 -54 It is. In certain embodiments, at a fraction of 0.1%, the P-value is 3.3x10 -26 It is. In certain embodiments, at a fraction of 0.01%, the P-value is 1.1x10 -5 It is.

[0070] In certain embodiments, the cfDNA sample contains 0.01% to 0.1% tumor DNA. In certain embodiments, the cfDNA sample contains 0.01% tumor DNA. In certain embodiments, the cfDNA sample contains 0.02% tumor DNA. In certain embodiments, the cfDNA sample contains 0.03% tumor DNA. In certain embodiments, the cfDNA sample contains 0.04% tumor DNA. In certain embodiments, the cfDNA sample contains 0.05% tumor DNA. In certain embodiments, the cfDNA sample contains 0.06% tumor DNA. In certain embodiments, the cfDNA sample contains 0.07% tumor DNA. In certain embodiments, the cfDNA sample contains 0.08% tumor DNA. In certain embodiments, the cfDNA sample contains 0.09% tumor DNA. In certain embodiments, the cfDNA sample contains 0.1% tumor DNA. In certain embodiments, the cfDNA sample contains 0.15% tumor DNA. In certain embodiments, the cfDNA sample contains 0.2% tumor DNA. In certain embodiments, the cfDNA sample contains 0.25% tumor DNA. In certain embodiments, the cfDNA sample contains 0.3% tumor DNA. In certain embodiments, the cfDNA sample contains 0.35% tumor DNA. In certain embodiments, the cfDNA sample contains 0.25% tumor DNA. In certain embodiments, the cfDNA sample contains 0.3% tumor DNA. In certain embodiments, the cfDNA contains 0.4% tumor DNA. In certain embodiments, the cfDNA contains 0.5% or more tumor DNA. In certain embodiments, the cfDNA contains 1% or more tumor DNA. In certain embodiments, the cfDNA contains 1.5% or more tumor DNA. In certain embodiments, the cfDNA contains 2% or more tumor DNA. In certain embodiments, the cfDNA contains 3% or more tumor DNA. In certain embodiments, the cfDNA contains 4% or more tumor DNA. In certain embodiments, cfDNA contains 5% or more tumor DNA.

[0071] In certain embodiments, the sequencing data includes sequence information for less than 0.01% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.05% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.1% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.2% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.3% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.4% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.5% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.6% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.7% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.8% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 0.9% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.1% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.2% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.3% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.4% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.5% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.6% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.7% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.8% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.9% of the target genome. In certain embodiments, sequencing data may contain less than 2% of the sequence information of the target genome.In certain embodiments, the sequencing data may contain sequence information for less than 5% of the target genome. In certain embodiments, the sequencing data may contain sequence information for less than 10% of the target genome.

[0072] In certain embodiments, each haplotype contains five CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains four CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains three CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains two CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains one CGI that is methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains six CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains seven CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains eight CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains nine CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains ten CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue).

[0073] In certain embodiments, the sequencing data includes sequence information substantially limited to one or more regions of the target genome having multiple CGIs that are methylated in the ExE genome but not methylated in the corresponding epiblast or adult tissue. In certain embodiments, one or more regions of the target genome are approximately 1200 CGIs as a pan-oncological methylation signature (e.g., shown in Table 3). In certain embodiments, one or more regions are 1 to 5 CGI patterns representing individual DNA methylation haplotypes. In certain embodiments, the region is an 8-megabase region. In certain embodiments, the 8-megabase region includes CHR14:57,258,577~57,282,337. In certain embodiments, the genomic region includes one or more sequences provided in Table 3.

[0074] In certain embodiments, fully methylated haplotypes are compared to one or more pre-established fully methylated haplotype signatures. The cfDNA sample is further characterized as corresponding to or not corresponding to a pre-established fully methylated haplotype signature. In some embodiments, fully methylated haplotypes are generally normalized (i.e., an NMR spectrum is obtained) by the total number of haplotypes across the entire region.

[0075] In certain embodiments, pre-established fully methylated haplotype signatures are identified by methods including random forests, support vector machines, or deep learning analysis. As used herein, the random forest algorithm operates by constructing a large number of decision trees during training time and outputting the classification or mean / mean predictor / regression of the individual trees.

[0076] As used herein, a support vector machine is a machine learning method that constructs a set of hyperplanes that can be used for classification, regression, or detection of multidimensional data. As used herein, deep learning analysis refers to a class of machine learning algorithms that use multiple layers to gradually extract higher-level features from raw input.

[0077] In certain embodiments, the sequencing data includes methylated sequence reads for a genomic sequence from a cfDNA sample enriched with methylated sequences. In certain embodiments, the enrichment includes a methyl-DNA binding protein-based enrichment method. In certain embodiments, the methyl-DNA binding protein of the enrichment method is a methyl-binding domain (MBD) selected from MBD1, MBD2, MBD3, and MBD4.

[0078] As used herein, “sample” is not limited and may be any suitable fluid disclosed herein. In some embodiments, the sample is blood, serum, plasma, urine, feces, menstrual fluid, lymph, and other bodily fluids.

[0079] As used herein, "CpG" and "CpG dinucleotide" are used interchangeably and refer to a dinucleotide sequence containing adjacent guanine and cytosine, with cytosine located at the 5' position of guanine.

[0080] As used herein, “CpG island” or “CGI” refers to a region with a high frequency of CpG sites. This region is at least 200 bp, has a GC percentage greater than 50%, and an observed-to-expected CpG ratio greater than 60%.

[0081] As used herein, “haplotype” refers to a combination of CpG sites found on the same chromosome. Similarly, “DNA methylation haplotype” describes the DNA methylation status of CpG sites on the same chromosome.

[0082] In certain embodiments, a sample (e.g., a fluid sample) is screened using whole-genome bisulfite sequencing (WGBS), TCGA Illumina Infinium Human Methylation 450K BeadChip sequencing (TCGA), and / or reductive bisulfite sequencing (RRBS), or by other suitable methylation detection assays known in the Art.

[0083] In certain embodiments, the invention disclosed herein relates to a method for detecting circulating tumor DNA (ctDNA) in a sample using the proportion of matched methylated reads (PMR) (i.e., fully methylated haplotype). In certain embodiments, a methylated sequence is obtained for the sample, and at least one CpG island (CGI) is identified in that methylated sequence. The PMR of the identified CpG island is calculated and then compared to a control background of normal tissue or epiblast. If the PMR of the sample is greater than that of the control background (e.g., a higher signal by bank sum test), the presence of ctDNA is detected in the sample.

[0084] The presence of ctDNA can be detected in cfDNA with higher sensitivity and specificity than methods previously known to those skilled in the art. For example, ctDNA can be detected in a sample using PMR with sensitivity exceeding 75%, 80%, 85%, 90%, 95%, or 99%. In certain embodiments, ctDNA is detected in a sample using PMR with 100% sensitivity. ctDNA can be detected in a sample using PMR with specificity exceeding 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%. In certain embodiments, ctDNA is detected in a sample using PMR with 95% specificity. In some embodiments, ctDNA is detected in a sample using PMR with at least 90% sensitivity and at least 90% specificity. In some embodiments, ctDNA is detected in a sample using PMR with at least 100% sensitivity and at least 95% specificity.

[0085] As used herein, “sensitivity” measures the percentage of positives correctly identified in cfDNA (i.e., the presence of ctDNA).

[0086] As used herein, “specificity” measures the proportion of negative (i.e., non-ctDNA) that are correctly identified in cfDNA.

[0087] The amount of ctDNA detected in the sample can be measured and quantified. In some embodiments, the sample contains 0.005% to 1.5% ctDNA, 0.01% to 1% ctDNA, 0.05% to 0.5% ctDNA, and 0.1% to 0.3% ctDNA. In some embodiments, the sample contains 0.01% ctDNA. In certain embodiments, the presence of 0.01% ctDNA can be measured using PMR with approximately 100% sensitivity and approximately 95% specificity. -4 It is detected in cfDNA at the p-value cutoff.

[0088] In some embodiments, the invention disclosed herein relates to a method for screening for cancer by using PMR to detect ctDNA in a sample described herein, wherein the presence of ctDNA in the sample indicates that the subject has cancer.

[0089] The methods described herein may be applied to subjects at risk of cancer or at risk of cancer recurrence. Subjects are not limited to any suitable subject. In some embodiments, subjects are individuals diagnosed with cancer, currently suffering from cancer, at risk of developing cancer, or suspected of having cancer. In some embodiments, subjects are human. In some embodiments, subjects are non-human mammals. In some embodiments, subjects are non-mammalian vertebrates. In some embodiments, subjects are common laboratory animals. Subjects at risk of cancer may, for example, be subjects who have not been diagnosed with cancer but are at high risk of developing cancer. Determining whether a subject is considered to be at “high risk” of cancer is within the scope of the art. Any suitable test(s) and / or criteria may be used. For example, a subject may be considered to be at “high risk” of developing cancer if any one or more of the following apply: (i) the subject has a hereditary mutation or genetic polymorphism associated with an increased risk of developing or having cancer (compared to other members of the general population who do not have such mutation or genetic polymorphism) (e.g., hereditary mutations in certain TSGs are known to be associated with an increased risk of cancer); (ii) the subject has a gene or protein expression profile and / or the presence of certain substances in a sample (e.g., blood) taken from the subject that is associated with an increased risk of developing or having cancer compared to the general population; (iii) the subject has one or more risk factors such as a family history of cancer or exposure to tumor stimulants or carcinogens (e.g., physical carcinogens such as ultraviolet or ionizing radiation; chemical carcinogens such as asbestos, tobacco or smoke components, aflatoxins, arsenic; biological carcinogens such as certain viruses or parasites); (iv) the subject is of a certain age, e.g., over 60 years. A subject suspected of having cancer may be one or more subjects exhibiting one or more symptoms of cancer, or a subject who has undergone a diagnostic procedure that suggests the possibility of cancer or is consistent with such a procedure.Individuals at risk of cancer recurrence may be those who have been treated for cancer and, for example, appear to be cancer-free based on appropriate assessment methods.

[0090] As used herein, the term "cancer" is intended to be used broadly to refer to any cancerous condition.

[0091] In certain embodiments, cancer is stage I, stage II, stage III, or stage IV. In certain embodiments, cancerous cells are present, but the cancer has not spread to nearby tissues.

[0092] Examples of cancers include adrenal cancer, adrenocortical carcinoma, anal cancer, appendiceal cancer, astrocytoma, atypical teratomatoid / rhabdoid tumor, basal cell carcinoma, cholangiocarcinoma, bladder cancer, bone cancer, brain / CNS cancer, breast cancer, bronchial tumor, cardiac tumor, cervical cancer, cholangiocarcinoma, chondrosarcoma, chordoma, colon cancer, colorectal cancer, craniopharyngioma, ductal carcinoma in situ (DCIS), endometrial cancer, ependymoma, esophageal cancer, sensory neuroblastoma, and euthyroid cancer. Gung's sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, eye cancer, fallopian tube cancer, fibrous histiosarcoma, fibrosarcoma, gallbladder cancer, stomach cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), germ cell tumor, glioma, glioblastoma, head and neck cancer, hemangioblastoma, hepatocellular carcinoma, hypopharyngeal cancer, intraocular melanoma, Kaposi's sarcoma, kidney cancer, laryngeal cancer, leiomyosarcoma, lip cancer, liposarcoma, liver cancer, lung cancer, non-small cell lung cancer, pulmonary carcinoid Tumors, malignant mesothelioma, medullary carcinoma, medulloblastoma, meningioma, melanoma, Merkel cell carcinoma, median ductal carcinoma, oral cancer, myxosarcoma, myelodysplastic syndrome, myeloproliferative neoplasms, nasal and paranasal sinus cancer, nasopharyngeal cancer, neuroblastoma, oligodendroglioma, oral cancer, oral cancer, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, islet cell tumors of the pancreas, papillary carcinoma, paraganglioma, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal gland tumor, pituitary tumor Examples of cancers that fall under this category include, but are not limited to, pleuropulmonary blastoma, primary colorectal cancer, prostate cancer, rectal cancer, retinoblastoma, renal cell carcinoma, renal pelvis and ureteral cancer, rhabdomyosarcoma, salivary gland cancer, sebaceous gland cancer, skin cancer, soft tissue sarcoma, squamous cell carcinoma, small cell lung cancer, small intestine cancer, gastric cancer, sweat gland cancer, synoviomas, testicular cancer, throat cancer, thymic cancer, thyroid cancer, urethral cancer, uterine cancer, uterine sarcoma, vaginal cancer, vascular cancer, vulvar cancer, and Wilms' tumor.In some embodiments of the methods described herein, cancers include adrenocortical carcinoma, urothelial carcinoma of the bladder, invasive breast cancer, cervical and endometrial cancer, cholangiocarcinoma, colonic adenocarcinoma, colorectal adenocarcinoma, lymphoid neoplasm diffuse large B-cell lymphoma, esophageal cancer, FFPE pilot phase II, glioblastoma multiforme, glioma, squamous cell carcinoma of the head and neck, renal pigmentophobe lentigines, and panrenal cohort (KICH). These include KIRC+KIRP), clear cell carcinoma of the kidney, papillary cell carcinoma of the kidney, acute myeloid leukemia, low-grade glioma of the brain, hepatocellular carcinoma of the liver, adenocarcinoma of the lung, squamous cell carcinoma of the lung, mesothelioma, serous cystadenocarcinoma of the ovary, adenocarcinoma of the pancreas, pheochromocytoma and paraganglioma, adenocarcinoma of the prostate, adenocarcinoma of the rectum, sarcoma, melanoma of the skin, adenocarcinoma of the stomach, germ cell tumor of the testis, thyroid cancer, thymic cancer, endometrial cancer of the uterine body, carcinosarcoma of the uterus, and uveal melanoma. In other embodiments, the present invention provides a method for treating subjects requiring treatment for cancer.

[0093] In some embodiments, PMR is used to detect ctDNA in the sample described herein, and the presence of ctDNA indicates that the subject has cancer. The individual is then treated for cancer using any treatment method (e.g., a therapeutic agent or procedure) that is generally known to those skilled in the art.

[0094] For example, treatments or anticancer agents that may be used to treat a subject include anticancer agents, chemotherapeutic agents, surgery, radiotherapy (e.g., gamma radiation, neutron therapy, electron beam therapy, proton therapy, close-range radiotherapy, and whole-body radioisotopes) that are useful in treating a subject requiring treatment for cancer, endocrine therapy, biological response modifiers (e.g., interferon, interleukin), hyperthermia, cryotherapy, agents that reduce any adverse effects, or combinations thereof. Non-limiting examples of cancer chemotherapeutic agents that may be used include, for example, alkylating agents and alkylating agent-like agents, e.g., nitrogen mustard (e.g., chlorambucil, chlormethine, cyclophosphamide, ifosfamide, and melphalan), nitrosourea (e.g., carmustine, fotemustine, lomustine, streptozocin); platinum agents (e.g., alkylating-like agents such as carboplatin, cisplatin, oxaliplatin, BBR3464, satoraplatin), busulfan, dacarbazine, procarbazine, temozolomide, thioTEPA, treosulfan, and uramustine; antimetabolites such as folic acid (e.g., aminopterin, methotrexate, pemetrexed, larcitrexed); purines, e.g., cladribine, clopharabine, fludarabine, mercaptopurine, pentostatin, thioguanine; Pyrimidines such as pecitabine, cytarabine, fluorouracil, phloxuridine, and gemcitabine; spindle toxins / mitotic inhibitors, e.g., taxanes (e.g., docetaxel, paclitaxel), vinca (e.g., vinblastine, vincristine, vindesine, and vinorelbine), and epothilons; cytotoxic / antinum antibiotics, e.g., anthracyclines (e.g., daunorubicin, doxorubicin, epirubicin, idarubicin, mitoxantrone, pixantrone, and barurubicin), compounds naturally produced by various species of Streptomyces (e.g., actinomycin, bleomycin, mitomycin, plicamycin), and hydroxyureas; topoisomerase inhibitors such as Camptotheca (e.g., camptothecin, topotecan, irinotecan) and podophyllum (e.g., etoposide, teniposide);Anti-receptor tyrosine kinases (e.g., cetuximab, panitumumab, trastuzumab), anti-CD20 (e.g., rituximab and tositumomab), and other monoclonal antibodies for cancer treatment such as alemtuzumab, aevatizumab, gemtuzumab; photosensitizers such as aminolevulinic acid, methyl aminolevulinic acid, sodium porfimer, and verteporfin; tyrosine and / or serine / threonine kinase inhibitors, e.g., Abl, Kit, insulin receptor family members(s), V EGF receptor family members (multiple), PDGF receptor family members (multiple), FGF receptor family members (multiple), mTOR, Raf kinase family, phosphatidylinositol (PI) kinases such as PI3 kinase, PI kinase-like kinase family members, cyclin-dependent kinase (CDK) family members, aurora kinase family members (for example, kinase inhibitors such as cediranib, crizoti Nib, dasatinib, erlotinib, gefitinib, imatinib, lapatinib, nilotinib, sorafenib, sunitinib, vandetanib (those already on the market or with efficacy demonstrated in at least one Phase III trial in tumors), growth factor receptor antagonists, other retinoids (such as alitretinoin and tretinoin), altretamine, amsacrin, anagrelide, arsenic trioxide, asparaginase (e.g., pegaparagase), bexarotene, bortezomib, denileukin difutite Examples include cucumbers, estramustine, ixabepyrone, masopropylmethylamine, mitotane, and testolactone; Hsp90 inhibitors; proteasome inhibitors (e.g., bortezomib); angiogenesis inhibitors (e.g., anti-vascular endothelial growth agents); anti-vascular endothelial growth factor agents or VEGF receptor antagonists such as bevacizumab (Avastin); matrix metalloproteinase inhibitors; various apoptosis-promoting agents (such as apoptosis-inducing agents); Ras inhibitors; anti-inflammatory drugs; cancer vaccines; and other immunomodulatory therapies. It should be understood that the above classification is not limiting.

[0095] In some embodiments, the method further includes the step of determining the origin of the tissue from the sequencing data.

[0096] Methods for detecting cancer

[0097] In another aspect, the method described herein relates to a method for detecting cancer in a subject, comprising the steps of: receiving sequencing data including methylation sequence reads for a genomic sequence from a cfDNA sample derived from the subject, wherein the genomic sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not in the corresponding epiblast or adult tissue; determining the proportion of fully methylated haplotypes in the genomic sequence; and detecting cancer in the subject if the proportion of fully methylated haplotypes is greater than a significance threshold.

[0098] Cancer is not limited to any cancer described herein. In certain embodiments, cancer is selected from acute myeloid leukemia, bladder cancer, breast cancer, colon cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, and stomach cancer.

[0099] In certain embodiments, each haplotype contains five CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains four CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains three CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains two CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains one CGI that is methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains six CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains seven CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains eight CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains nine CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue). In certain embodiments, each haplotype contains ten CGIs that are methylated in the ExE genome (but not methylated in the corresponding epiblast or adult tissue).

[0100] In certain embodiments, the cfDNA sample contains 0.01% to 0.1% tumor DNA. In certain embodiments, the cfDNA sample contains 0.01% tumor DNA. In certain embodiments, the cfDNA sample contains 0.02% tumor DNA. In certain embodiments, the cfDNA sample contains 0.03% tumor DNA. In certain embodiments, the cfDNA sample contains 0.04% tumor DNA. In certain embodiments, the cfDNA sample contains 0.05% tumor DNA. In certain embodiments, the cfDNA sample contains 0.06% tumor DNA. In certain embodiments, the cfDNA sample contains 0.07% tumor DNA. In certain embodiments, the cfDNA sample contains 0.08% tumor DNA. In certain embodiments, the cfDNA sample contains 0.09% tumor DNA. In certain embodiments, the cfDNA sample contains 0.1% tumor DNA.

[0101] In certain embodiments, the sequencing data includes sequence information for less than 0.1% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.2% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.3% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.4% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.5% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.6% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.7% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.8% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 0.9% of the target genome. In certain embodiments, the sequencing data includes sequence information for less than 1% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.1% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.2% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.3% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.4% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.5% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.6% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.7% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.8% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 1.9% of the target genome. In certain embodiments, the sequencing data includes sequence information of less than 2% of the target genome.

[0102] In certain embodiments, the sequencing data includes sequence information substantially limited to one or more regions of the genome of a subject having multiple CGIs that are methylated in the ExE genome but not methylated in the corresponding epiblast or adult tissue.

[0103] In certain embodiments, a fully methylated haplotype is compared to one or more pre-established fully methylated haplotype signatures corresponding to one or more tumor types. This method includes determining the presence or absence of one or more tumor types detected in a subject.

[0104] In certain embodiments, pre-established fully methylated haplotype signatures corresponding to one or more tumor types have been identified by methods including random forests, support vector machines, or deep learning analysis.

[0105] In certain embodiments, the sequencing data includes reads of methylated sequences to genomic sequences from a cfDNA sample enriched with methylated sequences. In certain embodiments, the enrichment includes a methyl-DNA binding protein-based enrichment method. In certain embodiments, the methyl-DNA binding protein of the enrichment method is a methyl-binding domain (MBD) selected from MBD1, MBD2, MBD3, and MBD4. In certain embodiments, the enrichment method further includes targeted bisulfite sequencing (targeted BS). In certain embodiments, up to 6.2 Mb of ExE hyper-CGI is enriched. In certain embodiments, the enrichment method achieves enrichment of more than 50 times compared to whole-genome bisulfite sequencing (WGBS). In certain embodiments, the enrichment method achieves enrichment of more than 100 times compared to WGBS. In certain embodiments, the enrichment method achieves enrichment of more than 400 times compared to WGBS.

[0106] In certain embodiments, the cfDNA sample is obtained from plasma, urine, feces, menstrual fluid, or lymph.

[0107] In certain embodiments, the presence of cancer is detected in the sample with 100% sensitivity and 95% specificity. The presence of ctDNA can be detected in cfDNA with higher sensitivity and specificity than methods previously known to those skilled in the art. For example, ctDNA can be detected in the sample using PMR with sensitivity exceeding 75%, 80%, 85%, 90%, 95%, or 99%. In certain embodiments, ctDNA is detected in the sample using PMR with 100% sensitivity. ctDNA can be detected in the sample using PMR with specificity exceeding 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%. In certain embodiments, ctDNA is detected in the sample using PMR with 95% specificity. In some embodiments, ctDNA is detected in the sample using PMR with at least 90% sensitivity and at least 90% specificity. In some embodiments, ctDNA is detected in the sample using PMR with at least 100% sensitivity and at least 95% specificity.

[0108] In certain embodiments, the method further includes the step of treating the subject for cancer if cancer is detected in the subject. The treatment method is not limited and may be any method described herein. In some embodiments, the treatment method is by a chemotherapeutic agent. In some embodiments, the method further includes the step of determining the tissue of origin from sequencing data.

[0109] Methods for detecting cancer eradication

[0110] In another embodiment, the method described herein is intended to detect the eradication of cancer from a subject and includes the steps of: receiving sequencing data, including methylation sequence reads, of a genomic sequence from a cfDNA sample derived from the subject after cancer treatment, wherein the genomic sequence includes multiple CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not in the corresponding epiblast or adult tissue; determining the proportion of fully methylated haplotypes of the genomic sequence; and detecting cancer in the subject if the proportion of fully methylated haplotypes is greater than a significance threshold, wherein if no cancer is detected in the subject, the cancer is eradicated from the subject. The cancer is not limited and may be any suitable cancer described herein. The subject is not limited and may be any subject described herein. In some embodiments, the subject is human.

[0111] In certain embodiments, the genome sequence contains 1 to 1300 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 1 to 25 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 25 to 50 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 50 to 75 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 75 to 100 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 100 to 200 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 200 to 300 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 300 to 400 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 400-500 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 500-600 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 600-700 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 700-800 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 800-900 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 900-1000 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 1000-1100 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 1100-1200 CGIs methylated in the ExE genome. In certain embodiments, the genome sequence contains 1200–1300 methylated CGIs in the ExE genome.

[0112] As used herein, cancer eradication refers to a substantial reduction in cancer cells compared to the original sample. In certain embodiments, substantial reduction means a reduction of 90% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 95% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 98% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 99% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 99.5% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 99.9% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 99.99% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 99.999% or more of cancer cells. In certain embodiments, substantial reduction means a reduction of 100% of cancer cells. In certain embodiments, substantial reduction means the presence of only a small number of cancer cells.

[0113] Method for determining probability distributions

[0114] In another aspect, the present invention relates to a method for determining the probability distribution of haplotypes, comprising the steps of: receiving sequencing data comprising methylated sequence reads for a genomic sequence from a cfDNA sample, wherein the genomic sequence comprises a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome and not methylated in the corresponding epiblast or adult tissue; assigning a training or validation set based on the methylated ExE CGI data; applying a machine learning method to estimate the probability distribution of all haplotypes across the ExE sites; and determining one or more classifications of tumor samples versus normal samples based on a predictive score (P score) used herein, obtained from the machine learning method.

[0115] In certain aspects, the machine learning method is a random forest. In certain aspects, the machine learning method is a support vector machine. In certain aspects, the machine learning method is deep learning.

[0116] In certain embodiments, the method further includes a method for evaluating the predictive performance, which includes performing an in silico simulation by comparing sequencing reads randomly sampled from epiblast or adult tissue with ExE reads. In certain embodiments, the method further includes a step of determining the tissue of origin from the sequencing data.

[0117] Determining the origin of the organization

[0118] Some aspects of this disclosure relate to a method for determining tissue origin, comprising the steps of: a) receiving targeted bisulfite sequencing data, which includes methylated sequence reads for a genomic sequence from a cfDNA sample, wherein the genomic sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome and not methylated in the corresponding epiblast or adult tissue; and b) determining the tissue of origin by calculating the relative abundance of a haplotype from a methylated genomic region by defining a tissue-specific index (TSI) for each haplotype. In certain embodiments, the TSI is calculated by the following formula:

number

[0119] The description of embodiments of this disclosure is not intended to be exhaustive or to limit the disclosure to the exact form disclosed. While specific embodiments and examples of this disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of this disclosure, as will be recognized by those skilled in the art. For example, while a method, step, or function is presented in a given order, alternative embodiments may perform the functions in a different order, or the functions may be performed substantially simultaneously. The teachings of the disclosure provided herein may be applied to other procedures or methods as needed. Further embodiments may be provided by combining the various embodiments described herein. The aspects of this disclosure may be modified as needed to use the composition, functions, and concepts of the above-mentioned references and applications to provide further embodiments of this disclosure. These and other modifications may be made to this disclosure in light of the detailed description.

[0120] Certain elements of any of the embodiments described above can be combined with or replaced by elements of other embodiments. Furthermore, while the advantages related to specific embodiments of this disclosure have been described in the context of those embodiments, other embodiments may also demonstrate such advantages, and not all embodiments are required to demonstrate such advantages in order to fall within the scope of this disclosure.

[0121] All identified patents and other publications are expressly incorporated herein by reference, for example, for the purpose of explaining and disclosing methodologies described in such publications that may be used in connection with the present invention. These publications are provided solely for their disclosure prior to the filing date of this application. Nothing in this regard should be construed as the inventors acknowledging that they do not have prior rights to such disclosure by prior invention or prior publication, or for any other reason. All statements or expressions relating to the dates or contents of these documents are based on information available to the applicant and do not constitute any endorsement of the accuracy of the dates or contents of these documents.

[0122] Those skilled in the art will readily understand that the present invention is well suited to performing its objectives and obtaining the objectives and benefits mentioned, as well as those specific to them. The details of the description and examples herein are representative and illustrative of specific embodiments and do not limit the scope of the invention. Modifications and other uses therein will be conceivable to those skilled in the art. These modifications are included within the spirit of the invention. It will be readily apparent to those skilled in the art that various substitutions and modifications can be made to the inventions disclosed herein without departing from the scope and spirit of the invention.

[0123] The articles “a” and “an” used herein and in the claims should be understood to include multiple referents unless explicitly stated otherwise. Any claim or statement containing “or” between one or more members of a group is considered satisfied unless otherwise stated or the context makes otherwise clear, if one, more than one, or all of the group members are present, used, or otherwise related to a given product or process. The present invention includes embodiments in which exactly one member of a group is present, used, or otherwise related to a given product or process. The present invention also includes embodiments in which more than one or all of the group members are present, used, or otherwise related to a given product or process. Furthermore, it should be understood that the present invention provides all variations, combinations, and substitutions in which one or more limitations, elements, clauses, descriptive terms, etc., from one or more of the enumerated claims are introduced into another claim dependent on the same basic claim (or any other related claim). All embodiments described herein are intended to be applicable to all different aspects of the Invention where appropriate. Any embodiment or aspect may be freely combined with one or more other such embodiments or aspects as needed. Where elements are presented as a list, for example, in the form of a Markush group or similar, each subgroup of the elements is also disclosed, and it should be understood that any element(s) may be removed from the group. In general, where the Invention or an aspect of the Invention is referred to as including certain elements, features, etc., it should be understood that a particular embodiment of the Invention or an aspect of the Invention consists of or is essentially such elements, features, etc. For simplicity, these embodiments are not, in all cases, specifically described herein in so many words.It should also be understood that any embodiment or aspect of the present invention may be expressly excluded from the claims, regardless of whether certain exclusions are described herein. For example, any one or more activators, additives, components, any selection of agents, types of organisms, disorders, subjects, or combinations thereof may be excluded.

[0124] Where the claims or description relate to a composition of a substance, a method of producing or using a composition of a substance according to any of the methods disclosed herein, and a method of using a composition of a substance for any of the purposes disclosed herein, should be understood as an embodiment of the invention unless otherwise indicated, or unless it is obvious to a person skilled in the art that this would result in a contradiction or inconsistency. Where the claims or description relate to a method, for example, a composition useful for carrying out that method and a method of producing a product manufactured according to that method, should be understood as an embodiment of the invention unless otherwise indicated, or unless it is obvious to a person skilled in the art that this would result in a contradiction or inconsistency.

[0125] Where a range is given herein, the present invention includes embodiments in which an endpoint is included, embodiments in which both endpoints are excluded, and embodiments in which one endpoint is included but the other is excluded. Unless otherwise specified, it should be assumed that both endpoints are included. Furthermore, unless otherwise indicated by the context and the understanding of those skilled in the art, or unless otherwise evident, it should be understood that a value expressed as a range may be any particular value or subrange within the range described in different embodiments of the present invention, up to one-tenth of the lower limit unit of the range, unless the context clearly indicates otherwise. Also, where a set of numerical values ​​is described herein, it should be understood that the present invention includes embodiments similarly related to any intervening value or range defined by any two values ​​in the set, with the smallest value being the minimum and the largest value being the maximum. Numerical values ​​used herein include values ​​expressed as percentages. For any embodiment of the present invention preceded by “about” or “approximately”, the present invention includes embodiments in which exact values ​​are described. With respect to any embodiment of the present invention in which the words "about" or "approximately" are not placed before the numerical value, the present invention also includes embodiments in which the words "about" or "approximately" are placed before the value.

[0126] "Approximately" or "about" generally means, unless otherwise stated in the context or otherwise evident, a number that falls within 1% of the number in either direction, or within 5% in some embodiments, or within 10% in some embodiments (greater than or less than that number) (unless such number inevitably exceeds 100% of the possible value). Unless otherwise expressly indicated, in any method claimed herein that includes more than one act, the order of the acts of the method is not necessarily limited to the order in which the acts of the method are enumerated, but it should be understood that the present invention includes embodiments in which the order is thus limited. It should also be understood that, unless otherwise specifically indicated or evident in the context, any product or composition described herein may be considered "isolated". [Examples]

[0127] Examples

[0128] Introduction

[0129] Recent discoveries of genetic alterations involved in the onset and progression of human cancer have established a new generation of biomarkers. These alterations include single nucleotide substitutions, insertions, deletions, and translocations. These somatic mutations can also be detected in cell-free circulating tumor DNA (cfDNA) [6]. The development of non-invasive fluid biopsy methods based on ctDNA analysis offers opportunities for a new generation of diagnostic approaches. A recently developed blood test can detect eight common cancer types through the assessment of mutation levels in circulating proteins and cfDNA, with a sensitivity ranging from 69 to 98% and a specificity higher than 99% [7]. However, mutation-based fluid biopsy tests are less sensitive due to intratumoral and intertumoral heterogeneity [8], as not all samples of a single cancer type contain the same genetic driver alterations. For example, analysis of lung adenocarcinoma samples identified 22 drivers [9], but up to 25% of patients do not have genetic alterations in any of these genes [10,11]. Furthermore, the presence of low-frequency subclones further complicates mutation-based diagnosis: in stage I disease, the fraction of cfDNA is approximately 0.1%

[12] , and therefore detecting subclonal mutations with a frequency of 5% in early-stage disease challenges the detection limits of current sequencing techniques

[13] .

[0130] In recent years, DNA methylation profiling has been adopted as a promising approach for liquid biopsies

[14] . Abnormal DNA methylation is ubiquitous in human cancers and has been shown to occur early in carcinogenesis, thus offering an attractive potential biomarker for early detection of cancer

[15] . Compared to normal genomes, cancer genomes are generally hypomethylated and locally hypermethylated in CpG islands (CGIs) [16,17]. Markers associated with these two features have been widely used for methylation-based ctDNA detection [18,19]. For example, FBN1, FBN2, HLTF, PHACTR3, SEPT9, SNCA, SST, TAC1, and VIM have been used individually for colorectal cancer (CRC) detection

[20] . However, single-gene-based diagnostics have low accuracy due to tumor heterogeneity. Therefore, genome-wide assays such as whole-genome bisulfite sequencing (WGBS) and reduced-expression bisulfite sequencing (RRBS) are being tested to improve predictive performance. For example, plasma hypomethylation provided 74% and 94% sensitivity and specificity, respectively, for detecting non-metastatic cancer cases when an average of 93 million WGBS reads per case were obtained

[18] . More recently, methylated DNA immunoprecipitation sequencing (MeDIP-seq), a genome-wide assay, has demonstrated high sensitivity for tumor detection and classification using plasma cell-free DNA methylomes

[21] . Regarding analytical methods, since CpG mean methylation-based methods are insufficiently sensitive for early cancer detection, methylated haplotype blocks (MHBs; i.e., comethylation stretches of DNA) have been used instead, which can detect 2% tumor DNA

[22] . This approach led to the development of CancerDetector, a novel methylated haplotype analysis tool that can detect 0.1% tumor DNA, as demonstrated by spike-in experiments

[23] . While genome-wide assays are promising in terms of both high sensitivity for early cancer detection and cancer type classification, they generally suffer from higher costs and longer turnaround times.Targeted assays that examine only a predetermined set of genomic regions represent a solution that balances the information gained with the cost. For example, padlock-based targeted sequencing

[24] has been evaluated for non-invasive detection of hepatocellular carcinoma (HCC) with 83.3% sensitivity and 90.5% specificity using only 10 markers

[25] . Detection of HCC is relatively easy compared to other cancer types, as up to 20% of cfDNA originates from liver tissue even in normal controls

[26] . Recently, a marker with four consecutive CpG sites was characterized in breast cancer by amplicon-based bisulfite sequencing, identifying a complete methylation pattern for early identification of metastasis

[27] . Although the sensitivity is low at 25%, this method represents a novel approach for the collaborative analysis of multiple CpG sites at a single locus. Published studies using targeted sequencing have primarily addressed the detection of single cancer types, and therefore, ultra-sensitive methods for non-invasive detection of multiple cancer types remain undeveloped. Epigenetic restrictions in extraembryonic lineages reflect somatic migration to cancer

[28] . Extraembryonic methylation signatures were found to distinguish cancer samples from matching normal tissue for nearly all cancer types tested. Based on these findings, extraembryonic signatures, combined with DNA methylation haplotype analysis, represent a universal framework for highly sensitive, non-invasive early cancer diagnosis.

[0131] result

[0132] Extraembryonic hypermethylation CGI provides a universal cancer signature.

[0133] The placenta has long been considered a pseudo-malignant tumor tissue with several phenotypes reminiscent of human cancer, such as its angiogenic, immunosuppressive, and invasive capabilities

[29] . The DNA methylation landscape of the extraembryonic ectoderm (ExE), the progenitor of the placenta, was compared to the DNA methylation landscape of the epiblast of mouse E6.5 conception product

[28] (Figure 1A). Using this data, we identified ExE hypermethylated CGI (ExE hyper-CGI) as a DNA methylation signature that can distinguish these two tissue types. Interestingly, ExE hyper-CGI is more conserved at the sequence level than the genomic background (Figure 1B), and the majority of mouse ExE hyper-CGI has human orthologues localized near CGI (Figure 1C). Surprisingly, the ExE hyper-CGI signature was found to be hypermethylated in 14 cancer types profiled within the Cancer Genome Atlas (TCGA) project, including matched normal tissues

[28] . The only exception was thyroid cancer, which may be explained by the observation that FGF and WNT pathways are shared during histological identification of ExE and normal thyroid epithelium

[30] . Next, we tested the performance of ExE hyperCGI in cancer prediction using the TCGA pan-cancer dataset. When TCGA samples were randomly assigned to training and validation sets, ExE hyperCGI was able to classify tumor samples versus normal samples with high sensitivity and specificity using a support vector machine (SVM) classification method (Methods, AUC=0.98, Figure 1D). Similar results were obtained when a random forest of an independent method was applied to the same dataset (AUC=0.98, Methods and Figure 6). This observation suggests that when using ExE hyperCGI, the majority of cases of each tumor type can be accurately identified, and that human cancer types are significantly more homogeneous when analyzed for the methylation status of ExE hyperCGI than when profiling for the mutation status of any driver gene (Figure 1E).For example, somatic mutations in TP53 represent the most frequent genetic alteration in human cancers, while many cancer types, such as renal papillary cell carcinoma (KIRP) and renal clear cell carcinoma (KIRC), exhibit a low mutation frequency in TP53 (Figure 1E). Therefore, ExE hyperCGI represents a foundation for developing novel DNA methylation signatures for pan-cancer diagnosis and this non-invasive liquid biopsy platform (f).

[0134] DNA methylation haplotypes improve detection sensitivity.

[0135] The development of non-invasive fluid biopsy methods based on ctDNA DNA methylation has revolutionized cancer diagnosis.

[21] However, several challenges remain. First, disordered methylation is frequently observed in cancer.

[31] This is one reason why single CpG-based diagnostic platforms suffer from low sensitivity. For example, the overall sensitivity of SEPT9 is only 60% for colorectal cancer (CRC) detection.

[32] Second, the fraction of ctDNA in cell-free DNA is as low as 0.01% in early-stage disease.

[33] Therefore, background by normal cells must be nearly zero to enable the detection of tumor cells. However, normal cells acquire low levels of methylation (about 1%) when measured at a single CpG site due to noise, aging.

[34] and other stochastic processes.

[35] To overcome these problems, a novel approach has been developed based on the observation that DNA methylation haplotypes measured stepwise on the same molecule offer a better selection for diagnostic purposes. Even when measured from bulk data, DNA methylation information obtained from a single sequencing fragment is guaranteed to originate from a single chromosome and a single cell. Therefore, the CpG methylation pattern of each fragment represents an individual DNA methylation haplotype (Figure 2A). In normal somatic cell tissues, fully methylated reads are extremely rare when analyzing ExE hyper-CGI. Therefore, the percentage of fully methylated reads (PMRs) calculated from sequencing data represents a novel method for quantifying the degree of DNA methylation (Figures 7 and 8). This method significantly reduces background noise compared to standard methods. For example, OTX2 is a developmental regulator that is hypermethylated in ExE and the placenta and also acts as one of the ExE hyper-CGI markers. When its mean methylation level was used, a considerable amount of background noise was observed in normal samples. In contrast, PMR-based quantification at this locus significantly reduced background noise (Figure 2B).

[0136] To evaluate the performance of PMR, silico simulations were performed by randomly sampling sequencing reads from normal-like tissue epiblasts and tumor-like tissue ExE as spike-ins. The spike-in fractions ranged from 0.01% to 1%, which coincided with the fraction of ctDNA in cell-free DNA (Methods). In addition to mean methylation and PMR, DNA methylation haplotype loading (MHL)

[22] to quantify the level of comethylation was also included for comparison (Figures 9, 10, and 11). Using this approach, all three methods had significant predictive power in both the 1% and 0.1% spike-in groups. However, when the spike-in fraction decreased to 0.01%, only PMR-based predictions became significant when the mean spike-in coverage was greater than 5-fold (Figure 2C). Note that PMR is a k-mer-based approach, and when tested in the simulated 0.01% spike-in group, the highest sensitivity was achieved at k = 5 (Figure 12).

[0137] Efficient workflow for enriching DNA methylation haplotypes

[0138] Several recent studies have profiled cell-free DNA using one of the following approaches: reduced-expression bisulfite sequencing (RRBS)

[22] , whole-genome bisulfite sequencing (WGBS)

[23] , or methylated DNA immunoprecipitation sequencing (MeDIP-seq)

[21] , all of which suffer from insufficient coverage in the region of interest in exchange for the availability of genome-wide information. Instead of these approaches, we used targeted bisulfite sequencing (targeted BS). This is because this assay generates data with stronger signals from the region of interest, which is associated with lower cost compared to the other methods. For this purpose, a highly specific target capture pipeline was established using the SeqCap Epi technology

[36] , which can enrich ExE hyper-CGI (total 6.2 Mb; method) with an on-target rate of approximately 80%. Given the small fraction of tumor-derived DNA in plasma, most sequencing reads obtained from plasma samples originate from normal DNA that is largely unmethylated in the target region. Tumor-derived DNA was analyzed by further specific enrichment of methylated DNA fragments using the MBD2 protein, followed by targeted BS (Figure 3A). The customized probe set performed similarly to commercially available probe sets in terms of enrichment uniformity. Specifically, 80% of the loci had higher coverage than the 60% with central coverage (Figure 3B). When tested on biopsy samples from both tumor and normal tissues, the targeted BS approach achieved enrichment of over 400-fold compared to WGBS. Even with challenging samples such as cell-free DNA, enrichment of over 100-fold was observed (Figure 3C). When this workflow was combined with MBD enrichment before bisulfite conversion, high specificity was achieved, with over 90% of reads partially or completely methylated on average (Figure 3D).

[0139] Unbiased measurement of DNA methylation across assays

[0140] By definition, PMR is the number of fully methylated k-mer haplotypes divided by the total number of k-mers in each genomic feature, such as a CpG island, and was set to 5 to maximize sensitivity (Figure 12). Similarly, MHL is a normalized PMR with different k-mer lengths (method, k=1 to 10). Thus, although both PMR and MHL are locally normalized haplotype-based methods, neither could be applied without bias between assays, and when the same sample was profiled by targeted BS with or without MBD enrichment, neither PMR nor MHL were equivalent between these two assays (Figures 13 and 14). An alternative to global normalization is to normalize the number of haplotypes within a region by the total number of haplotypes across the entire region. For a given haplotype width k (i.e., k=5), the globally normalized coverage of each type of DNA methylated haplotype was compared for the same sample profiled by both assays with and without MBD enrichment. This approach was used to profile two cell lines (HuES64 and HCT116) and two primary tissues (normal uterine and uterine cancer). The highest Pearson correlation coefficient (PCC) was observed between these two approaches when using the number of fully methylated DNA methylation haplotypes (mean PCC = 0.998) (Figure 4A). For example, when normalized coverage (NMR) of fully methylated reads was evaluated for normal uterine and uterine cancer, nearly perfect correlation was observed between assays with and without MBD enrichment (PCC > 0.99, p < 10). -16 (Figures 4B and 15). As expected, unbiased measurements were observed when comparing targeted BS and WGBS, but greater variability was observed in WGBS assay samples due to the lower sequencing depth (PCC = 0.958 for uterine cancer, PCC = 0.979 for normal uterus, p-value < 10). -16(Figure 4C). In summary, NMR is an unbiased metric for quantifying haplotype-level DNA methylation across WGBS and targeted BS approaches, with or without MBD enrichment. Methodological improvements allowed for the development of markers from existing data and their validation with new data.

[0141] Ultra-sensitive cancer detection using DNA methylation haplotypes

[0142] Because ctDNA levels are very low in most early-stage and many advanced-stage cancer patients [6], the main challenge is how to identify the trace amounts of ctDNA within total cfDNA. To test the sensitivity of an MBD enrichment-based workflow, we initially performed experiments mixing DNA from ES cells (HuES64) with DNA from a colon cancer cell line (HCT116) as the spike-in. An NMR-based method reliably predicted a 0.01% spike-in when using at least 1 μg of total input DNA (Figure 16A). However, when analyzing 50 ng of total input DNA, the prediction limit dropped to 0.1% (Figure 16B). Novel analytical techniques such as NMR can improve sensitivity to targeted BS data even without MBD enrichment, which works well with lower input DNA. When the targeted BS workflow was tested without MBD enrichment using 50 ng of DNA as input, conditions with a 0.01% spike-in were correctly identified with only 50 CGIs (Figures 5A and 17). In contrast, mean methylation and MHL-based methods were only able to accurately identify tumor signatures when the fraction of spike-in DNA was greater than 0.1% (Figure 20A). Detection of HCT116 DNA was easier than detection of other samples because its genome is almost completely methylated, and similar dilution experiments were then performed using primary colon cancer tissue as the spike-in. Here again, the NMR-based method reliably detected 0.01% of spike-in cancer DNA (Figures 5B and 18), while the mean methylation and MHL-based methods detected only 1% of cancer DNA spike-in (Figure 20B). Furthermore, detection sensitivity depends on background noise originating from normal cells. For example, when uterine cancer DNA was spike-in along with normal uterine DNA, the NMR-based method was able to detect 0.1% of cancer DNA (Figure 5C), while both the mean methylation and MHL-based methods detected only 1% of cancer DNA (Figure 20C). Detection sensitivity also depends on the selection of parameters. For example, in NMR spectroscopy, the highest sensitivity was obtained when the k-mer length was set to 5 (Figure 19).

[0143] Finally, experimental and computational pipelines were tested on plasma samples obtained from colon adenocarcinoma patients, using age-matched normal individuals as negative controls. Each cohort included two samples from patients with stages I, II, and III cancer. The platform was able to detect all cancers, including stage I cancer, with high reliability (FDR < 1%), and no false positives were observed (Table 1A). To further evaluate the sensitivity of this method, the fraction of reads predicted to originate from tumor cells was estimated. In the colon cancer cohort, the estimated fraction of cancer DNA ranged from 0.05% to 20% (Methods; Figure 21), suggesting a predictive resolution of 0.05% for colon cancer. Next, a breast cancer patient cohort (invasive ductal carcinoma) was tested, including two cases each for stages I, II, and III. The NMR-based method detected five out of six cancer samples, with one stage II sample being false negative, which was CDX171 (FDR < 1%, Table 1B). However, the mean methylation and MHL-based methods each accurately identified only one sample. The estimated tumor fraction of CDX171 was approximately 0.03%, which is similar to background noise, suggesting that the false negatives were likely due to a low tumor DNA fraction (Methods and Figure 22).

[0144] Machine learning methods

[0145] We developed a wide range of predictive models using machine learning approaches (random forests, support vector machines, and deep learning) to estimate the total probability distribution of all haplotypes across ExE sites for each tumor type. These methods will improve the accuracy of predicting cell type origin based on cfDNA samples.

[0146] Table 3 shows the pan-cancer-related methylation sites.

[0147] Consideration

[0148] DNA methylation haplotypes have been used for many years, but only recently have they been shown to be useful in cancer diagnosis. For example, Guo et al. demonstrated MHL, a DNA methylation haplotype-based metric combined with methylation haplotype block (MHB). An experimental and computational framework for ultra-sensitive non-invasive early cancer detection using fully methylated DNA methylation haplotypes was proposed. As demonstrated by dilution experiments, this framework outperformed mean methylation and MHL-based methods, and was able to detect 0.01% of colon cancer spike-in with just 50 CGIs. When tested with human plasma samples, both colon cancer and breast cancer samples were correctly detected at early stages, with a detection limit of 0.05%. This threshold has sufficient sensitivity to detect most stage I tumors. This is the first study to utilize a universal cancer signature for non-invasive pan-cancer diagnosis that is potentially more cost-effective compared to genome-wide assays

[21] .

[0149] cohort

[0150] Tumor and normal samples from 12 cancer types were included, with the exception of bladder and prostate cancer, which contained only normal samples, as described below. For cancer types, different major subtypes characterized by invasive breast cancer were included where possible. All samples were uniformly processed at the Broad Institute and profiled by targeted bisulfite sequencing using a customized probe design covering the 8M genomic region primarily hypermethylated in human cancers. [Table 4]

[0151] Origin organization

[0152] A highly sensitive method was developed based on DNA methylation haplotypes of extraembryonic methylated CpG islands. This method was able to detect 0.05% of tumor DNA from cell-free DNA in patient plasma. To further develop this method and predict the tissue of origin with high sensitivity, the method includes identifying cancer-specific DNA methylation haplotypes. For each CpG position in the designed region, the relative abundance of all possible k-mer haplotypes (k=5) was calculated across all tissue samples, including tumor and normal samples. A tissue-specific index (TSI) was then defined for each k-mer as follows:

[0153]

number

[0154] If n represents the number of tissues, PKR(j) represents the fraction of a specific k-mer in tissue j, and PKR max represents the PKR of the tissue with the highest methylation. Cancer-specific DNA methylation haplotypes were selected by TSI with a cutoff of 0.6. By adding cancer-specific DNA methylation haplotypes to the original signature, highly sensitive prediction of the origin tissue becomes possible.

[0155] Table 2 provides the identified regions of cancer-specific DNA methylation.

[0156] method

[0157] Targeted BS and MBD enrichment

[0158] Genomic DNA was extracted from cultured cells using the Genomic DNA Clean & Concentrator Kit (Zymo Research). Human tumor DNA was purchased from OriGene Technologies or BioChain Institute. Genomic DNA was sheared to an average fragment size of 180–220 bp in a 130 μl microtube using an S2 focused sonication system (Covaris) at an intensity of 5 per burst, a duty cycle of 10 and 200 cycles for 300 seconds. The sheared DNA was enriched with 1.8 volumes of Agencourt AMPure XP beads (Beckman Coulter) before bisulfite conversion. Purified human cell-free DNA and frozen human plasma from cancer patients were obtained from BioChain Institute. Free circulating DNA was isolated from 4 ml of human plasma using the QIAamp MinElute ccfDNA Mini Kit (Qiagen), which scales up the reaction as described in the manufacturer's manual. To enrich methylated DNA, selected samples were treated with the MethylMiner Methylated DNA Enrichment Kit (Thermo Fisher Scientific). DNA bound to the MBD2 protein coupled to streptavidin beads was eluted in a single elution step with the provided high-salt buffer, and the DNA was precipitated with ethanol. The pellet was dissolved in 20 μl of water. Shear genomic DNA, cfDNA, and MBD-enriched DNA were bisulfite-converted using the EpiTect Fast Bisulfite Conversion Kit (Qiagen) with two 60°C cycles extended to 20 minutes each, according to the kit instructions. Illumina library construction was performed after bisulfite conversion using the Accel-NGS Methyl-Seq Kit (Swift Biosciences), according to the manufacturer's recommendations for NimbleGen SepCap Epi Hybridization Capture (Appendix Section A).The libraries were amplified by 8–14 cycles of PCR using Accel-NGS Methyl-Seq Unique Dual Indexing primers (Swift Biosciences). The SeqCap Epi hybridization reaction involved a pool of 3–4 PCR-amplified pre-capture libraries totaling 1 μg, 2 μl of xGen Universal BlockersTS Mix (Integrated DNA Technologies) blocking oligonucleotides, and a custom SeqCap probe pool. After hybridization at 47°C (typically about 70 hours), streptavidin pull-down, and washing, the entire bead-bound capture material was amplified by 9–10 cycles of PCR. The hybrid-selected libraries were sequenced on an Illumina HiSeq 2500 instrument in high-speed mode, along with a 10% spike-in of an unindexed PhiX174 library.

[0159] Targeted BS probe set design

[0160] For targeted bisulfite sequencing, 1,265 CGIs hypermethylated in extraembryonic tissues were selected

[28] . Specifically, 473 CGIs were hypermethylated in the mouse ectoderm and liftover to the human genome. The remainder were hypermethylated in 8 of 14 TCGA cancer types and also in the human placenta. To cover loci with multiple hypermethylated CGIs, such as the OTX2 locus, 20kbp-separated CGIs were merged. The resulting regions were extended 2k upstream and 2k downstream, respectively, to cover the CpG Shore. The probe was designed by NimbleDesign using default parameters (design.nimblegen.com). The resulting design covers 6.1Mbp with estimated coverage of 98.2%.

[0161] Data processing

[0162] Raw sequencing reads were preprocessed with "trim_galore(v0.4.4)" using the following parameters: "--clip_R1 5--three_prime_clip_R1 2--clip_R2 10--three_prime_clip_R2 2". Low-quality base calls and adapters were truncated from the 3' end of the reads by default. Trimmed reads were aligned to the human reference genome GRCh37 using Bismark(v 0.19.0)

[37] with default parameters. Duplicate reads were identified and removed using Bismark's tools. DNA methylation haplotypes were extracted using an in-house tool called mHaplotype(github.com / JiantaoShi / mHaplotype). Reads with methylated cytosines in non-CpG contexts (CHG, CHH) were removed to eliminate potential bias caused by incomplete bisulfite conversion.

[0163] In silico simulation

[0164] ExE and epiblast represent typical tumor-like and normal-like genomes, respectively, with respect to the DNA methylation landscape. To evaluate the performance of different cancer prediction methods, in silico simulations were performed by randomly sampling sequencing reads from ExE and epiblast samples. Briefly, ExE and epiblast RRBS data were obtained from the publicly available dataset GSE98963, which contains four biological replicates for each tissue. DNA methylation haplotypes were extracted using the in-house tool mHaplotype, and biological replicates were pooled. Sequencing reads were randomly sampled as spike-ins from epiblast and ExE, representing 1%, 0.1%, and 0.01% of the total reads in the three simulation groups, respectively. In each group, the mean coverage of spike-in DNA ranged from 1 to 20, with 10 replicates each. Spike-in reads were sampled from epiblast, including negative controls.

[0165] Estimation of methylation levels

[0166] The mean methylation level was estimated as the number of C-reporting sites, obtained by dividing the mean methylation level by the total number of C or T-reporting sites. The methylation pattern of CpGs in each fragment represents an individual DNA methylation haplotype. The methylation haplotype loading (MHL), which is a normalized fraction of methylation haplotypes of various lengths, was calculated as previously described

[22] :

number

[0167] Here, k is the length of the haplotype, and for a haplotype of length L, all substrings with lengths from 1 to a maximum of 10 were considered in this calculation. k This represents the weight of the k-mer haplotype. In this study, w k =k was applied. PMR k k is the fraction of fully continuous methylated CpGs for a haplotype of length k (k-mer) (Figure 8). In this study, k was set to 5 to maximize detection sensitivity (Figure 12). To calculate the normalized coverage of fully methylated reads (NMR), the number of fully methylated k-mers was determined for each CGI, then divided by the total number of fully methylated k-mers in all designed regions, and then averaged. Here again, k was set to 5 to maximize detection sensitivity (Figure 19).

[0168] Prediction of the presence of cancer DNA

[0169] The presence of cancer-specific DNA methylation suggests the presence of cancer DNA in the mixture. As described above, four metrics—mean methylation, MHL, PMR, and NMR—were used for DNA methylation quantification and cancer prediction. Four types of samples—tumor tissue samples, normal tissue samples, normal cfDNA samples, and patient cfDNA samples—were used for prediction. For a given CGI, DNA methylation in these groups was measured using Me, respectively. (t) Me(n) Me (f) Me (p) This was expressed as follows. Regardless of the metrics used, the general process for cancer prediction is very similar.

[0170] Marker identification

[0171] ExE HyperCGI is largely hypermethylated in cancer compared to normal tissue. Markers were redefined for each cancer type and metric used to maximize detection sensitivity. Specifically, tumor tissue samples were compared to normal tissue samples with a threshold of 0.1 (Me). (t) -Me (n) In >0.1), we defined markers that are hypermethylated in tumors.

[0172] Marker improvements

[0173] Next, the selected markers were ranked in descending order based on the difference in methylation (Me(t)-Me(f)) between tumor samples and normal cfDNA. The top 200 regions were selected as cancer prediction markers.

[0174] Significance Test

[0175] The test sample was compared to a normal cfDNA sample using the cancer marker defined above, and the resulting methylation difference was calculated as ΔMe = Me (p) -Me (f) This was defined as follows. Instead of using the actual values ​​of the methylation difference, the number of markers with increased methylation (ΔMe>0) and markers with decreased methylation (ΔMe<0) was counted. A higher number of markers with increased methylation indicates a higher likelihood of detecting cancer in the sample. The p-value was calculated using a one-sided binomial test and corrected for multiple tests using the Benjamini-Hochberg procedure.

[0176] Prediction of tumor DNA fraction

[0177] The fraction of tumor DNA was predicted by comparing the observed data with simulated normal cfDNA data that had tumor DNA as a spike-in, and the fraction ranged from 0.01% to 100%. NMR was performed on the observed samples (NMP) using predefined markers for each cancer type. o ) and simulated sample (NMP s ) is compared with ΔNMR = NMR s -NMR o This was shown. Next, the distance metric was calculated as follows.

number

[0178] The predicted tumor fraction was defined as the value that minimizes the distance d.

[0179] Cancer prediction using TCGA 450K array data

[0180] To evaluate the performance of ExE HyperCGI in cancer prediction, 14 TCGA cancer types containing matched normal tissue in TCGA were tested. Since thyroid cancer and normal thyroid tissue cannot be distinguished by ExE HyperCGI, samples from the thyroid cancer dataset were removed

[28] . This pan-cancer cohort consisted of 685 tumor samples and 710 normal samples.

[0181] Half of the samples were randomly selected as the training set, and the remainder were used for validation. A support vector machine (SVM) with a Gaussian kernel from the R package kernlab was used for classification. To resolve dependencies between ExE hyperCGIs, 50 CGIs were randomly selected for classification, and this process was repeated 200 times, with the resulting prediction scores averaged as the final concentration scores. Receiver operating characteristic (ROC) curves were generated using the R package ROCR.

[0182] Similarly, Random Forest (RF) was implemented using the `randomForest` function from the `randomForest` R package, with default parameter settings. Classification accuracy was calculated as the percentage of samples in the validation set that the trained model correctly classified. False positive and true positive rates were calculated using the `roc` function from the `pROC` R package, based on `out-of-bag` votes for the training data. The area under the ROC curve (AUC) was calculated based on these values ​​using the `auc` function from the `pROC` package.

[0183] Data availability

[0184] All datasets are deposited in the Gene Expression Omnibus and are accessible under GSE84236. Additional data includes TCGA DNA methylation, mutation data, and full tumor type names from Broad Firehose (gdac.broadinstitute.org). [Table 5]

[0185] References

[0186] 1. McGranahan, N. and C. Swanton, Clonal Heterogeneity and Tumor Evolution: Past, Present, and the Future. Cell, 2017. 168(4): p. 613-628.

[0187] 2. Winawer, SJ, et al., Prevention of colorectal cancer by colonoscopic polypectomy. The National Polyp Study Workgroup. N Engl J Med, 1993. 329(27): p. 1977-81.

[0188] 3. Karam, A.K. and B.Y. Karlan, Ovarian cancer: the duplicity of CA125 measurement. Nat Rev Clin Oncol, 2010. 7(6): p. 335-9.

[0189] 4. Gao, Y., et al., Evaluation of Serum CEA, CA19-9, CA72-4, CA125 and Ferritin as Diagnostic Markers and Factors of Clinical Parameters for Colorectal Cancer. Sci Rep, 2018. 8(1): p. 2732.

[0190] 5. Nordstrom, T., et al., Prostate-specific antigen (PSA) density in the diagnostic algorithm of prostate cancer. Prostate Cancer Prostatic Dis, 2018. 21(1): p. 57-63.

[0191] 6. Bettegowda, C., et al., Detection of circulating tumor DNA in early- and late-stage human malignancies. Sci Transl Med, 2014. 6(224): p. 224ra24.

[0192] 7. Cohen, J.D., et al., Detection and localization of surgically resectable cancers with a multi-analyte blood test. Science, 2018.

[0193] 8. Yates, L.R. and P.J. Campbell, Evolution of the cancer genome. Nat Rev Genet, 2012.

[0194] 13(11): p. 795-806.

[0195] 9. Lawrence, M.S., et al., Discovery and saturation analysis of cancer genes across 21 tumour types. Nature, 2014. 505(7484): p. 495-501.

[0196] 10. Pao, W. and K.E. Hutchinson, Chipping away at the lung cancer genome. Nat Med, 2012. 18(3): p. 349-51.

[0197] 11. Cancer Genome Atlas Research, N., Comprehensive molecular profiling of lung adenocarcinoma. Nature, 2014. 511(7511): p. 543-50.

[0198] 12. Phallen, J., et al., Direct detection of early-stage cancers using circulating tumor DNA. Sci Transl Med, 2017. 9(403).

[0199] 13. Corcoran, R.B. and B.A. Chabner, Application of Cell-free DNA Analysis to Cancer Treatment. N Engl J Med, 2018. 379(18): p. 1754-1765.

[0200] 14. Laird, P.W., The power and the promise of DNA methylation markers. Nat Rev Cancer, 2003. 3(4): p. 253-66.

[0201] 15. Baylin, S.B., et al., Aberrant patterns of DNA methylation, chromatin formation and gene expression in cancer. Hum Mol Genet, 2001. 10(7): p. 687-92.

[0202] 16. Berman, B.P., et al., Regions of focal DNA hypermethylation and long-range hypomethylation in colorectal cancer coincide with nuclear lamina-associated domains. Nat Genet, 2011. 44(1): p. 40-6.

[0203] 17. Zhou, W., et al., DNA methylation loss in late-replicating domains is linked to mitotic cell division. Nat Genet, 2018. 50(4): p. 591-602.

[0204] 18. Chan, K.C., et al., Noninvasive detection of cancer-associated genome-wide hypomethylation and copy number aberrations by plasma DNA bisulfite sequencing. Proc Natl Acad Sci U S A, 2013. 110(47): p. 18761-8.

[0205] 19. Kang, S., et al., CancerLocator: non-invasive cancer diagnosis and tissue-of-origin prediction using methylation profiles of cell-free DNA. Genome Biol, 2017. 18(1): p. 53.

[0206] 20. Leygo, C., et al., DNA Methylation as a Noninvasive Epigenetic Biomarker for the Detection of Cancer. Dis Markers, 2017. 2017: p. 3726595.

[0207] 21. Shen, S.Y., et al., Sensitive tumour detection and classification using plasma cell-free DNA methylomes. Nature, 2018.

[0208] 22. Guo, S., et al., Identification of methylation haplotype blocks aids in deconvolution of heterogeneous tissue samples and tumor tissue-of-origin mapping from plasma DNA. Nat Genet, 2017. 49(4): p. 635-642.

[0209] 23. Li, W., et al., CancerDetector: ultrasensitive and non-invasive cancer detection at the resolution of individual reads using cell-free DNA methylation sequencing data. Nucleic Acids Res, 2018.

[0210] 24. Diep, D., et al., Library-free methylation sequencing with bisulfite padlock probes. Nat Methods, 2012. 9(3): p. 270-2.

[0211] 25. Xu, R.H., et al., Circulating tumour DNA methylation markers for diagnosis and prognosis of hepatocellular carcinoma. Nat Mater, 2017. 16(11): p. 1155-1161.

[0212] 26. Sun, K., et al., Plasma DNA tissue mapping by genome-wide methylation sequencing for noninvasive prenatal, cancer, and transplantation assessments. Proc Natl Acad Sci U S A, 2015. 112(40): p. E5503-12.

[0213] 27. Widschwendter, M., et al., Methylation patterns in serum DNA for early identification of disseminated breast cancer. Genome Med, 2017. 9(1): p. 115.

[0214] 28. Smith, Z.D., et al., Epigenetic restriction of extraembryonic lineages mirrors the somatic transition to cancer. Nature, 2017. 549(7673): p. 543-547.

[0215] 29. Novakovic, B. and R. Saffery, Placental pseudo-malignancy from a DNA methylation perspective: unanswered questions and future directions. Front Genet, 2013. 4: p. 285.

[0216] 30. Kurmann, A.A., et al., Regeneration of Thyroid Function by Transplantation of Differentiated Pluripotent Stem Cells. Cell Stem Cell, 2015. 17(5): p. 527-42.

[0217] 31. Landau, D.A., et al., Locally disordered methylation forms the basis of intratumor methylome variation in chronic lymphocytic leukemia. Cancer Cell, 2014. 26(6): p. 813- 825.

[0218] 32. Nian, J., et al., Diagnostic Accuracy of Methylated SEPT9 for Blood-based Colorectal Cancer Detection: A Systematic Review and Meta-Analysis. Clin Transl Gastroenterol, 2017. 8(1): p. e216.

[0219] 33. Aravanis, A.M., M. Lee, and R.D. Klausner, Next-Generation Sequencing of Circulating Tumor DNA for Early Cancer Detection. Cell, 2017. 168(4): p. 571-574.

[0220] 34. Gentilini, D., et al., Stochastic epigenetic mutations (DNA methylation) increase exponentially in human aging and correlate with X chromosome inactivation 女性中的偏斜。《衰老》(纽约奥尔巴尼),2015年。7(8):第568 - 578页。

[0221] 35. 瓦尔贝里,P.等人,急性淋巴细胞白血病细胞的DNA甲基化组 分析揭示了CpG岛中随机的从头DNA甲基化。《表观基因组学》,2016年。8(10):第1367 - 1387页。

[0222] 36. 李,Q.等人,哺乳动物和植物基因组中修饰胞嘧啶的转换后靶向捕获。《核酸研究》,2015年。43(12):第e81页。

[0223] 37. 克鲁格,F.和S.R.安德鲁斯,Bismark:一种用于亚硫酸氢盐测序应用的灵活比对器和甲基化调用器。《生物信息学》,2011年。27(11):第1571 - 1572页。

[0224]

Table 1A

[0225]

Table 1B

[0226]

Table 2 - 1

Table 2 - 2

[0227] Table 3-1 Table 3-2 Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11

Claims

1. A method for characterizing cell-free DNA (cfDNA) samples derived from a target, a) A step of subjecting the cfDNA sample to whole-genome bisulfite sequencing, reductive bisulfite sequencing, or targeted bisulfite sequencing to prepare sequencing data including methylated sequence reads for the genomic sequence from the cfDNA sample, wherein the genomic sequence includes one or more CpG islands (CGIs) selected from Table 2 or Table 3 that are methylated in the extraembryonic ectoderm (ExE) genome and not methylated in the corresponding epiblast or adult tissue. b) A step of analyzing sequencing data including methylated sequence reads for the genome sequence containing one or more CGIs selected from Table 2 or Table 3, and c) A method comprising the step of classifying the cfDNA sample using the analyzed sequencing data.

2. The method according to claim 1, wherein the sequencing data includes sequence information for less than 0.3% of the target genome.

3. The method according to claim 1, wherein the sequencing data includes sequence information substantially limited to one or more regions of the target genome having multiple CGIs that are methylated in the genome of the ExE and not methylated in the corresponding epiblast or adult tissue.

4. The method according to claim 1, wherein the sequencing data, which includes reads of methylated sequences for the genome sequence from the cfDNA sample, is enriched with respect to sequences containing methylation.

5. The method according to claim 4, wherein the concentration comprises an MBD2 protein-based concentration method.

6. The genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple methylated CpG islands (CGIs) in the extraembryonic ectoderm (ExE) genome and / or one or more regions identified in Table 3, or The genome sequence includes 50 to 75 methylated CGIs in the genome of ExE, The method according to claim 1, wherein "approximately" means ±10% of the following number.

7. The method according to claim 1, wherein the cfDNA is classified as originating from a tumor.

8. The method according to claim 1, wherein the cfDNA is classified as originating from one or more tumor types or tissues.

9. The method according to claim 8, wherein the one or more tumor types include one or more of acute myeloid leukemia, bladder cancer, breast cancer, colon cancer, esophageal cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, or stomach cancer.

10. The method according to claim 1, wherein the cfDNA is classified as comprising fully methylated reads.

11. A method for characterizing cell-free DNA (cfDNA) samples derived from a target, a) A step of preparing sequencing data including methylated sequence reads for the genome sequence from the cfDNA sample by subjecting a cfDNA sample from the subject to whole-genome bisulfite sequencing, reductive bisulfite sequencing, or targeted bisulfite sequencing, wherein the genome sequence includes a plurality of CpG islands (CGIs) that are methylated in the extraembryonic ectoderm (ExE) genome but not in the corresponding epiblast or adult tissue. b) A step of receiving the sequencing data, which includes reads of methylated sequences for the genome sequence from the cfDNA sample. c) A step of applying a machine learning method to estimate the probability distribution of haplotypes across multiple CGIs in the genome of ExE, and d) A step of determining the classification of tumor versus normal based on the predicted score obtained from the machine learning method, Methods that include...

12. The method according to claim 11, wherein the machine learning method is a random forest, deep learning, or a support vector machine.

13. A method for characterizing cell-free DNA (cfDNA) samples derived from a target, a) A step of preparing sequencing data including methylation sequence reads for the genome sequence from the cfDNA sample obtained from the subject after cancer treatment by subjecting it to whole-genome bisulfite sequencing, reduced expression bisulfite sequencing, or targeted bisulfite sequencing. b) A step of receiving the sequencing data, which includes reads of methylated sequences for the genome sequence from the cfDNA sample. c) A step of determining the amount of haplotypes containing five methylated CpG sites, and d) If the amount of the haplotype is greater than the significance threshold, the cfDNA sample is characterized as containing fully methylated cfDNA. Methods that include...

14. A step of detecting cancer in the subject when the cfDNA sample is characterized as containing fully methylated cfDNA, or a step of detecting the eradication of cancer in the subject when the cfDNA sample is characterized as not containing fully methylated cfDNA. The method according to claim 13, further comprising:

15. The method according to claim 13, wherein the sequencing data includes sequence information for less than 0.3% of the target genome.

16. The method according to claim 13, wherein the sequencing data includes sequence information substantially limited to one or more regions of the target genome having multiple CGIs that are methylated in the genome of the ExE and not methylated in the corresponding epiblast or adult tissue.

17. The method according to claim 13, wherein the sequencing data, which includes reads of methylated sequences for the genomic sequence from the cfDNA sample, is enriched with respect to the methylated sequences, and optionally the enrichment includes an MBD2 protein-based enrichment method.

18. The method according to claim 13, wherein the five CpG sites are five consecutive CpG sites.

19. The genome sequence includes a continuous sequence of approximately 8 megabases of the human genome containing multiple methylated CpG islands (CGIs) in the extraembryonic ectoderm (ExE) genome and / or one or more regions identified in Table 3, or The genome sequence includes 50 to 75 methylated CGIs in the genome of ExE, The method according to claim 13, wherein "approximately" means ±10% of the following number.

20. The method according to claim 13, further comprising the step of determining from the sequencing data that the origin of the tissue is a tumor or cancer.