Methods for assaying circulating tumor DNA

EP4743587A1Pending Publication Date: 2026-05-20NATERA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
NATERA INC
Filing Date
2024-07-12
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Current methods for detecting circulating tumor DNA (ctDNA) in blood samples face challenges due to the low levels of tumor DNA present in early-stage and advanced-stage cancer patients, and the difficulty in distinguishing it from normal DNA.

Method used

A method involving the enrichment of cell-free DNA (cfDNA) subsets with differentially methylated regions, followed by sequencing and analysis of fragment length distribution to identify ctDNA.

Benefits of technology

This approach enhances the specificity and sensitivity of ctDNA detection, allowing for the identification of cancer presence even at low tumor DNA fractions, while reducing false positives from biological noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000048_0001
    Figure IMGF000048_0001
  • Figure IMGF000049_0001
    Figure IMGF000049_0001
  • Figure 00000056_0000
    Figure 00000056_0000
Patent Text Reader

Abstract

The present disclosure includes a method of comprising (a) obtaining cell-free DNA (cfDNA) from a sample from a subject; (b) selectively enriching subsets of the cfDNA from (a) or derivatives therefrom having one or more target regions to obtain enriched DNA, wherein the target regions are differentially methylated in cancer; (c) sequencing the enriched DNA from (b) to obtain sequence reads; and (d) partitioning a plurality of the sequence reads into two or more groups based on their methylation status and determining the fragment length distribution of the sequence reads in at least one of the groups.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS FOR ASSAYING CIRCULATING TUMOR DNACROSS-REFERENCE TO RELATED APPLICATION[1] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 526,943, filed July 14, 2023, the content of which is hereby incorporated by reference in its entirety.BACKGROUND[2] Unlike traditional biopsies involving invasive surgery, liquid biopsy is a non-invasive method that utilizes blood samples to diagnose and monitor cancers. Cell-free circulating DNA (“cfDNA”) in the blood comprises degraded DNA fragments that arc released from apoptotic or necrotic cells in many organs, and are regarded as a mixture of DNA from many normal tissue cells and diseased cells (e.g., cancerous tumor cells). Therefore, cfDNA are one of the best sources for blood-based cancer detection and have become of major interest for bloodbased cancer detection.[3] Early detection of cancer (e.g., before it has had a chance to metastasize) presents the best strategy for increasing cancer survival. Cancer detection using cfDNA from blood has attracted significant interest due to its non-invasive nature. However, tumor cfDNA levels are very low in most early-stage and many advanced-stage cancer patients (Bettegowda et al., Sci. Tansl. Med. 6(224): 224ra24, 2014; Newman et al., Nat. Med. 20(5): 548-54, 2014). Therefore, one major challenge in cfDNA-based early cancer diagnostics is how to identify the tiny amount of tumor cfDNA out of total cfDNA in blood. The mainstream approach to address this challenge is mutation-based, i.e., using targeted deep sequencing (>5,000x coverage), combined with errorsuppression techniques, to call cfDNA mutations, either specific mutations identified from a patient’s tumor biopsy or common tumor mutations, in a relatively small gene panel (Bettegowda et al., Sci. Tansl. Med. 6(224): 224ra24, 2014; Newman et al., Nat. Med. 20(5): 548-54, 2014; Newman et al., Nat. Biotechnol. 34:547-555, 2016). While this approach provides a sensitive and specific way to monitor cancer recurrence when the mutations are known, a small gene panel could not serve diagnostic purposes because mutations can be wide-spread and very heterogeneous, evenin the same type of cancer (Burrell et al., Nature 501(7467): 338-345, 2013; Turner et al., Lancet Oncol. 13(4): el78-185, 2012; Greenman et al., Natyre 446: 153-158, 2007; Schmitt et al., Ann. N. Y. Acad. Sci. 1267: 110-116, 2012). Unlike in a post-operative setting, there is no prior knowledge of what specific mutations might be present in a patient’ s tumor. However, enlarging the gene panel, while maintaining the sequencing depth, is cost-prohibitive. In addition, biological noise such as mutations from benign lesions or from the bone marrow through clonal hematopoiesis of indeterminate potential (CHIP), which increases with age, can result in false positives. Therefore, there remains the challenge of detecting the trace amount of tumor cfDNA, and alternative and / or additional multiomic approaches, including using the cfDNA methylation patterns, have been investigated.[4] A recent discovery by the inventors has established a new approach for detecting the presence of tumor cfDNA in a blood sample with increased specificity based on the fragment length distribution of cfDNA fragments comprising differentially methylated sequences.SUMMARY[5] In one aspect, the present disclosure provides a method of preparing deoxyribonucleic acid (DNA) useful for detecting circulating tumor DNA (ctDNA), said method comprising: (a) obtaining cell-free DNA (cfDNA) from a sample from a subject; (b) selectively enriching subsets of the cfDNA from (a) or derivatives therefrom having one or more target regions to obtain enriched DNA, wherein the target regions are differentially methylated in cancer; (c) sequencing the enriched DNA from (b) to obtain sequence reads; and (d) partitioning a plurality of the sequence reads into two or more groups based on their methylation status and determining the fragment length distribution of the sequence reads in at least one of the groups.[6] In some embodiments, the two or more groups comprises a hypermethylated group and an undermethylated group. In some embodiments, the two or more groups comprises a hypomethylated group and an overmethylated group.[7] In some embodiments of the methods described herein, the selectively enriching subsets of the cfDNA is performed by hybrid capture using a set of hybrid capture probes. In some embodiments, each probe of the set is designed to hybridize to one or more target regions.[8] Some embodiments of the methods described herein further comprise ligating adapters to the cfDNA to obtain adapter-ligated DNA. In some embodiments, the adapters are methylated adapters. In some embodiments, the adapters are Y adapters. In some embodiments, the adapters each comprise a universal priming site. In some embodiments, the adapters each further comprise a molecular barcode. In some embodiments, the number of adapters having different molecular barcodes is between 10 to 1,000, and wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1,000:1. In some embodiments, the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1.[9] Some embodiments of the methods described herein further comprises performing bisulfite conversion on the adapter-ligated DNA, thereby generating bisulfite-converted DNA. Some embodiments comprise treating the adapter-ligated DNA with sodium bisulfate, thereby generating bisulfite-converted DNA.

[0010] Some embodiments of the methods described herein further comprise performing enzymatic conversion on the adapter-ligated DNA, thereby generating enzymatic-converted DNA. In some embodiments, enzymatic conversion comprises treating the adapter- ligated DNA with TET2, T4-bGT, and APOBEC.

[0011] Some embodiments of the methods described herein further comprise amplifying the bisulfite- or enzyme-converted DNA. In some embodiments, the amplifying comprises one or more PCRs using universal primers. In some embodiments, amplification comprises barcoding PCR.

[0012] In some embodiments of the methods described herein, the hypermethylated group comprises sequence reads of cfDNA fragments each having 2 or more CpG sites, wherein 70% or more of the CpG sites are methylated. In some embodiments, the hypomethylated group comprises sequence reads of cfDNA fragments each having 2 or more CpG sites, wherein 30% or less of the CpG sites are methylated.

[0013] In some embodiments, a fragment length distribution of the sequence reads in the hypermethylated group of at least 5 bases shorter than the fragment length distribution of thesequence reads in the undermethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject. In some embodiments, a fragment length distribution of the sequence reads in the hypermethylated group of at least 10 bases shorter than the fragment length distribution of the sequence reads in the undcrmcthylatcd group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject. In some embodiments, a fragment length distribution of the sequence reads in the hypomethylated group of at least 10 bases shorter than the fragment length distribution of the sequence reads in the overmethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject. In some embodiments, a fragment length distribution of the sequence reads in the hypomethylated group of at least 20 bases shorter than the fragment length distribution of the sequence reads in the overmethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject.

[0014] In some embodiments, a length of the sequence reads in the hypermethylated group statistically not significantly different than the length of the sequence reads in the undermethylated group indicates the absence of ctDNA in the sample and / or the absence of cancer in the subject. In some embodiments, length of the sequence reads in the hypomethylated group approximately equal to the length of the sequence reads in the overmethylated group indicates the absence of ctDNA in the sample and / or the absence of cancer in the subject.

[0015] In one aspect, the present disclosure provides a method of preparing deoxyribonucleic acid (DNA) useful for detecting circulating tumor DNA (ctDNA), said method comprising: (a) obtaining cell-free DNA (cfDNA) from a sample from a subject; (b) contacting the cfDNA from (a) or derivatives therefrom with one or more restriction enzymes capable of binding to one or more DNA sequences containing one or more CpG sites and thereby generating treated cfDNA; (c) selectively enriching subsets of the treated cfDNA from (b) or derivatives therefrom having one or more target regions to obtain enriched DNA, wherein the target regions are differentially methylated in cancer; (d) sequencing the enriched DNA from (c) to obtain sequence reads; and (e) determining the fragment length distribution of the sequence reads in the enriched DNA.

[0016] Some embodiments further comprise ligating adapters to the cfDNA to obtain adapter- ligated DNA. In some embodiments, the adapters are methylated adapters. In some embodiments,the adapters are Y adapters. In some embodiments, the adapters each comprise a universal priming site. In some embodiments, the adapters each further comprise a molecular barcode. In some embodiments, the number of adapters having different molecular barcodes is between 10 to 1 ,000, and wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1,000:1. In some embodiments, the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1.

[0017] In some embodiments, the one or more restriction enzymes are one or more methylation sensitive restriction enzyme (MSREs). In some embodiments, the one or more restriction enzymes are one or more methylation dependent restriction enzyme (MDREs).

[0018] Some embodiments of the methods described herein further comprise amplifying the treated DNA. In some embodiments, amplifying comprises one or more PCRs using universal primers. In some embodiments, amplification comprises barcoding PCR.

[0019] In some embodiments, a fragment length distribution of the sequence reads relative to a threshold is used to classify the likelihood of presence or absence of cancer in the subject.

[0020] Some embodiments of the methods described herein, further comprise performing size selection on the cfDNA. In some embodiments, size selection comprises enriching the cfDNA for molecules that are between 70 and 500 base pairs in length. In some embodiments, size selection comprises enriching the cfDNA for molecules that arc between 100 and 200 base pairs in length. In some embodiments, size selection comprises enriching the cfDNA for molecules that are between 130 and 170 base pairs in length.

[0021] In some embodiments, the fragment length distribution is calculated as an average or mean length of the fragments in each group.

[0022] In some embodiments, wherein the sample is a liquid sample. In some embodiments, the sample is a blood, plasma, serum, or urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample. In some embodiments, the sample comprises DNA from a tumor.

[0023] Some embodiments of the methods described herein, further comprises repeating the method on a second sample. In some embodiments, the first sample is collected at a first time point and the second sample is collected at a second time point. In some embodiments, the first sample and the second sample arc collected from the same subject. In some embodiments, the subject is suspected or at risk of having a disease. In some embodiments, the first sample and the second sample are from different subjects. In some embodiments, the first subject is suspected or at risk of having a disease, and wherein the second subject is not suspected or at risk of having the disease. In some embodiments, the disease is a cancer. In some embodiments,

[0024] the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer.

[0025] In some embodiments, the method is capable of detecting fully methylated DNA molecules present at 1.0% (by mass) or more in a mixture of DNA molecules. In some embodiments, circulating tumor DNA (ctDNA) is present in 1% or more of the total circulating free DNA (cfDNA) in the sample. In some embodiments, circulating tumor DNA (ctDNA) is present in 0.1% or more of the total circulating free DNA (cfDNA) in the sample. In some embodiments, circulating tumor DNA (ctDNA) is present in 0.01% or more of the total circulating free DNA (cfDNA) in the sample.BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1. Shows an exemplary, non-limiting workflow of the present invention. As indicated in the Figure, a deamination agent can be used to convert unmethylated cytosine residues in adapted cfDNA fragments into uracils. Enrichment may be performed following library amplification using target capture methods that preserve the entire length of the targeted cfDNA fragments.

[0027] Figure 2. Shows an exemplary, non-limiting workflow of the present invention. As indicated in the Figure, adapted cfDNA fragments may be converted by either bisulfite or enzymatic conversion prior to amplification. Enrichment of the amplified fragments may be performed using hybrid capture with probes designed to target one or more differentially methylated sequences.

[0028] Figure 3. Shows size distribution graph for three representative samples (A-C) from healthy subjects for hypermethylated and undermethylated fragments comprising target differentially methylated regions (“DMRs”).

[0029] Figure 4. Shows size distribution graph for two representative samples (A-B) from colorectal cancer (“CRC”) subjects for hypermethylated and undermethylated fragments comprising target DMRs.DETAILED DESCRIPTION

[0030] The present invention generally relates to the field of cancer detection. More particularly, the present invention relates to detection and analysis of differentially methylated cell-free DNA (cfDNA) fragments in cancer.

[0031] Both global and local epigenetic changes are widely regarded as a hallmark of cancer. Alterations include a global decrease in overall CpG methylation levels coupled with discrete regions of hypermethylation, typically in CpG islands located in the promoter regions of tumor suppressor genes. Hypermethylation has been associated with cancer progression and the silencing of growth regulating genes and tumor suppressor genes. As a result, a growing number of DNA methylation biomarkers are being utilized in the development of novel assays for monitoring cancer progression, treatment response and early detection. It is possible to detect the presence of cancer through the analysis of the methylation status of specific CpG sites that arc hypermethylated or hypomethylated in DNA from tumor cells, including circulating tumor DNA (ctDNA) released from tumor cells, as compared to DNA from non-tumor cells, including cfDNA released from nontumor cells.

[0032] In addition, cfDNA fragmentation reflects the nucleosomal organization in the originating cells. Tumor-derived DNA fragments, such as ctDNA, tend to be shorter than non-tumor derived DNA fragments. Fragment length distribution of plasma DNA can therefore be used to determine whether a subject has cancer. However, because the fragment length distribution is generally wide and noisy, detecting DNA fragmentation changes requires millions of fragments for a low fractionof tumor-derived fragments. For example, in Stage 1 CRC samples, less than 1 in 1,000 cfDNA fragments is a tumor-derived fragment.

[0033] If cfDNA fragments can be classified as more likely tumor-derived or non-tumor-dcrivcd, DNA fragmentation changes would be easier to detect. As described herein, the inventors have developed a novel method to first classify cfDNA fragments based on methylation patterns. For example, fragments with hypermethylated CpG patterns are more likely to be tumor-derived and hence observing DNA fragmentation changes becomes much easier even for low tumor DNA fractions. Combining fragment length distribution with changes in methylation patterns also reduces false-positives from analyzing methylation patterns alone as fragments differentially methylated due to other confounding diseases or conditions have different fragmentation patterns.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which the present technology belongs.

[0035] As used herein, unless otherwise stated, the singular forms “a,” “an,” and “the” include plural reference. Thus, for example, a reference to “an oligonucleotide” includes a plurality of oligonucleotide molecules, and a reference to “a nucleic acid” is a reference to one or more nucleic acids.

[0036] As used herein, the term “about” in reference to a number is generally taken to include numbers that fall within a range of 1 %— 10% in cither direction (greater than or less than) of the number unless otherwise stated or otherwise evident from the context.

[0037] As used herein, the terms “individual”, “patient”, or “subject” can be an individual organism, a vertebrate, a mammal, or a human. In a preferred embodiment, the individual, patient or subject is a human.

[0038] As used herein, the term “library” refers to a collection of nucleic acid sequences, e.g., a collection of nucleic acids derived from whole genomic, sub-genomic fragments, cell-free DNA, cell-free DNA fragments, cDNA, cDNA fragments, RNA, RNA fragments, or a combination thereof. In one embodiment, a portion or all of the library nucleic acid sequences comprises an adapter sequence. The adapter sequence can be located at one or both ends. The adapter sequencecan be useful, e.g., for a sequencing method (e.g., an NGS method), for amplification, for reverse transcription, or for cloning into a vector. In some embodiments described here, the libraries are constructed from cell-free DNA and / or gDNA.

[0039] “Next generation sequencing” or “NGS” as used herein, refers to any sequencing method that determines the nucleotide sequence of either individual nucleic acid molecules (e.g., in single molecule sequencing) or clonally expanded proxies for individual nucleic acid molecules in a high throughput parallel fashion (e.g., greater than 103, 104, 105or more molecules are sequenced simultaneously). In one embodiment, the relative abundance of the nucleic acid species in the library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment. Next generation sequencing methods are known in the art, and are described, e.g., in Metzker, M. Nature Biotechnology Reviews 11 :31-46 (2010).

[0040] As used herein, the term “nucleic acid” refers to a deoxyribonucleotide (DNA), ribonucleotide polymer (RNA), RNA / DNA hybrids and polyamide nucleic acids (PNAs) in either single- or double-stranded form, and unless otherwise limited, would encompass known analogs of natural nucleotides that can function in a similar manner as naturally occurring nucleotides.

[0041] As used herein, the term “cell-free DNA” or “cfDNA” refers to any free-floating DNA existing in a sample, such as the blood plasma of a cancer patient, pregnant person, or a transplant recipient. Cell-free DNA found in a pregnant woman's blood may contain DNA originating from both the mother and the fetus. Cell free DNA found in a transplant recipient may contain DNA originating from the recipient and the donor. The term “circulating tumor DNA” or “ctDNA” as used herein refers to tumor DNA circulating freely in the blood of a cancer patient, which carries the genetic and epigenetic changes specific to tumors. The ctDNA may be sampled by venipuncture on the subject and provides the basis for non-invasive cancer diagnosis and testing.

[0042] As used herein, “oligonucleotide” refers to a molecule that has a sequence of nucleic acid bases on a backbone comprised mainly of identical monomer units at defined intervals. The bases are arranged on the backbone in such a way that they can bind with a nucleic acid having a sequence of bases that are complementary to the bases of the oligonucleotide. The most common oligonucleotides have a backbone of sugar phosphate units. A distinction may be made betweenoligodeoxyribonucleotides that do not have a hydroxyl group at the 2' position and oligoribonucleotides that have a hydroxyl group at the 2' position. Oligonucleotides may also include derivatives, in which the hydrogen of the hydroxyl group is replaced with organic groups, e.g., an allyl group. Oligonucleotides that function as primers or probes arc generally at least about 10-15 nucleotides in length or up to about 70, 100, 110, 150 or 200 nucleotides in length, and more preferably at least about 15 to 25 nucleotides in length. Oligonucleotides used as primers or probes for specifically amplifying, enriching, or detecting a particular target nucleic acid generally are capable of specifically hybridizing to the target nucleic acid.Subjects

[0027] Subjects involved in methods herein can be virtually any animal, in some embodiments a mammal, and in further embodiments, a human. In some embodiments, the subject is suspected or at risk of having a disease, in some embodiments, cancer. In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject comprising an organ from another individual.

[0028] In embodiments where the subject has cancer, the cancer can be any type of cancer provided that the genome of cancerous cells of the subject have portions of their genome that are differentially methylated compared to non-cancerous cells of the subject. Typically, some, most, almost all or all cancer cells of the subject have region(s) of their genome that are methylated that are not methylated in non-cancerous cells of the subject, or that are more methylated than non- cancerous cells, or vice versa. Thus, in some embodiments, the subject has one or more cancers (e.g. one cancer) including ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficialbasal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.

[0029] In certain embodiments of methods herein, the subject has a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovaries, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whippie resection. In some embodiments, the cancer is selected from nasopharyngeal carcinoma, hepatocellular carcinoma, breast cancer, ovarian cancer, pancreatic cancer, colorectal cancer, lung cancer, oesophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. In some embodiments, the cancer is colorectal cancer.Biological Sample Collection and Preparation

[0043] The cell-free nucleic acid may be isolated from any type of suitable liquid biological specimen or sample (e.g., a test sample). A sample or test sample can be any specimen that is isolated or obtained from a subject or pail thereof (e.g., a human subject, a pregnant female, afetus). Non-limiting examples of specimens include fluid from a subject, including, without limitation, blood or a blood product (e.g., serum, plasma, or the like), umbilical cord blood, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneal, ductal, car, arthroscopic), washings of female reproductive tract, urine, fcccs, sputum, saliva, nasal mucous, prostate fluid, lavage, semen, lymphatic fluid, bile, tears, sweat, breast milk, breast fluid, the like or combinations thereof. In some embodiments, a liquid biological sample is a blood plasma or serum sample. The term “blood” as used herein refers to a blood sample or preparation from a subject being tested for the presence of cancer. The term encompasses whole blood, blood product or any fraction of blood, such as serum, plasma, buffy coat, or the like as conventionally defined. Blood or fractions thereof often comprise nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes comprise nucleic acids and are sometimes cell-free or intracellular. Blood also comprises buffy coats. Buffy coats are sometimes isolated by utilizing a ficoll gradient. Buffy coats can comprise white blood cells (e.g., leukocytes, T-cells, B-cells, platelets, and the like). Blood plasma refers to the fraction of whole blood resulting from centrifugation of blood treated with anticoagulants. Blood serum refers to the watery portion of fluid remaining after a blood sample has coagulated. Fluid samples often are collected in accordance with standard protocols hospitals or clinics generally follow. For blood, an appropriate amount of peripheral blood (e.g., between 3-40 milliliters) often is collected and can be stored according to standard procedures prior to or after preparation. A fluid sample from which nucleic acid is extracted may be acellular (e.g., cell-free). In some embodiments, a fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancer cells may be included in the sample. In some embodiments the cells, cellular elements or cellular remnants are removed from the liquid sample prior to nucleic acid extraction.

[0044] A sample may be heterogeneous, by which is meant that more than one type of nucleic acid species is present in the sample. For example, heterogeneous nucleic acids can include, but are not limited to, (i) fetal derived and maternal derived nucleic acids, (ii) nucleic acids from cancerous and non-cancerous cells, (iii) pathogen and host nucleic acids, (iv) transplant donor and recipient nucleic acids, or (v) mutated and wild-type nucleic acids.

[0045] To prepare serum or plasma from blood, blood can be placed for example in a tube containing EDTA or a specialized commercial product such as Streck BCT collection tubes(Streck). Seram may be obtained with or without centrifugation-following blood clotting. If centrifugation is used then it is typically, though not exclusively, conducted at an appropriate speed, e.g., 1 ,500-3,000 x g. Plasma or serum may be subjected to additional centrifugation steps before being transferred to a fresh tube for cfDNA extraction. In some embodiments, cfDNA can be extracted using a kit. For example, the Qiagen Circulating Nuclei Acid Kit can be used. In some embodiments, the QIAsymphony circulating DNA Kit may be used for automated cfDNA extraction. In some embodiments, gDNA may be extracted using the QIAgen DNeasy Blood and Tissue Kit.Size Selection

[0046] In some embodiments, size selection may be used to enrich extracted cfDNA molecules of certain sizes before library preparation according to methods described herein.

[0047] Size selection may be performed for example by gel electrophoresis, paramagnetic beads, spin column, salt precipitation, or biased amplification to isolate DNA of a particular size. For example, size selection may be performed to enrich cfDNA molecules that are from 50 to 1200 base pairs in length, or from 70 to 800 base pairs in length, or from 100 to 200 base pairs in length, or from 130 to 180 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length. In some embodiments, the enriched cfDNA molecules range from 100 to 200 bp, from 120 to 180bp, from 140 to 160bp, from 150 to 170bp, from 160 to 190 bp, or from 170 to 220 bp. In some embodiments, the enriched cfDNA molecules are less than 500, 400, 200, 150, 100, 90, 75, or 50 bp in length.

[0048] In some embodiments, the size exclusion may be performed by using gel electrophoresis to separate the cfDNA samples according to size and a determined size range was selected. For example, gel electrophoresis separates DNA molecules based on their size by applying an electric field to a gel, such as an agarose gel, upon which DNA molecules will move through the gel towards the positively charged anode. The size of the DNA molecules determines the speed by which the DNA molecules migrate through the gel. A standard mixture of DNA molecules with predetermined sizes may be applied to the gel to identify the size of the DNA. The DNA moleculesof desired size were then extracted and purified. In some embodiments, the size selection was performed on an automated high-throughput gel electrophoresis system such as Pippin or Costal Genomics systems.

[0049] In some embodiments, the size exclusion step of the methods disclosed herein may be performed by using paramagnetic beads. The use of paramagnetic beads for size selection of DNA fragments is described in DeAngelis et al., Solid-Phase Reversible Immobilization for the Isolation of PCR Products, Nucleic Acid Research, 23(22): 4742-4743 (1995), incorporated herein by reference in its entirety. In brief, this method is based on that DNA fragment size affects the total charge per molecule with larger DNAs having larger charges, which promotes their electrostatic interaction with the beads and displaces smaller DNA fragments. Thus, by manipulating the composition of the buffer solution used to mix beads and DNA, the beads can be made to bind DNA within specific size ranges. The most famous and highly applied approach is Solid Phase Reversible Isolation (SPRI) selection which utilizes carboxyl coated paramagnetic beads in the presence of high salt and the crowding agent polyethylene glycol (PEG), to promote controlled adsorption, configure to bind DNA molecules within a certain molecular weight ranges by varying PEG concentrations. DNA molecules of differing length can be partitioned by subjecting source DNA to various binding and elution schemes in the presence of different amounts of PEG. In some embodiments, AMPURE™ beads are used for the size exclusion step.

[0050] In some embodiments, the size exclusion step of the methods disclosed herein may be performed by using spin columns. A spin column contains material that will absorb molecules based on the size of the molecules. The spin column material contains pores of defined sizes and molecules with a size above a cutoff size determined by the pore size will not enter the pores, and are eluted with the column's void volume. Different types of column material can be chosen to achieve absorption or exclusion of DNA molecules within various size ranges. In some embodiments, the spin column material comprises siliceous materials, silica gel, glass, glass fiber, zeolite, aluminum oxide, titanium dioxide, zirconium dioxide, kaolin, gelatinous silica, magnetic particles, ceramics, polymeric supporting materials, or a combination thereof. In a particular embodiment, the spin column material comprises glass fiber.

[0051] In some embodiments, spin columns may be used for size exclusion by using different binding buffers configured to provide low or high stringency binding conditions when applying the DNA samples to the spin column, as described in PCT patent application No. PCT / US2019 / 018274 filed on Feb. 15, 2019, which is incorporated herein by reference in its entirety. Under low stringency binding conditions, the spin column material be configured to restrict binding of DNA fragments of high molecular weights, whereas high stringency binding conditions will configure the spin column to facilitate binding of DNA fragments with low molecular weights.

[0052] In some embodiments, the low and / or high stringency binding buffer comprises a nitrile compound selected from acetonitrile (ACN), propionitrile (PCN), butyronitrile (BCN), isobutylnitrile (IBCN), or a combination thereof. The first and / or second binding buffer can comprise, for example, about 15% to about 35%, or about 20% to about 30%, or about 25% of the nitrile compound (e.g., ACN).

[0053] In some embodiments, the low and / or high stringency binding buffer comprises a chaotropic compound selected from GnCl, urea, thiourea, guanidine thiocyanate, Nal, guanidine isothiocyanate, D- / L-arginine, a perchlorate or perchlorate salt of Li+, Na+, K+, or a combination thereof. The low and / or high stringency binding buffer can comprise, for example, about 5 M to about 8 M, or about 5.6 M to about 7.2 M, or about 6 M of the chaotropic compound (e.g., GnCl).

[0054] The binding buffers may also comprise an alcohol, a chelating agent, and a detergent. In some embodiments, the alcohol is propanol. In some embodiments, the chelating compound comprises ethylenediaminetetraccetic (EDTA), ethyleneglycol-bis(2-aminoethylether)- N,N,N',N'-tetraacetic acid (EGTA), citric acid, N,N,N',N'-Tetrakis(2- pyridylmethyl)ethylenediamine (TPEN), 2,2'-Bipyridyl, deferoxamine methanesulfonate salt (DFOM), 2,3-Dihydroxybutanedioic acid (tartaric acid), or a combination thereof. In some embodiments, the detergent may be Triton X-100, Tween 20, N-lauroyl sarcosine, sodium dodecylsulfate (SDS), dodecyldimethylphosphine oxide, sorbitan monopalmitate, decylhexaglycol, 4-nonylphenyl-polyethylene glycol, or a combination thereof. In a particular embodiment, the detergent is Triton X-100.

[0055] In some embodiments, the size selection step of the methods disclosed herein can be performed by using salt precipitation. Larger DNA molecules will precipitate at lower salt concentrations than smaller DNA molecules. By varying the concentration of salt in the precipitation buffer, DNA molecules in different size ranges can be separated.

[0056] In some embodiments, the size selection step can be performed by biased PCR. In some embodiments, biased PCR can enrich for shorter DNA molecules by using shorter time for DNA extension in the PCR cycle protocol. If desired, the extension step of the PCR amplification may be limited from a time standpoint to reduce amplification from fragments longer than 200 nucleotides, 300 nucleotides, 400 nucleotides, 500 nucleotides or 1,000 nucleotides. This may result in the enrichment of fragmented or shorter DNA (such as fetal DNA or DNA from cancer cells that have undergone apoptosis or necrosis) and improvement of test performance. In some embodiments, biased PCR can enrich for shorter DNA molecules by using a polymerase with low processivity.Library Preparation

[0057] Typically, methods herein include a step of ligating nucleic acid adapters to sample DNA molecules, or nucleic acid derivatives generated therefrom. In some embodiments, before such ligation, extracted or isolated DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adapter ligation. For example, sample DNA molecules can be end repaired (e.g. blunt end repaired), nucleotides can be added to sample DNA molecules or bluntcd-cndcd derivative therefrom, and / or phosphate moictics can be added to or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, sample DNA molecules may be blunt end repaired, and then a single adenosine base can be added to the 3’ end, i.e. A-tailed. In some embodiments, during ligation the 3’ adenosine of the sample fragments and the complementary 3’ tyrosine overhang of an adapter can enhance ligation efficiency. Typically, adapter ligation is performed using a T4 ligase.

[0058] In some embodiments, methylated adapters are utilized in methods herein. In some embodiments, adapters containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adapters are Y adapters, for example in methods inwhich targeted amplicons are sequenced using NGS. In some embodiments, the adapters each comprises a universal priming site. In some embodiments, the adaptors further comprise a sequence that can be utilized in subsequent sequencing step, such as NGS. In some embodiments, the adapters further comprise a sample barcode such that multiple samples can be pooled and analyzed in the same sequencing reaction. The sample barcode can be used to process data according to the sample from which the data was generated.

[0059] In some embodiments, the adapters each further comprises a molecular barcode. In some embodiments, the number of adapters having different molecular barcodes is between 10 to 1,000, and wherein the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1 ,000: 1. The number of different molecular barcodes in the ligation reaction, in certain embodiments, ranges from 10 to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different molecular barcodes in the ligation reaction. In some embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1. In some embodiments, the ratio of the total number of sample nucleic acid or cfDNA molecules to the number of different molecular barcodes in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000: 1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50:1. In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1 ,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different molecular barcodes to each one sample nucleic acid or cfDNA molecules.Amplifications

[0060] Methods in some aspects herein include performing one, two, or more amplifications. In some embodiments, methods herein include one or more universal amplifications using primers that hybridize to universal priming sequences. In some embodiments, sample barcodes are added during one or more amplifications.

[0061] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g. recombinase polymerase amplification(RPA) (Kersting et al., Microchim. Acta 181(13-14): 1715-1723, 2014; incorporated by reference in its entirety), a ligase-based amplification, PCR, or a combination thereof (e.g. ligation-mediated PCR).

[0062] As discussed above, in some methods herein, a universal amplification(s) can be performed before and / or after target enrichment, following bisulfite or enzymatic treatment of the DNA molecules. Such universal amplification can be performed for example using a primer pair that binds universal priming sites in the adapter. In some embodiments, the methods herein include performing a universal PCR using a plurality of adapted DNA molecules or a plurality of adapted cfDNA, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters, to generate amplified, adapted, DNA or cfDNA molecules, before performing target enrichment. In some embodiments, the methods herein include performing a universal PCR using a plurality of enriched, adapted DNA molecules or a plurality of enriched, adapted cfDNA, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters or introduced via an initial amplification, to generate amplified, enriched, adapted, DNA or cfDNA molecules, after performing target enrichment.

[0063] In some embodiments, performing a PCR further comprises using primers comprising a sequence that can be utilized in subsequent sequencing step, such as NGS. For example, such sequences can include NGS flow cell binding sites (e.g. Illumina P5 and P7 sequences) and / or NGS sequencing primer binding sites. In some embodiments, performing a PCR further comprises using primers comprising a sample index.

[0064] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g. multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more cycles. In some embodiments, amplification cycles can include at least 5, 6, 7, 8, 9, or 10 cycles. In some embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.

[0065] Typically, in embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., adapted sample DNA or cfDNA) followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, dcoxynuclcotidcs (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is between from 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.

[0066] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgCh), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between of 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.

[0067] In some embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England BioLabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England BioLabs, Inc.).

[0068] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High- Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3'— 5' exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5'— > 3 'exonuclease activity and strand displacement activity.

[0069] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5'— > 3' direction and requires the presence of template and primer. This enzyme has a 3'— > 5' exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5'— ► 3' exonuclease activity and strand displacement activity.

[0070] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers arc between 5 and 50 nucleotides in length, between 10 and 40 nucleotides in length, between 15 and 30 nucleotides in length, between 15 and 25 nucleotides in length, between 20 and 40 nucleotides in length, between 25 and 50 nucleotides in length, or between 30 and 50 nucleotides in length. In some embodiments, the primers are between 25 and 100 nucleotides in length, between 35 and 100 nucleotides in length, between 45 and 100 nucleotides in length, between 55 and 100 nucleotides in length, between 65 and 100 nucleotides in length, or between 75 and 100 nucleotides in length.Detection of Differential Methylation

[0071] DNA methylation occurs predominantly on cytosine residues in CpG dinucleotide pairs. Both global and local changes in methylation levels are widely regarded as a hallmark of cancer. Alterations include a global decrease in overall CpG methylation (hypomethylation) coupled with an increase in methylation (hypcrmcthylation) in discrete CpG-rich regions called CpG islands, typically in CpG islands located in the promoter regions of tumor suppressor genes. Hypermethylation has been associated with cancer progression and the silencing of growth regulating genes and tumor suppressor genes. As a result, a growing number of DNA methylation biomarkers are being utilized in the development of novel assays for monitoring cancer progression, treatment response and early detection. It is possible to detect the presence of cancer through the analysis of the methylation status of specific and of combination of CpG sites that are predominantly methylated in DNA from tumor cells, including circulating tumor DNA (ctDNA) released from tumor cells and / or from cells within the tumor’ s microenvironment.

[0072] As used herein, the term “differentially methylated” refers to a region or fragment or CpG dinucleotide of DNA that is differentially methylated (either more methylated or less methylated) in a cancer sample compared to the same region or fragment or CpG in a non- cancerous sample. Genomic regions comprising a plurality of differentially methylated CpG sites can be referred to as differentially methylated regions or DMRs. Typically, some, most, almost all or all cancer cells have region(s) of their genome that are methylated that are not methylated in non-cancerous cells, or that are more methylated than non-cancerous cells.

[0073] To determine the methylation status of CpG dinucleotides in a DNA fragment, the methods described herein may utilize several approaches, including but not limited to bisulfite conversion and enzymatic (EM) conversion.

[0074] The term “bisulfite” as used herein encompasses all types of bisulfites, such as sodium bisulfite, that are capable of chemically converting an unmethylated cytosine (C) to a uracil (U) through deamination, but not a methylated cytosine, which are resistant to the deamination process. Bisulfite conversion therefore can be used to differentiate and detect methylated versus unmethylated cytosines in a DNA fragment. Commercial bisulfite conversion kits, such as EpiMark bisulfite conversion Kit (NEB) and EZ DNA Methylation Kit (Zymo Research), are available. Subsequent analysis of the bisulfite converted DNA can be used to quantify methylation,for example by DNA sequencing, PCR (for sequence specific amplification), Southern bolt analysis, and use of methylation- sensitive restriction enzymes. Genomic sequencing is a technique that has been simplified for analysis of DNA methylation patterns and 5-methylcytosine distribution by using bisulfite treatment (Frommcr et al., Proc. Natl. Acad. Sci. USA 89:1827- 1831 (1992), incorporated herein by reference in its entire). Additionally, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA may be used, e.g., the method described by Sadri & Hornsby (Nucl. Acids Res. 24:5058-5059, 1996, incorporated herein by reference in its entire), or COBRA (Combined Bisulfite Restriction Analysis) (Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997, incorporated herein by reference in its entire).

[0075] Enzymatic methyl sequencing (“EMseq”) may also be used to characterize DNA methylation states. This method relies on enzymatic-conversion, using two sets of enzymatic reactions. In the first reaction, TET2 and T4-bGT convert 5mC and 5hmC into substrates that cannot be deaminated by APOBEC3A. In the second reaction, APOBEC3A deaminates unmodified cytosines converting them to uracils. The protection of 5mC and 5hmC permits the discrimination of cytosines from 5mC and 5hmC. EM-seq libraries may be prepared using commercially available kits, for example the NEBNext Enzymatic Methyl-seq Kit (NEB).

[0076] In addition, methylation sensitive restriction enzymes (MSREs) can be used to determine the methylation status of DNA fragments and / or to enrich for fully methylated DNA fragments. MSREs are DNA restriction endonucleases that are dependent on the methylation state of their DNA recognition site for activity, and selectively cleave a DNA fragment when the MSRE recognition site is unmcthylatcd, but not when the MSRE recognition site is methylated. An “isoschizomer” of an MSRE is a restriction enzyme that recognizes the same recognition site as a methylation sensitive restriction enzyme but cleaves both methylated CGs and unmethylated CGs. Isoschizomer of the selected MSREs may be used in control reactions. Non-limiting examples of methylation sensitive restriction enzyme include, and thus in some embodiments, the one or more MSREs can include, Aatll, Acc65I, AccI, Acil, Acll, Afel, Agel, AgeLHF®, AhdI, Alel-v2, Apal, ApaLI ApeKI, Asci, AsiSI, Aval, Avail, Bael, BanI, BbvCI, BceAI,, Bcgl, BcoDI, BfuAI, Bgll, BmgBI, BsaAI, BsaBI, BsaHI, BsaI-HF®v2, BseYI, BsiE, BsiWI, BsiWLHF®, BslI, BsmAI, BsmBI-v2, BsmFI, BspDI, BspEI, BsrBI, BsrFLv2, BssHII, BstAPI, BstBI, BstUI, BstZ17LHF®, BtgZI, Cac8I, Clal, Dpnl, Dralll-HF®, DrdI, Eael, Eagl-HF®, Earl, Ecil, Eco53kl, EcoRI, EcoRI-HF®,EcoRV, EcoRV-HF®, Esp3I, Faul, Fnu4HI, FokI, Fsel, FspI, Haell, Hgal, Hhal, HinPlI, Hindi, Hinfl, Hpal, Hpall, Hpyl66II, Hpyl88III, Hpy99I, HpyAV, HpyCH4IV, KasI, Mbol, Mini, MluI-HF®, Mmel, MspAlI, Mwol, Nad, Narl, Neil, NgoMIV, Nhel-HF®, NlalV, Notl, Notl-HF®, Nrul, NruI-HF®, Nt.BbvCI, Nt.BsmAI, Nt.CviPII, PacR7I, PaqCI, Pld, PluTI, Pmd, Pmll, PshAI, PspOMI, PspXI, Pvul, PvuI-HF®, Rsal, RsrII, Sad-HF®, Sadi, Sall, Sall-HF®, Sau3AI, Sau96I, ScrFI, SfaNI, Sfil, Sfol, SgrAI, Smal, SnaBI, Srfl, StyD4I, Tfil, Tsd, TspMI, Xhol, Xmal, and / or Zral. In some embodiments, the one or more MSREs can include Hhal, Hpall, BstUl, and / or HpyCH41V.

[0077] In some embodiments disclosed herein, one or more methylation dependent restriction enzymes (MDREs) can be used to enrich for unmethylated DNA fragments. MDREs selectively cleave a sample nucleic acid when one or more nucleotides in the MDRE recognition site is methylated. A skilled artisan will understand how to modify the methods provided herein to include MDREs, or to replace elements recited as MSREs with MDREs. Thus, in some embodiments, one or more MDREs can be used instead of or in the absence of MSREs. In some embodiments, one or more MDREs can be used in combination with one or more MSREs. In some embodiments, the one or more MDREs can be AbaSI, AoxI, BisI, BlsI, Dpnl, FspEI, Glal, Glul, Krol, LpnPI, Mall, MspJI, Mtel, Pcsl, PkrI, or Sgel.

[0078] The MSRE (or MDRE) can be selected based on differentially methylated CpG sites in target DNA molecules, such as tumors or ctDNA from specific cancer targets, or a diverse spectrum of tumors. Further criteria for selection may include low background methylation in normal tissues, size and number of cleavage fragments, number of base pairs of recognition sequence, whether the cleavage results in blunt vs. tailed end fragments, and whether the enzymes have the same or similar reaction conditions such that the contacting step can be done under the same set of conditions and / or in a single reaction.

[0079] In some embodiments, more than one, a plurality, or a set of MSREs (and / or MDREs) can be used to contact target DNA molecules comprising one or more CpG sites of interest. The number of selected MSREs in certain embodiments is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; from 2 to 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; or from 5 to 6, 7, 8, 9,10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 10, 1 to 5, from 2 to 5, from 3 to 5, from 4 to 5, 3 to 7, from 5 to 10, from 6 to 10, or from 7 to 10. Thus, criteria such as target site, target site methylation status, number of base pairs in tailed end fragments, reaction conditions including but not limited to buffer conditions, incubation time and temperature for optimal activity, as well as deactivation time and temperature for each MSRE can be used for selecting the plurality or set of MSREs and / or MDREs to include in the method. As activity measured by units is specific to each enzyme, the target DNA molecules can be contacted with from 1 to 5 Units (U) of each MSREs in the reaction sample. In some embodiments, the target DNA molecules are contacted with from 1 to 5 U, from 1.5 to 4 U, from 2 to 3 U, from 2.5 to 4 U, or from 3 to 5 U of each MSRE. In some embodiments, one or more of the MSRE is selected from Hpall, Sall, Bbel, Notl, Smal, Xmal, Mbol, BstUI, BstBI, Clal, Mini, Nael, Narl, Pvul, Sadi, HpyCH41V, Hhal, and combinations thereof. In exemplary embodiments, the one or more MSREs is selected from one or more of Hpall, Hhal, HpyCH41V, and BstUI. In some embodiments, the one or more MSREs are selected from one or more of Hpall, Hhal, HpyCH41 V, or BstU 1. In some embodiments of the methods as described herein, the contacting comprises contacting with two or more MSREs. In some embodiments, the contacting comprises contacting with three or more MSREs. In some embodiments, the contacting comprises contacting with four or more MSREs. In some embodiments, the contacting comprises contacting with the two or more MSREs in a single reaction. MSREs and methods of using MSREs are described in detail in U.S. Application No. 63 / 437016, incorporated by references herein.Target Enrichment

[0080] In methods described herein, cfDNA fragments and full-length amplicons derived therefrom comprising differentially methylated targets (e.g. having 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more 45 or more, 50 or more differentially methylated CpG sites) can be identified and enriched for using a custom panel containing probes targeting the differentially methylated sequences. In some embodiments, 75% or more of targets include 10 or more CpGs per target, with a median of 15-17 CpGs and a mean of 18-20 CpGs per target. In some embodiments, 85% or more of targets include 10 or more CpGs per target, with a median of 16 CpGs and a mean of 19.6 CpGs per target. In some embodiments, some probes may target multiple different differentially methylatedsequences. In some embodiments, a differentially methylated sequence may be targeted by multiple probes. In some embodiments, probes can be designed for both the methylated and unmethylated versions of the target sequence. In some embodiments, probes can be designed for both the top and bottom strands of the target sequence. In some embodiments, subsets of targets can be selected to improve signal to noise ratio.

[0081] In some embodiments, the panel targets up to 100 (e.g., 10-100, or 20-100, or 50-100) differentially methylated regions. In some embodiments, the panel targets up to 200 (e.g., 20-200, or 50-200, or 100-200) differentially methylated regions. In some embodiments, the panel targets up to 500 (e.g., 50-500, or 100-500, or 200-500) differentially methylated regions. In some embodiments, the panel targets up to 800 (e.g., 50-800, or 100-800, or 200-800) differentially methylated regions. In some embodiments, the panel targets up to 1,000 (e.g., 50-1,000, or 100- 1,000, or 200-1,000, or 500-1,000) differentially methylated regions. In some embodiments, the panel targets up to 2,000 (e.g., 50-2,000, or 100-2,000, or 200-2,000, or 500-2,000) differentially methylated regions. In some embodiments, the panel targets up to 5,000 (e.g., 50-5,000, or 100- 5,000, or 200-5,000, or 500-5,000) differentially methylated regions. In some embodiments, the panel targets up to 10,000 (e.g., 50-10,000, or 100-10,000, or 200-10,000, 500-10,000, or 1,000- 10,000) differentially methylated regions. In some embodiments, the panel targets up to 50,000 (e.g., 50-50,000, or 100-50,000, or 200-50,000, 500-50,000, 1,000-50,000, or 10,000-50,000) differentially methylated regions. In some embodiments, the panel targets up to 100,000 (e.g., 50- 100,000, or 100-100,000, or 200-100,000, 500-100,000, 1,000-100,000, or 10,000-100,000) differentially methylated regions. In some embodiments, the panel targets up to 150,000 (e.g., 50- 150,000, or 100-150,000, or 200-150,000, 500-150,000, 1,000-150,000, or 10,000-150,000) differentially methylated regions. The panel may optionally include probes targeting methylated and / or unmethylated sequences from a control plasmid or phage (e.g. pUC19, lambda).

[0082] In some embodiments, the panel may provide sequence coverage of at least about 50 kb, at least about lOOkb, at least about 150 kb, at least about 200kb, at least about 500kb, at least about 1 Mb, at least about 5 Mb, at least about 10 Mb, or at least about 20 Mb.

[0083] In some embodiments the panel may comprise up to 100 (e.g., 10-100, or 20-100, or 50-100) different probes. In some embodiments the panel may comprise up to 200 (e.g., 20-200, or50-200, or 100-200) different probes. In some embodiments the panel may comprise up to 500 (e.g., 50-500, or 100-500, or 200-500) different probes. In some embodiments the panel may comprise up to 800 (e.g., 50-800, or 100-800, or 200-800) different probes. In some embodiments the panel may comprise up to 1,000 (e.g., 50-1,000, or 100-1,000, or 200-1,000, or 500-1,000) different probes. In some embodiments the panel may comprise up to 2,000 (e.g., 50-2,000, or 100-2,000, or 200-2,000, or 500-2,000) different probes. In some embodiments the panel may comprise up to 5,000 (e.g., 50-5,000, or 100-5,000, or 200-5,000, or 500-5,000) different probes. In some embodiments the panel may comprise up to 10,000 (e.g., 100-10,000, 200-10,000, or 500- 10,000, or 1,000-10,000) different probes. In some embodiments the panel may comprise up to 50,000 different probes. In some embodiments the panel may comprise up to 80,000 different probes. In some embodiments the panel may comprise up to 100,000 different probes.

[0084] In some embodiments, target enrichment may be performed by linked target capture (LTC) using probe-dependent primers (PDPs). PDPs and LTC methods have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2020 / 039261 “Linked target capture and ligation”, which are hereby incorporated by reference in their entirety). Briefly, in an LTC method, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base to inhibit polymerase extension. In some embodiments, probes are designed to cover the desired region with zero gap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the primers of a PDP pair comprises a sample index. In some embodiments, a sample index is added via barcoding PCR following LTC enrichment.

[0085] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, atleast one of the probe binding regions can include a differentially methylated site. In some embodiments, one of the probe binding regions can include a differentially methylated site. In some embodiments, both of the probe binding sites of the probe binding regions can include one a differentially methylated site. In some embodiments, neither of the probe binding sites of the probe binding region comprises a differentially methylated site.

[0086] Typically, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on a ligated adapter or introduced through optional pre-capture universal PCR amplification following ligation. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In some embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adapter.

[0087] The primer can be tailed or untailed on either side of the forward and reverse side depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides in length. In certain embodiments, the PDP primers arc between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.

[0088] Typically, probe dependent primers comprise a linker between the probe and the primer. Probe and primer portions of the PDP are typically linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. Insome embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof.

[0089] In some embodiments, target enrichment may involve fragment capture by hybridization (i.e. hybrid capture; HC). Although any hybrid capture method can be used to perform methods herein that include a selective enrichment step, in some embodiments, a method of the present disclosure may involve using any of the hybrid capture methods disclosed herein to selectively enrich DNA, for example, cfDNA. Such selective enrichment step can follow an amplification step, typically follows a universal pre-amplification step, sometimes immediately following a universal pre-amplification step, sometimes following an additional barcoding PCR amplification step wherein a sample index is introduced.

[0090] In capture by hybridization, hybrid capture oligonucleotide probes complementary to a specific nucleic acid sequence are utilized to capture nucleic acid fragments, such as DNA or cfDNA fragments in a sample or DNA or cfDNA derived therefrom, comprising at least a portion of the sequence. In some embodiments, a hybrid capture panel includes one or more probe sets each having at least one hybrid capture probe, wherein each set is designed to bind to a different DMR. In some embodiments, one or more hybrid capture probes in the one or more probe sets can bind to two or more DMRs. In illustrative embodiments disclosed herein, probe sets are designed to ensure full coverage of the entire target DMR (differentially methylated region). For example, in some embodiments, a single hybrid capture probe in a set is designed to bind the top strand of the target sequence (i.e., IT). In some embodiments, the set includes two hybrid capture probes that bind each target region on cither the top strand (i.e., 2T) or the bottom strand (i.e., 2B). In some embodiments, the set includes at least one hybrid capture probe that binds the top strand of the target sequence and at least one hybrid capture probe that binds the bottom strand of the hybrid capture sequence (i.e., 1T1B). In some embodiments, the set includes at least two hybrid capture probes that bind the top strand of the target sequence and at least two hybrid capture probes that bind the bottom strand of the hybrid capture sequence (i.e., 2T2B). Sets can also be designed as 3T3B, 4T4B, etc. In some embodiments, at least two hybrid capture probes are tiled to ensure coverage of the target sequence. For example, the probes may overlap by about 5 nucleotides, by about 10 nucleotides, by about 15 nucleotides, by about 20 nucleotides, by about 25 nucleotides, by about 30 nucleotides, by about 40 nucleotides, or by about 50 nucleotides or more. In someembodiments, the set includes both hybrid capture probe(s) designed to bind the methylated version of the target sequence(s) and hybrid capture probe(s) designed to bind the unmethylated version of the target sequence(s).

[0091] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 20 bases to 120 bases, 20 bases to 110 bases, 20 bases to 100 bases, 20 bases to 90 bases, 20 bases to 80 bases, 20 bases to 70 bases, 20 bases to 60 bases, 20 bases to 50 bases, 20 bases to 40 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases, 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 10 bases to 120 bases, 20 to 120 bases, 30 bases to 120 bases, 40 bases to 120 bases, 50 bases to 120 bases, 60 bases to 120 bases, 70 bases to 120 bases, 80 bases to 120 bases, or 90 bases to 120 bases.

[0092] Hybrid capture probes may be added to a prepared sample and hybridized through a denature -reannealing process to form duplexes of exogenous-endogenous fragments (e.g. hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes are by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURE SELECT (Agilent). Thus, hybrid capture probes can be covalently attached to a biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solid support with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in some embodiments, the hybrid capture probes are immobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin.

[0093] If hybrid capture is used upstream of a next-generation sequencing reaction, one way to increase the number of reads that interrogate the position of interest is to decrease the length of the hybrid capture probe, as long as it does not result in bias in the underlying enriched alleles. The length of the hybrid capture probe should be long enough such that two hybrid capture probes designed to bind to two different target DNA sequences within the same DMR hybridize with near equal affinity to the target sequences. In certain embodiments, the use of shorter probes results in a greater chance that the hybrid capture probes bind to DNA molecular fragments from liquid samples, such as cfDNA. Thus, using hybrid capture, DNA molecules that include targeted DMRs in the DNA sample can be selectively enriched.

[0094] In some embodiments, target enrichment may involve Primer Extension Target Enrichment (PETE), an NGS hybridization capture technology that uses primer extension reactions to specifically capture and release target library molecules for sequencing. Commercial kits are available for performing PETE, including the KAPA HyperPETE kit form Roche.

[0095] In some embodiments, an optional PCR can be performed to amplify the captured DNA molecules comprising target DMRs.

[0096] In some embodiments, additional methods, including but not limited to MSRE methods, can be used to further enrich DNA fragments that are fully (i.e. 100%) methylated. In some embodiments, such further enrichment can be performed before a target enrichment by capture method disclosed herein. In some embodiments, such further enrichment can be performed after a target enrichment by capture method disclosed herein.Detection and Analysis

[0097] Methods as described herein typically include detecting and optionally quantifying nucleic acids, including DNA, cfDNA, adapted, bisulfite- or enzyme-converted, amplified cfDNA, enriched subsets of cfDNA having target DMRs, and in some embodiments, amplicons derived therefrom. Before detection and analysis of the sequences, the method described herein includes multiple steps. Non-limiting embodiments of the methods described herein are summarized in FIGs. 1 and 2. In some embodiments, cfDNA from a liquid sample (e.g., a blood sample) from the individual is analyzed. Not to be limited by theory, cfDNA is believed to be released fromcertain cells, such as cancer cells, for example when they undergo necrosis or apoptosis. In some embodiments, methods herein can be used to detect differentially methylated target regions or nucleic acid sequence of interest that is present in a small percentage of DNA in a sample, such as cfDNA, for example from a fetus, a cell from a donated organ, or in some embodiments, a cancer cell.

[0098] In methods herein, at least some of the subsets of adapted, converted, amplified cfDNA can be enriched using a set of hybrid capture probes to form at least some of the enriched subsets of adapted, converted, amplified cfDNA before detecting or quantifying an amount for at least some of the enriched subsets. In methods that include detecting or quantifying cfDNA fragments or their amplicons comprising target DMRs, such target fragments or their amplicons can be enriched using a set of hybrid capture probes or LTC PDP probes before detecting or quantifying.

[0099] In some embodiments, the method further comprises an additional amplification reaction that amplifies at least some of the enriched subsets of adapted, converted, amplified cfDNA comprising target DMRs for detecting or quantifying. The additional amplification reaction can be, for example, a quantitative PCR (qPCR reaction) such as TAQMAN assay (LIFE TECHNOLOGIES), or an INVADER assay (THIRD WAVE TECHNOLOGIES), a digital PCR, or any other method for detecting and / or quantifying a target DNA, which typically herein comprise a target region that includes one or more differentially methylated sequences.

[0100] In some embodiments, the additional amplification reaction is a clonal amplification reaction to form clonally amplified enriched subsets of adapted, converted, amplified cfDNA comprising target DMRs. For detecting or quantifying at least some of the amplified and / or selectively enriched target cfDNA fragments or their amplicons, methods herein can comprise i) performing an additional amplification reaction that amplifies at least some of the amplified target amplicons, wherein the additional amplification reaction is a clonal amplification reaction to form clonally amplified target amplicons, and ii) performing a next-generation sequencing reaction on the clonally amplified target amplicons. For detecting or quantifying at least some of the enriched subsets of adapted, converted, amplified cfDNA, methods herein can comprise i) performing an additional amplification reaction that amplifies at least some of the enriched subsets of adapted, converted, amplified cfDNA, wherein the additional amplification reaction is a clonalamplification reaction to form clonally amplified enriched subsets of adapted, converted, amplified cfDNA, and ii) performing a next-generation sequencing reaction on the clonally amplified enriched subsets of adapted, converted, amplified cfDNA. The detecting or quantifying included in methods herein, can comprise performing a sequencing reaction on the clonally amplified target region amplicons.

[0101] In some embodiments, the sequencing reaction is a next-generation sequencing (NGS) reaction. In methods herein, detecting or quantifying comprises counting sequence reads generated from clonally amplified target amplicons. Quantifying can also comprise determining a depth of read per target region for at least some of the target regions. Depth of read for each of the target region can be normalized relative to a depth of read for a normalization sequence. DNA sequences used for normalization can be derived from genomic DNA or from control plasmids, and can contain non-methylated, partially methylated or fully methylated sequences. DNA sequences used for normalization for quantitative methods herein, such as NGS, can be non-methylated, partially methylated or fully methylated sequences, and is typically fully methylated. The normalization sequence can be a 100% methylated contrived DNA molecule derived from a control genomic DNA sample or a pUC19 control plasmid. The normalization sequence can be derived from a lambda control plasmid having 100% methylated DNA. A fully methylated (100% methylated) synthetic sequence, for example, can also be considered as a normalization sequence. The normalization sequence can be a spike-in control DNA sample that is not subject to a conversion step. A spike-in control DNA sample can be a fully methylated genomic, plasmid, or synthetic DNA sample. A spike-in control sample used for normalization, in some embodiments, can be 0.05%, 0.1%, 0.2%, 0.5%, 0.7%, 0.8%, 1.0%, 1.2%, 1.4%, 1.6%, 1.8, 2%, or more fully methylated DNA by mass. A spike-in control sample can be fully methylated or non-methylated DNA that range between 0.005 to 2%, 0.01 to 2%, 0.05 to 2%, 0.1 to 2%, 0.5 to 2%, 1 to 2%, 0.005 to 1.5%, 0.005 to 1.2%, 0.005 to 1%, 0.005 to 0.8%, 0.005 to 0.5%, 0.005 to 0.3%, or 0.005 to 0.2% by weight. Methods herein, can include more than one, for example 2, 3, 4, 5, or more control or spike-in control sample. DNA sequencing techniques, particularly high throughput nextgeneration sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed in MYSEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LIFE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS FLEX+(ROCHE 454) etc.,can be used for quantitative measurements of the number of copies of a target region present after conversion, for example, but not limiting to, clonally amplified target region amplicons or clonally amplified enriched subsets of adapted, converted, amplified cfDNA, and thus provide quantitative information regarding the number and / or amount of methylation in sample DNA molecules, for example, cfDNA. High throughput genetic sequencers are amenable to the use of sample barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. Methods as described herein that utilize NGS detection, in some embodiments can have an average depth of read of at least 30x, 50x, lOOx, 200x, 500x, lOOOx, 2000x, 2900x, 3000x, 3500x, 4000x, 5000x, 10,000x, 50,000x, 75,000x, 100,000x, 130,000x, 150,000x, 175,000x, or 200,000x. In some embodiments, methods as described herein that utilize NGS detection can have an average depth of read of between 30x and 100,000x, between 50x and 50,000x, between lOOx and 10,000x, between lOOx and 5,000x, between 200x and 5,000x, or between 200x and l,000x.

[0102] In some embodiments, the adaptors or primers describe herein may comprise one or more molecular barcodes. Molecular barcodes or molecular indexing sequences may be used in next generation sequencing to reduce quantitative bias introduced by replication. In next generation sequencing, each nucleic acid fragment may be tagged with a molecular barcode or molecular indexing sequence. Sequence reads that have different molecular barcodes or molecular indexing sequences represent different original nucleic acid molecules. By referencing the molecular barcodes or molecular indexing sequences, PCR artifacts, such as sequence changes generated by polymerase errors that arc not present in the original nucleic acid molecules can be identified and separated from real variants / mutations present in the original nucleic acid molecules.

[0103] In some embodiments, molecular barcodes are introduced by ligating adaptors carrying the molecular barcodes to the isolated cfDNA to obtain adaptor-ligated and molecular barcoded DNA. In some embodiments, molecular barcodes are introduced by amplifying the adaptor- ligated DNA with primers carrying the molecular barcodes to obtain amplified adaptor- ligated and molecular barcoded DNA.

[0104] In some embodiments, the molecular barcoding adaptor or primers may comprise a universal sequence, followed by a molecular barcode region, optionally followed by a target specific sequence in the case of a primer. The sequence 5’ of molecular barcode may be used for subsequent PCR amplification or sequencing and may comprise sequences useful in the conversion of the amplicon to a library for sequencing. The random molecular barcode sequence could be generated in a multitude of ways. The preferred method synthesizes the molecule tagging adaptor or primer in such a way as to include all four bases to the reaction during synthesis of the barcode region. All or various combinations of bases may be specified using the IUPAC DNA ambiguity codes. In this manner the synthesized collection of molecules will contain a random mixture of sequences in the molecular’ barcode region. The length of the barcode region will determine how many adaptors or primers will contain unique barcodes. The number of unique sequences is related to the length of the barcode region as NLwhere N is the number of bases, typically 4, and L is the length of the barcode. A barcode of five bases can yield up to 1024 unique sequences; a barcode of eight bases can yield 65536 unique barcodes. In an embodiment, the DNA can be measured by a sequencing method, where the sequence data represents the sequence of a single molecule. This can include methods in which single molecules are sequenced directly or methods in which single molecules are amplified to form clones detectable by the sequence instrument, but that still represent single molecules, herein called clonal sequencing.

[0105] In some embodiments, the molecular barcodes described herein are Molecular Index Tags (“MITs”), which are attached to a population of nucleic acid molecules from a sample to identify individual sample nucleic acid molecules from the population of nucleic acid molecules (i.c. members of the population) after sample processing for a sequencing reaction. MITs are described in detail in U.S. Pat. No. 10,011,870 to Zimmermann et al., which is incorporated herein by reference in its entirety. Unlike prior art methods that relate to unique identifiers and teach having a diversity of unique identifiers that is greater than the number of sample nucleic acid molecules in a sample in order to tag each sample nucleic acid molecule with a unique identifier, the present disclosure typically involves many more sample nucleic acid molecules than the diversity of MITs in a set of MITs. In fact, methods and compositions herein can include more than 1,000, IxlO6, IxlO9, or even more starting molecules for each different MIT in a set ofMITs. Yet the methods can still identify individual sample nucleic acid molecules that give rise to a tagged nucleic acid molecule after amplification.

[0106] Methods herein can include analyzing data obtained from NGS or other sequencing techniques. In some embodiments of methods herein, clonally amplified target amplicons or clonally amplified enriched subsets of adapted, converted, amplified cfDNA can be subjected to sequencing using next-generation sequencing techniques. For a skilled artisan, algorithm design tools are available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.

[0107] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Fast and accurate long-read alignment with Burrows- Wheeler Transform. Bioinformatics.) on single end mode using pear merged reads to a reference, such as the hgl9 genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.

[0108] Methods herein can include a background error model that can be constructed using normal, or healthy liquid samples, in illustrative embodiments, normal, or healthy plasma samples, which arc scqucnccd on the same sequencing run to account for run-specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, or healthy liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g. plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of everygenomic loci, the depth of read (DOR) weighted mean and standard deviation of the error can be calculated.

[0109] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200- nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).

[0110] In some embodiments, the uniformity in DOR can be measured using standard methods such as, but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90tll-95thpercentile. For this measurement, the loci arc sorted in descending DOR order. In some embodiments, a DOR distribution using the 90th-95t11percentile contains 5 percent of reads. The reads of all loci between the 90thpercentile and 95thpercentile can be counted and divided by the total reads for all loci.

[0111] hr some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90th-95thpercentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1 .0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method or the amplification steps in the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90th-95thpercentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.

[0112] In some embodiments of methods herein, in addition to detecting differentially methylated sequences, the length of the cfDNA fragments and their amplicons comprising differentially methylated target sequences can be measured. For example, in some embodiments, the length of the cfDNA fragments and their amplicons comprising differentially methylated sequences may be measured using gel electrophoresis to separate the cfDNA samples according to size. For example, gel electrophoresis separates DNA molecules based on their size by applying an electric field to a gel, such as an agarose gel, upon which DNA molecules will move through the gel towards the positively charged anode. The size of the DNA molecules determines the speed by which the DNA molecules migrate through the gel. A standard mixture of DNA molecules with predetermined sizes is applied to the gel to identify the size of the DNA. In some embodiments, the size determination may be performed on an automated high-throughput gel electrophoresis system such as Pippin or Costal Genomics systems.

[0113] In some embodiments, the length of the cfDNA fragments and their amplicons comprising differentially methylated target sequences can be determined by analyzing the data generated from NGS. In some embodiments, the length is the average length of all fragments in a particular group. In some embodiments, the length is the mean length of fragments in a particular group. In some embodiments, the fragments median length is calculated for each group. In some embodiments, fragment length distribution is calculated for each group. In some embodiments, the mode of thefragment lengths is calculated for each group. In some embodiments, the interquartile range (IQR) of the fragment lengths is calculated for each group. In some embodiments, statistical learning methods are used to compare groups of hypermethylated fragments with undermethylated fragments for hypermethylated target regions as well as groups of hypomcthylatcd fragments with overmethylated fragments for hypomethylated target regions. In some embodiments, the fragment lengths of likely tumor-derived fragments are compared to the fragment lengths of likely nontumor derived fragments.

[0114] In some embodiments, cfDNA fragments in each sample or their amplicons comprising one or more target DMRs are classified into different groups based on their methylation levels. In some embodiments, a cfDNA fragment comprising one or more hypermethylated target regions having 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, or 12 or more CpG sites is classified as a “hypermethylated fragment” wherein 70% or more, 75% or more, 80% or more, 85% or more, or 90% or more of the CpG sites are methylated, and classified as an “undermethylated fragment” wherein 40% or less, 30% or less, 25% or less, 20% or less, or 15% or less of the CpG sites are methylated. In some embodiments, a cfDNA fragment comprising one or more hypomethylated target regions having 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, or 12 or more CpG sites is classified as an “overmethylated fragment” wherein 60% or more, 70% or more, 75% or more, 80% or more, or 85% or more of the CpG sites are methylated, and classified as a “hypomethylated fragment” wherein 30% or less, 25% or less, 20% or less, 15% or less, or 10% or less of the CpG sites are methylated. In some embodiments, hypermethylated fragments comprising hypermethylated target regions arc likely derived from tumor cells, whereas undermethylated fragments comprising hypermethylated target regions are likely derived from non-tumor cells. In some embodiments, hypomethylated fragments comprising hypomethylated target regions are likely derived from tumor cells, whereas overmethylated fragments comprising hypomethylated target regions are likely derived from non-tumor cells.

[0115] In some embodiments, statistical or machine learning methods are employed to use the fragment length difference as a feature for classifying samples. In some embodiments, a sample is classified as having ctDNA when the length of the hypermethylated fragments comprising hypermethylated target regions in or derived from the sample is significantly shorter than thelength of the undermethylated fragments in or derived from the sample comprising these target regions. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the length of the hypermethylated fragments in or derived from the sample comprising hypermethylated target regions is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the length of the undermethylated fragments in or derived from the sample comprising these target regions. In some embodiments, a sample is classified as having ctDNA when the length of the hypomethylated fragments in or derived from the sample comprising hypomethylated target regions is significantly shorter than the length of the overmethylated fragments in or derived from the sample comprising these target regions. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the length of the hypomethylated fragments in or derived from the sample comprising hypomethylated target regions is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the length of the overmethylated fragments in or derived from the sample comprising these target regions. In some embodiments, fragment lengths are calculated and compared in mean, median, and / or mode values. In some embodiments, fragment lengths are compared using statistical or machine learning methods.

[0116] In some embodiments, ctDNA is detected when the average or mean length of the hypermethylated sequences is shorter than the average or mean length of the undermethylated sequences. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypermethylated sequences is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides,or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the average or mean length of the undcrmcthylatcd sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypermethylated sequences is at least 10 nucleotides shorter than the average or mean length of the undermethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypermethylated sequences is at least 15 nucleotides shorter than the average or mean length of the undermethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypermethylated fragments is at least 20 nucleotides shorter than the average or mean length of the undermethylated sequences. In some embodiments, ctDNA is detected when the average or mean length of the hypomethylated fragments in or derived from the sample comprising hypomethylated target regions is significantly shorter than the length of the overmethylated fragments in or derived from the sample comprising these target regions. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypomethylated sequences is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the average or mean length of the overmethylated fragments. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypomethylated sequences is at least 10 nucleotides shorter than the average or mean length of the overmethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypomethylated sequences is at least 15 nucleotides shorter than the average or mean length of the overmethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the average or mean length of the hypomethylated fragments is at least 20 nucleotides shorter than the average or mean length of the overmethylated sequences.

[0117] In some embodiments, ctDNA is detected when the mode length of the hypermethylated sequences is shorter than the mode length of the undermethylated sequences. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypermethylated sequences is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the mode length of the undermethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypermethylated sequences is at least 10 nucleotides shorter than the mode length of the undermethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypermethylated sequences is at least 15 nucleotides shorter than the mode length of the undermethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypermethylated fragments is at least 20 nucleotides shorter than the mode length of the undermethylated sequences. In some embodiments, ctDNA is detected when the mode length of the hypomethylated fragments in or derived from the sample comprising hypomethylated target regions is significantly shorter than the length of the overmethylated fragments in or derived from the sample comprising these target regions. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypomethylated sequences is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 1 1 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than the mode length of the overmethylated fragments. In some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypomethylated sequences is at least 10 nucleotides shorter than the mode length of the overmethylated sequences. In some embodiments, a patient sample is identified as positive for ctDNA when the mode length of the hypomethylated sequences is at least 15 nucleotides shorter than the mode length of the overmethylated sequences. In some embodiments,a patient sample is identified as positive for ctDNA when the mode length of the hypomethylated fragments is at least 20 nucleotides shorter than the mode length of the overmethylated sequences.

[0118] In some embodiments, a sample is classified as having ctDNA when the length of the hypermethylated fragments comprising hypermethylated target regions in or derived from the sample is significantly shorter than the length of fragments comprising these target regions derived from one or more healthy samples. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the length of the hypermethylated fragments in or derived from the sample comprising hypermethylated target regions is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 11 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than a reference value representing or reflecting the length of fragments comprising these target regions derived from one or more healthy samples. In some embodiments, a sample is classified as having ctDNA when the length of the hypomethylated fragments comprising hypomethylated target regions in or derived from the sample is significantly shorter than the length of fragments comprising these target regions derived from one or more healthy samples. For example, in some embodiments, a patient sample is identified as positive for ctDNA when the length of the hypomethylated fragments in or derived from the sample comprising hypomethylated target regions is at least 8 nucleotides, or at least 9 nucleotides, or at least 10 nucleotides, or at least 1 1 nucleotides, or at least 12 nucleotides, or at least 13 nucleotides, or at least 14 nucleotides, or at least 15 nucleotides, or at least 16 nucleotides, or at least 17 nucleotides, or at least 18 nucleotides, or at least 19 nucleotides, or at least 20 nucleotides, or at least 21 nucleotides, or at least 22 nucleotides, or at least 23 nucleotides, or at least 24 nucleotides, or at least 25 nucleotides, shorter than a reference value representing or reflecting the length of fragments comprising these target regions derived from one or more healthy samples. In some embodiments, fragment lengths are calculated and compared in mean, median, and / or mode values. In some embodiments, fragment lengths are compared using statistical or machine learning methods.

[0119] In some embodiments of methods herein, in addition, or, in some embodiments, as an alternative to analyzing an altered (increased or decreased) methylation levels in a sample, and length of differentially methylated samples, one or more other factors can be analyzed if desired. These factors can be used to increase the accuracy of the diagnosis (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer, or staging the cancer) or prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject.

[0120] In some embodiments, the methods described herein can be combined with other various detection methods, such as genetic mutation (SNV, indel, CNV, fusion, etc.) based panels to increase specificity.Limits of Detection

[0121] Exemplary methods herein are to detect small amounts of ctDNA fragments in a sample based on the methylation level and fragment length of the cfDNA fragments from the sample comprising target DMRs. cfDNA fragments comprising the target DMRs can be selectively enriched using methods herein. Exemplary methods herein, in some embodiments, have a limit of detection of as low as 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% wherein the method is capable of detecting fully methylated DNA molecules present at 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% (by mass) or more in a mixture of DNA molecules. In some embodiments, methods herein are capable of differentiating samples having 1.0%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, or 0.001% (by mass) or more contrived fully methylated DNA molecules from samples having 0% of contrived fully methylated DNA molecules.

[0122] In some embodiments, measurements can be adjusted for bias, such as bias due to differences in amplification efficiency or adjusted for sequencing errors. In some embodiments, differentiation between methylated and unmethylated samples can be analyzed using a control sample that is fully methylated, for example a fully methylated plasmid control (e.g. pUC19) for normalization of quantitative results. In some embodiments, the differentiation between the methylated samples recited in this paragraph can be achieved after normalization of a detected and typically quantified signal using one or more (e.g. 2, 3, 4, 5, or 6) controls that are not methylated.

[0123] In certain embodiments, ctDNA is detected when hypermethylated fragments comprising hypermethylated target regions are detected and determined to be shorter than undermethylated fragments comprising these target regions from the same sample or a reference value representing or reflecting the length of fragments comprising these target regions derived from one or more healthy samples. In certain embodiments, ctDNA is detected when hypomethylated fragments comprising hypomethylated target regions are detected and determined to be shorter than overmethylated fragments comprising these target regions from the same sample or a reference value representing or reflecting the length of fragments comprising these target regions derived from one or more healthy samples. In certain embodiments, the method is capable of detecting ctDNA when it is present in 50%, 45%, 40% 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2.5%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001% or less of the circulating free DNA in a sample. In some embodiments, methods herein detect or are capable of detecting ctDNA from a sample when it is present at a range between 0.001% on the low end and 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2.5%, 1%, or 0.5% on the high end of the range of total cfDNA in the sample. Methods herein, in some embodiments, detect or are capable of detecting 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001% or less circulating tumor DNA, as a percent of total circulating free DNA (cfDNA) from a sample. The percentage of ctDNA in a cfDNA sample can be determined or approximated by variant allele frequency (VAF) of one or more DNA variants, such as single nucleotide variants (“SNVs”), present in tumor cells but not in normal cells. (See e.g., W02019200228, incorporated by reference herein in its entirety). VAF can be used as a surrogate measure of the percentage of ctDNA in the cfDNA sample.Comparison of Two or More Samples

[0124] In some embodiments, methods described herein can include comparing two or more samples. In some embodiments, methods as described herein can include comparing 3, 4, 5, 6, 7, 8, 9, 10 or more samples. In some embodiments, methods as described herein can include comparing two samples, such methods further comprise performing the method on a second sample, and wherein comparing the length of the fragments in the differentially methylated fractions further comprises comparing the length of the fragments in the differentially methylated fractions from the first sample to the length of the fragments in the differentially methylated fractions from a second sample. In some embodiments, the second sample is from a secondsubject. In some embodiments, the second sample is derived from a cell line. In some embodiments, the method is performed on the first and second samples simultaneously. In some embodiments, the first sample is from a subject suspected or at risk of having a disease, and wherein the second sample is from a subject that is not suspected or at risk of having the disease. In some embodiments, the second sample is collected at a first time point and the first sample is collected at a second time point. In some embodiments, the first time point precedes the second time point. In some embodiments, the determining comprises comparing the length of the fragments in the differentially methylated fractions derived from the first sample to the length of the fragments in the differentially methylated fractions derived from a second sample using the same method.Diagnosis and Monitoring of Cancer Patients

[0125] The methods described herein are useful for both the diagnosis and monitoring of cancer patients. Blood samples may be taken from subjects having an unknown cancer status and processed and analyzed using methods described herein. Subjects having a size differential between the length of hypermethylated or hypomethylated fragments, compared with the undermethylated or overmethylated, respectfully, are diagnosed as at risk for or as having cancer.

[0126] Further samples may be taken from subjects diagnosed as at risk for or as having cancer at a later time point to monitor the subject post treatment.Kits

[0127] The present disclosure also provides kits for determining the presence of ctDNA in one or more samples.

[0128] The kit may further comprise one or more of: wash buffers and / or reagents, hybridization buffers and / or reagents, labeling buffers and / or reagents, dilution buffers and / or reagents, and detection means. The buffers and / or reagents are usually optimized for the particular amplification / detection technique for which the kit is intended. Protocols for using these buffers and reagents for performing different steps of the procedure may also be included in the kit.

[0129] The kit additionally may comprise an assay definition scan card and / or instructions such as printed or electronic instructions for using the oligonucleotides in an assay. In some embodiments, a kit comprises an amplification reaction mixture or an amplification master mix. Reagents included in the kit may be contained in one or more containers, such as a vial.WORKING EXAMPLES

[0130] In general, the following Examples were performed using the workflow described in FIG. 1 and FIG. 2, the individual steps of which are described below.

[0131] Briefly, samples were collected and processed as described previously to purify cfDNA. Size selection was performed and the size-selected cfDNA was blunted, A-tailed, and ligated with a fully methylated adaptor comprising a universal binding site. After ligation with methylated adapters, adapter-ligated DNA were subjected to bisulfite conversion using the Zymo EZ DNA Methylation Gold (BS-G) kit. Following the conversion, the library was universally amplified and purified. Thereafter, barcoding PCR was performed to add sample barcodes.

[0132] Target DMRs were selected from The Cancer Genome Atlas (TCGA) consortium and inhouse discovery prioritizing differentially methylated CpG targets that can distinguish between colorectal cancer (CRC) and normal samples. A custom hybrid capture panel containing probe sets targeting 860 human DMR targets for CRC spanning 112 kb was used. The probes were designed to perfectly complement either completely methylated targets or completely unmethylated targets. Hybrid capture was performed using IDT xGcn hybridization capture protocol with a 20-25% tolerance for imperfect complementarity to capture fully methylated, fully unmethylated, and partially methylated targets.

[0133] Post capture PCR was then performed to re-amplify captured material. Amplified, captured DNA was then purified with and the concentration quantified. Captured libraries were paired-end sequenced at 2x150 bp on Novaseq platform.EXAMPLE 1 - Analysis of Non-Tumor Sample

[0134] Differentially methylated target sequences were analyzed in blood samples taken from 50 healthy donors. Samples were processed and sequenced using the methods described herein.

[0135] Data for three exemplary samples is shown in FIGs. 3A-3C. Fragment length distribution is shown for both hypermethylated and undermethylated sequences. Table 1 below shows the mode, median and interquartile range of the hypermethylated and undermethylated sequences for each of the three samples.Table 1. Length distribution of hypermethylated vs. undermethylated fragments of hypermethylated targets in healthy samples.

[0136] For this analysis, only sequence reads comprising targets known to be hypermethylated in tumor cells and having 8 or more CpGs are included. A read is classified as a “hypermethylated fragment” if 90% or more CpGs are methylated. A read is classified as an “undermethylated fragment” if 20% or less CpGs are methylated.

[0137] Data shows that hypermethylated fragments typically have a similar fragment length distribution as compared to undermethylated fragments in healthy samples.EXAMPLE 2 - Analysis of Tumor Samples

[0138] Differentially methylated target sequences were analyzed in blood samples taken from 60 donors having a colorectal cancer diagnosis. The samples were processed and sequenced using the methods described herein.

[0139] Data for two exemplary samples are shown in FIGs. 4A-4B. Fragment length distribution is shown for both hypermethylated and undermethylated sequences. Table 2 below shows the mode, median and interquartile range of the hypermethylated and undermethylated sequences for each of the samples.Table 2. Length distribution of hypermethylated vs. undermethylated fragments of hypermethylated targets in CRC samples.

[0140] For this analysis, only sequence reads comprising targets known to be hypermethylated in tumor cells and having 8 or more CpGs are included. A read is classified as a “hypermethylated fragment” if 90% or more CpGs are methylated. A read is classified as an “undermethylated fragment” if 20% or less CpGs are methylated.

[0141] Data shows that hypermethylated fragments are typically shorter than undermethylated fragments in cancer derived samples.

[0142] Comparing data from healthy and cancer samples revealed that a size differential exists between the length of hypermethylated and undermethylated fragments in cancer samples that does not exist in healthy samples. For example, 24% of healthy samples have a HyMeFF (hypermethylated fragment fraction, i.e., the fraction of hypermethylated fragments across all hypermethylated targets in a sample) greater than 2e-4 but only 10% also have median fragment length difference of greater than 10 bases. On the other hand, 96% of CRC samples have a HyMeFF greater than 2e-4 and 84% also have a median fragment length difference of greater than 10 bases.* * * *

Claims

CLAIMSWhat is claimed is:

1. A method of preparing deoxyribonucleic acid (DNA) useful for detecting circulating tumor DNA (ctDNA), said method comprising:(a) obtaining cell-free DNA (cfDNA) from a sample from a subject;(b) selectively enriching subsets of the cfDNA from (a) or derivatives therefrom having one or more target regions to obtain enriched DNA, wherein the target regions are differentially methylated in cancer;(c) sequencing the enriched DNA from (b) to obtain sequence reads; and(d) partitioning a plurality of the sequence reads into two or more groups based on their methylation status and determining the fragment length distribution of the sequence reads in at least one of the groups.

2. The method of claim 1, wherein the two or more groups comprises a hypermethylated group and an undermethylated group.

3. The method of claim 1 or claim 2, wherein the two or more groups comprises a hypomethylated group and an overmethylated group.

4. The method of any of claims 1-3, wherein selectively enriching subsets of the cfDNA is performed by hybrid capture using a set of hybrid capture probes.

5. The method of claim 4, wherein each probe of the set is designed to hybridize to one or more target regions.

6. The method of any of claims 1-5, further comprising ligating adapters to the cfDNA to obtain adapter-ligated DNA.

7. The method of claim 6, wherein the adapters are methylated adapters.

8. The method of claim 6 or 7, wherein the adapters arc Y adapters.

9. The method of any one of claims 6-8, wherein the adapters each comprise a universal priming site.

10. The method of any of claims 6-9, wherein the adapters each further comprise a molecular barcode.

11. The method of claim 10, wherein the number of adapters having different molecular barcodes is between 10 to 1,000, and wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1,000:1.

12. The method of claim 11, wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1.

13. The method of any of claims 6-12, further comprising performing bisulfite conversion on the adapter-ligated DNA, thereby generating bisulfite-converted DNA.

14. The method of any of claims 6-12, further comprising treating the adapter-ligated DNA with sodium bisulfate, thereby generating bisulfite-converted DNA.

15. The method of any of claims 6-12, further comprising performing enzymatic conversion on the adapter-ligated DNA, thereby generating enzymatic-converted DNA.

16. The method of claim 15, wherein enzymatic conversion comprises treating the adapter- ligated DNA with TET2, T4-bGT, and APOBEC.

17. The method of any of claims 13-16, further comprising amplifying the bisulfite- or enzyme-converted DNA.

18. The method of claim 17, wherein the amplifying comprises one or more PCRs using universal primers.

19. The method of any of claim 15 or 18, where amplification comprises barcoding PCR.

20. The method of any of claims 2-19, wherein the hypermethylated group comprises sequence reads of cfDNA fragments each having 2 or more CpG sites, wherein 70% or more of the CpG sites are methylated.

21. The method of any of claims 3-20, wherein the hypomethylated group comprises sequence reads of cfDNA fragments each having 2 or more CpG sites, wherein 30% or less of the CpG sites are methylated.

22. The method of any of claims 2-21, wherein a fragment length distribution of the sequence reads in the hypermethylated group of at least 5 bases shorter than the fragment length distribution of the sequence reads in the undermethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject.

23. The method of any of claims 2-22, wherein a fragment length distribution of the sequence reads in the hypermethylated group of at least 10 bases shorter than the fragment length distribution of the sequence reads in the undermethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject.

24. The method of any of claims 3-23, wherein a fragment length distribution of the sequence reads in the hypomethylated group of at least 5 bases shorter than the fragment length distribution of the sequence reads in the overmethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject.

25. The method of any of claims 3-24, wherein a fragment length distribution of the sequence reads in the hypomethylated group of at least 10 bases shorter than the fragment length distribution of the sequence reads in the overmethylated group indicates the presence of ctDNA in the sample and / or the presence of cancer in the subject.

26. The method of any of claims 2-25, wherein a length of the sequence reads in the hypermethylated group approximately equal to the length of the sequence reads in the undermethylated group indicates the absence of ctDNA in the sample and / or the absence of cancer in the subject.

27. The method of any of claims 3-26, wherein a length of the sequence reads in the hypomethylated group approximately equal to the length of the sequence reads in the overmethylated group indicates the absence of ctDNA in the sample and / or the absence of cancer in the subject.

28. A method of preparing deoxyribonucleic acid (DNA) useful for detecting circulating tumor DNA (ctDNA), said method comprising:(a) obtaining cell-free DNA (cfDNA) from a sample from a subject;(b) contacting the cfDNA from (a) or derivatives therefrom with one or more restriction enzymes capable of binding to one or more DNA sequences containing one or more CpG sites and thereby generating treated cfDNA;(c) selectively enriching subsets of the treated cfDNA from (b) or derivatives therefrom having one or more target regions to obtain enriched DNA, wherein the target regions are differentially methylated in cancer;(d) sequencing the enriched DNA from (c) to obtain sequence reads; and(e) determining the fragment length distribution of the sequence reads in the enriched DNA.

29. The method of claim 28, further comprising ligating adapters to the cfDNA to obtain adapter-ligated DNA.

30. The method of claim 29, wherein the adapters are methylated adapters.

31. The method of claim 29 or 30, wherein the adapters are Y adapters.

32. The method of any one of claims 29-31, wherein the adapters each comprise a universal priming site.

33. The method of any of claims 29-32, wherein the adapters each further comprise a molecular barcode.

34. The method of claim 33, wherein the number of adapters having different molecular barcodes is between 10 to 1,000, and wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 1,000:1.

35. The method of claim 34, wherein the ratio of the total number of cfDNA molecules to the number of different molecular barcodes in the ligation reaction is at least 10,000:1.

36. The method of any of claims 29-35, wherein the one or more restriction enzymes are one or more methylation sensitive restriction enzyme (MSREs).

37. The method of any of claims 29-35, wherein the one or more restriction enzymes arc one or more methylation dependent restriction enzyme (MDREs).

38. The method of any of claim 36 or 37, further comprising amplifying the treated DNA.

39. The method of claim 38, wherein the amplifying comprises one or more PCRs using universal primers.

40. The method of any of claim 38 or 40, where amplification comprises barcoding PCR.

41. The method of any of claims 36-40, wherein a fragment length distribution of the sequence reads relative to a threshold is used to classify the likelihood of presence or absence of cancer in the subject.

42. The method of any of claims 1-41, further comprising performing size selection on the cfDNA.

43. The method of claim 42, wherein size selection comprises enriching the cfDNA for molecules that are between 70 and 500 base pairs in length.

44. The method of claim 42 or 43, wherein size selection comprises enriching the cfDNA for molecules that are between 100 and 200 base pairs in length.

45. The method of any of claims 42-44, wherein size selection comprises enriching the cfDNA for molecules that are between 130 and 170 base pairs in length.

46. The method of any of claims 1-45, wherein the fragment length distribution is calculated as an average or mean length of the fragments in each group.

47. The method of any of claims 1-46, wherein the sample is a liquid sample.

48. The method of any of claims 1-47, wherein the sample is a blood, plasma, serum, or urine , vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample.

49. The method of any of claims 1-48, wherein the sample comprises DNA from a tumor.

50. The method of any of claims 1-49, further comprising repeating the method on a second sample.

51. The method of claim 50, wherein the first sample is collected at a first time point and the second sample is collected at a second time point.

52. The method of claim 52 or 53, wherein the first sample and the second sample are collected from the same subject.

53. The method of claim 54, wherein the subject is suspected or at risk of having a disease.

54. The method of claim 52 or 53, wherein the first sample and the second sample are from different subjects.

55. The method of claim 54, wherein the first subject is suspected or at risk of having a disease, and wherein the second subject is not suspected or at risk of having the disease.

56. The method of any of claims 53-55, wherein the disease is a cancer.

57. The method of claim 56, wherein the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer.

58. The method of any of claims 1-57, wherein the method is capable of detecting fully methylated DNA molecules present at 1 .0% (by mass) or more in a mixture of DNA molecules.

59. The method of any of claims 1-58, wherein circulating tumor DNA (ctDNA) is present in 1% or more of the total circulating free DNA (cfDNA) in the sample.

60. The method of any of claims 1-58, wherein circulating tumor DNA (ctDNA) is present in 0.1% or more of the total circulating free DNA (cfDNA) in the sample.

61. The method of any of claims 1-58, wherein circulating tumor DNA (ctDNA) is present in 0.01% or more of the total circulating free DNA (cfDNA) in the sample.