Methods for maintaining methylation in amplification of methylated nucleic acid molecules
The method of extracting and copying nucleic acid molecules with methyltransferase preserves methylation status, addressing the loss of signatures during amplification, enabling sensitive and efficient quantification and analysis.
Patent Information
- Application Number
- PCT/US2025/036655
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-07
- Publication Date
- 2026-01-15
AI Technical Summary
Current methods for preparing and quantifying small amounts of nucleic acid molecules, such as circulating free DNA (cfDNA) or circulating free RNA (cfRNA), result in the loss of methylation signatures due to degradation during amplification, hindering the preservation of methylation status for diagnostic and research purposes.
A method for preparing nucleic acid molecules that involves extracting DNA, copying template molecules, transferring methylation status using a methyltransferase, and repeating these steps to generate a library of amplified DNA molecules with preserved methylation signatures, utilizing isothermal extension and adapters with universal priming sites.
Maintains methylation signatures during amplification, enabling multiple downstream analyses like methylation analysis and cancer detection, while improving the sensitivity and efficiency of quantifying differentially methylated nucleic acid molecules.
Smart Images

Figure US2025036655_15012026_PF_FP_ABST
Abstract
Description
METHODS FOR MAINTAINING METHYLATION IN AMPLIFICATION OF METHYLATED NUCLEIC ACID MOLECULESTECHNICAL FIELD
[0001] The present disclosure relates generally to methods for preparing and analyzing nucleic acid molecules, such as deoxyribonucleic acid (DNA) molecules, and, more particularly, to methods for preparing and analyzing nucleic acid molecules that include one or more methylation sites.BACKGROUND
[0002] Currently there is a need for improved methods for preparing and quantifying small amounts of nucleic acid molecules in samples, such as circulating free DNA (cfDNA) or circulating free RNA (cfRNA) carrying epigenetic alterations, such as changes in methylation status or patterns. There is a need for sensitive and efficient methods for preparing and quantifying small amounts of differentially methylated nucleic acid molecules of interest in a sample. Unfortunately, amplification of the nucleic acid molecules results in the loss of methylation signatures. While the nucleic acid molecules can be treated chemically (such as with bisulfite (BS)) or enzymatically before amplification to generate replicable information reflecting methylation status, such treatments tend to degrade a large fraction of input molecules prior to amplification, resulting in low molecular recovery of these molecules.
[0003] Patient derived cfDNA sample libraries are an invaluable source of highly annotated samples that are available for screening or diagnostics as well as research and development. These libraries in essence immortalize the original cfDNA samples, so that they can be available for additional analysis and new uses. Furthermore, current and future discovery projects focused on methylation signatures will compete for precious plasma samples with other analysis modalities and other projects. Sample availability is one of the most severe bottlenecks to the development of molecular screening or diagnostic tests, such as early cancer detection and tumor-naive MRD tests, and the ability to preserve such sample cohorts for multiple uses is imperative for phenotype screening, diagnostics and research.
[0004] Accordingly, there is a need for improved methods for preparing and amplifying differentially methylated DNA, including differentially methylated cfDNA, whilst maintaining the methylation status.SUMMARY
[0005] To overcome the above-mentioned and additional problems in the art, the present disclosure provides methods for preparing a preparation of nucleic acid molecules (such as DNA) from a subject useful for determining a methylation status of a genomic region of interest by generating copies of the nucleic acid molecules including the transfer of the methylation status of the parent to daughter DNA or cDNA molecules.
[0006] Provided herein in one aspect is a method for preparing a non-naturally occurring preparation of DNA molecules useful for determining a methylation status of a genomic region of interest, said method comprising: (a) extracting DNA from a sample of a subject; (b) copying one or more template DNA molecules from the extracted DNA or its derivative, thereby generating one or more copied DNA molecules; (c) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules, thereby generating one or more methylation-transferred DNA molecules; and (d) optionally repeating steps (b) and (c) one or more times, thereby generating a library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules.
[0007] In some embodiments of the methods described herein, copying one or more template DNA molecules from the extracted DNA or its derivative, thereby generating one or more copied DNA molecules is performed by isothermal extension. In some embodiments, the isothermal extension is performed at about 58-64 °C. In some embodiments, the isothermal extension is performed at about 60 °C.
[0008] In some embodiments of the methods described herein, the methylation status is transferred using a methyltransferase. In some embodiments the methylation status of at least one of the one or more template DNA molecules is preserved in one or more of the methylation- transferred DNA molecules. In some embodiments, the template and copied DNA molecules are treated with the methyltransferase at about 30-40 °C. In some embodiments, the template and copied DNA molecules are treated with the methyltransferase at about 35 °C. In someembodiments, steps (b) and (c) are repeated for 1-20 cycles. In some embodiments, steps (b) and (c) are repeated for 1-10 cycles. In some embodiments, wherein steps (b) and (c) are repeated for 1-5 cycles. In some embodiments, the copying is performed using a DNA polymerase having reduced activity below 37 °C. In some embodiments, the reduced activity is below 15%. In some embodiments, the polymerase is BST polymerase or derivative thereof. In some embodiments, is a Taq polymerase or derivative thereof. In some embodiments, the methyl transferase is DNMT1 or derivative thereof. In some embodiments, step (b) is performed for a shorter time than step (c). In some embodiments, step (b) is performed for between 1 and 10 minutes. In some embodiments, step (c) is performed for between 5 and 60 minutes.
[0009] In some embodiments of the methods described herein, the method further comprises appending adapters to the extracted DNA prior to the copying in step (b). In some embodiments, the adapters are Y-adapters. In some embodiments, the adapters comprise a universal priming site. In some embodiments, the adapters comprise methylated cytosines. In some embodiments, the adapters comprise at least one methylated cytosine and at least one unmethylated cytosine. In some embodiments, the adapters comprise a molecular barcode. In some embodiments, the molecular barcode is used to construct a methylation consensus. In some embodiments, the molecular barcode is used to identify an error in isothermal amplification, methyltransferase treatment, or bisulfite or enzymatic conversion. In some embodiments, the adapters do not include any recognition site for any methylation sensitive restriction enzymes (MSRE) or methylation dependent restriction enzymes (MDRE). In some embodiments, the isothermal amplification is performed with at least one primer that binds to the universal priming site. In some embodiments, the primer comprises locked nucleic acid (LN A) oligonucleotides. In some embodiments, the primer comprises 5’ linked primers. In some embodiments, the sample DNA molecules are fragmented to form fragmented DNA molecules before appending the adapters.
[0010] In some embodiments of the methods described herein, the method further comprises contacting at least a portion of the library of amplified DNA molecules or their derivatives with a deaminating agent to generate treated DNA molecules. In some embodiments, the deaminating agent is sodium bisulfite. In some embodiments, the deaminating agent is an enzyme that converts cytosine to uracil.
[0011] In some embodiments of the methods described herein, the method further comprises contacting at least a portion of the library of amplified DNA molecules with one or more MSREs and / or MDREs to generate treated DNA. In some embodiments, the one or more MSREs are selected from one or more of Hpall, Hhal, HpyCH41V, or BstUl. In some embodiments, the one or more MDREs are selected from one or more of AbaSI, Glal, McrBC, or MspII. In some embodiments, the contacting comprises contacting with two or more MSREs and / or MDREs. In some embodiments, the contacting comprises contacting with two or more MSREs and / or MDREs in a single reaction. In some embodiments, the contacting comprises contacting with three or more MSREs and / or MDREs. In some embodiments, the contacting comprises contacting with four or more MSREs and / or MDREs.
[0012] In some embodiments of the methods described herein, the method further comprises enriching a plurality of target loci in the converted or treated DNA. In some embodiments, the converted or treated DNA is enriched using hybrid capture probes to generate enriched DNA. In some embodiments, the converted or treated DNA is enriched using PCR to generate enriched DNA. In some embodiments, the PCR is targeted multiplex PCR. In some embodiments, the target loci comprise 50-50,000 target loci. In some embodiments, the target loci are differentially methylated in cancer.
[0013] In some embodiments of the methods described herein, the method further comprises performing high-throughput sequencing on the enriched DNA to generate sequencing data. In some embodiments the method further comprises using the sequencing data to determine the methylation status of the plurality of target loci. In some embodiments the method further comprises using the sequencing data to determine the presence of somatic mutations.
[0014] In some embodiments of the methods described herein, the sample is a liquid sample. In some embodiments, the liquid sample is a blood, serum, plasma, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample. In some embodiments, the liquid sample is a blood, plasma, serum, or urine sample. In some embodiments, the sample is a blood sample or a derivative sample thereof. In some embodiments, the sample is a plasma sample. In some embodiments, the sample comprises DNA from a tumor. In some embodiments,the sample is a sample from a cancer tissue. In some embodiments, the sample comprises DNA from a transplanted organ. In some embodiments, the sample comprises DNA from a fetus.
[0015] In some embodiments, the subject is suspected or at risk of having a disease. In some embodiments, the disease is a cancer. In some embodiments, the cancer is selected from ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, nonsmall cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular’ lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low- grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia. In some embodiments, the cancer is selected from a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node,malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whippie resection. In some embodiments, the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer. In some embodiments, the methylation status of the plurality of target loci is indicative of the presence or absence of the cancer. In some embodiments, the presence of somatic mutations is indicative of the presence or absence of the cancer.
[0016] In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject comprising an organ from another individual.
[0017] In some embodiments of the methods described herein, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 70 and 500 base pairs in length. In some embodiments, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 100 and 200 base pairs in length. In some embodiments, the method further comprises enriching the sample DNA molecules for DNA molecules that are between 130 and 170 base pairs in length.
[0018] In some embodiments of the methods described herein, the plurality of target loci each comprises a set of two or more CpG sites that are differentially methylated in one or more cancers. In some embodiments, the target loci each comprises three or more CpG sites that are differentially methylated in a cancer.
[0019] In some embodiments of the methods described herein, one or more steps arc performed in a microfluidic device or in an emulsion-split.
[0020] Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are for ease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, and kits or functional elements therein across sections. Further details regarding aspects and embodiments of the present disclosure are provided throughout this patent application. Sections and section headers are forease of reading and are not intended to limit combinations of disclosure, such as methods, compositions, or other functional elements therein across sections.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG. 1 illustrates an exemplary, non-limiting workflow for preparing methylated targeted DNA molecules, comprising isothermal amplification and methylation transfer.
[0022] FIG. 2 illustrates an exemplary, non-limiting workflow for preparing methylated targeted DNA molecules, utilizing Y-adapters.
[0023] FIG. 3 illustrates an exemplary, non-limiting workflow for preparing methylated targeted DNA molecules, utilizing Y-adapters and strand invasion.
[0024] FIG. 4 illustrates an exemplary, non-limiting workflow for preparing methylated targeted DNA molecules, utilizing linked primers and enhanced strand invasion.
[0025] FIG. 5 illustrates an exemplary, non-limiting workflow for preparing methylated targeted single stranded DNA (ssDNA) molecules, comprising isothermal amplification and methylation transfer.
[0026] FIGs. 6A and B illustrate the propagation of errors after conversion and amplification events using methods that do not copy the template DNA molecules prior to conversion. (A) shows the process in the absence of conversion errors and (B) shows the process with conversion and PCR errors occurring. Y and Z represent molecular indexes (MITs) attached to two different target DNA molecules. C represents unmethylated cytosine and mC represents methylated cytosine. Red residues represent errors made during conversion and amplification.
[0027] FIGs. 7 A and B illustrate the propagation of errors after conversion and amplification events using methods of the invention described herein that copy the template DNA molecules prior to conversion. (A) shows the process in the absence of conversion / amplification errors and (B) shows the process with conversion / amplification errors occurring. Y and Z represent MITs attached to two different target DNA molecules. C represents unmethylated cytosine and mCrepresents methylated cytosine. Red residues represent errors made during conversion and amplification.
[0028] FIG. 8 shows the improvement on library conversion resulting from the methods described herein.
[0029] FIG. 9 shows a simulation of the methylation conversion rate using the MTase Amplification Model described herein.
[0030] FIG. 10 shows a simulation of the protection rate (a measure of methylation conversion specificity) using the MTase Amplification Model described herein.
[0031] FIG. 11 illustrates the digestion of adapted unmethylated (top), hemimethylated (middle), and methylated (bottom) DNA with the MSRE Sall.
[0032] FIG. 12 illustrates amplification by Bst2.0 WarmStart polymerase at various time / temperature combinations. 150 bp constructs with Y-adapters were amplified using universal primers in the presence of Bst2.0 WarmStart polymerase. A band at 150 bp demonstrates successful amplification.
[0033] FIG. 13. illustrates amplification by Bst2.0 WarmStart polymerase in the presence of DNMT1 to confirm Bst activity in a “1-pot” reaction. 150 bp constructs with Y-adapters were amplified using universal primers in the presence of Bst2.0 WarmStart polymerase. A band at 150 bp demonstrates successful amplification.
[0034] FIG. 14 illustrates Sall digestion after a 1-pot Bst2.0 WarmStart polymerase and DNMT1 reaction to assess DNMT1 activity. The assay was performed on both methylated and unmethylated 150 bp constructs with Y-adapters.
[0035] FIG. 15 summarizes results of the 1-pot reaction after Sall digestion. Positive and negative controls arc shown to demonstrate successful methyl transfer of the methylated template and protection of the sample from Sall digestion.
[0036] While the above- identified drawings set forth presently disclosed embodiments, other embodiments are also contemplated, as noted in the discussion. This disclosure presentsillustrative embodiments by way of representation and not limitation. Numerous other modifications and embodiments can be devised by those skilled in the art which fall within the scope and spirit of the principles of the presently disclosed embodiments.DETAILED DESCRIPTION
[0037] Both global and local epigenetic changes are widely regarded as a hallmark of diseases, including cancer, or other phenotypes (such as age) and their progression. In cancer, for example, alterations include a global decrease in overall CpG methylation levels coupled with discrete regions of hypermethylation, typically in CpG islands located in the promoter regions of tumor suppressor genes. Hypermethylation has been associated with cancer progression and the silencing of growth regulating genes and tumor suppressor genes. As a result, a growing number of DNA methylation biomarkers are being utilized in the development of novel assays for monitoring cancer progression, treatment response and early detection. It is possible to detect the presence of cancer through the analysis of the methylation status of specific CpG sites that are predominantly methylated in DNA from tumor cells, including circulating tumor DNA (ctDNA) released from tumor cells.
[0038] The present disclosure addresses many long-felt needs and long-standing problems in the ail, such as, but not limited to, those mentioned in the Background section herein. For example, methods are provided herein for preparing DNA molecules useful for determining a methylation status of a genomic region of interest. Such methods can be used to amplify methylated DNA molecules of interest, whilst maintaining the methylation signatures of the original template. For example, the methods provided herein comprise steps including a) extracting template DNA molecules from a subject; b) copying said template DNA molecules; c) transferring the methylation status from the template DNA molecules to the copied DNA molecules; and d) repeating steps b) and c) to generate a library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules. The libraries can then be used for multiple downstream analyses, including methylation analysis, somatic mutation analysis, fragmentomics, non-invasive prenatal testing (NIPT), cancer detection / diagnosis, and transplant health diagnosis, for example. Accordingly, as illustrated in FIGS. 1-4, provided herein are methods for preparinglibraries of amplified DNA molecules comprising one or more methylation-transferred DNA molecules.
[0039] Methods herein can optionally include extracting or isolating DNA molecules from a subject, for example from a liquid sample in a subject, which in illustrative embodiments comprises cfDNA. Alternatively, DNA molecules extracted or derived from a subject can be fragmented, for example by shearing, before they are processed in exemplary methods herein. In certain illustrative embodiments, cfDNA is extracted or derived from a subject having, suspected of having, or who had cancer. In some embodiments, methylation of genomic DNA is analyzed in a tumor of such subject, optionally using methods herein, and methylation of cfDNA is analyzed from that same subject, typically using methods herein, to identify ctDNA in the cfDNA sample.
[0040] In some embodiments, such as but not limited to those shown in FIGs. 1-4, methods herein can include modifying extracted or isolated DNA molecules (e.g., cfDNA) to produce DNA molecules derived from template DNA molecules (e.g., cfDNA) (e.g., modified DNA molecules such as modified cfDNA), before adapters are ligated to the modified DNA molecules (e.g., modified cfDNA). Such steps can include steps to prepare sample DNA molecules for adapter ligation, such as steps that can be performed during NGS library preparation. For example, such steps can include blunt-end repair and / or A-tailing, including addition of a poly- or single A tail. The adapters used for ligation in illustrative embodiments are Y adapters, such as those used in NGS library preparation. In some embodiments, molecular index tags (MITs) may be added to the template DNA molecules. In some embodiments, as described herein, the MITs can be used to identify errors caused by downstream applications, such as amplification, methyl transfer, and bisulfite or enzymatic conversion. In some embodiments, sample barcodes may be added to the template DNA molecules, which will allow for sample pooling in downstream applications.
[0041] In certain embodiments, adapted template DNA molecules have two, three, four or more methylated nucleic acid residues. For example, in some embodiments, the adapted template DNA molecules may have 3, 4, 5, 6, 10, 15, 20, 25, 30, 35, 40 or more methylated nucleic acid residues. The adapted template DNA molecules of FIGs. 2-4 each includes a methylated nucleic acid residue. The methylated nucleic acid of the adapted template DNA molecule includes a CpG site with a methylated cytosine residue. In some embodiments, the adapter sequences comprisemethylated cytosines. In some embodiments, the adapter sequences comprise at least one methylated nucleotide (such as methylated cytosine) and at least one unmethylated nucleotide (such as unmethylated cytosine).
[0042] In some embodiments, the adapted template DNA molecules are copied, for example by primer extension. In some embodiments, universal primers bind to adapter sequences and are extended using the methods described herein to generate a copy of the template DNA molecule. In some embodiments, target specific primers are hybridized at or near the genomic region of interest and are extended to using the methods described herein to generate a copy of the template DNA molecule. After the template is copied, the methylation pattern is transferred, for example using a methyltransferase. The copying / methyltransferase steps can be repeated for multiple cycles to generate a library of amplified DNA molecules comprising the one or more methylation- transferred DNA molecules. In some embodiments, the library may then be split for use in multiple different downstream applications. For example, the library can be used for methylation analysis, somatic variant analysis, fragmentomics, non-invasive prenatal testing (NIPT), cancer detection / diagnosis, and / or transplant health diagnosis.
[0043] In some embodiments, the methods herein include performing one or more steps to analyze the methylation signatures of the one or more methylation-transferred DNA molecules or their derivatives. In some embodiments, the methylation signatures are ascertained by treating the one or more methylation-transferred DNA molecules with an agent that discriminates between methylated and non-methylated cytosines. In some embodiments, the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules, or a portion thereof or their derivatives, may be treated with a chemical agent (such as bisulfite) or an enzyme that converts unmcthylatcd cytosine to uracil, while maintaining methylated cytosine (such as 5- methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC)). In some embodiments, the methods herein include contacting the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules, or a portion thereof or their derivatives, with one or more restriction enzymes that discriminate between methylated and unmethylated CpG dinucleotides, such as methylation sensitive restriction enzymes (MSREs), which can specifically cleave unmethylated CpG, but not methylated CpG, within the enzymes’ recognition sites, and methylation dependent restriction enzymes (MDREs), which can specifically cleave methylatedCpG, but not unmethylated CpG, within the enzymes’ recognition sites. The methylation status of the genomic regions of interest can then be determined using methods described herein.
[0044] Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The singular terms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. “Comprising A or B” means including A, or B, or A and B. It is further to be understood that all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for DNA molecules or polypeptides are approximate, and are provided for description.
[0045] Further, ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 1 to 49, 1 to 25, 1.7 to 31.9, and so forth (as well as fractions thereof unless the context clearly dictates otherwise). Any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated. Also, any number range recited herein relating to any physical feature, such as polymer subunits, size or thickness, are to be understood to include any integer within the recited range, unless otherwise indicated. When multiple low and multiple high values for ranges are given that overlap, a skilled artisan will recognize that a selected range will include a low value that is less than the high value.
[0046] As used herein, “about” or “consisting essentially of’ mean ± 10% of the indicated range, value, or structure, unless otherwise indicated. As used herein, the terms “include” and “comprise” are open ended and are used synonymously. As used herein, “comprising” is synonymous with "including," "containing," or "characterized by," and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. As used herein, "consisting of" excludes any element, step, or ingredient not specified in the claim element. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novelcharacteristics of the claim. In each instance herein any of the terms "comprising", "consisting essentially of" and "consisting of" may be replaced with either of the other two terms. The invention illustratively described herein suitably may be practiced in the absence of any element or elements, limitation or limitations which is not specifically disclosed herein.
[0047] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entireties. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0048] It is appreciated that certain features of aspects and embodiments herein, which are, for clarity, discussed in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various aspects and embodiments, which are, for brevity, discussed in the context of a single aspect or embodiment, may also be provided separately or in any suitable sub-combination. All combinations of aspects and embodiments are specifically embraced herein and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various aspects and embodiments and elements thereof are also specifically disclosed herein even if each and every such subcombination is not individually and explicitly disclosed herein.Sample Extraction and Enrichment of DNA Molecules
[0049] Samples that arc useful for methods herein can be virtually any nucleic acid sample. In illustrative embodiments, the nucleic acid sample is extracted or isolated from a subject. In illustrative examples, DNA molecules are extracted or isolated from a sample such as a tissue sample, and for illustrative embodiments herein, a liquid sample. Methods that are particularly useful in exemplary embodiments include methods for isolating circulating free DNA (cfDNA) from a liquid sample, and in illustrative embodiments from a blood, serum, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample, and in further illustrative embodiments, a plasma sample.
[0050] DNA methylation biomarkers are increasingly being utilized in the development of novel assays for use in, for example, cancer and other disease detection, women’s health, organ health, and veterinary health. For example, methylation profiling is being used in non-invasive prenatal testing (NIPT) for monitoring of placental and fetal epigenomic changes, pre- symptomatic detection of preterm birth, preeclampsia, placental insufficiency, and fetal growth restriction. Methylation markers can also be indicative of congenital diseases of the fetus. Furthermore, methylation biomarkers are also used for monitoring of organ health such as predicting and monitoring organ rejection in transplant patients, monitoring of immune changes in transplant rejection, and for monitoring of organ health in high risk or predisposed individuals. In addition, methylation biomarkers can be used to determine biological age or detect age-related diseases. Accordingly, the sample described herein may be a maternal sample, a fetal sample, a sample from a transplant patient, or a sample from a subject having, had, suspected of having, or at risk of having a certain phenotype.
[0051] In certain illustrative embodiments, isolation of cfDNA from a liquid (e.g., blood or blood derivative sample such as a serum or plasma sample) can involve binding DNA molecules from a sample to a matrix and isolating the DNA molecules in the presence of a solvent. In some embodiments, the method further comprises incubating the biological sample comprising DNA molecules with a protease, prior to contacting the DNA molecules to the matrix. In some embodiments, the method can further include the steps of washing the matrix with a wash buffer to remove impurities, and optionally, drying the matrix. Enriched nucleic acid samples can be eluted from the matrix with an elution buffer.
[0052] Other methods for nucleic acid isolation, for example cfDNA isolation, and optional enrichment of certain cfDNA can include ion exchange columns, or microfluidic devices, such as solid phase isolation, based on DNA capture by immobilized beads or functionalized surface. Additional methods include liquid phase isolation, utilizing an electric field, or chemical reagents, instead of a functionalized surface. In illustrative embodiments herein, isolation of cfDNA from a patient sample is performed using a DNA isolation kit (e.g., QIAamp Circulating Nucleic Acid kit (Qiagen)).
[0053] In some embodiments, cfDNA or their derivatives of certain sizes can be enriched before or after subjecting the cfDNA to methods herein. In some embodiments, size selection can be performed before the sequencing library preparation. In some embodiments, size selection can be performed after the sequencing library preparation and before sequencing. In some embodiments, size selection is performed on a sequencing-ready pool. Enriched cfDNA molecules can be, for example 50 to 1200 base pairs in length, 70 to 500 base pairs in length, 100 to 200 base pairs in length, or 130 to 170 base pairs in length. In some embodiments, the enriched cfDNA molecules are from 50 to 200 bp in length. In some embodiments, the enriched cfDNA molecules are between 60 and 200 bp in length, between 60 and 150 bp in length, or between 60 and 100 bp in length before the enriched cfDNA molecules, or derivatives thereof, are ligated to adapters in methods herein. In some embodiments, the enriched cfDNA molecules are less than 150, 100, 90, 75, or 50 bp in length before they are ligated to adapters. Such enrichment methods can be performed for example using the methods of WO2018156418 A l , Stray, et al. (incorporated herein by reference in its entirety).
[0054] In some embodiments, the sample is enriched for tumor DNA molecules, which are typically less than 160 bp and have a peak at about 145 bp in length. In illustrative embodiments, the enriched nucleic acid sample is circulating tumor DNA (ctDNA) or amplicons thereof. In such embodiments, the enriched nucleic acid is less than 160, 150, 145, 120, 100, 90, 75, or 50 bp in length. In some embodiments, size selection is used to filter out cfDNA molecules carrying mutations derived from clonal hematopoiesis of indeterminate potential (CHIP) but not tumor- derived mutations, which are typically longer than ctDNA molecules, at about 165bp.
[0055] In some embodiments, the sample is enriched for fetal DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is fetal circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of fetal cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170 bp, from 160 to 190 bp, or from 170 bp to 220 bp.
[0056] In some embodiments, the sample is enriched for transplant donor DNA molecules. In illustrative embodiments of such embodiments, the enriched nucleic acid sample is transplantdonor circulating free DNA or amplicons thereof. In such embodiments, the enriched nucleic acid is between 100 bp and 220 bp in length. In some embodiments, the length of the nucleic acid sample of transplant donor cfDNA ranges from 100 bp to 200 bp, from 120 bp to 180 bp, from 140 bp to 160 bp, from 150 bp to 170bp, from 160 bp to 190 bp, or from 170 bp to 220 bp.Subjects
[0057] Subjects in methods herein can be virtually any animal, in illustrative embodiments a mammal, and in further illustrative embodiments, a human. In some embodiments, the subject has, had, or is suspected or at risk of having a disease, in some embodiments, cancer. In some embodiments, the subject has received or is receiving a treatment for cancer and is being tested for minimal residual disease (MRD). In some embodiments, the subject is a pregnant female. In some embodiments, the subject is a subject comprising an organ from another individual.
[0058] In embodiments where the subject has, had, or is suspected or at risk of having cancer, the cancer can be any type of cancer provided that the genome of cancerous cells of the subject have portions of their genome that are differentially methylated compared to non-cancerous cells of the subject. Typically, some, most, almost all or all cancer cells of the subject have region(s) of their genome that are methylated that are not methylated in non-cancerous cells of the subject, or that a e more methylated than non-cancerous cells. Thus, in some embodiments, the subject has one or more cancers (e.g., one cancer) including ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lung carcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2- expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma,biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.
[0059] In certain embodiments of methods herein, the subject has, had, or is suspected or at risk of having a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, blood, bone, bone marrow, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastro-esophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovaries, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whippie resection. In illustrative embodiments, the cancer is selected from nasopharyngeal carcinoma, hepatocellular carcinoma, breast cancer, ovarian cancer, pancreatic cancer, colorectal cancer, lung cancer, oesophageal cancer, prostate cancer, bladder cancer, melanoma, and acute leukemia. In illustrative embodiments, the cancer is selected from colorectal cancer.Nucleic Acid Processing
[0060] Typically, methods herein include a step of appending nucleic acid adapters to template DNA molecules or to nucleic acid derivatives generated therefrom. For example, adapters may be appended on to the DNA molecules by ligation or PCR. The template DNA molecules in illustrative embodiments are from a subject. In some embodiments, appending nucleic acidadapters is preformed after template DNA molecules are fragmented to form fragmented DNA molecules. Typically, methods include exposing template DNA molecules to one or more polymerases or kinases, such as Klenow Large Fragment Polymerase and T4 polynucleotide kinase (PNK), as well as a ligase, such as T4 ligase. In some embodiments, template DNA molecules or the fragmented DNA molecules are exposed to one or more polymerases and / or kinases to generate the nucleic acid derivatives. In some embodiments, the method further comprises appending adapters to the nucleic acid derivatives generated therefrom. In some embodiments, template DNA molecules are not fragmented prior to appending nucleic acid adapters thereto. In some embodiments, template DNA molecules are cfDNA molecules.
[0061] In some embodiments, adapters are ligated to template DNA molecules. Before such ligation, extracted or isolated DNA molecules can be modified to form sample nucleic acid derivatives, for example to make them more amenable to adapter ligation. For example, template DNA molecules can be blunt ended, nucleotides can be added to template DNA molecules or blunted-ended derivative therefrom, and / or phosphate moieties can be added or removed from the ends of sample DNA molecules or derivatives thereof. In some embodiments, prior to ligation, template DNA molecules may be blunt ended, and then a single adenosine base can be added to the 3’ end. Prior to ligation the DNA may be cleaved using a restriction enzyme or some other cleavage method. During ligation the 3’ adenosine of the template fragments and the complementary 3’ thymidine overhang of an adapter can enhance ligation efficiency. In some embodiments, adapter ligation is performed using a T4 ligase.
[0062] In some embodiments, adapters containing one or more universal priming sequences are utilized in methods herein. In some embodiments, the adapters are Y adapters, for example in methods in which target region amplicons arc sequenced using NGS. In some embodiments, the adapters each comprises a universal priming site. The adapters may or may not include methylated cytosine residues.
[0063] In some embodiments, the adapters each further comprises a sample barcode. Thus, multiple samples can be analyzed in the same sequencing reaction. The sample barcode can be used to process data according to the sample from which the data was generated.
[0064] In some embodiments, the adapters each further comprises an MIT. In some embodiments, the number of adapters having different MITs is between 10 to 1,000, and wherein the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 1,000:1. The number of different MITs in the ligation reaction, in certain embodiments, ranges fromlO to 50, 10 to 100, 50 to 200, 100 to 300, 200 to 500, 300 to 600, 500 to 700, 600 to 800 or 700 to 1,000. In some embodiments, there are at least 1, 10, 20, 30, 40, 50, or at least 100; 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different MITs in the ligation reaction. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction is at least 10,000:1. In some embodiments, the ratio of the total number of template nucleic acid or cfDNA molecules to the number of different MITs in the ligation reaction ranges from 50,000: 1 to 50:1, from 25,000:1 to 100:1, from 10,000:1 to 100:1, from 10:000:1 to 8,000:1 to 500:1, from 5,000:1 to 200:1, from 10,000:1 to 50: 1 . In some embodiments, the methods disclosed herein result in at least 100; 200; 500; 750; 1,000; 2,000; 5,000; 7,500; 10,000; 20,000; 25,000; 30,000; 40,000; 50,000 different MITs to each one template nucleic acid or cfDNA molecules.Template Copying
[0065] Methods described herein include copying of the template DNA molecules for example by primer extension, as shown in FIGs. 2-5. In some embodiments, adapters may be appended to the 5’ and / or 3’ ends of the template DNA molecule, and primers specific for an adapter sequence, such as a universal priming sequence, hybridize to the adapter sequence and are extended using a polymerase enzyme, for example, to generate a copied DNA molecule.
[0066] In some embodiments of the methods described herein, primer extension is carried out under isothermal conditions, such as is the embodiment of FIG. 1. In some embodiments, thermal cycling is employed. In some embodiments, primer extension is performed at between about 37 °C and 64 °C, between about 40 °C and 64 °C, between about 50 °C and 64 °C, or between about 58 °C and 64 °C. In some embodiments, primer extension is performed at about 60 °C. Methods that can be carried out under such isothermal conditions include adapter priming, as illustrated in FIGs.2 and 5, strand invasion, as illustrated in FIG. 3, and enhanced strand invasion, as illustrated in FIG. 4, rolling circle amplification, and variations thereof.Adapter priming
[0067] In some embodiments, adapters comprising primer binding sequences may be appended to the ends of the template DNA molecules. In some embodiments, Y-adapters having unique tail sequences and comprising primer binding sequences may be appended to the ends of the template DNA molecules. In some embodiments, primers hybridize to a single stranded “Y” portion of the adapter, as shown in FIG. 2. This allows for single strand extension and avoids the need for heat denaturation of the template DNA molecule. Using this method and variations thereof, a maximum of 2-fold amplification can be achieved.Strand invasion
[0068] In some embodiments, as illustrated in FIG. 3, adapters can be designed to be more susceptible to fraying at specific temperatures. For example, adapters having higher AT content will be prone to fraying at around 60 °C. In such methods, primer hybridization can take advantage of fraying DNA ends, allowing strand invasion to occur. In some embodiments, primers are designed to improve strand invasion capability. For example, primers comprising locked nucleic acids (LNAs) may be used. In some embodiments described herein, recombinase polymerase amplification (RPA) may be used to enhance primer binding. Using these methods and variations thereof, 10-fold or more amplification can be achieved.Enhanced strand invasion
[0069] In some embodiments, as illustrated in FIG. 4, the primers used in the methods described herein are physically linked (5 ’-5’) to increase primer concentration and on-rate. One challenge with isothermal amplification is creating single- stranded regions where primers can bind. This is particularly challenging since only the 5’ end of a primer extension product is ‘controlled’ i.e., can be determined by a primer sequence (and priming occurs on the 3’ end of a template strand). By utilizing 5’ linked primers, the local concentration of primers can be dramatically increased for the subsequent cycle, which leads to a higher rate of strand invasion.Extension reaction mixture
[0070] Typically, in embodiments described herein, primer extension is performed by adding an extension reaction mixture to the template DNA (e.g., adapted sample DNA, such as adapted sample cfDNA) followed by addition of a polymerase enzyme. In some embodiments, the extension reaction mixture contains one or more primers, deoxynucleotides (dNTPs), reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from 0.1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is between 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0071] In illustrative embodiments, the reaction buffer is a BST polymerase buffer such as the ThermoPol® Reaction Buffer (B9004S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is an isothermal amplification buffer (e.g., B0537S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is Q5® Reaction Buffer (B9027S, New England BioLabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England BioLabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England BioLabs, Inc.).
[0072] In some embodiments, the polymerase is a BST DNA polymerase, such as BST DNA Polymerase (M0328S, New England BioLabs, Inc.) or BST 2.0 DNA Polymerase (M0537S, New England BioLabs, Inc.). BST DNA polymerases have 5 ^ 3" polymerase and double-strand specific 5' — 3" exonuclease activity, but lacks 3' -- 5' exonuclease activity. In some embodiments, a warm start BST DNA polymerase may be used to help reduce non-specific activity. For example, in some embodiments, the polymerase is Bst 2.0 WarmStart® DNA Polymerase (M0538S, New England BioLabs, Inc.).
[0073] In some embodiments, the polymerase is selected based on the particular activity levels at specific temperatures. This is important when using the polymerase for primer extension prior tomethyl transfer, as explained in more detail below. For example, BST DNA polymerase provides 100% activity at 60-65 °C, but low activity (10-15%) at 37 °C. Accordingly, there will be little amplification occurring when the methyl transfer step is being performed. In some embodiments, isothermal conditions are used for primer extension, such that the reaction is performed at around 60 °C for example. In some embodiments, primer extension may be performed for a short amount of time. For example, in some embodiments, primer extension is performed for between 1 and 5 mins, between 1 and 10 mins, between 1 and 20 mins, or between 1 and 30 mins.
[0074] Buffer solution creates a suitable environment for the polymerase enzyme and can contain many different components, including magnesium chloride (MgCh), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.
[0075] In some embodiments, the primers are universal primers designed to hybridize to sequences on the appended adapters. In some embodiments, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length,between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.Single-stranded DNA (ssDNA) molecules
[0076] In some embodiments, the template DNA molecules are single-stranded DNA (ssDNA). In some embodiments, as illustrated in FIG. 5, at least one adapter having a primer binding sequence is appended to one or more ssDNA template molecules. In some embodiments, only a single primer is needed for primer extension, which does not require heat-denaturing. The primer binds to and extends from the primer binding sequence in the adapter. In some embodiments, primers and adapters in the methods disclosed herein may include nucleotides with methylation modifications (e.g. methylated cytosines) for use in downstream methylation detection assays. ssDNA can be prepared using single stranded library preparation methods such as IDT xGen™ ssDNA & Low-Input DNA Library Preparation Kit or ClaretBIO SRSLY kit. In some embodiments, a second adapter is appended before or after the primer extension and methylation transfer steps. In some embodiments, after the first primer extension, other amplification methods disclosed herein (e.g., strand invasion, enhanced strand invasion) can be used for additional amplification.Methylation Transfer
[0077] As mentioned previously, template copying or amplification fails to preserve the methylation status of the template DNA molecule on the copied DNA molecule. Accordingly, after the template DNA molecule is copied, the methods described herein include a step of methylation transfer to generate methylation-transferred DNA molecules.
[0078] In the methods described herein, one or more methylating agents (for example, methyltransferase) are used to preserve the methylation pattern of the template DNA molecule on the copied DNA molecule. In some illustrative embodiments, the methylating agent is DNMT1. DNMT1 is a maintenance DNA methyltransferase that propagates the CpG DNA methylation pattern in dividing cells; DNMT1 acts on hemi-methylated DNA, where one strand (the original, template strand) is methylated and the other strand (the newly synthesized strand) is not methylated, and copies the methylation pattern from the old strand to the new strand with highefficiency and specificity. DNMT1 is not thermostable, and accordingly in some embodiments the method described herein comprises adding DNMT1 fresh after each cycle of template amplification or copying. In some embodiments, other methylating agents may be used, such as mammalian methyltransferases DNMT3a and DNMT3b, plant methyltransferases DRM2, MET1, and CMT3, and bacterial methyltransferase Dam. In some embodiments, methyltransferases such as DNMT1 or other suitable methyltransferases are used with one or more sources of methyl groups, such as S-adenosylmethionine (SAM), and may be used with or without cofactors such as NP95 (Uhrfl). In some embodiments, a DNMT1 reaction buffer may be used with the methyltransferase that includes DNMT1 reaction buffer. For example, the DNMT1 reaction buffer may comprise 200 mM NaCl, 50 mM Tris-HCl, 1 mM EDTA, 1 mM DTT, and 50% glycerol. In some embodiments, the DNMT1 reaction buffer may further comprise BSA and S- adenosylmethionine (SAM).
[0079] In some embodiments, a methyltransferase, such as DNMT1, may require conditions different to those present in the template copying reactions. For example, methyl transfer reactions may require buffer conditions that do not include ions (such as cations), such as magnesium ions or manganese ions which may be a component of a primer extension reaction. According, in some embodiments, a chelating agent such as EDTA is required after the primer extension step to chelate ions, such as magnesium ions in order for the methylation step to be carried out. In some embodiments, magnesium is replenished back into the primer extension reaction mixture for the next round of primer extension after the completion of methyl transfer reaction.
[0080] As explained above, the primer extension step is generally performed at about 60 °C, when the polymerase activity is highest. Methyl transferase on the other hand, such as DNMT 1 has a recommended temperature of 37 °C and is deactivated at 65 °C. In some embodiments, the methyl transfer step is performed for between 5 and 120 mins, for between 10 and 90 mins, for between 20 and 60 mins, or for between 30 and 45 mins. Accordingly, in illustrative examples, a two-step process of primer extension at approximately 60 °C (or in some embodiments, between about 37 °C and 64 °C, between about 40 °C and 64 °C, between about 50 °C and 64 °C, or between about 58 °C and 64 °C) to copy template DNA molecules, and DNA methyl transfer at about 37 °C (in some embodiments, between about 30 °C and 45 °C, between about 33 °C and 42 °C, or between about 35 °C and 40 °C) to copy the methylation pattern on to the copied DNA molecule can berepeated multiple times, as shown in FIGs. 1-5, to generate a library of amplified DNA molecules comprising one or more methylation-transferred DNA molecules, during which the two enzymes will not be active at the same time. In some embodiments, the primer extension step is performed for a shorter amount of time than the methyl transfer step.
[0081] In some embodiments, the two-step process of primer extension and methylation transfer may be repeated for at least 1 cycle, at least two cycles, at least three cycles, at least four cycles, at least five cycles, at least 10 cycles, at least 15 cycles, at least 20 cycles, at least 30 cycles, at least 40 cycles, or at least 50 cycles. In some embodiments, the two-step process of primer extension and methylation transfer may be repeated for more than 50 cycles. In some embodiments, the two- step process of primer extension and methylation transfer may be repeated for between 1 and 50 cycles, between 1 and 40 cycles, between 1 and 30 cycles, between 1 and 20 cycles, between 1 and 15 cycles, between 1 and 10 cycles, or between 1 and 5 cycles.Methylation Assays
[0082] Libraries of methylation-transferred DNA molecules generated using the methods described herein or portions thereof can be used for many downstream applications. One of such applications is for the analysis of methylation biomarkers, including for non-invasive prenatal testing (NIPT), cancer detection or diagnosis, and transplant health diagnosis. For purposes of illustration only, some of the methods that can be used for the analysis of methylated DNA are described below.Conversion-based methylation detection
[0083] In some embodiments, the methods described herein include treating the methylation- transferred DNA molecules, or their derivatives, with a deaminating agent. In some embodiments, a portion of a I ibrary of amplified DNA molecules comprising the one or more methylation- transferred DNA molecules is treated with a deamination agent. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with a deamination agent. In some embodiments, the entire library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules is treated with a deamination agent.
[0084] In some embodiments, the deaminating agent is an enzyme such as a deaminase. In some embodiments, the deaminating agent comprises a bisulfite reagent. In some embodiments, the bisulfite reagent may be sodium bisulfite, potassium bisulfite, ammonium bisulfite, magnesium bisulfite, sodium metabisulfite, potassium metabisulfite, ammonium metabisulfite, magnesium metabisulfite, and equivalents. In some embodiments, a combination of one or more chemical reagents and / or one or more enzymes can be used to discriminate between various forms methylated cytosines (e.g., mC or hmC) and unmethylated cytosines.
[0085] Bisulfite reagents, such as sodium bisulfite, convert cytosine to uracil and leave the 5- methylcytosine (mC) unchanged. Therefore, after bisulfite treatment, 5-mC in the DNA remains as cytosine and unmodified cytosine will be changed to uracil. The bisulfite treatment can be performed by commercial kits such as the Imprint DNA Modification Kit (Sigma), EZ DNA Methylation- DirectTM Kit ( ZYMO), and the EZ DNA Methylation-Gold Kit (ZYMO). After DNA bisulfite conversion, single stranded DNA is captured, desulphonated and cleaned. The bisulfite-treated DNA can be captured by purification columns or magnetic beads and eluted. Bisulfite-treated single stranded DNA can be converted into dsDNA through an enzyme-catalyzed DNA strand synthesis with appropriate primers and polymerase. The polymerase will recognize the uracil in the ssDNA template as thymine and add an adenine to the complimentary strand.Further polymerase extension on the complimentary strand will result in replication of the original bisulfite treated ssDNA template, substituting uracil with thymine. Identification of cytosine to thymine conversion and guanine to adenine conversions (complementary strand) through comparing to the reference genome, will determine all unmodified cytosines, while the remaining cytosines are considered to be methylated. In some embodiments, after bisulfite conversion, the methylation-transferred and bisulfite-converted molecules are subject to target enrichment as described herein followed by high-throughput sequencing.
[0086] Other conversion-based methyl detection methods, including EM-seq, oxidative bisulfite sequencing (oxBS-seq), TET-assisted bisulfite sequencing (TAB-seq), TET-assisted pyridine borane sequencing (TAPS), TAPS with T4- GT protection sequencing (TAPS ), chemical- assisted pyridine borane sequencing (CAPS), and many variants of these that can discriminate between C, mC, and / or hmC. These methods can be performed with commercial kits such asNEBNext® Enzymatic Methyl-seq (EM-seq™) (New England Biolabs, Inc.), the 5hmC TAB-Seq Kit (WsseGene), and the EpiTect® Bisulfite Kit (Qiagen), for example.Restriction enzyme-based methylation detection methods
[0087] In some embodiments, the methods described herein include treating the methylation- transferred DNA molecules, or their derivatives, with one or more restriction enzymes. In some embodiments, the one or more restriction enzymes are one or more methylation sensitive restriction enzymes (MSREs). In some embodiments, a portion of a library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules is treated with one or more MSREs. For example, in some embodiments, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the library is treated with one or more MSREs. In some embodiments, the entire library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules is treated with one or more MSREs.
[0088] MSREs are useful in analyzing the methylation status of cytosine residues in CpG sites. As the name implies, these enzymes are not able to cleave their palindromic target sites when the cytosine residues are methylated. The size of the MSRE targets sites range from 4 bp up to 76bp, but typically are in the 4 to 8 bp range.
[0089] In one aspect, methods herein typically include contacting methylation-transferred DNA molecules, or their derivatives, with one or more MSREs. MSREs selectively cleave a sample nucleic acid when the MSRE recognition site is unmethylated, but not when the MSRE recognition site is methylated. An “isoschizomer” of an MSRE is a restriction enzyme that recognizes the same recognition site as a methylation sensitive restriction enzyme but cleaves both methylated CGs and unmethylated CGs. Isoschizomer of the selected MSREs may be used in control reactions. Non-limiting examples of methylation sensitive restriction enzyme include, and thus in some embodiments, the one or more MSREs can include, Aatll, Acc65I, AccI, Acil, Acll, Afel, Agel, Agel-HF®, AhdI, Alel-v2, Apal, ApaLI ApeKI, Asci, AsiSI, Aval, Avail, Bael, BanI, BbvCI, BceAI,, Bcgl, BcoDI, BfuAI, Bgll, BmgBI, BsaAI, BsaBI, BsaHI, BsaI-HF®v2, BseYI, BsiE, BsiWI, BsiWI-HF®, BslI, BsmAI, BsmBI-v2, BsmFI, BspDI, BspEI, BsrBI, BsrFI-v2,BssHII, BstAPI, BstBI, BstUI, BstZ17I-HF®, BtgZI, Cac8I, Clal, Dpnl, Dralll-HF®, DrdI, Eael, Eagl-HF®, Earl, Ecil, Eco53kl, EcoRI, EcoRI-HF®,EcoRV, EcoRV-HF®, Esp3I, Faul, Fnu4HI,FokI, Fsel, FspI, Haell, Hgal, Hhal, HinPlI, HincII, Hinll, Hpal, Hpall, Hpyl66II, Hpyl88III, Hpy99I, HpyAV, HpyCH4IV, KasI, Mbol, Mini, MluI-HF®, Mmel, MspAlI, Mwol, Na , Narl, Neil, NgoMIV, Nhd-HF®, NlalV, Notl, Notl-HF®, Nrul, NruI-HF®, Nt.BbvCI, Nt.BsmAI, Nt.CviPII, PaeR7I, PaqCI, Pld, PluTI, Pmel, Pmll, PshAI, PspOMI, PspXI, Pvul, PvuI-HF®, Rsal, RsrII, Sad-HF®, Sadi, Sall, Sall-HF®, Sau3AI, Sau96I, ScrFI, SfaNI, Sfil, Sfol, SgrAI, Smal, SnaBI, Srfl, StyD4I, Tfil, Tsd, TspMI, Xhol, Xmal, and / or Zral. In illustrative embodiments, the one or more MSREs can include Hhal, Hpall, BstUI, and / or HpyCH4IV.
[0090] In some embodiments, the one or more restriction enzymes are one or more methylation dependent restriction enzymes (MDREs). MDREs selectively cleave a sample nucleic acid when one or more nucleotides in the MDRE recognition site is methylated. In some embodiments, one or more MDREs can be used in combination with one or more MSREs. In some embodiments, the one or more MDREs can be AbaSI, AoxI, BisI, BlsI, Dpnl, FspEI, Glal, Glul, Krol, LpnPI, Mall, MspJI, Mtd, Pcsl, PkrI, or Sgel.
[0091] The MSRE or MDRE can be selected based on differentially methylated CpG sites in target DNA molecules, such as tumors or ctDNA from specific cancer targets, or a diverse spectrum of tumors. Further criteria for selection may include low background methylation in normal tissues, size and number of cleavage fragments, number of base pairs of recognition sequence, whether the cleavage results in blunt vs. tailed end fragments, and whether the enzymes have the same or similar reaction conditions such that the contacting step can be done under the same set of conditions and / or in a single reaction.
[0092] In some embodiments, more than one, a plurality, or a set of MSREs (and / or MDREs) can be used to contact methylation-transferred DNA molecules, or their derivatives, comprising one or more CpG sites of interest. The number of selected MSREs in certain embodiments is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; from 2 to 3, 4, 5, 6, 7, 8, 9, 10, or 20, 25, 50 or 100; or from 5 to 6, 7, 8, 9, 10, or 20, 25, 50 or 100. In some embodiments, the number of MSREs range from 1 to 10, 1 to 5, from 2 to 5, from 3 to 5, from 4 to 5, 3 to 7, from 5 to 10, from 6 to 10, or from 7 to 10. Thus, criteria such as target site, target site methylation status, number of base pairs in tailed end fragments, reaction conditions including but not limited to bufferconditions, incubation time and temperature for optimal activity, as well as deactivation time and temperature for each MSRE can be used for selecting the plurality or set of MSREs and / or MDREs to include in the method. As activity measured by units is specific to each enzyme, the target DNA molecules can be contacted with from 1 to 5 Units (U) of each MSREs in the reaction sample. In some embodiments, the methylation-transferred DNA molecules, or their derivatives, are contacted with from 1 to 5 U, from 1.5 to 4 U, from 2 to 3 U, from 2.5 to 4 U, or from 3 to 5 U of each MSRE. In some embodiments, one or more of the MSRE is selected from Hpall, Sall, Bbel, Notl, Smal, Xmal, Mbol, BstUI, BstBI, Clal, Mini, Nael, Narl, Pvul, SacII, HpyCH41V, Hhal, and combinations thereof. In exemplary embodiments, the one or more MSREs is selected from one or more of Hpall, Hhal, HpyCH41 V, and BstUI. In some embodiments, the one or more MSREs are selected from one or more of Hpall, Hhal, HpyCH41V, or BstUI. In some embodiments of the methods as described herein, the contacting comprises contacting with two or more MSREs. In some embodiments, the contacting comprises contacting with three or more MSREs. In some embodiments, the contacting comprises contacting with four or more MSREs. In some embodiments, the contacting comprises contacting with the two or more MSREs in a single reaction.
[0093] In some embodiments, MSREs and / or MDREs are included that are able to cleave the MSRE sites and / or MDRE sites in the same cleavage buffer conditions. In some embodiments, the cleavage buffer conditions include 5-500 mM potassium acetate, for example 25-100 mM potassium acetate. In some embodiments, the cleavage buffer conditions include 2-200 mM Trisacetate, for example 10-40 mM Tris-acetate. In some embodiments, the cleavage buffer conditions include 1-100 mM magnesium acetate, for example 5-20 mM magnesium acetate. In some embodiments, the cleavage buffer conditions include 10-1000 pg / ml recombinant albumin, for example 50-200 pg / ml recombinant albumin. In some embodiments, the pH is between 6.9 and 8.9 at 25 °C, for example, between 7.4 and 8.4, 7.5 and 8.3, 7.6 and 8.2, 7.7 and 8.1, or 7.8 and 8, or about 7.9 at 25 °C. In some embodiments, the cleavage buffer conditions include 1-100 mM bis- tris-propane-HCl, for example 5-20 mM bis-tris-propane-HCl. In some embodiments, the cleavage buffer conditions include 1-100 mM MgCh. for example 5-20 mM MgCh. In some embodiments, the pH is between 6 and 8 at 25 °C, for example, between 6.5 and 7.5, 6.6 and 7.4, 6.7 and 7.3, 6.8 and 7.2, or 6.9 and 7.1, or about 7.0 at 25 °C. In some embodiments, the cleavage bufferconditions include 5-500 mM NaCl, for example 25-100 mM NaCl. In some embodiments, the cleavage buffer conditions include 1-100 mM Tris-HCl, for example 5-20 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 10-1000 mM NaCl, for example 50-200 mM NaCl. In some embodiments, the cleavage buffer conditions include 5-500 mM Tris-HCl, for example 25-100 mM Tris-HCl. In some embodiments, the cleavage buffer conditions include 25- 100 mM potassium acetate, 10-40 mM Tris-acetate, 5-20 mM magnesium acetate, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 5-20 mM bis-tris-propane-HCl, 5-20 mM MgCh, and 50- 200 p.g / ml recombinant albumin, and the pH is between 6.7 and 7.3at 25 °C. In some embodiments, the cleavage buffer conditions include 25-100 mM NaCl, 5-20 mM Tris-HCl, 5-20 mM MgCh, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C. In some embodiments, the cleavage buffer conditions include 50-200 mM NaCl 25-100 mM Tris- HCl, 5-20 mM MgCh, and 50-200 pg / ml recombinant albumin, and the pH is between 7.6 and 8.2 at 25 °C.Other Downstream Assays
[0094] The methylated DNA libraries prepared using the methods herein or portions thereof can be used for additional downstream assays, such as the detection of one or more somatic mutations. Detection of somatic mutations are useful for non-invasive prenatal testing (NIPT), cancer detection / diagnosis, and transplant health diagnosis.
[0095] In some embodiments, methods of detecting somatic mutations may include for example, PCR assays such as ddPCR, or NGS, including PCR-targeted NGS, LTC-targeted NGS, or hybrid capture-targeted NGS. Methods of detecting somatic variants are described in e.g., W02019200228, WO2017168329, WO2017168332, and W02012108920, incorporated herein by reference. In some embodiments, the library of methylation-transferred DNA molecules generated using the methods described herein, or a portion thereof, are used to detect somatic variants (e.g., single nucleotide variants (SNVs), multi-nucleotide variants (MN Vs), InDeis, gene fusions, structural variants, or a combination thereof) in one or more genomic regions of interest.
[0096] The methylated DNA libraries prepared using the methods herein or portions thereof can also be used for fragmentomics studies. The fragment characteristics of cfDNA, such as fragmentlength, end motifs and frequencies, nucleosome footprints, can be used to detect ctDNA and determine tissue of origin, and can be particularly useful in early detection of multiple cancers. cfDNA fragmentation patterns can also be used for assessing graft suitability or transplant health in a transplant patient. (See e.g., WO2021055968, incorporated by reference herein in its entirety).Error Correction using MITs
[0097] Some embodiments of the methods described herein use molecular indexes (MITs) for identifying and removing errors produced at different stages of the methods.
[0098] In previous methods of conversion based detection of a methylation pattern of a target region, amplification of a template is not performed prior to the conversion, for example chemical and / or enzymatic conversion. Accordingly, as illustrated in FIG. 6, an MIT would only represent a single conversion event, since there are no molecular copies to be converted. The MITs are represented in FIG. 6 by “Y” and “Z” on the end of the templates. If a mistake is made during conversion, as illustrated in FIG. 6B, this mistake will be propagated to any molecular copies made post-conversion, resulting in an incorrect consensus sequence. Likewise, in previous restriction-enzyme-based methylation detection methods, a mistake during the restriction enzyme digestion would be propagated to any molecular copies made post-restriction digest.
[0099] Using the methods described herein, as illustrated in FIGs. 7A-B, because MITs are appended to the template DNA molecules before template copying and methylation transfer, which in turn are performed prior to bisulfite or enzymatic conversion, MITs represent multiple conversion events, and thus can be used to correct errors in conversion and / or amplification. Consensus can be made which preserve the initial state of the template DNA molecule (both sequence and methylation). Unlike consensus sequences often used for determining genetic variants, the consensus required to detect a methylated base does not need to be over 50% of the bases at a given position for a given MIT. For example, as long as the methylated base signal is above the background methylation non-specific activity, library preparation, PCR, sequencing errors and other error and non-specific activity at that position, a methylated base may be called. Background errors can be modeled and / or measured experimentally. The use of methylation amplification in conjugation with MITs to create a methylation consensus from multipleconversion events was previously unknown and represents a significant improvement over existing technology.Enrichment
[0100] In some embodiments, the genomic regions of interest may be selectively enriched either before or after preparation of the methylation-transferred DNA molecules. For example, in some embodiments, after the DNA is extracted from the subject, the DNA may be selectively enriched for genomic regions of interest using probes or primers specific for the regions of interest. In some embodiments, after the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules, the DNA library may be selectively enriched for genomic regions of interest using probes or primers specific for the regions of interest. In some embodiments, after the library has been treated with a deaminating agent to generate converted DNA molecules, the DNA library may be selectively enriched for genomic regions of interest using probes or primers specific for the regions of interest. In some embodiments, after the library has been treated with one or more MSREs or MDREs to generate treated DNA molecules, the DNA library may be selectively enriched for genomic regions of interest using probes specific for the regions of interest.Hybrid Capture
[0101] In some embodiments, the selective enrichment technique can involve fragment capture by hybridization (i.e., hybrid capture). Although any hybrid capture method can be used to perform methods herein that include a selective enrichment step, in some embodiments, a method of the present disclosure may involve using any of the hybrid capture methods disclosed herein to selectively enrich DNA, for example, cfDNA. In some embodiments described herein, such selective enrichment steps can be performed after the template DNA molecules are extracted. In some embodiments described herein, such selective enrichment steps can be performed following the addition of adapters to the template DNA molecules. In some embodiments described herein, such selective enrichment steps can follow generation of the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules.
[0102] In capture by hybridization, hybrid capture oligonucleotide probes complementary to one or both strands of specific target DNA sequences, or DNA derived therefrom in a sample, are utilized, i.e., the probes may be strand specific. In some embodiments, the probes may be complementary to the target sequence in either a 100% or 0% methylated state. Accordingly, each strand specific probe may have two designs, assuming both a completely methylated and a completely unmethylated status of each target DNA sequence, yielding a total of four unique probes per target. The specific target DNA sequence in illustrative embodiments overlaps with or is found within a target region of a sample DNA molecule such as a cfDNA. Thus, hybrid capture probes when used in methods herein can be designed to bind to a DNA molecule that contains at least one target region or a portion thereof. In some embodiments, the hybrid capture probes can be designed to bind to a target DNA sequence within or overlapping a target region, which in illustrative embodiments contains one or more methylated nucleic acids. In other examples, the hybrid capture probes can be designed to bind to a common region that is flanking but not overlapping the target region and that can be a common region that was added to some, most, almost all or all of the DNA in a sample, or added to all amplicons using a common sequence on at least one primer of a primer pair. In illustrative embodiments, a hybrid capture probe or set thereof, are designed to bind to a target DNA sequence within target region, or set of target regions, respectively.
[0103] Hybrid capture probes may be added to a prepared sample and hybridized through a denature -reannealing process to form duplexes of exogenous-endogenous fragments (e.g., hybrid capture probes bound to sample DNA molecules, or DNA derived therefrom). These duplexes may then be physically separated from the sample by various means. In some embodiments, once the hybrid capture probes are removed, the sample DNA molecules, or DNA derived therefrom can be amplified. Some ways to physically remove the hybrid capture probes arc by covalently bonding the hybrid capture probes to a solid support, for example a magnetic bead, or a chip. Another way to physically remove the hybrid capture probes is by covalently bonding them to a molecular moiety with a strong affinity for another molecular moiety. An example of such a molecular pair is biotin and streptavidin, such as is used in SURE SELECT (Agilent). Thus, hybrid capture probes, for example that bind to a target DNA sequence within or overlapping a target region of a DNA molecule obtained or derived from a sample, can be covalently attached toa biotin molecule, and after hybridization with sample DNA or DNA derived therefrom, a solid support with streptavidin affixed can be used to pull down the biotinylated hybrid capture probes, which are hybridized to DNA molecules obtained or derived from a sample that include a target region that includes the target DNA sequence recognized by the hybrid capture probes. Thus, in some embodiments, the hybrid capture probes are immobilized, directly or indirectly to a solid support. In some embodiments, the hybrid capture probes include a binding partner, for example biotin.
[0104] In some embodiments of any of the aspects herein, the hybrid capture probes can be a part of a set of at least two hybrid capture probes. In some embodiments, the set includes at least one hybrid capture probe for each target region. In some embodiments, the set includes two or more hybrid capture probes for each target region.
[0105] In some embodiments of any of the aspects herein, the hybrid capture probes can have a length in the range of 30 bases to 170 bases, 30 bases to 160 bases, 30 bases to 150 bases, 30 bases to 140 bases, 30 bases to 130 bases, 30 bases to 120 bases, 30 bases to 110 bases, 30 bases to 100 bases, 30 bases to 90 bases, 30 bases to 80 bases, 30 bases to 70 bases, 30 bases to 60 bases, 30 bases to 50 bases, 40 bases to 160 bases, 40 bases to 150 bases, 40 bases to 140 bases, 40 bases to 130 bases, 40 bases to 120 bases, 40 bases to 110 bases, 40 bases to 100 bases, 40 bases to 90 bases, 40 bases to 80 bases, 40 bases to 70 bases, 40 bases to 60 bases, 50 bases to 150 bases, 50 bases to 140 bases, 50 bases to 130 bases, 50 bases to 120 bases, 50 bases to 110 bases, 50 bases to 100 bases, 50 bases to 90 bases, 50 bases to 80 bases, 50 bases to 70 bases, 60 bases to 140 bases, 60 bases to 130 bases, 60 bases to 120 bases, 60 bases to 110 bases, 60 bases to 100 bases, 60 bases to 90 bases, 60 bases to 80 bases, 70 bases to 130 bases, 70 bases to 120 bases, 70 bases to 110 bases, 70 bases to 100 bases, 70 bases to 90 bases, 80 bases to 120 bases, 80 bases to 110 bases, 80 bases to 100 bases, 90 bases to 120 bases, 90 bases to 110 bases, 100 bases to 165 bases, 100 bases to 150 bases, 100 bases to 140 bases, 100 bases to 130 bases, 100 bases to 120 bases, 110 bases to 150 bases, 110 bases to 140 bases, 110 bases to 130 bases, 120 bases to 150 bases, or 130 bases to 160 bases.Targeted amplification
[0106] In some embodiments, at least a portion of the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules may enriched for specific target regions by targeted amplification. The enriched library may then be used for downstream applications such as somatic mutation analysis, fragmentomics, non-invasive prenatal testing (NIPT), cancer detection / diagnosis, and transplant health diagnosis.
[0107] Methods in some aspects herein include performing one or, in some embodiments, two or more amplifications. Such amplifications in certain illustrative embodiments include at least one targeted amplification wherein at least one primer and in certain embodiments both primers of a primer pair, one or more primer pairs, or a set of primer pairs used for the amplification are each designed to bind to a specific nucleic acid sequence at or near, typically within, a genomic region of interest (i.e. are target- specific primers) to generate target region amplicons. In some embodiments, methods herein include one or more universal amplifications.
[0108] A number of amplification technologies can be used with methods herein. For example, such amplification can be an isothermal amplification (e.g., recombinase polymerase amplification (RPA) (Kersting et al. 2014 Microchim Acta 181 (13-14), 1715-1723, (incorporated by reference in its entirety)), a ligase-based amplification, PCR, or a combination thereof (e.g., ligation- mediated PCR). In some illustrative embodiments, the targeted amplification is a targeted PCR(s) that is performed using a PCR reaction mixture that includes one primer pair, or in illustrative embodiments a set of primer pairs, and at least a portion of the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules or DNA molecules generated therefrom.
[0109] Typically, at least one primer of a primer pair used for targeted amplification herein, is a target- specific primer designed to bind to a specific nucleic acid sequence in or near, typically within, a genomic region of interest, which in illustrative examples can be genomic regions where epigenetic changes, such as changes in DNA methylation, are associated with or indicative of formation or presence of cancer, and can, for example, include promoter regions of tumor suppressor genes, and in some embodiments DNA methylation markers for a specific cancer type. A target- specific primer can be designed to bind to any sequence within or near a target region for amplification of the target region or a portion of the target region. In some embodiments, a target-specific primer can be designed to bind to a sequence that includes a CpG site. In other embodiments, a target- specific primer can be designed to bind to a sequence that is upstream or downstream to one or more CpG sites. One of the advantages of the methods described herein is increased flexibility in primer / probe design for targeted amplification or enrichment. In some embodiments, one primer of the one or more primer pairs or the set of primer pairs in the reaction mixture used for a targeted amplification is a universal primer and binds to a primer binding site on an adapter. Thus, for example, in such embodiments a universal primer that binds an adapter primer binding site can be used for an amplification reaction along with a target- specific primer that binds a primer binding site on a sample DNA region.
[0110] Target-specific primers typically define the ends of target region amplicons, which typically encompass at least a portion of the target region. In some embodiments, a PCR can be performed using two target-specific primers. The target region amplicon in such embodiment would extend from the sample DNA region bound by target- specific primer on a 5’ end to the sample DNA region bound by primer on the 3’ end. In some embodiments, a PCR can be performed using a universal primer and a target- specific primer. The target region amplicon in such embodiment would extend from the sample DNA region bound by target- specific primer on a 3’ end of one strand to the end of the sample DNA fragment on the 5’ end of that strand.
[0111] In some methods herein, a universal amplification of the library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules can be performed before the targeted amplification. Such universal amplification can be performed for example using a universal primer pair that binds primer binding sites in the adapter. Thus, in some embodiments, the methods herein include performing a universal PCR using the plurality of adapted methylation- transferred DNA molecules, and a universal PCR primer pair comprising primers designed to bind universal primer binding sequences on the adapters, before performing one or more targeted PCRs.
[0112] The one or more primer pairs in illustrative embodiments is a set of primer pairs. In some embodiments, the set of primer pairs is a set of between 2 and 1,000, 2 and 500, 2 and 250, 2 and 200, 2and 150, 2 and 100, 2 and 50 or 2 and 10 primer pairs, or between 5 and 1,000, 5 and 500, 5and 250, 5 and 200, 5 and 150, 5 and 100, 5 and 50 or 5 and 10 primer pairs, or between 50 and 250 or between 100 and 200 primer pars.
[0113] In some embodiments, at least one of the primer pairs comprises a universal primer and a target- specific primer. In some embodiments, at least one of the primer pairs comprises two targetspecific primers. In some embodiments, at least one of the primers comprises a sequencing tag. In some embodiments, at least one of the primers comprises a sample index. In some embodiments, at least one of the primers comprises biotin modification. In some embodiments, performing a PCR further comprises using primers comprising a sequencing tag. In some embodiments, performing a PCR further comprises using primers comprising a sample index. In some embodiments, the primers of the primer pairs are probe-dependent primers and the amplification is a target capture polymerase chain reaction. In some embodiments, the target regions each comprises one or more CpG sites differentially methylated in one or more cancers. In some embodiments, the set of loci comprises two or more loci that are differentially methylated in a different type of cancer from each other.
[0114] In some embodiments, one or both of the primer binding sites of a primer pair can include at least a portion of one of the adapter sequences. In some embodiments, one of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, both of the primer binding sites of a primer pair can include at least a portion of the adapter sequences. In some embodiments, neither of the primer binding sites of a primer pair include any of the adapter sequences.
[0115] Methods as described herein, in some embodiments, can include multiple amplification cycles (e.g., multiple PCR temperature cycles), and in some embodiments can include several sequential PCR reactions performed during the same set of temperature cycles. In some embodiments, amplification cycles can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 cycles. In some embodiments, amplification cycles can include at least 7, 8, 9, or 10 cycles. In illustrative embodiments, amplification cycles can include at least 11, 12, 13, 14, 15, 16, or 17 cycles.
[0116] Typically, in embodiments described herein, PCR amplification is performed by adding a PCR reaction mixture to the DNA template (e.g., methylation-transferred DNA molecules)followed by addition of a polymerase enzyme, and then amplified through multiple amplification cycles. In some embodiments, the PCR reaction mixture contains one or more primer pairs, deoxynucleotides (dNTPs), PCR reaction buffer, and deionized water. In some embodiments, the dNTPs can comprise a mixture of dATP, dCTP, dGTP and dTTP. In some embodiments, the final concentration of each dNTP in the reaction mixture can range from 0.05 mM to 0.5 mM dNTPs, for example, from 0.05 mM to 0.5 mM, 0.05 mM to 0.1 mM, 0.05 mM to 0.15 mM, 0.05 to 0.2 mM, 0.05 mM to 0.25 mM, 0.05 mM to 0.3 mM, 0.05 mM to 0.35 mM, 0.05 to 0.4 mM 0.05 mM to 0.45 mM, from .1 mM to 0.5 mM, 0.15 mM to 0.5 mM, or 0.2 mM to 0.5 mM, 0.25 to 0.5 mM, 0.3 mM to 0.5 mM, 0.35 mM to 0.5 mM, 0.4 mM 0.5 mM, or from 0.45 mM to 0.5 mM. In illustrative embodiments, the final concentration of each dNTP in the reaction mixture is between 0.15 mM and 0.25 mM. In some embodiments, the final concentration of each dNTP in the reaction mixture is 0.2 mM.
[0117] PCR buffer solution creates a suitable environment for the polymerase chain reaction and can contain many different components, including magnesium chloride (MgCh), potassium chloride (KC1), dimethyl sulfoxide (DMSO), and glycerin or bovine serum albumin (BSA). In some embodiments, the concentration of KC1 can be between 25 and 50 mM, between 25 and 75 mM, between 25 and 100 mM, between 30 and 100 mM, between 50 and 100 mM, or between 70 and 100 mM. In some embodiments, the MgC12 concentration can be in the range of 0.5 mM to 5 mM, between 0.5 mM to 4.5 mM, 0.5 to 4.0 mM, 0.5 to 3.5 mM, 0.5 to 3.0 mM, 0.5 to 2.5 mM, 0.5 to 2.0 mM, 0.5 to 1.5 mM, 1.0 to 5 mM, 1.5 to 5mM, 2.0 to 5 mM, 2.5 to 5 mM, 3.0 to 5 mM, 3.5 to 5 mM, 4.0 to 5 mM, or 4.5 to 5 mM. In some embodiments, the concentration of MgC12 is 2.0 mM.
[0118] In illustrative embodiments, the buffer solution is a Q5® Reaction Buffer (B9027S, New England Biolabs, Inc.). In some embodiments, the reaction buffer is Standard Taq Reaction Buffer (B9014S, New England Biolabs, Inc). In some embodiments, the reaction buffer is a Standard Taq (Mg-free) Reaction Buffer (B9015S, New England Biolabs, Inc.).
[0119] In some embodiments, a DNA polymerase is used to produce DNA amplicons using DNA as a template. In some embodiments, the polymerase is a Q5® DNA Polymerase, such as Q5® High-Fidelity DNA Polymerase (M0491S, New England BioLabs, Inc.) or Q5® Hot Start High-Fidelity DNA Polymerase (M0493S, New England BioLabs, Inc.). Q5® High-Fidelity DNA polymerase is a high-fidelity, thermostable, DNA polymerase with 3'— > 5' exonuclease activity, fused to a processivity-enhancing Sso7d domain. Q5® High-Fidelity DNA polymerase lacks 5'—> 3 'exonuclease activity and strand displacement activity.
[0120] In some embodiments, the polymerase is a T4 DNA polymerase (M0203S, New England BioLabs, Inc.). T4 DNA Polymerase catalyzes the synthesis of DNA in the 5'— 3' direction and requires the presence of template and primer. This enzyme has a 35 " exonuclease activity which is much more active than that found in DNA Polymerase I. T4 DNA polymerase lacks 5'— > 3' exonuclease activity and strand displacement activity.
[0121] In some embodiments of any of the aspects herein, the length of the primers can be between 10 to 100 nucleotides, such as between 10 to 75 nucleotides, 10 to 40 nucleotides, 10 to 35 nucleotides, 10 to 30 nucleotides, 10 to 20 nucleotides, 15 to 100 nucleotides, 20 to 100 nucleotides, from 25 to 100 nucleotides, from 30 to 100 nucleotides from 35 to 100 nucleotides, from 40 to 100 nucleotides, from 45 to 100 nucleotides, from 50 to 100 nucleotides, from 55 to 100 nucleotides, from 60 to 100 nucleotides, from 65 to 100 nucleotides, from 70 to 100 nucleotides, or from 75 to 100 nucleotides. In some embodiments, the range of the length of the primers is between 5 to 50 nucleotides, such as 5 to 40 nucleotides, 5 to 20 nucleotides, or 5 to 10 nucleotides. In some embodiments, the primers are between 5 and 50 bp in length, between 10 and 40 bp in length, between 15 and 30 bp in length, between 15 and 25 bp in length, between 20 and 40 bp in length, between 25 and 50 bp in length, or between 30 and 50 bp in length. In some embodiments, the primers are between 25 and 100 bp in length, between 35 and 100 bp in length, between 45 and 100 bp in length, between 55 and 100 bp in length, between 65 and 100 bp in length, or between 75 and 100 bp in length.
[0122] In some embodiments of any of the aspects or embodiments herein, the number of primer pairs can range from 1 to 100,000 primer pairs that each bind to one or more primer binding sequences. In some embodiments, the primer pairs are a part of a set of primer pairs. In some embodiments, the set of primers range from 2 to 100,000, from 2 to 10,000, from 2 to 1,000, from 2 to 100, from 2 to 50, from 10 to 100, from 50 to 100, from 100 to 200, from 100 to 500, from 100 to 1,000, from 100 to 10,000, from 100 to 100,000, from 1,000 to 100,00, or from 10,000 to100,000 primer pairs. In some embodiments, the number of primer pairs can range from 10 to 10,000, 10 to 1,000, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 15 to 30, or 15 to 25 primer pairs.
[0123] In some embodiments, PCR is used to generate very short amplicons. cfDNA (such as fetal cfDNA in maternal serum or necrotically- or apoptotically-released cancer cfDNA) is highly fragmented. For fetal cfDNA, the fragment sizes are distributed in approximately a Gaussian fashion with a mean of 160 bp, a standard deviation of 15 bp, a minimum size of about 100 bp, and a maximum size of about 220 bp. Methylation site(s) of interest may occupy any position from the start to the end among the various fragments originating from a particular locus. Because cfDNA fragments are short, the likelihood of a fragment of length L comprising both the forward and reverse primers sites is the ratio of the length of the amplicon to the length of the fragment. Under ideal conditions, assays in which the amplicon is 45, 50, 55, 60, 65, or 70 bp will successfully amplify from 72%, 69%, 66%, 63%, 59%, or 56%, respectively, of available template fragment molecules. Thus, in some embodiments target amplicons generated in method herein are between 40 and 100, 40 and 75, or 45 and 70 bp in length. In certain embodiments that relate most preferably to cfDNA from samples of individuals suspected of having cancer, the cfDNA is amplified using primers that yield a maximum amplicon length of 85, 80, 75 or 70 bp, and in certain preferred embodiments 75 bp, and that have a melting temperature between 50 and 65 °C, and in certain preferred embodiments, between 54-60.5°C. The amplicon length is the distance between the 5-prime ends of the forward and reverse priming sites. Amplicon length that is shorter than typically used by those known in the art may result in more efficient measurements of the desired methylation sites by only requiring short sequence reads. In an embodiment, a substantial fraction of the amplicons are between 25 on the low end of the range, and 100 bp, 90 bp, 80 bp, 70 bp, 65 bp, 60 bp, 55 bp, 50 bp, or 45 bp on the high end of the range.Probe-dependent primers
[0124] In some embodiments, one or more primers herein are probe-dependent primers. Probedependent primers (PDPs) have been disclosed (Pel, et al. “Rapid and highly-specific generation of targeted DNA sequencing libraries enabled by linking capture probes with universal primers” PLoS ONE 13(12):e0208283 (2018); WO 2017168332A1 “Linked duplex target capture”, which are hereby incorporated by reference in their entirety). Such embodiments can be considered LTCmethods. Briefly, in an LTC method, PDPs are designed to incorporate non-extendable capture probes linked 5’ to 5’ with a primer. Multiple linker types are possible as discussed below. Typically, probes of PDPs can be between 30 to 70 nucleotides in length, and include or comprise a 3’ inverted dT base to inhibit polymerase extension. In some embodiments, probes arc designed to cover the desired region with zero gap between forward and reverse probes. In some embodiments, the probes are between 20 and 100 nucleotides in length. In some embodiments, the size of the probe can be between 20 and 40 nucleotides, between 30 and 50 nucleotides, between 40 and 60, between 50 and 70 between 60 and 80, between 70 and 90, 80 and 100, 90 and 110, 100 and 120 nucleotides in length. In some embodiments, at least one of the probes of a PDP pair comprises a sample index.
[0125] In PDPs, forward and reverse probes can be designed to bind to nucleic acid sequences within or near a genomic region of interest on a sample DNA molecule to enrich nucleic acid molecules comprising the genomic region of interest or copies thereof. In some embodiments, at least one of the probe binding regions can include one or more CpG sites. In some embodiments, both of the probe binding sites of the probe binding regions can include one or more CpG sites.
[0126] Typically, the primer portion of a PDP is a universal primer designed to bind to a universal primer site on the appended adapter. In some embodiments, the PDP is designed with a sequencer binding sequence, such as an Illumina flow cell binding sequence, incorporated therein. In some embodiments, the sequencer flow cell binding sequence is between the probe and universal primer, and adjacent to the primer. Linked primers of the invention may also include sequencing tags to ensure that all cluster reads originate from the same linked template molecule. The lengths of the primers can be extended or shortened at the 5' end or the 3' end to produce primers with desired melting temperatures. Also, the annealing position of each primer pair can be designed such that the sequence and length of the primer pairs yield the desired melting temperature. In illustrative embodiments, the primer is a low melting temperature universal primer complementary to a portion of the ligated adapter.
[0127] The primer can be tailed or untailed depending on the specific requirements. In some embodiments, the universal primer comprises an A tail. In some embodiments, the universal primer is blunt ended. The length of the primers of the PDP can range from 5 to 40 nucleotides inlength. In certain embodiments, the PDP primers are between 10 and 25 nucleotides long. In embodiments, the primers of the PDP can range from 5 to 15 nucleotides, from 10 to 25 nucleotides, from 15 to 35 nucleotides, or from 25 to 40 nucleotides in length.
[0128] Typically, probe dependent primers comprise a linker between the probe and the primer. Probe and primer portions of the PDP are typically linked by a polyethylene glycol derivative, an oligosaccharide, a lipid, a hydrocarbon, a polymer, or a protein. In some embodiments, the linker is a PEG molecule, or derivative thereof. In some embodiments, the linker is an oligosaccharide. In some embodiments, the linker is a lipid. In some embodiments, the linker is a hydrocarbon. In some embodiments, the linker is a polymer. In some embodiments, the linker is a protein, or portion thereof.Detection & Analysis
[0129] Methods as described herein typically include determining the methylation status, somatic mutation status, fragmentomics, and / or quantifying nucleic acids, including DNA, cfDNA, methylation-transferred DNA molecules, enriched subsets of DNA having target regions, and target region amplicons or amplicons derived therefrom. In some embodiments, methylation- transferred DNA molecules prepared from the cfDNA from a blood sample from the individual is analyzed. Not to be limited by theory, cfDNA is believed to be released from certain cells, such as cancer cells, for example when they undergo necrosis or apoptosis. In some embodiments, methods herein can be used to detect methylation in target regions or nucleic acid sequences of interest, such as somatic mutations, that are present in a small percentage of DNA in a sample, such as cfDNA, for example from a fetus, a cell from a donated organ, or in illustrative embodiments, a cancer cell.
[0130] In methods described herein, determining the methylation status, somatic mutation status, fragmentomics, and / or quantifying nucleic acids is performed after treatment of methylation- transferred DNA molecules with a deaminating agent or MSRE / MDRE. In some embodiments, determining the methylation status, somatic mutation status, fragmentomics, and / or quantifying nucleic acids is performed after enrichment of the deaminated, or MSRE / MDRE treated, methylation-transferred DNA molecules, at genomic regions of interest. In some embodiments, the determining the methylation status, somatic mutation status, fragmentomics, and / orquantifying nucleic acids included in methods herein, can comprise performing a sequencing reaction on the deaminated, or MSRE / MDRE treated, methylation-transferred DNA molecules. In some embodiments, the sequencing reaction is a next-generation sequencing reaction.
[0131] DNA sequencing techniques, particularly high throughput next-generation sequencing techniques (often referred to as massively parallel sequencing techniques) such as those employed NOVASEQ (ILLUMINA), MISEQ (ILLUMINA), HISEQ (ILLUMINA), ION TORRENT (LITE TECHNOLOGIES), GENOME ANALYZER ILX (ILLUMINA), GS ELEX+(ROCHE 454) etc., can be used for determining the methylation status, somatic mutation status, fragmentomics, and / or quantifying nucleic acids prepared by the methods described herein. High throughput genetic sequencers are amenable to the use of barcoding (i.e., sample tagging with distinctive nucleic acid sequences) so as to identify specific samples from individuals thereby permitting the simultaneous analysis of multiple samples in a single run of the DNA sequencer. Methods as described herein that utilize NGS detection, in some embodiments can have an average depth of read of at least 0.1, 0.5, 1, 10, 50, 100, 200, 500, 1000, 2000, 2900, 3000, 3500, 4000, 5000, 10,000, 50,000, 75,000, 100,000, 130,000, 150,000, 175,000, or 200,000.
[0132] Methods herein can include analyzing data obtained from next-generation sequencing techniques. In some embodiments of methods herein, deaminated, or MSRE treated, methylation- transferred DNA molecules can be subjected to sequencing using next-generation sequencing techniques. Nucleic acid sequencing data can be generated for amplicons created by PCR, for example a multiplex targeted PCR (mPCR). In some embodiments, the multiplex PCR can be a tiled multiplex PCR. Lor a skilled artisan, algorithm design tools arc available that can be used and / or adapted to analyze the sequencing data. In addition, those skilled in the art can determine appropriate parameters for measuring alignment to a consensus sequence and / or to a known target region sequence, including any algorithms needed to achieve maximal alignment over the length of the sequences being compared.
[0133] Sequencing reads can be demultiplexed using an in-house tool and mapped using the Burrows-Wheeler alignment software, Bwa mem function (BWA, Burrows-Wheeler Alignment Software (see Li H. and Durbin R. (2010) Past and accurate long-read alignment with Burrows- Wheeler Transform. Bioinformatics.) on single end mode using pear merged reads to the hgl9genome. Amplification statistics QC can be performed by analyzing one or more of, but not limiting to, total reads, number of mapped reads, number of mapped reads on target, and number of reads counted.
[0134] Methods herein can include a background error model that can be constructed using normal, or healthy liquid samples, in illustrative embodiments, normal, or healthy plasma samples, which arc sequenced on the same sequencing run to account for run- specific artifacts. In some embodiments, 5, 10, 15, 20, 25, 30, 40, 50, 100, 150, 200, 250, or more than 250 normal, or healthy liquid samples, in illustrative embodiments, plasma samples can be analyzed on the same sequencing run. The number of samples that can be sequenced on the same sequencing run can be in the range of 5 to 500, 5 to 400, 5 to 300, 5 to 250, 20 to 250, 30 to 250, 50 to 250, 75 to 250, 100 to 250, 50 to 500, or 100 to 500. Sample barcodes are used in illustrative embodiments. In some illustrative embodiments, 20, 25, 40, or 50 normal samples (e.g., plasma samples) can be analyzed on the same sequencing run. Outlier samples can be iteratively removed from the model to account for noise and contamination. In some embodiments, samples with a Z score of greater than 5, 6, 7, 8, 9, or 10 are removed from the data analysis. For each base substitution of every genomic loci, the DOR weighted mean and standard deviation of the error can be calculated.
[0135] Methods herein can include calculating percent identity that can be calculated by determining the number of matched positions in aligned DNA sequences, dividing the number of matched positions by the total number of aligned DNA sequences, and multiplying by 100. A matched position refers to a position in which identical nucleotides occur at the same position in aligned DNA sequences. The percent identity over a particular length can be determined by counting the number of matched positions over that length and dividing that number by the length followed by multiplying the resulting value by 100. A non-limiting example for calculating the percent identity, can be, if (i) a 500-nucleotide DNA target sequence is compared to a subject DNA sequence, (ii) an alignment program presents 200 nucleotides from the target DNA sequence aligned with a region of the subject DNA sequence where the first and last nucleotides of that 200- nucleotide region are matches, and (iii) the number of matches over those 200 aligned nucleotides is 180, then the 500-nucleotide nucleic acid target sequence contains a length of 200 and a sequence identity over that length of 90 percent (i.e., 180, 200x100=90).
[0136] In some embodiments, the uniformity in DOR can be measured using standard methods such as, but not limiting to, DOR slope, normalized median depth of read (nmDOR), or breadth of read (BOR). DOR slope represents the slope of the line in the linear portion of a list of loci sorted in descending DOR order. Closer to zero is better, as it represents a flat line. In some embodiments, the uniformity in DOR can be measured using the percent of reads in the 90lh-95thpercentile. For this measurement, the loci are sorted in descending DOR order. In illustrative embodiments, a DOR distribution using the 90lh-95lhpercentile contains 5 percent of reads. The reads of all loci between the 90thpercentile and 95thpercentile can be counted and divided by the total reads for all loci.
[0137] In some embodiments, the magnitude of the DOR slope can be less than 0.005, 0.001, 0.0005, 0.0001, 0.00005, 0.00001, 0.000005, or 0.000001. The magnitude of the DOR slope can be between 0 and 0.005, such as 0.000001 to 0.005, such as between 0.000005 to 0.00001, 0.00001 to 0.00005, 0.00005 to 0.0001, 0.0001 to 0.0005, 0.0005 to 0.001, or 0.001 to 0.005. The percent of reads in the 90dl-95thpercentile can be between 0.2 and 9 percent, such as between 0.2 to 8 percent, 0.2 to 7 percent, 0.2 to 6 percent, 0.4 to 9 percent, 0.4 to 8 percent, 0.4 to 7 percent, 0.4 to 6 percent, 1 to 9 percent, 1 to 8 percent, 1 to 7 percent, 1 to 6 percent, 2 to 9 percent, 2 to 8 percent, 2 to 7 percent, 2 to 6 percent, 3 to 9 percent, 3 to 8 percent, 3 to 7 percent, 3 to 6 percent, 0.2 to 1.0 percent, 1 to 2 percent, 2 to 3 percent, 2 to 4 percent, 3 to 4 percent, 4 to 5 percent, 5 to 6 percent, or 6 to 8 percent, or 7 to 9 percent. In some embodiments of methods herein, the method or the amplification steps in the method can produce a composition comprising at least 100 different amplicons (e.g., at least 300, 500, 750, 1,000, 2,000, 5,000, 7,500, 10,000, 15,000, 19,000, 20,000, 25,000, 27,000, 28,000, 30,000, 40,000, 50,000, 75,000, or 100,000 non-identical amplicons) with the magnitude of the DOR slope in any of the ranges herein, or with a percent of reads in the 90lll-95lhpercentile in any of the ranges herein. In some embodiments, different amplicons can range in between 100 to 500,000, 100 to 400,000, 100 to 300,000, 100 to 200,000, 100 to 100,000, 100 to 75,000, 100 to 50,000, 100 to 40,000, 100 to 30,000, 100 to 25,000, 100 to 20,000, or 100 to 15,000 non-identical amplicons.
[0138] In some embodiments of methods herein, in addition, or, in some embodiments, as an alternative to analyzing an altered (increased or decreased) methylation levels in a sample, one or more other factors can be analyzed if desired. These factors can be used to increase the accuracyof the diagnosis (such as determining the presence or absence of cancer or an increased risk for cancer, classifying the cancer, or staging the cancer) or prognosis. These factors can also be used to select a particular therapy or treatment regimen that is likely to be effective in the subject.Error ModelsMethylation Amplification Error Model
[0139] In order to improve methylation detection from sequencing data, an error model may be incorporated into the methylation calling procedure that incorporates element specific priors from one or more steps of the workflow, for example from methylation preserving amplification and / or methylation conversion. In some embodiments, the element specific priors described herein may be determined based on published data. In some embodiments, the element specific priors described herein may be determined based on experimental analysis.
[0140] In some embodiments, the error model may incorporate priors based on particular motifs where DNMT1 may transfer methylation with a lower or higher likelihood. In some embodiments, the error model may incorporate priors based on motifs where DNMT1 may add a methyl group where a methyl group did not exist on the other strand. In some embodiments, the error model may incorporate priors based on the overall efficiency of DNMT1 (e.g., the fraction of mC that remain as mC after DNMT1 treatment, per cycle, regardless of motif). In some embodiments, the error model may incorporate priors based on bias of DNMT1 with respect to fragment length and number of CpG sites. In some embodiments, the error model may incorporate priors based on bias of DNMT1 with respect to mismatches (e.g., due to polymerase errors). In some embodiments, the error model may incorporate priors based on the overall non-specific activity of DNMT1 (e.g., the fraction of C that DNMT1 converts from C to mC inadvertently). In some embodiments, the error model may incorporate priors based on motifs where BST has lower copying efficiency, including methylation motifs (e.g., in instances where CmCGA is challenging for BST to amplify with a methyl group present). In some embodiments, the error model may incorporate priors based on motifs where BST make incorrect base insertions. In some embodiments, the error model may incorporate priors based on performance as a function of the methyl amplification time and temperature (e.g., long time and higher temperature may result in higher / different efficiency and motifs). In some embodiments, motifs may be 1 bp, 2, bp, 3 bp, 4 bp, 5 bp, or longer.
[0141] In some embodiments, the error model may incorporate any combination of the priors described herein depending on the workflow implemented. For example, EM-Seq and BS-Seq error models may each have bias towards specific sequences or template length.
[0142] In some embodiments, the methods described herein may include a position-based methylation error model. A cohort of healthy plasma samples can be used for error model construction. In some embodiments, 10 healthy samples, 50 healthy samples, 100 healthy samples, 200 healthy samples, 500 healthy samples, or 1000 healthy samples may be used for error model construction.
[0143] In some embodiments, methylation transfer efficiency and error rates (spontaneous methylation) are estimated for each target in the panel of differentially methylated regions. Using a set of normal samples that are not expected to have any phenotype-related differential methylation, the per position efficiency and error rate per cycle can be estimated. The error model accounts for DNA input amount by adjusting the number of library prep cycles based on input amount.Mean Sample Methylated Allele Frequency (MAF)
[0144] Methods herein can include quantification of ctDNA in a sample, such as an MRD sample, by calculating a sample MAF. In some embodiments, only highly methylated CpG sites are used for calculating a mean sample MAF. In some embodiments, statistical weighting is further applied. In some embodiments, mean sample MAF may be determined by computing the weighted average of MAFs, where each differentially methylated site’s contribution is weighted by its corresponding prior probability, incorporating prior knowledge or assumptions about the likelihood of differentially methylated site’s presence in specific genes. In some embodiments, mean sample MAF may be determined by computing the weighted average of MAFs, where each differentially methylated site’s contribution is weighted by its corresponding gene posterior probability, reflecting the confidence in the variant calls based on posterior distribution. In some embodiments, the observed MAF for highly methylated CpG sites may undergo adjustment to account for the background error rate. In some embodiments, the background error rate may be derived from a position-based error model. In some embodiments, the average observed error issubtracted from the observed MAF to obtain an adjusted value. In some embodiments, the mean sample MAF is only calculated if a positive call is made for the sample.Methylation Calling
[0145] In some embodiments, a methylation caller algorithm calculates a confidence score using the likelihoods generated from the error model and combining them with uniform priors across a grid of ctDNA amounts. In some embodiments, the methylation caller calculates a likelihood for each of the targets in a beta binomial model. The confidence score is then calculated by getting the maximum likelihood across that grid, and dividing it by that (max likelihood+negative likelihood).
[0146] Training Data: Di k=denotes
[0147] Test Data:
[0148] Result: Mutation call confidence scores for non-reference alleles in the test set for all bases 1,2,...,B.
[0149] For i = 1,2, ..., B do1. Estimate efficiency and error from training data for base z, using the data Di k. Estimate methylation transfer and error rate parameters using training data: Transfer rate (p); Error rate (pe).2. Estimate starting copy for base z for test data at base i. Xo, the total number of starting fragments at a given base.
[0150] In some embodiments, an improved methylation caller algorithm as disclosed herein can be used. This algorithm provides an improved way of calculating confidence scores, rescaling it to significantly differentiate between positive and negative confidences, and allows the addition of non-uniform priors. This would result in more stable and more accurate confidences.
[0151] In some embodiments, in the improved variant caller algorithm, likelihoods are combined with priors, derived from training data ahead of time as described above, for smoother, more realistic and easily adjusted confidences.
[0152] Using the same error model as previously described, for each differentially methylated allele, at each position p, the likelihood of data at allele a, position p, MAF v L!K(D(a, p) |v).
[0153] Smoothed allele likelihood ratio: Computed by summing over MAF likelihoods scaled by MRD MAF priors.
[0154] Position likelihood ratio: Compute the weighted position likelihood ratio rat(p) assuming only one of the alleles at a position is the true differentially methylated allele.
[0155] Sample positive probability: Recursively compute the sample positive probability, as the probability of at least one target being positive, and find combined probability of sample being positive, taking into account methylation hotspots, via methylation priors, and control sample outcome via sample prior.
[0156] In some embodiments, sample prior is input to the algorithm and will be adjusted to achieve desired false positive, false negative and no call rate. Allele and position priors are calculated from MRD data positive prevalence and can always be adjusted given additional information. MAF priors are calculated from MAF rates of positive MRD samples.SNV Error Model
[0157] In some embodiments, wherein a methyl amplified library has been divided into two aliquots - one for methylation detection, and the other for SNV detection, the methods described herein may also include a position-based somatic variant error model, e.g., an SNV / indel error model. A cohort of healthy plasma samples can be used for error model construction. In some embodiments, 10 healthy samples, 50 healthy samples, 100 healthy samples, 200 healthy samples, 500 healthy samples, or 1000 healthy samples may be used for error model construction.
[0158] In some embodiments, per cycle PCR efficiency and error rates are estimated for each target in the panel. Using a set of normal samples that are not expected to have any phenotype related mutations, the per position efficiency and error rate per cycle can be estimated. The error model accounts for DNA input amount by adjusting the number of library prep cycles based on input amount.Mean Sample VAF
[0159] Methods herein can include quantification of cfDNA such as ctDNA in a sample, such as an MRD sample, by calculating a sample VAF. In some embodiments, only consensus variants are used for calculating a mean sample VAF. In some embodiments, statistical weighting is further applied. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant’s contribution is weighted by its corresponding gene prior probability, incorporating prior knowledge or assumptions about the likelihood of variant presence in specific genes. In some embodiments, mean sample VAF may be determined by computing the weighted average of variant VAFs, where each variant's contribution is weighted by its corresponding gene posterior probability, reflecting the confidence in the variant calls based on posterior distribution. In some embodiments, the observed VAF for consensus SNVs may undergo adjustment to account for the background error rate. In some embodiments, the background error rate may be derived from a position-based error model. In some embodiments, the average observed error is subtracted from the observed VAF to obtain an adjusted value. In some embodiments, the mean sample VAF is only calculated if a positive call is made for the sample.SNV Calling
[0160] In some embodiments, a variant caller algorithm calculates a confidence score using the likelihoods generated from the error model and combining them with uniform priors across a grid of ctDNA amounts. In some embodiments, the variant caller calculates a likelihood for each of the targets in a beta binomial model. The confidence score is then calculated by getting the maximum likelihood across that grid, and dividing it by that (max likelihood+negative likelihood).
[0161] Training Data: Di k=denotes
[0162] Test Data: Di k= (R^ RefAllele^A^^ C^^ G^^ Tl^for i = 1, 2, ..., B
[0163] Result: Mutation call confidence scores for non-reference alleles in the test set for all bases 1,2,...,B.
[0164] For i = 1,2, ..., B do
[0165] 1. Estimate efficiency and error from training data for base i, using the data Di k. Estimate PCR parameters using training data: Replication rate (p); Error rate (pP).2. Estimate starting copy for base z for test data at base i. Xo, the total number of starting fragments at a given base.
[0166] In some embodiments, an improved variant caller algorithm as disclosed herein can be used. This algorithm provides an improved way of calculating confidence scores, rescaling it to significantly differentiate between positive and negative confidences, and allows the addition of non-uniform priors. This would result in more stable and more accurate confidences.
[0167] In some embodiments, in the improved variant caller algorithm, likelihoods are combined with priors, derived from training data ahead of time as described above, for smoother, more realistic and easily adjusted confidences.
[0168] Using the same error model as previously described, for each mutation allele, at each position p, the likelihood of data at allele a, position p, VAF v LIK D(a, p)|v).
[0169] Smoothed allele likelihood ratio: Computed by summing over VAF likelihoods scaled by MRD VAF priors.
[0170] Position likelihood ratio: Compute the weighted position likelihood ratio rat{p) assuming only one of the alleles at a position is the true mutation allele.
[0171] Sample positive probability: Recursively compute the sample positive probability, as the probability of at least one target being positive, and find combined probability of sample being positive, taking into account hotspots, via position priors, and control sample outcome via sample prior.
[0172] In some embodiments, sample prior is input to the algorithm and will be adjusted to achieve desired false positive, false negative and no call rate. Allele and position priors are calculated from MRD data positive prevalence and can always be adjusted given additional information. VAF priors are calculated from VAF rates of positive MRD samples.InDei Calling
[0173] In some embodiments, an InDel-based sample caller combines position-based InDei calls with context-based InDei calls to generate InDel-based sample calls. In some embodiments, the sample calls generated by this caller are combined with the SNV calls to generate a combined sample-level call.Position-based InDei caller
[0174] In some embodiments, position-based calling involves training a target specific error model for each InDei target that is encountered in the test sample. In one example, the error model is trained using a cohort of 200 healthy donor samples.
[0175] In some embodiments, the position-based caller comprises (1) target consolidation, (2) outlier detection using a beta binominal distribution, and (3) likelihood from tail probability.
[0176] In some embodiments, targets are consolidated to account for alignment differences that may result in incorrect characterization of target error rates, and outliers are detected using a beta binomial distribution.
[0177] The probability mass function (PMF) for the binomial distribution is:for k E {0,1, ... , n], n > 0, a > 0, b > 0, where B(a, b~)is the beta function.
[0178] where the beta-binomial distribution is a binomial distribution, whose probability of success p follows a beta distribution B(a, 0). A Beta distribution is fitted to the error model target VAFs, by computing the maximum likelihood estimates for a and 0.
[0179] Log likelihood ratio for the target is then computed using the following formula:LogLikelihoodRatio = -log(test tail probability / train tail probability)
[0180] Target posterior probability can then be calculated as:PosteriorProbability = eLLR / ( 1 + eLLR)
[0181] Finally, in some embodiments, only replicate consensus targets that are found to have sufficient support in two plasma replicates are called.Context-based InDei caller
[0182] Insertions and deletions that do not occur in a repeat context are likely to have very low rates of background error. The relatively small number of training samples used to fit the error model is unable to characterize the background error profile for such targets, resulting in poor fit, or fit failure. For such targets a context-based calling strategy is used.
[0183] In some embodiments, the context-based caller comprises features to fit a Beta-Binomial regression model, which is used to compute target confidence.
[0184] The context-based caller is primarily used to call less noisy targets where position-based calling is not feasible.Ensemble InDei caller or combined position and context InDei caller
[0185] In some embodiments, the ensemble caller combines position- and context-based InDei calls to generate a single InDei target call.
[0186] In some embodiments, the ensemble caller uses the position-based call for InDei targets when the Beta-Binomial background error fit is successful and meets goodness of fit (GoF) criteria.
[0187] In some embodiments, the ensemble caller uses the context-based call for InDei targets when Beta-Binomial background error fit fails, Beta-Binomial fit does not meet GoF criteria, or the number of non-zero mutant DOR error models samples is below threshold, i.e., only use context-based calling for low noise targets.Sample InDei calling
[0188] In some embodiments, InDcl sample calling uses an approach very similar to SNV sample caller, with primary difference being InDei specific WES / W GS priors, which are generated using filtered InDei targets for the same CRC sample cohort as was used for SNVs.
[0189] The sample caller algorithm recursively computes the sample likelihood ratio incorporating target likelihoods and target WES / WGS priors.Consensus Calling
[0190] PCR and sequencing artifacts are a result of random process noise and are expected to be discordant between library pools. Accordingly, consensus calling helps filter a significant fraction of such discordant targets.
[0191] In some embodiments, the methods described herein are used to filter targets that are likely a result of process noise and generate a candidate target set that will be input to the sample caller to compute the sample posterior positive probability. Because the sample level call is made by the sample caller by integrating target WES / WGS prior data, this version of the target caller is intentionally tuned to be more permissive than previous versions. The primary objective here is to filter targets lacking sufficient support or likely to be a result of random process noise.
[0192] In some embodiments, the variant caller described herein uses data likelihoods, target VAFs and MRD VAF priors, to first call targets separately for each library pool, followed by a combined call using both library pools.Comparison of Two or More Samples
[0193] Methods herein, in some embodiments, can include preparing 2 or more libraries of methylation-transferred DNA molecules. In some embodiments, methods as described herein can include preparing 3, 4, 5, 6, 7, 8, 9, 10 or more libraries of methylation-transferred DNA molecules. In some embodiments, methods as described herein can include preparing 2 libraries of methylation-transferred DNA molecules, such methods further comprise performing the method on a second sample. In some embodiments, the second sample is from a second subject. In some embodiments, the second sample is derived from a cell line. In some embodiments, the method is performed on the first and second samples simultaneously. In some embodiments, methods as described herein can include comparing 2 libraries of methylation-transferred DNA molecules. In some embodiments, such methods further comprise determining and comparing the methylation status of a genomic region of interest in each of the libraries. In some embodiments, such methodsfurther comprise determining and comparing the somatic mutation status in each of the libraries. In some embodiments, a first library is derived from a subject suspected or at risk of having a disease, and a second library is derived from a subject that is not suspected or at risk of having the disease. In some embodiments, the first and second libraries are derived from the same subject. In some embodiments, the library is derived from a sample collected at a first timepoint and the second library is derived from a sample collected at a second timepoint. In some embodiments, the first timepoint precedes the second timepoint. In some embodiments, the method comprises comparing the methylation status and / or somatic mutation status of a genomic region of interest derived from the first sample to that from a second sample using the same method. In some embodiments, the determining comprises comparing the methylation status and / or somatic mutation status of a genomic region of interest to a preset threshold amount.Kits
[0194] Additional embodiments of the invention described herein relate to kits for preparing nucleic acids from a biological sample using the methods described herein. In some embodiments, the kits may comprise one or more of nucleotides, e.g., dNTPs, polymerase, e.g., BST polymerase, and a methyl transfer agent, e.g., DNMT1, necessary to carry out the claimed invention. In some embodiments, the kit may comprise primers, such as universal primers, or primers specific for target regions.
[0195] In some embodiments, the kit may further comprise reagents necessary for nucleic acid processing. For example, the kit may comprise adapters, barcodes, molecular indexes, sample indexes. In some embodiments, the kit may comprise ligases and or primers sued to append the barcodes and indexes to the template DNA.
[0196] In some embodiments, the kit may further comprise reagents necessary for downstream applications. For example, in some embodiments, the kits may comprise reagents necessary for enrichment of target regions, such as hybrid capture probes or target specific oligonucleotides. In some embodiments, the kit may comprise reagents required for methylation analysis such as deaminating reagents or MSREs. In some embodiments, the kits may comprise reagents necessary for somatic mutation analysis, such as probes and primers. In some embodiments, the kits may comprise reagents necessary for sequencing reactions.
[0197] In some embodiments, the kit may further comprise tubes, tools, or devices necessary for performing the methods described herein. In some embodiments, the kit may further comprise instructions for performing the methods described herein. In some embodiments, the kits may include one or more devices, such as a microfluidic device, on which the method can be performed. In some embodiments, the kits may be used in combination with the diagnostic box, as disclosed herein.Diagnostic Box
[0198] In an embodiment, the present disclosure comprises a diagnostic box that is capable of partly or completely carrying out any of the methods described in this disclosure. In an embodiment, the diagnostic box may be located at a physician’s office, a hospital laboratory, or any suitable location reasonably proximal to the point of patient care. The box may be able to run the entire method in a wholly automated fashion, or the box may require one or a number of steps to be completed manually by a technician. In an embodiment, the box may be able to analyze at least the genotypic data measured on the samples from the subject. In an embodiment, the box may be linked to means to transmit the genotypic data measured on the diagnostic box to an external computation facility which may then analyze the genotypic data, and possibly also generate a report. The diagnostic box may include a robotic unit that is capable of transferring aqueous or liquid samples from one container to another. It may comprise a number of reagents, both solid and liquid. It may comprise a high throughput sequencer. It may comprise a computer.
[0199] While the embodiments of the present disclosure are amenable to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail below. The intention, however, is not to limit the disclosure to the particular embodiments described. On the contrary, the disclosure is intended to cover all modifications, equivalents, and alternatives falling within the scope of the disclosure as defined by the appended claims.
[0200] The terms and expressions which have been employed herein are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed.Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects, exemplary aspects and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims. The specific aspects provided herein are examples of useful aspects of the present invention and it will be apparent to one skilled in the art that the present invention may be carried out using a large number of variations of the devices, device components, methods steps set forth in the present description. As will be obvious to one of skill in the art, methods and devices useful for the present methods can include a large number of optional composition and processing elements and steps.
[0201] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the invention pertains. References cited herein are incorporated by reference herein in their entirety to indicate the state of the art as of their publication or filing date and it is intended that this information can be employed herein, if needed, to exclude specific aspects that are in the prior art. For example, when composition of matter are claimed, it should be understood that compounds known and available in the art prior to Applicant's invention, including compounds for which an enabling disclosure is provided in the references cited herein, are not intended to be included in the composition of matter claims herein.
[0202] One of ordinary skill in the art will appreciate that starting materials, biological materials, reagents, synthetic methods, purification methods, analytical methods, assay methods, and biological methods other than those specifically exemplified can be employed in the practice of the invention without resort to undue experimentation. All art-known functional equivalents of any such materials and methods arc intended to be included in this invention. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by illustrative aspects and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled inthe ai t, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[0203] The disclosed embodiments, examples and experiments are not intended to limit the scope of the disclosure or to represent that the experiments below are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperature, etc.) but some experimental errors and deviations should be accounted for. It should be understood that variations in the methods as described may be made without changing the fundamental aspects that the experiments are meant to illustrate.
[0204] Those skilled in the art can devise many modifications and other embodiments within the scope and spirit of the present disclosure. Indeed, variations in the materials, methods, drawings, experiments, examples, and embodiments described may be made by skilled artisans without changing the fundamental aspects of the present disclosure. Any of the disclosed embodiments can be used in combination with any other disclosed embodiment.
[0205] In some instances, some concepts have been described with reference to specific embodiments. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the invention as set forth in the claims below. Accordingly, the specification and figures are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of invention.
[0206] The following non-limiting examples are provided purely by way of illustration of exemplary embodiments, and in no way limit the scope and spirit of the present disclosure. Furthermore, it is to be understood that any inventions disclosed or claimed herein encompass all variations, combinations, and permutations of any one or more features described herein. Any one or more features may be explicitly excluded from the claims even if the specific exclusion is not set forth explicitly herein. It should also be understood that disclosure of a reagent for use in a method is intended to be synonymous with (and provide support for) that method involving the use of that reagent, according either to the specific methods disclosed herein, or other methods known in the art unless one of ordinary skill in the art would understand otherwise. In addition, where the specification and / or claims disclose a method, any one or more of the reagents disclosedherein may be used in the method, unless one of ordinary skill in the ail would understand otherwise.EXAMPLESExample 1. Primer extension and methyltransfer
[0207] As shown in FIG. 2, for example, adapters (204) are appended to the ends of the template DNA molecules (206, 208), which include methylated cysteine residues. The adapters may contain a sample barcode (210) and a universal binding sequences (212) The top (206) and bottom (208) strands of the template DNA molecules are copied by primer extension using a universal primer (202) which hybridizes to a universal binding sequence (212) in the adapter. Primer extension generates two duplex DNA molecules, each comprising a methylated strand of the original template DNA molecule (218, 220) and an unmethylated copy of the template DNA molecule (214, 216). The methylated strand of the original template DNA molecule (218, 220) of each duplex DNA molecule is then used as the template in the methyltransferase reaction to transfer the methylation status from the template DNA molecules to the copied DNA molecules (222, 224).Example 2. MTase Amplification Model
[0208] Molecular recovery vs amplification was modelled as a Poisson process (i.e., molecular sampling during the conversion step). As demonstrated in FIG. 8, even a 2-fold amplification of the template DNA molecules results in increased recovery of greater than 30%. A 10-fold amplification increases molecular recovery to approximately 60%.
[0209] A simple python model was created to estimate the conversion rate and protection rate of the amplification and methyltransferase step of the methods described herein, prior to methylation detection (by, for example, bisulfite or enzymatic conversion), or other subsequent analysis (therefore, chemical or enzymatic conversion / protection rates are additive to modeled rates). The false positive rate (FPR) is the rate at which the sequence consensus is ‘mC’ when it should actually be ‘C’ . The false negative rate (FNR) is the rate at which the sequence consensus is ‘C’ when it should actually be ‘mC’. Conversion rate = 1-FPR. Protection rate = 1-FNR. Mtase_eff = fraction ‘mC’ that methyltransferase copies as ‘mC’ correctly (i.e., methyltransferase efficiency).Mtase_ns = fraction of ‘C’ that methyltransferase converts from ‘C’ to ‘mC’ inadvertently (i.e., non-specific methyltransferase efficiency). While a python model was used for the analysis shown in FIG. 8, any other programming languages can be used.
[0210] For the conversion rate model, the conditions used are: 6 amplification cycles, 50% amplification efficiency per cycle (which results in ~11 fold amplification), and 0% methylated starting template. As shown in FIG. 9, conversion rate should be >99% to match typical conversion rate with bisulfite treatment. This is expected to be possible if the non-specific activity rate (i.e., mistakenly converting ‘C’ to ‘mC’) of MTase is <~3%.
[0211] For the protection rate model, the conditions used are: 6 amplification cycles, 50% amplification efficiency per cycle (which results in ~11 fold amplification), and 100% methylated starting template. As shown in FIG. 10, the protection rate should be >96% to match typical protection rate with BS treatment. This is accepted to be possible if the transfer rate (i.e., transfer of ‘mC’ to ‘mC’) of MTase is >95%.Example 3. Temperature compatibility testing for Bst2.0 polymerase and DNMT1
[0212] As explained above, it is important to select a polymerase based on the particular activity levels at specific temperatures, such that there will be little amplification occurring when the methyl transfer step is being performed. Additionally, conditions during primer extension should not inactivate the DNMT 1. It is also important to select a buffer in which both the polymerase and methyltransferase will function.
[0213] In this Example, inventors analyzed various incubation times and temperatures at which Bst2.0 polymerase functions while DNMT1 is not simultaneously inactivated. A buffer composition in which both DNMT1 and Bst2.0 Warmstart polymerase work was determined to be DNMT1 buffer with 2 mM added MgSO4 and 10 mM added KC1. Note that while DNMT1 was found to function in this buffer, it had between 25% and 50% lower activity at 1 hour than in its recommended buffer (0 magnesium, 0 potassium).
[0214] Each enzyme was first tested separately under different conditions. Hemimethylated (DNMT1 test) and unmethylated (Bst Polymerase test) 150 bp constructs with “stubby” Y- adaptors as shown in FIG. 11 were used.Bst2.0 WarmStart Activity! Test
[0215] Bst Polymerase activity was determined by TapeStation and gelshift band analysis. The standard protocol calls for Bst2.0 WarmStart polymerase to be incubated at 55°C for 5 min. Here, Bst2.0 WarmStart activity was tested at 30 sec, 2 min, 5 min, and 10 min at 40°C, 45°C, 50°C, and 55°C. Bst2.0 WarmStart is not supposed to be active below 45°C due to a bound aptamer, however, residual activity was found at 37°C over long periods of time. As DNMT1 cannot survive above 55 °C, higher temperatures were not tested.
[0216] As shown in FIG. 12, Bst2.0 WarmStart was incubated for the stated times at the stated temperatures. A band at -150 bp indicates successful amplification. Unamplified products should run at between 200 and 300 bp, but because the Y-adaptors cause them to run in a diffuse band, it appears that they were below the detection threshold for the instrument. The results show that Bst2.0 WarmStart is capable of complete amplification of the 150 bp template in 5 min at any of the tested temperatures. However, it does not amplify effectively in 30 sec at less than 55°C.Bst2.0 WarmStart Activity 1-Pot Test
[0217] The assay was then repeated as a “1-pot” reaction, i.e., with Bst2.0 WarmStart and DNMT1 in the same reaction volume. The reactions were prepared as stated in Table 1 below. The reactions were incubated at the stated times and temperatures for Bst2.0 WarmStart activity, then placed at 37°C for 2 hours to promote DNMT1 activity. DNMT1 from two different sources, Active Motif and Sigma- Aldrich, were tested. Samples were then cleaned up using QiaQuick spin columns. One of the two replicates at each condition were ran on the TapeStation to confirm single-cycle strand displacement primer extension (FIG. 13). The presence of a 150 bp band indicates Bst2.0 WarmStart activity. All temperature conditions appear to allow for complete replication of the template by Bst2.0 WarmStart, including the sample placed directly at 37 °C. This indicates that some Bst2.0 WarmStart activity is occurring during DNMT1 incubation, even though Bst2.0 WarmStart is supposed to have an aptamer that prevents its activity below 45°C.Table 1. Conditions used in 1-pot testDNMT1 Activity 1-Pot Test
[0218] The Assay was repeated as above, and all samples were then incubated for 4 days with Sall in CutSmart buffer. The reactions were then run on the TapeStation (FIG. 14). Although the TapeStation experienced errors in detecting the internal standards for some reactions, making absolute band quantification unreliable, the relative within-well concentration values should be consistent, therefore the percent cut could be calculated for almost all wells (Table 2). DNMT1 activity is indicated by a lower ratio of the lower -75 bp band to the upper -150 bp band (lower percentage cut). Samples in which the DNMT1 was added after Bst incubation was completed are the positive control for DNMT1 activity (“Post”), while the samples without DNMT1 are negative controls for DNMT 1 activity.
[0219] Results for incubation at 45°C for either 5 or 10 min before DNMT1 incubation at 37°C (# 5*-9*, Table 2) were indistinguishable from the results for the reaction where DNMT1 was added after the reaction had cooled from a 5 min Bst incubation at 55°C (#19pand 20p, Table 2), i.e., thepositive controls. Interestingly, even though the Bst reaction appears to have gone to completion even with low or no Bst-specific incubation, DNMT1 activity was lower without a Bst incubation at least 45°C. It is hypothesized that at very low temperatures, the Bst reaction is very slow and / or that the Bst does not detach from the template efficiently. Both of these situations could result in DNMT1 having less time with full access to a finished hemimethylated target thus reducing DNMT1 activity.
[0220] It was also noted that the Sigma- Aldrich DNMT1 (S-A in FIG. 13) did not perform well in any of the tested conditions. This could be because it is a less active enzyme, or that it requires a different buffer.Table 2. Quantification of 1-pot Bst2.0 WarmStart and DNMT1 reaction after Sall digestion to assess DNMT1 activity.p= Positive control; * = most effective test conditions.Conclusion
[0221] As shown in summary FIG. 15, the present assay demonstrates a buffer and incubation conditions at which both amplification and methyl transfer can occur. The positive control and test reactions highlighted in FIG. 15 show that DNMT was able to propagate methylation to a newly synthesized DNA strand in the shared buffer. The results also demonstrate that DNMT1 did not methylate unmethylated CpG sites, as shown by the highlighted negative controls in FIG. 15, which were not protected from Sall digestion.
Claims
What is claimed is:
1. A method of preparing a non-naturally occurring preparation of DNA from a subject useful for determining a methylation status of a genomic region of interest, comprising:(a) extracting DNA from a sample of a subject;(b) copying one or more template DNA molecules from the extracted DNA or its derivative, thereby generating one or more copied DNA molecules;(c) transferring methylation status from the one or more template DNA molecules to the one or more copied DNA molecules, thereby generating one or more methylation-transferred DNA molecules; and(d) optionally repeating steps (b) and (c) one or more times, thereby generating a library of amplified DNA molecules comprising the one or more methylation-transferred DNA molecules.
2. The method of claim 1, wherein the copying in step (b) is by isothermal extension.
3. The method of claim 2, wherein the isothermal extension is performed at about 58-64 °C.
4. The method of claim 2 or 3, wherein the isothermal extension is performed at about 60 °C.
5. The method of any of claims 1-4, wherein the methylation status is transferred using a methyltransferase.
6. The method of any of claims 1-5, wherein the methylation status of at least one of the one or more template DNA molecules is preserved in one or more of the methylation-transferred DNA molecules.
7. The method of any of claims 5-6, wherein the template and copied DNA molecules are treated with the methyltransferase at about 30-40 °C.
8. The method of any of claims 5-7, wherein the template and copied DNA molecules are treated with the methyltransferase at about 35 °C.
9. The method of any of claims 1-8, wherein steps (b) and (c) are repeated for 1-20 cycles.
10. The method of any of claims 1-8, wherein steps (b) and (c) are repeated for 1-10 cycles.
11. The method of any of claims 1-6, wherein steps (b) and (c) are repeated for 1-5 cycles.
12. The method of any of claims 1-11, wherein the copying is performed using a DNA polymerase having reduced activity below 37 °C.
13. The method of claim 12, wherein the reduced activity is below 15%.
14. The method of claim 12 or 13, wherein the polymerase is BST polymerase or derivative thereof.
15. The method of claim 12 or 13, wherein the polymerase is a Taq polymerase or derivative thereof.
16. The method of any of claims 1-15, wherein the methyl transferase is DNMT1 or derivative thereof.
17. The method of any of claims 1-16, wherein step (b) is performed for a shorter time than step (c).
18. The method of any of claims 1-17, wherein step (b) is performed for between 1 and 10 minutes.
19. The method of cany of claims 1-18, wherein step (c) is performed for between 5 and 60 minutes.
20. The method of any of claims 1-20, further comprising appending adapters to the extracted DNA prior to the copying in step (b).
21. The method of claim 20, wherein the adapters are Y-adapters.
22. The method of claim 20 or 21, wherein the adapters comprise a universal priming site.
23. The method of any of claims 20-22, wherein the adapters comprise methylated cytosines.
24. The method of any of claims 20-22, wherein the adapters comprise at least one methylated cytosine and at least one unmethylated cytosine.
25. The method of any of claims 20-24, wherein the adapters comprise a molecular barcode.
26. The method of claim 25, wherein the molecular barcode is used to construct a methylation consensus.
27. The method of claim 25 or 26, wherein the molecular barcode is used to identify an error in isothermal amplification, methyltransferase treatment, or bisulfite or enzymatic conversion.
28. The method of any of claims 20-27, wherein the adapters do not include any recognition site for any methylation sensitive restriction enzymes (MSRE) or methylation dependent restriction enzymes (MDRE).
29. The method of any of claims 22-28, wherein the isothermal amplification is performed with at least one primer that binds to the universal priming site.
30. The method of claim 29, wherein the primer comprises locked nucleic acid (LNA) oligonucleotides .
31. The method of claim 29, wherein the primer comprises 5’ linked primers.
32. The method of any of claims 1-31, further comprising contacting at least a portion of the library of amplified DNA molecules or their derivatives with a deaminating agent to generate treated DNA molecules.
33. The method of claim 32, wherein the deaminating agent is sodium bisulfite.
34. The method of claim 32, wherein the deaminating agent is an enzyme that converts cytosine to uracil.
35. The method of any of claims 1-31, further comprising contacting at least a portion of the library of amplified DNA molecules with one or more MSREs and / or MDREs to generate treated DNA.
36. The method of claim 35, wherein the one or more MSREs are selected from one or more of Hpall, Hhal, HpyCH41V, or BstUl.
37. The method of claim 35, wherein the one or more MDREs are selected from one or more of AbaSI, Glal, McrBC, or MspJI.
38. The method of claim 35 or 36, wherein the contacting comprises contacting with two or more MSREs and / or MDREs.
39. The method of any of claims 35-38, wherein the contacting comprises contacting with two or more MSREs and / or MDREs in a single reaction.
40. The method of any of claims 35-39, wherein the contacting comprises contacting with three or more MSREs and / or MDREs.41 . The method of any of claims 35-40, wherein the contacting comprises contacting with four or more MSREs and / or MDREs.
42. The method of any of claims 34-41, further comprising enriching a plurality of target loci in the treated DNA.
43. The method of claim 42, wherein the treated DNA is enriched using hybrid capture probes to generate enriched DNA.
44. The method of claim 42, wherein the treated DNA is enriched using PCR to generate enriched DNA.
45. The method of claim 44, wherein the PCR is targeted multiplex PCR.
46. The method of any of claims 42-45, wherein the target loci comprise 50-50,000 target loci.
47. The method of any of claims 42-46, wherein the target loci are differentially methylated in cancer.
48. The method of any of claims 1 -47, further comprising performing high-throughput sequencing on the enriched DNA to generate sequencing data.
49. The method of claim 48, further comprising using the sequencing data to determine the methylation status of the plurality of target loci.
50. The method of claim 48 or 49, further comprising using the sequencing data to determine the presence of somatic mutations.
51. The method of any one of claims 1 to 50, wherein the sample is a liquid sample.
52. The method of claim 51, wherein the liquid sample is a blood, serum, plasma, urine, vitreous, sputum, saliva, tears, perspiration, feces, bile, lymph, cervical mucus, or semen sample.
53. The method of claim 51 or 44, wherein the liquid sample is a blood, plasma, serum, or urine sample.
54. The method of any one of claims 1 to 53, wherein the sample is a blood sample or a derivative sample thereof.
55. The method of any one of claims 1 to 54, wherein the sample is a plasma sample.
56. The method of any one of claims 1 to 55, wherein the sample comprises DNA from a tumor.
57. The method of any one of claims 1 to 55, wherein the sample is a sample from a cancer tissue.
58. The method of any one of claims 1 to 55, wherein the sample comprises DNA from a transplanted organ.
59. The method of any one of claims 1 to 55, wherein the sample comprises DNA from a fetus.
60. The method of any one of claims 1 to 55, wherein the subject is suspected or at risk of having a disease.
61. The method of claim 60, wherein the disease is a cancer.
62. The method of claim 61, wherein the cancer is selected from ovarian cancer, soft tissue sarcoma, peripheral T cell cancer, colorectal cancer, intrahepatic cholangiocarcinoma, glioblastoma, esophageal cancer, cutaneous T cell lymphoma, non-Hodgkin lymphoma, urothelial cancer, basal cell carcinoma, epithelioid sarcoma, pancreatic cancer, non-small cell lungcarcinoma, Hodgkin lymphoma, renal cell carcinoma, mesothelioma, metastatic uveal melanoma, kidney cancer, blood cancer, HER2-expressing cancers, non-melanoma skin cancer, liposarcoma, hepatocellular carcinoma, small lymphocytic lymphoma, prostate cancer, breast cancer, anal cancer, marginal zone lymphoma, cutaneous squamous cell carcinoma, thyroid cancer, medullary thyroid cancer, triple-negative breast cancer, neuroendocrine prostate cancer, bladder cancer, paraganglioma, medulloblastoma, superficial basal cell carcinoma, head and neck squamous cell carcinoma, hematologic malignancies, melanoma, B-cell lymphoma, relapsed / refractory acute myeloid leukemia, angiosarcoma, bone sarcoma, refractory cervical cancer, cholangiocarcinoma, osteosarcoma, biliary tract cancer, castration-resistant prostate cancer, gastroesophageal adenocarcinomas, rhabdomyosarcoma, carcinoma, non-muscle invasive bladder cancer, uveal melanoma, small cell lung cancer, cervical cancer, primary open angle glaucoma, follicular lymphoma, synovial sarcoma, liver cancer, carcinosarcoma, leptomeningeal brain tumors, T-cell lymphoma, lymphoma, small cell lung cancer, mantle cell lymphoma, B-cell malignancies, endometrial cancer, myxoid / round cell liposarcoma, metastatic Merkel cell carcinoma, neuroblastoma, chronic lymphocytic leukemia, tenosynovial giant cell tumors, sarcoma, acute myeloid leukemia, skin cancer, nasopharyngeal carcinoma, relapsed / refractory Ewing sarcoma, bone cancer, glioma, salivary gland carcinoma, gastric cancer, benign tumor, low-grade serous ovarian cancer, metastatic breast cancer, multiple myeloma, diffuse large B cell lymphoma, relapsed / refractory lymphoma, metastatic colorectal cancer, advanced malignancies, and acute lymphoblastic leukemia.
63. The method of claim 61, wherein the cancer is selected from a cancer of the abdomen or abdominal wall, adrenal gland, anus, appendix, bladder, bone, brain, breast, cervix, chest wall, colon, diaphragm, duodenum, ear, endometrium, esophagus, fallopian tube, gallbladder, gastroesophageal junction, head and neck, kidney, larynx, liver, lung, lymph node, malignant effusions, mediastinum, nasal cavity, omentum, ovarian, pancreas, pancreatobiliary, parotid gland, pelvis, penis, pericardium, peritoneum, pleura, prostate, rectum, salivary gland, skin, small intestine, soft tissue, spleen, stomach, thyroid, tongue, trachea, ureter, uterus, vagina, vulva, and whippie resection.
64. The method of claim 61 , wherein the cancer is selected from lung cancer, breast cancer, bladder cancer, and colorectal cancer.
65. The method of claim 49, wherein the methylation status of the plurality of target loci is indicative of the presence or absence of the cancer.
66. the method of claim 50, wherein the presence of somatic mutations is indicative of the presence or absence of the cancer.
67. The method of any one of claims 1-66, wherein the subject is a pregnant female.
68. The method of any one of claims 1-66, wherein the subject is a subject comprising an organ from another individual.
69. The method of any of claims 1-68, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 70 and 500 base pairs in length.
70. The method of any of claims 1-68, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 100 and 200 base pairs in length.
71. The method of any of claims 1-68, wherein the method further comprises enriching the sample DNA molecules for DNA molecules that are between 130 and 170 base pairs in length.
72. The method of any one of claims 20-71, wherein before appending the adapters, the sample DNA molecules are fragmented to form fragmented DNA molecules.
73. The method of any of claims 1-72, wherein the plurality of target loci each comprises a set of two or more CpG sites that are differentially methylated in one or more cancers.
74. The method of any of claims 1-73, wherein the target loci each comprises three or more CpG sites that are differentially methylated in a cancer.
75. The method of any of claims 1-73, wherein one or more steps are performed in a microfluidic device or emulsion- split.