Primers for selectively amplifying GC-rich sequences

Primers designed to amplify GC-rich regions of genomic DNA fragments through deamination and PCR enhance cancer detection by resolving methylation at the single-nucleotide level, addressing inefficiencies in current methods.

WO2026003796A1PCT designated stage Publication Date: 2026-01-02DANA FARBER CANCER INSTITUTE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056549
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-24
Filing Date
2025-06-27
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Current methods for detecting aberrant methylation in genomic DNA, particularly in cancer diagnosis, are limited by their inability to resolve methylation at the single-nucleotide level and are inefficient due to the dominance of non-tumor DNA in liquid biopsies, masking hypermethylated gene regulatory regions.

Method used

The development of primers that selectively amplify GC-rich regions of genomic DNA fragments by hybridizing to adaptors and utilizing deamination to convert methylated or unmethylated cytosines to uracils, followed by PCR to enrich for these regions, allowing high-resolution methylation analysis.

Benefits of technology

Enables efficient and precise identification of aberrantly methylated or demethylated gene regulatory sequences, enhancing cancer detection and diagnosis by selectively amplifying GC-rich regions, even in low-tumor DNA samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025056549_02012026_PF_FP_ABST
    Figure IB2025056549_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Primers are provided that are capable of (i) specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, and (ii) selectively amplifying GC-rich sequences of a genome. The primers can be employed to interrogate the methylation status of multiple gene regulatory sequences in parallel, e.g., to identify aberrantly methylated or demethylated gene regulatory sequences on a global genomic scale. Methods that utilize, and compositions and kits that include, such primers are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PRIMERS FOR SELECTIVELY AMPLIFYING GC-RICH SEQUENCES

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] [1] This application claims benefit of U.S. Provisional Application No. 63 / 665,402 filed June 28, 2024 and U.S. Provisional Application No. 63 / 793,730 filed April 24, 2025, the entire contents of each is incorporated herein by reference.

[0004] FIELD

[0005] [2] Primers are provided that are capable of (i) specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, and (ii) selectively amplifying GC-rich sequences of the genome. The primers can be employed to interrogate the methylation status of multiple gene regulatory sequences in parallel, e.g., to identify aberrantly methylated or demethylated gene regulatory sequences on a global genomic scale.

[0006] SEQUENCE LISTING

[0007] [3] The present specification makes reference to a Sequence Listing, submitted electronically as an .xml file name “3433W01WO_Sequence Listing” on June 27, 2025. The .xml file was generated on June 17, 2025 and is 53,248 bytes in size. The entire contents of the Sequence Listing are herein incorporated by reference.

[0008] BACKGROUND

[0009] [4] Plasma-circulating, cell-free genomic DNA bearing aberrant methylation can be a strong prognostic factor for cancer development and can enable early cancer detection. Cancerous cells contain densely methylated promoters in tumor suppressor genes (aberrant hypermethylation), as well as significant de-methylation (global hypomethylation) in oncogenes, repeat sequences and single copy sequences. Identification of traces of aberrant hypermethylation or hypomethylation in plasma circulating DNA by employing, e.g., next generation sequencing (NGS) can provide a strong prognostic biomarker for predicting cancer development and directing cancer treatment accordingly. Methods that allow the rapid and efficient detection of aberrant methylation both in tissue biopsies and in circulating DNA obtained from plasma are in demand. [5] Current methods to enrich DNA for methylated sequences prior to sequencing include (a) antibody-based methods (e.g., MeDIP or others); and (b) methylation-sensitive enzyme-based methods. These approaches require protracted protocols and have limited flexibility as to the sites chosen for screening. Antibody-based methods extract information only from sites where the antibody can bind, while enzyme-based methods are restricted to sites containing the enzymatic restriction sites. Neither of these methods can resolve methylation at the single-nucleotide level, which is desirable to identify methylated DNA targets within a genome with high resolution that may, e.g., reveal cancer type and tissue of origin.

[0010] [6] Alternative approaches to detecting methylation by sequencing include de-amination- based methods, e.g., bisulfite or apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC) conversion (which converts unmethylated cytosines to uracils), or TET- Assisted Pyridine Borane Sequencing (TAPS, which converts methylated cytosines to uracil), thereby enabling sequencing to identify sites of methylation at the single base level. A common technical difficulty for all these methods is that DNA originating from tumors is only a small proportion of the overall genomic DNA in circulation. Hence, in liquid biopsies, hypermethylated gene regulatory regions (e.g., promoters and CpG islands) become ‘masked’ by presence of circulating DNA from normal cells that do not contain such aberrant methylation and the effectiveness of the hypermethylated gene regulatory regions as biomarkers of cancer is reduced. When tumors are small, as in early stages, the amount of tumor circulating DNA is also small compared to overall circulating DNA. Although methods like RRBS and Heat-rich can select GC-rich portions of the genome and enrich gene regulatory regions, this approach requires extensive sequencing and is thus inefficient.

[0011] [7] Beyond cancer, there are other medical conditions that can generate unusually high levels of tissue-specific methylation in bodily fluids (e.g., blood), and this may also be used as an early biomarker for such conditions. For example, early stages of stroke and heart disease can generate increased levels of CNS-specific or heart-specific methylated or unmethylated DNA in blood. Similarly, at early stages of bone marrow transplantation rejection, there are increased levels of DNA originating from the donor that circulate in the recipient’s blood. [8] Methods that address the shortcoming of existing methods and that are capable of identifying aberrantly methylated / de-methylated DNA in liquid and tissue biopsies, e.g., for the purpose of improved cancer diagnosis, would be welcome in the art.

[0012] SUMMARY

[0013] [9] The primers, methods of using them and compositions and kits comprising the same can be used to assess the methylation status of gene regulatory regions (e.g., promoters and / or CpG islands) within a genome. Methylated cytosine (C) is present only in CpG dinucleotides in the genome. C that is not in CpG context is not methylated by methyltransferases. The primers are designed to selectively amplify GC-rich regions of adaptor-flanked genomic DNA fragments from a mixture of adaptor-flanked genomic DNA fragments prepared from a sample (e.g., cell- free genomic DNA present in a blood biopsy). Identification of small proportions of the overall genomic DNA that may be aberrantly methylated, e.g., using sequencing, is thus possible.

[0014]

[0010] The primers include a nucleotide sequence A that is complementary to at least 10 consecutive nucleotides comprised in the adaptor. In addition, the primers comprise a probe nucleotide sequence P that is capable of binding to GC-rich regions in genomic DNA. The primers may further comprise a bridging nucleotide sequence B that is not complementary to a nucleotide sequence comprised in the adaptor. The presence of a bridging nucleotide sequence may be advantageous, e.g., when longer adaptor-flanked genomic DNA fragments are employed in the methods described herein, as it may provide the primer with more flexibility to allow nucleotide sequence A to specifically hybridize to the at least 10 consecutive nucleotides comprised in the adaptor and probe nucleotide sequence P to specifically hybridize to the at least one CG dinucleotide that may be comprised in the genomic DNA fragment. In some instances, the genomic DNA fragments may be PCR-amplified such that the primers described herein may be employed on amplicons of adaptor-flanked genomic DNA fragments.

[0015]

[0011] Accordingly, in one aspect, provided herein is a primer capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, the primer having the following structure: 5’-A-B-P-3’, wherein:

[0016] A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor; B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; and

[0017] P is a probe nucleotide sequence comprising n nucleic acids including at least one 5’-CG-3’ dinucleotide, wherein and n=4-18.

[0018]

[0012] In some aspects, methylated Cs in GC-rich regions of the genomic DNA fragments are deaminated to Us. Alternatively, unmethylated Cs in GC-rich sequences of the genomic DNA fragments are converted to Us. Accordingly, also provided herein is a primer capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor the primer having the following structure: 5’-A-B-P-3’, wherein:

[0019] A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;

[0020] B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; and

[0021] P is a probe nucleotide sequence comprising n nucleic acids including at least one 5’-CA-3’ dinucleotide, wherein and n=4-18.

[0022] After the de-amination step, e.g., using APOBEC conversion, the genomic DNA fragments (which may now comprise Us) may be subjected to a PCR amplification, which produces amplicons of the genomic DNA fragments (which may now comprise A or T in place of the methylated / unmethylated C originally present in the genomic DNA fragment). The amplicons of the genomic DNA fragments comprising 5’-TG-3’ or 5’-CA-3’ may be subsequently selectively amplified. Accordingly, provided herein is a primer capable of specifically hybridizing to an amplicon of a genomic DNA fragment comprising an adaptor, the primer having the following structure: 5’-A-B-P-3’, wherein:

[0023] A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;

[0024] B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; and P is nucleotide sequence comprising n nucleic acids including at least one 5’-TG-3’ dinucleotide or at least one 5’-CA-3’ dinucleotide, wherein and n=4-18.

[0025]

[0013] Also provided are methods that utilize the primer described herein to selectively amplify one or more adaptor-flanked genomic DNA fragment(s) of interest from a mixture of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils or wherein methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the first primer has the structure 5’-A-B-P-3’ as described above (where probe nucleotide sequence P depends on the adaptor-flanked genomic DNA fragments of interest and the deamination of unmethylated or methylated cytosine nucleotides); and c. performing the PCR to obtain PCR products.

[0026]

[0014] In step b, the one or more adaptor-flanked genomic DNA fragment(s) may comprise at least one 5’-CG-3’ and the first primer comprises a nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide. Or, in step b, the one or more adaptor-flanked genomic DNA fragment(s) may comprise at least one 5’-UG-3’ and the first primer comprises a nucleotide sequence P comprising at least one 5’-CA-3’ dinucleotide.

[0027]

[0015] Similarly, the method may be used to selectively amplify one or more amplicons of the one or more adaptor-flanked genomic DNA fragments from a mixture of amplicons of adaptor-flanked genomic DNA fragments.

[0028]

[0016] Accordingly, further provided herein is a method of selectively amplifying one or more adaptor-flanked genomic DNA fragment(s), or amplicons thereof, from a mixture of adaptor-flanked genomic DNA fragments, or amplicons thereof, prepared from a sample, the method comprising: a. obtaining adaptor-flanked genomic DNA fragments, or amplicons thereof, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils or wherein methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragment(s) or amplicons with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the first primer has structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’ dinucleotide, at least one 5’-TG-3’ dinucleotide or at least one 5’-CA-3’ dinucleotide; and c. performing the PCR to obtain PCR products.

[0029]

[0017] Provided herein is a method of selectively amplifying one or more amplicons of adaptor- flanked genomic DNA fragment(s) from a mixture of amplicons of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining one or more amplicons of adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils; b. contacting the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a PCR, wherein: the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-CG-3’ and the first primer has structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’, or the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise 5’-CA-3’ and / or 5’-TG-3’ and the first primer has structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-TG-3’ or 5’-CA-3’, respectively; and c. performing the PCR to obtain PCR products.

[0018] In another aspect, provided herein is a method of determining methylation status of genomic DNA of a subject, comprising: a. providing PCR products obtained from a method of selectively amplifying one or more adaptor-flanked genomic DNA fragment(s) as described above and b. sequencing the PCR products.

[0030]

[0019] Similarly, the PCR products may be obtained from a method of selectively amplifying one or more amplicons of the one or more adaptor-flanked genomic DNA fragments.

[0031]

[0020] Compositions and kits comprising the primer described herein are also provided.

[0032]

[0021] In some aspects, provided is a method for determining whether one or more GC-rich regions of a genome of a subject is methylated, wherein the method comprises: a. adding deamination-resistant first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject (e.g., wherein the deamination-resistant first and second adaptors comprise methylated cytosine nucleotides); b. deaminating unmethylated cytosines comprised in the adaptor-flanked genomic DNA fragments (e.g., by using TET2 oxidation followed by APOBEC) c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer with structure 5’-A-B-P-3’ as described herein and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein:

[0033] (i) nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA,

[0034] (ii) probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and

[0035] (iii) the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, and d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were methylated.

[0036]

[0022] In other aspects, provided is a method for determining whether one or more GC-rich regions of a genome of a subject is unmethylated, wherein the method comprises: a. adding first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject; b. deaminating methylated cytosines comprised in the adaptor-flanked genomic DNA fragments (e.g., by using TET1 and pyridine borane); c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer with structure 5’-A-B-P-3’ as described herein and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein:

[0037] (i) nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA,

[0038] (ii) probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and

[0039] (iii) the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; and f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were unmethylated.

[0040] BRIEF DESCRIPTION OF THE DRAWINGS

[0041]

[0023] Drawings are for illustration purposes only.

[0024] FIG. 1 is a schematic depiction of a forward primer comprising a first portion that is complementary to a first region of a template DNA molecule and a second portion that is complementary to a second region of the template DNA molecule. The first and second regions of the template DNA are typically non-consecutive, e.g., they may be separated by an intermediate region. The hybridization of the primer to the template DNA is also schematically depicted.

[0042]

[0025] FIG. 2A shows accumulation plots obtained from qPCRs using primers A-F (shown in Table 4C) to amplify a portion of KRAS, from a genomic DNA fragment. The first portion of each primer specifically hybridized to the same first target sequence in KRAS. The second portions of each primer specifically hybridized to different second target sequences in KRAS. The second target of primer A was closest to the first target, and the second target of primer F was furthest from the first target.

[0043]

[0026] FIG. 2B shows representative melt curves of the different PCR products obtained from each qPCR shown in FIG. 2A.

[0044]

[0027] FIG. 2C shows a representative graph showing the correlation between cycle threshold (Ct) value and number of intermediate nucleotides between the target of the first portion of the primer and the target of the second portion of the primer for the qPCRs shown in FIG. 2A.

[0045]

[0028] FIG. 3 is a schematic depiction of a forward primer capable of specifically hybridizing to two non-consecutive regions of an adaptor-flanked genomic DNA fragment. The primer comprises (i) a nucleotide sequence A that is complementary to consecutive nucleotides comprised in the adaptor, and (ii) a probe nucleotide sequence P that comprises a nucleotide sequence complementary to a portion of a genomic DNA sequence of interest.

[0046]

[0029] FIG. 4 is a schematic depiction of a primer capable of specifically hybridizing to two non- consecutive regions of an adaptor-flanked genomic DNA fragment. The primer comprises (i) a nucleotide sequence A that is complementary to consecutive nucleotides comprised in the adaptor, (ii) a bridging nucleotide sequence B that is not complementary to a nucleotide sequence comprised in the adaptor; and (iii) a probe nucleotide sequence P that comprises a nucleotide sequence complementary to a portion of a genomic DNA sequence of interest.

[0030] FIG. 5A shows representative qPCR amplification plots showing accumulation of the PCR products obtained from a methylated (M) or unmethylated (U) adaptor-RASSFl ultramer amplicon template, demonstrating that primers as disclosed herein comprising the probe sequence 5’-CGCCCG-3’ as set out in Table 5 can be used to enrich for the methylated form of the ultramer.

[0047]

[0031] FIG. 5B shows representative melt curves of the PCR products obtained from the qPCRs shown in FIG. 5 A.

[0048]

[0032] FIG. 6A shows representative qPCR amplification plots showing accumulation of the PCR products obtained from a methylated (M) or unmethylated (U) adaptor-RASSFl ultramer amplicon template, demonstrating that primers as disclosed herein comprising the probe sequence 5’-CCCGCG-3’ as set out in Table 5 can be used to enrich for the methylated form of the ultramer.

[0049]

[0033] FIG. 6B shows representative melt curves of the PCR products obtained from the qPCR shown in FIG. 6A.

[0050]

[0034] FIG. 7A shows representative qPCR amplification plots showing accumulation of the PCR products obtained from a methylated (M) or unmethylated (U) adaptor-flanked RASSF1 ultramer amplicon template, demonstrating that primers as disclosed herein comprising the probe sequence 5’-GACCCGCG-3’ as set out in Table 5 can be used to enrich for the methylated form of the ultramer.

[0051]

[0035] FIG. 7B shows representative melt curves of the PCR products obtained from the qPCRs shown in FIG. 7 A.

[0052]

[0036] FIG. 8A shows representative qPCR amplification plots showing accumulation of the PCR products obtained from a methylated (M) or unmethylated (U) adaptor-RASSFl ultramer amplicon template, demonstrating that primers as disclosed herein comprising the probe sequence 5’-CGACCCGCG-3’ as set out in Table 5 can be used to enrich for the methylated form of the ultramer.

[0053]

[0037] FIG. 8B shows representative melt curves of the PCR products obtained from the qPCRs shown in FIG. 8A.

[0038] FIG. 9A shows representative qPCR amplification plots showing the PCR products obtained from adaptor-flanked human genomic DNA fragments. The three samples, #1, #2, and #3, differed in the proportion of methylated DNA present in each, with #1 having the most and #3 having the least methylated DNA. The primer (primer N) comprised the probe sequence 5’-CGACCCGCG-3’, as set out in Table 5.

[0054]

[0039] FIG. 9B shows representative melt curves of the different PCR products obtained from the qPCRs shown in FIG. 9A.

[0055]

[0040] FIG. 10A shows a representative bar graph showing that the qPCR shown in FIG. 9A selectively amplified amplicons of adaptor-flanked genomic DNA fragments comprising 5 -CG-3’.

[0056]

[0041] FIG. 10B shows a representative bar graph showing the chromosomal position of individual promoters that were specifically amplified by primer N in a qPCR using sample #1 as template as shown in FIG. 9A.

[0057]

[0042] FIG. 10C shows a representative bar graph showing the chromosomal position of individual promoters that were specifically amplified by primer N in a qPCR of sample #2 as template as shown in FIG. 9A.

[0058]

[0043] FIG. 11A shows a representative bar graph showing the chromosomal position of individual promoters that were specifically amplified by primer H2 in a qPCR using sample #1 as template.

[0059]

[0044] FIG. 11B shows a representative bar graph showing the chromosomal position of individual promoters that were specifically amplified by primer H2 in a qPCR using sample #2 as template.

[0060]

[0045] FIG. 12 is a schematic showing the number of CpG islands sequenced with a coverage of 5 reads or greater for each template sample using primer O, as set out on Table 5, with probe nucleotide sequence P being 5’-GACCCGCG-3’.

[0061]

[0046] FIG. 13 is a schematic showing the number of CpG islands detected in the sequencing data of FIG. 12 that are known to be hypermethylated in colon cancer.

[0047] FIG. 14A shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using a standard emSEQ protocol. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0062]

[0048] FIG. 14B shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using Heatrich enrichment followed by a standard emSEQ protocol. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0063]

[0049] FIG. 14C shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the octamer 5’-GACCCGCG-3’, and a second (universal) primer. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0064]

[0050] FIG. 14D shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using Heatrich enrichment followed by EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the octamer 5’-GACCCGCG-3’, and a second (universal) primer. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0065]

[0051] FIG. 14E shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using EM-conversion and subsequent amplification using three first primers with structure 5’-A-P-3’, with probe nucleotide sequence P being the hexamer 5’-AACGCG-3’, the dodecamer 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), or the nonamer 5’-CGAACGCGA-3’, and a second (universal) primer. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0066]

[0052] FIG. 14F shows a representative pie chart showing the percentage of sequencing reads comprising CpG islands obtained using Heatrich enrichment followed by EM-conversion and subsequent amplification using three first primers with structure 5’-A-P-3’, with probe nucleotide sequence P being the hexamer 5’-AACGCG-3’, the dodecamer 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), or the nonamer 5’-CGAACGCGA-3’ and a second (universal) primer. Dark gray indicates the percentage of reads comprising a CpG island. Light gray indicates the percentage of reads not comprising a CpG island.

[0067]

[0053] FIG. 14G shows a representative bar chart showing the number of CpG islands covered and coverage depth for the experiments performed and summarized in FIG. 14A-F. Using a first primer with structure 5’-A-P-3’ resulted in the coverage of a significantly higher percentage of hypermethylated CpG islands than the use of a standard emSEQ protocol.

[0068]

[0054] FIG. 15A shows a representative bar chart showing the number of hypermethylated CpG islands out of the top 20 known hypermethylated CpG islands in human colon cancer gDNA that were captured using (left to right) (i) EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the decamer 5’-CGAACGCGAA-3’ (SEQ ID NO: 2) and a second (universal) primer, (ii) EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the decamer 5’-CGAACGCGAC-3’ (SEQ ID NO: 3) and a second (universal) primer, (iii) standard EMseq. Using a first primer with structure 5’-A-P-3’ resulted in the coverage of at least 70% of the known hypermethylated CpG islands compared to 20% with standard EMseq.

[0069]

[0055] FIG. 15B shows a representative bar chart showing the number of hypermethylated CpG islands out of the top 20 known hypermethylated CpG islands in human breast cancer gDNA that were captured using (left to right) (i) EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the decamer 5’-CGAACGCGAC-3’ (SEQ ID NO: 3) and a second (universal) primer, (ii) EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being the decamer 5’-CGAACGCGTA-3’ (SEQ ID NO: 4) and a second (universal) primer, (iii) standard EMseq. Using a first primer with structure 5’-A-P-3’ resulted in the coverage of at least 65% of the known hypermethylated CpG islands compared to 15% with standard EMseq.

[0056] FIG. 16 shows a representative bar chart showing CpG island (CGI) coverage of sequenced adaptor-flanked gDNA obtained from Human HCT116 DKO Methylated DNA (Zymo) using standard EMseq (denoted “Regular Illumina”), or EM-conversion and subsequent amplification using one first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being 5’-GACCCGCG-3’ (“Octamer probe 1”) or 5’-CGAACGCG-3’ (“Octamer probe 2”) and a second (universal) primer, or four first primers with structure 5’-A-P-3’, with probe nucleotide sequence P being 5’-CGMCGMCG-3’ (“Octamer probe 3”, where each M is independently selected from A or C), and a second (universal) primer.

[0070]

[0057] FIG. 17 shows a representative bar chart showing coverage of a set of 5 genes known to indicate CpG island methylator phenotype (CIMP) in colon cancer (i.e., the Weisenberger panel) (denoted “CIMP in colorectal cancer”) from adaptor-flanked gDNA obtained from human colon cancer patients using standard EMseq (denoted “Regular Illumina”), or EM-conversion and subsequent amplification using one first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being 5’-GACCCGCG-3’ (“Octamer probe 1”) or 5’-CGAACGCG-3’ (“Octamer probe 2”) and a second (universal) primer, or four first primers with structure 5’-A-P-3’, with probe nucleotide sequence P being 5’-CGMCGMCG-3’ (“Octamer probe 3”), where each M is independently selected from A or C) and a second (universal) primer.

[0071]

[0058] FIG. 18 shows a representative bar chart showing coverage of a set of 5 genes known to indicate a CpG island methylator phenotype (CIMP) in colon cancer (also referred to as the Weisenberger panel) from adaptor-flanked gDNA obtained from human colon cancer patients using standard EMseq (denoted “Regular Illumina”), or EM-conversion and subsequent amplification using a first primer with structure 5’-A-P-3’, with probe nucleotide sequence P being 5’-AACGCG-3’ (“Hexamer probe”), 5’-CGAACGCG-3’ (“Octamer probe”), 5’-CGAACGCGAA-3’ (“Decamer probe”) (SEQ ID NO: 2), or 5’-CGCGAACGCGAT-3’ (“Dodecamer probe”) (SEQ ID NO: 5), and a second (universal) primer.

[0072]

[0059] FIG. 19 shows a nucleotide sequence of a region on chromosome 17 comprising nucleic acids 77,373,475-77,373,62 (SEQ ID NO: 6). The positions of 5’-CG-3’ dinucleotides 1-13 (as numbered in Table 17) are shown in the nucleotide sequence, indicated by the number in the grey box shading each 5’-CG-3’ dinucleotide. Nucleotides represented in italic and underline are the target region that probe nucleotide sequence P of SEQ ID NO: 20 is capable of specifically hybridizing to.

[0073] DETAILED DESCRIPTION

[0074]

[0060] In order for the following description to be more readily understood, certain terms are first defined below. Additional definitions may be set forth throughout the specification.

[0075]

[0061] “Adaptor” and “adapter” are used interchangeably herein.

[0076]

[0062] Unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Thus, as used in this specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise.

[0077]

[0063] Unless specifically stated or obvious from context, as used herein, the term “or” is understood to be inclusive and covers both “or” and “and.” Furthermore, “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include “A and B”, “A or B”, “A” (alone), and “B” (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” is intended to include “A and / or B and / or C” and to thus encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0078]

[0064] It is understood that wherever aspects are described herein with the language “comprising,” otherwise analogous aspects described in terms of “consisting of’ and / or “consisting essentially of’ are also provided. In other words, if an aspect is described as comprising A, B and C, aspects consisting essentially of A, B and C are also contemplated, as are aspects consisting of A, B and C.

[0079]

[0065] The term “about” refers to an interval of accuracy that a person skilled in the art will understand to still ensure the technical effect of the feature in question. The term indicates a deviation from the indicated numerical value of ±10%, ±5%, or ±1% of the indicated numerical value.

[0080]

[0066] Unless otherwise stated, nucleotide sequences are presented in 5’ to 3’ direction.

[0067] An “adaptor” is an artificial DNA sequence that flanks genomic DNA fragments described herein. Various adaptor sequences that can be used in Next Generation Sequencing (NGS) methods and third-generation sequencing methods are described in the art, for example Illumina Inc.’s P5 adaptor: 5'-AATGATACGGCGACCACCGAGATCTACAC-3' (SEQ ID NO: 7) and P7 adaptor: 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 8).

[0081]

[0068] A “de-amination resistant” adaptor is an adaptor where its cytosines are resistant to conversion to uracil during the de-amination procedure selected. Thus, for de-amination that converts unmethylated cytosine to uracil (e.g., APOBEC-based de-amination), the cytosines of the adaptor are replaced with methyl-cytosines. For de-amination that converts methylated cytosine to uracil (e.g., TAPS-based de-amination), unmethylated cytosine is used.

[0082]

[0069] A “methylated cytosine” or “methyl-cytosine” typically refers to a cytosine in which a methyl group is present at the 5C position (5-methylcytosine). This term can further refer to oxidized derivatives of 5-methylcytosine, in particular, 5 -hydroxymethyl and 5- carboxy cytosine, and - in some instances - glycosylated 5-hydroxymethylcytosine. Glycosylated 5-hydroxymethylcytosine will not be converted by TAPS. The 4N position of a cytosine can also be methylated (N4-methylcytosine), and this modified base will also not be converted by TAPS. Accordingly, a de-amination resistant adaptor may comprise 5-methylcytosine, 5-hydroxymethylcytosine, glycosylated 5-hydroxymethylcytosine, N4-methylcytosine, or 5-carboxylcytosine in place of cytosine. The structures of each base are provided below:

[0083]

[0084] N4-methylcytosine 5-carboxy cytosine

[0085]

[0070] A phosphorothioate bond between two nucleotides is one in which a sulfur atom is replaced with a non-bridging oxygen in the phosphate backbone: phosphorothioate bond

[0086]

[0071] Unless otherwise defined herein, technical and scientific terms used herein have the same meaning as commonly used and / or understood by one of ordinary skill in the art to which this application belongs. In case of conflict, the present specification, including definitions, will control.

[0087]

[0072] Generally, nomenclature used in connection with, and techniques of, cell and tissue culture, molecular biology, virology, immunology, microbiology, genetics, analytical chemistry, synthetic organic chemistry, medicinal and pharmaceutical chemistry, and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art. Enzymatic reactions and purification techniques are performed according to manufacturer’s specifications, as commonly accomplished in the art or as described herein.

[0088] Sample

[0089]

[0073] The genomic DNA used with the primers and methods described herein is isolated from a sample obtained from a subject, e.g., a clinical sample from a subject suffering from a disease. Typically, isolation comprises one or more purification steps, e.g., to remove other nucleic acids and proteins found in the sample). The sample may be a liquid biopsy. For instance, the sample can be sputum, urine, stool, saliva, blood, or cerebrospinal fluid. Alternatively, the sample may be a tissue biopsy (e.g., a tumor biopsy).

[0090]

[0074] The genomic DNA may be fragmented to obtain the genomic DNA fragments (e.g., using sonication or enzymatic digestion). Or, the genomic DNA may already be present as genomic DNA fragments in the sample (e.g., circulating tumor DNA or cell-free DNA). Typically, a single strand of a genomic DNA fragment comprises less than 1,000 nucleotides, or 750 nucleotides or less. The genomic DNA fragment may comprise from about 50 nucleotides to about 750 nucleotides, or from about 100 to 500, or from 150 to 500, or from 200 to 500 nucleotides.

[0091]

[0075] Typically, the subject from which the sample is obtained has or is suspected of having a disease or disorder. The disease or disorder may be a neoplastic disease or disorder, an autoimmune disease or disorder, a metabolic disease or disorder, a cardiovascular disease or disorder, or a neurological disease or disorder. In particular, the neoplastic disease or disorder may be a cancer such as a neurological cancer, a gastrointestinal cancer (e.g., urological, colorectal, or colon cancer), a cardiovascular cancer, a hematological cancer (e.g., blood, lymphatic), a dermatological cancer, a pulmonary cancer, a musculoskeletal cancer, an endocrine cancer, a reproductive system cancer (e.g., cervical or ovarian cancer), or a breast cancer.

[0092]

[0076] The subject may be a mammal, e.g., a human, or a non-human mammal. Typically, the subject is human.

[0093]

[0077] For example, the subject is a human and accordingly the genomic DNA is human genomic DNA. The subject may have or be suspected of having a disease or disorder. The disease or disorder may be a neoplastic disease or disorder, e.g., a cancer.

[0094] Primer Structure

[0095]

[0078] The primers include a nucleotide sequence A that is complementary to at least 10 consecutive nucleotides comprised in an adaptor that is added to genomic DNA fragments in accordance with the disclosed methods. In addition, the primers comprise a probe nucleotide sequence P that is capable of specifically hybridizing to at least four consecutive nucleotides found in a genomic DNA sequence of interest (e.g., GC-rich regions of a genome). Nucleotide sequence A is capable of specifically hybridizing to any genomic DNA fragments comprising an adaptor with a complementary sequence. Stable recruitment of a DNA polymerase to the 3 ’ end of the primer described herein and subsequent DNA synthesis thereafter is dependent on the specific hybridization of probe nucleotide sequence P to a complementary sequence in the genomic DNA sequence of interest (e.g., a GC-rich region or an oncogene) of an adaptor-flanked genomic DNA fragment.

[0096]

[0079] Primers of this design allow a user to selectively amplify one or more adaptor-flanked genomic DNA fragments comprising a genomic DNA sequence of interest (e.g., a GC-rich region or an oncogene) from a sample comprising a mixture of adaptor-flanked genomic DNA fragments, thereby enriching the sample for nucleotide sequences deriving from genomic DNA fragments that comprise the sequence of interest (e.g., a GC-rich region or an oncogene). Such primers are capable of specifically hybridizing to non-consecutive first and second regions of an adaptor-flanked genomic DNA fragment. That is nucleotide sequence A specifically hybridizes to the at least 10 consecutive nucleotides comprised in the adaptor and probe nucleotide sequence P specifically hybridizes to the at least four consecutive nucleotides comprised in the genomic DNA sequence of interest. Under suitable conditions, such primers allow for polymerasedependent amplification of a genomic DNA sequence of interest (e.g., a GC-rich region or an oncogene).

[0097]

[0080] The primers may further comprise a bridging nucleotide sequence B that is not complementary to a nucleotide sequence comprised in the adaptor. Such an optional bridging nucleotide sequence B is considered to provide the disclosed primers with additional flexibility.

[0098]

[0081] Accordingly, provided herein is a primer that is capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, the primer having the following structure: 5’-A-B-P-3’, wherein:

[0099] A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;

[0100] B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; and

[0101] P is a probe nucleotide sequence that specifically hybridizes to a genomic DNA sequence of interest.

[0102]

[0082] In particular, P may be (i) a nucleotide sequence complementary to at least 4 consecutive nucleotides found in GC-rich regions in genomic DNA fragment, or a nucleotide sequence complementary to at least 4 consecutive nucleotides of a conversion product or amplicon of a GC-rich region of genomic DNA fragment.

[0103]

[0083] It will be understood that the primers provided herein can be designed to be used as a forward primer or a reverse primer during PCR amplification.

[0104]

[0084] It will also be understood that the primers provided herein may be utilized as a sequencing primer in a sequencing method that is sequencing-by-synthesis (SBS) based.

[0105] Nucleotide sequence A

[0106]

[0085] Nucleotide sequence A of the primers described herein is designed to be complementary and so capable of specifically hybridizing to at least 10 consecutive nucleotides comprised in an adaptor sequence. Nucleotide sequence A is designed to be complementary to the adaptor that is flanked to the 3 ’ terminus of the genomic DNA fragment, or amplicon thereof.

[0107]

[0086] Nucleotide sequence A may be capable of specifically hybridizing to 10-30 nucleotides consecutive nucleotides comprised in an adaptor sequence. Hence, sequence A comprises at least 10 nucleotides. Or, nucleotide sequence A may comprise 10-30 nucleotides, or 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides complementary to consecutive nucleotides comprised in an adaptor sequence. Nucleotide sequence A may alternatively comprise 15, 20, 25, 28 or 30 nucleotides complementary to consecutive nucleotides comprised in an adaptor sequence.

[0108]

[0087] Alternatively, nucleotide sequence A may consist of from 10-30 nucleotides, or 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides complementary to consecutive nucleotides comprised in an adaptor sequence. Nucleotide sequence A may alternatively consist of 15, 20, 25, 28 or 30 nucleotides complementary to consecutive nucleotides comprised in an adaptor sequence.

[0109]

[0088] The at least 10 consecutive nucleotides comprised in an adaptor sequence to which nucleotide sequence A hybridizes may be the 3 ’ terminal nucleotides of the adaptor flanking a genomic DNA fragment.

[0110]

[0089] Nucleotide sequence A can be adapted as appropriate to be complementary to the adaptor(s) that are used to flank the genomic DNA fragment. For example, the described methods may be used in conjunction with different sequencing platforms (Illumina, Pacific BioSciences, Oxford Nanopore). Each of these sequencing platforms uses different types of adaptor sequencing to flank genomic DNA fragments prior to sequencing.

[0111]

[0090] Next generation sequencing (NGS) platforms utilize adaptors to flank genomic DNA fragments. Typically, nucleotide sequence A of the primer described herein is designed to be complementary to an adaptor of an NGS platform. Various such adaptor sequences are described in the art.

[0112]

[0091] For example, the adaptor may be a Y-shaped adaptor, which comprises two nucleotide sequences, e g., 5’-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ ID NO: 9) and 5’-AGATCGGAAGAGCACACGTCTGAACTCCAGTCATTTAA-3’ (SEQ ID NO: 10).

[0092] Alternatively, the adaptor may comprise the nucleotide sequence 5'-AATGATACGGCGACCACCGAGATCTACAC-3' (SEQ ID NO: 7) or the nucleotide sequence 5'-CAAGCAGAAGACGGCATACGAGAT-3' (SEQ ID NO: 8).

[0113]

[0093] Nucleotide sequence A may be designed to be complementary to at least 10 of the 30 nucleotides that comprise the 5’ terminus of the 3’ adaptor.

[0114] Bridging nucleotide sequence B

[0115]

[0094] The bridging sequence is designed such that it is not complementary to the nucleotide sequence of the adaptor located 3’ to the binding site of nucleotide sequence A. Nucleotide sequence B typically consists of at least 5 nucleotides. It will be understood that nucleotide sequence B should be adapted to the adaptors that are used to flank the genomic DNA fragment.

[0116]

[0095] While nucleotide sequence B may be longer than 30 nucleotide sequences (e.g., up to 40 nucleotides), a length of 5-30 is adequate. For instance, nucleotide sequence B may consist of 5- 30 nucleotides, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some instances, nucleotide sequence B may consist of 5-15 nucleotides.

[0117]

[0096] Exemplary nucleotide sequences for use as nucleotide sequence B include 5’-AGTTAA- 3’ and 5’-AGTTAGAGTTGAA-3’ (SEQ ID NO: 11).

[0118] Probe nucleotide sequence P

[0119]

[0097] Probe nucleotide sequence P is designed to be complementary and so capable of specifically hybridizing to at least 4 consecutive nucleotides comprised in a genomic DNA sequence of interest. Hence, probe nucleotide sequence P comprises at least 4 nucleotides. Probe nucleotide sequence P may be capable of specifically hybridizing to 4-18 consecutive nucleotides comprised in a genomic DNA fragment. Or, probe nucleotide sequence P may be capable of specifically hybridizing to 19-30 consecutive nucleotides comprised in a genomic DNA fragment.

[0120]

[0098] The design of probe nucleotide sequence P depends on the specific application. As noted above, probe nucleotide sequence P can be designed to specifically hybridize to any genomic 1 DNA sequence of interest, including GC-rich regions and oncogenes (e.g., to detect a mutated version of an oncogene).

[0121]

[0099] Probe nucleotide sequence P can be designed in accordance with (i) the method used to differentiate methylated Cs from unmethylated Cs in the genomic DNA fragments and (ii) whether the user intends to amplify genomic DNA fragments that are methylated or unmethylated. For example, probe nucleotide sequence P may comprise at least 4 nucleotides, at least one or two of which are 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’. In other examples, probe nucleotide sequence P typically comprises (e.g., consists of) 4-18 nucleotides.

[0122]

[0100] Probe nucleotide sequence P may comprise or consist of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides. In particular, probe nucleotide sequence P may comprise or consist of 6, 7, 8, 9, 10, or 12 nucleotides.

[0123]

[0101] Positioning the at least one 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’ dinucleotide at the 3’ terminus of probe nucleotide sequence P may increase the stabilization of the hybridization of the primer to the adaptor-flanked genomic DNA fragment, or an amplicon thereof. Accordingly, probe nucleotide sequence ? may comprise 5’-NNCG-3’, 5’-NNCA-3’, or 5’-NNTG-3’, where each instance of N is independently selected from adenine, cytosine, guanine, or thymine. Alternatively, the at least one 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’ dinucleotide may not be positioned at the 3’ terminus of probe nucleotide sequence P. Or, the at least one 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’ dinucleotide may be positioned at the 3’ terminus of probe nucleotide sequence P and / or at the 5’ end terminus of probe nucleotide sequence P.

[0124]

[0102] Probe nucleotide sequence P may comprise two or more 5’-CG-3’ dinucleotides, two or more 5’-CA-3’ dinucleotides, or two or more 5’-TG-3’ dinucleotides, respectively.

[0125]

[0103] The two or more 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’ dinucleotides may be consecutive, e.g., 5’-CGCG-3’, 5’-CACA-3’, or 5’-TGTG-3’. Alternatively, the two or more 5’-CG-3’, 5’-CA-3’, or 5’-TG-3’ dinucleotides may be non-consecutive, e.g., 5’-CGNnCG-3’, 5’-CANnCA-3’, or 5’-TGNnTG-3’, where n is between 1 and 14 nucleotides and each instance of N is a nucleotide independently selected from adenine (A), cytosine (C), guanine (G), or thymine (T).

[0126]

[0104] Probe nucleotide sequence P may comprise at least one 5’-CG-3’ dinucleotide. For example, probe nucleotide sequence P may comprise or consist of 5’-CGCCCG-3’, 5 -CGCCCG-3’, 5 -CCCGCG-3’, 5 -CGACCCGCG-3’, 5 -GACCCGCG-3’, 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), 5’-CGAACGCGAA-3’ (SEQ ID NO: 2), 5’- AACGCG-3’, 5’-CGAACGCGAC-3’ (SEQ ID NO: 3), 5’-CGAACGCGTA-3’ (SEQ ID NO: 4), 5’-CGAACGCG-3’, 5’-CGMCGMCG-3’ where M = A or C, or 5’-CGCGAACGCGAT-3’ (SEQ ID NO: 5), 5’-CGCGAA-3’, 5’-ACGCGA-3’, 5’-CGAACG-3’, 5’-CGACGA-3’.

[0127]

[0105] Probe nucleotide sequence P may comprise at least one 5’-CG-3’ dinucleotide and at least 50% of the nucleic acids comprised in P are C. For example, probe nucleotide sequence P may comprise or consist of 5’-CGCCCG-3’, 5’-CGCCCG-3’, 5’-CCCGCG-3’, 5 -CGACCCGCG- 3’, or 5’-GACCCGCG-3’. Or, probe nucleotide sequence P may comprise at least one 5’-CG-3’ dinucleotide and less than 50% of the nucleic acids comprised in P are C. For example, probe nucleotide sequence P may comprise or consist of 5’-CGCGAA-3’, 5’-ACGCGA-3’, 5’- CGAACG-3’, 5’-CGACGA-3’, 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), 5’-CGAACGCGAA-3’ (SEQ ID NO: 2), 5’-AACGCG-3’, 5’-CGAACGCGAC-3’ (SEQ ID NO: 3), 5’-CGAACGCGTA-3’ (SEQ ID NO: 4), 5’-CGAACGCG-3’, or 5’-CGCGAACGCGAT-3’ (SEQ ID NO: 5).

[0128]

[0106] Probe nucleotide sequence P may comprise at least one 5’-CA-3’ dinucleotide. For example, probe nucleotide sequence P may comprise or consist of 5’-CACATT-3’, 5’-TTCACA-3’, 5’-CACACA-3’, 5’-CACACAT-3’, 5’-CACACAC-3’, or 5’-CACACAA-3’. Probe nucleotide sequence P may comprise at least one 5’-CA-3’ dinucleotide and at least 50% or at least 60% of the nucleic acids comprised in P are A. Or, probe nucleotide sequence P may comprise at least one 5’-CG-3’ dinucleotide and less than 50% of the nucleic acids comprised in P are A.

[0129]

[0107] Probe nucleotide sequence P may comprise at least one 5’-TG-3’ dinucleotide. For example, probe nucleotide sequence P may comprise or consist of 5’-TGCCTG-3’, 5’-CCTGTG-3’, 5’-TGACCTGTG-3’, or 5’-GACCTGTG-3’. Probe nucleotide sequence P may comprise at least one 5’-TG-3’ dinucleotide and at least 50% or at least 60% of the nucleic acids comprised in P are T. Or, probe nucleotide sequence P may comprise at least one 5’-CG-3’ dinucleotide and less than 50% of the nucleic acids comprised in P are T. Method of selective amplification

[0130]

[0108] The described primers can be used to enrich for one or more adaptor-flanked genomic DNA fragments, or one or more amplicons thereof, that comprise a sequence of interest, from a mixture of adaptor-flanked genomic DNA fragments, or amplicons thereof. Stated another way, the primers can be used to enrich at least a portion of the adaptor-flanked genomic DNA fragments, or amplicons thereof, in a sample.

[0131]

[0109] A suitable method to selectively amplify one or more adaptor-flanked genomic DNA fragment(s) wherein the genomic DNA fragment comprises a sequence of interest from a mixture of adaptor-flanked genomic DNA fragments the method comprises (a) contacting adaptor- flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the first primer has the structure 5’-A- B-P-3’ as described above, comprising a probe nucleotide sequence P that is complementary to 4-18 consecutive nucleotides of the sequence of interest in the genomic DNA; and (b) performing the PCR to obtain PCR products. The second primer is typically capable of specifically hybridizing to any of the adaptor-flanked genomic DNA fragments (and may be referred to as “universal primer”). The second primer may be designed to be complementary to at least 10 consecutive nucleotides of an adaptor of the adaptor-flanked genomic DNA fragments (i.e., capable of specifically hybridizing to at least 10 consecutive nucleotides of an adaptor of the adaptor-flanked genomic DNA fragments).

[0132] [HO] The method may be adapted to amplify adaptor-flanked genomic DNA fragments comprising different sequences of interest in parallel. Accordingly, the method may utilize two or more first primers in step (a), wherein the two or more first primers comprise different probe nucleotide sequence Ps that are each complementary to 4-30 consecutive nucleotides (e.g., 4-18 nucleotides) of each different sequence of interest.

[0133] Selective amplification regardless of cytosine methylation status

[0134]

[0111] Gene regulatory sequences such as promoters and CpG islands are enriched in 5’-CG-3’ dinucleotides.

[0135]

[0112] The genomic DNA fragments comprising GC-rich regions may be enriched using known methods, e.g., using Heatrich enrichment. Heatrich enrichment is typically performed before genomic DNA fragments are flanked with adaptors (i.e., before adaptor ligation, e.g., after end repair of the genomic DNA fragments).

[0136]

[0113] Alternatively, or in combination, adaptor-flanked genomic DNA fragments comprising GC-rich regions may be further enriched using a method of selective amplification using a primer as described herein. A suitable method to selectively amplify one or more adaptor-flanked genomic DNA fragment(s) comprising at least one 5’-CG-3’ dinucleotide from a mixture of adaptor-flanked genomic DNA fragments, regardless of the methylation status of the C, comprises (a) contacting adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the first primer has the structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’ dinucleotide; and (b) performing the PCR to obtain PCR products. The probe nucleotide sequence P may comprise two or more 5’-CG-3’ dinucleotides, optionally consecutively. This method does not differentiate an adaptor-flanked genomic DNA fragment comprising a 5’-CG-3’ comprising a methylated cytosine from an adaptor-flanked genomic DNA fragment comprising a 5’-CG-3’ comprising an unmethylated cytosine, because each is capable of being specifically hybridized to by probe nucleotide sequence P comprising at least one 5’-CG-3’.

[0137] Selective amplification of methylated or unmethylated GC-rich regions

[0138]

[0114] The C of a 5’-CG-3’ dinucleotide located in a GC-rich region, e.g., a gene promoter or CpG island, may be in a methylated or unmethylated state.

[0139]

[0115] The methylation status may be correlated with the expression (de-methylation) or inactivation (methylation) of the corresponding gene, respectively. Aberrant methylation of C nucleotides in 5’-CG-3’ genomic dinucleotides is known to be associated with numerous diseases or disorders including, but not limited to, cancer, autoimmune diseases, and psychiatric disorders.

[0140]

[0116] The methylation status of cytosines (Cs) can be delineated using genomic DNA preparation methods which either deaminate methylated Cs to uracil (U) (e.g., by using tet methylcytosine dioxygenase 1 (TET1) and pyridine borane) or deaminate unmethylated Cs to Us (e.g., by using chemical conversion, for example bisulphite conversion, or by using enzymatic conversion, for example TET2 oxidation followed by APOBEC deamination, e.g., EMseq).

[0141]

[0117] Deamination of unmethylated Cs to Us may be performed chemically or enzymatically. Deamination performed chemically may comprise sodium bisulphite treatment. Deamination performed enzymatically may make use of a bacterial DNA deaminase, e.g., as described in WO 2023 / 097226. For example, the bacterial DNA deaminase may be selected from CseDaOl, MGYPDaO6, and LbDaO2. The bacterial DNA deaminase may be selected from HcDaOl and StsDaOl. Alternatively, deamination can be performed enzymatically using apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC). Typically, performing deamination enzymatically with APOBEC is preceded by a TET2 oxidation step, which protects methylated cytosines from deamination by APOBEC.

[0142]

[0118] Deamination of methylated Cs to Us may be performed using tet methylcytosine dioxygenase 1 (TET1) and pyridine borane.

[0143]

[0119] Accordingly, a suitable method for selectively amplifying one or more adaptor flanked genomic DNA fragment(s) of interest from a mixture of adaptor flanked genomic DNA fragments comprises: a. obtaining adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils or wherein methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragments with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the first primer has the structure 5’-A-B-P-3’ as described above (where probe nucleotide sequence P depends on the adaptor-flanked genomic DNA fragments of interest and the deamination of unmethylated or methylated cytosine nucleotides); and c. performing the PCR to obtain PCR products.

[0144]

[0120] The adaptors may be added to the ends of the genomic DNA fragment before or after the genomic DNA is pre-treated.

[0121] The method may be adapted to amplify adaptor-flanked genomic DNA fragments comprising different sequences of interest in parallel. Accordingly, the method may utilize two or more first primers in step (b), wherein the two or more first primers comprise different probe nucleotide sequence Ps that are each complementary to 4-30 consecutive nucleotides (e.g., 4-18 nucleotides) of each different sequence of interest. As demonstrated herein, utilizing two or more first primers in step (b) can further increase the sequencing coverage of GC-rich regions (e.g., CpG islands), compared to using a single first primer in step (b).

[0145] Primer design considerations

[0146]

[0122] The probe nucleotide sequence P of the primer utilized for the selective amplification of methylated or unmethylated GC-rich regions depends on (i) the method used to differentiate methylated Cs from unmethylated Cs in the genomic DNA fragments and (ii) whether the user intends to amplify genomic DNA fragments that are methylated or unmethylated. Typically, the genomic DNA fragments are first pre-treated by (i) deaminating unmethylated cytosines or (ii) deaminating methylated cytosines.

[0147]

[0123] Table 1 summarizes the appropriate dinucleotide to include in probe nucleotide sequence P if the primer is designed to be complementary to an adaptor-flanked genomic DNA fragment. Scenarios 1-4 set out in this table describe how the sequence of the 5’-CG-3’ dinucleotide in the target sequence is altered as a result of the selected deamination method and what the appropriate design of the dinucleotide in probe nucleotide sequence P is in each scenario.

[0148] Table 1. Probe nucleotide sequence P design without amplification, C = unmethylated cytosine, mC = methylated cytosine.

[0149]

[0124] Table 2 summarizes the appropriate dinucleotide to include in probe nucleotide sequence P if the primer is designed to be complementary to an amplicon of an adaptor-flanked genomic DNA fragment. Scenarios 5-8 set out in this table describe how the sequence of the 5’-CG-3’ dinucleotide in the target sequence is altered as a result of the selected deamination method and subsequent amplification, which may employ a uracil-tolerant polymerase, and what the appropriate design of the dinucleotide in probe nucleotide sequence P is in each scenario. For example, following TAPS de-amination, methylated C is converted to dihydrouracil, which pairs with A. Upon PCR the dihydrouracil becomes T.

[0150] Table 2, Probe nucleotide sequence P design with amplification, C = unmethylated cytosine, mC = methylated cytosine.

[0151] ♦depending on the strand of the double stranded (ds) DNA amplicon that is selected for amplification

[0152] Scenario 1

[0153]

[0125] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA is first pre-treated to deaminate unmethylated cytosine nucleotides (e.g., via bisulfite or APOBEC) to uracil nucleotides and (ii) the aim of the PCR is to obtain products amplified from one or more adaptor-flanked genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was unmethylated.

[0154]

[0126] In this scenario, unmethylated cytosine nucleotides in genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the one or more genomic DNA fragments comprising an unmethylated C are converted to 5’-UG-3’. Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the adaptor-flanked genomic DNA fragments with the first primer with a probe nucleotide sequence P comprising at least one 5’-CA-3’.

[0127] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C is unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C is unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence is converted to 5’-UUUGUG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of the DNA sequence is converted to 3’-GGGUGU-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 3’-AAACAC-5’, can be used as described herein to selectively amplify the converted adaptor-flanked gDNA fragments comprising 5’-UUUGUG-3’. Likewise, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCACA-3’, can be used as described herein to selectively amplify the converted adaptor-flanked gDNA fragments comprising 3’-GGGUGU-5’.

[0155] Scenario 2

[0156]

[0128] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA is first pre-treated to deaminate methylated cytosine nucleotides to uracil nucleotides (e.g., via TAPS) and (ii) the aim of the PCR is to obtain products amplified from one or more adaptor-flanked genomic DNA fragments that originally comprised a 5 ’ -CG-3 ’ dinucleotide, wherein the C was unmethylated.

[0157]

[0129] In this scenario, methylated cytosine nucleotides in genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5 ’-CG-3’ dinucleotides of the one or more genomic DNA fragments comprising an unmethylated C remain as 5’ -CG-3 ’ . Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the adaptor- flanked genomic DNA fragments with the first primer, wherein probe nucleotide sequence P comprises at least one 5 ’-CG-3’.

[0158]

[0130] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C is unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C is unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence remains 5’-CCCGCG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of the DNA sequence remains 3’-GGGCGC-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 3’-GGGCGC-5’, can be used as described herein to selectively amplify the adaptor-flanked gDNA fragments comprising 5’-CCCGCG-3’. Likewise, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCGCG-3’, can be used as described herein to selectively amplify the adaptor- flanked gDNA fragments comprising 3’-GGGCGC-5’.

[0159] Scenario 3

[0160]

[0131] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA is first pre-treated to deaminate unmethylated cytosine nucleotides to uracil nucleotides and (ii) the aim of the PCR is to obtain products amplified from one or more adaptor-flanked genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was methylated.

[0161]

[0132] In this scenario, unmethylated cytosine nucleotides in genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the one or more genomic DNA fragments comprising a methylated C remain as 5’-CG-3’. Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the adaptor- flanked genomic DNA fragments with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’.

[0162]

[0133] The adaptors may be flanked to the genomic DNA fragment before the genomic DNA is pre-treated. Accordingly, the adaptors may be designed to be deamination-resistant, i.e., by incorporating methylated Cs into the adaptor instead of unmethylated Cs. Alternatively, the adaptors may be flanked to the genomic DNA after the genomic DNA is pre-treated. Accordingly, the adaptors may comprise unmethylated Cs.

[0163]

[0134] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C in each 5’-CG-3’ dinucleotide is methylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence is converted to 5’-UUCGCG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence remains 3’-GGGCGC-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 3’-AAGCGC-5’, can be used as described herein to selectively amplify the converted adaptor-flanked gDNA fragments comprising 5’-UUCGCG-3’. Likewise, a primer having structure 5’-A-B-P-3’ as described herein wherein B is optional and probe nucleotide sequence P comprises 5’-CCCGCG-3’ can be used as described herein to selectively amplify the adaptor-flanked gDNA fragments comprising 3 -GGGCGC-5’.

[0164] Scenario 4

[0165]

[0135] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA is first pre-treated to deaminate methylated cytosine nucleotides to uracil nucleotides and (ii) the aim of the PCR is to obtain products amplified from one or more adaptor-flanked genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was methylated.

[0166]

[0136] In this scenario, methylated cytosine nucleotides in genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the one or more genomic DNA fragments comprising a methylated C are converted to 5’-UG-3’. Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the adaptor-flanked genomic DNA fragments with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CA-3’.

[0167]

[0137] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C in each 5’-CG-3’ dinucleotide is methylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence is converted to 5’-CCUGUG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence is converted to 3’-GGGUGU-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 3’-GGACAC-5’, could be used as described herein to selectively amplify the converted adaptor-flanked gDNA fragments comprising 5’-CCUGUG-3’. Likewise, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCACA-3’, could be used as described herein to selectively amplify the converted adaptor-flanked gDNA fragments comprising 3’-GGGUGU-5’.

[0168] Scenario 5

[0169]

[0138] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA fragments are first pre-treated to deaminate unmethylated cytosine nucleotides to uracil nucleotides, (ii) the pre-treated genomic DNA fragments are amplified in a standard PCR prior to the selective amplification PCR, which results in the generation of amplicons of the genomic DNA fragments and (iii) the aim of the selective amplification PCR is to obtain products that are amplified from the amplicons of the genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was unmethylated, wherein the amplicons are obtained or obtainable from step (ii).

[0170]

[0139] In this scenario, unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the genomic DNA fragments comprising an unmethylated C are converted to 5’-UG-3’. Next, the genomic DNA fragments that were treated are amplified in a standard PCR, which results in the production of a double-stranded amplicon that comprises 5’-CA-3’ in one strand and 5’-TG-3’ in the other strand. Either strand of the amplicon can be used in the method of selective amplification. Accordingly, in this scenario, step (a) of the selective amplification method comprises (i) contacting the amplicons comprising 5’-CA-3’ dinucleotides with a first primer comprising structure 5’-A-B-P-3’, wherein probe nucleotide sequence P comprises at least one 5’-TG-3’, and / or (ii) contacting the amplicons comprising 5’-TG-3’ dinucleotides with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CA-3’.

[0140] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C is unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C is unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence is converted to 5’-UUUGUG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence is converted to 3’-GGGUGU-5’. Amplification of gDNA fragments obtained from a converted sense strand of DNA using PCR results in amplicons comprising 5’-TTTGTG-3’ and 3 ’ -AAACAC-5 ’ . Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-TTTGTG-3’ or 3’-AAACAC-5’, can be used as described herein to selectively amplify the amplicons of the adaptor-flanked converted gDNA fragments comprising 3 ’-AAACAC-5’ or 5’-TTTGTG-3’, respectively. Amplification of gDNA fragments obtained from a converted antisense strand of DNA using PCR results in amplicons comprising 5’-CCCACA-3’ and 3’-GGGTGT-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCACA-3’ or 3’-GGGTGT-5’, could be used as described herein to selectively amplify the amplicons of the adaptor-flanked converted gDNA fragments comprising 3’-GGGTGT-5’ or 5’-CCCACA-3’, respectively.

[0171] Scenario 6

[0172]

[0141] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA fragments are first pre-treated to deaminate methylated cytosine nucleotides to uracil nucleotides, (ii) the pre-treated genomic DNA fragments are amplified in a standard PCR prior to the selective amplification PCR, which results in the generation of amplicons of the genomic DNA fragments and (iii) the aim of the selective amplification PCR is to obtain products that are amplified from the amplicons of the genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was unmethylated, wherein the amplicons are obtained or obtainable from step (ii).

[0173]

[0142] In this scenario, unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the genomic DNA fragments comprising an unmethylated C remain as 5’-CG-3’. Next, the genomic DNA fragments that were treated are amplified in a standard PCR, which results in the production of a double-stranded amplicon that comprises 5’-CG-3’ in one strand and 5’-CG-3’ in the other strand. Either strand of the amplicon can be used in the method of selective amplification. Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the amplicons comprising 5’-CG-3’ dinucleotides with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’.

[0174]

[0143] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C is unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3’-GGGCGC-5’, wherein each C is unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence remains 5’-CCCGCG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence remains 3’-GGGCGC-5’. Amplification of gDNA fragments obtained from the sense strand of DNA using PCR results in amplicons comprising 5’-CCCGCG-3’ and 3’-GGGCGC-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCGCG-3’ or 3’-GGGCGC-5’, can be used as described herein to selectively amplify the amplicons of the adaptor-flanked gDNA fragments comprising 3’-GGGCGC-5’ or 5’-CCCGCG-3’, respectively. Amplification of gDNA fragments obtained from the antisense strand of DNA using PCR results in amplicons comprising 5’-CCCGCG-3’ and 3’-GGGCGC-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCGCG-3’ or 3’-GGGCGC-5’, can be used as described herein to selectively amplify the amplicons of the adaptor-flanked gDNA fragments comprising 3’-GGGCGC-5’ or 5’-CCCGCG-3’, respectively.

[0175] Scenario 7

[0176]

[0144] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA fragments are first pre-treated to deaminate unmethylated cytosine nucleotides to uracil nucleotides, (ii) the pre-treated genomic DNA fragments are amplified in a standard PCR prior to the selective amplification PCR, which results in the generation of amplicons of the genomic DNA fragments and (iii) the aim of the selective amplification PCR is to obtain products that are amplified from the amplicons of the genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was methylated, wherein the amplicons are obtained or obtainable from step (ii).

[0177]

[0145] In this scenario, unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the genomic DNA fragments comprising a methylated C remain as 5’-CG-3’. Next, the genomic DNA fragments that were treated are amplified in a standard PCR, which results in the production of a double-stranded amplicon that comprises 5’-CG-3’ in one strand and 5’-CG-3’ in the other strand. Either strand of the amplicon can be used in the method of selective amplification. Accordingly, in this scenario, step (a) of the selective amplification method comprises contacting the amplicons comprising 5’-CG-3’ dinucleotides with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’.

[0178]

[0146] The adaptors may be flanked to the genomic DNA fragment before the genomic DNA is pre-treated. Accordingly, the adaptors may be designed to be deamination-resistant, i.e., by incorporating methylated Cs into the adaptor instead of unmethylated Cs. Alternatively, the adaptors may be flanked to the genomic DNA after the genomic DNA is pre-treated. Accordingly, the adaptors may comprise unmethylated Cs.

[0179]

[0147] Deamination resistant adaptors are used when unmethylated cytosine are converted to uracil (e.g., APOBEC-based de-amination), as in Scenarios 1, 3, 5 and 7. In these scenarios, if the adaptors are flanked to the genomic DNA fragment before the genomic DNA is treated, then the adaptor is typically designed to be deamination-resistant. Deamination resistance can be achieved by replacing cytosine with a methylated cytosine, for example 5-methylcytosine, 5 -hydroxymethylcytosine, glycosylated 5-hydroxymethylcytosine, N4-methylcytosine, or 5- carboxylcytosine. It is not necessary to use deamination resistant adaptors when the deamination converts methylated cytosine to uracil (e.g., TAPS-based deamination), as in Scenarios 2, 4, 6 and 8.

[0180]

[0148] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA therefore correspondingly comprises the nucleotide sequence 3 ’-GGGCGC-5’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence is converted to 5’-UUCGCG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence remains 3’-GGGCGC-5’. Amplification of gDNA fragments obtained from a converted sense strand of DNA using PCR results in amplicons comprising 5’-TTCGCG-3’ and 3’-AAGCGC-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-TTCGCG-3’ or 3’-AAGCGC-5’, could be used as described herein to selectively amplify the amplicons of the adaptor-flanked converted gDNA fragments comprising 3’-AAGCGC-5’ or 5’-TTCGCG-3’, respectively. Amplification of gDNA fragments obtained from the antisense strand of DNA using PCR results in amplicons comprising 5’-CCCGCG-3’and 3 ’ -GGGCGC-5 ’ . Accordingly, a primer having structure 5’-A-B-P-3’ as described herein, wherein B is optional and probe nucleotide sequence P comprises 5’-CCCGCG-3’ or 3 ’-GGGCGC-5’, can be used in the selective amplification method described herein to selectively amplify the amplicons of the adaptor-flanked gDNA fragments comprising 3 ’-GGGCGC-5’ or 5’-CCCGCG-3’, respectively.

[0181] Scenario 8

[0182]

[0149] This scenario describes the appropriate dinucleotide to utilize in probe nucleotide sequence P if (i) the genomic DNA fragments are first pre-treated to deaminate methylated cytosine nucleotides to uracil nucleotides, (ii) the pre-treated genomic DNA fragments are amplified in a standard PCR prior to the selective amplification PCR, which results in the generation of amplicons of the genomic DNA fragments and (iii) the aim of the selective amplification PCR is to obtain products that are amplified from the amplicons of the genomic DNA fragments that originally comprised a 5’-CG-3’ dinucleotide, wherein the C was methylated, wherein the amplicons are obtained or obtainable from step (ii).

[0183]

[0150] In this scenario, methylated cytosine nucleotides in the genomic DNA fragments are deaminated before step (a) to uracil nucleotides, such that 5’-CG-3’ dinucleotides of the genomic DNA fragments comprising a methylated C are converted to 5’-UG-3’. Next, the genomic DNA fragments that were treated are amplified in a standard PCR, which results in the production of a double-stranded amplicon that comprises 5’-CA-3’ in one strand and 5’-TG-3’ in the other strand. Either strand of the amplicon can be used in the method of selective amplification. Accordingly, in this scenario, step (a) of the selective amplification method comprises (i) contacting the amplicons comprising 5’-CA-3’ dinucleotides with a first primer comprising structure 5’-A-B-P-3’, wherein probe nucleotide sequence P comprises at least one 5’-TG-3’, and / or (ii) contacting the amplicons comprising 5’-TG-3’ dinucleotides with the first primer, wherein probe nucleotide sequence P comprises at least one 5’-CA-3’.

[0184]

[0151] For example, a gDNA fragment obtained from a sense strand of DNA may comprise the nucleotide sequence 5’-CCCGCG-3’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated, and a gDNA fragment obtained from the corresponding antisense strand of DNA would therefore correspondingly comprise the nucleotide sequence 3’-GGGCGC-5’, wherein each C in each 5’-CG-3’ dinucleotide is methylated and remaining Cs are unmethylated. According to this scenario, the recited sequence of the gDNA fragment obtained from the sense strand of the DNA sequence would be converted to 5’-CCUGUG-3’ and the recited sequence of the gDNA fragment obtained from the antisense strand of DNA sequence would be converted to 3’-GGGUGU-5’. Amplification of gDNA fragments obtained from a converted sense strand of DNA using PCR would result in amplicons comprising 5’-CCTGTG-3’ and 3 ’ -GGACAC-5 ’ . Accordingly, a primer having structure 5’-A- B-P-3’ as described herein wherein B is optional and probe nucleotide sequence P comprises 5’-CCTGTG-3’ or 3 ’-GGACAC-5’ or could be used in the selective amplification method described herein to selectively amplify the adaptor-flanked amplicons of the converted gDNA fragments comprising 3 ’-GGACAC-5’ or 5’-CCTGTG-3’, respectively. Amplification of gDNA fragments obtained from a converted antisense strand of DNA using PCR would result in amplicons comprising 5’-CCCACA-3’ and 3’-GGGTGT-5’. Accordingly, a primer having structure 5’-A-B-P-3’ as described herein wherein B is optional and probe nucleotide sequence P comprises 5’-CCCACA-3’ or 3’-GGGTGT-5’ could be used in the selective amplification method described herein to selectively amplify the adaptor-flanked amplicons of the converted gDNA fragments comprising 3’-GGGTGT-5’ or 5’-CCCACA-3’, respectively.

[0152] From the foregoing scenarios and explanations, a person of ordinary skill in the art understands that the probe nucleotide sequence P of a primer having structure 5’-A-B-P-3’ as described herein is altered depending on the deamination method used and whether methylated or unmethylated CG dinucleotides are of interest. Although only some scenarios using specific primers are described in detail herein, it is understood that primers with accordingly altered probe nucleotide sequences P for use with alternative deamination methods are also encompassed by the present disclosure.

[0185] Exemplary primers

[0186] Exemplary primers are provided in Table 3.

[0187] Table 3, Exemplary primers

[0153] As shown in Table 3, the primer can have a structure 5’-A-P-3’. Exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least six nucleotides. For example, such primers may have the nucleotide sequence

[0188] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCCCG-3 ’ (SEQ ID NO: 13),

[0189] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCCCGCG-3 ’ (SEQ ID NO: 14),

[0190] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCGAA-3 ’ (SEQ ID NO: 15)

[0191] 5 ’ -CGTGTGCTCTTCCGATCTAATATTACGCGA-3 ’ (SEQ ID NO: 16),

[0192] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACG-3 ’ (SEQ ID NO: 17),

[0193] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGACGA-3 ’ (SEQ ID NO: 18), or

[0194] 5 ’ -CGTGTGCTCTTCCGATCTAATATTAACGCG-3 ’ (SEQ ID NO: 19). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least eight nucleotides. For example, such primers may have the nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTGACCCGCG-3 ’ (SEQ ID NO: 20), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCG-3 ’ (SEQ ID NO: 21), or 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGMCGMCG-3 ’ where each M is independently selected from A or C (SEQ ID NO: 22). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least nine nucleotides. For example, such primers may have the nucleotide sequence

[0195] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGACCCGCG-3 ’ (SEQ ID NO: 23). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least ten nucleotides. For example, such primers may have the nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCGAA-3 ’ (SEQ ID NO: 24), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCGAC-3 ’ (SEQ ID NO: 25), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCGTA-3 ’ (SEQ ID NO: 26). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least twelve nucleotides. For example, such primers may have the nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCGAACGCGTA-3 ’ (SEQ ID NO: 27) or 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCGAACGCGAT-3 ’ (SEQ ID NO: 28).

[0196]

[0154] As also shown in Table 3, the primer can also have a structure 5’-A-B-P-3’. Exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least six nucleotides. For example, such primers may for example have the nucleotide sequence 5’-CGTGTGCTCTTCCGATCTAATATTAGTTAGAGTTGAACGCCCG-3’ (SEQ ID NO: 29), 5’-CGTGTGCTCTTCCGATCTAATATTAGTTAGAGTTGAACCCGCG-3’ (SEQ ID NO: 30). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least eight nucleotides. For example, such primers may have the nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTAGTTAAGACCCGCG-3 ’ (SEQ ID NO: 31). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least nine nucleotides. For example, such primers may have the nucleotide sequence 5 ’-CGTGTGCTCTTCCGATCTAATATTAGTTAGAGTTGAACGACCCGCG-3 ’ (SEQ ID NO: 32). Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least ten nucleotides. Or, exemplary primers of this structure may comprise a probe nucleotide sequence P comprising at least twelve nucleotides.

[0197]

[0155] The universal primer may comprise or consist of the nucleotide sequence

[0198] 5’-ATACACTCTTTCCCTACACGACGC-3’ (SEQ ID NO: 43) or

[0199] 5’-CACTCTTTCCCTACACGA*C-3’, where A*C = an adenosine nucleotide and a cytosine nucleotide linked by a phosphorothioate bond (SEQ ID NO: 47).

[0200] Exemplary methods

[0201]

[0156] A method for selectively amplifying one or more adaptor flanked genomic DNA fragment(s) of interest from a mixture of adaptor flanked genomic DNA fragments may be utilized to selectively amplify adapter-flanked genomic DNA fragments comprising one or more CG dinucleotides (such as CpG islands), e.g., in a sample obtained from a subject having or suspected of having cancer. The method may comprise: a. providing adaptor-flanked genomic DNA fragments prepared from genomic DNA isolated from a sample obtained from a subject having or suspected of having cancer, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragments with at least one first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein the at least one first primer has the structure 5’-A-B-P-3’ as described above wherein probe nucleotide sequence P comprises at least one 5’-CG-3’ dinucleotide; and c. performing the PCR to obtain PCR products.

[0202]

[0157] Or, step a may comprise providing adaptor-flanked genomic DNA fragments prepared from genomic DNA isolated from a sample obtained from a subject having or suspected of having cancer, wherein methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils

[0203]

[0158] Alternatively, a C of one or more CG dinucleotides (such as CpG islands) in the genomic DNA fragment may have undergone conversion to U as a result of deamination in step a. Adaptor-flanked genomic DNA fragments comprising one or more UG dinucleotides may be of interest. Therefore, a method for selectively amplifying such adaptor-flanked genomic DNA fragments requires a different probe nucleotide sequence P. Accordingly, a method for selectively amplifying one or more adaptor flanked genomic DNA fragment(s) comprising one or more UG dinucleotides, e.g., in a sample obtained from a subject having or suspected of having cancer, comprises: a. obtaining adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragment(s) with at least one first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein: the one or more adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-UG-3’ and the at least one first primer has the structure 5’-A-B- P-3’ as described above wherein probe nucleotide sequence P comprises at least one 5’-CA-3’ dinucleotide; and c. performing the PCR to obtain PCR products.

[0204]

[0159] Amplicons of the adaptor-flanked genomic DNA fragments of interest that comprised one or more CG dinucleotides may be selectively amplified. Such amplicons comprise one or more CG dinucleotides as explained above. Accordingly, provided herein is a method of selectively amplifying one or more amplicons of adaptor-flanked genomic DNA fragment(s) from a mixture of amplicons of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining one or more amplicons of adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils; b. contacting the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a PCR, wherein: the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-CG-3’ and the first primer has structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-CG-3’; and c. performing the PCR to obtain PCR products

[0205]

[0160] Or, amplicons of the adaptor-flanked genomic DNA fragments of interest that comprised one or more UG dinucleotides as a result of conversion by deamination may be selectively amplified. Such amplicons comprise one or more TG dinucleotides or one or more CA dinucleotides as explained above. Therefore, a method for selectively amplifying such adaptor- flanked genomic DNA fragments requires a different probe nucleotide sequence P. Accordingly, provided herein is a method of selectively amplifying one or more amplicons of adaptor-flanked genomic DNA fragment(s) from a mixture of amplicons of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining one or more amplicons of adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils; b. contacting the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a PCR, wherein: the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise 5’-CA-3’ and / or 5’-TG-3’ and the first primer has structure 5’-A-B-P-3’ as described above, wherein probe nucleotide sequence P comprises at least one 5’-TG-3’ or 5’-CA-3’, respectively; and c. performing the PCR to obtain PCR products.

[0206]

[0161] In any of the above methods, the sample may be a biopsy from a tumor that is cancer or that is suspected of being cancer. Or the sample may be from a resected tumor. The cancer may be colon cancer or breast cancer.

[0207]

[0162] Probe nucleotide sequence P of the at least one first primer may comprise nucleotide sequence 5’-CGMCGMCG-3’ wherein M is independently selected from A or C. Alternatively, probe nucleotide sequence P of the at least one first primer may comprise nucleotide sequence 5 -GACCCGCG-3’, 5’-CGAACGCG-3’, 5’-AACGCG-3’, 5’-CGAACGCGAA-3’ (SEQ ID NO: 2), 5’-CGCGAACGCGAT-3’ (SEQ ID NO: 5), 5’-CGAACGCGAC-3’ (SEQ ID NO: 3), 5’-CGAACGCGTA-3’ (SEQ ID NO: 4), 5’-AACGCG-3’, 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), or 5’-CGMCGMCG-3’ wherein each M is independently selected from A or C. For example, first primers with an octamer or nonamer probe nucleotide sequence P comprising 5’- GACCCGCG-3’ and a decamer probe nucleotide sequence P comprising 5’-CGAACGCGAA- 3’ (SEQ ID NO: 2), respectively have been used herein to identify hypermethylated CpG islands associated with colon cancer, and first primers with decamer probe sequences P comprising 5’-CGAACGCGAC-3’ (SEQ ID NO: 3) and 5’-CGAACGCGAA-3’ (SEQ ID NO: 2), respectively, have been used herein to identify hypermethylated CpG islands associated with breast cancer.

[0208]

[0163] In some instances, step b. employs more than one first primer. For example, two, three, four, or five first primers with different probe sequences P may be used. Selective amplification of oncogenes

[0209]

[0164] It will be understood that probe nucleotide sequence P of a primer described herein could be adapted to be complementary to an oncogenic mutation of interest. Such a modified primer may be used to selectively amplify adaptor-flanked genomic DNA fragments comprising an oncogenic point mutation, e.g., BRAFV600E, from a mixture of adaptor-flanked genomic DNA fragments.

[0210] Sequencing

[0211]

[0165] PCR products obtained from a method of selective amplification described herein can be sequenced, e.g., to determine the methylation status of genomic DNA and / or to determine the presence of (a) mutation(s) in one or more oncogene(s).

[0212]

[0166] And so, there is also provided a method of determining methylation status of genomic DNA of a subject, comprising: a. obtaining PCR products obtained from a method of selective amplification described herein, and b. sequencing the PCR products.

[0213]

[0167] Typically, the sequencing of step (b) is performed with a next-generation sequencing (NGS) technology because the sequencing reads are short (e.g., less than 500 bases). Or, a third- generation single-molecule sequencing technology, e.g., polymerase-based single-molecule, real-time sequencing (e.g., Pacific Biosciences), or nanopore sequencing (e.g., Oxford Nanopore), may be used.

[0214]

[0168] Subsequently, the sequencing reads can be analyzed. For example, the method may further comprise aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived. The reference genome may be human.

[0215] Exemplary method for determining methylation status of genomic DNA

[0216]

[0169] The method may comprise selectively amplifying adaptor-flanked genomic DNA fragments comprising GC-rich regions comprising one or more 5’-CG-3’ dinucleotides wherein the C is methylated. In some instances, Heatrich enrichment may be performed before adding first and second adaptors to the ends of genomic DNA fragments.

[0217]

[0170] Accordingly, one exemplary method for determining whether one or more GC-rich regions of a genome of a subject is methylated comprises: a. adding deamination-resistant first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject (e.g., wherein the deamination-resistant first and second adaptors comprise methylated cytosine nucleotides); b. deaminating unmethylated cytosines comprised in the adaptor-flanked genomic DNA fragments (e.g., by using TET2 oxidation followed by APOBEC) c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer with structure 5’-A-B-P-3’ as described herein and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein:

[0218] (i) nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA,

[0219] (ii) probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and

[0220] (iii) the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, and d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were methylated, e.g., comprised one or more 5’-CG-3’ dinucleotides, wherein the C was methylated.

[0221]

[0171] In a further example, a method for determining whether one or more GC-rich regions of a genome of a subject is methylated comprises: a. adding first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject; b. deaminating methylated cytosines comprised in the adaptor-flanked genomic DNA fragments (e.g., by using TET1 and pyridine borane) c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer with structure 5’-A-B-P-3’ as described herein and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein:

[0222] (i) nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA,

[0223] (ii) probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CA-3’ dinucleotide present in one or more of the genomic DNA fragments, and

[0224] (iii) the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, and d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were methylated, e.g., comprised one or more 5’-CG-3’ dinucleotides, wherein the C was methylated.

[0225]

[0172] Or, a method for determining whether one or more GC-rich regions of a genome of a subject is unmethylated comprises: a. adding first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject; b. deaminating methylated cytosines comprised in the adaptor-flanked genomic DNA fragments (e.g., by using TET1 and pyridine borane); c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer with structure 5’-A-B-P-3’ as described herein wherein nucleotide sequence P comprises at least one 5’-CG-3’ dinucleotide and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein: i. nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA, ii. probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and iii. the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; and f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were unmethylated, e.g., comprised one or more 5’-CG-3’ dinucleotides, wherein the C was unmethylated.

[0226]

[0173] In each of the above exemplary methods, two or more first primers may be used in step b.

[0227]

[0174] In each of the above exemplary methods, step b may precede step a.

[0228] Compositions and Kits

[0229]

[0175] A composition that comprises one or more of the provided primer(s) - for instance, at least one primer having the structure 5’-A-B-P-3’ as described herein, wherein B is optional - is also provided, e.g., as a component of a kit. Accordingly, provided herein is a kit comprising a primer having the structure 5’-A-B-P-3’ as described herein, wherein B is optional. The primer may comprise a probe nucleotide sequence P described herein.

[0230]

[0176] The probe nucleotide sequence P of the primer may comprise or consist of 5’-CGAACGCGAA-3’ (SEQ ID NO: 2), 5’-CGAACGCGAC-3’ (SEQ ID NO: 3), 5’-AACGCG-3’, 5’-GACCCGCG-3’, 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), 5’-CGAACGCG-3’, 5’-CGMCGMCG-3’ wherein each M is independently selected from A or C, or 5’-CGAACGCGTA-3’ (SEQ ID NO: 4), or 5’-CGCGAACGCGAT-3’ (SEQ ID NO: 5). The nucleotide sequence A of the primer may comprise or consist of 5’- CGTGTGCTCTTCCGATCTAATATT-3’ (SEQ ID NO: 12).

[0231]

[0177] Accordingly, exemplary primers (the probe nucleotide sequence P is underlined) may have the nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCGAA-3 ’ (SEQ ID NO: 24), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCGAC-3 ’ (SEQ ID NO: 25), 5 ’ -CGTGTGCTCTTCCGATCTAATATTAACGCG-3 ’ (SEQ ID NO: 19),

[0232] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCGAACGCGTA-3 ’ (SEQ ID NO: 27), 5 ’ -CGTGTGCTCTTCCGATCTAATATTGACCCGCG-3 ’ (SEQ ID NO: 20), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGAACGCG-3 ’ (SEQ ID NO: 21), 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGMCGMCG-3 ’ . wherein each M is independently selected from A or C (SEQ ID NO: 22), or

[0233] 5 ’ -CGTGTGCTCTTCCGATCTAATATTCGCGAACGCGAT-3 ’ (SEQ ID NO: 28).

[0234]

[0178] In some instances, the kit comprises two, three or four primers with the structure 5’-A-B-P-3’, wherein B is optional, each primer having a different probe nucleotide sequence P.

[0235]

[0179] Such kits may be used for detecting CpG islands associated with a cancer (e.g., colon or breast cancer), for instance to determine whether the CpG islands are hypermethylated.

[0236]

[0180] The kit may further comprise PCR reagents such as a DNA polymerase, a suitable buffer, and / or free nucleotides as further components, wherein the free nucleotides comprise adenine, cytosine, guanine, and thymine.

[0237]

[0181] If the described primer is a forward primer, the kit may further comprise a reverse primer that is complementary to at least 10 consecutive nucleotides of the adaptor. If the described primer is a reverse primer, the kit may further comprise a forward primer that is complementary to at least 10 consecutive nucleotides of the adaptor.

[0238]

[0182] The kit may also comprise additional reagents, e.g., to perform end-repair of fragmented genomic DNA, and adaptor ligation. EXAMPLES

[0239]

[0183] Although methods and materials similar or equivalent to those described herein can be used, suitable methods and materials are described below. The following examples are included for illustrative purposes only and are not intended to be limiting.

[0240] Example 1. Primer complementary to two non-consecutive regions of a template DNA

[0241]

[0184] This example illustrates the design and use of a bipartite primer comprising a first portion that is complementary to a first region of a template DNA molecule and a second portion that is complementary to a second region of the template DNA molecule, wherein the first and second regions of the template DNA are non-consecutive.

[0242]

[0185] Hybridization of the first portion of the primer with the first region and hybridization of the second portion of the primer with the second region leaves an intermediate region of the template DNA that does not hybridize with the primer and therefore is not included in the resulting PCR product. This is shown schematically in FIG. 1. In some instances, the second region may be located in the genomic DNA fragment in such a manner that there is no intermediate region.

[0243]

[0186] Six forward primers were designed that comprised a first portion and a second portion. Each of the six primers included a first portion with sequence 5’-AAGGCCTGCTGAAAATGACTGAATATAA-3’ (SEQ ID NO: 33), which is complementary to a first region the KRAS gene. The second portion of each primer was designed to be complementary to a second region of the KRAS gene that was 5’ upstream and non- consecutive with the first region. Specifically, six different second portions, each consisting of a nonamer were designed. Each nonamer was complementary to nine nucleotides in the KRAS gene that were incrementally further away from the first region of the KRAS gene. The sequence of each forward primer tested and the number of nucleotides comprising the intermediate region (#IR) are shown in Table 4. The second portion is underlined.

[0244] Table 4, Primers used to amplify KRAS targets

[0245]

[0187] Six different quantitative PCRs (qPCRs) were performed using one of the forward primers in Table 4 and a reverse primer with sequence 5’-GTATCCTATCTGGACCTAAAAG-3’ (SEQ ID NO: 40). In these experiments, commercially available human male control DNA was used as the template DNA.

[0246]

[0188] Each primer was able to amplify its target, as shown in FIG. 2A. The specificity of the amplification of each of the intended targets is shown by the melt curves provided in FIG. 2B. The corresponding cycle threshold (Ct) values were obtained for each primer pair and plotted against the corresponding number of nucleotides in the intermediate region of the template DNA as shown in FIG. 2C. The figure shows that, the greater the number of nucleotides were in the intermediate region (i.e., the further the second region is from the first region), the greater the number of cycles was that was required to amplify the DNA.

[0247]

[0189] This example demonstrates that a bipartite primer design as described above can be used in a PCR to specifically amplify template DNA from a mixture of genomic DNA, even when the intended target sequence (i.e., the second region of the template DNA) that is complementary to the second portion of the primer is up to 226 nucleotides away from the first region of the template DNA (that is complementary to the first portion of the primer).

[0248] Example 2. Design of a primer that specifically hybridize to an adaptor and a sequence of interest

[0249]

[0190] The primer design described in Example 1 and depicted in FIG. 1 can be adapted to facilitate the detection of a sequence of interest within a genomic DNA fragment that is flanked to an adaptor at its 3’ terminus.

[0191] The “first portion” of the primer was changed so that it was complementary to at least 10 consecutive nucleotides comprised in the adaptor; this portion of the primer is referred to as “adaptor sequence A”. The first portion allows the primer to hybridize to the adaptor regardless of the genomic DNA fragment that it flanks. Polymerase extension occurs only if the “second portion” of the primer stably hybridizes with at least a region of the genomic DNA fragment.

[0250]

[0192] The “second portion” of the primer can be designed to be complementary to any given genomic DNA sequence of interest, such that use of this primer in a PCR results in specific amplification of one or more adaptor-flanked genomic DNA fragments that contained the genomic DNA sequence of interest. This portion of the primer is referred to as “probe nucleotide sequence P”.

[0251]

[0193] FIG. 3 depicts a primer that comprises (i) a nucleotide sequence A that is complementary to consecutive nucleotides comprised in the adaptor, and (ii) a probe nucleotide sequence P that comprises a nucleotide sequence complementary to a genomic DNA sequence of interest. Although an intermediate region is shown between the nucleotides hybridized by nucleotide sequence A and the nucleotides hybridized by probe nucleotide sequence P, in some instances, the sequence hybridized by probe nucleotide sequence P may be located in the genomic DNA fragment in such a manner that there is no intermediate region.

[0252]

[0194] A bridging nucleotide sequence B can be introduced between nucleotide sequence A and probe nucleotide sequence P. The bridging nucleotide sequence B is designed such that it is not complementary to a nucleotide sequence comprised in the adaptor.

[0253]

[0195] FIG. 4 depicts a primer that comprises (i) a nucleotide sequence A that is complementary to consecutive nucleotides comprised in the adaptor, (ii) a bridging nucleotide sequence B that is not complementary to a nucleotide sequence comprised in the adaptor; and (iii) a probe nucleotide sequence P that comprises a nucleotide sequence complementary to a genomic DNA sequence of interest. Although an intermediate region is shown between the nucleotides hybridized by nucleotide sequence A and the nucleotides hybridized by probe nucleotide sequence P, this does not have to be the case, and in some instances, the sequence hybridized by probe nucleotide sequence P may be located in the genomic DNA fragment in such a manner that there is no intermediate region. Example 3. Amplification of a single methylated promoter sequence

[0254]

[0196] This example illustrates that, by altering the design of probe nucleotide sequence P, the primer described in Example 2 can be used to selectively amplify a GC-rich region in an adaptor-flanked genomic DNA fragment.

[0255]

[0197] First, adaptor-flanked ultramers were prepared. The ultramers consisted of the anticipated sequence of the methylated RASSF1 promoter following deamination by emSEQ 5 ’ -TTTagtTtggatTT tggggg aggCGTtgaagtCGgggTTCGTTTtgtggTTTCGTTCGgTTCGCGTtt gTtagCGTTTaaagTTagCGaagTaCGggTTTaaTCGggTTatgtCGggggagTTtgagTtTattgagTtgGg gtTagaTTtaggaTTTTTtTTGCgaTttTaGCTTTggGCgggaTaTTgggGCggGCTggGCGCgaaTgat TGCgggtttTggtTGCt tTg tGCTTgggttgGCTTggtaTaGCTTTTtTggaTtTgagtaaTtTgaT-3’ (SEQ ID NO: 41) or the anticipated sequence of the unmethylated RASSF1 promoter following deamination by emSEQ 5’-TTTagtTtggatTTtgggggaggTGTtgaagtTGgggTTTGTTTtgtggTTTTGTTTGgTTTGTGTttg TtagTGTTTaaagTTagTGaagTaTGggTTTaaTTGggTTatgtTGggggagTTtgagTtTattgagTtgGggt TagaTTtaggaTTTTTtTTGTgaTttTaGTTTTggGTgggaTaTTgggGTggGTTggGTGTgaaTgatTG TgggtttTggtTGTt tTgtGTTTgggttgGTTTggtaTaGTTTTTtTggaTtTgagtaaTtTgaT-3’ (SEQ ID NO: 42). Deamination by emSEQ converts unmethylated cytosine to uracil, whereas methylated cytosine nucleotides are unchanged.

[0256]

[0198] Adaptors were ligated to each ultramer. The adaptors were 5’-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ ID NO: 9) and 5’-AGATCGGAAGAGCACACGTCTGAACTCCAGTCATTTAA-3’ (SEQ ID NO: 10). Pre-amplification polymerase chain reaction (PCR) was performed using primers which targeted the adaptors, using a commercially available uracil tolerant DNA polymerase. This causes the synthesized strand to incorporate adenine nucleotides where the template DNA provides uracil. Accordingly, a 5’-CG-3’ dinucleotide comprising an unmethylated cytosine that was present in the ultramer is converted to 5’-CA-3’ in the amplicon obtained from the pre-amplification PCR. By contrast, a methylated 5’-CG-3’ dinucleotide comprising a methylated cytosine that was present in the ultramer remains as 5’-CG-3’ in the amplicon obtained from the pre-amplification PCR.

[0199] In this experiment, forward primers were designed as described in Example 2. The sequence of each forward primer tested is provided in Table 5. Each primer comprised a nucleotide sequence A that specifically hybridized to 24 consecutive nucleotides comprised in the 3’ adaptor flanking the ultramer. The probe sequences P of the primers comprised 6- 9 nucleotides including at least one 5’-CG-3’ dinucleotide (underlined in the sequences shown in Table 2). Primers of this design can hybridize to amplicons comprising 5’-CG-3’ more stably than amplicons comprising 5’-CA-3’, thereby resulting in more efficient recruitment of DNA polymerase and selective amplification of the sequences of interest. Some of the primers also included a bridging nucleotide sequence B of 6-13 nucleic acids (shown in bold in Table 5). For each qPCR experiment, a reverse primer with sequence 5’-ATACACTCTTTCCCTACACGACGC-3’ (SEQ ID NO: 43) was used.

[0257] Table 5, Primers used to amplify methylated DNA

[0258] ♦abbreviation is used in FIGs 5-8

[0200] Each primer was able to amplify the methylated RASSF1 sequence (designated “M” in each figure) in considerably fewer cycles (as indicated by PCR threshold values) and to a greater extent than the unmethylated RASSF1 sequence (designated “U” in each figure), as shown in FIG. 5 A (primers Hl and Hl+B), FIG. 6A (primers H2 and H2+B), FIG. 7A (primers O and O+B), and FIG 8A (primers N and N+B). In each figure the negative controls “NC” 1 and 2 are marked. “NCI” refers to the negative control for primers Hl+B, H2+B, O+B, or N+B in FIG. 5, 6, 7, or 8, respectively. “NC2” refers to the negative control for primer Hl, H2, O, or N in FIG. 5, 6, 7, or 8, respectively. The delta-Ct indicated a 100-10,000 fold faster amplification of the methylated vs non-methylated form of the ultramers. The specificity of the amplification of each of the intended targets (i.e., methylated RASSF1) is shown by the melt curves provided in FIG. 5B, 6B, 7B and 8B respectively.

[0259]

[0201] This example demonstrates that primers with structure 5’-A-P-3’ or 5’-A-B-P-3’ and P comprising at least one CG dinucleotide can be used in PCR to selectively amplify an amplicon of an adaptor-flanked genomic DNA fragment wherein the genomic DNA fragment comprises at least one CG dinucleotide.

[0260] Example 4. Global amplification of methylated GC-rich genomic regions

[0261]

[0202] This example illustrates that a primer with the structure 5’-A-P-3’ with P comprising 9 nucleotides and at least one 5’-CG-3’ dinucleotide can be used to amplify adaptor-flanked genomic DNA fragments comprising at least one 5’-CG-3’ dinucleotide, such as genomic DNA fragments comprising GC-rich genomic regions. The example also illustrates that the methylation status of the enriched amplified regions can be determined by sequencing.

[0262]

[0203] Adaptor-flanked genomic DNA fragments were first prepared from two starting samples of commercially available human genomic DNA, methylated human genomic DNA obtained from Promega (#N1231) and human genomic DNA with low methylation of the order of 1-5% obtained from Zymo (#D5014-l). These were combined to produce a third sample with intermediate levels of methylation, as summarized in Table 6. The genomic DNA of each template sample was then fragmented. Table 6,

[0263]

[0204] Each of the template samples were then Heatrich enriched to enrich the genomic DNA sample for GC-rich regions including promoters and CpG islands and to remove regions that are low in CG content and do not convey methylation-sensitive information. While it is not necessary to perform the Heatrich enrichment, it was performed to reduce even further the required sequencing depth after amplification with the primer described herein.

[0264]

[0205] Deamination resistant adaptors were designed, based on Illumina’s TruSeq adaptor. Methylated Y-shaped adapter inspired by TruSeq Illumina adaptor design were prepared, to provide a methylated form of the adapter is resistant to cytosine deamination. The adaptor was prepared by hybridizing the following two oligonucleotides:

[0265] / 5Phos / ATAmCAmCTmCTTTmCmCmCTAmCAmCGAmCGmCTmCTTmCmCGATmCTA ATAT*T-3’ (SEQ ID NO: 44), where T*T=two thymidine nucleotides linked by a phosphorothioate bond and mC=methylated cytosine, in this instance, 5 -methylcytidine, and

[0266] / 5Phos / ATATTAGATmCGGAAGAGmCAmCAmCGTmCTGAAmCTmCmCAGTmCATTT A*A-3’ (SEQ ID NO: 45), where A*A=two adenosine nucleotides linked by a phosphorothioate bond and mC=methylated cytosine, in this instance, 5-methylcytidine.

[0267]

[0206] The deamination resistant adaptors were ligated to the genomic DNA fragments using NEBNext Ultra II Ligation Module (#E7645). The adaptor-flanked genomic DNA fragments were then purified with AMPure XP beads using 1.2X beads, followed by final elution in 28 pl of low IE buffer (low IE buffer was prepared using a 1 : 10 dilution of standard TE buffer using distilled water).

[0268]

[0207] The adaptor-flanked genomic DNA fragments were subjected to enzymatic conversion, to allow for methylated cytosines to be differentiated from unmethylated cytosines. Enzymatic conversion was performed by TET2 Oxidation, which protects methylated cytosines from APOBEC deamination, and subsequent APOBEC deamination of unmethylated cytosines. Specifically, the NEBNext Methyl-seq Conversion Module (NEB, Catalog no. E7125S) was used in accordance with the manufacturer’s protocol. Consequently, unmethylated cytosines were converted to uracils. The converted DNA was purified with 1 ,2X AMPure XP beads, eluted in 15 pl nuclease-free water. This process produced single-stranded adaptor-flanked genomic DNA fragments in which previously unmethylated cytosines are converted to uracil. For example, a genomic 5’-CG-3’ dinucleotide that comprised an unmethylated C was converted to 5’-UG-3’. Conversely, a genomic 5’-CG-3’ dinucleotide that comprised a methylated C remained unchanged.

[0269]

[0208] The single-stranded adaptor-flanked genomic DNA fragments were then subjected to a pre-amplification step using primers targeting the adapters and a uracil-tolerant polymerase. A 5’-UG-3’ in the genomic DNA of the converted single-stranded adaptor-flanked genomic DNA fragment corresponded to a 5’-CA-3’ in the newly-synthesized amplicon. Conversely, a methylated 5’-CG-3’ in the genomic DNA of the single-stranded adaptor-flanked genomic DNA fragment subjected to the conversion protocol remained 5’-CG-3’ in the newly-synthesized amplicon. Therefore, a primer targeting the adaptor-flanked genomic DNA fragment or the amplicon thereof required a probe nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide to specifically hybridize to an amplicon obtained from a genomic DNA fragment that comprised one or more 5’-CG-3’ dinucleotides that were methylated.

[0270]

[0209] Accordingly, to specifically amplify amplicons derived from genomic DNA fragments comprising at least one 5 ’ -CG-3 ’ dinucleotide wherein the C was methylated, primer N described in Table 5 was used as the forward primer for each experiment. The reverse primer described in Example 3 was used as the reverse primer in these experiments.

[0271]

[0210] The amount of PCR product resulting from each qPCR is shown in the amplification plot of FIG. 9A. The corresponding melt curves of the PCR products obtained from each qPCR are provided in FIG. 9B.

[0272]

[0211] The PCR products were then sequenced using nanopore sequencing. Specifically, the PCR products obtained from each qPCR were sequenced using a low-throughput Oxford Nanopore Flongle Flow Cell. This flow cell provides a maximum of 1,000,000 sequence reads, but most commonly about 500,000 sequence reads. The specifications are available on the company website.

[0273]

[0212] The enrichment of amplicons of adaptor-flanked genomic DNA fragments comprising 5’-CG-3’ dinucleotides (which were obtained from genomic DNA fragments that comprised one or more 5’-CG-3’ dinucleotides that were methylated) compared to the starting material is shown in FIG. 10A. In particular, 91.4% of PCR products obtained from sample #2 as the template comprised at least one 5’-CG-3’ dinucleotide, compared to 95.6% of PCR products obtained from sample #1 as the template. Correspondingly, 26.8% of PCR products obtained from sample #3 as the template comprised at least one 5’-CG-3’ dinucleotide.

[0274]

[0213] The chromosomal position of individual promoters specifically amplified by primer N from sample #1 and sample #2 are shown in FIG. 10B and FIG. 10C, respectively. 48 promoters were detected in the PCR products obtained from sample #1, whereas 16 promoters were detected in the PCR products obtained from sample #2. This was expected, given the reduced methylation of the input DNA used to create sample #2. 0 promoters were detected in the PCR products obtained from sample #3.

[0275]

[0214] This example demonstrates that a primer with structure 5’-A-P-3’ with P comprising 9 nucleotides including at least one 5’-CG-3’ dinucleotide can be used to selectively amplify amplicons of adaptor-flanked genomic DNA fragments comprising at least one 5’-CG-3’ dinucleotide (which were obtained from genomic DNA fragments that comprised one or more 5’-CG-3’ dinucleotides that were methylated), thereby enriching them from a mixture. The example further demonstrates that sequencing can be used to identify specific methylated regions of promoters.

[0276] Example 5. Global amplification of methylated GC-rich genomic regions

[0277]

[0215] This example demonstrates that a primer with structure 5’-A-P-3’ with P comprising 6 nucleotides and at least one 5’-CG-3’ dinucleotide can be used to amplify adaptor-flanked genomic DNA fragments comprising at least one 5’-CG-3’ dinucleotide.

[0216] Genomic DNA fragments were prepared as described in Example 4 and combined to provide a methylation level as shown in Table 7. Adaptor-flanked genomic DNA fragments were then prepared as described in Example 4.

[0278] Table 7,

[0279]

[0217] To specifically amplify amplicons derived from genomic DNA fragments comprising at least one 5’-CG-3’ dinucleotide wherein the C was methylated, the H2 primer described in Table 5 was used as the forward primer for each experiment. The reverse primer described in Example 3 was used as the reverse primer in these experiments.

[0280]

[0218] The PCR products obtained from each qPCR were then purified via Ampure purification kits, and ligated with a Nanopore ligation kit following the company’s specifications and then sequenced using a Nanopore Flongle Flow Cell as described in Example 4. As in Example 4, enrichment of adaptor-flanked genomic DNA fragments comprising 5’-CG-3’ dinucleotides was demonstrated. 45% of PCR products obtained from sample #4 as the template comprised at least one 5’-CG-3’ dinucleotide, compared to 84% of PCR products obtained from sample #1 as the template.

[0281]

[0219] The chromosomal position of individual promoters specifically amplified by primer N from sample #1 and sample #2 are shown in FIG. 11 A and FIG. 1 IB, respectively. 735 promoters were detected starting from sample #1, whereas 1535 promoters were detected starting from sample #4. This may be due to variation between Flow Cells.

[0282]

[0220] Compared to the experiment outlined in Example 4 that utilized a primer comprising a probe nucleotide sequence P comprising 9 nucleotides, an increased number of genomic sequences comprising at least one 5’-CG-3’ dinucleotides were amplified in this experiment which utilized a primer comprising a probe nucleotide sequence P comprising 6 nucleotides. This is a result of the shorter probe nucleotide sequence P sequence which allows the primer to specifically hybridize to a greater number of targets. This example demonstrates that more targets can be amplified by using a shorter probe nucleotide sequence P. Example 6. Global amplification of methylated GC-rich genomic regions

[0283]

[0221] This example demonstrates that a primer with structure 5’-A-P-3’ with P comprising 8 nucleotides and at least one 5 ’-CG-3’ dinucleotide can be used to amplify adaptor-flanked genomic DNA fragments comprising at least one 5 ’-CG-3’ dinucleotide.

[0284]

[0222] Genomic DNA fragments were prepared as described in Examples 4 and 5 and summarized in Tables 6 and 7. Adaptor-flanked genomic DNA fragments were prepared as described in Example 4.

[0285]

[0223] To specifically amplify amplicons derived from genomic DNA fragments comprising at least one 5 ’ -CG-3 ’ dinucleotide wherein the C was methylated, the O primer described in Table 5 was used as the forward primer. A forward primer with sequence 5’-ATTTACTGACCTCAAGTCTG-3’ (SEQ ID NO: 46), which was designed to be complementary to a portion of the adaptor only, was used as a control (i.e., control forward primer). The reverse primer described in Example 3 was used as the reverse primer in these experiments.

[0286]

[0224] PCR was performed using the primers as described, using samples #1, #2, #3, or #4 as template, to amplify adaptor-flanked genomic DNA fragments comprising at least one 5 ’-CG-3’ dinucleotide.

[0287]

[0225] The PCR products were then purified via Ampure kit, quantified via Qubit DNA measurement and submitted for EZ-amplicon sequencing on a Miseq (Azenta Inc.).

[0288]

[0226] PCR products were then sequenced using the Illumina sequencing platform, using a Miseq instrument. Samples were sequenced at 500,000 reads each. Coverage of sequencing reads was increased in comparison to the results obtained in Examples 4 and 5 with Nanopore.

[0289]

[0227] The number of CpG islands detected with high coverage per reaction using primer O by sequencing are shown in FIG. 12. With a sequencing coverage >5, more than 10,000 CpG islands were detected when sample #1 (100% methylation) was used as the template, including CpG islands. The lower the methylation levels (compare the results of sample #1 with those of samples #2, #3 and #4), the lower were the CpG islands with high coverage.

[0228] The sequencing reads were analyzed for known colon cancer methylation marks. Colon cancer methylation marks were identified by a database search, which revealed hypermethylation of 84 genes.

[0290]

[0229] The aligned reads obtained from each selective-amplification PCR were analyzed that had a coverage of 5 or greater were analyzed. 35 of the 84 genes were identified in the reads obtained from sample #1, as summarized in FIG. 13.

[0291]

[0230] The aligned reads obtained from each selective amplification PCR were then analyzed without the coverage restriction. Specifically, presence of a set of 5 genes known to indicate CpG island methylator phenotype (CIMP) in colon cancer (i.e., the Weisenberger panel) were analyzed; the analysis is summarized in Table 8. The complete Weisenberger panel was detectable when samples #1 and #2 were used as template. Surprisingly, even in the samples with the lowest methylation (samples #3 and #4), two of five of the genes were detected.

[0292] Table 8,

[0293]

[0231] Importantly, when amplification was performed with the control forward primer, CpG island detection was about 100 times less than when using primer O. This demonstrates the ability of a primer with a suitable probe nucleotide sequence P to specifically enrich for the intended target, i.e., adaptor-flanked genomic DNA fragments comprising at least one 5’-CG-3’ dinucleotide, thereby allowing greater sequencing coverage.

[0232] The number of methylated human tumor suppressor genes (TSGs) detected from each template sample, compared to the sequencing coverage, are summarized in Table 9. For coverage of >5 reads, the control forward primer amplification resulted in zero methylated targets. In contrast, 43 TSGs with coverage >5 were captured from template sample #4, indicating the strong enrichment of methylated targets achieved with primer O.

[0294] Table 9,

[0295]

[0233] This example demonstrates that a primer with structure 5’-A-P-3’ with P comprising at least one 5’-CG-3’ dinucleotide can be utilized to selectively amplify GC-rich regions known to be hypermethylated in a cancer cell genome.

[0296] Example 7. Selective amplification of methylated GC-rich genomic regions using ultra low- depth sequencing

[0297]

[0234] This example demonstrates primers having structure 5’-A-P-3’ with a probe nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide can be used to selectively amplify adaptor-flanked genomic DNA fragments comprising methylated GC-rich regions (e.g., CpG islands).

[0298]

[0235] Adaptor-flanked genomic DNA fragments were prepared from human genomic DNA (gDNA) with low methylation of the order of 1 -5% obtained from Zymo (#D5014- 1 ). The gDNA was first fragmented. End preparation was then performed on the gDNA fragments using NEBNext Ultra II DNA End Prep Kit for Illumina (Catalog No. E7645) according to the manufacturer’s protocol, with an input of 5-20 ng of human gDNA. The reaction was incubated at 20 °C for 30 minutes, followed by 65 °C for 30 minutes to deactivate the end preparation enzymes. Optionally, Heatrich enrichment was then performed by applying 88 °C for 5 minutes (Heatrich enrichment was used for FIGs. 14B, D, and F). Next, adaptor ligation was performed in the same reaction mix using the NEBNext Ultra II Ligation Module (Catalog No. E7645). This included 2.5 pl of 1.5 pM customized deamination-resistant Y-shaped adapter as provided in Example 4, in which all cytosines were methylated, along with 1 pl of ligation enhancer and 30 pl of ligation master mix. The reaction was incubated at 20 °C for 30 minutes. Following adapter ligation, purification was carried out using AMPure XP beads at a 1.2X bead-to-sample ratio, with final elution in 28 pl of low TE buffer as provided in Example 4.

[0299]

[0236] Next, unmethylated cytosines in the adapter-flanked gDNA fragments were deaminated via enzymatic conversion using the NEBNext Methyl-seq Conversion Module (Catalog no. E7125S, New England Biolabs) according to the manufacturer’s protocol. The protocol includes two main steps: TET2 oxidation and APOBEC deamination. The deaminated adaptor-flanked genomic DNA fragments were then cleaned up using 1.2X AMPure XP beads and eluted in 15 pl of nuclease-free water. Pre-amplification was performed using the NEBNext Q5U Master Mix (Catalog No. M0597) with primers targeting the adapters for 8 cycles, following the manufacturer’s protocol.

[0300]

[0237] For the samples used for FIG. 14A and FIG. 14B, sequencing was then performed using the MiSEQ platform with 2x250 bp paired-end reads and 100,000 reads per sample. This represents an “ultra low-depth sequencing”.

[0301]

[0238] For the experiments summarized in FIG. 14C-F, a further methylated-specific amplification step (i.e., a selective amplification step) using PCR was carried out prior to sequencing using at least one first primer having structure 5’-A-P-3’ (see Table 10 below) and a second (universal) primer having sequence 5’-CACTCTTTCCCTACACGA*C-3’, where A*C = an adenosine nucleotide and a cytosine nucleotide linked by a phosphorothioate bond (SEQ ID NO: 47). The first primer was capable of specifically hybridizing to adapter-flanked gDNA fragments comprising a sequence complementary to probe nucleotide sequence P, allowing DNA polymerase to initiate amplification. The first and second primers act as forward and / or reverse primers depending on whether the probe hybridizes to the top strand or bottom strand, or both.

[0239] For the experiments summarized in FIGs. 14C and 14D, a single first primer with structure 5’-A-P-3’ having an octamer probe nucleotide sequence P as shown in Table 10 (i.e., SEQ ID NO: 20) was used. For the experiment summarized in FIG. 14D, an additional Heatrich enrichment was performed after end preparation of gDNA fragments, as described above.

[0302]

[0240] For the experiments summarized in FIGs. 14E and 14F, three first primers with structure 5’-A-P-3’ and having a hexamer, decamer, dodecamer probe nucleotide sequences P, respectively, as shown in Table 10 (i.e., SEQ ID NOs: 19, 24, and 27, respectively) were used. For the experiment summarized in FIG. 14F, an additional Heatrich enrichment was performed after end preparation of gDNA fragments, as described above.

[0303] Table 10. Primers with structure 5’-A-P-3’ used to amplify methylated DNA

[0304]

[0241] After sequencing, the sequencing reads were aligned to the human reference genome to identify the genomic regions from which gDNA fragments were derived. To determine whether a read contained a CpG island (CGI), the obtained binary alignment map (BAM) file was intersected with a predefined set of CGIs, thereby allowing the number of sequencing reads that overlapped with CGIs to be counted.

[0305]

[0242] The following programs were used to analyze the sequencing reads: Bismark (Krueger F & Andrews SR, 2011. Bioinformatics. Jun 1 ;27(11 ): 1571 -2. doi: 10.1093 / bioinformatics / btrl67), Bedtools (Quinlan AR & Hall IM, 2010. Bioinformatics. 26, 6, pp. 841-842), Samtools (Danecek P, 2021. GigaScience, Volume 10, Issue 2, giab008, https: / / doi.org / 10.1093 / gigascience / giab008), Seqtk (https: / / github.com / lh3 / seqtk), BWA-MEM (Li H, 2013. arXiv.1303.3997v2 [q-bio.GN]), Picard Tools (“Picard Toolkit.” 2019. Broad Institute, GitHub Repository, https: / / broadinstitute.github.io / picard / ; Broad Institute), and MethylDackel (https: / / github.com / dpryan79 / MethylDackel).

[0243] The percentage of sequencing reads comprising a CGI out of the total number of obtained sequencing reads per experiment are shown in FIG. 14A-F. Use of a single first primer having the structure 5’-A-P-3’ and for selective amplification of GC-rich regions resulted in 33.5% coverage of known CGIs (FIG. 14C) compared to 2.4% coverage using standard EMseq (FIG. 14A) or 23% coverage using standard EMseq with Heatrich enrichment (FIG. 14B). Coverage was further improved to 40.8% (FIG. 14E) when three different first primers having the structure 5’-A-P-3’ were used to selectively amplify adaptor-flanked gDNA fragments. Even further improvements in coverage were achieved when Heatrich enrichment was used in addition to selective amplification. Using a single first primer having structure 5’-A-P-3’ and Heatrich enrichment resulted in a coverage of 69.6% of known CGIs (FIG. 14D) and using three different first primers of that structure resulted in coverage of 80.3% of known CGIs (FIG. 14F).

[0306]

[0244] Each of the experiments summarized in FIG. 14C-F resulted in marked improvements compared to prior art methods represented by the experiments summarized in FIG. 14A and FIG. 14B. To illustrate this point and facilitate a side-by-side comparison, the data obtained from these experiments is also summarized in FIG. 14G. Higher coverage of CGIs (i.e., at least 5x coverage per CGI) was observed in each instance when a first primer having structure 5’-A-P-3’ was used to selectively amplify CGIs.

[0307]

[0245] This example demonstrates that a primer having structure 5’-A-P-3’ with a probe nucleotides sequence P comprising at least one 5’-CG-3’ dinucleotide can be used to selectively amplify adapter-flanked genomic DNA fragments comprising methylated GC-rich regions of the genome such as CpG islands, resulting in improved coverage relative to prior art methods when using ultra low-depth sequencing.

[0308] Example 8. Selective amplification of hypermethylated GC-rich regions associated with human cancers

[0309]

[0246] This example demonstrates that a primer having structure 5’-A-P-3’ with a probe nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide can be used to selectively amplify adaptor-flanked gDNA fragments comprising GC-rich regions such as CpG islands associated with colon cancer or breast cancer to determine their methylation status using ultra low-depth sequencing at improved read coverage.

[0247] To confirm the findings in Example 7 in a diagnostically relevant setting, human genomic DNA obtained from Human HCT116 DKO Methylated DNA (Zymo #D5014-2) was prepared and analyzed by ultra low-depth sequencing as described in Example 7. Specifically, prior to sequencing, adaptor-flanked human gDNA fragments were processed using a standard EMseq protocol, either with or without performing a selective amplification step using PCR using one of the first primers listed in Table 11 and the second (universal) primer provided in Example 7 (SEQ ID NO: 47).

[0310] Table 11. Primers with structure 5’-A-P-3’ used to amplify methylated human gDNA

[0311]

[0248] After sequencing, the sequencing reads were aligned to the human reference genome to identify the genomic regions from which gDNA fragments were derived. To determine whether a read contained a CGI, the obtained binary alignment map (BAM) file was intersected with a predefined set of the top 20 hypermethylated CGIs in human colon cancer, thereby allowing the number of sequencing reads that overlapped with the CGIs to be counted. The top 20 hypermethylated CGIs were identified from the scientific literature, specifically Ashktorab H & Brim H, 2014. Curr Colorectal Cancer Rep. Dec l;10(4):425-430. doi: 10.1007 / sl l888-014- 0245-2; Lofton-Day C et al., 2008. Clin Chem. Feb; 54(2): 414-23. doi: 10.1373 / clinchem.2007.095992; Oh T et al. 2013. J Mol Diagn. Jul;15(4):498-507. doi: 10.1016 / j.jmoldx.2013.03.004; Chen W et al. 2013. Mol Biol Rep. May;40(5): 3457-64. doi: 10.1007 / sl 1033-012-2338-9; and Matthaios D et al., 2016. Oncol Lett. Jul;12(l):748-756. doi: 10.3892 / ol.2016.4649. The final list included SEPTIN9_CpG 151; MLHI CpG 93; CDKN2A_CpG 156; APC_CpG 50; MGMT CpG 98; SFRPI CpG 148; VIM CpG 195; RASSFI CpG 130; GATA5_CpG 257; GATA4_CpG 80; ICAM5_CpG 399; CDHI CpG 103; NDRG4_CpG 192; CACNAlG_CpG 241; NEUROGl_CpG 151; IGF2_CpG 330; SOCSl_CpG 226; RUNX3_CpG 357; AXIN2_CpG 215; SDC2_CpG 160. Regions that had at least one paired-end sequencing read were included in the analysis.

[0312]

[0249] The results of this experiment are shown in FIG. 15 A. Use of the primer having SEQ ID NO: 24 to selectively amplify adaptor-flanked gDNA fragments allowed detection of all top 20 hypermethylated CGIs associated with colon cancer in human genomic DNA obtained from Human HCT116 DKO Methylated DNA (Zymo #D5014-2). Use of the primer having SEQ ID NO: 25 to selectively amplify adaptor-flanked gDNA fragments allowed detection of 14 of the top 20 hypermethylated CGIs associated with colon cancer. In contrast, only 4 of the top 20 hypermethylated CGIs associated with colon cancer were detected using the standard EMseq protocol without the selective amplification step.

[0313]

[0250] To determine whether similar results could be achieved for known hypermethylated CGIs associated with human breast cancer, gDNA obtained from Human HCT116 DKO Methylated DNA (Zymo #D5014-2) was prepared and analyzed by ultra low-depth sequencing as described in Example 7. Prior to sequencing, adaptor-flanked human gDNA fragments were processed using a standard EMseq protocol, either with or without performing a selective amplification step using PCR using one of the primers listed in Table 12.

[0314] Table 12, Primers with structure 5’-A-P-3’ used to amplify methylated human gDNA

[0315]

[0251] After sequencing, the sequencing reads were aligned to the human reference genome to identify the genomic regions from which gDNA fragments were derived. To determine whether a read contained a CGI, the obtained binary alignment map (BAM) file was intersected with a predefined set of the top 20 hypermethylated CGIs in human breast cancer, thereby allowing the number of sequencing reads that overlapped with the CGIs to be counted. The top 20 hypermethylated CGIs were identified from the scientific literature, specifically Tang Q et al., 2016. Clin Epigenetics. Nov 14;8: 115. doi: 10.1186 / sl 3148-016-0282-6; Xiang TX et al., 2013. Chin J Cancer. Jan;32(l): 12-20. doi: 10.5732 / cjc.011.10344; Van De Voorde L et al. 2012. Mutat Res. Oct-Dec;751(2): 304-325. doi: 10.1016 / j.mrrev.2012.06.001; Wu Y et al., 2015. Methods Mol Biol. 1238:425-66. doi: 10.1007 / 978-1-4939-1804-l_23; and Kristiansen S et al., 2013. Int J Biol Markers. Apr- Jun;28(2): 141-50. doi: 10.5301 / jbm.5000009. Specifically, these are ATM CpG 75; SOX17_CpG215; RASSFI CpG 130; BRCAl_promoter; GSTPI CpG 96; HIN-l_CpG 167; APC_promoter; CCND2 CpG 379; CDHI CpG 103; TWISTI CpG 158; DAPKI CpG 145; ESRl_CpG 105; CDKN2A_CpG 156; HOXA5_CpG 202; TIMP3_CpG 81; SFRPI CpG 148; SLIT2_CpG 284; MGMT CpG 98; GATA3_CpG 504; FOXCl_CpG 584. Regions that had at least one paired-end sequencing read were included in the analysis.

[0316]

[0252] The results of this experiment are shown in FIG. 15B. Use of the primer having SEQ ID NO: 25 to selectively amplify adaptor-flanked gDNA fragments allowed detection of 18 of the top 20 hypermethylated CGIs associated with breast cancer in human genomic DNA obtained from Human HCT116 DKO Methylated DNA (Zymo #D5014-2). Use of the primer having SEQ ID NO: 26 to selectively amplify adaptor-flanked gDNA fragments allowed detection of 13 of the top 20 hypermethylated CGIs associated with breast cancer. In contrast, only 3 of the top 20 hypermethylated CGIs associated with breast cancer were detected using the standard EMseq protocol without the selective amplification step.

[0317]

[0253] Thus, this example demonstrates that adaptor-flanked gDNA fragments can be selectively amplified to enrich for GC-rich regions such as CGIs using a primer having the structure 5’-A-P-3’ with probe nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide. This enrichment step can markedly improve the coverage of cancer-associated hypermethylated CGIs when analyzing human gDNA by ultra-low depth sequencing.

[0318] Example 9. Detection of CIMP using ultra low-depth sequencing

[0319]

[0254] This example demonstrates that the CpG island methylator phenotype (CIMP) associated with colon cancer (i.e., the Weisenberger panel) can be detected in human genomic DNA when using ultra low-depth sequencing when a primer having structure 5’-A-P-3’ with a probe nucleotide sequence P comprising 6-12 nucleotides including at least one 5’-CG-3’ dinucleotide is used to selectively amplify GC-rich regions prior to sequencing.

[0320]

[0255] Example 8 confirmed the suitability of the described primers for selectively amplifying GC-regions comprising hypermethylated CGIs associated with human cancers. In this example, the efficiency of different primers with structure 5’-A-P-3’ to selectively amplify CGIs was compared.

[0321]

[0256] For FIG. 16, Human HCT116 DKO Methylated DNA (Zymo #D5014-2) was prepared and analyzed by ultra low-depth sequencing as described in Example 7. Specifically, prior to sequencing, adaptor-flanked human gDNA fragments were processed using a standard EMseq protocol, either with or without performing a selective amplification step using PCR using one of the first primers listed in Table 13 and the second (universal) primer provided in Example 7 (SEQ ID NO: 47). Sequencing was performed to achieve a read depth of 200,000 sequencing reads.

[0322] Table 13, Primers with structure 5’-A-P-3’ used to amplify methylated DNA

[0323] ♦abbreviation is used in FIGs 16 and 17. **For octamer 3, each M is independently selected from A or C.

[0324]

[0257] After sequencing, the sequencing reads were aligned to the human reference genome to identify the genomic regions from which gDNA fragments were derived. To determine whether a read contained a CpG island methylator phenotype (CIMP) in colon cancer, the obtained binary alignment map (BAM) file was intersected with the Weisenberger panel described in Example 6, thereby allowing the number of sequencing reads that overlapped with the CGIs to be counted. The results are summarized in FIG. 17. This figure shows that standard EMseq without selective amplification (denoted “Regular Illumina”) identified none of the five of CGIs in the Weisenberg panel. In contrast, when a selective amplification step with one of the primers comprising an octamer probe nucleotide sequence shown in Table 13 was included, all five CGIs in the Weisenberg panel could be detected using ultra-low depth sequencing.

[0325]

[0258] The experiment was repeated using first primers listed in Table 14 instead and a second (universal) primer with nucleotide sequence 5’-CACTCTTTCCCTACACGA*C-3’ (SEQ ID NO: 47), where A*C = an adenosine nucleotide and a cytosine nucleotide linked by a phosphorothioate bond. The length of the probe nucleotide sequence P ranged from 6 nucleotides to 12 nucleotides. The results of this experiment are summarized in FIG. 18.

[0326] Table 14, Primers with structure 5’-A-P-3’ used to amplify methylated DNA

[0327]

[0328] ♦abbreviation is used in FIG 18.

[0329]

[0259] As was observed for the primers listed in Table 13, inclusion of a selective amplification step using one of the primers in Table 14 also allowed detection of all five of CGIs in the Weisenberg panel using ultra-low depth sequencing, whereas standard EMseq without the amplification step (denoted “Regular Illumina”) did not allow detection of any of the five CGIs in the Weisenberg panel (FIG. 18). The results of these experiments indicate that selective amplification using the primers described herein can be used to reduce sequencing depth (thereby saving costs) whilst simultaneously increasing coverage of GC-rich regions of interest.

[0330]

[0260] This example demonstrates that the CIMP in colon cancer (i.e., the Weisenberger panel) can be detected when using ultra low-depth sequencing when a primer having structure 5’-A-P-3’ with a probe nucleotide sequence P comprising 6-12 nucleotides including at least one 5’-CG-3’ dinucleotide is used to selectively amplify GC-rich regions prior to sequencing.

[0331] Example 10. Selective amplification of methylated GC-rich regions from cancer tissue

[0332]

[0261] This example demonstrates that a primer having structure 5’-A-P-3’ with a probe nucleotide sequence P including at least one 5’-CG-3’ dinucleotide can selectively amplify adapter-flanked gDNA fragments comprising methylated GC-rich regions prior to sequencing, even at low concentrations, when the gDNA fragments are obtained from a cancer sample.

[0333]

[0262] Two samples of genomic DNA were used, a colon cancer sample designated CT14 obtained from a patient tumor biopsy and a matched normal control designated CN14 obtained from the same patient in tumor-adjacent normal tissue. The CpG methylation signature in colon cancer sample CT14 was assessed in an undiluted sample, or in samples serially diluted into CN14, as provided in Table 15. Undiluted CN14 was used as a control. Table 15,

[0334]

[0263] Each sample of Table 15 was prepared for sequencing according to the method set out in Example 7 and selective amplification was performed using PCR using a first primer having a probe nucleotide sequence P of 5’-GACCCGCG-3’ (the full sequence of the first primer is shown in Table 10 “Octamer”, i.e., SEQ ID NO: 20) and the second (universal) primer provided in Example 7 (SEQ ID NO: 47).

[0335]

[0264] The results of the sequencing analysis are shown in Table 16.

[0336] Table 16,

[0337]

[0265] As expected, 100% of CG dinucleotides in the CpG island of SEPTIN9 in the undiluted cancer sample CT 14 100% were determined as being methylated. In comparison, in the normal sample, i.e., CN14 100%, only 28.57% of CG dinucleotides in the CpG island of SEPTIN9 were determined as being methylated.

[0338]

[0266] Even at 0.01% dilution (see CT14 0.01%), 95.23% of CG dinucleotides in the CpG island of SEPTIN9 were determined as being methylated, demonstrating the effectiveness of selective amplification of adapter-flanked genomic DNA fragments comprising a CG dinucleotide using a primer having structure 5’-A-P-3’ having probe nucleotide sequence P of 5’-GACCCGCG-3’ (the full sequence of the first primer is shown in Table 10 “Octamer”, i.e., SEQ ID NO: 20).

[0339]

[0267] Next, other colon cancer samples and patient-matched normal samples were prepared, selectively amplified using a primer having structure 5’-A-P-3’ having probe nucleotide sequence P of 5’-GACCCGCG-3’ (the full sequence of the first primer is shown in Table 10 “Octamer”, i.e., SEQ ID NO: 20) in the same manner as CT14 and CN14. Higher depth sequencing was used, using a read depth of 10 million reads per sample for NovaSeq platform. The results of this analysis are shown in Table 17.

[0340]

[0268] Table 17 shows the coverage and methylation percentage of each of CpGs (CG dinucleotide) 1-13 in a region of chromosome 17 (chr.17: 77,373,475-77,373,62). The position of CpGs 1-13 in this region are shown in FIG. 19. In Table 17, coverage is shown by the “n” number for each entry in the table, which refers to the total number of paired-end reads which covered each of CpGs 1-13. Methylation percentage is shown by the “%” value, which refers to the percentage of paired-end reads for each of CpGs 1-13 that comprised a CG dinucleotide (i.e., that originally comprised a methylated C in the CG dinucleotide in the genomic DNA, that was not converted by TET2 oxidation and APOBEC deamination) at the position of CpGs 1-13.

[0341] Table 17.

[0342]

[0269] This example demonstrates that adaptor-flanked gDNA fragments obtained from human tissue (cancer or normal tissue) can be selectively amplified to enrich for GC-rich regions such as CGIs using a primer having the structure 5’-A-P-3’ with probe nucleotide sequence P comprising at least one 5’-CG-3’ dinucleotide. Selective amplification is effective even when starting from gDNA that is at a low concentration.

[0343] Example 11. Selective amplification of unmethylated GC-rich regions

[0344]

[0270] This example demonstrates that a primer having structure 5’-A-P-3’ with a probe nucleotide sequence P including at least one 5’-CG-3’ dinucleotide can selectively amplify adapter-flanked gDNA fragments comprising unmethylated GC-rich regions from a mixture of adapter-flanked gDNA fragments, even at low concentrations.

[0345]

[0271] To assess the percentage of unmethylated CpGs (i.e., CG dinucleotides) in CGIs of different genes, unmethylated gDNA or methylated gDNA (Zymo #D5014) was prepared for sequencing as follows. First, an undiluted unmethylated gDNA sample was prepared, or samples of unmethylated gDNA serially diluted into methylated gDNA were prepared, as provided in Table 18. gDNA was fragmented, ends were repaired, adaptors were ligated, and the resulting adaptor-flanked gDNA fragments were purified according to the method used in Example 7. In contrast to Example 7, the adaptors did not comprise methylated cytosine nucleotides.

[0346]

[0272] Next, methylated cytosines in the adapter-flanked gDNA fragments were converted via the TAPS protocol. Specifically, a TET oxidation step was carried out using NEB #E7120S according to the manufacturer’s instructions. Next, 30 pL of the mixture of adaptor-flanked gDNA fragments was combined with 10 pL of TET2 buffer, 1 pL of DTT, and 4 pL of TET2 enzyme. Add 5 pL of diluted Fe(II) solution and incubate at 37°C for 1 hour. Then, 1 pL of Stop Reagent was added and incubated at 37°C for 30 minutes. This step was followed with 1.8 AMPure purification and diluted in 35 pL of water. Next, Pic-Borane Reduction was performed as follows. 10 pL of 0.5 M sodium acetate buffer (pH 4.0) was added to 35 pL of the mixture of adaptor-flanked gDNA fragments in a 1.5 mL DNA LoBind Eppendorf tube and mixed by pipetting up and down. 5 pL of 1 M pic-borane was added to the 45 pL DNA sample. The sample was then incubated in a 50°C water bath for 1 hour. Purification using theZymo-IC column with oligo binding buffer (Zymo #D4060) was then performed.

[0347]

[0273] A further unmethylated-specific amplification step (i.e., a selective amplification step) using PCR was carried out prior to sequencing using a first primer having structure 5’-A-B-P-3’ having nucleotide sequence

[0348] 5 ’ -CGTGTGCTCTTCCGATCTAATATTAGTTAGAGTTGAACCCGCG-3 ’ (see Example 3, Table 5, Hexamer 2 with bridge, i.e., SEQ ID NO: 30) or a first primer having structure 5’-A-P-3’ having nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTCCCGCG-3 ’ (SEQ ID NO: 14) and a second (universal) primer having sequence 5’-ATACACTCTTTCCCTACACGACGC-3’ (SEQ ID NO: 43).

[0349]

[0274] The first primer was capable of specifically hybridizing to adapter-flanked gDNA fragments comprising a sequence complementary to probe nucleotide sequence P, allowing DNA polymerase to initiate amplification. The first and second primers act as forward and / or reverse primers depending on whether the probe hybridizes to the top strand or bottom strand, or both. Table 18.

[0350]

[0275] Sequencing was then performed using the MiSEQ platform with 2x250 bp paired-end reads and 200,000 reads per sample. This represents an “ultra low-depth sequencing”.

[0351]

[0276] The methylation status of the SDC2 gene CGI on chromosome 8 (chr8: 96,494,925 - 96,495,031) was assessed. The results are shown in Table 19.

[0352] Table 19,

[0353]

[0277] The experiment was repeated, instead using a first primer having structure 5’-A-B-P-3’ having nucleotide sequence

[0354] 5 ’ -CGTGTGCTCTTCCGATCTAATATTAGTTAAGACCCGCG-3 ’ (see Example 3, Table 5, Octamer with bridge, i.e., SEQ ID NO: 31) or a first primer having structure 5’-A-P-3’ having nucleotide sequence 5 ’ -CGTGTGCTCTTCCGATCTAATATTGACCCGCG-3 ’ (see Example 3, Table 5, Octamer, i.e., SEQ ID NO: 20) and a second (universal) primer having sequence 5’-ATACACTCTTTCCCTACACGACGC-3’ (SEQ ID NO: 43).

[0278] The methylation status of the CACNA2D2 gene CGI on chromosome 3 (chr3: 50,372,371 - 50,372,418) was assessed. The results are shown in Table 20.

[0355] Table 20,

[0356]

[0279] This example demonstrates that selective amplification combined with TAPS conversion and a first primer comprising a probe nucleotide sequence P having at least one CG dinucleotide enables enrichment of unmethylated CGI targets, even at low unmethylated DNA input levels.

[0357]

[0280] It should be understood that the details provided herein are given by way of illustration only, not limitation. Other features, objects, and advantages are apparent from the above detailed description, drawings and examples. Various changes and modifications will be apparent to those skilled in the art.

[0358]

[0281] All patents, patent publications, and non-patent publications referenced herein are indicative of the level of skill of those skilled in the art to which this invention pertains. All these publications are herein incorporated by reference to the same extent as if each individual publication were specifically and individually indicated as being incorporated by reference.

Claims

1. CLAIMS1. A primer capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, the primer having the following structure: 5’-A-B-P-3’, wherein:A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; andP is a probe nucleotide sequence comprising n nucleic acids including at least one 5’-CG-3’ dinucleotide, wherein and n=4-18.

2. The primer of claim 1, wherein P comprises two or more CG dinucleotides.

3. The primer of claim 2, wherein the two or more CG dinucleotides are consecutive.

4. The primer of any one of the preceding claims, wherein P comprises at least one CG dinucleotide at the 3’ end and / or at least one CG dinucleotide at the 5’ end.

5. The primer of any one of the preceding claims, wherein at least 50% or at least 60% of the nucleic acids comprised in P are cytosine (C).

6. The primer of any one of the preceding claims, wherein P comprises the nucleotide sequence 5 -CGCCCG-3’, 5 -CCCGCG-3 ’, 5 -CGACCCGCG-3 ’, 5 -GACCCGCG-3’, 5’-CGCGAA-3’, 5’-ACGCGA-3’, 5’-CGAACG-3’, 5’-CGACGA-3’, 5’- CGAACGCGAA-3 ’ (SEQ ID NO: 2), 5’-CGAACGCGAC-3’ (SEQ ID NO: 3), 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), 5’-AACGCG-3’, 5’-CGAACGCGTA-3’ (SEQ ID NO: 4), or 5’-CGMCGMCG-3’ wherein each M is independently selected from A or C, 5’-CGAACGCG-3’, or 5’-CGCGAACGCGAT-3’ (SEQ ID NO: 5).

7. A primer capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, the primer having the following structure: 5’-A-B-P-3’, wherein:A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; andP is a probe nucleotide sequence comprising n nucleic acids including at least one 5’-TG-3’ dinucleotide, wherein and n=4-18.

8. The primer of claim 7, wherein P comprises two or more TG dinucleotides.

9. The primer of claim 8, wherein the two or more TG dinucleotides are consecutive.

10. The primer of any one of claims 7-9, wherein P comprises at least one TG dinucleotide at the 3’ end and / or at least one TG dinucleotide at the 5’ end.

11. The primer of any one of claims 7-10, wherein at least 50% or at least 60% of the nucleic acids comprised in P are thymine (T).

12. A primer capable of specifically hybridizing to a genomic DNA fragment comprising an adaptor, or an amplicon thereof, the primer having the following structure: 5’-A-B-P-3’, wherein:A is a nucleotide sequence complementary to at least 10 consecutive nucleotides comprised in the adaptor;B is an optional bridging nucleotide sequence that is not complementary to a nucleotide sequence comprised in the adaptor; andP is a probe nucleotide sequence comprising n nucleic acids including at least one 5’-CA-3’ dinucleotide, wherein and n=4-18.

13. The primer of claim 12, wherein P comprises two or more CA dinucleotides.

14. The primer of claim 13, wherein the two or more CA dinucleotides are consecutive.

15. The primer of any one of claims 12-14, wherein P comprises at least one CA dinucleotide at the 3’ end and / or at least one CA dinucleotide at the 5’ end.

16. The primer of any one of claims 12-15, wherein at least 50% or at least 60% of the nucleic acids comprised in P are adenine (A).

17. The primer of any one of the preceding claims, wherein n is at least 5.

18. The primer of claim 17, wherein n is 6, 8, 9, 10, or 12.

19. The primer of any one of the preceding claims, wherein B comprises 5-30 nucleotides.

20. The primer of claim 19, wherein B comprises the nucleotide sequence 5’- AGTTAGAGTTGAA-3’ (SEQ ID NO: 11) or 5’-AGTTAA-3’.

21. A composition comprising the primer according to any one of the preceding claims.

22. A kit comprising the primer according to any one of claims 1-20 or the composition according to claim 21.

23. The kit of claim 22, further comprising a DNA polymerase.

24. The kit of claim 22 or claim 23, further comprising free nucleotides, wherein the free nucleotides comprise adenine, cytosine, guanine, and thymine.

25. A method of selectively amplifying one or more adaptor-flanked genomic DNA fragment(s) from a mixture of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments are deaminated to uracils; b. contacting the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein: the one or more adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-CG-3’ and the first primer is according to any one of claims 1-6 or 17-20, or the one or more adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-UG-3’ and the first primer is according to any one of claims 12- 20; and c. performing the PCR to obtain PCR products.

26. A method of selectively amplifying one or more amplicons of adaptor-flanked genomic DNA fragment(s) from a mixture of amplicons of adaptor-flanked genomic DNA fragments prepared from a sample, the method comprising: a. obtaining one or more amplicons of adaptor-flanked genomic DNA fragments, wherein unmethylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils or methylated cytosine nucleotides in the genomic DNA fragments were deaminated to uracils; b. contacting the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) with a first primer and a second primer under conditions suitable for a polymerase chain reaction (PCR), wherein: the one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise at least one 5’-CG-3’ and the first primer is according to any one of claims 1-6 or 17-20, orthe one or more amplicons of the adaptor-flanked genomic DNA fragment(s) comprise 5’-CA-3’ and / or 5’-TG-3’ and the first primer is according to any one of claims 7-20; and c. performing the PCR to obtain PCR products.

27. The method of claim 25 or 26, wherein deamination is performed by chemical or enzymatic treatment, or a combination of both.

28. The method of claim 27, wherein deamination of unmethylated cytosine nucleotides is performed chemically using sodium bisulfite treatment.

29. The method of claim 27, wherein deamination of unmethylated cytosine nucleotides is performed enzymatically using a eukaryotic deaminase such as apolipoprotein B mRNA editing catalytic polypeptide-like (APOBEC).

30. The method of claim 29, wherein the APOBEC is APOBEC3A.

31. The method of claim 25 or 26, wherein the deamination of methylated cytosine nucleotides is performed using TET1 and pyridine borane.

32. The method of any one of claims 25-31, wherein the adaptors flanking the one or more genomic DNA fragment(s) are deamination-resistant.

33. The method of claim 32, wherein the deamination-resistant adaptors comprise5 -methylcytosine, 5 -hydroxymethylcytosine, glycosylated 5-hydroxymethylcytosine, N4-methylcytosine, or 5-carboxylcytosine in place of cytosine.

34. The method of any one of claims 25-33, wherein the genomic DNA fragments are isolated from a sample obtained from a subject.

35. The method of claim 34, wherein the subject has or is suspected of having a disease or disorder.

36. The method of claim 35, wherein the disease or disorder is selected from a cancer, an autoimmune disease or disorder, a metabolic disease or disorder, a cardiovascular disease or disorder, and a neurological disease or disorder.

37. The method of any one of claims 34-36, wherein the sample is a sputum, urine, stool, saliva, blood, or cerebrospinal fluid sample.

38. The method of any one of claims 34-37, wherein the method comprises isolating genomic DNA from the sample, fragmenting the isolated genomic DNA to obtain the genomic DNA fragments, and adding adaptors to the ends of genomic DNA fragments.

39. The method of any one of claims 25-38, wherein the second primer is complementary to at least 10 nucleotides comprised in the adaptor.

40. The method of any one of claims 25-39, wherein the first primer is a forward primer and the second primer is a reverse primer; or the first primer is a reverse primer and the second primer is a forward primer.

41. The method of any one of claims 25-40, wherein the method comprises sequencing the PCR products to obtain sequencing reads.

42. A method of determining methylation status of genomic DNA of a subject, comprising: a. providing PCR products obtained from the method according to any of claims 25-40, and b. sequencing the PCR products.

43. The method of claim 41 or 42, further comprising aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived.

44. The method of claim 43, wherein the one or more genomic region(s) are one or more gene promoter(s) or CpG islands.

45. A method for determining whether one or more GC-rich regions of a genome of a subject is methylated, the method comprising: a. adding deamination-resistant first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject; b. deaminating unmethylated cytosines comprised in the adaptor-flanked genomic DNA fragments; c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer according to any one of claims 1-6 or 17-20 and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein: i. nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA, ii. probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and iii. the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; and f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were methylated.

46. A method for determining whether one or more GC-rich regions of a genome of a subject is unmethylated, the method comprising: a. adding first and second adaptors to the ends of genomic DNA fragments isolated from a sample obtained from a subject; b. deaminating methylated cytosines comprised in the adaptor-flanked genomic DNA fragments; c. contacting the adaptor-flanked genomic DNA fragments with a first forward primer according to any one of claims 1-6 or 17-20 and a reverse primer under conditions suitable for a polymerase chain reaction (PCR), wherein: i. nucleotide sequence A of the forward primer specifically hybridizes to at least 10 consecutive nucleotides of the first adaptor of the adaptor-flanked genomic DNA, ii. probe nucleotide sequence P of the forward primer specifically hybridizes to 4-18 consecutive nucleotides including a 5’-CG-3’ dinucleotide present in one or more of the genomic DNA fragments, and iii. the reverse primer specifically hybridizes to at least 10 consecutive nucleotides of the second adaptor, d. performing the PCR to obtain PCR products; e. sequencing the PCR products to obtain sequencing reads; and f. aligning the sequencing reads to a reference genome to identify the one or more genomic region(s) from which the one or more genomic DNA fragments are derived; wherein presence of sequencing reads of the one or more GC-rich regions indicates that the one or more GC-rich regions of the genome were unmethylated.

47. The method of claim 45 or claim 46, further comprising a second forward primer according to any one of claims 1-6 or 17-20, wherein the first forward primer and the second forward primer comprise different probe nucleotide sequence Ps.

48. The method of claim 47, wherein the first forward primer comprises a probe nucleotide sequence P comprising one 5’-CG-3’ dinucleotide and the second forward primer comprises a probe nucleotide sequence P comprising two 5’-CG-3’ dinucleotides.

49. The method of claim 48, wherein the first forward primer comprises a probe nucleotide sequence P comprising 5’-CGCCCG-3’ and the second forward primer comprises a probe nucleotide sequence P comprising 5’-CCCGCG-3’.

50. The method of claim 49, wherein the first forward primer comprises a probe nucleotide sequence P comprising two non-consecutive 5’-CG-3’ dinucleotides and the second forward primer comprises a probe nucleotide sequence P comprising two consecutive 5’-CG-3’ dinucleotides.

51. The method of claim 50, wherein the first forward primer comprises a probe nucleotide sequence P comprising 5 ’ -CGACCCGCG-3 ’ and the second forward primer comprises a probe nucleotide sequence P comprising 5’-GACCCGCG-3’.

52. The method of claim 51, wherein the first forward primer comprises a probe nucleotide sequence P comprising three 5’-CG-3’ dinucleotides and the second forward primer comprises a probe nucleotide sequence P comprising two consecutive 5’-CG-3’ dinucleotides.

53. The method of claim 52, further comprising a third primer according to any one of claims 1-6 or 17-20, wherein the first and second forward primers and the third forward primer comprise different probe nucleotide sequence Ps.

54. The method of claim 53, wherein the first forward primer comprises a probe nucleotide sequence P comprising two consecutive 5’-CG-3’ dinucleotides, the second primer comprises a probe nucleotide sequence P comprising three 5’-CG-3’ dinucleotides, and the third primer comprises a probe nucleotide sequence P comprising four 5’-CG-3’ dinucleotides.

5. The method of claim 54, wherein the first forward primer comprises a probe nucleotide sequence P comprising 5 ’ -AACGCG-3 ’ , the second forward primer comprises a probe nucleotide sequence P comprising 5’-CGCGAACGCGTA-3’ (SEQ ID NO: 1), and the third forward primer comprises a probe nucleotide sequence P comprising 5’- CGAACGCGAA-3 ’ (SEQ ID NO: 2).

Citation Information

Patent Citations

  • Normalization of NGS library concentration

    US10961562B2

  • Methods of enriching for and identifying polymorphisms

    US20140024027A1

  • Method and kits for preparing multicomponent nucleic acid constructs

    US7223539B2