Detection of Epigenetic Cytosine Modifications

Enzymatic reduction with enreductase addresses the toxicity issues of chemical methods by converting cytosine modifications to DHU, enhancing detection sensitivity and safety in epigenetic analysis.

JP2025520439APending Publication Date: 2025-07-03F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024573504
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-14
Filing Date
2023-06-12
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for detecting epigenetic cytosine modifications, such as the TAPS method, rely on toxic chemical reagents like pyridine borane for reducing the C5-C6 double bond of cytosine, posing environmental and safety concerns.

Method used

Employing enzymatic reduction using enreductase to convert cytosine modifications like 5-formylcytosine (5fC) and 5-carboxycytosine (5caC) to dihydrouracil (DHU), replacing the use of toxic chemical reagents.

Benefits of technology

Provides a safe and environmentally friendly method for detecting epigenetic cytosine modifications, reducing DNA degradation and enabling high-sensitivity sequencing without the use of harmful chemicals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520439000001
    Figure 2025520439000001
  • Figure 2025520439000002
    Figure 2025520439000002
  • Figure 2025520439000003
    Figure 2025520439000003
Patent Text Reader

Abstract

The present invention includes improved methods and compositions for reducing the C5-C6 double bond of cytosine. In particular, the improved methods and compositions for reducing the C5-C6 double bond of cytosine use enzymatic means rather than chemical means. In particular, the present disclosure relates to methods for converting 5,6-dihydro-fC (fC) and / or 5,6-dihydro-caC to 5,6-dihydro-U (DHU). In particular, the present disclosure relates to methods for converting 5fC and / or 5caC to DHU. Furthermore, the present disclosure relates to a method for detecting epigenetic cytosine modifications, particularly methylation of cytosine, using an enredactase that reduces the C5-C6 double bond of cytosine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of nucleic acid-based diagnostics. More specifically, the present invention relates to a method for detecting epigenetic cytosine modifications in nucleic acids, wherein the epigenetic modifications can have biological and / or clinical significance.

Background Art

[0002] Mapping of epigenetic modifications of DNA and RNA is becoming increasingly important as these modifications play roles in several biological processes and diseases, including development, aging, cancer, etc. The most appropriate for the identification of epigenetic modifications is detection by sequencing-based methods.

[0003] The discrimination between cytosine (C) and 5-methylcytosine (5mC) is accounted for by direct sequencing of DNA through nanopores (without pre-amplification), e.g., by Pacific Bioscience Single Molecule, Real-Time (SMRT) sequencing technology, thereby reading the kinetic differences in the incorporation of nucleotides opposite C versus mC by polymerase. Other techniques use methods to convert C or mC to T(U) equivalents and subsequent amplification to enable identification when comparing untreated sample DNA with the converted sample DNA (see Zhao, et al., ‘‘Mapping the epigenetic modifications of DNA and RNA,’’ Protein&Cell 11(11):792-808(2020)). The most important methods are: (a) bisulfite sequencing (conversion of C to uracil (U) by bisulfite treatment), (b) New England Biolab’s (NEB’s) EM-seq method (oxidation of mC by TET2 enzyme and glucosylation by β-glucosyltransferase (to block the enzymatic deaminase reaction) and conversion of C to U by APOBEC deaminase), (c) TAPS method (TET-assisted pyridine borane sequencing) that applies TET enzyme oxidation of mC and subsequent reduction of the oxidized mC species by pyridine borane to obtain dihydrocytosine nucleosides that are readily deaminated to give dihydrouracil (T equivalent), and (d) CLEVER method based on the subsequent reaction of 5-formyl-C with malononitrile to give an adduct that acts mainly as a T equivalent in subsequent PCR amplification after TET enzyme oxidation of mC (Zhu, et al., ‘‘Single-Cell 5-Formylcytosine Landscapes of Mammalian Early Embryos and ESCs at Single-Base Resolution,’’ Cell Stem Cell 20:720-731(2017)). Instead of enzymatic oxidation of mC by TET enzyme, oxidation can also be carried out by chemical means, e.g., using potassium perruthenate (KRuO4).There is also a method for distinguishing modified 5-hydroxymethyl-dC (5hmC), 5-formyl dC (5fC), or 5-carboxy-dC (5caC) by partially modifying the above method.

[0004] The TAPS method uses the oxidation of 5-methylcytosine to 5-formylcytosine and / or 5-carboxycytosine based on the TET enzyme, followed by the reduction of the C5-C6 double bond of cytosine under the respective deformation and decarboxylation of cytosine. The formed 5,6-dihydrocytosine is easily deaminated at the N-4 position to give 5,6-dihydrouracil (DHU). The polymerase reads DHU as T, and as a result, incorporates A opposite DHU, making it possible to distinguish epigenetic C modifications (e.g., mC, hmC, fC, or caC) from unmodified C (see Liu, et al., ‘‘Bisulfite-free direct detection of 5-methylcytosine and 5-hydroxymethylcytosine at base resolution,’’ Nat. Biotechnol. 37(4):424-429(2019)).

[0005] The above TAPS method utilizes the chemical reduction of the C5-C6 double bond of cytosine. Borane (especially pyridine borane or 2-picoline borane) is used as the reducing agent. However, these chemical reagents and chemical methods are disadvantageous because they are toxic. Therefore, in the art, there is a need for a safe and environmentally friendly method for reducing the C5-C6 double bond of cytosine.

[0006] This patent application relates to the use of enzymatic reduction of the C5-C6 double bond of cytosine instead of reduction by chemical means (such as the TAPS method). Therefore, the chemical reduction by toxic reagents (of the TAPS method) is replaced by an environmentally friendly enzymatic reduction process.

Summary of the Invention

[0007] Therefore, in the art, there is a need for a safe and environmentally friendly method for reducing the C5-C6 double bond of cytosine. This patent application relates to the use of enzymatic reduction of the C5-C6 double bond of cytosine instead of reduction by chemical means (such as the TAPS method). Thus, chemical reduction with toxic reagents (of the TAPS method) is replaced by an environmentally friendly enzymatic reduction process.

[0008] The enzyme used is an enzyme of the enreductase group that can reduce electron-deficient C-C double bonds, particularly double bonds having an electron-withdrawing substituent. In the case of fC or caC, both the formyl and carboxy groups are electron-withdrawing groups and can thus reduce fC and / or caC, but cannot reduce unsubstituted C and U or mC and T. Figure 1 shows the reduction of 5,6-dihydro-fC (fC) to 5,6-dihydro-U (DHU). Figure 2 shows the reduction of 5,6-dihydro-caC (caC) to 5,6-dihydro-U (DHU). Enreductase (ERED) can be purchased commercially (e.g., Codexis, Redwood City, California, USA). Figure 3 shows the reaction for the purpose of enreductase (ERED) reducing a double bond (cited from CODEXIS, Codex ERED Screening Kit, Screening Protocol, document #PRO-004-003, page 1 (accessed June 10, 2022, https: / / www.codexis-estore.com / _files / ugd / 5a7b2a_5fbf41f3ae2741a081894f64424552d1.pdf)).

[0009] One embodiment relates to a method for reducing the C5-C6 double bond of cytosine in a nucleic acid, which includes contacting cytosine in the nucleic acid with an enreductase. In one embodiment, the cytosine is 5-formylcytosine (5fC). In one embodiment, the cytosine is 5-carboxycytosine (5caC). In one embodiment, the cytosine is 5,6-dihydro-fC. In one embodiment, the cytosine is 5,6-dihydro-caC. In one embodiment, the cytosine is 2'-deoxy-5-formylcytidine. In one embodiment, the cytosine is 2'-deoxy-5-carboxycytidine. In one embodiment, the cytosine is 5-formyl-2'-deoxycytidine. In one embodiment, the cytosine is 5-formyl-dC. In one embodiment, the cytosine is 5-carboxy-2'-deoxycytidine. In one embodiment, the cytosine is 5-carboxy-dC. In one embodiment, when cytosine in the nucleic acid is contacted with an enreductase, the cytosine is reduced to 5,6-dihydro-U (DHU). In one embodiment, the DHU is then converted to thymine. In one embodiment, the reduction of the C5-C6 double bond of cytosine in the nucleic acid by contacting cytosine in the nucleic acid with an enreductase is part of a method for detecting epigenetic cytosine modifications.

[0010] Another embodiment relates to a method for detecting epigenetic cytosine modifications, the method comprising at least reducing the C5-C6 double bond of cytosine in a nucleic acid by contacting the cytosine in the nucleic acid with an enreductase. In one embodiment, the cytosine is 5-formylcytosine (5fC). In one embodiment, the cytosine is 5-carboxylcytosine (5caC). In one embodiment, the cytosine is 5,6-dihydro-fC. In one embodiment, the cytosine is 5,6-dihydro-caC. In one embodiment, the cytosine is 2'-deoxy-5-formylcytidine. In one embodiment, the cytosine is 2'-deoxy-5-carboxycytidine. In one embodiment, the cytosine is 5-formyl-2'-deoxycytidine. In one embodiment, the cytosine is 5-formyl-dC. In one embodiment, the cytosine is 5-carboxy-2'-deoxycytidine. In one embodiment, the cytosine is 5-carboxy-dC. In one embodiment, when the cytosine in the nucleic acid is contacted with an enreductase, the cytosine is reduced to 5,6-dihydro-U (DHU). In one embodiment, the DHU is subsequently converted to thymine.

[0011] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present subject matter, suitable methods and materials are described below. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0012] Details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the drawings and the detailed description of the embodiments for carrying out the invention, as well as from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] This patent or application file contains at least one drawing created in color. Copies containing color drawings of this patent or patent application publication will be provided by the Patent Office upon request and payment of the necessary fees.

Figure 1

Figure 2

Figure 3

Mode for Carrying Out the Invention

[0014] Abbreviations Some of the abbreviations used throughout this disclosure are listed below. C - Cytosine T - Thymine U - Uracil DHU - Dihydrouracil 5mC - 5-Methylcytosine 5hmC - 5-Hydroxymethylcytosine 5ghmC - 5-Glucosyl-hydroxymethylcytosine 5fC - 5-Formylcytosine 5caC - 5-Carboxycytosine TET - 10-11 translocation dioxygenase TAPS - TET-assisted pick balancing sequencing CAPS - Chemistry-assisted pick balancing sequencing oxBS or oxBS-Seq - Oxidative bisulfite sequencing

[0015] 5-Methylcytosine and 5-hydroxymethylcytosine (5mC and 5hmC) are important epigenetic biomarkers with many clinical applications in oncology, prenatal testing, and other fields. Until recently, base-level detection of methylation has been achieved by reacting unmethylated cytosine with bisulfite, followed by PCR, array hybridization, or sequencing. Unmethylated cytosine (C) is read as thymine (T) after reaction with bisulfite, and methylated cytosine (5mC and 5hmC) is read as C. Unfortunately, bisulfite treatment results in substantial degradation of nucleic acid samples and is not suitable for applications requiring high sensitivity. For example, it is not suitable for recent applications analyzing cell-free nucleic acids such as cell-free DNA.

[0016] In recent years, less harsh methods for detecting methylated cytosine have been disclosed. The latest methods involve modification of methylated cytosine instead of unmethylated cytosine as in the case of bisulfite treatment (Liu, et al., Nature Biotechnology 37:424-429 (2019) (hereinafter referred to as "Liu et al. (2019)")). Stepwise oxidation of methylcytosine (5mC) to formylcytosine (5fC) and carboxylcytosine (5caC) via 5-hydroxymethylcytosine (5hmC) is carried out using 10-11 translocation dioxygenase (TET) in the presence of Fe(II) ions and alpha-ketoglutaric acid.

[0017] Liu et al. (2019) further described reducing 5fC (and 5caC) to dihydrouracil (DHU) using borane derivatives (pyridine borane, picoline borane, etc.). Subsequently, DHU is read as T by uracil-tolerant nucleic acid polymerase in subsequent amplification and sequencing. As a result, methylated C is read as T, while unmethylated C remains unchanged. This TET- and picoline-borane-based method, called TAPS (TET-assisted picoline-borane sequencing), does not cause as much DNA degradation as bisulfite treatment and allows for direct detection of signals instead of subtracting background to obtain signals. Both advantages enable higher alignment speeds and, in some cases, lower sequencing depths, and recover higher molecular diversity from samples.

[0018] Another technique called CAPS (chemical-assisted picoline-borane sequencing) involves the selective conversion of 5hmC to 5fC using potassium perruthenate (KRuO4). The use of KRuO4 as a chemical alternative to TET is known from a technique called oxidative bisulfite sequencing or oxBS-seq (Booth, et al., Science 336(6083):934-937(2012) (hereinafter referred to as "Booth et al. (2012)")). The 5fC obtained by potassium perruthenate conversion is a preferred target for further treatment, for example, by borane treatment or any other downstream method.

[0019] Yet another sequencing technique is an alternative to the reduction of 5fC by borane. This method involves forming an adduct of 5fC that is recognized as T. The adduct is formed using malononitrile (see Zhu et al. Cell Stem Cell 20:720-731(2017) (hereinafter referred to as "Zhu et al. (2017)")). The TAPS, CAPS, and malononitrile methods of Zhu et al. (2017) are superior to the bisulfite method in that they avoid harsh chemical treatment and the resulting loss of nucleic acid samples.

[0020] However, these methods have the drawback that they often use chemical reduction of the C5-C6 double bond of cytosine. Boranes (especially pyridine borane or 2-picoline borane) are used as reducing agents. However, these chemical reagents and chemical methods are disadvantageous because they are toxic. Therefore, in the art, there is a need for a safe and environmentally friendly method for reducing the C5-C6 double bond of cytosine.

[0021] This patent application relates to the use of enzymatic reduction of the C5-C6 double bond of cytosine instead of reduction by chemical means (such as the TAPS method). Thus, chemical reduction with (the toxic reagents of the TAPS method) is replaced by an environmentally friendly enzymatic reduction process.

[0022] In some embodiments, the invention is a method for detecting epigenetic modifications, specifically epigenetic cytosine modifications (including but not limited to methylation of cytosine) in nucleic acids. State-of-the-art methods for detecting epigenetic cytosine modifications may involve using harsh chemical methods to reduce the C5-C6 double bond. However, this patent application provides an improvement in the art. This patent application relates to the use of enzymatic reduction (i.e., not chemical reduction) of the C5-C6 double bond of cytosine.

[0023] The invention includes methods of manipulating nucleic acids from a sample. In some embodiments, the sample is derived from a subject or patient. In some embodiments, the sample can include, for example, by biopsy, a fragment of solid tissue or solid tumor derived from a subject or patient. The sample can also include a body fluid (e.g., urine, sputum, serum, blood or blood fraction, i.e., plasma, lymph, saliva, sputum, sweat, tears, cerebrospinal fluid, amniotic fluid, synovial fluid, pericardial fluid, ascites, pleural effusion, cyst fluid, bile, gastric juice, intestinal juice, or fecal sample) that contains nucleic acids. In other embodiments, the sample is a culture sample, e.g., a tissue culture containing cells and fluids from which nucleic acids can be isolated. In some embodiments, the nucleic acid of interest in the sample is derived from an infectious pathogen such as a virus, bacterium, protozoan, or fungus.

[0024] The present invention involves manipulating isolated nucleic acids that have been isolated or extracted from a sample. Methods for nucleic acid extraction are well known in the art (see Sambrook et al., ''Molecular Cloning: A Laboratory Manual,'', 1989, 2nd Ed., Cold Spring Harbor Laboratory Press: New York, N.Y.). Various kits for extracting nucleic acids (DNA or RNA) from biological samples (e.g., KAPA Express Extract (Roche Sequencing Solutions, Pleasanton, Cal.)), as well as other similar products from BD Biosciences Clontech (Palo Alto, Cal.), Epicentre Technologies (Madison, Wisc.); Gentra Systems (Minneapolis, Minn.); and Qiagen (Valencia, Cal.), Ambion (Austin, Tex.); BioRad Laboratories (Hercules, Cal.) are commercially available.

[0025] In some embodiments, for example, as described in International Publication No. WO 2019 / 092269 and International Publication No. WO 2020 / 074742, nucleic acids are extracted, separated by size, and optionally concentrated by epitacophoresis.

[0026] The present invention involves detecting epigenetic modifications, specifically epigenetic cytosine modifications (including but not limited to methylation of cytosine) in nucleic acids. Nucleic acid sequences that undergo conditional epigenetic modifications are target sequences analyzed by the methods disclosed herein. The same nucleic acid sequence may or may not have an epigenetic modification characterized by methylation of cytosine at the 5-position (5mC or 5hmC). In some embodiments, a set or panel of target nucleic acids is probed for the presence of methylation. For example, as shown in Patai, et al.’’Comprehensive DNA Methylation Analysis Reveals a Common Ten-Gene Methylation Signature in Colorectal Adenomas and Carcinomas’’PLOS ONE 10(8):e0133836(2015), and Onwuka, et al.,’’A panel of DNA methylation signature from peripheral blood may predict colorectal cancer susceptibility,’’BMC Cancer 20,692(2020), methylation of biomarkers in a panel of methylation biomarkers indicates the presence of colorectal cancer in a patient. Thus, testing of any known or future panel of methylation biomarkers for prognostic or diagnostic purposes is envisioned by the methods disclosed herein.

[0027] In some embodiments, the entire genome of an organism is probed for the presence of methylation. The methods of the present invention involve detecting methylation at all sites throughout the genome of an organism and diagnosing a disease or condition or a predisposition to a disease or condition using, for example, sequence analysis and artificial intelligence tools as described in Shull, et al.’’Sequencing the cancer methylome,’’Methods Mol Biol.1238:627-635(2015).

[0028] In some embodiments, the present invention includes an improved process for detecting methylated cytosine in a nucleic acid by forming and detecting 5-carboxylcytosine (5caC) and 5-formylcytosine (5fC), wherein 5fC, 5caC, or a mixture of 5fC and 5caC is formed by one of the methods described hereinabove. The method includes contacting a sample containing a nucleic acid comprising 5fC and / or 5caC with an enzyme, particularly an enreductase, to form dihydrouracil (DHU).

[0029] After the formation of DHU, the nucleic acid containing DHU is subjected to sequencing. In some embodiments, the sequencing is by a next-generation massively parallel sequencing process. The sequencing results in a test sequence in which DHU is read as thymine (T), i.e., the sequencing polymerase can accommodate the adduct or DHU in the copied strand and incorporate an adenine (A) opposite the DHU. The method further includes comparing the test sequence to a reference sequence, and a change from cytosine (C) in the reference sequence to thymine (T) at the corresponding position in the test sequence indicates the presence of methylated cytosine in the test nucleic acid.

[0030] In some embodiments, the nucleic acid in the sample is amplified prior to sequencing. In some embodiments, the amplification utilizes a B-family polymerase that efficiently incorporates an adenine (A) nucleotide opposite the DHU. In this embodiment, since DHU has already been recognized as T by the amplification polymerase, the sequencing can be carried out with any polymerase suitable for the sequencing process.

[0031] In some embodiments, the nucleic acid in the sample is ligated to an adapter, and the adapter contains elements useful in amplification and sequencing. The adapter includes at least one of a barcode, a primer binding site, and a ligation site.

[0032] In some embodiments, the present invention is an improved method for detecting methylated cytosine nucleotides in nucleic acids, comprising: (i) ligating an adapter containing an amplification primer binding site to the nucleic acids in a sample; (ii) forming a reaction mixture by contacting the sample containing the adapter-ligated nucleic acids with TET, which can convert methylated cytosine in the nucleic acids to 5-carboxycytosine (5caC) or a mixture of 5-formylcytosine (5fC) and 5caC; (iii) contacting the reaction mixture with a borane derivative capable of reacting with 5fC and 5caC in the nucleic acids to form DHU; (iv) incubating the reaction mixture for 1 hour or less, such that at least 90% of 5fC and 5caC have formed DHU; (v) amplifying the adapted nucleic acids using a DNA polymerase and a primer capable of binding to the primer binding site, wherein the DNA polymerase reads DHU as thymine (T) during amplification; (vi) sequencing the amplified nucleic acids to obtain a test sequence; and (vii) comparing the test sequence with a reference sequence, wherein a change from cytosine (C) in the reference sequence to thymine (T) at the corresponding position in the test sequence indicates the presence of 5mC in the nucleic acids. In some embodiments, in steps (iii) and (iv), the borane derivative is present in a non-aqueous solvent, such as ethanol or methanol. In some embodiments, in steps (iii) and (iv), the borane derivative is present in a solution containing an organic acid, such as acetic acid.

[0033] In some embodiments, the present invention utilizes adapters added to one or both ends of a nucleic acid or nucleic acid strand. Adapters of various shapes and functions are known in the art (see, for example, International Patent Application No. PCT / EP2019 / 05515 filed on February 28, 2019, US Patent Nos. 8,822,150 and 8,455,193). In some embodiments, the function of the adapter is to introduce a desired element into the nucleic acid. Adapter-derived elements include at least one of a nucleic acid barcode, a primer binding site, or a ligation-capable site.

[0034] The adapter can be double-stranded, partially single-stranded, or single-stranded. In some embodiments, a Y-shaped, hairpin adapter, or stem-loop adapter is used, and the double-stranded portion of the adapter is ligated to the formed double-stranded nucleic acid as described herein.

[0035] In some embodiments, the adapter molecule is an artificially synthesized sequence in vitro. In other embodiments, the adapter molecule is a naturally occurring sequence synthesized in vitro. In still other embodiments, the adapter molecule is a separated natural molecule or a separated non-natural molecule.

[0036] Double-stranded or partially double-stranded adapter oligonucleotides can have overhangs or blunt ends. In some embodiments, the double-stranded DNA can include blunt ends to which blunt-ended adapters can be ligated by applying blunt-end ligation. In other embodiments, the blunt-end DNA is A-tailed and a single A nucleotide is added to the blunt end to match an adapter designed to have a single T nucleotide extending from the blunt end, facilitating ligation between the DNA and the adapter. Commercially available kits for adapter ligation include the AVENIO ctDNA Library Prep Kit or the KAPA HyperPrep and HyperPlus kits (Roche Sequencing Solutions, Pleasanton, CA). In some embodiments, adapter-ligated (adapted) DNA can be separated from excess adapter and unligated DNA.

[0037] In some embodiments, the invention includes the use of barcodes. In some embodiments, methods for detecting epigenetic modifications include sequencing. Nucleic acids processed as described herein are subjected to sequencing, preferably massively parallel single molecule sequencing. Analyzing individual molecules by massively parallel sequencing typically requires different levels of barcoding for sample identification and error correction. Use of molecular barcodes as described in U.S. Patent Nos. 7,393,665, 8,168,385, 8,481,292, 8,685,678, and 8,722,368. A unique molecular barcode is added to each molecule to be sequenced to mark the molecule and its progeny (e.g., the original molecule and its amplicons generated by PCR). Unique molecular barcodes (UIDs) have multiple uses including counting the number of original target molecules in a sample and error correction (Newman, et al., An ultrasensitive method for quantitating circulating tumor DNA with broad patient coverage, Nature Medicine doi:10.1038 / nm.3519(2014)).

[0038] In some embodiments, unique molecular barcodes (UIDs) are used for error correction of sequencing. All progeny of a single target molecule are labeled with the same barcode, forming a barcoded family. Variations within the sequence not shared by all members of the barcoded family are discarded as artifacts. Since the entire family represents a single molecule in the original sample, the barcode can also be used for locus duplication exclusion and target quantification (Newman, et al. Integrated digital error suppression for improved detection of circulating tumor DNA, Nature Biotechnology 34:547(2016)).

[0039] In some embodiments of the present invention, the adapter ligated to one or both ends of the barcoded target nucleic acid contains one or more barcodes used in sequencing. The barcode can be a UID or a multiplex sample ID (MID or SID) used to identify the source of the sample when samples are mixed (multiplexed). The barcode can also be a combination of a UID and an MID. In some embodiments, a single barcode is used as both a UID and an MID. In some embodiments, each barcode contains a predetermined sequence. In other embodiments, the barcode contains a random sequence. In some embodiments of the present invention, the barcode is between about 4 and 20 bases in length, and as a result, 96 to 384 different adapters (each having a different pair of the same barcode) are added to a human genomic sample. In some embodiments, the number of UIDs in the reaction can exceed the number of molecules to be labeled. Those skilled in the art will recognize that the number of barcodes depends on the complexity of the sample (i.e., the expected number of unique target molecules) and that the appropriate number of barcodes for each experiment can be created.

[0040] In some embodiments, the method includes forming a library containing nucleic acids from a sample. The library consists of a plurality of nucleic acids that are ready for sequencing or another type of detection method, such as PCR. The library can be stored and used multiple times for further processing such as amplification or sequencing of the nucleic acids in the library. In some embodiments, the library is the input nucleic acid in which methylation is detected by the method described herein. In other embodiments, the library is formed from nucleic acids that have undergone the methylation detection reaction described herein.

[0041] In some embodiments, the nucleic acids processed for detection of epigenetic modifications by the methods described herein are sequenced. Any of a number of sequencing techniques or sequencing assays can be utilized. As used herein, the term "next-generation sequencing (NGS)" refers to sequencing methods that enable the ultra-parallel sequencing of cloned amplified molecules and single nucleic acid molecules.

[0042] Non-limiting examples of sequencing assays suitable for use with the methods disclosed herein include nanopore sequencing (U.S. Patent Application Publication Nos. 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycle sequencing (Sears et al., Biotechniques, 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods Mol. Cell Biol., 3:39-42 (1992)), sequencing using mass spectrometry, such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nature Biotech., 16:381-384 (1998)), sequencing by hybridization (Drmanac et al., Nature Biotech., 16:54-58 (1998), and including but not limited to, sequencing by synthesis (e.g., HiSeq™, MiSeq™, or Genome Analyzer, each available from Illumina), sequencing by ligation (e.g., SOLiD™, Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent™, Life Technologies), and NGS methods including SMRT® sequencing (e.g., Pacific Biosciences).

[0043] Commercially available sequencing technologies include the sequencing platforms by hybridization of Affymetrix Inc. (Sunnyvale, California), the sequencing platforms by synthesis of Illumina / Solexa (San Diego, California) and Helicos Biosciences (Cambridge, Massachusetts), and the sequence platforms by ligation of Applied Biosystems (Foster City, California). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisher Scientific), and nanopore sequencing (Genia Technology of Roche Sequencing Solutions (Santa Clara, California) and Oxford Nanopore Technologies (Oxford, UK)).

[0044] In some embodiments, the sequencing step includes sequence alignment. In some embodiments, the alignment is used to derive a consensus sequence from a plurality of sequences, e.g., having the same unique molecular ID (UID). The molecular ID is a barcode that can be added to each molecule prior to sequencing or, if an amplification step is included, prior to the amplification step. In some embodiments, the UID is present in the 5' portion of the RT primer. Similarly, the UID can be present at the 5' end of the last barcode subunit added to the compound barcode. In other embodiments, the UID is present in an adapter and is added to one or both ends of the target nucleic acid by ligation.

[0045] In some embodiments, the consensus sequence is determined from a plurality of sequences all having the same UID. Sequences having the same UID are presumed to be derived from the same original molecule via amplification. In other embodiments, the UID is used to eliminate artifacts, i.e., variations present in the progeny of a single molecule (characterized by a particular UID). Such artifacts resulting from PCR errors or sequencing errors can be eliminated using the UID.

[0046] In some embodiments, the number of each sequence in the sample can be quantified by quantifying the relative number of sequences having each UID within a population having the same multiplex sample ID (MID). Since each UID represents a single molecule in the original sample, by counting the different UIDs associated with each sequence variant, the fraction of each sequence variant in the original sample where all molecules share the same MID can be determined. One of ordinary skill in the art can determine the number of sequence reads necessary to determine the consensus sequence. In some embodiments, a reasonable number is the number of reads per UID (the "sequence depth") necessary for accurate quantification results. In some embodiments, the desired depth is 5 to 50 reads per UID.

[0047] In some embodiments, the invention is a kit comprising components and tools for performing an improved method of detecting DNA methylation described herein. In some embodiments, the kit comprises components for detecting methylation of cytosine in a nucleic acid by detecting the product of in vitro oxidized 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC). In some embodiments, the product is 5-formylcytosine (5fC) or 5-carboxycytosine (5caC). In other embodiments, the kit further comprises components for performing an in vitro oxidation of 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) to 5-formylcytosine (5fC) or 5-carboxycytosine (5caC).

[0048] In some embodiments, the kit contains an enzyme. In some embodiments, the enzyme is an enreductase. In some embodiments, the kit further contains an organic acid. In some embodiments, the kit contains instructions regarding the use of an organic acid (such as acetic acid) in a method for detecting DNA methylation that includes a borane derivative in a non-aqueous solvent as described herein. In some embodiments, the kit further contains a buffer such as MES or TRIS.

[0049] In some embodiments, the kit contains malononitrile and a non-aqueous solvent. The non-aqueous solvent is selected from ethanol and methanol. In other embodiments, instead of including a non-aqueous solvent, the kit contains instructions regarding the use of a non-aqueous solvent (such as ethanol or methanol) in a method for detecting DNA methylation with malononitrile as described herein. In some embodiments, the kit further contains an organic acid and a primary, secondary, or tertiary amine. The organic acid may be acetic acid and the amine may be triethanolamine. In other embodiments, the kit contains instructions regarding the use of an organic acid and an amine (such as acetic acid and triethanolamine) in a method for detecting DNA methylation with malononitrile as described herein. In some embodiments, the kit further contains a buffer such as MES or TRIS.

[0050] In some embodiments, the kit further comprises a TET enzyme for the in vitro oxidation of 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) to 5-carboxycytosine (5caC). In some embodiments, the TET is selected from mouse TET1, TET2 or TET3 (mTET1, 2 or 3), human TET1, TET2 or TET3 (hTET1, 2 or 3), Naegleria TET (NgTET), Coprinopsis cinerea (CcTET). In some embodiments, the TET is Naegleria TET-like oxygenase (NgTET1). In some embodiments, the TET is a wild-type protein. In other embodiments, the TET is a mutant protein. In some embodiments, the kit further comprises one or more cofactors selected from α-ketoglutaric acid and a source of Fe(II) ions.

[0051] In some embodiments, as an alternative to TET, the kit comprises a chemical oxidant, such as potassium perruthenate (KRuO4) or potassium ruthenate (K2RuO4).

[0052] In some embodiments, the kit further comprises a reagent for chemically blocking 5hmC so that the reaction containing 5mC is not affected. In some embodiments, the kit comprises a glucose compound and a glucosyltransferase capable of transferring the glucose moiety to the 5-hydroxyl moiety of 5hmC. In some embodiments, the kit comprises beta-glucosyltransferase (BGT) and UDP-glucose. In some embodiments, the BGT is T4 BGT.

[0053] In some embodiments, the method further comprises assessing the state of a subject (e.g., a patient) based on the methylation status of one or more loci within the genome of the patient. In some embodiments, the method comprises determining, in a sample from the patient, a genomic location and optionally the amount of methylated cytosine (5mC and / or 5hmC and / or 5fC and / or 5caC) within the genome. In some embodiments, loci known to be disease biomarkers are evaluated for methylation. The method further comprises selecting or modifying a treatment based on the diagnosis of a disease or condition of the patient or the presence or amount of methylation in a nucleic acid isolated from the patient.

[0054] There are several methods for identifying disease- or condition-specific methylation loci that can be evaluated for methylation using the methods disclosed herein (see, e.g., U.S. Patent Application Publication No. 2020 / 0385813, “Systems and methods for estimating cell source fractions using methylation information,” U.S. Patent Application Publication No. 2020 / 0239965, “Source of origin deconvolution based on methylation fragments in cell-free DNA samples,” U.S. Patent Application Publication No. 2019 / 0287652, “Anomalous fragment detection and classification” (methylation markers indicative of a disease state), U.S. Patent Application Publication No. 2019 / 0316209, “Multi-assay prediction model for cancer detection,” U.S. Patent Application Publication No. 2019 / 0390257 A1, “Tissue-specific methylation marker,” International Publication No. 2011 / 070441, “Categorization of DNA samples,” International Publication No. 2011 / 101728, “Identification of source of DNA samples,” International Publication No. 2020 / 188561, “Methods and systems for detecting methylation changes in DNA samples”).

[0055] In some embodiments, the invention includes a method of detecting tissue-specific DNA methylation patterns using the methylation detection methods disclosed herein. In one aspect of this embodiment, the method may further include identifying the origin tissue of methylated DNA present in the sample. In some embodiments, the method further includes identifying the origin tissue of cell-free DNA isolated from blood. In another aspect of this embodiment, the invention includes detecting organ failure or organ injury, including organ transplant rejection, in a transplant recipient using the methylation pattern of cell-free DNA. The invention includes detecting circulating cell-free DNA having an organ-specific methylation pattern, the presence of such cell-free DNA indicating organ transplant rejection. In some embodiments, the invention includes periodically sampling circulating cell-free DNA and measuring changes in the levels of cell-free DNA by an organ-specific methylation pattern to monitor transplant rejection, an increase in the levels of such cell-free DNA indicating organ transplant rejection.

[0056] In some embodiments, the invention includes a method of diagnosing or screening for the presence of a cancerous tumor in a patient or subject. In some embodiments, the invention includes detecting a tumor using the methylation pattern of cell-free DNA using the methylation detection methods disclosed herein. In some embodiments, the invention includes detecting a tumor derived from a specific tissue or organ by detecting circulating cell-free DNA using a tissue- or organ-specific methylation pattern detected using the methylation detection methods disclosed herein, the presence of such cell-free DNA indicating the presence of a tumor derived from that tissue or organ. In some embodiments, the invention includes periodically sampling circulating cell-free DNA and measuring changes in the levels of cell-free DNA using a tumor-specific methylation pattern to monitor tumor growth or shrinkage, an increase in the levels of such cell-free DNA indicating tumor growth, while a decrease in the levels of such cell-free DNA indicating tumor shrinkage.

[0057] In some embodiments, the present invention includes a method of monitoring the effectiveness of cancer treatment in a patient or subject. In some embodiments, the present invention includes detecting tumor dynamics correlated with treatment using methylation patterns of cell-free DNA detected using the methylation detection methods disclosed herein. In some embodiments, the present invention includes periodically sampling circulating cell-free DNA and measuring changes in the levels of cell-free DNA using tissue- or organ-specific methylation patterns to detect the effect of treatment on tumors derived from a particular tissue or organ, an increase in the levels of such cell-free DNA indicating tumor growth and treatment ineffectiveness, while a decrease in the levels of such cell-free DNA indicating tumor shrinkage and treatment effectiveness, and a stable level of such cell-free DNA indicating stable disease and treatment effectiveness.

[0058] In some embodiments, the present invention includes a diagnostic method or minimal residual disease (MRD) in a cancer patient after treatment. The National Cancer Institute defines MRD as very few cancer cells remaining in the body during or after treatment when the patient has no signs or symptoms of the disease. In some embodiments, the present invention includes a method of detecting MRD using methylation patterns of cell-free DNA detected using the methylation detection methods disclosed herein. In some embodiments, the present invention includes detecting MRD from tumors derived from a particular tissue or organ by detecting circulating cell-free DNA with tissue- or organ-specific methylation patterns, the presence of such cell-free DNA indicating the presence of MRD from the tumor.

[0059] In some embodiments, the present invention includes a method for diagnosing or screening for the presence or condition of an autoimmune disease in a patient or subject. In some embodiments, the present invention includes the detection of an autoimmune disease using a methylation pattern of cell-free DNA detected using the methylation detection methods disclosed herein. In some embodiments, the present invention includes detecting an autoimmune disease characterized by damage to a particular tissue or organ by detecting circulating cell-free DNA in a tissue- or organ-specific methylation pattern, the presence of such cell-free DNA indicating organ damage resulting from an autoimmune disease and the presence of the autoimmune disease. In some embodiments, the present invention includes monitoring for recurrence or remission of an autoimmune disease by periodically sampling circulating cell-free DNA and measuring changes in the levels of cell-free DNA using a tissue- or organ-specific methylation pattern, an increase in the levels of such cell-free DNA indicating an increase in organ damage and recurrence of the autoimmune disease, while a decrease in the levels of such cell-free DNA indicating a decrease in organ damage and remission of the autoimmune disease.

Claims

**Claim 1** A method for reducing the C5-C6 double bond of cytosine in a nucleic acid, the method comprising contacting the cytosine in the nucleic acid with an enreductase. **Claim 2** The method according to claim 1, wherein the cytosine is selected from the group consisting of 5-formylcytosine (5fC), 5-carboxycytosine (5caC), 5,6-dihydro-fC, 5,6-dihydro-caC, 2'-deoxy-5-formylcytidine, 2'-deoxy-5-carboxycytidine, 5-formyl-2'-deoxycytidine, 5-formyl-dC, 5-carboxy-2'-deoxycytidine, and 5-carboxy-dC. **Claim 3** The method according to claim 1, wherein contacting the cytosine in the nucleic acid with an enreductase reduces the cytosine to 5,6-dihydro-U (DHU). **Claim 4** The method according to claim 3, wherein the DHU is subsequently converted to thymine. **Claim 5** The method according to claim 1, wherein the reduction of the C5-C6 double bond of the cytosine in the nucleic acid by contacting the cytosine in the nucleic acid with an enreductase is part of a method for detecting epigenetic cytosine modifications. **Claim 6** A method for detecting epigenetic cytosine modifications, the method comprising at least the step of reducing the C5-C6 double bond of cytosine in a nucleic acid by contacting the cytosine in the nucleic acid with an enreductase. **Claim 7** The method according to claim 6, wherein the cytosine is selected from the group consisting of 5-formylcytosine (5fC), 5-carboxycytosine (5caC), 5,6-dihydro-fC, 5,6-dihydro-caC, 2'-deoxy-5-formylcytidine, 2'-deoxy-5-carboxycytidine, 5-formyl-2'-deoxycytidine, 5-formyl-dC, 5-carboxy-2'-deoxycytidine, and 5-carboxy-dC. **Claim 8** The method according to claim 6, wherein contacting the cytosine in the nucleic acid with an enreductase reduces the cytosine to 5,6-dihydro-U (DHU). **Claim 9** The method according to claim 8, wherein the DHU is subsequently converted to thymine.