Enzymatic conversion of methylated nucleic acids for sequencing

By repairing and transforming DNA nicks using a multi-enzyme transformation method, the problems of low molecular recovery rate and methylation signal bias in existing methylation sequencing methods are solved, enabling efficient methylation detection in applications with low sample input.

CN121844059APending Publication Date: 2026-04-10HE SEQUENCING SOLUTIONS +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methylation sequencing methods, such as bisulfite treatment and enzymatic methylation sequencing (EM-seq), suffer from low molecular recovery rates and methylation signal bias in low-sample-input applications, especially in cell-free DNA samples where methylated cytosine is heavily biased toward the 3' end.

Method used

A multi-enzyme transformation method was adopted, including using enzymes such as Taq ligase, TET methylcytosine dioxygenase 2, T4-phage β-glucosyltransferase and APOBEC3A to contact double-stranded DNA, repairing the nick and transforming methylcytosine and cytosine into sequenceable modified forms, and then ligating them through sequencing adaptors to reduce methylation signal bias.

Benefits of technology

It improves the recovery rate of methylated molecules and the uniformity of sequencing coverage in low sample input applications, reduces methylation signal bias, and improves the accuracy of sequencing results, especially in cell-free DNA samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure relates generally to the enzymatic conversion of methylated nucleic acids in order to differentiate methylated cytosine from unmethylated cytosine in DNA, and more particularly to improved methods and compositions for enzymatic methylation sequencing. In one aspect, various compositions and methods for improving the recovery of methylation signals are provided. The methods include one or more of a nick repair step, methylation signal restoration, use of a modified methylcytosine nucleic acid adapter, and use of a helicase, an ssDNA binding protein, an engineered DNA ligase, or a combination thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This disclosure claims the benefit of U.S. Provisional Patent Application No. 63 / 466,190, filed May 12, 2023, the disclosure of which is incorporated herein by reference in its entirety.

[0003] Statement as to Federally Sponsored Research

[0004] Not Applicable. BACKGROUND

[0005] The present disclosure relates generally to enzymatic conversion of methylated nucleic acids in order to distinguish between methylated cytosines and unmethylated cytosines in DNA, and more particularly to improved methods and compositions for enzymatic methylation sequencing.

[0006] Methylation of DNA plays an important role in various biological processes, including, for example, regulation of gene expression, organismal development, X-chromosome inactivation, and genetic imprinting in vertebrates, as one of the major epigenetic mechanisms. Detecting changes in methylation status and methylation patterns in DNA is of great importance for many clinical applications, including examination of circulating tumor DNA for analysis of tissue of origin and disease state. Historically, nucleobase level detection of modified cytosines in DNA was achieved by deaminating unmodified / unmethylated cytosines via bisulfite treatment followed by amplification via polymerase chain reaction (PCR). Bisulfite treatment reactions are very efficient; however, the method has several drawbacks, including i) the need for large amounts of input DNA due to the harsh nature of the chemical treatment, and ii) the need for long incubation times in order to achieve full conversion of unmethylated cytosines. For liquid biopsy applications that are limited by the amount of sample DNA available for testing, the bisulfite treatment method is hindered by low unique molecular recovery and significant loss of useful information about the methylation status of the input DNA.

[0007] A recently developed enzymatic version of cytosine deamination, called “enzymatic methylation sequencing” or “EM-seq,” can overcome some of the shortcomings of DNA chemical degradation in bisulfite treatment (see Vaisvila et al., Genome Res. 2021. 31: 1280-1289). The method includes tet methylcytosine dioxygenase 2 (TET2 or TET) oxidation and T4-bacteriophage b-glucosyltransferase (T4-bGT) glucosylation to prevent 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) deamination. Subsequently, apolipoprotein B mRNA-editing enzyme catalytic subunit 3A (APOBEC3A or A3A) deaminates all non-modified cytosines to uracils. This enzymatic version of cytosine conversion for methylation detection by sequencing has been shown to recover more unique molecules and provide better coverage uniformity across the input DNA. However, EM-seq can not be suitable for low sample input applications (i.e., nanogram to picogram input DNA) due to inefficient unique molecule recovery and deamination reaction.

[0008] Other challenges observed in methylation sequencing workflows, including both EM-seq and bisulfite sequencing, are related to library preparation. Specifically, current methylation sequencing workflows include one or more library preparation steps that prepare the sample for sequencing and one or more conversion steps that enable nucleobase resolution methylation detection. Generally, the first step of library preparation is to repair and blunt end the termini of double-stranded DNA (dsDNA) to allow efficient ligation of nucleic acid adapters for sequencing. In this step, if nicks are present on the dsDNA, a nick translation process influenced by the repair enzymes present in this step can eliminate a significant amount of methylation signal and cause an unmethylated cytosine bias towards the 3’ end of the molecule. This methylation bias or M-bias is widely observed in samples such as cell-free DNA (cfDNA), where the dsDNA molecules contain multiple nicks, for example, due to the presence of endonucleases or oxidation processes in the plasma.

[0009] For at least the reasons described above, there is a need for improved methods for distinguishing between methylated and unmethylated cytosines in DNA samples for various applications. SUMMARY

[0010] The present invention overcomes the aforementioned disadvantages by providing methods for enzymatic conversion of methylated nucleic acids for sequencing.

[0011] According to one embodiment of the present disclosure, a method includes contacting a first enzyme having ligase activity with double-stranded DNA (dsDNA) comprising at least one nick, at least one cytosine, and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, thereby repairing at least one nick, converting at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base, and converting at least one cytosine into uracil.

[0012] In one aspect, the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

[0013] On the other hand, a second enzyme is used to convert at least one methylcytosine into a modified methylcytosine.

[0014] On the other hand, the second enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

[0015] On the other hand, at least one cytosine is converted into uracil using a third enzyme.

[0016] On the other hand, the third enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0017] On the other hand, the first enzyme is Taq ligase.

[0018] On the other hand, the method further includes ligating a sequencing adaptor to dsDNA.

[0019] In another aspect, the method further includes contacting the nick-repaired dsDNA with a fourth enzyme containing at least one of DNA end repair activity and A-tailing activity.

[0020] On the other hand, the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2, and the β-glucosyltransferase is T4-phage β-glucosyltransferase.

[0021] According to another embodiment of this disclosure, a method includes contacting a Taq ligase with a double-stranded DNA (dsDNA) containing at least one nick, at least one cytosine, and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine to repair at least one nick; enzymatically converting at least one methylcytosine to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET); enzymatically converting 5-hydroxymethylcytosine to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT); and enzymatically converting at least one cytosine to uracil using apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0022] In one aspect, the method further includes ligating a sequencing adaptor to dsDNA.

[0023] In another aspect, the method further includes contacting the nick-repaired dsDNA with an enzyme containing at least one of DNA end repair activity and A-tailing activity.

[0024] On the other hand, the method further includes contacting the dsDNA with a purine-free / pyrimidine-free endonuclease.

[0025] According to another embodiment of this disclosure, a method includes contacting a first enzyme having methyltransferase activity with a double-stranded DNA (dsDNA) containing at least one hemimethylated CpG site, at least one additional cytosine, and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, wherein the at least one additional cytosine remains unmethylated, converting the at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and converting the at least one additional cytosine into uracil.

[0026] In one aspect, the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

[0027] On the other hand, a second enzyme is used to convert at least one methylcytosine into a modified methylcytosine.

[0028] On the other hand, the second enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

[0029] On the other hand, at least one additional cytosine is converted into uracil using a third enzyme.

[0030] On the other hand, the third enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0031] On the other hand, the first enzyme is human DNA (cytosine-5) methyltransferase.

[0032] On the other hand, the method further includes ligating a sequencing adaptor to dsDNA.

[0033] In another aspect, the method further includes contacting the dsDNA with a fourth enzyme containing at least one of DNA end repair activity and A-tailing activity.

[0034] On the other hand, the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2, and the β-glucosyltransferase is T4-phage β-glucosyltransferase.

[0035] According to another embodiment of this disclosure, a method includes linking an adaptor to a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, the adaptor comprising a nucleic acid sequence comprising at least one methylcytosine, wherein each methylcytosine present in the adaptor is 5-hydroxymethylcytosine; converting at least one methylcytosine present in the dsDNA and at least one methylcytosine present in the adaptor into a modified methylcytosine capable of pairing with a guanine base; and converting at least one additional cytosine into uracil.

[0036] In one aspect, the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

[0037] On the other hand, at least one methylcytosine is converted into modified methylcytosine using a first enzyme.

[0038] On the other hand, the first enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

[0039] On the other hand, a second enzyme is used to convert at least one cytosine into uracil.

[0040] On the other hand, the second enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0041] In another aspect, the method further includes contacting the dsDNA with a third enzyme containing at least one of DNA end repair activity and A-tailing activity.

[0042] On the other hand, the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2, and the β-glucosyltransferase is T4-phage β-glucosyltransferase.

[0043] According to another embodiment of this disclosure, a method includes linking an adaptor to a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, the adaptor comprising a nucleic acid sequence comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adaptor is 5-(β-glucosyloxymethyl)cytosine, converting at least one methylcytosine present in the dsDNA into a modified methylcytosine capable of pairing with a guanine base, and converting at least one additional cytosine into uracil.

[0044] In one respect, the modified methylcytosine present in dsDNA is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

[0045] On the other hand, at least one methylcytosine in the dsDNA is converted into a modified methylcytosine using a first enzyme.

[0046] On the other hand, the first enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

[0047] On the other hand, a second enzyme is used to convert at least one cytosine into uracil.

[0048] On the other hand, the second enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0049] In another aspect, the method further includes contacting the dsDNA with a third enzyme containing at least one of DNA end repair activity and A-tailing activity.

[0050] On the other hand, the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2, and the β-glucosyltransferase is T4-phage β-glucosyltransferase.

[0051] According to another embodiment of this disclosure, the adapter for methyl sequencing comprises a first nucleic acid having a 3' end complementary to the 5' end of a second nucleic acid to form a complementary region, the complementary region comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adapter is 5-(β-glucosyloxymethyl)cytosine.

[0052] According to another embodiment of this disclosure, the adapter for methyl sequencing comprises a first nucleic acid having a 3' end complementary to the 5' end of a second nucleic acid to form a complementary region, the complementary region comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adapter is 5-hydroxymethylcytosine.

[0053] According to another embodiment of this disclosure, a method includes: combining the following: i) a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; ii) a cytidine deaminase; and iii) at least one of a single-stranded DNA binding protein and a helicase; converting at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and converting at least one cytosine into uracil using a cytidine deaminase.

[0054] In one aspect, the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

[0055] On the other hand, at least one methylcytosine is converted into modified methylcytosine using at least one of methylcytosine dioxygenase and β-glucosyltransferase.

[0056] On the other hand, cytidine deaminase is the catalytic subunit 3A (APOBEC3A) of apolipoprotein B mRNA editing enzyme.

[0057] On the other hand, the method further includes ligating a sequencing adaptor to dsDNA.

[0058] In another aspect, the method further includes contacting the dsDNA with an enzyme having at least one of DNA end repair activity and A-tailing activity.

[0059] On the other hand, the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2, and the β-glucosyltransferase is T4-phage β-glucosyltransferase.

[0060] According to another embodiment of this disclosure, a method includes: combining the following: i) a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; ii) a cytidine deaminase; and iii) at least one of a single-stranded DNA binding protein and a helicase; enzymatically converting at least one methylcytosine to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET); enzymatically converting 5-hydroxymethylcytosine to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT); and enzymatically converting at least one cytosine to uracil using apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

[0061] In one aspect, the method further includes ligating a sequencing adaptor to dsDNA.

[0062] In another aspect, the method further includes contacting the dsDNA with an enzyme having at least one of DNA end repair activity and A-tailing activity.

[0063] According to another embodiment of this disclosure, the composition comprises double-stranded DNA (dsDNA) containing at least one cytosine and at least one modified cytosine selected from 5-carboxycytosine and 5-(β-glucosyloxymethyl)cytosine, a buffer, and cytidine deaminase, wherein the amount of dsDNA is from about 1 picogram to about 10 ng, wherein the concentration of cytidine deaminase is from about 0.05 uM to about 0.5 uM, and wherein the total volume of the composition is less than about 100 uL.

[0064] According to another embodiment of this disclosure, a method includes providing a first composition comprising double-stranded DNA (dsDNA), the double-stranded DNA comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; enzymatically converting the at least one methylcytosine to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET) to form a first dsDNA product; enzymatically converting the 5-hydroxymethylcytosine in the first dsDNA product to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT) to form a second dsDNA product; and enzymatically converting at least one cytosine in the second dsDNA product to uracil using apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A) to form a third dsDNA comprising APOBEC3A and a third dsDNA. The product is a methyl sequencing product composition, which is combined with an amplification composition, and a third dsDNA product is amplified by polymerase chain reaction, wherein the amplification occurs in the presence of APOBEC3A.

[0065] In one respect, each of the above compositions and methods may further include an engineered DNA ligase for ligating the adaptor to dsDNA.

[0066] According to another embodiment of this disclosure, a method includes combining the following: a first enzyme having ligase activity, a second enzyme having 5' to 3' polymerase activity, a second enzyme lacking 5' to 3' exonuclease activity, a third enzyme having purine / pyrimidine endonuclease activity, a double-stranded DNA (dsDNA) comprising at least one single-stranded discontinuity selected from nicks and gaps, at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, and a plurality of deoxynucleoside triphosphates (dNTPs). The method further includes repairing at least one discontinuity in the dsDNA, converting at least one methylcytosine to a modified methylcytosine capable of pairing with a guanine base, and converting at least one cytosine to uracil.

[0067] The foregoing and other aspects and advantages of the invention will become apparent from the following description. In this specification, reference is made to the accompanying drawings, which form a part thereof, and preferred embodiments of the invention are illustrated by way of example in the drawings. Such embodiments do not necessarily represent the full scope of the invention; however, the scope of the invention is therefore defined by the claims herein. Attached Figure Description

[0068] Figure 1 is an example of a first method 100 and a second method 200 for detecting DNA methylation by enzyme-catalyzed methylation sequencing according to the present disclosure.

[0069] Figure 2a is a schematic diagram of the process for enzymatically converting methylated DNA for sequencing. Above the dashed lines, multiple enzymatic methylation sequencing reactions for converting unmethylated and methylated cytosine bases are illustrated. Arrows indicate enzymatic conversions that transform one molecule (shown as a rounded rectangle) into another molecule using the enzyme listed adjacent to the arrow. An "X" on the arrow indicates that the depicted enzyme exhibits low or no activity for the conversion of the indicated molecule. Below the dashed lines, as indicated by the dashed arrows, are shown the nucleobases of the corresponding products of the EM-seq conversion reactions detected by sequencing (e.g., uracil is ultimately detected as thymine by sequencing).

[0070] Figure 2b is an illustration of the selection of methylated and unmethylated nucleobases from Figure 2a. From left to right, the illustrated molecules include cytosine, 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), and 5-(β-glucosyloxymethyl)cytosine (5gmC).

[0071] Figure 3 is a bar chart illustrating the effect of ligase selection (engineered DNA ligase versus wild-type T4 DNA ligase) on the mean CpG modification rate (“initial”) of human cfDNA with unmethylated λ-spiked (“λ”) in the control and methylated pUC19 plasmid DNA-spiked (“pUC19”) in the control. Data from three different cfDNA samples are presented in the format “[sample number] - [replication number]”, with two replicates for each sample.

[0072] Figure 4 is a bar chart illustrating the effect of ligase selection (engineered DNA ligase versus wild-type T4 DNA ligase) on reducing replication rate as a measure of the recovery of unique molecules from human cfDNA samples. Data from three different cfDNA samples, each with two replicates, are shown in the format "[sample number] - [replication number]".

[0073] Figure 5 is a graph illustrating the effect of adding Taq ligase before end repair and A-tailing on the methylation signal at the 3' end of the sample DNA molecule as determined by nucleic acid sequencing.

[0074] Figure 6 is a bar graph illustrating the transformation efficiency of the EM-seq reaction of DNA samples linked to the adaptor, where each modified cytosine in the adaptor is 5 mC (λmC) or 5 hmC (initial hmC).

[0075] Figure 7 is a pair of graphs illustrating the effect of the addition of Taq DNA ligase and methyltransferase DNMT5 from Cryptococcus neoformans on the methylation signal at the 3' end of the sample DNA molecule, as determined by nucleic acid sequencing, before end repair and A-tailing. Detailed Implementation

[0076] I. Definition

[0077] In this application, unless the context clearly indicates otherwise, (i) the term “a” (a) may be understood to mean “at least one”; (ii) the term “or” may be understood to mean “and / or”; (iii) the terms “comprising” and “including” may be understood to encompass the listed components or steps, whether they are presented individually or together with one or more other components or steps; (iv) the terms “about” and “approximately” may be understood to allow for standard variations, as would be understood by one of ordinary skill in the art; and (v) the scope provided includes endpoints.

[0078] Approximately: As used herein, the term “approximately” or “about”, when applied to one or more values ​​of interest, refers to a value similar to the stated reference value. In some embodiments, the term “approximately” or “about” refers to a range of values ​​within or below 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% in any direction (greater or less) of the reference value, unless otherwise stated or apparent from the context (unless the number exceeds 100% of the possible value).

[0079] Related: Two events or entities are "related" to each other, as used herein, if the presence, level, and / or form of one event or entity is related to the presence, level, and / or form of another event or entity. For example, if the presence, level, and / or form of a particular entity (e.g., a polypeptide, genetic marker, metabolite, etc.) is related to the incidence and / or susceptibility to that disease, condition, or symptom, then that particular entity is considered related to that disease, condition, or symptom (e.g., in a relevant population). In some embodiments, two or more entities are physically "related" to each other if they interact directly or indirectly, such that they are physically close to each other and / or remain physically close to each other. In some embodiments, two or more entities that are physically bound to each other are covalently linked; in some embodiments, two or more entities that are physically bound to each other are not covalently linked, but are non-covalently linked, for example, through hydrogen bonds, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.

[0080] Biological Samples: As used herein, the term "biological sample" generally refers to a sample obtained or derived from a target biological source (e.g., a tissue or organism or cell culture), as described herein. In some embodiments, the target source includes an organism, such as an animal or a human, or is composed of such an organism. In some embodiments, a biological sample includes or is composed of biological tissues or fluids. In some embodiments, a biological sample may be or include bone marrow; blood; blood cells; ascites; tissue or fine-needle biopsy samples; cell-containing body fluids; free floating nucleic acids; sputum; saliva; urine; cerebrospinal fluid; peritoneal fluid; pleural fluid; feces; lymph; gynecological fluids; skin swabs; vaginal swabs; oral swabs; nasal swabs; washing or lavage fluids, such as catheter lavage fluid or bronchoalveolar lavage fluid; aspirates; scrapings; bone marrow specimens; tissue biopsy specimens; surgical specimens; other body fluids, secretions and / or excretions; and / or cells derived therefrom, etc. In some embodiments, a biological sample comprises or is composed of cells obtained from an individual. In some embodiments, the obtained cells are or include cells from the individual from whom the sample was obtained. In some embodiments, the sample is a “primary sample” obtained directly from the target source by any suitable means. For example, in some embodiments, the original biological sample is obtained by a method selected from the group consisting of: biopsy (e.g., fine-needle aspiration or tissue biopsy), surgery, body fluid collection (e.g., blood, lymph, feces, etc.). In some embodiments, as will be clearly apparent from the context, the term “sample” refers to a preparation obtained by processing the original sample (e.g., by removing one or more of its components and / or by adding one or more agents to it). For example, using semi-permeable membrane filtration. Such “processed samples” may include, for example, nucleic acids or proteins extracted from the sample, or nucleic acids or proteins obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, separation and / or purification of certain components.

[0081] The description herein of a composition or method “comprising” one or more named elements or steps is open-ended, meaning that the named element or step is essential, but other elements or steps may be added within the scope of the composition or method. It should be understood that a composition or method described as “comprising” (or “comprises”) one or more named elements or steps also describes a corresponding, more limited composition or method “consisting essentially of” (or “consists essentially of”) of the same named element or step, meaning that the composition or method includes the essential named element or step and may also include additional elements or steps that do not substantially affect the essential and novel characteristics of the composition or method. It should also be understood that any composition or method described herein as “comprising” or “essentially of” one or more named elements or steps also describes a corresponding, more limited, closed composition or method “consisting of” (or “consists of”) of named elements or steps to exclude any other unnamed elements or steps. In any composition or method disclosed herein, any known or disclosed equivalent of any essential named element or step may be substituted for that element or step.

[0082] Design: As used herein, the term “design” refers to a pharmaceutical agent that: (i) has a structure selected artificially; (ii) is produced by a method requiring artificial intervention; and / or (iii) is different from natural substances and other known pharmaceutical agents.

[0083] Determination: As will be understood by those of ordinary skill in the art who read this specification, "determination" can be accomplished using or by any of a variety of techniques available to those skilled in the art, including, for example, specific techniques explicitly mentioned herein. In some embodiments, determination involves manipulation of a physical sample. In some embodiments, determination involves consideration and / or manipulation of data or information, such as using a computer or other processing unit suitable for performing relevant analyses. In some embodiments, determination involves receiving relevant information and / or materials from a source. In some embodiments, determination involves comparing one or more features of a sample or entity with a comparable reference.

[0084] Identity: As used herein, the term "identity" refers to the overall relevance between polymer molecules, such as nucleic acid molecules (e.g., DNA and / or RNA molecules) and / or polypeptide molecules. In some embodiments, polymer molecules are considered "substantially identical" to each other if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. For example, for optimal comparison purposes, the percentage of identity between two nucleic acid or polypeptide sequences can be calculated by aligning the two sequences (e.g., gaps can be introduced in one or both of the first and second sequences to achieve optimal alignment, while dissimilar sequences can be ignored for comparison purposes). In some embodiments, the length of the sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or substantially 100% of the length of the reference sequence. Nucleotides at the corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., a nucleotide or amino acid) as the corresponding position in the second sequence, the molecules are identical at that position. The percentage of identity between two sequences is a function of the number of shared positions, taking into account the number and length of each gap, and this needs to be introduced to achieve optimal alignment of the two sequences. Sequence comparison and determination of the percentage of identity between two sequences can be accomplished using mathematical algorithms. For example, the algorithm of Meyers and Miller (CABIOS, 1989, 4: 11-17) can be used to determine the percentage of identity between two nucleotide sequences, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons performed using the ALIGN program use a PAM120 weighted residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the NWSgapdna.CMP matrix can be used, and the GAP program in the GCG software package can be used to determine the percentage of identity between two nucleotide sequences.

[0085] Sample: As used herein, the term "sample" refers to a substance or a substance containing a target composition for qualitative and / or quantitative evaluation. In some embodiments, the sample is a biological sample (i.e., derived from a biological source, such as a cell or organism). In some embodiments, the sample originates from geological, aquatic, astronomical, or agricultural sources. In some embodiments, the target source includes organisms, such as animals or humans, or is composed of organisms. In some embodiments, the sample for forensic analysis is or includes biological tissue, biological fluids, organic or inorganic substances, such as, for example, clothing, dirt, plastics, water. In some embodiments, agricultural samples include or consist of organic matter, such as leaves, petals, bark, wood, seeds, plants, fruits, etc.

[0086] Specificity: As used in this article, the term “specificity” refers to an enzyme’s preference for a particular substrate.

[0087] Selectivity: As used in this article, the term “selectivity” refers to an enzyme’s preference for one particular substrate over another.

[0088] Essentially: As used herein, the term “essentially” refers to qualitative conditions that exhibit all or nearly all of the range or degree of the target characteristic or property. Those skilled in the art of biology will understand that biological and chemical phenomena rarely (if at all) complete and / or continue to complete or achieve or avoid absolute results. Therefore, this article uses the term “essentially” to capture the inherent potential incompleteness in many biological and chemical phenomena.

[0089] Synthetic: As used in this article, the term “synthetic” means artificially produced and therefore produced in a form that does not exist in nature, either because it has a structure that does not exist in nature, or because it is combined with one or more other components (which are not combined with in nature), or not combined with one or more other components (which are combined with in nature).

[0090] Methylation: As used herein, the term "methylation" refers to a cytosine having a methyl or hydroxymethyl group at the C-5 position of the cytosine ring in DNA. Methylated DNA is DNA having at least one methylated cytosine. Methylated cytosines in DNA can naturally include 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC); however, 5mC can be present in methylated DNA in significantly greater abundance than 5hmC, and 5hmC can be present in methylated DNA in significantly greater abundance than 5fC and 5caC.

[0091] Unmethylated: As used herein, the terms “unmethylated” or “non-methylated” refer to cytosine lacking a methyl or hydroxymethyl group at the C-5 position of the cytosine ring in DNA. Unmethylated DNA is DNA lacking methylated cytosine.

[0092] Methylation signal: As used herein, the term "methylation signal" means a signal that serves as a representative measurement of cytosine methylation. In embodiments of this disclosure, the methylation signal may include a signal designated as cytosine, determined by nucleic acid sequencing.

[0093] End repair: As used herein, the term "end repair" refers to a method for repairing DNA (e.g., fragmented or damaged DNA, or DNA molecules incompatible with other DNA molecules). In some embodiments, the process involves two functions: 1) converting double-stranded DNA with single-stranded protrusions into double-stranded DNA without single-stranded protrusions by means of an enzyme (such as T4 DNA polymerase and / or Klenow fragment); and 2) adding phosphate ester groups to the 5' end of (single-stranded or double-stranded) DNA by means of an enzyme (such as polynucleotide kinase).

[0094] A-tailing: As used herein, the term "A-tailing" refers to the addition of a single deoxyadenosine residue to the end of a blunt-ended double-stranded DNA fragment to form a 3' deoxyadenosine single-base single-stranded overhang. A-tailed fragments are not suitable for self-ligation (i.e., self-circularization and tandemization of DNA), but they are compatible with 3' deoxythymidine single-stranded overhangs, such as those present on adaptors.

[0095] Conversion rate: As used in this article, “conversion” refers to the enzymatic conversion (or bioconversion) of the substrate to the corresponding product. “Percent conversion” refers to the percentage of substrate that is converted to the product under specified conditions over a period of time. Therefore, the “enzyme activity” or “activity” of an enzyme can be expressed as the “percent conversion” of the substrate to the product over a specified time period.

[0096] II. Detailed Description of Certain Embodiments

[0097] As stated above, in various situations, it may be useful to provide methods for converting methylated nucleic acids to distinguish between methylated and unmethylated cytosine in DNA, and more specifically, to provide improved methods and compositions for methylation sequencing, including EM-seq and bisulfite sequencing. On one hand, methylation sequencing may not be suitable for low-sample-input applications (i.e., nanograms to picograms of input DNA) due to the inefficiency of unique molecule recovery and deamination reactions. On the other hand, current methylation sequencing methods may be limited by methylation bias in samples such as cfDNA due to DNA damage caused by endonucleases, oxidation processes, etc.

[0098] These and other challenges can be overcome by compositions and methods according to this disclosure for the enzymatic conversion of methylated nucleic acids for sequencing. In one aspect, this disclosure provides improved recovery of methylation signals (i.e., accurate detection of methylated and unmethylated cytosine by sequencing). Improvements are achieved, at least in part, by various methods including repairing nicks present in sample DNA prior to the library preparation step, restoring methylation signals in the case of hemimethylated dsDNA, using nucleic acid adaptors comprising or composed of methylcytosine selected from 5mC, 5hmC, and 5-(β-glucosyloxymethyl)cytosine (5gmC), using helicases, ssDNA binding proteins, or combinations thereof, and improved compositions for library preparation.

[0099] This disclosure is based, at least in part, on the surprising finding that the choice of DNA ligase used for attaching nucleic acid adaptors has a significant impact on the downstream detection of methylation. In one aspect, the selection of the DNA ligase used for adaptor ligation improves conversion efficiency (i.e., the efficiency of target product ligation) even with very low sample input volumes and further improves ligation specificity, resulting in reduced sensitivity to adaptor concentration and a decrease in observed adaptor dimer formation. Furthermore, the selection of a specific DNA ligase reduces the ligation reaction time from 16 hours with standard DNA ligases to approximately 5 minutes with the ligase according to this disclosure.

[0100] Turning to Figure 1, an embodiment of method 100 for enzymatic methylation (EM) sequencing is illustrated. In one aspect, method 100 represents a current method known in the art for detecting methylation in enzymatically converted methylated DNA samples using next-generation sequencing technologies, such as flow cell-based sequencing-by-synthesis methods. Generally, methylation is detected by selectively deaminated unmethylated cytosine to uracil, leaving 5mC and 5hmC (and their derivatives) intact; specifically, the detection of 5mC and 5hmC, which can be naturally present in DNA. Each uracil is converted to thymine by PCR amplification of the product of the selective deamination step, while 5hmC and 5mC (and their derivatives) are converted to cytosine. As in the case of bisulfite sequencing, the amplified DNA retains only methylated cytosine, producing single-nucleotide resolution information about the methylation state of the DNA, which can be interpreted using various known informatics methods.

[0101] According to step 102 of method 100, sample DNA is prepared. The preparation of sample DNA may include collecting DNA or DNA-containing materials, such as whole blood, plasma, serum, tissue (including formalin-fixed paraffin-embedded or FFPE tissue), etc. Step 102 may further include recovering DNA (if applicable) from blood or tissue samples or other DNA-containing samples. In one aspect, the DNA is cfDNA. In another aspect, the DNA is circulating tumor DNA (ctDNA). In another aspect, the sample DNA is double-stranded DNA (dsDNA). In another aspect, the DNA is cleaved or otherwise fragmented to provide a uniform size or desired size distribution relative to the length of the nucleotides in the DNA. In another aspect, the DNA is purified using one or more purification methods to obtain isolated DNA. Step 102 may include any other necessary steps for preparing sample DNA for downstream steps of method 100. In one aspect, the sample DNA includes DNA damage, including but not limited to single-strand cuts, double-strand breaks, hemimethylation (i.e., partial methylation only at CpG sites on one strand of dsDNA), etc. For clarity, CpG or CG sites are defined as DNA regions on a linear sequence of bases in the 5' to 3' orientation, followed by guanine after cytosine. Alternatively, the sample DNA may contain cytosine bases, at least a portion of which are methylated. Furthermore, methylated cytosine bases may include at least one of 5mC and 5hmC.

[0102] In step 104, the DNA sample is treated with one or more enzymes to convert fragmented DNA into repaired DNA with 5' phosphorylated, 3' dA-tailed ends (also known as end repair and A-tailing). Suitable enzymes include DNA polymerases with 3' exonuclease activity, such as Taq DNA polymerase I from *Thermus aquaticus*, T4 DNA polymerase, T4 polynucleotide kinase, and the Klenow fragment. In one aspect, the sample DNA product of step 104 can be purified from any enzyme, salt, buffer, or other reagent used in step 104. In one instance, a suitable combination of enzymes for end repair and A-tailing includes a polynucleotide kinase, a first DNA polymerase, and a second DNA polymerase. In one aspect, the polynucleotide kinase has 5' phosphorylation and 3' phosphorylation removal activity (e.g., T4 polynucleotide kinase). On the other hand, the first DNA polymerase has the activity of nick filling and removing 3' single-stranded protrusions or filling 5' single-stranded protrusions to form blunt ends, but lacks strand substitution and 5' to 3' exonuclease activity (e.g., T4 DNA polymerase). On the other hand, the second DNA polymerase has the activity of adding at least 3' dA, also known as adding an A tail (e.g., Taq DNA polymerase).

[0103] In step 106, a nucleic acid adaptor may be added to enable downstream workflows or other processing steps. In one aspect, an adaptor is selected to enable a downstream sequencing step. In another aspect, an adaptor is selected to enable a downstream amplification step. Typically, adding an adaptor to the repaired DNA from step 104 involves an enzymatic ligation step. Therefore, the adaptor may include a 5' phosphorylated, 3' dA-tailed end compatible with the sample DNA produced in step 104; however, other methods known in the art may also be applied. Example adaptors include Y-shaped adaptors, dumbbell adaptors, blunt-end adaptors, single-stranded overhang adaptors, single-stranded adaptors, etc. The adaptor may further include a sample barcode / identification sequence for distinguishing one sample DNA from another, a unique molecular recognition sequence for distinguishing individual DNA molecules, etc. For EM-seq, bisulfite sequencing, and other methylation detection methods, it may be useful to provide an adaptor in which some or all of the cytosine present in the sample is converted to 5mC or is provided as originally 5mC. On the one hand, the sample DNA product of step 106 can be purified from any enzyme, salt, buffer or other reagent used in step 106.

[0104] In step 108, the repaired DNA, including the 5mC adaptor applied thereto, is treated to convert the 5mC and 5hmC present in the DNA into modified methylcytosine. As defined herein, the modified methylcytosine is a derivative or modified structure of 5mC or 5hmC that is not a substrate for cytidine deaminase. It should be understood that the purpose of step 108 is to convert 5mC and 5hmC into derivatives or modified structures that are not a substrate for the cytidine deaminase selected for step 110. Therefore, deamination of the converted 5mC and 5hmC is prevented in step 110. Furthermore, the converted 5mC and 5hmC are preferably capable of effectively base-pairing with guanine bases during PCR in step 112. As described above, the result of method 100 is to provide amplified DNA that retains only methylated cytosine, generating single-nucleotide resolution information about the methylation state of the DNA for downstream analysis.

[0105] Step 108 may include one or more enzymes to selectively convert 5mC and 5hmC while keeping cytosine intact or otherwise unmodified. Examples of enzymes known for use in EM-seq include methylcytosine dioxygenases and β-glucosyltransferases. One example methylcytosine dioxygenase is TET. One example β-glucosyltransferase is T4-βGT. Reference Figure 2aIn 2b, TET can enzymatically catalyze the formation of 5hmC from 5mC. Furthermore, TET can further act on 5hmC to produce 5-formylcytosine (5fC), which can then be converted to 5-carboxycytosine (5caC) by TET. In one aspect, T4-βGT can enzymatically catalyze the formation of 5gmC from 5hmC. On the other hand, it may be useful to further convert 5hmC or 5fC to the corresponding derivatives using TET, T4-βGT, or another enzyme, as cytidine deaminases with little or no activity towards 5gmC or 5caC can be selected, as will be discussed with reference to step 110. In contrast, as mentioned above, it may be necessary to avoid cytidine deaminases with some activity towards 5hmC and 5fC in order to achieve selective deamination of cytosine only or to distinguish it from methylated cytosine in downstream sequencing applications. In one aspect, step 108 ideally comprises the complete transformation of each 5 mC and 5 hmC present in the sample DNA. However, depending on the selected enzyme, reaction conditions, and other factors, it is reasonably possible to achieve a transformation percentage of less than 100% for each 5 mC and 5 hmC. Preferably, at least 70% transformation is achieved. More preferably, at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or at least 99.5% transformation is achieved. On the other hand, the sample DNA product of step 108 can be purified from any enzyme, salt, buffer, or other reagent used in step 108.

[0106] In step 110, the DNA product from step 108, including converted 5mC and converted 5hmC, is treated with cytidine deaminase to selectively convert unmethylated or otherwise unmodified cytosine bases present in the sample DNA into uracil. As defined herein, cytidine deaminase is any enzyme capable of converting cytosine bases present in nucleic acids into uracil. Therefore, cytidine deaminase can be effectively used to convert cytosine bases in both cytidine (e.g., in RNA) and deoxycytidine (e.g., in DNA).

[0107] As described above, step 110 is performed to chemically / structurally distinguish cytosine bases from methylated cytosine bases. PCR can be used to amplify the resulting product by selectively deamination of only unmethylated cytosine bases to uracil bases. In this case, when only deoxynucleoside triphosphates (dNTPs) adenine, cytosine, guanine, and thymine (i.e., dATP, dCTP, dGTP, and dTTP, respectively), any uracil base will produce a thymine base, while methylated cytosine bases in any of the aforementioned modified forms will produce a cytosine base. In one aspect, the cytidine deaminase is APOBEC3A; however, as those skilled in the art will understand, other cytidine deaminases can be selected. In one aspect, step 108 ideally involves the complete conversion of every unmethylated cytosine base present in the sample DNA. However, depending on the selected enzyme, reaction conditions, and other factors, it is reasonably possible to achieve a conversion percentage of less than 100% for each cytosine base. Preferably, at least 70% conversion is achieved. More preferably, at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or at least 99.5% conversion is achieved. Furthermore, ideally, the selectivity of cytosine deamination to cytosine relative to methylated cytosine is 100%, including any converted methylcytosine (e.g., 5hmC, 5fC, 5gmC, and 5caC). 100% selectivity means that the methylcytosine bases present in the sample DNA are not the substrate of the cytidine deaminase. However, depending on the selected enzyme, reaction conditions, etc., in step 110, the cytidine deaminase may exhibit some activity towards the methylcytosine present in the sample DNA (i.e., the cytidine deaminase exhibits less than 100% selectivity). Preferably, at least 70% selectivity is achieved. More preferably, at least 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or at least 99.5% selectivity is achieved. In one aspect, the above-mentioned selectivity is achieved with respect to at least one methylcytosine structure selected from 5hmC, 5fC, 5gmC, and 5caC. In another aspect, the above-mentioned selectivity is achieved with respect to at least one methylcytosine structure selected from 5gmC and 5caC. In one aspect, the sample DNA product of step 110 can be purified from any enzyme, salt, buffer, or other reagent used in step 110.

[0108] In step 112, the sample DNA generated in step 110 is amplified by PCR. Ideally, the sample DNA generated in step 110 comprises only methylated cytosine bases and excludes unmethylated cytosine bases, since each unmethylated cytosine base is preferably converted to uracil in step 110. On the other hand, the methylated cytosine bases comprise 5 gmC and 5 caC. Step 112 may include PCR amplification using dNTPs composed of dATP, dCTP, dGTP, and dTTP. Primers complementary to the adaptor nucleic acid sequence attached to the sample DNA in step 106 may be selected. The primers may further comprise the sample recognition sequence or unique molecular recognition sequence as described above. On one hand, the sample DNA product of step 112 can be purified from any enzyme, salt, buffer, or other reagent used in step 112. On one hand, the product of step 112 may be referred to as a library. The term "library" can also be reasonably applied to the product of any of the aforementioned steps 104, 106, 108, and 110.

[0109] In step 114, the sequence of the product of PCR amplification step 112 is determined. In one aspect, any suitable DNA sequencing technology can be chosen to determine the nucleic acid sequence of interest. For example, sequencing-by-synthesis can be used. The data collected from sequencing determines the identity of each base in the nucleic acid sequence present in the sequenced library. As mentioned above, each cytosine base detected by sequencing can be interpreted as a methylcytosine base in the original sample DNA sequence – such as 5mC or 5hmC – while each thymine base detected can be interpreted as either thymine or cytosine in the original sample DNA sequence. Using known informatics tools, the allocation of thymine or cytosine can be performed or predicted regarding the sequence identity of the original sample DNA.

[0110] The method 100 described herein is generally useful for EM-seq. However, this disclosure provides several novel improvements and modifications to method 100. Continuing to refer to Figure 1, method 200 illustrates at least some of the disclosed improvements to the EM-seq workflow. In step 202, sample DNA input is prepared. In some embodiments, step 202 may be similar to or identical to step 102 of method 100. In other embodiments, step 202 may be modified to provide sample DNA in the same format as subsequent steps in method 200.

[0111] In step 203 of method 200, the potential loss of methylation signal is mitigated by repairing one or both of the nicks and gaps that may be present in the dsDNA in the sample DNA. Generally, dsDNA contains nicks (i.e., discontinuities in the dsDNA molecule where there is no phosphodiester bond between adjacent nucleotides of one strand of the dsDNA). In contrast, gaps in dsDNA are discontinuities in the dsDNA molecule where one strand of the dsDNA is missing one or more consecutive nucleotides. Nicks and gaps in dsDNA can provide replication initiation sites for DNA polymerase, enabling a process called nick translation. When nick translation occurs, DNA polymerase extends the 3' hydroxyl terminus of the nick or gap site, removes nucleotides through 5' to 3' exonuclease activity, and replaces them with dNTPs. Therefore, the strand of dsDNA located at the 3' of the nick, including any methylated or otherwise modified bases, is not present in the product of nick translation, and the corresponding methylation signal is lost. Step 203 considers repairing the nick, gap, or both to mitigate the potentially detrimental effects of nick translation on methylation detection. It should be understood that not all DNA polymerases possess 5' to 3' exonuclease activity or the ability to otherwise catalyze nick translation. Therefore, step 203 may include one or more DNA polymerases with activities such as nick filling without strand displacement, removing 3' single-stranded overhangs, or filling 5' single-stranded overhangs to form blunt ends, provided that such DNA polymerases lack nick translation activity.

[0112] In some embodiments of step 203, the nick is repaired using an enzyme such as a ligase. In one aspect, the ligase has dsDNA nick repair activity. In one example, the ligase is a DNA ligase derived from a species of the genus *Thermus*, such as the Taq DNA ligase from *Thermus thermophilus* HB8. In one aspect, the ligase with dsDNA nick repair activity catalyzes the formation of a phosphodiester bond between the 5'-phosphate and 3'-hydroxyl groups of two adjacent DNA strands in the nicked dsDNA. Specifically, the ligase with dsDNA nick repair activity should preferentially catalyze bond formation when the strand to be ligated hybridizes with the complementary DNA strand and (nick-free) accurately pairs. In another aspect, the ligase with dsDNA nick repair activity should have limited or no blunt-end ligation activity to avoid ligating dissimilar DNA fragments.

[0113] In some embodiments, in addition to a DNA ligase suitable for nick repair, step 203 further includes treating the damaged dsDNA with a purine-free / pyrimidine-free endonuclease before end repair and A-tailing. One example of a purine-free / pyrimidine-free endonuclease is APEI. APEI catalyzes nick formation in the phosphodiester backbone of dsDNA at base-free sites. When a nick present on the backbone includes a hydroxyl group at its 5' end, there may be too much steric hindrance for polynucleotide kinases to phosphorylate the 5' end, thus preventing DNA ligases such as Taq DNA ligase from repairing the nick. Without being limited by any particular theory, treatment with APEI is expected to effectively remove the unphosphorylated nucleobases present at the 5' end of the nick site, resulting in a single nucleotide gap. This single nucleotide gap can then be filled using T4 DNA polymerase, and the resulting product can subsequently be repaired with Taq DNA ligase. In addition, in the case of gaps of more than one nucleotide, T4 DNA polymerase can be used to fill these larger gaps, and the resulting product is similarly repaired with Taq DNA ligase.

[0114] In step 204 of method 200, the sample DNA is treated with one or more enzymes to convert fragmented DNA into repaired DNA with 5' phosphorylation and a 3' dA tail. Suitable enzymes include DNA polymerases with 3' exonuclease activity, such as Taq DNA polymerase I derived from *Thermophyton aquaticus*, T4 DNA polymerase, T4 polynucleotide kinase, and the Klenow fragment.

[0115] In some embodiments, step 204 includes using a methyltransferase to restore the methylation signal. In one embodiment, the methyltransferase is DNMT1. Cytosine bases at CpG sites in human DNA are methylated on both strands (i.e., symmetrical methylation). During replication, methyltransferases such as DNMT1 can restore a specific methylation pattern consistent with the parental DNA on the daughter strand by methylating hemimethylated DNA. The methylation specificity of DNMT1 can be utilized to restore the methylation signal. In one aspect, the methylation signal may be lost due to damage to the sample DNA. In another aspect, the methylation signal may be lost during nick translation during step 204. In some embodiments, the DNMT1 enzyme may be added in combination with a composition for end repair and A-tailing. In some embodiments, the DNMT1 enzyme may be added after end repair and A-tailing are completed and before ligation of the nucleic acid adaptor in step 206. Therefore, step 204 may include two consecutive sub-steps (i.e., the end repair and A-tailing sub-step, followed by the methylation re-atom step). In some embodiments, the sample DNA product of step 204 can be purified from any enzyme, salt, buffer, or other reagent used in step 204. Furthermore, in some embodiments, the sample DNA product of one or both sub-steps of step 204 can be purified from any enzyme, salt, or other reagent used in the respective sub-steps.

[0116] In step 206, a nucleic acid adaptor may be added to enable a downstream sequencing workflow. Typically, adding an adaptor to the repaired sample DNA from step 204 involves an enzymatic ligation step. Therefore, the adaptor may include a 5' phosphorylated, 3' dA-tailed end compatible with the sample DNA; however, other methods known in the art may also be applied. Example adaptors include Y-shaped adaptors, dumbbell adaptors, blunt-end adaptors, single-stranded overhang adaptors, etc. The adaptor may further include a sample barcode or identification sequence for distinguishing one sample DNA from another, a unique molecular recognition sequence for distinguishing individual DNA molecules, etc. For EM-seq, bisulfite sequencing, and other methylation detection methods, it may be useful to provide an adaptor in which some or all of the cytosines present in the sample are converted to 5mC or are provided as originally 5mC. According to method 200 of this disclosure, it may be useful to provide an adaptor in which some or all of the cytosine bases present in the sample are converted from 5mC to their derivatives.

[0117] In one respect, different methylcytosine dioxygenases have been shown to exhibit sequence background bias unique to the enzyme. Therefore, oxidation efficiency can vary depending on the sequence identity of the input DNA. Consequently, methylcytosine dioxygenases exhibiting lower oxidation efficiency for their sequences will have less protection against deamination, for example, when treated with cytidine deaminases. To mitigate the potential sequence dependence of a given methylcytosine dioxygenase, sequencing adaptors can be provided in which each 5mC is converted to 5hmC or otherwise substituted with 5hmC. Adaptors in which each methylcytosine is 5hmC can allow for better protection against deamination and result in higher yields (e.g., due to improved consistency of the adaptor sequence) and better aggregation on flow cell-based sequencing instruments. Therefore, step 206 can include using an adaptor in which some or all of the cytosine bases present in the sample are converted to 5hmC or otherwise provided as 5hmC. In other embodiments of step 206, for the same reasons described above, it may be useful to provide an adaptor in which some or all of the cytosine bases present in the sample are converted or otherwise provided as 5 g mC. On the one hand, the sample DNA product of step 206 can be purified from any enzyme, salt, buffer, or other reagent used in step 206.

[0118] In step 208 of method 200, the sample DNA, including the adaptor applied thereto in step 206, is processed to convert 5mC and 5hmC present in the DNA and optionally in the adaptor nucleic acid molecule into modified methylcytosine. In one embodiment, step 208 is substantially similar to or identical to step 108 of method 100.

[0119] In step 210, the sample DNA product from step 208, including converted 5mC and converted 5hmC, is treated with cytidine deaminase to selectively convert unmethylated or otherwise unmodified cytosine bases present in the sample DNA into uracil. In some embodiments, step 210 is substantially similar to or identical to step 110 of method 100. In some embodiments, step 210 further includes enhancing the deamination process by adding a reagent for inducing or maintaining the formation of single-stranded DNA (ssDNA). In one example, the reagent is a helicase, a single-strand binding protein, etc. In one aspect, the single-strand binding protein is the T4 gene 32 protein. Without being limited by any particular theory, it has been observed that cytidine deaminase preferentially converts deoxycytidine in ssDNA to deoxyurea, while exhibiting low activity towards dsDNA. In one aspect, the addition of a helicase or single-strand protein during the deamination process in step 210 can ensure that the substrate maintains its single-strand conformation and drives the deamination reaction to completion.

[0120] In step 212, the sample DNA generated in step 210 is amplified using PCR. In one embodiment, step 212 is substantially similar to or identical to step 112 of method 100.

[0121] In step 214, the sequence of the product of PCR amplification step 212 is determined. In one embodiment, step 214 is substantially similar to or identical to step 114 of method 100.

[0122] It is worth noting that embodiments of method 200 according to this disclosure may include one or more additional steps or omit one or more described steps of method 200. Generally, method 200 can be modified in any suitable manner to still achieve the result of providing modified sample DNA capable of distinguishing between methylated and unmethylated cytosine. Other variations of method 200 falling within the scope of this disclosure will become apparent from the additional examples and descriptions included herein.

[0123] Example

[0124] The following examples are intended to be illustrative and are not intended to be limiting in any way.

[0125] Example 1:

[0126] In this example, the library preparation workflow is applied to ligating nucleic acid adaptors using engineered DNA ligase or wild-type T4 DNA ligase.

[0127] In some embodiments, in the methods described herein, engineered ligases can provide improvements over naturally occurring (i.e., wild-type) ligases, including improved transformation efficiency at very low sample input levels (i.e., ligation efficiency of the intron to the target in the formation of the target product during intron ligation), and improved ligation specificity, which results in reduced sensitivity to intron concentration and reduced observed intron dimers. Another advantage of using engineered ligases is that the ligation reaction time with naturally occurring DNA ligases is reduced from several hours (e.g., overnight ligation) to less than 1 hour (e.g., 5 minutes) with engineered DNA ligases.

[0128] In some embodiments, an engineered DNA ligase is selected to provide enhanced ligation activity. In one aspect, the ligase is engineered to produce samples with low concentrations of DNA for induction into the method according to this disclosure. Samples that can produce low concentrations of DNA include cell-free DNA, circulating tumor DNA, DNA isolated from circulating tumor cells, circulating fetal DNA, DNA isolated from virus-infected cells, fine needle aspirates, or single cells isolated by FACS (fluorescence-activated cell sorting), laser capture microscopy, or microfluidic devices.

[0129] In some embodiments, engineered DNA ligases are selected for use with lower concentrations of nucleic acid adaptors. In one aspect, providing a lower adaptor concentration in the ligation reaction mixture to minimize the generation of adaptor dimers (i.e., the first adaptor directly ligated to the second adaptor) may be useful. Exemplary engineered DNA ligases used in the methods according to this disclosure include those disclosed in U.S. Patent Application Publication No. 2018 / 0320162, filed May 7, 2018, belonging to Miller et al., which is incorporated herein by reference. In one aspect, engineered DNA ligases for ligating nucleic acid adaptors as disclosed herein differ from the DNA ligases used for nick / gap repair in step 203 of method 200. For example…

[0130] refer to Figure 3As shown in Figure 4, the use of engineered DNA ligase in the methylation sequencing workflow demonstrates that not only are the results for multiple sequencing metrics comparable to those of wild-type T4 DNA ligase, but the engineered DNA ligase also improves the conversion efficiency from 98.87% to 99.19% compared to the unmethylated λ DNA spiked control (conversion efficiency is calculated as follows: 100 – [percentage of methylation detected in unmethylated λ DNA]). Here, sequencing-ready libraries were prepared from 2 ng of purified cfDNA from each of three healthy donors using either the engineered DNA ligase or the wild-type T4 DNA ligase in the library preparation workflow. Otherwise, samples were prepared for sequencing using standard library preparation and the EM-seq workflow according to Method 100. Molecular recovery of samples prepared using engineered DNA ligase also demonstrated an improvement in repetition rate (i.e., the fraction of mapped reads where any two reads share the same 5' and 3' coordinates) compared to wild-type T4 DNA ligase.

[0131] Example 2:

[0132] This example demonstrates the application of Taq DNA ligase and purine-free / pyrimidine-free endonuclease I in the end-repair step of nucleic acid sequencing library preparation.

[0133] Generally, cfDNA contains a nick (i.e., a discontinuity in the dsDNA molecule where there are no phosphodiester bonds between adjacent nucleotides on one strand of dsDNA). The nick in dsDNA (including cfDNA) provides a replication initiation site for DNA polymerase, enabling a process called nick translation. When nick translation occurs, DNA polymerase extends the 3' hydroxyl terminus of the nick site, removes nucleotides via 5' to 3' exonuclease activity, and replaces them with dNTPs. Therefore, the strand of dsDNA located at the 3' of the nick, including any methylated or otherwise modified bases, is not present in the product of nick translation, and the corresponding methylation signal is lost.

[0134] According to this disclosure, modifications to the library preparation method for methyl sequencing enable the repair of dsDNA damage before nick translation occurs. In one aspect, initially excluding DNA polymerase from the end repair and A-tailing reactions, and treating the sample with at least one of a DNA ligase and a purine-free / pyrimidine-free endonuclease, can help preserve this methylation signal. After end repair, Taq DNA polymerase is added to complete A-tailing before adaptor ligation. In one aspect, the ligase preferably catalyzes the formation of a phosphodiester bond between the 5'-phosphate and 3'-hydroxyl groups of two adjacent DNA strands in the nicked dsDNA. Specifically, the ligase should preferentially catalyze bond formation when the strand to be ligated hybridizes with the complementary DNA strand and (without nicks) accurately pairs. An example of a suitable ligase for nick repair is a Taq DNA ligase from a species of the genus *Thermomyces*.

[0135] To demonstrate the application of nick repair prior to A-tailing, 2 ng of purified cfDNA from three healthy donors were treated with or without Taq DNA ligase derived from *Thermophilus HB8*. Specifically, end repair and A-tailing reactions were performed in the absence of DNA polymerase, followed by sample treatment with Taq DNA ligase. Subsequently, Taq DNA polymerase was added prior to adaptor ligation to complete A-tailing. Sequencing libraries were then prepared from the Taq DNA ligase-treated DNA using the method according to this disclosure. Specifically, as described with respect to method 100, sample DNA was further prepared for adaptor ligation in the end repair and A-tailing steps, followed by ligation of 5mC adaptors, 5mC / 5hmC transformation, cytosine deamination, and PCR amplification. Referring to Figure 5, the data show that the addition of Taq DNA ligase significantly improved the methylation signal at the 3' end of the sequencing template and reduced methylation bias.

[0136] As used herein, the term "methylation bias" or "M-bias" refers to the deviation of the expected average methylation level from each position in a sequencing read. Ideally, the average methylation level should remain constant across all positions within a sequencing read; however, various factors can cause variations in the average methylation level, resulting in M-bias. In one instance, methylation signals may be lost due to processing steps during the collection and preparation of nucleic acids for sequencing. This loss of methylation signal (also known as hypomethylation) can reduce the measured average methylation level. This effect is typically observed at the ends of sequencing reads, as shown in Figure 5, where the methylation frequency (plotted on the y-axis) decreases significantly towards the 3' end of the sequencing read (i.e., starting approximately from position 125 on the x-axis). For each sample in Figure 5, the addition of Taq DNA ligase (+Taq) resulted in an increase in the measured average methylation level at the 3' end of the sequencing read compared to the sample without Taq DNA ligase treatment (-Taq). Overall, the +Taq assay showed more uniform methylation frequencies and a more constant average methylation level in sequencing reads, indicating an overall improvement in M-bias.

[0137] On the other hand, in addition to ligases suitable for nick repair, apurine / pyrimidine-free endonucleases can be used to treat damaged dsDNA before end repair and A-tailing. One example of an apurine / pyrimidine-free endonuclease is APEI. APEI catalyzes nick formation in the phosphodiester backbone of dsDNA at nick sites. When a nick present on the backbone includes a hydroxyl group at the 5' end, there may be too much steric hindrance for polynucleotide kinases to phosphorylate that 5' end, thus preventing Taq DNA ligase from repairing the nick. Without being constrained by any particular theory, treatment with APEI is expected to effectively remove the unphosphorylated nucleobases present at the 5' end of the nick site. Subsequently, a one-nucleotide gap can be filled using T4 DNA polymerase, and the resulting product can then be repaired with Taq DNA ligase.

[0138] Example 3:

[0139] This example demonstrates the use of methyltransferases to restore methylation signals.

[0140] In human DNA, cytosine at CpG sites is methylated on both strands (i.e., symmetrical methylation). During replication, the methyltransferase DNMT1 restores a specific methylation pattern consistent with the parent DNA on the daughter strand by methylating hemimethylated DNA. The methylation specificity of DNMT1 can be utilized to restore methylation signals erased by nick translation during the end-repair step of sequencing library preparation. DNMT1 can be added after end-repair and A-tailing, before ligation into sequencing adaptors, to restore the symmetrical methylation pattern, thereby improving the methylation signal detected by sequencing.

[0141] Example 4:

[0142] This example demonstrates the use of hydroxymethylation adaptors for better prevention of deamination.

[0143] It has been demonstrated that different TET enzymes exhibit sequence background bias unique to that enzyme. Therefore, oxidation efficiency can vary depending on the sequence identity of the input DNA. Consequently, sequences to which TET exhibits lower oxidation efficiency will have less protection against deamination, for example, when treated with APOBEC3A. To mitigate the potential sequence dependence of a given TET enzyme, sequencing adaptors can be provided in which each 5mC is converted to 5hmC or otherwise replaced by 5hmC. Adaptors where each methylcytosine is 5hmC will allow for better protection against deamination and result in higher yields (e.g., due to improved consistency of the adaptor sequence) and better aggregation during sequencing on flow cell-based sequencing instruments (e.g., Illumina-based sequencing platforms).

[0144] Referring to Figure 6, the conversion efficiency between the 5mC sequencing adaptor (i.e., a sequencing adaptor where each methylcytosine is 5mC) and the 5hmC sequencing adaptor (i.e., a sequencing adaptor where each methylcytosine is 5hmC) demonstrates that, in the EM-seq workflow, using the 5hmC adaptor results in a higher overall conversion rate of methylated cytosine and a decrease in deamination to uracil via cytidine deaminase. Conversion efficiency was determined by analyzing cytosine-to-uracil conversion against unmethylated λ genomic DNA and the entire non-CpG background in the human genome (because in humans, cytosine methylation only occurs on the cytosine of CpG dinucleotides). For clarity, non-deamination cytosine will be detected as methylcytosine (i.e., false positives).

[0145] Example 5:

[0146] This example illustrates the use of glycosylated hydroxymethylation linkers during EM-seq library preparation to better prevent deamination.

[0147] In one respect, according to publicly available sequencing adaptors, deamination can be further prevented by replacing each methylcytosine with 5gmC (e.g., 5mC, 5hmC) through conversion or other means. For example, 5mC sequencing adaptors can be treated with T4-βGT to produce adaptors containing 5gmC. This ensures that the 5hmC present in the sequencing adaptor is not further oxidized to 5fC, which may not completely prevent deamination.

[0148] Example 6:

[0149] This example demonstrates the addition of at least one of a helicase and an ssDNA-binding protein, such as the T4 gene 32 protein, during deamination to provide a base molecule in single-stranded form.

[0150] Apolipoprotein B mRNA editing enzymes catalyze peptide-3 family enzymes, including APOBEC3A, which is a cytidine deaminase. It acts as an innate immune response factor by inhibiting the replication of viral and transposon elements through binding and deamination. Cytidine deaminases convert deoxycytidine in ssDNA to deoxyuridine, exhibiting lower activity towards dsDNA. The addition of helicases, ssDNA-binding proteins, or a combination thereof during deamination can help ensure the substrate maintains its ssDNA conformation and drive the deamination reaction to completion.

[0151] Example 7:

[0152] This example describes reducing the volume of the cytidine deamination reaction to improve deamination efficiency.

[0153] In some embodiments, reducing the volume of the deamination reaction to effectively increase the concentration of the reagent and drive the deamination reaction to completion may be useful. In one aspect, the total amount of dsDNA present in the reaction may be equal to or less than about 10 nanograms (ng). In another aspect, the total amount of dsDNA present in the reaction may be equal to or less than about 1 ng, 100 picograms (pg), 10 pg, or 1 pg. In another aspect, the concentration of dsDNA present in the reaction may be equal to or less than about 1 ng, 100 pg, 10 pg, or 1 pg / 100 microliters (μL). In another aspect, the concentration of cytidine deaminase is from about 0.05 micromoles per liter (μM) to about 0.5 μM. In another aspect, the concentration of cytidine deaminase is from about 0.1 μM to about 0.2 μM. In another aspect, the total volume of the composition is less than about 100 μL, 50 μL, 25, or 10 μL.

[0154] Example 8:

[0155] This example describes how to eliminate the nucleic acid purification step after cytidine deamination during sample preparation for sequencing to improve the recovery of unique library members.

[0156] On one hand, using the unpurified product of the deamination reaction as a substrate for PCR amplification can lead to better recovery of unique molecules for sequencing, since purification would otherwise result in the loss of some sample DNA.

[0157] Example 9:

[0158] This example demonstrates another use of methyltransferases to demethylate signals.

[0159] As described above, cytosine at CpG sites in human DNA is methylated on both strands (i.e., symmetrical methylation). During replication, DNMT1 (encoded by the DNMT1 gene in humans) restores a specific methylation pattern by methylating hemimethylated DNA. While DNMT1 is one example of a methyltransferase that can be used to methylate hemimethylated DNA, other methyltransferases can also be used to restore hemimethylated DNA, including in vitro. Another example of a suitable methyltransferase according to this disclosure is DNMT5 from Cryptococcus neoformans. In one aspect, Cryptococcus neoformans DNMT5 has been characterized as a highly specific, maintenance-type CpG methyltransferase mediating long-term epigenomic evolution. Cryptococcus neoformans DNMT5 comprises a DNMT domain and an SNF2 ATPase domain, the combination of which is believed to provide at least partially an enzyme with high specificity for hemimethylated DNA (relative to, for example, human DNMT1). Without being constrained by any specific theory, it is believed that hemimethylated DNA preferentially stimulates ATPase activity to trigger structural remodeling of the methyltransferase catalytic pocket, thereby enabling cofactor binding, cytosine base flipping, and catalysis. In contrast, bound unmethylated DNA does not open the catalytic pocket of the DNMT5 enzyme but is ejected upon ATP binding, thus driving high fidelity (i.e., specificity to hemimethylated CpG sites in DNA compared to unmethylated CpG sites in DNA).

[0160] In this example, 5 ng of purified cfDNA for sequencing was prepared according to one of three different methods. In the first method, standard library preparation was performed according to steps 102-112 of method 100. More specifically, 5 ng of purified cfDNA from two healthy donors underwent end repair and A-tailing reactions, followed by the addition of Taq DNA polymerase to complete the A-tailing before adaptor ligation. Then, as described with respect to method 100, a sequencing library was prepared from the resulting DNA via ligation with a 5mC adaptor, transformation with 5mC / 5hmC, deamination of cytosine, and PCR amplification. The second method is the same as the first method, but after end repair and A-tailing but before the addition of Taq DNA polymerase, a Taq DNA ligase derived from *Thermophilus HB8* was added for nick / gap repair. Finally, the third method is the same as the second method, but with the addition of *Cryptococcus neoformans* DNMT5 after the addition of Taq DNA ligase but before the addition of Taq DNA polymerase, as described in step 204 of method 200. Libraries prepared according to each of these three methods were then sequenced to determine the effect of *Cryptococcus neoformans* DNMT5 on methylation bias.

[0161] Referring to Figure 7, the data demonstrate that the addition of Cryptococcus neoformans DNMT5 had a positive effect on M-bias, even exceeding the effect observed in experiments described in Example 2 and Figure 5 by adding Taq DNA ligase. Specifically, for each sample in Figure 5, the addition of both Taq DNA ligase and Cryptococcus neoformans DNMT5 (+Taq +DNMT) resulted in an increase in the average methylation level measured at the 3' end of the sequencing reads compared to samples not treated with Taq DNA ligase (-Taq -DNMT) and samples treated with Taq DNA ligase but not with Cryptococcus neoformans DNMT5 (+Taq -DNMT). Overall, the +Taq +DNMT5 experiment showed more uniform methylation frequencies and a more uniform average methylation level in the sequencing reads, indicating an overall improvement in M-bias.

[0162] The schematic flowcharts shown in the accompanying drawings are typically presented in the form of logical flowcharts. Therefore, the depicted sequence and labeled steps indicate one embodiment of the presented method. Other steps and methods may be conceived that are functionally, logically, or effectively equivalent to one or more steps or portions thereof of the illustrated method. Furthermore, the formats and symbols used in the drawings are intended to explain the logical steps of the method and should be understood as not limiting the scope of the method. While various types of arrows and lines may be used, they should be understood as not limiting the scope of the corresponding method. In practice, some arrows or other connectors may be used only to indicate the logical flow of the method. For example, an arrow may indicate a wait or monitoring period between enumeration steps of the depicted method where no specified duration is specified. Moreover, the order in which a particular method occurs may or may not strictly adhere to the order of the corresponding steps shown.

[0163] The invention is presented in several different embodiments in the following description with reference to the accompanying drawings, wherein the same reference numerals denote the same or similar elements. Throughout this specification, references to "an embodiment," "embodiment," and similar language mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in one embodiment," "in an embodiment," and similar language appearing throughout this specification may, but not necessarily, refer to the same embodiment.

[0164] The features, structures, or characteristics described in this invention can be combined in any suitable manner in one or more embodiments. Numerous specific details are set forth in the following description to provide a thorough understanding of embodiments of the system. However, those skilled in the art will recognize that the system and method can be practiced without one or more of these specific details or using other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring various aspects of the invention. Therefore, the foregoing description is intended to be exemplary and does not limit the scope of the inventive concept.

[0165] Each reference identified in this application is incorporated herein by reference in its entirety.

Claims

1. A method comprising: (a) A first enzyme having double-stranded DNA (dsDNA) nick repair activity is brought into contact with dsDNA containing at least one nick, at least one cytosine, and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, thereby repairing the at least one nick; (b) Converting the at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and (c) Converting the at least one cytosine into uracil.

2. The method according to claim 1, wherein the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

3. The method of claim 1, wherein the at least one methylcytosine is converted into the modified methylcytosine using a second enzyme.

4. The method according to claim 3, wherein the second enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

5. The method of claim 1, wherein the at least one cytosine is converted to uracil using a third enzyme.

6. The method according to claim 5, wherein the third enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

7. The method according to claim 1, wherein the first enzyme is Taq DNA ligase.

8. The method of claim 1, further comprising ligating a sequencing adaptor to the dsDNA.

9. The method of claim 1, further comprising contacting the nick-repaired dsDNA with a fourth enzyme comprising at least one of DNA end repair activity and A-tailing activity.

10. The method according to claim 4, wherein the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2 and wherein the β-glucosyltransferase is T4-phage β-glucosyltransferase.

11. A method comprising: (a) Contacting Taq DNA ligase with a double-stranded DNA (dsDNA) containing at least one nick, at least one cytosine, and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, thereby repairing said at least one nick; (b) Enzymatically converting the at least one methylcytosine to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET); (c) Enzymatic conversion of 5-hydroxymethylcytosine to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT); and (d) The at least one cytosine is enzymatically converted to uracil using the apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

12. The method of claim 11, further comprising ligating a sequencing adaptor to the dsDNA.

13. The method of claim 11, further comprising contacting the nick-repaired dsDNA with an enzyme comprising at least one of DNA end repair activity and A-tailing activity.

14. The method according to claim 1 or 11, wherein step (a) further comprises contacting the dsDNA with a purine-free / pyrimidine-free endonuclease.

15. A method comprising: (a) Contacting a first enzyme having methyltransferase activity with a double-stranded DNA (dsDNA) containing at least one hemimethylated CpG site, at least one additional cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, thereby methylating the hemimethylated CpG site, wherein the at least one additional cytosine remains unmethylated; (b) Converting the at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and (c) Converting the at least one additional cytosine into uracil.

16. The method of claim 15, wherein the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

17. The method of claim 15, wherein the at least one methylcytosine is converted into the modified methylcytosine using a second enzyme.

18. The method of claim 17, wherein the second enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

19. The method of claim 15, wherein the at least one additional cytosine is converted to uracil using a third enzyme.

20. The method of claim 19, wherein the third enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

21. The method of claim 15, wherein the first enzyme is selected from human DNA (cytosine-5) methyltransferase (DNMT1) and Cryptococcus neoformans DNA methyltransferase (DNMT5).

22. The method of claim 15, further comprising ligating a sequencing adaptor to the dsDNA.

23. The method of claim 15, further comprising contacting the dsDNA with a fourth enzyme comprising at least one of DNA end repair activity and A-tailing activity.

24. The method of claim 18, wherein the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2 and wherein the β-glucosyltransferase is T4-phage β-glucosyltransferase.

25. A method comprising: (a) Connecting an adapter to a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, said adapter comprising a nucleic acid sequence comprising at least one methylcytosine, wherein each methylcytosine present in said adapter is 5-hydroxymethylcytosine; (b) Converting the at least one methylcytosine present in the dsDNA and the at least one methylcytosine present in the adaptor into a modified methylcytosine capable of pairing with a guanine base; and (c) Convert at least one additional cytosine into uracil.

26. The method of claim 25, wherein the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

27. The method of claim 25, wherein the at least one methylcytosine is converted into the modified methylcytosine using a first enzyme.

28. The method according to claim 27, wherein the first enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

29. The method of claim 25, wherein the at least one cytosine is converted to uracil using a second enzyme.

30. The method of claim 29, wherein the second enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

31. The method of claim 25, further comprising contacting the dsDNA with a third enzyme comprising at least one of DNA end repair activity and A-tailing activity.

32. The method of claim 28, wherein the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2 and wherein the β-glucosyltransferase is T4-phage β-glucosyltransferase.

33. A method comprising: (a) Connecting an adaptor to a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, the adaptor comprising a nucleic acid sequence comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adaptor is 5-(β-glucosyloxymethyl)cytosine. (b) Converting at least one methylcytosine present in the dsDNA into a modified methylcytosine capable of pairing with a guanine base; and (c) Convert at least one additional cytosine into uracil.

34. The method of claim 33, wherein the modified methylcytosine present in the dsDNA is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

35. The method of claim 33, wherein the at least one methylcytosine in the dsDNA is converted to the modified methylcytosine using a first enzyme.

36. The method according to claim 35, wherein the first enzyme is selected from methylcytosine dioxygenase and β-glucosyltransferase.

37. The method of claim 33, wherein the at least one cytosine is converted to uracil using a second enzyme.

38. The method of claim 37, wherein the second enzyme is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

39. The method of claim 33, further comprising contacting the dsDNA with a third enzyme comprising at least one of DNA end repair activity and A-tailing activity.

40. The method of claim 36, wherein the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2 and wherein the β-glucosyltransferase is T4-phage β-glucosyltransferase.

41. An adaptor for methyl sequencing, the adaptor comprising a first nucleic acid having a 3' end complementary to the 5' end of a second nucleic acid to form a complementary region, the complementary region comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adaptor is 5-(β-glucosyloxymethyl)cytosine.

42. An adaptor for methyl sequencing, the adaptor comprising a first nucleic acid having a 3' end complementary to the 5' end of a second nucleic acid to form a complementary region, the complementary region comprising at least one modified methylcytosine, wherein each modified methylcytosine present in the adaptor is 5-hydroxymethylcytosine.

43. A method comprising: (a) Combining the following: i) a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; ii) a cytidine deaminase; and iii) at least one of a single-stranded DNA binding protein and a helicase; (b) Converting the at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and (c) Converting the at least one cytosine to uracil using the cytidine deaminase.

44. The method of claim 43, wherein the modified methylcytosine is selected from 5-(β-glucosyloxymethyl)cytosine and 5-carboxycytosine (5caC).

45. The method of claim 43, wherein the at least one methylcytosine is converted into the modified methylcytosine using at least one of methylcytosine dioxygenase and β-glucosyltransferase.

46. ​​The method of claim 43, wherein the cytidine deaminase is apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

47. The method of claim 43, further comprising ligating a sequencing adaptor to the dsDNA.

48. The method of claim 43, further comprising contacting the dsDNA with an enzyme having at least one of DNA end repair activity and A-tailing activity.

49. The method according to claim 45, wherein the methylcytosine dioxygenase is TET methylcytosine dioxygenase 2 and wherein the β-glucosyltransferase is T4-phage β-glucosyltransferase.

50. A method comprising: (a) Combining the following: i) a double-stranded DNA (dsDNA) comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; ii) a cytidine deaminase; and iii) at least one of a single-stranded DNA binding protein and a helicase; (b) Enzymatically converting the at least one methylcytosine to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET); (c) Enzymatic conversion of 5-hydroxymethylcytosine to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT); and (d) The at least one cytosine is enzymatically converted to uracil using the apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A).

51. The method of claim 50, further comprising ligating a sequencing adaptor to the dsDNA.

52. The method of claim 50, further comprising contacting the dsDNA with an enzyme having at least one of DNA end repair activity and A-tailing activity.

53. A composition comprising: Double-stranded DNA (dsDNA) comprising at least one cytosine and at least one modified cytosine selected from 5-carboxycytosine and 5-(β-glucosyloxymethyl)cytosine; Buffer; and Cytidine deaminase, The amount of dsDNA mentioned therein is from about 1 picogram to about 10 ng. The concentration of the cytidine deaminase is from about 0.05 μM to about 0.5 μM, and The total volume of the composition is less than about 100 μL.

54. A method comprising: (a) A first composition comprising double-stranded DNA (dsDNA), said double-stranded DNA comprising at least one cytosine and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine; (b) The at least one methylcytosine is enzymatically converted to one of 5-hydroxymethylcytosine and 5-carboxycytosine using TET methylcytosine dioxygenase 2 (TET) to form a first dsDNA product; (c) The 5-hydroxymethylcytosine in the first dsDNA product is enzymatically converted to 5-(β-glucosyloxymethyl)cytosine using T4-phage β-glucosyltransferase (T4-BGT) to form the second dsDNA product; (d) At least one cytosine in the second dsDNA product is enzymatically converted to uracil by apolipoprotein B mRNA editing enzyme catalytic subunit 3A (APOBEC3A), thereby forming a methyl sequencing product composition comprising said APOBEC3A and the third dsDNA product. (e) Combining the methyl sequencing product composition with the amplification composition; and (f) The third dsDNA product is amplified by polymerase chain reaction, wherein the amplification occurs in the presence of APOBEC3A.

55. The method according to any one of claims 1 to 24, 43 to 52 and 54, further comprising ligating the adaptor to the dsDNA using an engineered DNA ligase.

56. The method according to any one of claims 25 to 40, further comprising ligating the adaptor to the dsDNA using an engineered DNA ligase.

57. A method comprising: (a) Combine the following items: The first enzyme with wound repair activity, A second enzyme possessing 5' to 3' polymerase activity, but lacking 5' to 3' exonuclease activity. A third enzyme with purine / pyrimidine endonuclease activity. Double-stranded DNA (dsDNA) comprising: at least one single-stranded discontinuity selected from nicks and gaps; at least one cytosine; and at least one methylcytosine selected from 5-methylcytosine and 5-hydroxymethylcytosine, and Multiple deoxynucleoside triphosphates (dNTPs); (b) Repair at least one discontinuity in the dsDNA; (c) Converting the at least one methylcytosine into a modified methylcytosine capable of pairing with a guanine base; and (d) Convert the at least one cytosine into uracil.

Citation Information

Patent Citations

  • Engineered ligase variants

    US20180320162A1