Detection of modified nucleobases in DNA samples
The method uses base excision repair enzymes to generate and repair gaps in DNA, addressing the limitations of current detection methods by providing a faster, more specific, and cost-effective way to identify modified nucleobases in DNA.
Patent Information
- Application Number
- JP2025540503
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-13
- Filing Date
- 2024-01-11
- Publication Date
- 2026-01-23
AI Technical Summary
Current methods for detecting modified nucleobases in DNA, such as methylation and DNA damage, are limited by requiring large amounts of material, being expensive, time-consuming, and lacking specificity, and are not easily applicable to a wide range of modifications.
A method involving base excision repair enzymes to remove nucleotides containing modified nucleobases, generating single-strand gaps, and using gap ligation or gap filling to repair these gaps, allowing for precise detection of modified nucleobases through sequencing.
The method is faster, more specific, and less degradative, enabling stable and efficient detection of modified nucleobases without the need for specialized reagents, and is compatible with PCR amplification.
Smart Images

Figure 2026502527000005 
Figure 2026502527000006 
Figure 2026502527000007
Abstract
Description
[Background technology]
[0001] Methylation and various forms of DNA damage products are involved in a variety of important biological processes. Changes in methylation patterns and the appearance of damaged DNA are often among the earliest events observed in various disease states.
[0002] Epigenetic modifications are essential for normal development. For example, methylcytosine, the most widely studied epigenetic modification, is associated with several important processes, including genomic imprinting, X-chromosome inactivation, repetitive element silencing, and carcinogenesis. For example, DNA methylation at the 5-position of cytosine has the specific effect of reducing gene expression and has been found in all vertebrates investigated. In many disease processes, such as cancer, gene promoter CpG islands acquire aberrant hypermethylation, resulting in transcriptional silencing that can be inherited by daughter cells after cell division. Furthermore, changes in DNA methylation have been recognized as a critical component of cancer development. While hypomethylation generally occurs earlier and is associated with chromosomal instability and loss of imprinting, hypermethylation is promoter-associated and can occur subsequent to gene (oncogene suppressor) silencing. Furthermore, hydroxymethylcytosine has also emerged as an important epigenetic modification with potential regulatory roles in gene expression ranging from development to aging. Various cancers show that hydroxymethylcytosine content is consistently and significantly reduced in malignant versus healthy tissue, even in early lesions.
[0003] DNA is constantly subjected to stress from both endogenous and exogenous sources. Bases exhibit limited chemical stability and are vulnerable to chemical modification by various types of damage, including oxidation, alkylation, radiation damage, and hydrolysis. Damage to DNA bases can affect their base-pairing properties and thus be mutagenic. DNA base modifications resulting from these types of DNA damage are widespread and play an important role in influencing physiological states and disease phenotypes. Examples include 7,8-dihydro-8-oxoguanine (8-oxoG) (oxidative damage), 8-oxoadenine (oxidative damage; aging, Alzheimer's disease, Parkinson's disease), 1-methyladenine, 0,6-methylguanine (alkylation; glioma and colorectal carcinoma), benzo[a]pyrene diol epoxide (BPDE), pyrimidine dimers (adduct formation; smoking, exposure to industrial chemicals, exposure to UV light; lung and skin cancer), and 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, and thymine glycol (ionizing radiation damage; chronic inflammatory diseases, prostate cancer, breast cancer, and colorectal cancer). For example, 8-oxoG is a frequent product of DNA oxidation. 8-oxoG has a tendency to base pair with adenine, resulting in G>>C T·A transversion mutations. Another example is the hydrolytic deamination of cytosine and 5-methylcytosine (5-meC), which mispairs with guanine to produce uracil and thymine, respectively, leading to C>>G to T·A transition mutations if unrepaired. In another example, alkylation can generate various DNA base lesions, including 6-meG, N7-methylguanine (7-meG), or N3-methyladenine (3-meA). While 6-meG is mutagenic due to its ability to pair with thymine, 7-meG and 3-meA block replicative DNA polymerases and are therefore cytotoxic. These and many other forms of DNA base damage occur in cells multiple times each day, and only the continuous action of specialized DNA repair systems can prevent the rapid decay of genetic information. In addition to damage to nuclear DNA, mitochondrial DNA also suffers significant oxidative damage, as well as damage from alkylation, hydrolysis, and adducts.For example, oxidative damage is the most common type of damage to mitochondrial DNA, primarily because mitochondria are the major cellular source of reactive oxygen species (ROS). Furthermore, mitochondria house approximately 30% of the cellular pool of S-adenosylmethionine, which can non-enzymatically methylate DNA. Furthermore, exposure to certain agents, such as estrogen, cigarette smoke, and certain chemicals, results in preferential damage to mitochondrial DNA.
[0004] Because DNA damage and epigenetic modifications can be the earliest signs of disease states, detecting epigenetic modifications and DNA damage patterns can be useful for early detection of disease and intervention. However, detection methods have limitations. For example, with regard to methylation status, spectrophotometry can be used to indicate the overall content of modifications in target DNA, but specificity is limited. High-performance liquid chromatography (HPLC) and mass spectrometry are also often used, but they are expensive, require significant amounts of material, and reduce DNA to its constituent nucleosides or nucleotides, thus destroying sequence information for downstream analysis. Immunoprecipitation (IP) using monoclonal antibodies can enrich DNA by targeted modifications, but has been noted to have limited specificity. Restriction digest profiling utilizes fragment analysis of DNA treated with modification-sensitive restriction endonucleases, but requires large amounts of material and is limited to sequences characterized by restriction sites with known sensitivity. Bisulfite sequencing is considered the "gold standard" technique for detecting DNA methylation, but it has important limitations. First, this approach requires a large amount of starting material because the chemical conversion process causes extensive nonspecific damage to DNA. Second, this method can be expensive and time-consuming, potentially requiring multiple sequencing runs. Finally, importantly, it is generally only applicable to the methylcytosine (mC) modification. While mutations that allow targeting a limited number of additional modification types (methylcytosine (mC) and hydroxymethylcytosine (hmC)) have been developed or suggested, these have low yields and still share the other limitations listed above. They are also not easily applicable to other modifications and are quite complex.
[0005] Thus, there is a need in the art for improved methods for detecting modified nucleobases in a DNA sample of interest. The present invention fulfills these needs and, as described below, provides further related advantages.
[0006] Not all of the subject matter described in the Background Art section is necessarily prior art, and it should not be assumed to be prior art merely as a result of its description in the Background Art section. Along these lines, awareness of prior art problems described in the Background Art section or related to such subject matter should not be treated as prior art unless expressly stated to be prior art. Instead, the discussion of any subject matter in the Background Art section should be treated as part of the inventor's approach to a particular problem, which may itself also be inventive. Summary of the Invention
[0007] Embodiments of the present invention include the detection of modified nucleobases, such as epigenetic changes and DNA damage, in DNA samples.
[0008] In one aspect, the present invention provides a method for detecting modified nucleobases in a plurality of nucleic acids, the method comprising: providing a sample containing a plurality of DNA templates; generating complementary copies of the DNA templates, the generation being directed by oligonucleotide primers using a DNA polymerase in the presence of natural dNTPs, such that each complementary copy hybridizes to one of the DNA templates; subjecting the DNA templates and the complementary copies to base excision repair enzymatic treatment, wherein the base excision repair enzyme specifically removes nucleotides containing the modified nucleobase from the DNA template to generate single-stranded gaps at the positions of the modified nucleobases, and the complementary copies are resistant to treatment with the base excision repair enzyme; repairing the single-stranded gaps in the DNA templates to generate continuous DNA template strands and determining the nucleotide sequence of the continuous DNA template strands; and comparing the nucleotide sequences of the continuous DNA template strands and the complementary copies, thereby determining the positions of the modified nucleobases in the DNA templates prior to base excision repair enzymatic treatment.
[0009] In one embodiment, repairing the gaps to form a continuous full-length DNA target fragment strand comprises treating the double-stranded DNA fragment with a DNA ligase enzyme, thereby generating deletions in the DNA target fragment strand at each nucleotide position containing a modified nucleobase of interest. In some embodiments, the DNA ligase enzyme is T4 DNA ligase.
[0010] In another embodiment, repairing the gap to form a continuous full-length DNA target fragment comprises treating the double-stranded DNA fragment strand with a DNA polymerase in the presence of unnatural nucleotides and a DNA ligase, thereby generating nucleotide substitutions in the DNA target fragment strand at each nucleotide position containing the desired modified nucleobase. In one embodiment, the DNA polymerase does not exhibit exonuclease or strand displacement activity, and the DNA ligase enzyme is unable to ligate across single-stranded gaps. In yet another embodiment, the DNA polymerase is Klenow exo or T4 DNA polymerase, and the DNA ligase is E. coli DNA ligase.
[0011] In some embodiments, the step of comparing the nucleotide sequences of the DNA target fragment and the complementary copy strand identifies one or more differences in the sequence of the DNA target fragment strand compared to the sequence of the complementary copy strand, and the location of the one or more differences identifies the location of the modified nucleobase of interest in the DNA target fragment. In certain embodiments, the one or more differences in the sequence of the DNA target fragment strand relative to the sequence of the complementary copy strand are one or more mutations, one or more deletions, or one or more substitutions.
[0012] In some embodiments, the base excision repair enzyme is selected from the group of enzymes shown in Table 1.
[0013] In some embodiments, the base excision repair enzyme is N-methylpurine DNA glycosylase (MPG), MutY homolog (MUTYH), Nth-like DNA glycosylase 1 (NTHL1), Nei-like DNA glycosylase 1 (NEIL1), Nei-like DNA glycosylase 2 (NEIL2), Nei-like DNA glycosylase 3 (NEIL3), 8-oxoguanine DNA glycosylase (OGG1), uracil DNA glycosylase 1 (UNG1), uracil DNA glycosylase 2 (UG2), single-strand selective monofunctional The base excision repair enzyme is selected from the group consisting of soluble uracil glycosylase (SMUG1), thymine DNA glycosylase (TDG), methyl-binding domain 4 (MBD4), FPG, UNG, demeter (DME), demeter-like protein 2 (DMEL-2), demeter-like protein 3 (DMEL-3), ROS1, UDG, apurinic endonuclease (APE1), DNA polymerase β (POLB), XRCCC1, DNA ligase 1 (LIG1), DNA ligase 3 (LIG3), and DNA polymerase γ (POLG). In certain embodiments, the base excision repair enzyme comprises a multifunctional DNA glycosylase enzyme, which exhibits both glycosylase activity and lyase activity. In some embodiments, the multifunctional DNA glycosylase enzyme is FPG, DME, ROS1, DMEL-2, or DMEL-3. In another embodiment, the base excision repair enzyme comprises a first enzyme exhibiting glycosylase activity and a second enzyme exhibiting lyase activity, in one embodiment, the first enzyme is TDG or UDG, and the second enzyme is FPG, DME, ROS1, DMEL-2, or DMEL-3.
[0014] In some embodiments, the DNA polymerase is a high fidelity DNA polymerase.
[0015] In some embodiments, the double-stranded DNA target fragment is genomic DNA, mitochondrial DNA, cell-free DNA, circulating tumor DNA, or a combination thereof.
[0016] In some embodiments, the modified bases of interest are 5-mC, 5-hmC, 5-fC, and / or 5-caC.
[0017] In one embodiment, a single-stranded adaptor-ligated DNA target fragment is immobilized on a solid support, while in another embodiment, a complementary copy strand is immobilized on a solid support.
[0018] In some embodiments, the method further comprises polishing the single-stranded gap with one or more enzymes to generate a free 3' hydroxyl and a free 5' phosphate group at each position of the gap. In one embodiment, the one or more enzymes comprise APE1, endonuclease B, PolB, and / or PNK.
[0019] In some embodiments, the non-natural nucleotide is dZTP, dPTP, dSTP, or dBTP.
[0020] In some embodiments, the DNA template comprises a first adaptor joined to the 5' end of the DNA template and a second adaptor joined to the 3' end of the DNA template. In certain embodiments, the first adaptor is a Y adaptor, and the second adaptor is a Y adaptor or a hairpin adaptor. In some embodiments, at least one of the first adaptor and the second adaptor comprises a unique molecular identifier barcode (UMI). In further embodiments, comparing the sequences of the consecutive DNA template strands and the complementary copy comprises bioinformatically pairing sequences comprising the same unique molecular barcode (UMI). [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 1 is a schematic diagram summarizing an alternative embodiment of the method of the present invention. [Figure 2] FIG. 1 is a schematic diagram summarizing an alternative embodiment of the method of the present invention. [Figure 3]1 shows the structures of two embodiments of unnatural nucleobases. [Figure 4A] 1 is a diagram summarizing one embodiment of a workflow for generating gaps in a DNA target fragment at modified nucleobase positions of interest, and the subsequent detection steps. [Figure 4B] 1 is a diagram summarizing one embodiment of a workflow for generating gaps in a DNA target fragment at modified nucleobase positions of interest, and the subsequent detection steps. [Figure 5A] FIG. 1 is a schematic diagram showing an alternative embodiment of solid-state synthesis of a primer extension reaction. [Figure 5B] FIG. 1 is a schematic diagram showing an alternative embodiment of solid-state synthesis of a primer extension reaction. [Figure 6] The generalized structure of XNTP is shown in more detail. DETAILED DESCRIPTION OF THE INVENTION
[0022] The present invention may be more readily understood by reference to the following detailed description of preferred embodiments of the invention and the examples contained herein. Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0023] References throughout this specification to "one embodiment" or "an embodiment" and variations thereof mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0024] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents, i.e., one or more, unless the content and context clearly dictate otherwise. It should also be noted that the conjunctions "and" and "or" are generally used in their broadest sense to include "and / or," unless the content and context clearly dictate inclusiveness or exclusiveness, as the case may be. Thus, the use of alternatives (e.g., "or") should be understood to mean either one, both, or any combination thereof of the alternatives. Furthermore, when described herein as "and / or," the combination of "and" and "or" is intended to encompass embodiments that include all of the associated items or ideas, as well as one or more other alternative embodiments that include fewer than all of the associated items or ideas.
[0025] Unless the context requires otherwise, throughout the specification and claims that follow, the word "comprise," as well as its synonyms and variations, such as "have" and "include," and variations thereof, such as "comprises" and "comprising," are to be construed in an open and inclusive sense, e.g., "including, but not limited to." The term "consisting essentially of" limits the scope of a claim to particular materials or steps, or those that do not materially affect the basic and novel characteristics of the claimed invention.
[0026] The abbreviation "eg," comes from the Latin exempli gratia, and is used herein to indicate a non-limiting example. Thus, the abbreviation "eg," is synonymous with the term "for example." As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise, and the term "X and / or Y" means "X" or "Y," or both "X" and "Y," and it should also be understood that the letter "s" following a noun denotes both the plural and the singular form of the noun. Furthermore, where features or aspects of the invention are described in terms of a Markush group, the invention is intended to encompass and be described in terms of any individual member of the Markush group and any subgroup of members, as would be recognized by one of skill in the art, and applicant reserves the right to amend this application or claims to specifically refer to any individual member or any subgroup of members of the Markush group.
[0027] Any headings used within this document are merely utilized to facilitate the reader's review and should not be construed as limiting the scope of the invention or the claims. Accordingly, the headings and abstracts of the disclosure provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.
[0028] Where a range of values is provided herein, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value in the stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0029] For example, any concentration range, percentage range, ratio range, or integer range provided herein should be understood to include any integer within the stated range, and, where appropriate, fractions thereof (e.g., tenths and hundredths of integers), unless specifically stated otherwise. Also, any numerical range described herein for any physical characteristic, such as polymer subunits, size, or thickness, should be understood to include any integer within the stated range, unless specifically stated otherwise. As used herein, the term "about" means ±20% of the indicated range, value, or structure, unless specifically stated otherwise.
[0030] Methods for detecting modified DNA nucleobases Described herein are alternative general methods for determining the location and identity of modified DNA nucleobases, such as those resulting from epigenetic modifications or DNA damage, in a DNA target fragment template. These methods are outlined in Figure 1, where in this embodiment, the modified nucleobase of interest is 5-methylcytosine (5-mC). As depicted in this exemplary embodiment, the top strand of a double-stranded DNA target fragment contains one 5-mC residue base-paired with G (represented by the hatched portion of the top strand), while the bottom strand contains no 5-mC residues. Both methods are based on specifically removing nucleotides containing the modified nucleobases of interest (i.e., nucleotides of interest), thereby creating single-stranded gaps at each position of the modified nucleobases of interest ("Step 1" depicted in Figure 1). The identity of the modified nucleobases is determined by the enzyme or chemistry used to specifically remove the nucleotides containing the modified nucleobases. After removal of the nucleotides of interest, the location of the resulting single-stranded gaps can be assessed using two methods for repairing the gaps created in Step 1. The first method involves ligating across the gaps to create a continuous DNA template strand with a deletion at each position of the gap created in step 1 ("gap ligation," see "Step 2A" in Figure 1 ). The second method involves filling the gaps with nucleotides containing alternative, e.g., unnatural DNA bases, to generate a continuous DNA target strand (gap fill, see "Step 2B" in Figure 1 ).
[0031] These methods offer improvements over state-of-the-art workflows for epigenetic detection, for example, those based on bisulfite conversion, overcoming well-known technical drawbacks in the art, such as DNA degradation and reduced genomic complexity. Because the methods disclosed herein are based on enzymatic removal of the modified nucleotides of interest, they are faster, more specific, less degradative, and do not require specialized reagents. Furthermore, the modified (i.e., "converted") DNA templates are stable and easily amplified, for example, by PCR.
[0032] As mentioned above, the method disclosed herein comprises enzymatically removing the nucleotides containing the target modified base in double-stranded DNA target fragment template, so that a single gap is generated at each position where the target nucleotide occurs in the nucleic acid sequence of DNA template.Subsequently, the single-strand gap is repaired, and a continuous DNA template strand is generated by gap ligation or gap filling process.The position of the repaired gap can be identified by multiple DNA sequencing methods as described herein.
[0033] Further details of the methods disclosed herein are shown in Figure 2. As explained, both methods share an upstream workflow ("Step 1" depicted in Figure 2) that involves generating a single-stranded gap by specific removal of the nucleotides that constitute the modified base of interest. The specificity of the removal allows for reliable detection of the nucleobase of interest by pinpointing the location of the newly created gap.
[0034] One method for specifically generating single-stranded, single-nucleotide gaps in DNA target fragments is to utilize DNA glycosylases. DNA glycosylases are a family of enzymes also known in the art as "base excision repair" enzymes. The diversity of DNA glycosylases allows many different nucleobase modifications to be assayed using the methods disclosed herein. Enzymes that only remove nucleobases that generate abasic sites can also be used, as these sites can then be further reacted to form single-nucleotide gaps.
[0035] In the embodiment depicted in Figure 2, epigenetic methylation of cytosine (e.g., 5-mC) can be assayed by treating DNA target fragments with members of the Demeter / ROS1 family of glycosylases, which act directly on 5-mC by removing the nucleotide containing this epigenetic mark. Alternatively, 5-mC can be converted to 5-formylcytosine (5-fC) or 5-carboxylcytosine (5-caC) by oxidation via ten-to-eleven translocation (TET) methylcytosine dioxygenase. 5-fC and 5-caC can then be specifically removed, for example, by thymine DNA glycosylase (TDG).
[0036] One method for repairing the gap formed in step 1 is via "gap ligation," as shown in step 2A ("2A" depicted in Figure 2). In certain embodiments, this method exploits the ability of T4 DNA ligase to ligate across small single-stranded gaps in otherwise continuous double-stranded DNA. This cross-gap ligation creates a double-stranded DNA with a "bulged" base on the opposite side of the ligation site (i.e., a deletion of the strand containing the targeted modified nucleobase and the intact opposite strand). To repair the gap and generate a continuous DNA template strand, a PCR reaction can amplify both strands of the double-stranded target fragment. When sequenced, the gap-joined DNA strand is read as containing a deletion at each gap site when compared to a reference sequence.
[0037] When employing this method, it is advantageous to utilize a sequencing technology with a low deletion error rate to minimize sequencing errors that can generate false signals. To reduce Type 1 (false positive) errors, several methods are known in the art. Additional information, such as sequence context, the identity of the deleted base, and the generation of multiple sequence reads, can identify the removed modified base with high confidence. Because the chemical reaction used to remove the nucleotide of interest is controlled, this method can predict which of the four DNA bases will be detected. The context of the base sequence, such as CpG in DNA methylation, can further verify that the detected deletion is not a base sequence error. Because this method is compatible with PCR amplification, unique molecular identifiers (UMIs) can be utilized to provide consensus sequences before identifying gap ligation events.
[0038] Another way to repair the gap formed in step 1 is via "gap filling," as shown in step 2B ("2B" depicted in Figure 2). In this embodiment, a nucleotide containing an alternative, e.g., non-natural or non-standard, nucleotide base can be incorporated into the gap by a DNA polymerase. The polymerase's incorporation, i.e., gap filling, leaves a nick in the DNA backbone, which can then be sealed by DNA ligase. By providing a DNA polymerase with a single non-natural nucleotide, a single-nucleotide gap can be filled even if the polymerase incorporates poorly. Examples of non-standard nucleotides are disclosed, for example, in U.S. Pat. No. 9,334,534, which is incorporated herein by reference in its entirety.
[0039] The choice of enzymes used in gap-fill protocols is important to ensure that the method is specific for incorporating and ligating non-natural nucleotides without disrupting DNA templates containing nucleotide gaps or nicks in the backbone. Preferred DNA polymerases are those that lack robust exonuclease / proofreading or strand-displacement activity. In certain embodiments, a suitable DNA polymerase may be Klenow exopolymerase or T4 DNA polymerase. In other embodiments, a suitable DNA ligase is one that cannot ligate across DNA gaps, such as E. coli DNA ligase.
[0040] In certain embodiments, PCR can be used to amplify DNA template strands repaired with nucleotides containing unnatural bases (as used herein, the terms "unnatural," "unnatural," and "nonstandard" are used interchangeably). For example, two unnatural bases that specifically and precisely base pair can be used to increase the "DNA alphabet" from four bases to six bases. One unnatural base is incorporated into the DNA target fragment during the gap-fill repair process in step 2B of FIG. 2, while a pair of unnatural bases is incorporated during subsequent PCR amplification of the repaired DNA target fragment. One exemplary pair of unnatural bases that can be used in accordance with the present invention is dZTP and dPTP (e.g., available from Firebird Biomolecular Sciences, Ltd.), shown in FIG. 3.
[0041] In certain embodiments, DNA template strands repaired by gap filling with unnatural bases can be directly sequenced by modifying existing DNA sequencing techniques to include reagents that specifically base pair with unnatural bases. For example, in certain embodiments, fluorescent nucleosides that pair with unnatural bases allow detection by synthetic optical sequencing. In other embodiments, extendible nucleotides whose bases pair with unnatural bases allow sequencing by extension, allowing unnatural bases to be directly detected.
[0042] In some embodiments, the methods disclosed herein may also include additional steps. For example, the methods may include a step of repairing or "polishing" the DNA target fragment prior to step 1 outlined in Figures 1 and 2. Such treatment may ensure that there is no pre-existing damage to the DNA target fragment, e.g., strand nicks, breaks, etc., that could cause false positive errors in downstream analysis.
[0043] In other embodiments, the DNA target fragment can be optionally modified to facilitate the removal of specific DNA nucleobases. As described herein, in one embodiment, specific DNA nucleobases can be oxidized, for example, with ten-eleven translocation (Tet) methylcytosine dioxygenase, which oxidizes 5-mC to 5f-C and 5-caC. In other embodiments, 8-oxo-G lesions can be specifically excised by DNA-formamidopyrimidine glycosylase.
[0044] In other embodiments, the method may include a "polishing" step after step 1 and before step 2. It is known in the art that DNA glycosylases can create a variety of functional groups after cleavage or removal of targeted modified nucleobases. Before repairing a single-stranded gap, the 5' and 3' ends of the gap must be treated to provide the correct chemical sites (e.g., 5' hydroxyl and 3' phosphate groups) for gap ligation or gap filling. In one embodiment, the gap in the DNA target fragment may be treated with polynucleotide kinase (PNK) to generate the necessary 5' and 3' functional groups. In other embodiments, the polishing step may include treatment with a cocktail of DNA repair enzymes, such as a mixture of APE, phosphatase, and kinase.
[0045] In certain embodiments, both the gap ligation reaction and the gap filling reaction can be combined into a single multi-enzyme reaction for the purpose of simplicity and to reduce reaction time and potential sources of error. In one embodiment, the following four steps can be combined: 1) modified nucleotide removal; 2) gap filling; 3) end polishing; and 4) ligation. This approach minimizes the lifespan of unstable single-base gaps that lead to double-strand breaks. By combining these four steps into a "one-pot" reaction, the various reactions can proceed rapidly through unstable intermediates to obtain a stable, continuous DNA template strand.
[0046] In certain embodiments, the methods of the present invention also include a workflow for generating a complementary copy (i.e., a "daughter" strand) of a DNA template (i.e., a "parent" strand). Importantly, the complementary copy is generated prior to the step of enzymatic removal of nucleotides containing the modified nucleobase of interest. Thus, the daughter strand encodes the genetic information of the DNA template, thereby serving as a reference sequence, while the parent strand encodes the epigenetic information through enzymatic conversion. The sequence information obtained from the complementary copy strand and the template strand can be bioinformatically paired and compared to identify the location of the modified nucleobase of interest in the nucleic acid sequence of the original DNA target fragment.
[0047] Overview of the "Parent-Child" Library Workflow For each of the methods described herein, the modified nucleobase of interest may be at least one of 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxycytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (*-oxoG), uracil, 6-methyladenine (6-mA) or 8-oxoadenine, O-6-methylguanine, 1-methyladenine, O-4-methylthymine, 5-hydroxycytosine, 5-hydroxyuracil, or a thymine dimer. In some examples, any combination of multiple modified nucleobases of these types may be detected.
[0048] In one aspect, a method for detecting modified DNA nucleobases in a DNA sample is provided. An exemplary schematic overview of the method is provided in Figures 4A and 4B. The method may include obtaining a DNA sample and fragmenting the DNA to generate a sample of DNA target fragments (Step A). As used herein, the term "target fragment" means that the corresponding DNA fragment is derived from a biological sample and serves as a template for the methods described herein, which interrogate nucleic acid sequences for the presence of specific modified nucleobases. In this non-limiting exemplary depiction, the modified nucleobase of interest is methylated cytosine (5-mC), and each strand of the DNA target fragment (i.e., the sense strand "+" and the antisense strand "-") contains one 5-mC residue.
[0049] In some examples, the DNA sample is genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof obtained from a biological sample.
[0050] The method may then include ligating adapters to the ends of the double-stranded DNA target fragments to create adapter-ligated DNA target fragments (Step B). The adapters may include a region of double-stranded DNA and a region of single-stranded DNA. In the example shown in FIG. 4A, the adapter includes two regions of single-stranded DNA. In other embodiments, the adapters may have any suitable configuration for a downstream step of a particular workflow. For example, in one embodiment, one of the adapters may be a hairpin adapter. The adapters may also include sequences or other features that mediate downstream steps of the workflow, such as sequences for immobilizing the adapter-ligated DNA target fragments on a solid support, sequences for hybridization of oligonucleotide primers, sequences that enable bioinformatic analysis of DNA sequence information (e.g., unique molecular identifiers [UMIs]), etc.
[0051] The method may then include a step of denaturing the adaptor-ligated double-stranded DNA target fragment to generate a single-stranded target fragment sense (+) strand and a single-stranded target fragment antisense (-) strand (Step C).
[0052] The method may then include step D, in which a first primer extension reaction is performed. The primer extension reaction is directed by an extension oligonucleotide (i.e., an oligonucleotide primer) hybridized to the single-stranded DNA target fragment using a DNA polymerase. The primer extension reaction results in a sample of double-stranded DNA fragments containing complementary copy strands hybridized to the DNA target fragment strands (step D). In some examples, the DNA polymerase is a high-fidelity DNA polymerase. In this step, the sample of double-stranded DNA fragments differs from the sample of DNA target fragments in step 1A in that the former contains complementary copy strands synthesized in vitro, while the latter contains two target strands each derived from a biological sample. The primer extension reaction is performed under conditions in which the resulting complementary copy strands do not contain a "native" strand, e.g., a modified nucleobase of interest present in the target strand. For example, in this figure, the first complementary copy strand incorporates a native cytosine residue in place of a methylated cytosine residue in the target strand.
[0053] In some examples, the single-stranded target fragment is immobilized on a solid support prior to performing a first primer extension reaction, as depicted in Figure 5A. As shown here, the complementary copy strand (i.e., the daughter strand) is not immobilized on the solid support and can be physically separated from the immobilized "parent" strand upon denaturation of the double-stranded DNA fragment. In other examples, as depicted in Figure 5B, oligonucleotides complementary to the single-stranded adaptor-ligated target fragment can be immobilized on a solid support, allowing the single-stranded target fragment to be "captured" on the solid support. After capture of the target fragment, a first primer extension reaction can be performed to generate a first complementary copy strand immobilized on the solid support. In this example, denaturation of the double-stranded DNA fragment releases the single-stranded target fragment from the solid support.
[0054] The method then treats the sample of double-stranded DNA fragments with a DNA glycosylase enzyme to remove nucleotides containing the modified nucleobase of interest (e.g., 5-mC in this illustration). It is well recognized in the art that glycosylase enzymes can also be classified as base excision repair enzymes, and both classes of enzymes can be used to practice the methods disclosed herein. Removal of nucleotides containing the modified nucleobase of interest (i.e., the modified nucleotides of interest) generates single-stranded, single-nucleotide gaps in the DNA target fragment at each position of the modified nucleotide of interest (step E in Figure 4A and depicted in Figure 4B). In some instances, multiple DNA glycosylases or other enzymes can be used to generate single-stranded gaps. An appropriate combination of enzymes provides both base-specific glycosylase activity and lyase activity to completely remove the modified nucleotide of interest from the DNA target fragment strand. A non-limiting list of exemplary DNA glycosylase enzymes is provided in Table 1. Notably, according to the present invention, the complementary copy strand remains resistant to DNA glycosylase treatment such that natural nucleotides in the complementary copy strand are not converted into single-stranded, single-base gaps. As used herein, the term "converted," when used in reference to a DNA target fragment, refers to a DNA target fragment or portion thereof that has been treated under conditions sufficient to convert a modified nucleotide of interest into a single-stranded gap.
[0055] This method can then include a subsequent step of repairing the single-strand gap in the DNA target strand by either gap ligation or gap filling, as described herein. In either case, gap repair generates a continuous DNA target fragment strand (i.e., a continuous template strand derived from the parent target fragment). The repaired DNA parent strand and unconverted daughter strand can optionally be amplified by conventional PCR technology.
[0056] This method can then comprise determining the nucleotide sequence of the parent strand and the daughter strand of DNA.A variety of sequencing platforms and methodologies are suitable for carrying out the present invention.In a preferred embodiment, the sequencing method is the Sequencing by Expansion (SBX) protocol developed by the present inventors, see, for example, U.S. Patent No. 7,939,259 and U.S. Patent No. 10,301,345, and International Publication No. 2020 / 172,479 and International Publication No. 2020 / 236,526.
[0057] The method may then include a step of bioinformatically analyzing the sequence data to compare the sequences of the parent and daughter strands and determine the location of the modified nucleobase of interest in the DNA target fragment before enzymatic conversion. The daughter strand (complementary strand) is used as a reference sequence because it encodes the genetic information of the original DNA target fragment. The parent strand encodes the epigenetic information of the DNA target fragment. Differences in the base sequences of the parent and daughter strands at specific positions (e.g., the presence of a mutation in the base sequence of the parent strand) indicate the location of the modified nucleobase of interest in the DNA target fragment. As disclosed herein, gap ligation protocols result in deletions in the sequence of the parent strand relative to the daughter strand at the position of the modified nucleotide of interest, while gap fill protocols result in base substitutions at the same positions.
[0058] Therefore, in one embodiment of the method of the present invention, the step of comparing the nucleotide sequences of the DNA target fragment and the complementary copy strand identifies differences in the sequence of the DNA target fragment strand relative to the sequence of the complementary copy strand, and the location of one or more differences identifies the location of the modified nucleobase of interest in the DNA target fragment. In one embodiment, the difference is a mutation, for example, one or more mutations. In another embodiment, the difference is a deletion, such as one or more deletions. In another embodiment, the difference is a substitution, such as one or more substitutions. In another embodiment, the difference is a mutation, deletion, and / or substitution.
[0059] The method provided herein is particularly useful in multiplex format, in which multiple DNA target fragments with different sequences and / or different nucleobase modification patterns are assayed in a common sample or pool.Therefore, the method described herein can provide the advantage of avoiding the need to separate different target fragments into separate vessels during one or more steps of nucleobase modification detection assay.
[0060] Further details regarding the above method are provided below.
[0061] The practice of the present invention employs, unless otherwise indicated, conventional techniques of molecular biology, microbiology, recombinant DNA, and the like, which are within the skill of the art. Such techniques are fully explained in the literature. See, for example, Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, Second Edition (1989), OLIGONUCLEOTIDE SYNTHESIS (M.J. Gait Ed., 1984), the series METHODS IN ENZYMOLOGY (Academic Press, Inc.), and CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F.M. Ausubel, R. Brent, R.E. Kingston, D.D. Moore, J.G. Siedman, J.A. Smith, and K. Struhl, eds., 1987).
[0062] DNA sample / target fragment In one embodiment, the DNA is obtained or provided from biological sample.The DNA obtained or provided from biological sample can be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof.
[0063] DNA sample can be obtained from patient or subject, from environmental sample or from organism of interest.In some embodiments, DNA sample is extracted, purified or derived from cell or cell cluster, body fluid, tissue sample, organ and / or organelle.In preferred embodiment, sample DNA is total genomic DNA.
[0064] In some instances, genomic DNA and mitochondrial DNA may be obtained separately from the same biological sample or source. Many different methods and techniques are available for isolating genomic DNA and mitochondrial DNA. Generally, such methods involve disruption and lysis of the starting material, followed by removal of proteins and other contaminants, and finally recovery of DNA. Protein removal can be achieved, for example, by digestion with proteinase K, followed by salting out, organic extraction, gradient separation, or binding of DNA to a solid support (either anion exchange or silica techniques). Mitochondrial DNA can be similarly isolated after initial isolation of mitochondria. DNA can be recovered by precipitation with ethanol or isopropanol. Commercially available kits are also available for isolating nuclear DNA or mitochondrial DNA. The choice of method depends on many factors, including, for example, the amount of sample, the required amount and molecular weight of DNA, the purity required for downstream applications, and time and cost.
[0065] The disclosed methods utilize mild enzymatic and chemical reactions that avoid the substantial degradation associated with methods such as bisulfite sequencing, and are therefore useful for the analysis of low-input samples such as circulating cell-free DNA, circulating tumor DNA, and single-cell analysis.
[0066] In some embodiments, DNA sample is the DNA found in blood, and is the circulating cell-free DNA (cfDNA) that does not exist in cells.cfDNA can be isolated from blood or plasma using methods known in the art.For example, commercially available kits are available for isolating cfDNA, including circulating DNA kit (Qiagen).DNA sample can be obtained from enrichment step, including but not limited to antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0067] In some instances, the isolated DNA is fragmented into multiple shorter double-stranded DNA pieces. Generally, fragmentation of DNA can be performed physically or enzymatically.
[0068] For example, physical fragmentation can be achieved by acoustic shearing, sonication, microwave irradiation, or hydrodynamic shearing. Acoustic shearing and sonication are the primary physical methods used to shear DNA. For example, the Covaris® instrument (Woburn, Massachusetts) is an acoustic device for shearing DNA to lengths ranging from 100 bp to 5 kb. Covaris also manufactures tubes (gTubes) for processing 6-20 kb samples for mate-pair libraries. Another example is the Bioruptor® (Denville, New Jersey), an ultrasonic device used to shear chromatin, DNA, and disrupt tissue. It can shear small amounts of DNA to lengths ranging from 150 bp to 1 kb. Digilab's Hydroshear® (Marlborough, Massachusetts) is another example, using hydrodynamic forces to shear DNA. Nebulizers, such as those manufactured by Life Technologies (Grand Island, NY), can be used to atomize liquids using compressed air, shearing DNA into fragments ranging from 100 bp to 3 kb in a matter of seconds. Nebulization can result in sample loss, so in some instances, it may not be the desired fragmentation method for samples with limited volume. Sonication and acoustic shearing may be better fragmentation methods for smaller sample volumes, as the total amount of DNA from the sample can be more efficiently retained. Other physical fragmentation devices and methods, known or to be developed, may also be used.
[0069] DNA can be fragmented using a variety of enzymatic methods. For example, DNA can be treated with DNase I or a combination of maltose-binding protein (MBP)-T7 Endo I and a nonspecific nuclease such as Vibrio vulnificus nuclease (Vvn). The combination of the nonspecific nuclease and T7 Endo acts synergistically to generate nonspecific nicks, neutralize the nicks, and generate fragments with 8 nucleotides or less dissociated from the nick site. In another example, DNA can be treated with NEBNext® dsDNA Fragmentase® (NEB, Ipswich, MA). NEBNext® dsDNA Fragmentase generates dsDNA breaks in a time-dependent manner to yield DNA fragments of 50–1,000 bp, depending on the reaction time. NEBNext dsDNA Fragmentase contains two enzymes: one randomly generates nicks in dsDNA, and the other recognizes the nicked site and cleaves the DNA strand opposite the nick, generating dsDNA breaks. The resulting DNA fragments contain short overhangs, a 5'-phosphate and a 3'-hydroxyl group.
[0070] In some examples, the DNA sample is fragmented into a specific size range. For example, the DNA sample can be fragmented into fragments of about 25-100 bp, about 25-150 bp, about 50-200 bp, about 25-200 bp, about 50-250 bp, about 25-250 bp, about 50-300 bp, about 25-300 bp, about 50-500 bp, about 25-500 bp, about 150-250 bp, about 100-500 bp, about 200-800 bp, about 500-1300 bp, about 750-2500 bp, about 1000-2800 bp, about 500-3000 bp, about 800-5000 bp, or any other size range within these ranges. For example, the DNA sample can be fragmented into fragments of about 50-250 bp. In some instances, the fragments may be larger or smaller than about 25 bp.
[0071] A DNA target fragment can be any DNA fragment derived from a biological sample that has a sequence of interest, which may or may not contain epigenetic modifications or DNA damage to one or more nucleic acid bases. In some embodiments, the DNA target fragment can contain cytosine modifications (i.e., 5mC, 5hmC, 5fC, and / or 5caC). A DNA target fragment can be a single DNA molecule in a sample, or the entire population of DNA molecules in a sample (or a subset thereof) that have, for example, cytosine modifications. A DNA target fragment can also contain multiple DNA sequences, so that the methods described herein can be used to generate a library of DNA target fragments that can be analyzed individually (e.g., by determining the sequence of each target) or in groups (e.g., by multiplexed DNA sequencing).
[0072] In embodiments, the methods described herein include adding an adapter DNA molecule to a double-stranded DNA target fragment. The adapter DNA or DNA linker is a short, chemically synthesized, single-stranded or double-stranded oligonucleotide that can be ligated to one or both ends of another DNA molecule. The double-stranded adapter can be synthesized so that each end of the adapter has a blunt end or a 5' or 3' overhang (i.e., a sticky end). The DNA adapter is ligated to the DNA target fragment to provide sequences for, for example, a primer extension reaction and a sequencing reaction using complementary primers and / or bioinformatics analysis (e.g., clustering related sequences into families based on shared unique molecular identifiers, UMIs).
[0073] Before ligating the adapter, the ends of the DNA fragments can be prepared for ligation. For example, by end repair and creating blunt ends with 5' phosphate groups. Fragmented DNA can be blunt-ended by several methods known to those skilled in the art. In a specific method, the ends of the fragmented DNA are "polished" with T4 DNA polymerase and Klenow polymerase, a procedure well known to those skilled in the art, and then phosphorylated with polynucleotide kinase enzyme. Then, using Taq polymerase or Klenow exo minus polymerase enzyme, a single "A" deoxynucleotide is added to both 3' ends of the DNA molecule, creating a one-base 3' overhang complementary to the one-base 3' T' overhang at the double-stranded end of the adapter.
[0074] In some instances, an adapter may contain two partially complementary oligonucleotides that hybridize to form a region of double-stranded sequence but also retain a region of single-stranded, non-hybridizing sequence. The region of single-stranded sequence may contain a "universal" oligonucleotide binding sequence, which allows all target fragments in a library to bind to the same oligonucleotide, which may be a capture oligonucleotide, allowing the target fragments to be localized to a solid support; an oligonucleotide primer for a primer extension reaction; a PCR primer; a sequencing primer; or a combination thereof. In certain cases, an adapter may contain two regions of single-stranded, non-hybridizing sequence (i.e., a first 5' single-stranded region and a second 3' single-stranded region). This configuration is known in the art as a "Y" adapter. The first and second single-stranded regions of the Y adapter are not complementary and may contain different primer hybridization sequences and other features.
[0075] The two single-stranded regions of the adapter typically comprise at least 10, 15, or 20 consecutive nucleotides on each strand. The lower limit of the length of the single-stranded region is typically determined by the need to provide a sequence suitable for function, such as primer binding for primer extension, PCR, and / or sequencing. Theoretically, there is no upper limit to the length of the single-stranded region, except that it is generally advantageous to minimize the total length of the adapter, for example, to facilitate separation of unbound adapters from adapter-ligated double-stranded DNA target fragments after the ligation step. Therefore, it is preferred that the single-stranded region be less than 50, 40, 30, or 25 consecutive nucleotides long on each strand.
[0076] The double-stranded region of an adapter is a short double-stranded region, typically containing five or more consecutive base pairs, formed by the annealing of two partially complementary polynucleotide strands. Generally, it is advantageous for the double-stranded region to be as short as possible without losing functionality. In this context, "functional" means that the double-stranded region forms a stable duplex under standard reaction conditions for enzyme-catalyzed nucleic acid ligation reactions.
[0077] The exact nucleotide sequence of the adapter is generally not a material of the present invention, and can be selected by the user so that the desired sequence element is ultimately included in the consensus sequence of a library of adapter-ligated double-stranded DNA target fragments. For example, additional sequence elements can be included to provide binding sites for primers that will ultimately be used to sequence complementary copy strands of the DNA target fragments. The adapters can further include "tag" sequences, unique molecular identifiers (UMIs), and / or sample identifier sequences that can be used to tag, track, and distinguish target fragments and their complementary copies from a particular source. The general characteristics and uses of such sequences are well known in the art.
[0078] The terminus of the single-stranded region of the adapter may be biotinylated or may have another functionality that allows for capture or immobilization on a surface, such as a solid support. Alternative functionality other than biotin is known in the art, as described in applicant's WO 2020 / 172479, entitled "Methods and Devices for Solid-Phase Synthesis of Xpandomers for use in Single Molecule Sequencing," which is incorporated herein by reference in its entirety.
[0079] "Ligation" of an adapter to the 5' and 3' ends of each fragmented double-stranded nucleic acid target fragment involves joining the two polynucleotide strands of the adapter to the double-stranded target polynucleotide such that a covalent bond is formed between both strands of the two double-stranded molecules. Preferably, such covalent bonding occurs via the formation of a phosphodiester bond between the two polynucleotide strands, although other covalent bonding means (e.g., non-phosphodiester backbone bonds) may be used. However, it is essential that the covalent linkage formed in the ligation reaction allows polymerase readthrough so that the resulting construct can be copied in a primer extension reaction using a primer that binds to a sequence in the region of the adapter-target construct derived from the adapter molecule.
[0080] In some instances, the adaptor and DNA target fragment may be incubated with a ligase to covalently link the adaptor and DNA target fragment. Ligase catalyzes the formation of a phosphodiester bond between juxtaposed 5' phosphate and 3' hydroxyl ends in double-stranded DNA or RNA. The enzyme joins blunt and cohesive ends and repairs single-stranded nicks in double-stranded DNA. An exemplary ligase is T4 ligase, the enzyme most frequently used for cloning. Another ligase that can be used is Escherichia coli (E. coli) DNA ligase, which preferentially ligates cohesive double-stranded DNA ends but is also active on blunt-ended DNA in the presence of Ficoll or polyethylene glycol. Another ligase that can be used is DNA ligase Ilia, which is known to function in mitochondria.
[0081] Before the adaptor-target construct is further processed, the product of the ligation reaction may be subjected to a purification step to remove unbound adaptor molecules.
[0082] Ligating adapters to both ends of double-stranded DNA target fragments results in a pool of adapter-ligated double-stranded DNA target fragments with adapters on both ends of the target.
[0083] There are several standard methods for separating the strands of adapter-ligated double-stranded DNA target fragments by denaturation, including heat or chemical denaturation with either 100 mM sodium hydroxide or formamide solutions. The pH of the solution of single-stranded DNA fragments can be neutralized by adjusting the pH with an appropriate solution of acid, or by buffer exchange through a size-exclusion chromatography column, preferably pre-equilibrated in a buffer solution.
[0084] Complementary copy strand (i.e., "daughter" strand) In embodiments disclosed herein, a single-stranded DNA target fragment provides a template nucleic acid (i.e., "parent" strand) for generating a complementary copy strand (i.e., "daughter" strand) of the target fragment via a primer extension reaction. As used herein, the term "primer extension reaction" is used interchangeably with "nucleic acid polymerization reaction" and refers to an in vitro method for creating a new strand of nucleic acid or extending an existing nucleic acid (e.g., DNA or RNA) in a template-dependent manner. The first complementary copy strand is generated by extending an oligonucleotide primer with a first DNA polymerase such that a first complementary copy of the template strand is extended in the 3' direction of the oligonucleotide primer.
[0085] In the embodiment where DNA target fragment is double-stranded, one or both strands can serve as template strand of primer extension reaction.For example, when one strand (" sense " strand) serves as template, it generates complementary copy that is complementary to sense strand.Similarly, when antisense strand serves as template, it generates complementary copy that is complementary to antisense strand.When both strands serve as template, it generates separate complementary copies for each of sense strand and antisense strand.In a preferred embodiment, each strand of double-stranded DNA target fragment is template nucleic acid.
[0086] As used herein, the term "complementary" refers to a nucleic acid sequence capable of forming Watson-Crick base pairs. For example, the complement of a first sequence is a sequence capable of forming Watson-Crick base pairs with the first sequence. The term "complementary" does not necessarily mean that a sequence is complementary to the entire length of its complementary strand, but the term can mean that a sequence is complementary to a portion of it. Thus, in some embodiments, complementarity encompasses sequences that are complementary along the entire length or a portion of the sequence. For example, two sequences may be complementary to each other along at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length of the sequences. Herein, the term "sequence" encompasses, but is not limited to, nucleic acid sequences, polynucleotides, oligonucleotides, probes, primers, primer-specific regions, and target-specific regions. Despite the mismatch, the two sequences should be capable of selectively hybridizing to one another under appropriate conditions.
[0087] Primer extension can be performed by any method that allows polymerase-based extension of a primer annealed (i.e., hybridized) to a single-stranded DNA target fragment. In some embodiments, simple primer extension involves adding a primer and a first DNA polymerase to the target DNA fragment under conditions that allow primer hybridization and primer extension by the polymerase. Of course, such a reaction includes the nucleotides, buffers, and other reagents required for primer extension that are known in the art. Importantly, the nucleotides included in the primer extension reaction are "native," i.e., unmodified, nucleotides, and therefore, the complementary copy strand does not contain modifications to the target nucleic acid base. The complementary copy strand is generated to encode and preserve the genetic sequence of the DNA target strand.
[0088] Any number of methods are known for detecting primer extension products. In some embodiments, the primer is detectably labeled (e.g., at its 5' end or otherwise positioned so as not to interfere with 3' extension of the primer), and after primer extension, the length and / or amount of the labeled extension product is detected by detecting the label.
[0089] In certain embodiments, the primer used in primer extension reaction is annealed to the (single-stranded) primer binding sequence in the single-stranded region of adapter.The term " annealing " used in this context refers to the sequence-specific binding / hybridization of primer to the primer binding sequence in the adapter region of adapter-linked DNA target fragment under the conditions used in the primer annealing step of initial primer extension reaction.Primer annealing conditions are well known in the art (see, for example, Sambrook et al., 2001, Molecular Cloning, A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor Laboratory Press, NY; Current Protocols, eds. Ausubel et al.).
[0090] In a preferred embodiment, the DNA polymerase is a high-fidelity DNA polymerase. The fidelity of a DNA polymerase is the result of accurate replication of the desired template. Specifically, this involves multiple steps, including the ability to read the template strand, select the appropriate nucleoside triphosphate, and insert the correct nucleotide at the 3' primer end so that Watson-Crick base pairing is maintained. In addition to effectively distinguishing between correct and incorrect nucleotide incorporation, some DNA polymerases possess 3'→5' exonuclease activity. This activity, known as "proofreading," is used to remove the incorrectly incorporated mononucleotide and then replace it with the correct nucleotide.
[0091] In certain embodiments, high-fidelity DNA polymerases suitable for practicing the present invention include KAPA HiFi DNA Polymerase commercially available from Roche Diagnostics Corp., Q5® High-Fidelity DNA Polymerase commercially available from New England Biolabs, Inc., and engineered Pfu DNA polymerases such as Pfu-X commercially available from Jena Biosciences.
[0092] solid phase synthesis In certain embodiments, the primer extension reaction can be carried out on a solid support. Thus, in a further aspect, the present invention provides methods for solid-phase nucleic acid synthesis using adapter-ligated DNA target fragments with known sequences at their 5' and 3' ends (e.g., sequence features designed into the adapters).
[0093] The terms "solid support," "solid phase," and "substrate" are used interchangeably herein and refer to a material or group of materials having one or more surfaces that are rigid or semi-rigid. In many embodiments, at least one surface of the solid support is substantially flat, e.g., the surface of a polymeric microfluidic card or chip. In some embodiments, it may be desirable to physically separate regions of the card or chip for different reactions, e.g., with etched channels, trenches, wells, raised areas, pins, etc. According to other embodiments, the solid support will take the form of insoluble beads, resins, gels, membranes, microspheres, or other geometric configurations composed, e.g., of controlled pore glass (CPG) and / or polystyrene.
[0094] The present invention encompasses solid-phase synthesis methods in which a capture moiety is immobilized on a solid support. In certain cases, the capture moiety comprises a first end covalently attached to the solid support and a second end providing a functional group capable of binding to the 5' end of a single-stranded adaptor-ligated DNA target fragment. In this case, the single-stranded DNA target fragment is immobilized on the solid support, and the complementary copy strand is not immobilized on the solid support. In other examples, the capture moiety comprises an extender oligonucleotide capable of hybridizing to the 3' end of the single-stranded adaptor-ligated target fragment. The single-stranded adaptor-ligated DNA target fragment is hybridized to the extender oligonucleotide, and a primer extension reaction is performed. In this case, only the complementary copy strand is immobilized on the solid support. These alternative solid-phase synthesis configurations are illustrated in Figures 5A and 5B.
[0095] As used herein, the term "immobilization" refers to the association, attachment, or bonding between a molecule (e.g., a linker, adapter, or oligonucleotide) and a support in a manner that provides a stable association under the conditions of extension, amplification, ligation, and other processes described herein. Such bonding can be covalent or non-covalent. Non-covalent bonding includes electrostatic, hydrophilic, and hydrophobic interactions. Covalent bonding is the formation of a covalent bond characterized by the sharing of electron pairs between atoms. Such covalent bonding can be directly between the molecule and the support, or can be formed by a cross-linking agent or by including specific reactive groups on the support, the molecule, or both. Covalent attachment of a molecule can be achieved using a binding partner such as avidin or streptavidin immobilized on the support and non-covalent binding of a biotinylated molecule to the avidin or streptavidin. Immobilization can also involve a combination of covalent and non-covalent interactions.
[0096] Any suitable covalent attachment means known in the art can be used for these purposes. The attachment chemistry selected will depend on the nature of the solid support and any derivatization or functional groups applied to the solid support. The extended oligonucleotide may contain a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. A specific exemplary embodiment of a suitable surface chemistry includes conventional streptavidin / biotin interaction chemistry, such as functionalizing the solid support with a linker moiety containing a terminal biotin moiety. In this embodiment, the 5' end of the single-stranded DNA fragment (or oligonucleotide) is attached to the linker moiety. Attachment is mediated by a streptavidin moiety provided by the 5' end of the single-stranded DNA fragment. The linker moieties disclosed herein can be of sufficient length to link the single-stranded DNA fragment to the support so that the support does not significantly interfere with the primer extension reaction.
[0097] Alternatively, immobilization of a capture moiety or oligonucleotide (e.g., an extension oligonucleotide) to a solid support can be achieved by covalently linking the capture oligonucleotide to the solid support via a click reaction. In this embodiment, the covalent linkage can be mediated by a maleimide-PEG-alkyne linker crosslinked to the solid support. The alkyne moiety provided by the end of the linker distal to the substrate can react with the azide group provided by the 5' end of the capture oligonucleotide. Methods for functionalizing solid supports with maleimide linker polymers are provided in applicant's published patent application WO 2020 / 172479, entitled "Methods and Devices for Solid-Phase Synthesis of Xpandomers for Use in Single Molecule Sequencing," which is incorporated herein by reference in its entirety.
[0098] In certain cases, the link between the capture moiety and the solid support is cleavable, allowing the primer extension product to be released from the support after synthesis.Cleaving linkers and methods for cleaving such linkers are known and can be adopted in the provided method using the knowledge of those skilled in the art.For example, the cleavable linker can be cleaved by an enzyme, a catalyst, a chemical compound, temperature, electromagnetic radiation, or light.Optionally, the cleavable linker includes a moiety that can be hydrolyzed by beta-elimination, a moiety that can be cleaved by acid hydrolysis, an enzymatically cleavable moiety, or a photocleavable moiety.In some embodiments, a suitable cleavable moiety is a photocleavable (PC) spacer or linker phosphoramidite available from Glen Research.
[0099] Glycosylase-mediated removal of modified nucleotides In one embodiment, the method comprises incubating the double-stranded DNA fragment product of the primer extension reaction with a DNA glycosylase enzyme to specifically remove the desired modified nucleotide.Many DNA glycosylases have been identified that target a wide range of specific modified nucleobases and DNA damage elements, including sequence mismatches and a wide range of epigenetic modifications.Exemplary genetic modifications that can be detected by the described method include, but are not limited to, 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxycytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (oxoG), uracil, methyladenine (mA), etc.
[0100] There are two major classes of DNA glycosylases: monofunctional and bifunctional. Monofunctional glycosylases have only glycosylase activity and cleave the N-glycosidic bond that connects damaged or modified nucleobases to the sugar-phosphate backbone of DNA. All DNA glycosylases cleave glycosidic bonds, but they differ in their base substrate specificity and their reaction mechanisms. Bifunctional glycosylases also have apurinic or apyrimidinic site (AP) lyase activity, which allows them to cleave the phosphodiester bond of DNA at the base lesion and create a single-strand break. The methods disclosed herein require a DNA glycosylase or a combination of glycosylases to provide both glycosylase and lyase activities to completely excise the modified nucleotide of interest from the DNA target strand.
[0101] Exemplary DNA glycosylases that can be used in the described methods are listed in Table 1. In some examples, one or more of the DNA glycosylases listed in Table 1 can be used in the described methods to remove modified nucleotides of interest from DNA target fragments. While selected DNA glycosylases are specifically identified in the present disclosure, it is understood that any DNA glycosylase can be used in performing the removal step of the described methods.
[0102] [Table 1] TIFF2026502527000002.tif68157
[0103] In one example, a suitable DNA glycosylase that directly removes 5-mC can be a member of the DEMETER (DME) family of DNA glycosylases, such as DME, ROS1, and DEMETER-like protein 2 (DMEL-2, DML2) and DEMETER-like protein 3 (DMEL-3, DML3). The Arabidopsis DME gene encodes a 1,729-amino acid protein with a central DNA glycosylase domain (amino acids 1167-1368) containing a helix-hairpin-helix (HhH) motif. The HhH motif in DME catalyzes the removal of 5-mC (see, e.g., Choi et al., 2002. Cell 110:33-42).
[0104] In some instances, a suitable DNA glycosylase that acts directly on 5mC may be an ortholog of DME. As used herein, the term "ortholog" refers to one of two or more homologous gene sequences found in different species. Table 2 provides an exemplary list of DME orthologs that can be used in accordance with the present invention.
[0105] [Table 2] TIFF2026502527000004.tif194157
[0106] When a DNA glycosylase is a bifunctional enzyme, the DNA glycosylase (e.g., DME, or its orthologue) exhibits both glycosylase and lyase activity. The reaction mechanisms of bifunctional DNA glycosylases are well known in the art (see, e.g., Scharer and Jiricny, 2001, Bioessays 23:270-281). In some cases, a conserved aspartic acid acquires a proton from a conserved lysine residue, attacking the C1' carbon of the deoxyribose ring, generating a covalent DNA-enzyme intermediate. A beta-elimination or gamma-elimination reaction releases the enzyme from DNA and cleaves one of the phosphodiester bonds.
[0107] Mutations that inactivate or optimize the appropriate characteristics of DNA glycosylases are also contemplated by the present invention. For example, DNA glycosylases can be engineered to increase their stability and / or solubility. DNA glycosylases can also be engineered to optimize desired substrate specificity.
[0108] In one embodiment, thymine DNA glycosylase (TDG) can be used to remove its known targets, 5-carboxycytosine (5-caC) and 5-formylcytosine (5-fC), and, with the additional step of modifying bases in a DNA sample, can be used to identify modified bases it does not specifically recognize, 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC). For example, DNA target fragments can be treated with ten-eleven translocation (TET) enzymes prior to treatment with TDG. The TET family of proteins includes three human proteins (TET1, TET2, and TET3) and is a cytosine oxygenase that catalyzes the conversion of 5-methylcytosine (5-mC) to 5-hydroxymethylcytosine (5-hmC). 5-hmC can be further oxidized to 5-formylcytosine (5-fC) and 5-carboxylcytosine (5-caC) by TET proteins (see, e.g., Parker, et al. 2019. Biochemistry 58:450-467). In another example, a suitable TET enzyme can be "nTET" (i.e., "ngTET") isolated from Naegleria (see, e.g., Hashimoto, et al. 2014. Nature 506(7488):391-395). TDG can be used to remove existing 5-caC and 5-fC modified bases present in DNA target fragments.
[0109] Other equivalent methods that alter the selective removal of modified nucleotides are possible. For example, a similar method can be performed to detect the same bases using thymine DNA glycosylase (TDG) and uracil DNA glycosylase (UDG).
[0110] The nucleotide removal process described above is performed using purified enzymes, which may be recombinant enzymes containing heterologous tags to facilitate purification. Protein tags are well known in the art and include, for example, terminal polyhistidine tags that allow purification by immobilized metal affinity chromatography (IMAC). In certain cases, it may be desirable to include two or more protein purification steps. For example, glycosylase enzymes used in the methods disclosed herein preferably should be free of contaminating nucleic acids. In some examples, the protein purification steps include one or more of size exclusion chromatography, ion exchange chromatography, affinity chromatography, etc.
[0111] In some embodiments, the nucleotide removal reaction comprises a suitable buffer, suitable cofactors, additives, and a sufficient amount of purified DNA glycosylase to achieve the desired base removal reaction such that all of the desired modified nucleobases in the DNA target fragment are removed to generate abasic sites.
[0112] After treatment with DNA glycosylase, the double-stranded DNA fragment is asymmetrically modified. Importantly, the DNA target strand (i.e., parent strand) contains a single-stranded gap at the original position of the desired modified base. In contrast, the complementary copy strand (i.e., daughter strand) remains unchanged (i.e., "unconverted") because the natural nucleotide incorporated during the initial primer extension reaction is resistant to conversion to a single-stranded gap via glycosylation.
[0113] According to the present invention, the DNA target strand and complementary copy strand may be assessed by a number of established and emerging nucleic acid sequencing techniques, including, but not limited to, deep sequencing, next generation sequencing, and nanopore sequencing.
[0114] Sequencing by extension One preferred nanopore sequencing technology for the methods disclosed herein is the "Sequencing by Expansion" (SBX) protocol developed by Stratos Genomics (see, e.g., Kokoris et al., U.S. Patent No. 7,939,259, "High Throughput Nucleic Acid Sequencing by Expansion"). SBX is based on the polymerization of highly modified, unnatural nucleotide analogs called "XNTPs." In general terms, SBX utilizes biochemical polymerization to transcribe the sequence of a DNA template (e.g., the first and second complementary copy strands of a DNA target fragment) into measurable polymers called "Xpandomers." The transcribed sequences are encoded along the Xpandomer backbone in high-signal-to-noise reporters spaced approximately 10 nm apart, designed for high signal-to-noise and highly differentiated response. These differences result in significant performance improvements in the sequence read efficiency and accuracy of Xpandomers compared to natural DNA.
[0115] XNTPs are extendable 5' triphosphate-modified unnatural nucleotide analogs compatible with template-dependent enzymatic polymerization. XNTPs have two distinct functional regions: a selectively cleavable phosphoramidate bond connecting the 5' α-phosphate to the nucleobase, and a symmetrically synthesized reporter tether (SSRT) attached within the nucleoside triphosphoramidate at a position that allows controlled extension by cleavage of the phosphoramidate bond. SSRTs contain linkers separated by selectively cleavable phosphoramidate bonds. Each linker is attached to one end of a reporter code. The XNTP substrate and its incorporation into the daughter strand product of template-dependent polymerization are in a "constrained" configuration. The constrained configuration of polymerized XNTPs is the precursor to the extended configuration, as seen in the Xpandomer product. The transition from the constrained to the extended configuration occurs upon cleavage of the phosphoramidate P-N bond within the primary backbone of the daughter strand.
[0116] The transition from the constrained to the extended configuration occurs through the cleavage of selectively cleavable phosphoramidate bonds in the primary backbone of the daughter strand. In this embodiment, the SSRTs contain one or more reporters or reporter codes specific to the nucleobases to which they are linked, thereby encoding the sequence information of the template. In this way, the SSRTs provide a means to extend the length of the Xpandomer and reduce the linear density of the sequence information of the parent strand.
[0117] The SSRT (i.e., "tether") of XNTP contains several functional elements, or "features," including a polymerase enhancer region, a reporter code, and a translational control element (TCE). Each of these features performs a specific function during Xpandomer translocation through the nanopore, producing a series of unique and reproducible electronic signals. The SSRT is designed to control the rate of Xpandomer translocation by the TCE through a combination of steric and / or electrical repulsion. Different reporter codes are sized to block ion flow through the nanopore at different, measurable levels. Specific SSRT polymer sequences can be efficiently synthesized using phosphoramidite chemistry, typically used in oligonucleotide synthesis. Reporter codes and other features can be designed by selecting specific phosphoramidite sequences from commercially available and / or proprietary libraries. Such libraries include, but are not limited to, polyethylene glycols with lengths of 1 to 12 or more ethylene glycol units and aliphatic polymers with lengths of 1 to 12 or more carbon units. In certain embodiments, the SSRT comprises a feature called a "polymerase-enhancing region" at the end of the SSRT proximal to the nucleotide triphosphoramidate diester. The polymerase-enhancing region may comprise a positively charged polyamine spacer (e.g., a primary, secondary, tertiary, or quaternary amine) or a triamine spacer (three secondary amines separated by three carbons) that promotes incorporation of the XNTP structure by a nucleic acid polymerase. In certain embodiments, the polymerase-enhancing region comprises two repeating units of spermine.
[0118] As used throughout this disclosure, the term "reporter construct" refers to an element of an SSRT that includes a reporter code, a symmetric chemical branch, and a migration control element. In certain embodiments, the reporter construct is a polymer that includes, consecutively from a first end to a second end, a first reporter code, a symmetric chemical branch carrying a migration control element, and a second reporter code. The term "carrying" refers to the covalent bond between the symmetric branch and the migration control element, which results in a favorable orientation of the migration control element relative to the two reporter codes. As further described herein, the symmetric chemical branch can be represented by the letter "Y," in which two reporter codes are linked to the arms of the Y and a translocation control element is linked to the stem of the Y. Thus, the two reporter codes are linked in-line by a branch, which carries the translocation control element in a linear, in-line, perpendicular orientation to the SSRT.
[0119] As used throughout this disclosure, the terms "linker A" and "linker B" refer to regions of the SSRT that comprise a polymerase-enhancing region and one or more translocation slowing features or regions, respectively, and, in certain embodiments, a spacer region comprising a polymer of, for example, PEG6, that can be customized to regulate the length of the SSRT that traverses within the nanopore.
[0120] In certain embodiments, the XNTP can be a compound having the following generalized structure, as depicted in FIG.
[0121] In one embodiment, for example, when the compound is used to sequence a DNA template, R can be H.
[0122] In certain embodiments, the nucleobase is adenine, cytosine, guanine, thymine, uracil, or a nucleobase analog. As will be understood by those skilled in the art, adenine, cytosine, guanine, thymine, and uracil are naturally occurring nucleobases. As used herein, the term "nucleobase analog" refers to a non-naturally occurring nucleobase that can form a Watson-Crick base pair with a complementary nucleobase on an adjacent single-stranded nucleic acid template.
[0123] As discussed herein, a reporter construct is a polymer having a first end and a second end, and comprises, consecutively from the first end to the second end, a first reporter code, a symmetric chemical branching chain carrying a migration-controlling element, and a second reporter code. This series of features reflects the symmetric structure of the reporter construct (and the entire SSRT including the symmetric linkers, linker A and linker B), in which the sequences of the two reporter codes are identical and are linked in-line in reverse by the symmetric chemical branching chain.
[0124] To obtain sequence information, Xpandomers are translocated from the cis reservoir to the trans reservoir through a nanopore. As the Xpandomers translocate, the reporters enter the stem until their translocation control elements stop them at the stem entrance. The reporters are held at the base until TCE enters and allows them to pass through the base, after which translocation proceeds to the next reporter. Upon passing through the nanopore, each reporter code on the linearized Xpandomers generates a distinct and reproducible electronic signal specific to the nucleobase to which it is linked.
[0125] In certain embodiments, Xpandomer produced by SBX chemistry can be analyzed using a nanopore-based sequencing chip. The nanopore-based sequencing chip can incorporate a large number of sensor cells configured as an array. For example, the chip can include an array of 1 million cells, configured with 1,000 rows and 1,000 columns of cells. Each cell in the array can include control circuitry integrated on a silicon substrate. Such nanopore-based sequencing chips, devices, and systems are described, for example, in the applicant's published patent application WO 2021 / 219795, the entire contents of which are incorporated herein by reference.
[0126] Diagnostic and prognostic methods In certain embodiments, the methods may relate to diagnosing an individual having a condition characterized by a methylation level and / or pattern at a particular locus in a test sample that differs from the methylation level and / or pattern at the same locus in a sample considered normal or in the absence of the condition. The methods may also be used to predict an individual's susceptibility to a condition characterized by a level and / or pattern of methylation at a locus that differs from the level and / or pattern of methylation at the locus exhibited in the absence of the condition.
[0127] Exemplary conditions suitable for analysis using the methods described herein may include, for example, cell proliferative disorders or predispositions to cell proliferative disorders; metabolic dysfunctions or disorders; immune dysfunctions, damage or disorders; central nervous system dysfunctions, damage or disease; symptoms of aggression or behavioral disorders; clinical, psychological and social consequences of brain injury; psychotic disorders and personality disorders; dementia or related syndromes; cardiovascular disease, dysfunction and damage, gastrointestinal dysfunction, damage or disease, respiratory system dysfunction, damage or disease, lesions, inflammation, infection, immunity and / or recovery, bodily dysfunction, damage or disease such as abnormalities in developmental processes, skin, muscle, connective tissue or bone dysfunction, damage or disease, endocrine and metabolic dysfunction, damage or disease, headache or sexual dysfunction, and combinations thereof.
[0128] Abnormal methylation of CpG islands associated with tumor suppressor genes can cause decreased gene expression. Increased methylation of such regions can result in a gradual decrease in normal gene expression, leading to the selection of cell populations with selective growth advantages. Conversely, decreased methylation (hypomethylation) of cancer genes can regulate normal gene expression, resulting in the selection of cell populations with selective growth advantages.
[0129] Thus, in certain embodiments, the disease or condition analyzed for methylation levels is cancer. Exemplary cancers that can be evaluated using the methods of the present invention can include, but are not limited to, cancers of soft tissues such as breast, prostate, lung, bronchus, colon, rectum, bladder, kidney, renal pelvis, pancreas, oral cavity or pharynx (head and neck), ovary, thyroid, stomach, brain, esophagus, liver, intrahepatic bile duct, cervix, larynx, heart, testis, gastrointestinal stroma, pleura, small intestine, anus, anal canal and anorectum, vulva, gallbladder, bone, joint, hypopharynx, eyeball or orbit, nose, nasal cavity, middle ear, nasopharynx, ureter, peritoneum, amniotic membrane or mesentery. Other cancers that can be evaluated include, for example, chronic myeloid leukemia, acute lymphocytic leukemia, malignant mesothelioma, acute myeloid leukemia, chronic lymphocytic leukemia, multiple myeloma, gastrointestinal carcinoid tumor, non-Hodgkin's lymphoma, Hodgkin's lymphoma, melanoma of the skin, etc.
[0130] With particular regard to cancer, altered DNA methylation has been recognized as one of the most common molecular alterations in human neoplasia. Hypermethylation of CpG islands located in the promoter regions of tumor suppressor genes is a well-established common mechanism for gene inactivation in cancer (Esteller, Oncogene 21(35):5427-40(2002)). In contrast, global hypomethylation and increased gene expression have been reported for many cancer genes (Feinberg, Nature 301(5895):89-92(1983), Hanada, et al., Blood 82(6):1820-8(1993)). Cancer diagnosis or prognosis can be performed using the methods described herein based on the methylation status of specific sequence regions of genes, including, but not limited to, coding sequences, 5' regulatory regions, or other regulatory regions that affect transcription efficiency.
[0131] The reference genomic DNA (e.g., gDNA considered "normal") and the test genomic DNA compared in the diagnostic or prognostic method can be obtained from different individuals, different tissues, and / or different cell types.In certain embodiments, the genomic DNA samples compared can be from the same individual, but from different tissues or different cell types, or from tissues or cell types that are differentially affected by disease or symptoms.Similarly, the genomic DNA samples compared can be derived from the same tissue or the same cell type, and the cells or tissues are differentially affected by disease or symptoms.
Claims
1. 1. A method for detecting modified nucleobases in a plurality of nucleic acids, said method comprising: providing a sample comprising a plurality of DNA templates; generating complementary copies of the DNA templates, said generating being directed by oligonucleotide primers using a DNA polymerase in the presence of native dNTPs, said generating generating complementary copies of each of the DNA templates such that each complementary copy hybridizes to one of the DNA templates; subjecting the DNA template and the complementary copy to base excision repair enzyme treatment, wherein the base excision repair enzyme specifically excises a nucleotide containing the modified nucleobase from the DNA template to generate a single-stranded gap at the position of the modified nucleobase, and wherein the complementary copy is resistant to treatment with the base excision repair enzyme; repairing the single-stranded gap in the DNA template to generate a continuous DNA template strand; determining the nucleotide sequence of said contiguous DNA template strand and said complementary copy; comparing the nucleotide sequences of the continuous DNA template strand and the complementary copy, thereby determining the location of the modified nucleobase within the DNA template prior to base excision repair enzymatic treatment; A method comprising:
2. 2. The method of claim 1, wherein repairing the single-stranded gaps in the DNA template to generate the continuous DNA template strand comprises treating the DNA template with a DNA ligase enzyme to create a deletion in the continuous DNA template strand at each position of the nucleotide in the DNA template that contains the modified nucleobase.
3. 2. The method of claim 1, wherein the step of repairing the single-stranded gap in the DNA template to generate the continuous DNA template strand comprises treating the DNA template with a DNA polymerase enzyme and a DNA ligase enzyme in the presence of unnatural nucleotides, thereby resulting in a nucleotide substitution in the continuous DNA template strand at each position of the nucleotide that contains the modified nucleobase in the DNA template.
4. 4. The method of any of claims 1 to 3, wherein said step of comparing the nucleotide sequences of said contiguous DNA template strand and said complementary copy identifies one or more differences in the sequence of said contiguous DNA template strand relative to the sequence of said complementary copy, and the positions of said differences identify the positions of modified nucleobases in said DNA template.
5. 5. The method of claim 4, wherein the one or more differences in the sequence of the DNA target fragment strand relative to the sequence of the complementary copy strand are one or more mutations, one or more deletions, or one or more substitutions.
6. The base excision repair enzyme is N-methylpurine DNA glycosylase (MPG), MutY homolog (MUTYH), Nth-like DNA glycosylase 1 (NTHL1), Nei-like DNA glycosylase 1 (NEIL1), Nei-like DNA glycosylase 2 (NEIL2), Nei-like DNA glycosylase 3 (NEIL3), 8-oxoguanine DNA glycosylase (OGG1), uracil DNA glycosylase 1 (Ung1), uracil DNA glycosylase 2 (Ung2), single-strand selection 6. The method of claim 1, wherein the target polypeptide is selected from the group consisting of specific monofunctional uracil glycosyltransferase (SMUG1), thymine DNA glycosylase (TDG), methyl-binding domain 4 (MBD4), FPG, Ung, Demeter (DME), DME-2, DME-3, ROS1, UDG, apurinic endonuclease (APE1), DNA polymerase β (POLB), XRCC1, DNA ligase 1 (LIG1), DNA ligase 3 (LIG3), and DNA polymerase γ (POLG).
7. The method of any one of claims 1 to 6, wherein the base excision repair enzyme comprises a DNA glycosylase enzyme, and the DNA glycosylase enzyme exhibits glycosylase activity and lyase activity.
8. 8. The method of claim 6 or 7, wherein the DNA glycosylase enzyme is selected from the group consisting of FPG, DME, ROS1, DEL-2, and DEL-3.
9. The method according to any one of claims 1 to 8, wherein the base excision repair enzymes comprise a first enzyme exhibiting glycosylase activity and a second enzyme exhibiting lyase activity.
10. 10. The method of claim 9, wherein the first enzyme is TDG or UDG, and the second enzyme is selected from the group consisting of FPG, DME, ROS1, DEL-2, and DEL-3.
11. The method according to any one of claims 1 to 10, wherein the DNA polymerase is a high-fidelity DNA polymerase.
12. The method of any of claims 1 to 11, wherein the DNA template comprises genomic DNA, mitochondrial DNA, cell-free DNA, circulating tumor DNA, or a combination thereof.
13. 13. The method of any of claims 1 to 12, wherein the modified nucleobase is selected from the group consisting of 5-mC, 5-hmC, 5-fC and 5-caC.
14. The method of any one of claims 1 to 13, wherein the DNA template is immobilized on a solid support.
15. The method of any one of claims 1 to 14, wherein the complementary copy is immobilized on a solid support.
16. 16. The method of any of claims 1 to 15, further comprising polishing the single-stranded gap with one or more enzymes to generate a free 3' hydroxyl group and a free 5' phosphate group at each position of the gap.
17. 17. The method of claim 16, wherein the one or more enzymes are selected from the group consisting of APE1, endonuclease B, Pol B, and PNK.
18. The method of any one of claims 3 to 17, wherein the non-natural nucleotide is selected from the group consisting of dZTP, dPTP, dSTP, and dBTP.
19. 18. The method of any one of claims 3 to 17, wherein the DNA polymerase enzyme does not exhibit exonuclease or strand displacement activity and the DNA ligase enzyme does not have the ability to ligate across single-stranded gaps.
20. 20. The method of claim 19, wherein the DNA polymerase enzyme is Klenow exopolymerase or T4 DNA polymerase and the DNA ligase enzyme is E. coli DNA ligase.
21. The method according to any one of claims 2 and 4 to 17, wherein the DNA ligase enzyme is T4 DNA ligase.
22. 22. The method of any one of claims 1 to 21, wherein the DNA template comprises a first adaptor joined to the 5' end of the DNA template and a second adaptor joined to the 3' end of the DNA template.
23. 23. The method of claim 22, wherein the first adaptor is a Y adaptor and the second adaptor is a Y adaptor or a hairpin adaptor.
24. 24. The method of claim 23, wherein at least one of the first adapter and the second adapter comprises a unique molecular identifier barcode (UMI).
25. 25. The method of Claim 24, wherein said comparing the sequences of said contiguous DNA template strand and said complementary copy comprises bioinformatically pairing sequences that contain the same unique molecular barcode (UMI).