Detection of modified nucleobases in nucleic acid samples
Patent Information
- Application Number
- JP2025522532
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-19
- Publication Date
- 2026-09-03
AI Technical Summary
Current methods for detecting modified nucleobases in DNA, such as epigenetic modifications and DNA damage, are limited by requiring large amounts of material, being expensive, time-consuming, and lacking specificity, and are not easily applicable to a wide range of modifications.
A method involving DNA glycosylase treatment to convert modified nucleobases to abasic sites, followed by generating complementary copies with abasic bypass DNA polymerases to identify the modified nucleobases through sequence comparison, using chemoenzymatic nucleobase conversion reactions.
Enables sensitive and specific detection of modified nucleobases with reduced DNA damage and lower material requirements, allowing for early disease detection and intervention.
Smart Images

Figure 00000038_0000 
Figure 00000044_0000 
Figure 00000044_0001
Abstract
Description
[Technical Field]
[0001] Incorporation by Reference of Sequence Listing This application has been filed electronically in xml format and contains a Sequence Listing, which is incorporated herein by reference in its entirety. The xml copy has the filename P37864-WO_Sequence_Listing, was created on October 13, 2023, and is 5kB in size. [Background technology]
[0002] background The products of methylation and various forms of DNA damage are involved in a variety of important biological processes. Changes in methylation patterns and the appearance of damaged DNA are often among the earliest events observed in various disease states.
[0003] Epigenetic modifications are essential for normal development. For example, methylcytosine, the most widely studied epigenetic modification, is associated with several important processes, including genomic imprinting, X-chromosome inactivation, repetitive element silencing, and carcinogenesis. For example, DNA methylation at the 5-position of cytosine has the specific effect of reducing gene expression and has been found in all vertebrates investigated. In many disease processes, such as cancer, gene promoter CpG islands acquire abnormal hypermethylation, resulting in transcriptional silencing that can be inherited by daughter cells after cell division. Furthermore, changes in DNA methylation have been recognized as a critical component of cancer development. While hypomethylation generally occurs earlier and is associated with chromosomal instability and loss of imprinting, hypermethylation is promoter-associated and can occur subsequent to gene (oncogene suppressor) silencing. Furthermore, hydroxymethylcytosine has also emerged as an important epigenetic modification with potential regulatory roles in gene expression ranging from development to aging. Various cancers show that hydroxymethylcytosine content is consistently and significantly reduced in malignant versus healthy tissue, even in early lesions.
[0004] DNA is constantly subjected to stress from both endogenous and exogenous sources. Bases exhibit limited chemical stability and are vulnerable to chemical modification by various types of damage, including oxidation, alkylation, radiation damage, and hydrolysis. Damage to DNA bases can affect their base-pairing properties and thus be mutagenic. DNA base modifications resulting from these types of DNA damage are widespread and play an important role in influencing physiological states and disease phenotypes. Examples include 7,8-dihydro-8-oxoguanine (8-oxoG) (oxidative damage), 8-oxoadenine (oxidative damage; aging, Alzheimer's disease, Parkinson's disease), 1-methyladenine, 6-methylguanine (alkylation; glioma and colorectal cancer), benzo[a]pyrene diol epoxide (BPDE), pyrimidine dimers (adduct formation; smoking, exposure to industrial chemicals, exposure to UV light; lung and skin cancer), and 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, and thymine glycol (ionizing radiation damage; chronic inflammatory diseases, prostate cancer, breast cancer, and colorectal cancer). For example, 8-oxoG is a frequent product of DNA oxidation. 8-oxoG has a tendency to base pair with adenine, resulting in G>>C T·A transversion mutations. Another example is the hydrolytic deamination of cytosine and 5-methylcytosine (5-meC), which mispairs with guanine to produce uracil and thymine, respectively, leading to C>>G to T·A transition mutations if unrepaired. In another example, alkylation can generate various DNA base lesions, including 6-meG, N7-methylguanine (7-meG), or N3-methyladenine (3-meA). While 6-meG is mutagenic due to its ability to pair with thymine, 7-meG and 3-meA block replicative DNA polymerases and are therefore cytotoxic. These and many other forms of DNA base damage occur in cells multiple times each day, and only the continuous action of specialized DNA repair systems can prevent the rapid decay of genetic information. In addition to damage to nuclear DNA, mitochondrial DNA also suffers significant oxidative damage, as well as damage from alkylation, hydrolysis, and adducts.For example, oxidative damage is the most common type of damage to mitochondrial DNA, primarily because mitochondria are the major cellular source of reactive oxygen species (ROS). Furthermore, mitochondria house approximately 30% of the cellular pool of S-adenosylmethionine, which can non-enzymatically methylate DNA. Furthermore, exposure to certain agents, such as estrogen, cigarette smoke, and certain chemicals, results in preferential damage to mitochondrial DNA.
[0005] Because DNA damage and epigenetic modifications can be the earliest signs of disease states, detecting epigenetic modifications and DNA damage patterns can be useful for early detection of disease and intervention. However, detection methods have limitations. For example, with regard to methylation status, spectrophotometry can be used to indicate the overall content of modifications in target DNA, but specificity is limited. High-performance liquid chromatography (HPLC) and mass spectrometry are also often used, but they are expensive, require significant amounts of material, and reduce DNA to its constituent nucleosides or nucleotides, thus destroying sequence information for downstream analysis. Immunoprecipitation (IP) using monoclonal antibodies can enrich DNA by targeted modifications, but has been identified to have limited specificity. Restriction digest profiling utilizes fragment analysis of DNA treated with modification-sensitive restriction endonucleases, but requires large amounts of material and is limited to sequences characterized by restriction sites with known sensitivity. Bisulfite sequencing is considered the "gold standard" technique for detecting DNA methylation but has important limitations. First, this approach requires large amounts of starting material because the chemical conversion process causes extensive, nonspecific damage to DNA. Second, this method is expensive, time-consuming, and may require multiple sequencing runs. Finally, importantly, it is generally only applicable to the methylcytosine (mC) modification. While mutations that allow targeting a limited number of additional modification types (methylcytosine (mC) and hydroxymethylcytosine (hmC)) have been developed or suggested, these have low yields and still share the other limitations listed above. They are also not easily applicable to other modifications and are quite complex.
[0006] Thus, there is a need in the art for improved methods for detecting modified nucleobases in a DNA sample of interest. The present invention fulfills these needs and, as described below, provides further related advantages.
[0007] Not all of the subject matter described in the Background Art section is necessarily prior art, and it should not be assumed to be prior art merely as a result of its description in the Background Art section. Along these lines, awareness of prior art problems described in the Background Art section or related to such subject matter should not be treated as prior art unless expressly stated to be prior art. Instead, the discussion of any subject matter in the Background Art section should be treated as part of the inventor's approach to a particular problem, which may itself also be inventive. Summary of the Invention
[0008] Brief Summary of the Invention Embodiments of the present invention include the detection of modified nucleobases, such as epigenetic changes and DNA damage, in DNA samples.
[0009] In one aspect, the present invention provides a method for identifying modified nucleobases in a plurality of nucleic acids, the method comprising: providing a sample comprising a plurality of DNA templates; generating first complementary copies of the DNA templates, the generating being directed by oligonucleotide primers using a first DNA polymerase in the presence of native dNTPs, generating each complementary copy of the DNA templates such that each complementary copy comprises a native dNTP, each complementary copy hybridizing to one of the DNA templates; and subjecting the DNA templates and the first complementary copies to DNA glycosylase treatment, wherein the DNA glycosylase specifically removes modified nucleobases in the DNA templates and converts positions of the modified nucleobases to abasic sites, such that each glycosylase-converted DNA template is complementary to an unconverted complementary copy. the first complementary copy hybridizing to the modified nucleobase; generating a second complementary copy of the glycosylase-converted DNA template, wherein the generating is directed by a second DNA polymerase, the second DNA polymerase being capable of incorporating a nucleotide opposite the abasic site of the converted DNA template, wherein the nucleotide does not form a Watson-Crick base pair with the modified nucleobase; determining the nucleotide sequences of the first and second complementary copies; and comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy of each of the DNA glycosylase-converted DNA templates, thereby determining the position of the modified nucleobase in the DNA template prior to DNA glycosylase conversion.
[0010] In one embodiment, comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each DNA glycosylase-converted DNA template identifies nucleotide substitutions in the sequence of the second complementary copy compared to the first complementary copy, and the location of the nucleotide substitution identifies the location of the modified base in the DNA template. In another embodiment, the modified nucleobase is selected from 5-mC, 5-hmC, 5-fC, and 5-caC. In another embodiment, the DNA glycosylase is a monofunctional DNA glycosylase. In yet another embodiment, the monofunctional DNA glycosylase is thymine DNA glycosylase (TDG) or a variant thereof. In another embodiment, subjecting the DNA template and the first complementary copy to DNA glycosylase treatment further comprises subjecting the DNA template and the first complementary copy to treatment with a ten-eleven translocation (TET) enzyme or a variant thereof. In yet another embodiment, the ten-eleven translocation (TET) enzyme or variant thereof is ngTET. In another embodiment, the DNA glycosylase is a bifunctional DNA glycosylase. In yet another embodiment, the bifunctional DNA glycosylase is a member of the DEMETER (DME) family of DNA glycosylases or a variant thereof. In yet another embodiment, the member of the DEMETER (DME) family of DNA glycosylases or a variant thereof is a variant engineered to inactivate lyase activity. In another embodiment, the second DNA polymerase is an abasic bypass DNA polymerase. In yet another embodiment, the abasic bypass DNA polymerase is DPO4 polymerase or a variant thereof. In yet another embodiment, the DPO4 polymerase or variant thereof is a variant comprising the following mutations: M76W, K78E, E79P, Q82W, Q83G, and S86E (SEQ ID NO: 3). In another embodiment, the abasic bypass DNA polymerase incorporates dATP into a second, complementary copy of the glycosylase-converted DNA template at a position opposite the abasic site. In another embodiment, the abasic bypass DNA polymerase further comprises a third DNA polymerase, wherein the third DNA polymerase has exonuclease activity.In yet another embodiment, the third DNA polymerase is DPO1 polymerase. In another embodiment, the first DNA polymerase is a high-fidelity DNA polymerase. In another embodiment, the method further comprises treating the glycosylase-converted DNA template with a stabilizing agent prior to generating a second complementary copy of the glycosylase-converted DNA template. In another embodiment, the stabilizing agent comprises an aldehyde-reactive compound that forms a stable adduct with the abasic site. In yet another embodiment, the stabilizing agent is selected from O-hydroxylamines, acylhydrazines, tryptamines, beta-aminothiols, alkylhydrazines, hydrazino-iso-picteth-Spenglerindoles, and methylaminooxy-iso-picteth-Spenglerindoles. In another embodiment, the stabilizing agent comprises an aminoxyalkyl group that can form an oxime adduct with the abasic site. In yet another embodiment, the stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]uracil, 1-[3-(aminoxy)propyl]uracil, 1-[4-(aminoxy)butyl]uracil, 1-[5-(aminoxy)pentyl]uracil, 1-[2-(aminoxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminoxy)ethyl]-thymine. In yet another embodiment, the stabilizer is 1-[2-(amino)ethyl]uracil. In another embodiment, the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment and the step of treating the glycosylase-converted DNA template with a stabilizing agent before generating the second complementary copy are performed in the same step. In another embodiment, the DNA template is selected from the group consisting of genomic DNA, mitochondrial DNA, cell-free DNA, circulating tumor DNA, or a combination thereof. In another embodiment, the DNA template is immobilized on a solid support. In yet another embodiment, the first or second complementary copy is immobilized on a solid support.In another embodiment, determining the nucleotide sequence of the first and second complementary copies comprises synthesizing Xpandomer copies of the first and second complementary copies and passing the Xpandomer copies of the first and second complementary copies through a nanopore. In another embodiment, the DNA template comprises a first adapter joined to the 5' end of the DNA template and a second adapter joined to the 3' end of the template. In yet another embodiment, the first or second adapter is a Y adapter. In yet another embodiment, at least one of the first and second adapters comprises a unique molecular identifier barcode (UMI). In another embodiment, comparing the sequences of the first and second complementary copies comprises bioinformatically pairing sequences comprising the same unique molecular identifier barcode (UMI).
[0011] In another aspect, the present invention provides a chemoenzymatic nucleobase conversion reaction mixture comprising a DNA glycosylase enzyme, a chemical stabilizer, and a suitable buffer. In one embodiment, the chemoenzymatic nucleobase conversion reaction mixture further comprises a DNA template strand hybridized to a first complementary copy strand, wherein the DNA template strand comprises modified nucleobases and the first complementary copy strand comprises native nucleobases. In another embodiment, the chemical stabilizer comprises an aminoxyalkyl group, which can react with an abasic nucleotide containing a ring-opened aldehyde moiety to form a stable oxime adduct. In another embodiment, the chemical stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminoxy)propyl]-uracil, 1-[4-(aminoxy)butyl]-uracil, 1-[5-(aminoxy)pentyl]-uracil, 1-[2-(aminoxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminoxy)ethyl]-thymine. In another embodiment, the chemical stabilizer is selected from O-hydroxylamines, acylhydrazines, tryptamines, beta-aminothiols, alkylhydrazines, hydrazino-iso-picteth-Spenglerindole, and methylaminooxy-iso-picteth-Spenglerindole. In another embodiment, the DNA glycosylase is selected from the DNA glycosylases listed in Table 1. In yet another embodiment, the reaction mixture comprises two or more DNA glycosylases listed in Table 1. In yet another embodiment, the DNA glycosylase is TDG or a variant thereof. In another embodiment, the reaction mixture further comprises a TET enzyme.
[0012] In another aspect, the present invention provides a kit for detecting modified nucleobases in a DNA sample, comprising the above-described chemoenzymatic nucleobase conversion reaction mixture, any one of an enzyme selected from at least one of a high-fidelity DNA polymerase, an abasic-bypass DNA polymerase, and a DNA polymerase with exonuclease activity, and an appropriate mixture of dNTPs or their analogs. In one embodiment, the kit further comprises one or more buffers for the enzyme. [Brief explanation of the drawings]
[0013] [Figure 1A] 1A, 1B, 1C and 1D are schematic diagrams summarizing one embodiment of the method of the present invention. [Figure 1B] 1A, 1B, 1C and 1D are schematic diagrams summarizing one embodiment of the method of the present invention. [Figure 1C] 1A, 1B, 1C and 1D are schematic diagrams summarizing one embodiment of the method of the present invention. [Figure 1D] 1A, 1B, 1C and 1D are schematic diagrams summarizing one embodiment of the method of the present invention.
[0014] [Figure 2A] 2A and 2B are schematic diagrams showing an alternative embodiment of solid phase synthesis of primer extension products. [Figure 2B] 2A and 2B are schematic diagrams showing an alternative embodiment of solid phase synthesis of primer extension products.
[0015] [Figure 3A] 3A and 3B are chemical schemes illustrating one embodiment of the instability of an abasic site and a means to stabilize the abasic site. [Figure 3B] 3A and 3B are chemical schemes illustrating one embodiment of the instability of an abasic site and a means to stabilize the abasic site.
[0016] [Figure 4A]4A and 4B are schemes illustrating two exemplary enzymatic reactions for removing a 5-mC nucleobase to generate an abasic site in a DNA target fragment. [Figure 4B] 4A and 4B are schemes illustrating two exemplary enzymatic reactions for removing a 5-mC nucleobase to generate an abasic site in a DNA target fragment.
[0017] [Figure 5A] 5A and 5B provide an exemplary embodiment of a chemical scheme for stabilization of abasic sites in converted DNA target fragments. [Figure 5B] 5A and 5B provide an exemplary embodiment of a chemical scheme for stabilization of abasic sites in converted DNA target fragments.
[0018] [Figure 6A] 6A and 6B are schemes illustrating one embodiment of the aminoxyalkyl-mediated "hijacking" of DNA lyase activity. [Figure 6B] 6A and 6B are schemes illustrating one embodiment of the aminoxyalkyl-mediated "hijacking" of DNA lyase activity.
[0019] [Figure 7] FIG. 7 provides the chemical structures of certain exemplary nucleotide analogs for practicing the methods of the invention.
[0020] [Figure 8] FIG. 8 provides the chemical structures of other exemplary nucleotide analogs for practicing the methods of the invention.
[0021] [Figure 9] FIG. 9 provides the chemical structures of other exemplary nucleotide analogs for practicing the methods of the invention.
[0022] [Figure 10]FIG. 10 provides the chemical structures of other exemplary nucleotide analogs for practicing the methods of the present invention.
[0023] [Figure 11] FIG. 11 provides the chemical structures of certain exemplary generic nucleotide analogs for practicing the methods of the invention.
[0024] [Figure 12A] 12A and 12B are schemes illustrating one embodiment of chemoenzymatic conversion of a DNA target fragment containing a modified nucleobase of interest (5-mC) using an exemplary aminoxyalkyluracil mimic to stabilize an abasic site in the DNA target fragment, and its use for identification of the modified nucleobase in the DNA target fragment by sequencing by extension. [Figure 12B] 12A and 12B are schemes illustrating one embodiment of chemoenzymatic conversion of a DNA target fragment containing a modified nucleobase of interest (5-mC) using an exemplary aminoxyalkyluracil mimic to stabilize an abasic site in the DNA target fragment, and its use for identification of the modified nucleobase in the DNA target fragment by sequencing by extension.
[0025] [Figure 13] FIG. 13 provides the chemical structures of certain exemplary aminoxyalkyl nucleobase mimetics for practicing the methods of the invention.
[0026] [Figure 14] FIG. 14 is a gel showing the DNA products of certain chemoenzymatic nucleobase conversion reactions.
[0027] [Figure 15] FIG. 15 is a gel showing the DNA products of specific primer extension reactions of an abasic DNA template using DPO4 polymerase.
[0028] [Figure 16]FIG. 16 is a gel showing the DNA products of specific primer extension reactions of an abasic template using various DNA polymerases and their combinations.
[0029] [Figure 17A] Figures 17A and 17B are graphs showing the nucleotide incorporated by a DNA polymerase opposite the abasic site in a converted DNA template stabilized by a uracil nucleobase mimic and opposite the 5-mC in an unconverted template, respectively, as determined by DNA sequencing of the template. [Figure 17B] Figures 17A and 17B are graphs showing the nucleotide incorporated by a DNA polymerase opposite the abasic site in a converted DNA template stabilized by a uracil nucleobase mimic and opposite the 5-mC in an unconverted template, respectively, as determined by DNA sequencing of the template. DETAILED DESCRIPTION OF THE INVENTION
[0030] Detailed Description of the Invention The present invention may be more readily understood by reference to the following detailed description of preferred embodiments of the invention and the examples contained herein. Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0031] References throughout this specification to "one embodiment" or "an embodiment" and variations thereof mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0032] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents, i.e., one or more, unless the content and context clearly dictate otherwise. It should also be noted that the conjunctive terms "and" and "or" are generally used in their broadest sense to include "and / or," unless the content and context clearly dictate inclusiveness or exclusiveness, as the case may be. Thus, the use of alternatives (e.g., "or") should be understood to mean either one, both, or any combination thereof of the alternatives. Furthermore, when described herein as "and / or," the "and" and "or" constructions are intended to encompass embodiments that include all associated items or ideas, as well as one or more other alternative embodiments that include fewer than all associated items or ideas.
[0033] Unless the context requires otherwise, throughout the specification and claims that follow, the word "comprise," as well as its cognate words and variations, such as "have" and "include," and variations thereof, such as "comprises" and "comprising," are to be construed in an open and inclusive sense, e.g., "including, but not limited to." The term "consisting essentially of" limits the scope of a claim to particular materials or steps, or those that do not materially affect the basic and novel characteristics of the claimed invention.
[0034] The abbreviation "for example" is derived from the Latin exempli gratia and is used herein to indicate a non-limiting example. Thus, the abbreviation "for example" is synonymous with the term "for example." As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise; the term "X and / or Y" refers to "X" or "Y," or both "X" and "Y," and it is also understood that the letter "s" following a noun refers to both the plural and the singular form of the noun. Furthermore, when features or aspects of the invention are described in terms of a Markush group, the invention is intended to encompass and be described in terms of any individual members and any subgroups of members of the Markush group, as will be recognized by those skilled in the art, and applicant reserves the right to amend this application or claims to specifically refer to any individual member or any subgroup of members of the Markush group.
[0035] Any headings used within this document are merely utilized to facilitate the reader's review thereof and should in no way be construed as limiting the scope of the invention or claims. Accordingly, the headings and abstracts of the disclosure provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.
[0036] Where a range of values is provided herein, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value in the stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the invention.
[0037] For example, any concentration range, percentage range, ratio range, or integer range provided herein should be understood to include any integer within the stated range, and, where appropriate, fractions thereof (e.g., tenths and hundredths of integers), unless specifically stated otherwise. Also, any numerical range described herein for any physical characteristic, such as polymer subunits, size, or thickness, should be understood to include any integer within the stated range, unless specifically stated otherwise. As used herein, the term "about" means ±20% of the indicated range, value, or structure, unless specifically stated otherwise.
[0038] Methods for detecting modified DNA nucleobases For example, methods and compositions for detecting modified nucleobases in DNA samples, which reflect epigenetic modifications and DNA damage, are described herein. The method involves enzymatically removing the modified nucleobases of interest in DNA target fragments to generate free base sites at each position where the modified nucleobases of interest occur in the nucleic acid sequence of the DNA target fragment. The positions of the free base sites can be identified by DNA sequencing methods as described herein. The method of the present invention also includes a workflow for generating a first complementary copy and a second complementary copy (i.e., a first daughter strand and a second daughter strand) of the DNA target fragment template. The first complementary copy is generated before the enzymatic removal of the modified nucleobases of interest, and the second complementary copy is generated after the enzymatic removal of the modified nucleobases of interest. Thus, the first and second complementary copies encode the genetic information and, for example, epigenetic information of the DNA target fragment, respectively. The sequence information obtained from the first and second complementary copies can be compared to identify the position of the modified nucleobases of interest in the nucleic acid sequence of the original DNA target fragment.
[0039] overview According to the methods described herein, the modified nucleobase of interest may include, but is not necessarily limited to, one or more of the following: 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxycytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (*-oxoG), uracil (U), 6-methyladenine (6-mA), 8-oxoadenine, O-6-methylguanine, 1-methyladenine, O-4-methylthymine, 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, or thymine dimer.In some examples, any combination of these exemplary nucleobases and other modified nucleobases can be detected by the methods of the present invention.
[0040] In one embodiment, a method for detecting modified nucleobases, e.g., modified nucleobases of interest, in a nucleic acid sample is provided. A schematic diagram of one exemplary method is shown in Figures 1A-1D. The method may include step A: obtaining a sample of nucleic acid and fragmenting the nucleic acid to produce a sample containing DNA target fragments 100. As used herein, the term "target fragment" means that the corresponding nucleic acid fragment is derived from a biological sample and is the target of the methods described herein, which interrogate nucleic acid sequences for the presence of specific modified nucleobases. In this non-limiting depiction, the modified nucleobase of interest is methylated cytosine (5-mC), and the DNA target fragment is a double-stranded nucleic acid fragment. Here, the strands of the DNA target fragment are designated "parent (+)" 100a (i.e., sense strand) and "parent (-)" 100b (i.e., antisense strand). For simplicity, each strand of the DNA target fragment in this example contains a single 5-mC residue.
[0041] In some examples, the DNA target fragment may be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof obtained from a biological sample.
[0042] In certain embodiments, the method may then include step B, in which adapters 101 and 103 are ligated (i.e., spliced) to the 5' and 3' ends of the DNA target fragments to produce adapter-ligated DNA target fragments. The adapters may include a region of double-stranded DNA and a region of single-stranded DNA. In the example shown in FIG. 1A, the adapter is a Y-adapter (YAD), which includes two regions: a double-stranded region and a single-stranded DNA. The adapter may also include sequences or other features that mediate downstream steps in the workflow. For example, in certain embodiments, the adapter may include sequences for immobilizing the adapter-ligated DNA target fragments on a solid support, sequences for hybridization of oligonucleotide primer(s), sequences that enable bioinformatic analysis of DNA sequence information (e.g., unique molecular identifier barcodes [UMI], sample identifiers [SID]), chemical moieties for solid-phase immobilization, etc. In certain embodiments, the structures of adapters 101 and 103 may be identical or different, depending on the particular application.
[0043] The method may then include a step C of denaturing the DNA target fragment to produce a single-stranded parent (+) strand 105a and a single-stranded parent (-) strand 105b. As used herein, the terms "target" and "parent" are used interchangeably as they relate to strands of nucleic acid. Furthermore, as used herein, a single-stranded DNA target fragment may be interchangeably referred to as a "DNA template," which refers to a strand of polynucleotide to which a complementary polynucleotide can be hybridized or synthesized by a nucleic acid polymerase, for example, in a primer extension reaction.
[0044] The method may then include step D, in which a first primer extension reaction is performed. The first primer extension reaction is guided by an extension oligonucleotide (i.e., a primer) hybridized to a DNA template using a first DNA polymerase. In some instances, the extension oligonucleotide may hybridize to a region in an adapter sequence. The first primer extension reaction produces a sample of double-stranded DNA fragments, each containing a newly synthesized first complementary copy strand (i.e., first daughter strand 107b and 107b) hybridized (i.e., coupled) to a target fragment template (i.e., parent strand 105a and 105b). In some instances, the first DNA polymerase is a high-fidelity DNA polymerase. In this step, the sample of double-stranded DNA fragments is distinguished from the sample of DNA target fragments in step A in that it contains a complementary copy strand synthesized in vitro. Primer extension reaction can be carried out under conditions in which the resulting complementary copy strand is a "native" strand, in that it does not contain the desired modified nucleobase(s) present in the target strand. For example, in this figure, the first complementary copy strand incorporates a native cytosine residue in the position of the methylated cytosine residue in the corresponding target strand. As used herein, the term "native" refers to a nucleobase, nucleotide, or polynucleotide that is similar to the related modified nucleobase, nucleotide, or polynucleotide, except for the specific modification of the modified nucleobase, nucleotide, or polynucleotide. Thus, in certain embodiments, each modified nucleobase, nucleotide, or polynucleotide can have a similar native nucleobase, nucleotide, or polynucleotide, and vice versa.
[0045] In some instances, the target fragment template is immobilized on a solid support prior to step (D) of performing the first primer extension reaction, as shown in Figure 2A. As shown here, the newly synthesized complementary copy strand is not immobilized on the solid support and can be physically separated from the immobilized template strand upon denaturation of the double-stranded DNA fragment. In other instances, as shown in Figure 2B, an oligonucleotide complementary to the template strand, e.g., an adapter sequence, can be immobilized on the solid support and "captured" by hybridization. Following capture of the target fragment, a first primer extension reaction can be performed using the hybridized oligonucleotide as a primer to produce a first complementary copy strand, also immobilized on the solid support. In this case, denaturation of the resulting double-stranded DNA fragment releases the template strand from the solid support while retaining the complementary copy.
[0046] The method may then include step E, treating the sample of double-stranded DNA fragments with a DNA glycosylase enzyme capable of removing a modified nucleobase of interest (e.g., 5-mC in this depiction). As used herein, the term "remove" refers to cleaving the N-glycosidic bond between the sugar and base of a nucleotide. Removal of the modified nucleobase of interest generates an abasic site (e.g., an apurinic or apyrimidinic AP site) in the DNA target fragment at each position of the modified nucleobase of interest. In some instances, two or more DNA glycosylases or other enzyme(s) may be used to generate the abasic sites. DNA glycosylase enzymes can also be engineered to inactivate functions that are not suitable for the desired result. For example, the lyase activity of the enzyme can be selectively inactivated while the glycosylase activity is maintained. Notably, the first complementary copy strand remains resistant to DNA glycosylase treatment, and such native nucleobase sites are not converted to abasic sites.
[0047] As used herein, the term "converted," when used in reference to a DNA target fragment, refers to a DNA target fragment or a portion thereof that has been treated under conditions sufficient to remove the modified nucleobase of interest to generate an abasic site in an otherwise continuous polynucleotide chain. This process may also be referred to herein as "conversion of a modified nucleobase to an abasic site." In contrast to prior art methods of epigenetic detection that rely on chemical conversion of native nucleobases to distinguish between native and modified bases (e.g., bisulfite conversion of natural cytosine), the method of the present invention offers the advantage of selective enzymatic removal of modified nucleobases, while leaving the native nucleobase unchanged. Therefore, the overall damage to the DNA target fragment is less extensive than in methods based on bisulfite conversion, and the complexity of the genetic code is not as dramatically reduced.
[0048] The method may then include step F, in which the sample of double-stranded DNA fragments is denatured to release the converted parent DNA template strands 105a and 105b. As described, in some instances, the DNA template is immobilized on a solid support prior to the first primer extension reaction, allowing it to be separated from the first complementary copy strand, which is partitioned into solution after denaturation. In other instances, the first complementary copy strand is retained on the solid support, allowing the DNA target fragments to partition into solution after denaturation. After step F, the DNA template strand and the first complementary copy strand are no longer coupled. As used herein, the term "coupled" is well known to those skilled in the art and refers to a process in which two nucleic acid strands are held together. Coupling is achieved, for example, by the formation of hydrogen bonds between the DNA template strand and its complementary copy strand. Therefore, in the context of the present disclosure, the terms "hybridized" and "hybridization" are defined as "coupled" and "coupling," respectively. That is, for example, a complementary copy of a DNA template may be coupled to the template by hybridization.
[0049] The method may then include step G, in which a second primer extension reaction is performed. The second primer extension reaction is guided, for example, by an extension oligonucleotide that hybridizes to a region in the adapter sequence of the DNA template using a second DNA polymerase to produce second complementary copies 109a and 109b of the DNA target strand template. The second DNA polymerase is selected for its ability to synthesize a complementary copy strand through (e.g., through and beyond) the location of the abasic site in the target fragment template. DNA polymerases that exhibit this property are sometimes referred to as "bypass polymerases" and may include translesion DNA polymerases. As discussed with reference to step D, in certain embodiments, either the DNA template strand or the second complementary strand may be selectively immobilized on a solid support, allowing for purification of the second complementary strand from the template strand.
[0050] According to the present invention, the nucleobase incorporated into the daughter strand at the position opposite the base-free site of the parent template does not form a standard Watson-Crick base pair with the original unconverted nucleobase under the extension conditions used in this step. In the example shown in Figure 1C, the nucleotide incorporated opposite the base-free site of the template strand is identified as "not G" because G normally base pairs with 5-mC (in this case, the desired converted nucleobase). Therefore, "not G" can be any nucleobase other than G, such as adenine (A), cytosine (C), or thymine (T).
[0051] In some instances, the second DNA polymerase can be selected based on its substrate specificity and preferred nucleotide incorporation at the position opposite the abasic site of the converted template strand. For example, a DNA polymerase known to preferentially incorporate an abasic site opposite dATP into the template, since "A" does not normally base pair with "C," is suitable for detecting modified cytosines in target fragments. As discussed further herein, it is known in the art that some DNA polymerases exhibit specific preferences for nucleotide incorporation at abasic sites.
[0052] The method can then comprise step H of determining the nucleotide sequence of the first and second complementary copy strands. A variety of sequencing platforms and methodologies are suitable for carrying out the present invention. In one example, the sequencing method is nanopore-based "Sequencing by Expansion" (SBX®), see for example, applicant's U.S. Patent Nos. 7,939,259 and 10,301,345 and published applications WO2020 / 172,479 and WO2020 / 236,526 (which are incorporated herein by reference in their entirety).
[0053] The method may then include step I, comparing the sequence reads of the first and second complementary copy strands to identify the location of the modified nucleobase of interest in the original DNA target fragment (e.g., using art-recognized bioinformatics analysis tools). The first complementary strand is used as a reference sequence because it encodes the genetic information of the DNA target fragment. In contrast, the second complementary strand encodes the epigenetic information of the DNA target fragment. A difference in the sequences of the first and second complementary copy strands at a specific position (e.g., a base substitution) indicates the location of the modified nucleobase of interest in the sequence of the DNA target fragment. In the example shown in Figure 1D, the detection of a "not G" in the second complementary strand at the same position as a "G" in the first complementary strand indicates that the DNA target fragment originally contained a 5-mC residue at this position on the opposite strand.
[0054] In some instances, the methods of the present invention may include an additional step to stabilize the abasic site generated in the converted DNA template before generating a second, complementary copy (Step G). As summarized in the diagram in Figure 3A, abasic sites in DNA are known in the art to exist as an equilibrated mixture of two structural forms: (I) a closed-ring hemiacetal (301) and (II) an open-ring aldehyde alcohol (303). The open-ring aldehyde 303 is a highly reactive compound. Thus, the abasic residue in the DNA fragment is converted to a strand break via a β-elimination reaction, in which the 3' phosphodiester bond of the open-ring aldehyde form is hydrolyzed to generate a 3'-terminal unsaturated sugar and a terminal 5' phosphate. The presence of nucleophilic molecules in the environment, including thiols, amines, polyamines, and basic proteins, further promotes this undesirable reaction. As is readily apparent to those skilled in the art, strand breaks are detrimental in that they prevent replication of the target fragment and result in information loss.
[0055] To overcome this problem, in certain embodiments, the methods disclosed herein may include the use of a stabilizer that prevents chemical decomposition of the ring-opened aldehyde 303 and subsequent chain scission. As shown in Figure 3B, in one example, the stabilizer may be a chemical that covalently reacts with the abasic site to form a stable adduct 305. As used herein, the term "adduct" refers to the product of the direct covalent addition of two or more different molecules, resulting in a single reaction product that contains all atoms of all components and is therefore a distinct molecular species. In other examples, the stabilizer may be a soluble buffer additive or other physicochemical reaction condition that does not covalently react with the abasic site.
[0056] Further details regarding the above methods and embodiments are provided below.
[0057] The practice of the present invention employs, unless otherwise indicated, conventional techniques of molecular biology, microbiology, recombinant DNA, and the like, which are within the skill of the art. Such techniques are fully explained in the literature. See, e.g., Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, Second Edition (1989); OLIGONUCLEOTIDE SYNTHESIS (M.J. Gait Ed., 1984); METHODS IN ENZYMOLOGY (Academic Press, Inc.) series; CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F.M.A. Usubel, R. Brent, R.E. Kingston, D.D. Moore, J.G. Siedman, J.A. Smith, and K. Struhl, eds., 1987).
[0058] DNA sample / DNA target fragment In one embodiment, the DNA is obtained or provided from biological sample.The DNA obtained or provided from biological sample can be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof.
[0059] DNA sample can be obtained from patient or subject, from environmental sample, or from organism of interest.In embodiments, DNA sample is extracted, purified, or derived from cell or cell cluster, body fluid, tissue sample, organ, and / or organelle.In some embodiments, sample DNA is total genomic DNA.
[0060] In some instances, genomic DNA and mitochondrial DNA may be obtained separately from the same biological sample or source. Many different methods and techniques are available for isolating genomic DNA and mitochondrial DNA. Generally, such methods involve disruption and lysis of the starting material, followed by removal of proteins and other contaminants, and finally recovery of DNA. Protein removal can be achieved, for example, by digestion with proteinase K, followed by salting out, organic extraction, gradient separation, or binding of DNA to a solid support (either anion exchange or silica techniques). Mitochondrial DNA can be similarly isolated after initial isolation of mitochondria. DNA can be recovered by precipitation with ethanol or isopropanol. Commercially available kits are also available for isolating nuclear DNA or mitochondrial DNA. The choice of method depends on many factors, including, for example, the amount of sample, the required amount and molecular weight of DNA, the purity required for downstream applications, and time and cost.
[0061] The disclosed methods, in certain embodiments, utilize mild enzymatic and chemical reactions that avoid the substantial degradation associated with methods such as bisulfite sequencing. Thus, the methods are useful for the analysis of low-input samples, such as circulating cell-free DNA, circulating tumor DNA, and single-cell analysis.
[0062] In some embodiments, DNA sample is circulating cell-free DNA (cfDNA), which is the DNA found in blood and not present in cells.cfDNA can be isolated from blood or plasma using methods known in the art.For example, commercially available kits are available for isolating cfDNA, including circulating DNA kit (Qiagen).DNA sample can be obtained from enrichment process, including but not limited to antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0063] In some instances, the isolated DNA is fragmented into multiple shorter double-stranded DNA target fragments. Generally, DNA fragmentation can be performed physically or enzymatically.
[0064] For example, physical fragmentation can be achieved by acoustic shearing, sonication, microwave irradiation, or hydrodynamic shearing. Acoustic shearing and sonication are the primary physical methods used to shear DNA. For example, the Covaris® instrument (Woburn, Massachusetts) is an acoustic device for shearing DNA into fragments ranging from 100 bp to 5 kb. Covaris also manufactures tubes (gTubes) that process samples ranging from 6 to 20 kb for Mate-Pair libraries. Another example is the Bioruptor® (Denville, NJ), an ultrasonicator utilized to shear chromatin, DNA, and disrupt tissue. It can shear small amounts of DNA to lengths ranging from 150 bp to 1 kb. Digilab's Hydroshear® (Marlborough, MA) is another example, utilizing hydrodynamic forces to shear DNA. Nebulizers, such as those manufactured by Life Technologies (Grand Island, NY), can also be used to atomize liquids using compressed air, shearing DNA into fragments ranging from 100 bp to 3 kb in a matter of seconds. Nebulization can result in sample loss, and therefore, in some instances, may not be a desirable fragmentation method for samples with limited volume. Sonication and acoustic shearing may be better fragmentation methods for smaller sample volumes, as the total amount of DNA from the sample may be more efficiently retained. Other physical fragmentation devices and methods, known or to be developed, may also be used.
[0065] DNA can be fragmented using various enzymatic methods. For example, DNA can be treated with DNase I or a combination of maltose-binding protein (MBP)-T7 Endo I and a nonspecific nuclease such as Vibrio vulnificus nuclease (Vvn). The combination of the nonspecific nuclease and T7 Endo acts synergistically to generate nonspecific nicks, neutralize the nicks, and generate fragments with 8 nucleotides or less dissociated from the nick site. In another example, DNA can be treated with NEBNext® dsDNA Fragmentase® (NEB, Ipswich, MA). NEBNext® dsDNA Fragmentase generates dsDNA breaks in a time-dependent manner to yield DNA fragments of 50–1,000 bp, depending on the reaction time. NEBNext dsDNA Fragmentase contains two enzymes: one randomly generates nicks in dsDNA, and the other recognizes the nicked site and cleaves the DNA strand opposite the nick, generating dsDNA breaks. The resulting DNA fragments contain short overhangs, a 5'-phosphate and a 3'-hydroxyl group.
[0066] In some examples, the DNA sample is fragmented into a specific size range of target fragments. For example, the DNA sample can be fragmented into fragments of about 25-100 bp, about 25-150 bp, about 50-200 bp, about 25-200 bp, about 50-250 bp, about 25-250 bp, about 50-300 bp, about 25-300 bp, about 50-500 bp, about 25-500 bp, about 150-250 bp, about 100-500 bp, about 200-800 bp, about 500-1300 bp, about 750-2500 bp, about 1000-2800 bp, about 500-3000 bp, about 800-5000 bp, or any other size range within these ranges. For example, the DNA sample can be fragmented into fragments of about 50-250 bp. In some instances, the fragments may be larger or smaller than about 25 bp.
[0067] A DNA target fragment may be any DNA fragment derived from a biological sample that has a sequence of interest, which may or may not contain epigenetic modifications or DNA damage to one or more nucleic acid bases. In some embodiments, the DNA target fragment may contain cytosine modifications (i.e., 5-mC, 5-hmC, 5-fC, and / or 5-caC). A DNA target fragment may be a single DNA molecule in a sample, or may be the entire population of DNA molecules in a sample (or a subset thereof) that have, for example, cytosine modifications. A DNA target fragment may contain multiple DNA sequences, so that the methods described herein can be used to generate a library of DNA target fragments that can be analyzed individually (e.g., by determining the sequence of each target) or in groups (e.g., by multiplexed DNA sequencing).
[0068] In embodiments, the methods described herein include adding an adapter DNA molecule to a double-stranded DNA target fragment. The adapter DNA or DNA linker is a short, chemically synthesized, single-stranded or double-stranded oligonucleotide that can be ligated to one or both ends of another DNA molecule. The double-stranded adapter can be synthesized so that each end of the adapter has a blunt end or a 5' or 3' overhang (i.e., sticky end). The DNA adapter is ligated to the DNA target fragment to provide sequences for, for example, a primer extension reaction and a sequencing reaction using complementary primers and / or bioinformatics analysis (e.g., clustering related sequences into families based on shared unique molecular identifier barcodes, UMIs).
[0069] Prior to adapter ligation, the ends of the DNA fragments can be prepared for ligation. For example, by end repair and creating blunt ends with 5' phosphate groups. Fragmented DNA can be blunt-ended by several methods known to those skilled in the art. In a specific method, the ends of the fragmented DNA are "polished" with T4 DNA polymerase and Klenow polymerase, a procedure well known to those skilled in the art, and then phosphorylated with polynucleotide kinase enzyme. Then, using Taq polymerase or Klenow exo-minus polymerase enzyme, a single "A" deoxynucleotide is added to both 3' ends of the DNA molecule, creating a one-base 3' overhang complementary to the one-base 3" T' overhang at the double-stranded end of the adapter.
[0070] In some instances, an adapter may contain two partially complementary oligonucleotides that hybridize to form a region of double-stranded sequence but also retain a region of single-stranded, non-hybridizing sequence. The region of single-stranded sequence may contain a "universal" oligonucleotide binding sequence, which allows all target fragments in a library to bind to the same oligonucleotide, which may be a capture oligonucleotide, allowing the target fragments to be localized to a solid support; an oligonucleotide primer for a primer extension reaction; a PCR primer; a sequencing primer; or a combination thereof. In certain cases, an adapter may contain two regions of single-stranded, non-hybridizing sequence (i.e., a first 5' single-stranded region and a second 3' single-stranded region). This configuration is known in the art as a "Y" adapter. The first and second single-stranded regions of the Y adapter are not complementary and may contain different primer hybridization sequences and other features.
[0071] The two single-stranded regions of the adapter typically comprise at least 10, 15, or 20 consecutive nucleotides on each strand. The lower limit of the length of the single-stranded region is typically determined by the need to provide a sequence suitable for function, such as primer binding for primer extension, PCR, and / or sequencing. Theoretically, there is no upper limit to the length of the single-stranded region, except that it is generally advantageous to minimize the total length of the adapter, for example, to facilitate separation of unbound adapters from adapter-ligated double-stranded DNA target fragments after the ligation step. Therefore, it is preferred that the single-stranded region be less than 50, 40, 30, or 25 consecutive nucleotides long on each strand.
[0072] The double-stranded region of an adapter is a short double-stranded region, typically containing five or more consecutive base pairs, formed by the annealing of two partially complementary polynucleotide strands. Generally, it is advantageous for the double-stranded region to be as short as possible without losing functionality. In this context, "functional" means that the double-stranded region forms a stable duplex under standard reaction conditions for enzyme-catalyzed nucleic acid ligation reactions.
[0073] The exact nucleotide sequence of the adapter is generally not a material of the present invention, and can be selected by the user so that the desired sequence element is ultimately included in the consensus sequence of a library of adapter-ligated double-stranded DNA target fragments. For example, additional sequence elements can be included to provide binding sites for primers that will ultimately be used to sequence complementary copy strands of the DNA target fragments. The adapters can further include "tag" sequences, unique molecular identifiers (UMIs), and / or sample identifier sequences that can be used to tag, track, and distinguish target fragments and their complementary copies from a particular source. The general characteristics and uses of such sequences are well known in the art.
[0074] The terminus of the single-stranded region of the adapter may be biotinylated or may have another functionality that allows for capture or immobilization on a surface, such as a solid support. Alternative functionality other than biotin is known in the art, as described, for example, in applicant's published patent application WO2020 / 172479, entitled "Methods and Devices for Solid-Phase Synthesis of Xpandomers for use in Single Molecule Sequencing," which is incorporated herein by reference in its entirety.
[0075] "Ligation" of an adapter to the 5' and 3' ends of each fragmented double-stranded nucleic acid target fragment involves joining the two polynucleotide strands of the adapter to the double-stranded target polynucleotide such that a covalent bond is formed between both strands of the two double-stranded molecules. Preferably, such covalent bonding occurs via the formation of a phosphodiester bond between the two polynucleotide strands, although other covalent bonding means (e.g., non-phosphodiester backbone linkages) may also be used. However, it is essential that the covalent linkage formed in the ligation reaction allows polymerase readthrough so that the resulting construct can be copied in a primer extension reaction using a primer that binds to a sequence in the region of the adapter-target construct derived from the adapter molecule.
[0076] In some instances, the adaptor and DNA target fragment may be incubated with a ligase to covalently link the adaptor and DNA target fragment. Ligase catalyzes the formation of a phosphodiester bond between juxtaposed 5' phosphate and 3' hydroxyl ends in double-stranded DNA or RNA. The enzyme joins blunt and cohesive ends and repairs single-stranded nicks in double-stranded DNA. An exemplary ligase is T4 ligase, the enzyme most frequently used for cloning. Another ligase that can be used is Escherichia coli (E. coli) DNA ligase, which preferentially ligates cohesive double-stranded DNA ends but is also active on blunt-ended DNA in the presence of Ficoll or polyethylene glycol. Another ligase that can be used is DNA ligase Ilia, which is known to function in mitochondria.
[0077] Before the adaptor-target construct is further processed, the product of the ligation reaction may be subjected to a purification step to remove unbound adaptor molecules.
[0078] Ligation of adapters to both free ends of double-stranded DNA target fragments generates a pool of adapter-ligated double-stranded DNA target fragments bearing adapters at the 5' and 3' ends of the target.
[0079] There are several standard methods for separating the strands of adapter-ligated double-stranded DNA target fragments by denaturation, including thermal or chemical denaturation with either 100 mM sodium hydroxide or formamide solutions. The pH of the solution of single-stranded DNA fragments can be neutralized by adjusting the pH with an appropriate solution of acid or by buffer exchange through a size-exclusion chromatography column, preferably pre-equilibrated in a buffer solution.
[0080] First complementary copy strand In the embodiments disclosed herein, a single-stranded DNA target fragment (i.e., the parent strand) provides a template for synthesizing a first complementary copy of the target fragment (i.e., the first daughter strand) via a primer extension reaction. The term "primer extension reaction" is used interchangeably with the term "nucleic acid polymerization reaction" herein and refers to an in vitro method for creating a new strand of nucleic acid or extending an existing nucleic acid in a template-dependent manner. The first complementary copy strand is synthesized by extending an oligonucleotide primer with a first DNA polymerase such that the first complementary copy of the template strand is extended in the 3' direction of the oligonucleotide primer.
[0081] In the embodiment where DNA target fragment is double-stranded, one or both strands can serve as template for primer extension reaction.For example, when one strand (" sense " strand) serves as template, it generates a complementary copy that is complementary to sense strand.Similarly, when antisense strand serves as template, it generates a complementary copy that is complementary to antisense strand.When both strands serve as template, it generates separate complementary copies for each of sense strand and antisense strand.In a preferred embodiment, each strand of double-stranded DNA target fragment is template nucleic acid.
[0082] As used herein, the term "complementary" refers to a nucleic acid sequence capable of forming Watson-Crick base pairs. For example, the complement of a first sequence is a sequence capable of forming Watson-Crick base pairs with the first sequence. The term "complementary" does not necessarily mean that a sequence is complementary to the entire length of its complementary strand, but the term can mean that a sequence is complementary to a portion of it. Thus, in some embodiments, complementarity encompasses sequences that are complementary along the entire length or a portion of the sequence. For example, two sequences may be complementary to each other along at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length of the sequences. Herein, the term "sequence" encompasses, but is not limited to, nucleic acid sequences, polynucleotides, oligonucleotides, probes, primers, primer-specific regions, and target-specific regions. Despite any mismatches, the two sequences should be capable of selectively hybridizing to one another under appropriate conditions.
[0083] Primer extension can be performed by any method that allows polymerase-based extension of a primer annealed (i.e., hybridized) to a single-stranded DNA target fragment. In some embodiments, simple primer extension involves adding a primer and a first DNA polymerase to the target DNA fragment under conditions that allow primer hybridization and primer extension by the polymerase. Of course, such a reaction includes the nucleotides, buffers, and other reagents required for primer extension that are known in the art. Importantly, the nucleotides included in the primer extension reaction are "native," i.e., unmodified, nucleotides; therefore, the first complementary copy strand does not contain modifications to the target nucleobase. The first complementary copy strand is generated to encode and preserve the genetic sequence of the DNA target strand.
[0084] Any number of methods are known for detecting primer extension products. In some embodiments, the primer is detectably labeled (e.g., at its 5' end or otherwise positioned so as not to interfere with 3' extension of the primer), and after primer extension, the length and / or amount of the labeled extension product is detected by detecting the label.
[0085] In certain embodiments, the primer used in primer extension reaction is annealed to the (single-stranded) primer binding sequence in the single-stranded region of adapter.The term " annealing " used in this context refers to the sequence-specific binding / hybridization of primer to the primer binding sequence in the adapter region of adapter-linked DNA target fragment under the conditions used in the primer annealing step of initial primer extension reaction.Primer annealing conditions are well known in the art (see, for example, Sambrook et al., 2001, Molecular Cloning, A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor Laboratory Press, NY; Current Protocols, eds. Ausubel et al.).
[0086] In a preferred embodiment, the first DNA polymerase is a high-fidelity DNA polymerase. The fidelity of a DNA polymerase is the result of accurate replication of the desired template. Specifically, this involves multiple steps, including the ability to read the template strand, select the appropriate nucleoside triphosphate, and insert the correct nucleotide at the 3' primer end so that Watson-Crick base pairing is maintained. In addition to effectively distinguishing between correct and incorrect nucleotide incorporation, some DNA polymerases possess 3'→5' exonuclease activity. This activity, known as "proofreading," is used to remove the incorrectly incorporated mononucleotide and then replace it with the correct nucleotide.
[0087] In certain embodiments, high-fidelity DNA polymerases suitable for practicing the present invention include KAPA HiFi DNA Polymerase commercially available from Roche Diagnostics Corp., Q5® High-Fidelity DNA Polymerase commercially available from New England Biolabs, Inc., and engineered Pfu DNA polymerases such as Pfu-X commercially available from Jena Biosciences.
[0088] solid phase synthesis In certain embodiments, the first primer extension reaction can be carried out on a solid support. Thus, in a further aspect, the present invention provides methods for solid-phase nucleic acid synthesis using adapter-ligated DNA target fragments having known sequences at their 5' and 3' ends (e.g., sequence features designed into the adapters).
[0089] The terms "solid support," "solid phase," and "substrate" are used interchangeably herein and refer to a material or group of materials having one or more surfaces that are rigid or semi-rigid. In many embodiments, at least one surface of the solid support is substantially flat, e.g., the surface of a polymeric microfluidic card or chip. In some embodiments, it may be desirable to physically separate regions of the card or chip for different reactions, e.g., with etched channels, trenches, wells, raised areas, pins, etc. According to other embodiments, the solid support(s) will take the form of insoluble beads, resins, gels, membranes, microspheres, or other geometric configurations composed, e.g., of controlled pore glass (CPG) and / or polystyrene.
[0090] The present invention encompasses solid-phase synthesis methods in which a capture moiety is immobilized on a solid support. In certain cases, the capture moiety comprises a first end covalently attached to the solid support and a second end providing a functional group capable of binding to the 5' end of a single-stranded adaptor-ligated DNA target fragment. In this case, the single-stranded DNA target fragment is immobilized on the solid support, and the complementary copy strand is not immobilized on the solid support. In other examples, the capture moiety comprises an extender oligonucleotide capable of hybridizing to the 3' end of the single-stranded adaptor-ligated target fragment. The single-stranded adaptor-ligated DNA target fragment is hybridized to the extender oligonucleotide, and a primer extension reaction is performed. In this case, only the complementary copy strand is immobilized on the solid support. These alternative solid-phase synthesis configurations are illustrated in Figures 2A and 2B.
[0091] As used herein, the term "immobilization" refers to the association, attachment, or bonding between a molecule (e.g., a linker, adapter, or oligonucleotide) and a support in a manner that provides a stable association under the conditions of extension, amplification, ligation, and other processes described herein. Such bonding can be covalent or non-covalent. Non-covalent bonding includes electrostatic, hydrophilic, and hydrophobic interactions. Covalent bonding is the formation of a covalent bond characterized by the sharing of electron pairs between atoms. Such covalent bonding can be directly between the molecule and the support, or can be formed by a cross-linking agent or by the inclusion of specific reactive groups on the support, the molecule, or both. Covalent attachment of a molecule can be achieved using a binding partner such as avidin or streptavidin immobilized on the support and non-covalent binding of a biotinylated molecule to the avidin or streptavidin. Immobilization can also involve a combination of covalent and non-covalent interactions.
[0092] Any suitable covalent attachment means known in the art can be used for these purposes. The attachment chemistry selected will depend on the nature of the solid support and any derivatization or functional groups applied to the solid support. The extended oligonucleotide may contain a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. A specific exemplary embodiment of a suitable surface chemistry includes conventional streptavidin / biotin interaction chemistry, such as functionalizing the solid support with a linker moiety containing a terminal biotin moiety. In this embodiment, the 5' end of the single-stranded DNA fragment (or oligonucleotide) is attached to the linker moiety. Attachment is mediated by a streptavidin moiety provided by the 5' end of the single-stranded DNA fragment. The linker moieties disclosed herein can be of sufficient length to link the single-stranded DNA fragment to the support so that the support does not significantly interfere with the primer extension reaction.
[0093] Alternatively, immobilization of a capture moiety or an oligonucleotide (e.g., an extended oligonucleotide) to a solid support can be achieved by covalently linking the capture oligonucleotide to the solid support via a click reaction. In this embodiment, the covalent linkage can be mediated by a maleimide-PEG-alkyne linker crosslinked to the solid support. The alkyne moiety provided by the end of the linker distal to the substrate can react with the azide group provided by the 5' end of the capture oligonucleotide. Methods for functionalizing solid supports with maleimide linker polymers are provided in applicant's published patent application WO2020 / 172479, the entire contents of which are incorporated herein by reference.
[0094] In certain cases, the link between the capture moiety and the solid support is cleavable, allowing the primer extension product to be released from the support after synthesis.Cleavable linkers and methods for cleaving such linkers are known and can be used in the provided method using the knowledge of those skilled in the art.For example, the cleavable linker can be cleaved by an enzyme, a catalyst, a chemical compound, temperature, electromagnetic radiation, or light.Optionally, the cleavable linker includes a moiety that can be hydrolyzed by beta-elimination, a moiety that can be cleaved by acid hydrolysis, an enzymatically cleavable moiety, or a photocleavable moiety.In some embodiments, a suitable cleavable moiety is a photocleavable (PC) spacer or linker phosphoramidite available from Glen Research.
[0095] Glycosylase-mediated removal of modified nucleobases In one embodiment, the method of the present invention comprises treating the double-stranded DNA product of the first primer extension reaction with a DNA glycosylase enzyme to specifically remove the modified base of interest.Many DNA glycosylases are known in the art that target a wide range of specifically modified nucleic acid bases and DNA damage elements, including sequence mismatches and a wide range of epigenetic modifications.Exemplary epigenetic modifications that can be detected by the described method include, but are not limited to, 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxycytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (oxoG), uracil, methyladenine (mA), etc.
[0096] There are two major classes of DNA glycosylases: monofunctional and bifunctional. Monofunctional glycosylases have only glycosylase activity and cleave the N-glycosidic bond that connects damaged or modified nucleobases to the sugar-phosphate backbone of DNA. Although all DNA glycosylases cleave glycosidic bonds, they differ in their base substrate specificity and their reaction mechanism; bifunctional glycosylases also have apurinic or apyrimidinic (AP) lyase activity, which allows them to cleave the phosphodiester bond of DNA at base lesions to create single-strand breaks.
[0097] A non-limiting list of exemplary DNA glycosylases useful in the methods of the invention is provided in Table 1. In some instances, one or more of the DNA glycosylases listed in Table 1 can be used in the described methods to remove a modified base of interest from a DNA target fragment. While selected DNA glycosylases are specifically identified in this disclosure, it is understood that any suitable DNA glycosylase can be used in performing the base removal step of the described methods. [Table 1]
[0098] In one embodiment, the method utilizes a DNA glycosylase that acts directly on 5-mC, i.e., a glycosylase that can hydrolyze the glycosidic bond between the 5-mC residue and the sugar-phosphate backbone. For example, a suitable DNA glycosylase that directly removes 5-mC can be a member of the DEMETER (DME) family of DNA glycosylases, such as DME, ROS1, or DMEL. The Arabidopsis DME gene encodes a 1,729-amino acid protein with a centrally located DNA glycosylase domain (amino acids 1167-1368) that contains a helix-hairpin-helix (HhH) motif. The HhH motif in DME catalyzes the removal of 5-mC (see, e.g., Choi et al., 2002, Cell 110:33-42). In certain embodiments, the DME glycosylase can be a variant that includes amino acids 1167-1368 but lacks certain other regions of the protein.
[0099] In some instances, a suitable DNA glycosylase that acts directly on 5-mC may be an ortholog of DME. As used herein, the term "ortholog" refers to one of two or more homologous gene sequences found in different species. Table 2 provides an exemplary list of DME orthologs that can be used in accordance with the present invention. [Table 2]
[0100] If the DNA glycosylase is a bifunctional enzyme, the glycosylase (e.g., DME or its orthologue) can be mutated to inactivate the lyase activity while still retaining glycosylase activity, as shown in Figure 4A. The reaction mechanisms of bifunctional DNA glycosylases are well known in the art (see, e.g., Scharer and Jiricny, 2001, Bioessays 23:270-281). In some cases, a conserved aspartic acid acquires a proton from a conserved lysine residue, attacking the C1' carbon of the deoxyribose ring, generating a covalent DNA-enzyme intermediate. A beta- or gamma-elimination reaction releases the enzyme from the DNA and cleaves one of the phosphodiester bonds. Mutant forms of DME in which the invariant aspartic acid at position 1304 or the lysine at position 1286 are altered (e.g., variants D1304N or K1286Q) have been shown to reduce DNA glycosylase activity while maintaining the structure and stability of the enzyme (see, e.g., Fromme et al. 2004 Nature 427:652-656).
[0101] Other mutations that inactivate or optimize the appropriate characteristics of DNA glycosylases are also contemplated by the present invention. For example, DNA glycosylases can be engineered to increase their stability and / or solubility. DNA glycosylases can also be engineered to optimize desired substrate specificity.
[0102] In certain embodiments, thymine DNA glycosylase (TDG) can be used to remove its known targets, 5-carboxycytosine (5-caC) and 5-formylcytosine (5-fC). In a further embodiment, as shown in Figure 4B, TDG can be used to identify the modified bases it does not specifically recognize, 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC). For example, DNA target fragments can be treated with ten-eleven translocation (TET) enzymes prior to treatment with TDG. The TET family of proteins includes three human proteins (TET1, TET2, and TET3) and is a cytosine oxygenase that catalyzes the conversion of 5-methylcytosine (5-mC) to 5-hydroxymethylcytosine (5-hmC). 5-hmC can be further oxidized to 5-formylcytosine (5-fC) and 5-carboxycytosine (5-caC) by TET proteins (see, e.g., Parker, et al. 2019. Biochemistry 58:450-467). In another example, a suitable TET enzyme can be any TET ortholog isolated from Naegleria, such as ngTET (see, e.g., Hashimoto, et al. 2014. Nature 506(7488):391-395). Thus, in certain embodiments, TDG can be used to remove any existing 5-caC and 5-fC modified bases present in DNA target fragments that have also been treated with a TET enzyme.
[0103] Other equivalent methods for altering the selective removal of modified bases are possible according to the present invention. For example, similar methods can be performed to detect the same bases as above using thymine DNA glycosylase (TDG) and uracil DNA glycosylase (UDG).
[0104] The base removal process discussed herein can be performed using purified enzymes, which can be recombinant enzymes containing heterologous tags to facilitate purification. Protein tags are well known in the art and include, for example, terminal polyhistidine tags that allow purification by immobilized metal affinity chromatography (IMAC). In certain cases, it may be desirable to include two or more protein purification steps. For example, glycosylase enzymes used in the methods disclosed herein preferably should be free of contaminating nucleic acids. In some examples, protein purification steps can include one or more of size exclusion chromatography, ion exchange chromatography, affinity chromatography, heparin adsorption chromatography, etc.
[0105] Of course, the nucleobase removal reaction includes appropriate buffers, cofactors, additives, and a sufficient amount of purified DNA glycosylase to achieve the desired base removal reaction, resulting in the removal of the desired modified nucleobase in the DNA target fragment and the generation of an abasic site. An exemplary nucleobase removal reaction is described in Example 1.
[0106] After treatment with DNA glycosylase, the double-stranded DNA fragment is asymmetrically altered. Notably, the DNA template strand lacks a nucleobase at the original modified target base position. In contrast, the first complementary copy strand remains unchanged (i.e., "unconverted") because the native nucleobase incorporated during the first primer extension reaction is resistant to glycosylation-mediated conversion to an abasic site.
[0107] Stabilization of abasic sites in DNA target fragments Advantageously, according to the methods of the present invention, the abasic sites generated in the DNA target fragments can be protected from further degradation by a stabilizing agent. In certain embodiments, a suitable stabilizing agent can be a chemical that covalently binds to the abasic site to form a stable abasic adduct. As discussed with reference to Figure 3B, certain aldehyde-reactive compounds are known to react with the ring-opened aldehyde form (II) of the abasic site to generate a stable open-ring structure, referred to herein as the abasic adduct. The abasic adduct is resistant to degradation-inducing chemical conditions, such as enzymatic activity (e.g., lyase-mediated degradation) or high pH. Some exemplary, non-limiting structural classes of aldehyde-reactive stabilizers are shown in Figures 5A and 5B and described below. Each class differs in the kinetics, stability, and size of the resulting protected adduct. The chemical properties of each abasic adduct product offer different chemoenzymatic properties with respect to the duration of stabilization and suitability as a template for extension by DNA polymerase.
[0108] As shown in Figure 5A, in one embodiment, a suitable stabilizer may be from the group of O-hydroxylamines (compounds IIIa), a class of compounds known to react with the aldehyde group of the ring-opened form of the abasic site (II) to generate highly stable oxime structures (compounds IVa) that are resistant to β-elimination by enzymatic activity (e.g., AP or dRp lyase) or high pH.
[0109] In another embodiment, suitable stabilizers may be from the group of acylhydrazines (compounds IIIb), which are a class of compounds that react with aldehydes (II) to form acylhydrazones (compounds IVb).
[0110] In another embodiment, a suitable stabilizer may be from the group of tryptamines (compound IIIc), which react with aldehydes (II) via a Pictet-Spengler ring formation reaction to form a tricyclic heterocycle (compound IVc).
[0111] As shown in Figure 5B, in another embodiment, a suitable stabilizer may be from the group of beta-aminothiols (compound IIId) (e.g., cysteine), which is a class of compounds that react with aldehydes (II) to form cyclic thiazolidines (compound IVd).
[0112] In another embodiment, suitable stabilizers may be from the group of alkylhydrazines (group IIIe), which is a class of compounds that react with aldehydes (II) to form alkylhydrazones (compounds IVe).
[0113] In another embodiment, suitable stabilizers may be from the group of hydrazino-iso-picteth-spenglerindoles (compounds IIIf) which react with the abasic aldehyde (II) form to form a tricyclic structure (compounds IVf).
[0114] In another embodiment, suitable stabilizers may be from the group of methylaminooxy-iso-picteth-spenglerindoles (group IIIg) which react with abasic aldehydes (II) to form tricyclic structures (compounds IVg).
[0115] In other examples, the stabilizer may be an agent that does not covalently react with the abasic sites in the DNA target fragment, such as a reaction additive or other physicochemical reaction conditions. The following is a non-limiting list of exemplary stabilizers: 1. Aqueous buffers lacking salt (e.g., water); 2. Various concentrations of basic buffers (e.g., ammonia, NaOH, or other hydroxide-based buffers); 3. Various concentrations of acidic buffers (e.g., acetic acid, HCl, or nitric acid-based buffers); 4. Urea; 5. Detergents (e.g., SDS, Tween®, or Triton®); 6. Solvents (e.g., acetonitrile, DMSO, formamide, DMF, or glycerol); 7. PEG and PEG variants; 8. Guanidine salts; and 9. Electric current or current pulses applied to the reaction. In some examples, any suitable combination of the aforementioned stabilizers may be used.
[0116] In certain embodiments, the chemistry described herein can be used to form stable abasic adducts during treatment of DNA target fragments with one or more monofunctional DNA glycosylases, bifunctional DNA glycosylases, or bifunctional DNA glycosylases engineered to inactivate lyase activity.
[0117] In certain embodiments, the methods of the present invention can utilize bifunctional DNA glycosylases to generate abasic sites that are stable and resistant to lyase-mediated backbone cleavage. In other words, glycosylase activity can be separated from the lyase activity of the bifunctional glycosylase by chemically "knockouting" the latter. In some embodiments, this can be achieved by including one or more abasic stabilizing agents disclosed herein in the glycosylase reaction. As described, the stabilizers form stable adducts at the abasic sites after removal of the modified nucleobase. Such abasic adducts are resistant to further lyase activity, preventing strand removal at these sites. This phenomenon is referred to herein as biochemical knockout or "hijacking" of DNA lyase activity.
[0118] The biochemical hijacking of DNA lyase activity is shown in simplified form in Figures 6A and 6B. Figure 6A shows the native activity of an exemplary bifunctional DNA glycosylase acting on 5-mC (e.g., DEMETER). After cleaving the N-glycosidic bond to release the methylated base, the enzyme forms a Schiff base intermediate (I) with the open ribose moiety and cleaves the phosphodiester bond in the DNA backbone via a β-elimination reaction to produce strand breaks (II). Figure 6B shows knockout of lyase activity with an aminoxyalkyl compound. As used herein, the term "aminoxyalkyl" is used to denote an O-alkylated derivative of hydroxylamine, a structure with the general formula H2N-OR, where R is an alkyl group. Here, an exemplary aminoxyalkyl, shown as "H2N-OR," is added during processing of a DNA substrate by a DNA glycosylase. Following enzyme-mediated cleavage of the N-glycosidic bond to release the modified nucleobase, the aminoxyalkyl reacts with the abasic site (I) to form a stable adduct (III) that prevents the enzyme from further interacting with the DNA substrate and, for example, cleaving the phosphodiester backbone.
[0119] Second complementary copy strand The method described herein includes carrying out a second primer extension reaction to generate a second complementary copy of the parent DNA template (i.e., a second daughter strand). This step is carried out after enzymatic removal of the modified nucleobase. Thus, the second complementary copy of the DNA template retains at least part of the epigenetic information encoded in the original DNA target fragment.
[0120] After glycosylase treatment, the asymmetrically altered DNA fragments are denatured using any suitable art-recognized method, including acid-base denaturation (e.g., using acetic acid, HCl, or nitric acid), base denaturation (e.g., using NaOH), solvent-based denaturation (e.g., using DMSO, formamide, guanidine, sodium salicylate, propylene glycol, or urea), or physical denaturation (e.g., using heat, beads, sonication, or radiation). The resulting single-stranded DNA template strand is then purified from the first complementary target strand. Purification of the population of converted template strands is facilitated by the solid-phase synthesis method described herein, in which one of the two populations of parent and daughter strands is selectively immobilized on a solid support.
[0121] The second primer extension reaction is directed by an extension oligonucleotide that hybridizes to the DNA target template using a second DNA polymerase to produce a second double-stranded DNA fragment containing a second complementary copy strand hybridized to the parent template strand. The second primer extension reaction can be performed on a solid support, as described herein, where either the parent template strand or the second daughter strand is selectively immobilized on the support.
[0122] The second DNA polymerase is selected for its ability to synthesize a second, complementary copy beyond the location of the abasic site in the converted parent template. DNA polymerases that exhibit this property are known in the art and are referred to, for example, as "bypass" or "translesion" polymerases.
[0123] In some instances, the second DNA polymerase can be selected based on its activity of preferentially incorporating a specific nucleotide opposite the abasic site in the template.The objective of the present invention is to generate a second complementary copy strand so that the nucleobase incorporated into the opposite abasic site in the template does not form Watson base pairs or Crick base pairs with the modified nucleobase previously removed from the template.For example, in some instances, the modified target base is 5-mC.In this case, the second DNA polymerase is selected based on its preference for incorporating any nucleotide except dGTP (i.e., "not G") opposite the position where 5-mC is converted into an abasic site; for example, the polymerase can preferentially incorporate dATP, dTTP, or dCTP at these sites.
[0124] It is known in the art that abasic sites represent the most frequent DNA damage in genomes, have high mutagenic potential, and lead to mutations commonly seen in human cancers.Although these damages lack genetic information, it has been observed that adenine is the most efficiently inserted nucleobase during bypass of abasic sites by DNA polymerase, a phenomenon known as the "A-rule."A strong preference for adenine (i.e., dATP) incorporation by DNA polymerases from family A (including human DNA polymerase γ and θ) and family B (including human DNA polymerase α, ε, and δ) has been observed (see, for example, Obeid, et al. 2010. EMBO J. 29(10):1738-1747).In a preferred embodiment of the present invention, the second DNA polymerase preferably incorporates A opposite the abasic site into the template, especially when the desired modified nucleobase is a derivative of C (e.g., 5-mC).
[0125] A non-limiting list of exemplary second DNA polymerases is shown in Table 3. [Table 3]
[0126] In some instances, the second DNA polymerase may comprise a mixture of two or more DNA polymerases. For example, the mixture may comprise a DNA polymerase that can incorporate a nucleotide opposite the abasic site but cannot further extend the daughter strand, and another DNA polymerase that has the ability to extend the daughter strand beyond the abasic site of the parent strand. In another example, the mixture may comprise a DNA polymerase with exonuclease activity. The combination of a bypass polymerase (e.g., DPO4 or a variant thereof) with a polymerase with exonuclease activity (e.g., DPO1) may provide several advantages. For example, the exonuclease may provide error-correcting activity, and the combination may result in more efficient and accurate incorporation of the desired nucleotide, for example, by minimizing polymerase stalls and errors.
[0127] In some instances, the substrate preference of a bypassing DNA polymerase at an abasic site can be optimized or directed by additional methods of the invention. For example, the DNA polymerase can be an engineered variant with mutations that increase its bypassing activity or preference for incorporating a specific nucleotide opposite the abasic site.
[0128] In one embodiment, the engineered variant is a variant of DPO4 DNA polymerase (SEQ ID NO: 1). DPO4 is a DNA polymerase naturally expressed by the archaeon Sulfolobus solfataricus and is a Y-family DNA polymerase that functions in the replication of damaged DNA through a process commonly known as translesion synthesis (TLS). Advantages of DPO4 include its monomeric structure, open conformation, lack of an exonuclease domain, and ability to bypass abasic sites. The crystal structure of DPO4 is available to guide protein engineering; see, e.g., Ling et al. (2001) "Crystal Structure of a Y-Family DNA Polymerase in Action: A Mechanism for Error-Prone and Lesion-Bypass Replication" Cell 107:91-102. As described in the art, the present inventors have engineered thousands of variants of DPO4 that are optimized for, for example, the ability to utilize non-conventional nucleotide analogs as substrates. A non-limiting list of DPO4 variants and screening methods that can be used in accordance with the present invention is disclosed in applicant's issued U.S. Patent Nos. 11,299,725, 11,530,392, and 11,708,566, the contents of which are incorporated herein by reference in their entireties.
[0129] We previously identified a region of DPO4 polymerase corresponding to amino acids 76-86 that was a key target for modifying and optimizing the polymerase's substrate specificity. Therefore, several variants with mutations in this region were screened for abasic bypass activity via dATP incorporation in an otherwise wild-type background. From this screen, one particular DPO4 polymerase variant was identified that exhibited robust abasic bypass activity, referred to herein as "C9110." This variant contains the following mutations compared to the wild-type polymerase: M76W_K78E_E79P_Q82W_Q83G_S86E and a deletion of amino acids 341-352 (SEQ ID NO: 3).
[0130] Nucleotide Analogues In certain cases, the substrate preference of a bypass DNA polymerase can be modified or directed by utilizing an alternative nucleotide (i.e., a nucleotide analog) in the second primer extension reaction. For example, if the modified nucleobase of interest is 5-mC, the primer extension reaction can include an analog of dATP, a specific example of which is shown in Figure 7. For example, the dATP can be one or more analogs (A) of 7-deaza dATP, such as DAP (diaminopurine), an alkynyl C8, C10, phenyl, or other 7-position substituents. Other exemplary dATP analogs include 7-deaza dATP with an iodo group (B) or bromo group attached to the C-7 atom analog (C), or a chloro group (D) attached to the C-2 atom analog (D). In another example, dATP can be modified with a 6-substituent such as N6-methyl dATP, analog (A), N6 aminohexa, analog (B), or 8-bromo group, analog (C), as shown in Figure 8. In one embodiment, N6-methyl dATP is utilized in the second primer extension reaction.
[0131] Designed Nucleotide Analogues In further aspects, the methods of the present invention may involve nucleotide analogs whose nucleobases are designed to introduce specific structural and / or chemical features that promote incorporation by bypassing DNA polymerases. Exemplary nucleobase features include an overall geometry that spatially matches the empty "pocket" left by removal of the nucleobase. For example, nucleotide analogs with the size and geometry of two bases (e.g., base pairs) may be advantageous. Other beneficial features may include an overall increase in hydrophobicity or the introduction of a moiety known to enhance incorporation by bypassing polymerases, such as spermine. In certain embodiments, designed nucleotide analogs may contain two or more such features; for example, they may contain both a polymerase-enhancing feature and a "bulky" hydrophobic feature.
[0132] Certain exemplary designed nucleotide analogs include, but are not limited to, the alkyl analogs, N6-ethyl-2'-dATP, analog (A), 2-methyl-2'-dATP, analog (B), 2-ethyl-2'-dATP, analog (C), and protected analogs, N6-benzoyl-2'dATP, analog (D), and N6-phenoxyacetal-2'dATP, analog (E), shown in Figure 9.
[0133] Other exemplary nucleotide analogs include, but are not limited to, the following shown in Figure 10: 7-ethynylphenyl-7-deaza-2'-ATP, analog (A), N6-trifluoroacetamido-2'-dATP, analog (B), and N6-ethoxyacetyl-2'dATP, analog (C).
[0134] The design of nucleotides, e.g., dATP analogs, suitable for practicing the present invention can be guided by the general structures shown in Figure 11, including: N6-(alkyl or aryl)-2'-dATP, Compound (A), N6-(alkyl or aryl)-2-alkyl-2'-dATP, Compound (B), N6,N6-(alkyl or aryl)-2-alkyl-2'-dATP, Compound (C), N6,N6-(alkyl or aryl)-2-alkyl-7-deaza -2'-dATP, Compound (D), N6,N6-(alkyl or aryl)-2-alkyl-7-alkynyl-7-deaza-2'-dATP, Compound (E), N6,N6-(alkyl or aryl)-2-alkyl-7-alkynyl-3,7-dideaza-2'-dATP, Compound (F), and gamma-O-alkyl-N6,N6-(alkyl or aryl)-2-alkyl-7-alkynyl-3,7-dideaza-2'-dATP, Compound (G).
[0135] In another embodiment, the use of a dGTP analog, such as 7-deaza-dGTP, which is a less preferred polymerase substrate, can influence which nucleotide is incorporated into the opposing abasic site during an abasic bypass primer extension reaction. In other embodiments, additional components of the primer extension reaction, such as buffer pH, solvent composition, relative ratios of dNTPs, etc., can be optimized to influence the substrate preference of the bypass DNA polymerase. In some instances, the amount of polymerase protein can be limited in the reaction, thereby minimizing the synthesis of undesired primer extension by-products.
[0136] Aminoxyalkyl nucleobase mimetics As discussed herein and with reference to Figure 5, certain chemical stabilizers react with abasic sites in DNA to form stable oxime adducts that prevent subsequent degradation of the phosphodiester backbone. As used herein, the term "oxime" refers to organic compounds belonging to the imine family, which has the general formula RR'C=N-OH, where R is an organic side chain and R' can be hydrogen, forming an aldoxime or another organic group, forming a ketoxime. O-substituted oximes form a closely related family of compounds. One particularly useful class of stabilizers used to form oxime adducts is one with the generalized aminoxyalkyl structure H2N-OR, as disclosed herein. Advantageously, the inventors have discovered that certain oximes have the additional ability to biologically mimic the Watson-Crick base-pairing activity of natural nucleobases. Thus, they not only stabilize abasic sites but also stabilize the direct incorporation of specific nucleotides at the opposing site during daughter strand synthesis. Such aminoxyalkyl-based stabilizing reagents and their corresponding oxime adduct products may alternatively be referred to herein in certain embodiments as "nucleobase mimetics," "aminoxyalkyl nucleobase mimetics," or "nucleobase oxime mimetics."
[0137] In one embodiment, the uracil mimetic 1-[2-(amino)ethyl]-uracil is used to stabilize abasic sites because the aminoxyalkyl moiety of the mimetic compound reacts with the abasic site to form a stable oxime adduct. Advantageously, the heterocyclic component of the compound can bind adenine from a Watson-Crick base pair, thus directing the incorporation of dATP during daughter strand synthesis.
[0138] Figure 12A shows an example of the conversion of 5-mC to a uracil oxime mimic. Here, as previously described, a DNA target molecule containing a 5-mC residue is treated with TET (I) to convert the 5-mC to 5-caC and then with TDG (II) to remove the 5-caC nucleobase and generate an abasic site. In this example, the DNA target is also treated with an aminoxyalkyluracil mimic (III), which chemically reacts with the abasic site to form a stable oxime mimic adduct (IV). In this embodiment, the aminoxyalkyluracil mimic (III) is 1-[2-(aminooxy)ethyl]-uracil, available from Enamine, Ltd. (Kyiv, Ukraine). Advantageously, we have found that both the enzymatic conversion and removal of 5-mC by TET and TDG and the chemical conversion of the abasic nucleotide to a stable oxime adduct can be performed in a single reaction, i.e., a "one-pot" reaction. This one-pot reaction is also referred to herein as a "chemoenzymatic nucleobase conversion reaction." Importantly, the oxime mimetic adduct (IV) can base pair with adenine and is therefore read as uracil during daughter strand synthesis.
[0139] Figure 12B illustrates how chemoenzymatic conversion of 5-mC to a uracil oxime mimic can be used to detect 5-mC in DNA target fragments. Here, a parent DNA template is subjected to steps (I) through (IV) to chemoenzymatically convert 5-mC to a uracil oxime mimic, as discussed with reference to Figure 12A. Prior to this conversion, a first daughter strand copy of the template is synthesized (V), as discussed with reference to Figure 1B. This reaction is performed using a native nucleotide, such that a native G is incorporated into the daughter strand opposite the 5-mC in the parent template. After chemoenzymatic conversion of the parent template, a second primer extension reaction generates a second daughter strand copy (VI). This reaction can also be performed using a native nucleotide, such that a native A is incorporated opposite the uracil oxime mimic. For sequence comparison analysis, both the first and second daughter strand copies serve as templates for the Sequencing by Expansion (SBX®) protocol (VII), as further described herein. The resulting sequencing reads of the first daughter strand copy show a "C" at each 5-mC position of the original parental template, while the sequencing reads of the second daughter strand copy show a "T" at each 5-mC position of the parental template. Thus, a "C→T" substitution in the sequence of the second daughter strand Xpandomer copy reveals the location of the 5-mC in the target fragment.
[0140] Other exemplary aminoxyalkyl nucleobase mimetics suitable for the methods of the present invention include 1-[3-(aminoxy)propyl]-uracil, 1-[4-(aminoxy)butyl]-uracil, and 1-[5-(aminoxy)pentyl]-uracil, commercially available from, for example, Enamine Ltd. In other embodiments, the present invention contemplates new aminoxyalkyl nucleobase mimetics with specific chemical features optimized for specific applications. For example, the mimetics can include heterocycles other than uracil, such as thymine, cytosine, guanine, or adenine. In other embodiments, the mimetics can include alternative atom distances between the oxime and the heterocycle, for example, from two carbons to three, four, or five carbons. Certain exemplary aminoxyalkyl nucleobase mimetics are shown in Figure 13 and include: 1-[2-(aminoxy)ethyl]-2,4-diiodo-5-methylbenzene, Compound (A); 1-[2-(aminoxy)ethyl]-2,4-dibromo-5-methylbenzene, Compound (B); 1-[2-(aminoxy)ethyl]-2,4-dichloro-5-methylbenzene, Compound (C); 1-[2-(aminoxy)ethyl]-2,4-difluoro-5-methylbenzene, Compound (D); 1-[2-(aminoxy)ethyl]-thymine, Compound (E); and additional predicted pseudouridine analogs, Compounds (F) and (G).
[0141] According to the present invention, when a nucleotide is incorporated into a second complementary strand opposite the abasic site of a DNA target strand, it is unable to form a Watson-Crick base pair with the original removed modified nucleobase under the primer extension conditions described herein. For example, if the modified nucleobase of interest is a cytosine derivative, the nucleotide incorporated opposite the removed base is not dGTP, but rather dATP, dCTP, or dTTP, or a derivative thereof; if the modified nucleobase of interest is a guanine derivative, the nucleotide incorporated opposite the removed base is not dCTP, but rather dATP, dGTP, or dTTP, or a derivative thereof; if the modified nucleobase of interest is an adenine derivative, the nucleotide incorporated opposite the removed base is not dTTP, but rather dATP, dCTP, or dGTP, or a derivative thereof; and if the modified nucleobase of interest is a thymine derivative, the nucleotide incorporated opposite the removed base is not dATP, but rather dCTP, dGTP, or dTTP, or a derivative thereof. In a preferred embodiment, as discussed herein, native dATP or a derivative thereof is the nucleotide incorporated into the opposing abasic site resulting from the removal (i.e., conversion) of a modified cytosine (e.g., 5-mC) in the original DNA target fragment.
[0142] In some examples, the yield of the desired incorporated nucleotide is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or approximately 100% of the total number of incorporation events for each second complementary copy strand produced. For example, the yield of the desired incorporated nucleotide can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or approximately 100% of all events in each second primer extension reaction. In one example, the yield of the desired incorporated product can be at least 80%. In one example, the yield of the desired incorporated nucleotide can be at least 85%. In another example, the yield of the desired incorporated nucleotide can be at least 90%. In another example, the yield of the desired incorporated nucleotide can be at least 95%. In another example, the yield of the desired incorporated nucleotide can be approximately 100%.
[0143] In certain cases, the second DNA polymerase may "skip" the base-free site during the second primer extension reaction, resulting in a deletion in the second complementary copy opposite the position of the missing site.In yet other cases, the second DNA polymerase may incorporate two or more nucleotides at the position opposite the base-free site of the target DNA polymerase, thus creating an insertion in the second complementary copy.In either case, the sequence of the second daughter strand contains a difference from the sequence of the first daughter strand, which indicates the position of the modified nucleobase in the target fragment.
[0144] In some examples, once the first and second complementary copy strands of a DNA target fragment are produced as described above, they can be evaluated by several established and emerging nucleic acid sequencing technologies, including, but not limited to, deep sequencing, next generation sequencing, and nanopore sequencing.
[0145] Chemoenzymatic nucleic acid base conversion reaction mixture In certain embodiments, a chemo-enzymatic nucleobase conversion reaction mixture according to the present invention may comprise at least one DNA glycosylase enzyme, a chemical stabilizer, and a suitable buffer.
[0146] Each DNA glycosylase may have specificity for one or more different types of modified nucleobases or one or more types of nucleobase modifications. In some embodiments, the DNA glycosylase enzyme comprises one of the glycosylase enzymes listed in Table 1. In other embodiments, the chemoenzymatic nucleobase conversion reaction mixture may contain an additional enzyme, such as a TET enzyme, that chemically converts the modified nucleobase of interest without removing the nucleobase from the DNA fragment. In some embodiments, the amount of DNA glycosylase enzyme in the nucleobase conversion reaction mixture is sufficient to completely remove most of the modified nucleobases of interest from the DNA target fragments. For example, the amount of DNA glycosylase enzyme can be about 0.1 μg of purified enzyme protein / pmol DNA template, about 0.15 μg of purified enzyme protein / pmol DNA template, about 0.2 μg of purified enzyme protein / pmol DNA template, about 0.3 μg of purified enzyme protein / pmol DNA template, about 0.5 μg of purified enzyme protein / pmol DNA template, about 0.7 μg of purified enzyme protein / pmol DNA template, about 1.0 μg of purified enzyme protein / pmol DNA template, about 1.5 μg of purified enzyme protein / pmol DNA template, about 2 μg of purified enzyme protein / pmol DNA template, or more than 2 μg of purified enzyme protein / pmol DNA template.
[0147] In some embodiments, the chemical stabilizer may be selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminoxy)propyl]-uracil, 1-[4-(aminoxy)butyl]-uracil, 1-[5-(aminoxy)pentyl]-uracil, 1-[2-(aminoxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminoxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminoxy)ethyl]-thymine. In some embodiments, the chemical stabilizer may be present in the nucleobase conversion reaction mixture at a final molar concentration of about 1 mM, about 5 mM, about 10 mM, about 15 mM, about 20 mM, about 25 mM, about 30 mM, up to 50 mM, up to 75 mM, up to 100 mM, or greater than 100 mM.
[0148] In some embodiments, a suitable buffer may be selected from the group consisting of MES, Tris-HCl, HEPES, etc. In further embodiments, a suitable buffer may contain additional excipients, such as a salt (e.g., NaCl or NaOAc), DTT, MgCl, DTT, PEG, etc. In other embodiments, the nucleobase conversion reaction may include one or more cofactors suitable for the particular DNA glycosylase or other conversion enzyme, such as ammonium iron(II) sulfate, α-ketoglutarate, and sodium ascorbate. In some embodiments, the final pH of the nucleobase conversion reaction mixture may be about pH 4, about pH 5, about pH 6, about pH 7, or greater than pH 7. Of course, one of skill in the art will understand that the final pH will depend on the particular stabilizers, DNA glycosylase, and other enzymes present in the reaction mixture.
[0149] In certain embodiments, the chemo-enzymatic nucleobase conversion reaction mixture can be a liquid, a frozen liquid, a dried liquid, a lyophilized liquid, or a partially lyophilized liquid.
[0150] kit In another aspect, a kit is provided that includes reagents for carrying out the methods described herein. In certain embodiments, the kit can include the chemoenzymatic nucleic acid base conversion reaction mixture described herein. Various other enzymes can be included in the kit. For example, the kit can include one or more of a high-fidelity DNA polymerase, an abasic bypass DNA polymerase, and a DNA polymerase with exonuclease activity. The kit can also include a DNA ligase for library preparation, such as a DNA ligase for ligating adapters to DNA target fragments to create a library of adapter-linked DNA target fragments.
[0151] In some examples, the kit may include one or more buffers and / or reaction components for carrying out the first primer extension reaction, nucleobase removal reaction, abasic stabilization reaction, and second primer extension reaction steps of the method. For example, the kit may include one or more of a DNA polymerase buffer, a DNA glycosylase buffer, a DNA ligase buffer, or any combination thereof. The kit may also include other reagents such as salts, cations, or detergents.
[0152] In some examples, the kit includes reagents and instructions for fragmenting a DNA sample and ligating adapters. For example, the kit can include one or more enzymes for fragmenting DNA and ligating adapters.
[0153] In some examples, the kit may further include a control DNA oligonucleotide containing one or more modified nucleobases of interest. The control oligonucleotide may be provided at a known concentration or may have a known amount of modified nucleobases per DNA molecule or concentration. In some examples, the control DNA oligonucleotide may be in a specific size range. For example, the control DNA oligonucleotide may be in a range of 25-100 bp, 25-150 bp, 50-200 bp, 50-300 bp, 25-500 bp, etc. In some examples, the control DNA oligonucleotide may be in the same approximate size range as the DNA molecules to be analyzed using the kit.
[0154] In some examples, the kit may further include instructions. The instructions may specify how to perform one or more of the following steps: DNA isolation, DNA fragmentation, adapter ligation, first primer extension, glycosylase treatment, abasic site stabilization, and second primer extension. Instructions describing how to use the control DNA oligonucleotide may also be included in the kit.
[0155] Sequencing by extension One nucleic acid sequencing methodology of the present invention is "Sequencing by Expansion" (SBX®), developed by Stratos Genomics (see, e.g., Kokoris et al., U.S. Pat. No. 7,939,259, "High Throughput Nucleic Acid Sequencing by Expansion," incorporated herein by reference in its entirety). SBX is based on the polymerization of highly modified, unnatural nucleotide analogs called "XNTPs." Generally, SBX uses biochemical polymerization to transcribe the sequence of a DNA template (e.g., first and second complementary copies of a DNA target fragment as described herein) onto measurable polymers called "Xpandomers." The transcribed sequences are encoded along the Xpandomer backbone in high-signal-to-noise reporters spaced approximately 10 nm apart, designed for high signal-to-noise, highly differentiated response. These differences result in significant performance improvements in sequence read efficiency and accuracy of Xpandomers compared to natural DNA.
[0156] XNTPs are extendable 5' triphosphate-modified unnatural nucleotide analogs compatible with template-dependent enzymatic polymerization. XNTPs have two distinct functional regions: a selectively cleavable phosphoramidate bond that connects the 5' α-phosphate to the nucleobase, and a symmetrically synthesized reporter tether (SSRT) attached to the nucleoside triphosphoramidate at a position that allows controlled extension by cleavage of the phosphoramidate bond. SSRTs contain linkers separated by selectively cleavable phosphoramidate bonds. Each linker is attached to one end of a reporter code. The XNTP substrate incorporated into the daughter strand product of template-dependent polymerization is in a "constrained" configuration. The constrained configuration of polymerized XNTPs is a precursor to the extended configuration, as seen in the Xpandomer product.
[0157] The transition from the constrained to the extended configuration occurs via cleavage of selectively cleavable phosphoramidate bonds within the primary backbone of the daughter strand. In this embodiment, the SSRTs contain one or more reporters or reporter codes specific to the nucleobases to which they are linked, thereby encoding the sequence information of the template. In this way, the SSRTs provide a means to extend the length of the Xpandomer and reduce the linear density of the sequence information of the parent strand.
[0158] The SSRT (i.e., "tether") of XNTP contains several distinct functional elements or features, such as a polymerase enhancer region, a reporter code, and a translational control element (TCE). Each of these features performs a specific function during Xpandomer translocation through the nanopore, producing a series of unique and reproducible electronic signals. The SSRT is designed to control the rate of Xpandomer translocation by the TCE through a combination of steric and / or electrical repulsion, and different reporter codes are sized to block ion flow through the nanopore at different, measurable levels.
[0159] Specific SSRT polymer sequences can be efficiently synthesized using phosphoramidite chemistry typically used in oligonucleotide synthesis. Reporter codes and other features can be designed by selecting specific phosphoramidite sequences from commercially available and / or proprietary libraries. Such libraries include, but are not limited to, polyethylene glycols having lengths of 1 to 12 or more ethylene glycol units and aliphatic polymers having lengths of 1 to 12 or more carbon units. In certain embodiments, the SSRT contains a feature called a "polymerase-enhancing region" at the end of the SSRT proximal to the nucleotide triphosphoramidate diester. The polymerase-enhancing region may contain a positively charged polyamine spacer (e.g., a primary, secondary, tertiary, or quaternary amine) or a triamine spacer (three secondary amines separated by three carbons) that promotes incorporation of XNTP structures by nucleic acid polymerases. In certain embodiments, the polymerase-enhancing region contains two repeating units of spermine.
[0160] As used throughout this disclosure, the terms "linker A" and "linker B" refer to regions of the SSRT that comprise a polymerase-enhancing region and one or more translocation slowing features or regions, respectively, and, in certain embodiments, a spacer region comprising a polymer of, for example, PEG6, that can be customized to modulate the length of the SSRT traversing the nanopore.
[0161] In certain embodiments, the XNTP may be a compound having the following generalized structure: [ka]
[0162] In one embodiment, R can be H, for example, when the compound is used to sequence a DNA template.
[0163] In certain embodiments, the nucleobase is adenine, cytosine, guanine, thymine, uracil, or a nucleobase analog. As will be understood by those skilled in the art, adenine, cytosine, guanine, thymine, and uracil are naturally occurring nucleobases. As used herein, the term "nucleobase analog" refers to a non-naturally occurring nucleobase that can form a Watson-Crick base pair with a complementary nucleobase on an adjacent single-stranded nucleic acid template.
[0164] To obtain sequence information, Xpandomers are translocated from the cis reservoir to the trans reservoir through a nanopore. As the Xpandomers translocate, the reporters enter the stem until their translocation control elements stop them at the stem entrance. The reporters are held at the base until TCE enters and allows them to pass through the base, after which translocation proceeds to the next reporter. Upon passing through the nanopore, each reporter code on the linearized Xpandomers generates a distinct and reproducible electronic signal specific to the nucleobase to which it is linked.
[0165] In certain embodiments, Xpandomer produced by SBX chemistry can be analyzed using a nanopore-based sequencing chip. The nanopore-based sequencing chip can incorporate a large number of sensor cells configured as an array. For example, the chip can include an array of 1 million cells, configured with 1,000 rows and 1,000 columns of cells. Each cell in the array can include control circuitry integrated on a silicon substrate. Such nanopore-based sequencing chips, devices, and systems are described, for example, in the applicant's published patent application WO2021 / 219795, the entire contents of which are incorporated herein by reference.
[0166] A proprietary in-house bioinformatics pipeline is typically used to process sequencing reads. The method disclosed herein utilizes UMI to allow pairing of first and second complementary copy reads. Read pairs can be quality filtered and trimmed of adapter and primer sequences. UMI sequences can be clustered together to define UMI families (all reads originating from a single DNA template).
[0167] Diagnostic and prognostic methods In certain embodiments, the methods can be directed to diagnosing an individual with a condition characterized by a methylation level and / or pattern at a particular locus in a test sample that differs from the methylation level and / or pattern at the same locus in a sample considered normal or in the absence of the condition. The methods can also be used to predict an individual's susceptibility to a condition characterized by a level and / or pattern of methylation at a locus that differs from the level and / or pattern of methylation at the locus exhibited in the absence of the condition.
[0168] With particular regard to cancer, DNA methylation changes have been recognized as one of the most common molecular alterations in human neoplasms. Hypermethylation of CpG islands located in the promoter regions of tumor suppressor genes is a well-established common mechanism for gene inactivation in cancer (Esteller, Oncogene 21(35):5427-40(2002)). In contrast, global hypomethylation of genomic DNA has been observed in tumor cells, and correlations between hypomethylation and increased gene expression have been reported for many cancer genes (Feinberg, Nature 301(5895):89-92(1983), Hanada, et al., Blood 82(6):1820-8(1993)). Cancer diagnosis or prognosis can be performed using the methods described herein based on the methylation status of specific sequence regions of genes, including, but not limited to, coding sequences, 5' regulatory regions, or other regulatory regions that affect transcription efficiency.
[0169] The reference genomic DNA (e.g., gDNA considered "normal") and the test genomic DNA compared in the diagnostic or prognostic method can be obtained from different individuals, different tissues, and / or different cell types.In certain embodiments, the genomic DNA samples compared can be from the same individual, but from different tissues or different cell types, or from tissues or cell types that are differentially affected by disease or symptoms.Similarly, the genomic DNA samples compared can be derived from the same tissue or the same cell type, and the cells or tissues are differentially affected by disease or symptoms. [Example]
[0170] Example 1 One-pot chemoenzymatic conversion reactions This example demonstrates the glycosylase-mediated removal of 5-mC from double-stranded DNA target fragments and the chemical conversion of the resulting abasic site to a stable oxime adduct, utilizing an aminoxyalkyluracil mimic. Advantageously, the enzymatic and chemical conversion reactions were carried out simultaneously in a single reaction vessel (i.e., a "one-pot" reaction).
[0171] For this experiment, a single-stranded DNA target fragment (80-mer) was designed to contain three spaced 5-mC residues. The 5' end of the target strand was covalently modified with biotin to facilitate physical manipulation of the strand with streptavidin-coated beads. The target strand was hybridized to a complementary oligonucleotide strand containing native nucleotides at a molar ratio of 5:7.5 pmol to produce a double-stranded fragment. A 21-mer oligonucleotide primer was designed to hybridize to the 3' end of the template.
[0172] The "one-pot" conversion reaction contained the following reagents: double-stranded DNA fragment, 3 μg purified ngTET protein, 8 μg purified TDG protein, 50 mM MES buffer, pH 6, 50 mM NaCl, 1 mM alpha-ketoglutarate (TET cofactor), 2 mM sodium ascorbate (TET cofactor), 1 mM DTT, 20% PEG, 0.1 mM ammonium iron(II) sulfate (Mohs salt, TET cofactor), and either 10 mM or 26 mM of the aminoxyalkyluracil mimetic, 1-[2-(aminooxy)ethyl]-4-hydroxy-1,2-dihydropyrimidin-2-one (C6H9N3O3) (commercially available from Enamine, Ltd., Kyiv, Ukraine). The final reaction (50 μL) was diluted to 28 μl. o C for 3 hours. Controls included a similar one-pot reaction but omitting the uracil mimetic and also omitting the TET and TDG enzymes.
[0173] To allow for detection of the chemoenzymatic conversion of the target strand, the reaction products were subjected to mildly basic conditions (100 mM NaOH for 20 min) to selectively cleave the target strand at the newly generated abasic site. The reaction products were analyzed by gel electrophoresis and visualized by Syber staining.
[0174] A representative gel is shown in Figure 14. Lane 1 shows the product of a control reaction lacking TET and TDG proteins. The larger band corresponds to the longer target strand, and the smaller band corresponds to the shorter complementary strand. As expected, no degradation of the target strand was observed in the absence of DNA glycosylase enzymes. In contrast, lane 2 shows degradation of the target strand in the presence of TET and TDG proteins, demonstrating the removal of the 5-mC residue, generating an unstable abasic site susceptible to base-mediated strand degradation. Importantly, lanes 3 and 4 show that including an aminoxyalkyluracil mimic in the conversion reaction prevents target strand degradation. This observation is consistent with a mechanism in which the mimic forms a stable oxime adduct at the abasic site created by nucleobase removal that is resistant to further degradation.
[0175] These results demonstrate the successful chemoenzymatic conversion of 5-mC residues in DNA target fragments to stable oxime adducts and provide proof-of-concept support that these separate reactions can be carried out in a single, one-pot reaction.
[0176] Example 2 DPO4 polymerase exhibits abasic bypass activity This example demonstrates that DPO4, a class Y DNA polymerase isolated from S. solfataricus, can successfully synthesize full-length copies of DNA templates containing several abasic challenges.
[0177] For this experiment, a single-stranded 80-mer template was designed to contain three abasic (AP) sites. The 5' end of the template was covalently modified with a biotin moiety for immobilization on streptavidin-coated beads. A 21-mer extension oligonucleotide (EO) was designed to hybridize to the 3' end of the template. The 5' end of the EO was covalently modified with a SIMA dye for fluorescent detection of primer extension products.
[0178] Prior to the primer extension reaction, the template was prepared by incubating 75 pmol of template with 100 pmol of EO and 50 μl (10 mg / ml) of beads (Dynabeads™ MyOne™ Streptavidin C1, Thermofisher, Inc.) and incubated at room temperature for 10 minutes.
[0179] For a given primer extension reaction, 10 μl of DNA-bead complex was used to provide the template and EO. The primer extension reaction contained the following reagents: 20 mM Tris-HCl, pH 8.8, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1% Triton X-100, 200 μM dNTPs, 1 mM MnCl, and 2 μg purified DPO4 polymerase. The total reaction volume was 20 μl. The reaction was carried out at 37°C for 1 hour. After elution of the products from the beads with a buffer containing NaOH, the primer extension products were analyzed by gel electrophoresis.
[0180] A representative gel is shown in Figure 15. Lane 1 shows the product of a primer extension reaction lacking DNA polymerase. As expected, no extension product is observed. Lanes 2-4 show the product of primer extension reactions containing no additional additives (lane 2), 50% 7-deaza-dGTP (lane 3), or 100% 7-deaza-dGTP (lane 4). As shown, DPO4 polymerase can effectively synthesize a full-length copy of the 80-mer template, demonstrating, surprisingly, that DPO4 polymerase can bypass all three abasic sites in the DNA template.
[0181] These results demonstrate that DPO4 is capable of synthesizing daughter strands across several abasic sites in the parental template, thus validating this enzyme as a potentially suitable polymerase for carrying out the methods disclosed herein.
[0182] Example 3 Improved bypass activity in DNA templates with stabilized abasic sites This example demonstrates that the combination of an engineered DPO4 variant with wild-type DPO1 polymerase can synthesize a full-length copy of a DNA template containing three abasic sites stabilized as uracil oxime mimics. Furthermore, this example demonstrates that stabilization of the abasic sites as uracil oxime mimics leads to efficient incorporation of dATP at the opposite site of the newly synthesized daughter strand.
[0183] For this experiment, a single-stranded DNA template (80-mer) was designed to contain three abasic (AP) sites spaced relatively evenly along the length of the template. Abasic oligonucleotides were synthesized using conventional phosphoramidite chemistry, for example, using Abasic II phosphoramidite (5-O-dimethoxytrityl-1-O-tert-butyldimethylsilyl-2-deoxyribose-3-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramidite) available from Glen Research, Sterling, Virginia, following the manufacturer's recommended protocol. The abasic oligonucleotides were treated with 100 mM aminoxyalkyl at pH 4-5 to generate oxime adducts at the abasic sites, which were purified by gel electrophoresis. This experiment utilized the aminoxyalkyl uracil mimics described in Example 1.
[0184] The 5' end of the template was conjugated with biotin to allow for physical manipulation of the strand. A 21-mer extension oligonucleotide (EO) was designed to hybridize to the 3' end of the template. The 5' end of the EO was covalently modified with SIMA dye for fluorescent detection of primer extension products.
[0185] The following primer extension reactions were performed using the abasic oligonucleotide as a template: A) extension with KAPA DNA polymerase, B) extension with wild-type DPO4 polymerase, C) extension with the DPO4 polymerase variant, C9110, and D) extension with a combination of the DPO4 variant polymerase, C9110, and DPO1 polymerase.
[0186] Primer extension reaction A contained the following reagents: 3 pmol abasic template, 2 pmol extension oligo primer, KAPA HiFi buffer, and polymerase available from Roche Sequencing Solutions. The total reaction volume was 10 μl. Reactions were performed at 55°C for 30 minutes according to the manufacturer's instructions. As a control, an identical primer extension reaction was performed using a native template lacking the abasic site. Primer extension reaction B contained the following reagents: 3 pmol abasic template, 2 pmol extension oligo primer, 20 mM Tris-HCl, pH 8.8, 10 mM (NH4)2SO4, 10 mM KCl, 2 mM MgSO4, 0.1% Triton X-100, 200 μM dNTPs, and 2 μg purified DPO4 polymerase. The total reaction volume was 10 μl. The reactions were performed at 37°C for 1 hour. Primer extension reaction C contained the following reagents: 3 pmol abasic template, 2 pmol extension oligo primer, 20 mM Tris-HCl, pH 8.8, 100 mM NaCl, 20 μM dNTP / 1000 μM dATP, 1 μg purified DPO4 polymerase variant C9110, 4 mM MgCl, 10% PEG, 10% BHA NMP, 150 mM betaine, 1 mM spermine, 0.15 mM HMP, and 1 mM PEM. The total reaction volume was 10 μl. The reaction was carried out at 55°C for 14 hours. Primer extension reaction D contained the following reagents: 3 pmol abasic template, 2 pmol extension oligo primer, 20 mM Tris-HCl, pH 8.8, 100 mM NaCl, 20 μM dNTP / 1000 μM dATP, 1 μg purified DPO4 variant C9110, 25 nM Dpo1, 4 mM MgCl2, 10% PEG, 10% BHA NMP, 150 mM betaine, 1 mM spermine, 0.15 mM HMP, 1 mM PEM. The total reaction volume was 10 μl. The reaction was carried out at 55°C for 14 h. Primer extension products were analyzed by gel electrophoresis and visualized by excitation of the SIMA(HEX) dye linked to the extension oligo.
[0187] A representative gel is shown in Figure 16. As shown in gel (A), KAPA polymerase is able to synthesize a full-length (FL) copy of the native 80-mer template (lane "C"); however, the small fluorescent band observed in the gel indicates that the polymerase stops at the first abasic site in the template; therefore, this polymerase is unable to extend the extension oligo hybridized to the abasic template (lane "AP"). In contrast, as shown in gel (B), wild-type DPO4 polymerase is able to synthesize a full-length copy of the abasic template, as evidenced by the large band corresponding to the full-length product in the gel. However, the wild-type polymerase also stops at the abasic site in the template, as demonstrated by the smear of incomplete extension products in the gel. As shown in gel (C), the DPO4 variant C9110 exhibits improved extension activity compared to the wild-type polymerase, indicating more efficient synthesis of a full-length copy of the abasic template. As shown in gel (D), the combination of the DPO4 variant, C9110, and DPO1 polymerase demonstrates the most significant improvement in primer extension activity, as the majority of extension products observed by the gel are full-length. Without being bound by theory, we speculate that the exonuclease activity of DPO1 may function as a "correction factor," for example, by reversing misincorporations performed by DPO4 and allowing the polymerase to resume extension with higher fidelity.
[0188] These results indicate that the combination of DPO4 variant, C9110, and DPO1 polymerase can synthesize full-length daughter strands over several abasic adducts in the parental template with improved efficiency compared to wild-type DPO4 or DPO4 variant polymerase alone.
[0189] The products of this primer extension reaction were subjected to DNA sequence analysis to identify the nucleotides incorporated by the DPO4 variant and DPO1 polymerase at the sites opposite the abasic site in reaction D above. The particular DNA sequencing method utilized was nanopore-based sequencing by extension method developed by the inventors and described in more detail above.
[0190] To synthesize Xpandomer copies of the primer extension products, SBX reactions were performed containing the following reagents: a 2:1 molar ratio of single-stranded DNA template to SBX extension oligonucleotide, 0.07 μg / μL DNA polymerase (DPO4 variant C7326, SEQ ID NO: 2), 15 mM AZ-43, 43 PEM (i.e., compound 73 disclosed in applicant's published PCT application no. WO2019 / 135975, which is incorporated herein by reference in its entirety), 100 μM XNTPS (disclosed in applicant's published PCT application no. WO2020 / 236526, which is incorporated herein by reference in its entirety), 0.2 mM HMP, 0.6 mM MnCl, 50 mM Tris HCl, 175 mM NaCl, 200 mM imidazole, 350 mM betaine, 20% PEG, 7% NMP, 3% DMSO. The reaction was carried out at 37°C for 2 hours. The resulting Xpandomer sample was treated with acid (7.5M DCl) to cleave the phosphoramidate bond within the XNMP subunit and generate an extended form of Xpandomer. Xpandomers were sequenced using a Roche HTP High Throughput Nanpore Sequencing Platform, as described, for example, in the applicant's published PCT application no. PCT / EP2019 / 084581, the entire contents of which are incorporated herein by reference.
[0191] For this experiment, 10 6Over 100 individual full-length Xpandomer sequences were obtained and analyzed. The results of these analyses are shown in Figures 17A and 17B, which are graphs depicting the percentage of all sequences exhibiting a specific nucleotide incorporation at each of the three abasic sites in the parental DNA template. Importantly, as shown in Figure 17A (corresponding to primer extension reaction D), dATP was overwhelmingly most efficiently incorporated at the nucleotide opposite each abasic site in the template, with over 90% of primer extension product sequences exhibiting an A at each of these three positions. Furthermore, dGTP incorporation at any of these positions was observed to be an extremely rare event. In contrast, as shown in Figure 17B (corresponding to primer extension reaction A), dGTP was by far the most efficiently incorporated nucleotide opposite each of the 5-mC residues in the native template, as expected.
[0192] In summary, these results demonstrate the chemoenzymatic conversion of 5-mC into uracil-mimetic adducts in DNA templates and the efficient incorporation of dATP opposite these sites by combining an engineered DPO4 variant with DPO1 polymerase. This novel conversion strategy allows for the identification of G→A substitutions when sequence reads from the first and second daughter strand copies of unconverted and converted DNA templates, respectively, are compared, thus providing an improved alternative approach to identifying epigenetic information in DNA samples.
Claims
1. A method for identifying modified nucleic acid bases in multiple nucleic acids, wherein the method is To provide a sample containing multiple DNA templates, The process involves generating a first complementary copy of the DNA template, wherein the generation is directed by an oligonucleotide primer using a first DNA polymerase in the presence of native dNTPs, and the generation produces a complementary copy of each of the DNA templates such that each complementary copy contains native dNTPs, and each complementary copy hybridizes into one of the DNA templates. The DNA template and the first complementary copy are subjected to DNA glycosylase treatment, wherein the DNA glycosylase specifically removes the modified nucleic acid bases in the DNA template, converts the positions of the modified nucleic acid bases into debase sites, and as a result, each glycosylase-converted DNA template hybridizes to the unconverted complementary copy. The process involves generating a second complementary copy of the glycosylase-converted DNA template, wherein the generation is directed by a second DNA polymerase, the second DNA polymerase is capable of incorporating nucleotides opposite to the debasement sites of the converted DNA template, and the nucleotides do not form Watson-Crick base pairs with the modified nucleic acid bases, Determining the nucleotide sequences of the first and second complementary copies, For each of the DNA glycosylase-converted DNA templates, the nucleotide sequence of the second complementary copy is compared with the nucleotide sequence of the first complementary copy, thereby determining the position of the modified nucleic acid base in the DNA template before DNA glycosylase conversion. Methods that include...
2. The method according to claim 1, wherein the step of comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each of the DNA glycosylase-converted DNA templates identifies nucleotide substitutions in the sequence of the second complementary copy relative to the first complementary copy, and the location of the nucleotide substitutions identifies the location of the modified base in the DNA template.
3. The method according to claim 1, wherein the modified nucleic acid base is selected from the group consisting of 5-mC, 5-hmC, 5-fC, and 5-caC.
4. The method according to claim 1, wherein the DNA glycosylase is a monofunctional DNA glycosylase.
5. The method according to claim 4, wherein the monofunctional DNA glycosylase is thymine DNA glycosylase (TDG) or a variant thereof.
6. The method according to claim 5, wherein the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment further comprises subjecting the DNA template and the first complementary copy to treatment with a 10,11 translocation (TET) enzyme or a variant thereof.
7. The method according to claim 6, wherein the 10,11 translocation (TET) enzyme or its variant is ngTET.
8. The method according to claim 1, wherein the DNA glycosylase is a bifunctional DNA glycosylase.
9. The method according to claim 8, wherein the bifunctional DNA glycosylase is a member of the DEMETER (DME) family of DNA glycosylases or a variant thereof.
10. The method according to claim 9, wherein the member of the DEMETER (DME) family of the DNA glycosylase or a variant thereof is a variant that has been manipulated to inactivate lyase activity.
11. The method according to any one of claims 1 to 10, wherein the second DNA polymerase is a debase bypass DNA polymerase.
12. The method according to claim 11, wherein the debase bypass DNA polymerase is DPO4 polymerase or a variant thereof.
13. The method according to claim 12, wherein the DPO4 polymerase or its variant is a variant comprising the following mutations: M76W, K78E, E79P, Q82W, Q83G, and S86E (SEQ ID NO: 3).
14. The method according to any one of claims 11, wherein the debase bypass DNA polymerase incorporates dATP into the second complementary copy of the glycosylase-converted DNA template at a position opposite to the debase site.
15. The method according to claim 11, wherein the debase bypass DNA polymerase further comprises a third DNA polymerase, and the third DNA polymerase has exonuclease activity.
16. The method according to claim 15, wherein the third DNA polymerase is DPO1 polymerase.
17. The method according to any one of claims 1 to 10, wherein the first DNA polymerase is a high-fidelity DNA polymerase.
18. The method according to any one of claims 1 to 10, further comprising the step of treating the glycosylase-converted DNA template with a stabilizer before the step of generating the second complementary copy of the glycosylase-converted DNA template.
19. The method according to claim 18, wherein the stabilizer comprises an aldehyde-reactive compound that forms a stable adduct with the debasement site.
20. The method according to claim 19, wherein the stabilizer is selected from the group consisting of O-hydroxylamine, acylhydrazine, tryptamine, beta-aminothiol, alkylhydrazine, hydrazino-iso-picte-spenglarindole, and methylaminooxy-iso-picte-spenglarindole.
21. The method according to claim 19, wherein the stabilizer comprises an aminooxyalkyl group that can form an oxime adduct with the debasement site.
22. The method according to claim 21, wherein the stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine.
23. The method according to claim 22, wherein the stabilizer is 1-[2-(amino)ethyl]-uracil.
24. The method according to claim 18, wherein the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment and the step of treating the glycosylase-converted DNA template with a stabilizer before generating the second complementary copy are performed in the same step.
25. The method according to any one of claims 1 to 10, wherein the DNA template is selected from the group consisting of genomic DNA, mitochondrial DNA, cell-free DNA, circulating tumor DNA, or a combination thereof.
26. The method according to any one of claims 1 to 10, wherein the DNA template is immobilized on a solid support.
27. The method according to any one of claims 1 to 10, wherein the first or second complementary copy is immobilized on a solid support.
28. The method according to any one of claims 1 to 10, wherein the step of determining the nucleotide sequences of the first and second complementary copies includes the steps of synthesizing Xpandomer copies of the first and second complementary copies and passing the Xpandomer copies of the first and second complementary copies through a nanopore.
29. The method according to any one of claims 1 to 10, wherein the DNA template includes a first adapter attached to the 5' end of the DNA template and a second adapter attached to the 3' end of the template.
30. The method according to claim 29, wherein the first or second adapter is a Y-adapter.
31. The method according to claim 29, wherein at least one of the first and second adapters includes a unique molecular identifier barcode (UMI).
32. The method according to claim 31, wherein the step of comparing the sequences of the first and second complementary copies includes bioinformatics pairing sequences that include the same unique molecular identifier barcode (UMI).
33. A chemical enzymatic nucleic acid base conversion reaction mixture comprising a DNA glycosylase enzyme, a chemical stabilizer, and a suitable buffer.
34. A chemical enzymatic nucleic acid base conversion reaction mixture according to claim 33, further comprising a DNA template strand hybridized to a first complementary copy strand, wherein the DNA template strand comprises modified nucleic acid bases and the first complementary copy strand comprises native nucleic acid bases.
35. The chemically enzymatic nucleic acid base conversion reaction mixture according to claim 33, wherein the chemical stabilizer comprises an aminooxyalkyl group, and the aminooxyalkyl group can react with a debasic nucleotide containing a ring-opening aldehyde moiety to form a stable oxime adduct.
36. The chemical stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine, as described in claim 35.
37. The chemically enzymatic nucleic acid base conversion reaction mixture according to claim 35, wherein the chemical stabilizer is selected from the group consisting of O-hydroxylamine, acylhydrazine, tryptamine, beta-aminothiol, alkylhydrazine, hydrazino-iso-picte-spenglarindole, and methylaminooxy-iso-picte-spenglarindole.
38. The chemical enzymatic nucleic acid base conversion reaction mixture according to any one of claims 33 to 37, wherein the DNA glycosylase is selected from the group consisting of N-methylpurine DNA glycosylase (MPG), MutY homolog (MUTYH), Nth-like DNA glycosylase 1 (NTHL1), Nei-like DNA glycosylase 1 (NEIL1), Nei-like DNA glycosylase 2 (NEIL2), Nei-like DNA glycosylase 3 (NEIL3), 8-oxoguanine DNA glycosylase (OGG1), uracil DNA glycosylase 1 (Ung1), uracil DNA glycosylase 2 (Ung2), single-strand selective monofunctional uracil glycosylase (SMUG1), thymine DNA glycosylase (TDG), methyl-binding domain 4 (MBD4), Fpg, Ung, Demeter (DME), and ROS1.
39. The chemical enzymatic nucleic acid base conversion reaction mixture according to claim 38, wherein the reaction mixture comprises two or more DNA glycosylases.
40. The chemical enzymatic nucleic acid base conversion reaction mixture according to any one of claims 33 to 37, wherein the DNA glycosylase is TDG or a variant thereof.
41. A chemical enzymatic nucleic acid base conversion reaction mixture according to claim 40, further comprising a TET enzyme.
42. A kit for detecting modified nucleic acid bases in a DNA sample, comprising a chemical enzymatic nucleic acid base conversion reaction mixture according to any one of claims 33 to 37, an enzyme selected from at least one of high fidelity DNA polymerase, debase bypass DNA polymerase, and DNA polymerase having exonuclease activity, and a suitable mixture of dNTPs or an analogue thereof.
43. The kit according to claim 42, further comprising one or more buffers for the enzyme.