Detection of modified nucleobases in nucleic acid samples
The method of determining the modified nucleobase position by generating and comparing complementary copies of DNA templates solves the limitations of detecting modified nucleobases in the prior art, and achieves efficient, selective and economical detection effects.
Patent Information
- Application Number
- CN202380073085.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has limitations in detecting modified nucleobases in DNA samples, including high cost, high material consumption, complex methods and difficult to be easily applicable to a variety of modification types.
By a method, the method includes providing a sample comprising a plurality of DNA templates, generating a first complementary copy of the DNA template, cleavage the modified nucleobase in the DNA template using a DNA glycosylate enzyme, generate a second complementary copy, and determining the position of the modified nucleobase by comparing the nucleotide sequences of the first complementary copy and the second complementary copy.
This method can efficiently and selectively detect modified nucleobases in DNA samples, avoid extensive nonspecific damage to DNA, reduce costs and material consumption, and is suitable for a variety of modification types.
Smart Images

Figure BDA0005359372590000241 
Figure BDA0005359372590000251 
Figure BDA0005359372590000252
Abstract
Description
[0001] The Sequence Listing is incorporated by reference
[0002] This application contains a Sequence Listing that has been electronically submitted in XML format and is incorporated herein by reference in its entirety. The file name of the XML copy is P37864-WO_Sequence_Listing, created on October 13, 2023, and is 5 kB in size. Background Art
[0003] Products of methylation and various forms of DNA damage are associated with a variety of important biological processes. Changes in methylation patterns and the presence of damaged DNA are often among the earliest events observed for various disease states
[0004] Epigenetic modifications are crucial for normal development. For example, 5-methylcytosine is the most widely studied epigenetic modification and is associated with many key processes, including genomic imprinting, X-chromosome inactivation, repetitive element suppression, and carcinogenesis. For example, DNA methylation at the 5-position of cytosine has a specific effect of reducing gene expression and has been found in every vertebrate examined. In many disease processes, such as cancer, gene promoter CpG islands acquire abnormal hypermethylation, resulting in transcriptional silencing, which can be inherited by daughter cells after cell division. In addition, alterations in DNA methylation have been considered an important component of cancer development. Generally, hypomethylation occurs earlier and is associated with chromosomal instability and loss of imprinting, while hypermethylation is promoter-related and can be secondary to gene (tumor suppressor) silencing. In addition, 5-hydroxymethylcytosine has also emerged as an important epigenetic modification, playing a potential regulatory role in gene expression from development to aging. Various cancers have shown that even in early lesions, the content of 5-hydroxymethylcytosine in malignant tissues is continuously and significantly reduced compared to healthy tissues.
[0005] DNA is under continuous stress from both endogenous and exogenous sources. Bases exhibit limited chemical stability and are readily chemically modified by different types of damage, including oxidation, alkylation, radiation damage, and hydrolysis. Damage to DNA bases can affect their base-pairing properties and can thus be mutagenic. DNA base modifications resulting from these types of DNA damage are very common and play important roles in influencing physiological states and disease phenotypes. Examples include 7,8-dihydro-8-oxoguanine (8-oxoG) (oxidative damage), 8-oxoadenine (oxidative damage; aging, Alzheimer's disease, Parkinson's disease), 1-methyladenine, O6-methylguanine (alkylation; glioma and colorectal cancer), benzo[a]pyrene diol epoxide (BPDE), pyrimidine dimers (adduct formation; smoking, industrial chemical exposure, ultraviolet exposure; lung cancer and skin cancer), and 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, and thymine glycol (ionizing radiation damage; chronic inflammatory diseases, prostate cancer, breast cancer, and colorectal cancer). For example, 8-oxoG is a common product of DNA oxidation. 8-oxoG has a tendency to pair with adenine bases, resulting in G>>C to T·A transversion mutations. Another example is the hydrolytic deamination of cytosine and 5-methylcytosine (5-meC) that leads to the mispairing of uracil and thymine with guanine, respectively, which, if unrepaired, causes C>>G to T·A transition mutations. In another example, alkylation can generate multiple DNA base damages, including 6-meG, N7-methylguanine (7-meG), or N3-methyladenine (3-meA). While 6-meG is mutagenic due to its property of pairing with thymine, 7-meG and 3-meA block replicative DNA polymerases and are thus cytotoxic. These and many other forms of DNA base damage occur multiple times per day in cells, and only the continuous action of specialized DNA repair systems can prevent the rapid decay of genetic information. In addition to nuclear DNA being damaged, mitochondrial DNA also suffers severe oxidative damage, as well as damage caused by alkylation, hydrolysis, and adducts. For example, oxidative damage is the most common type of damage in mitochondrial DNA, mainly because mitochondria are the major source of reactive oxygen species (ROS) in cells. In addition, mitochondria contain approximately 30% of the cellular S-adenosylmethionine pool, which can non-enzymatically methylate DNA. Furthermore, exposure to certain substances, such as estrogen, tobacco smoke, and certain chemicals, also leads to preferential damage to mitochondrial DNA.
[0006] Because DNA damage and epigenetic modifications can be the earliest signs of disease states, detecting epigenetic modifications and DNA damage patterns can be useful for the early detection and intervention of diseases. However, detection methods have limitations. For example, for methylation status, spectrophotometry can be used to indicate the overall content of modifications in the target DNA, but has limited specificity. High-performance liquid chromatography (HPLC) and mass spectrometry are also often used, but are costly, require large amounts of material, and break down DNA into its constituent nucleosides or nucleotides, thus destroying sequence information for downstream analysis. Immunoprecipitation (IP) using monoclonal antibodies can enrich DNA with the target modification, but has been found to have limitations in specificity. Restriction digestion analysis utilizes fragment analysis of DNA treated with restriction endonucleases sensitive to modifications, but requires large amounts of material and is limited to sequences with restriction sites of known sensitivity. While bisulfite sequencing is considered the "gold standard" technique for detecting DNA methylation, it also has important limitations. First, the chemical conversion process causes extensive non-specific damage to DNA, so the method requires large amounts of starting material. Second, the method is costly and time-consuming, requiring multiple sequencing runs. Finally, and importantly, it is generally only applicable to methylcytosine (mC) modifications. A number of variations have been developed or proposed that allow for targeting a limited number of additional modification types (methylcytosine (mC) and hydroxymethylcytosine (hmC)), but these variations have low yields and still have the other limitations described above. They are also not easily adaptable to other modifications and are rather complex.
[0007] Accordingly, there is a need in the art for improved methods for detecting modified nucleobases in a DNA sample of interest. The present invention meets these needs and provides further related advantages as described below.
[0008] All topics discussed in the background section are not necessarily prior art and should not be assumed to be prior art merely because of their discussion in the background section. Along these lines, unless expressly stated to be prior art, any recognition of problems in the prior art discussed in the background section or related to such topics should not be regarded as prior art. Instead, the discussion of any topic in the background section should be considered as part of the inventors' approach to solving a particular problem, which may itself be inventive. SUMMARY OF THE INVENTION
[0009] Aspects of the present invention encompass the detection of modified nucleobases (such as epigenetic changes and DNA damage) in a DNA sample.
[0010] In one aspect, the present invention provides a method for identifying modified nucleobases in a plurality of nucleic acids, the method comprising the steps of: providing a sample comprising a plurality of DNA templates; generating a first complementary copy of the DNA templates, the generating being guided by an oligonucleotide primer in the presence of native dNTPs using a first DNA polymerase, wherein the generating produces a complementary copy of each DNA in the DNA templates such that each complementary copy comprises native dNTPs and wherein each complementary copy hybridizes to one of the DNA templates in the DNA templates; subjecting the DNA templates and the first complementary copies to DNA glycosylase treatment, wherein the DNA glycosylase specifically excises the modified nucleobases in the DNA templates to convert the positions of the modified nucleobases to abasic sites, resulting in each DNA template converted by the glycosylase hybridizing to the unconverted complementary copy; generating a second complementary copy of the DNA templates converted by the glycosylase, the generating being guided by a second DNA polymerase, wherein the second DNA polymerase is capable of incorporating nucleotides opposite the abasic sites in the converted DNA templates, wherein the nucleotides do not Watson-Crick base pair with the modified nucleobases; determining the nucleotide sequences of the first complementary copies and the second complementary copies; and comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each of the DNA templates converted by the DNA glycosylase, thereby determining the positions of the modified nucleobases in the DNA templates prior to conversion by the DNA glycosylase.
[0011] In one embodiment, the step of comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each of the DNA templates converted by the DNA glycosylase identifies nucleotide substitutions in the sequence of the second complementary copy relative to the first complementary copy, wherein the positions of the nucleotide substitutions identify the positions of the modified bases in the DNA template. In another embodiment, the modified nucleobase is selected from 5-mC, 5-hmC, 5-fC, and 5-caC. In another embodiment, the DNA glycosylase is a monofunctional DNA glycosylase. In another embodiment, the monofunctional DNA glycosylase is thymine DNA glycosylase (TDG) or a variant thereof. In another embodiment, the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment further comprises subjecting the DNA template and the first complementary copy to treatment with a ten-eleven translocation (TET) enzyme or a variant thereof. In yet another embodiment, the ten-eleven translocation (TET) enzyme or a variant thereof is ngTET. In another embodiment, the DNA glycosylase is a bifunctional DNA glycosylase. In another embodiment, the bifunctional DNA glycosylase is a member of the DEMETER (DME) family of DNA glycosylases or a variant thereof. In yet another embodiment, the member of the DNA glycosylase DEMETER (DME) family or a variant thereof is a variant engineered to inactivate the lyase activity. In another embodiment, the second DNA polymerase is a base excision bypass DNA polymerase. In yet another embodiment, the base excision bypass DNA polymerase is DPO4 polymerase or a variant thereof. In yet another embodiment, the DPO4 polymerase or a variant thereof is a variant comprising the following mutations: M76W, K78E, E79P, Q82W, Q83G, and S86E (SEQ ID NO:3). In another embodiment, the base excision bypass DNA polymerase incorporates dATP into the second complementary copy at a position opposite the abasic site in the glycosylase-converted DNA template. In another embodiment, the base excision bypass DNA polymerase further comprises a third DNA polymerase, wherein the third DNA polymerase has exonuclease activity. In yet another embodiment, the third DNA polymerase is DPO1 polymerase. In another embodiment, the first DNA polymerase is a high-fidelity DNA polymerase. In another embodiment, the method further comprises the step of treating the glycosylase-converted DNA template with a stabilizer prior to the step of generating the second complementary copy of the glycosylase-converted DNA template. In another embodiment, the stabilizer comprises an aldehyde-reactive compound that forms a stable adduct with the abasic site.In yet another embodiment, the stabilizer is selected from O - hydroxylamine, hydrazide, tryptamine, β - aminothiol, alkyl hydrazine, hydrazino - iso - Pictet - Spengler indole, and methylaminooxy - iso - Pictet - Spengler indole. In another embodiment, the stabilizer comprises an aminooxyalkyl group capable of forming an oxime adduct with the abasic site. In yet another embodiment, the stabilizer is selected from 1 - [2 - (amino)ethyl] - uracil, 1 - [3 - (aminooxy)propyl] - uracil, 1 - [4 - (aminooxy)butyl] - uracil, 1 - [5 - (aminooxy)pentyl] - uracil, 1 - [2 - (aminooxy)ethyl] - 2,4 - diiodo - 5 - methylbenzene, 1 - [2 - (aminooxy)ethyl] - 2,4 - dibromo - 5 - methylbenzene, 1 - [2 - (aminooxy)ethyl] - 2,4 - dichloro - 5 - methylbenzene, 1 - [2 - (aminooxy)ethyl] - 2,4 - difluoro - 5 - methylbenzene, and 1 - [2 - (aminooxy)ethyl] - thymine. In yet another embodiment, the stabilizer is 1 - [2 - (amino)ethyl] - uracil. In another embodiment, the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment and the step of treating the glycosylase - converted DNA template with a stabilizer before generating the second complementary copy occur in the same step. In another embodiment, the DNA template is selected from the group consisting of genomic DNA, mitochondrial DNA, cell - free DNA, circulating tumor DNA, or a combination thereof. In another embodiment, the DNA template is immobilized on a solid support. In another embodiment, the first complementary copy or the second complementary copy is immobilized on a solid support. In another embodiment, the step of determining the nucleotide sequences of the first complementary copy and the second complementary copy includes synthesizing Xpandomer copies of the first complementary copy and the second complementary copy and passing the Xpandomer copies of the first complementary copy and the second complementary copy through a nanopore. In another embodiment, the DNA template includes a first adaptor ligated to the 5' end of the DNA template and a second adaptor ligated to the 3' end of the template. In yet another embodiment, the first adaptor or the second adaptor is a Y - shaped adaptor. In yet another embodiment, at least one of the first adaptor or the second adaptor comprises a unique molecular identifier barcode (UMI). In another embodiment, the step of comparing the sequences of the first complementary copy and the second complementary copy includes bioinformatics pairing of sequences containing the same unique molecular identifier barcode (UMI).
[0012] In another aspect, the present invention provides a chemo-enzymatic nucleobase conversion reaction mixture comprising a DNA glycosylase, a chemical stabilizer, and a suitable buffer. In one embodiment, the chemo-enzymatic nucleobase conversion reaction mixture further comprises a DNA template strand hybridized to a first complementary copy strand, wherein the DNA template strand comprises a modified nucleobase and the first complementary copy strand comprises a natural nucleobase. In another embodiment, the chemical stabilizer comprises an aminooxyalkyl group, wherein the aminooxyalkyl group is capable of reacting with an abasic nucleotide comprising a ring-opened aldehyde moiety to form a stable oxime adduct. In another embodiment, the chemical stabilizer is selected from 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine. In another embodiment, the chemical stabilizer is selected from O-hydroxylamine, hydrazide, tryptamine, β-aminothiol, alkylhydrazine, hydrazino-iso-Pictet–Spengler indole, and methylaminooxy-iso-Pictet–Spengler indole. In another embodiment, the DNA glycosylase is selected from the DNA glycosylases listed in Table 1. In yet another embodiment, the reaction mixture comprises more than one DNA glycosylase listed in Table 1. In another embodiment, the DNA glycosylase is TDG or a variant thereof. In another embodiment, the reaction mixture further comprises a TET enzyme.
[0013] In another aspect, the present invention provides a kit for the detection of modified nucleobases in a DNA sample, the kit comprising any one of the above-described chemo-enzymatic nucleobase conversion reaction mixtures; an enzyme selected from at least one of a high-fidelity DNA polymerase, an abasic site bypass DNA polymerase, and a DNA polymerase having exonuclease activity; and a suitable mixture of dNTPs or analogs thereof. In one embodiment, the kit further comprises one or more buffers for the enzyme in a buffer. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1A 、 1B Figures 1C and 1D are schematic diagrams outlining one embodiment of the method of the present invention.
[0015] Figure 2A and 2BIt is a schematic diagram showing an alternative embodiment of the solid-phase synthesis of primer extension products.
[0016] Figure 3A and 3B It is a chemical scheme showing an embodiment of demonstrating the instability of abasic sites and a method for stabilizing abasic sites.
[0017] Figure 4A and 4B It is a scheme showing two exemplary enzymatic reactions for excising 5-mC nucleobases to generate abasic sites in a DNA target fragment.
[0018] Figure 5A and 5B Provide exemplary embodiments of chemical schemes for stabilizing abasic sites in the converted DNA target fragment.
[0019] Figure 6A and 6B It is a scheme showing an embodiment of demonstrating the "hijacking" of aminooxyalkyl-mediated DNA lyase activity.
[0020] Figure 7 Provide the chemical structures of certain exemplary nucleotide analogs for practicing the methods of the present invention.
[0021] Figure 8 Provide the chemical structures of other exemplary nucleotide analogs for practicing the methods of the present invention.
[0022] Figure 9 Provide the chemical structures of other exemplary nucleotide analogs for practicing the methods of the present invention.
[0023] Figure 10 Provide the chemical structures of other exemplary nucleotide analogs for practicing the methods of the present invention.
[0024] Figure 11 Provide the chemical structures of certain exemplary general nucleotide analogs for practicing the methods of the present invention.
[0025] Figure 12A and 12B It is a scheme showing an embodiment of demonstrating the chemoenzymatic conversion of a DNA target fragment containing a modified nucleobase of interest (5-mC) using an exemplary aminooxyalkyl uracil mimic to stabilize the abasic site in the DNA target fragment and its use for identifying the modified nucleobase in the DNA target fragment by amplicon sequencing.
[0026] Figure 13 Provide the chemical structures of certain exemplary aminooxyalkyl nucleobase mimics for practicing the methods of the present invention.
[0027] Figure 14 is a gel showing DNA products of certain chemoenzymatic nucleobase conversion reactions.
[0028] Figure 15 is a gel showing DNA products of certain primer extension reactions of abasic DNA templates with DPO4 polymerase.
[0029] Figure 16 is a gel showing DNA products of certain primer extension reactions of abasic templates using various DNA polymerases and their combinations.
[0030] Figure 17A and 17B are graphs showing nucleotides incorporated by DNA polymerase at positions opposite abasic sites in transformed DNA templates stabilized by uracil nucleobase analogs and nucleotides incorporated relative to 5-mC in untransformed templates, respectively, as determined by DNA sequencing of the templates. DETAILED DESCRIPTION
[0031] The present invention can be more readily understood by reference to the following detailed description of the preferred embodiments of the invention and the examples included herein. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0032] References throughout this specification to "one embodiment," "an embodiment," and variations thereof, mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0033] As used in this specification and the appended claims, the singular forms "a," "an," "the," and "said" include plural referents, i.e., one or more, unless the context clearly dictates otherwise. It should also be noted that the conjunctive terms "and" and "or" are generally used in their broadest sense to include "and / or," unless the content and context clearly dictate otherwise in terms of inclusivity or exclusivity. Thus, the use of an alternative (e.g., "or") should be understood to mean any one of the alternatives, both, or any combination thereof. Additionally, the compositions of "and" and "or" when referred to herein as "and / or" are intended to cover embodiments that include all relevant items or ideas, as well as one or more other alternative embodiments that include less than all relevant items or ideas.
[0034] Unless the context otherwise requires, throughout the specification and the subsequent claims, the word "comprising" and its synonyms and variations, such as "having" and "including", and its variations, such as "containing", are to be interpreted in an open, inclusive sense, e.g., "including but not limited to". The term "consisting essentially of" limits the scope of a claim to the specified materials or steps, or those that do not materially affect the basic and novel characteristics of the claimed invention.
[0035] The abbreviation "e.g." is derived from the Latin exempli gratia and is used herein to indicate non-limiting examples. Thus, the abbreviation "e.g." is synonymous with the term "for example". It should also be understood that the singular forms "a", "an", and "the" used herein and in the appended claims include the plural forms, and unless the context clearly dictates otherwise, the term "X and / or Y" means "X" or "Y" or "X" and "Y", and the letter "s" following a noun denotes both the plural and singular forms of that noun. Further, in the case of describing the features or aspects of the invention in terms of a Markush group, it is intended and will be recognized by those skilled in the art that the invention includes and is thus also described in terms of any individual member and any subgroup of members of the Markush group, and the applicant reserves the right to amend the application or claims to expressly refer to any individual member or any subgroup of members of the Markush group.
[0036] Any headings used in this document are for the sole purpose of expediting the reader's review and should not be construed as limiting the invention or the claims in any way. Accordingly, the headings and abstracts of the present disclosure provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.
[0037] Where a numerical range is provided herein, it should be understood that each intermediate value between the upper and lower limits of that range (to the extent of one-tenth of the unit of the lower limit, unless the context clearly dictates otherwise) and any other stated or intermediate value within the stated range is included in the invention. The upper and lower limits of these smaller ranges may independently be included within the smaller ranges and are also covered by the invention, subject to any expressly excluded limitations within the stated range. When the stated range includes one or both of the limits, ranges excluding one or both of those included limits are also included in the invention.
[0038] For example, unless otherwise specified, any concentration range, percentage range, ratio range, or integer range provided herein should be understood to include any integral values within the stated range and, where appropriate, fractional values thereof (e.g., one-tenth and one-hundredth of an integer). Additionally, unless otherwise specified, any numerical range recited herein with respect to any physical feature (such as polymer subunits, dimensions, or thickness) should be understood to include any integer within the stated range. As used herein, unless otherwise specified, the term "about" means ±20% of the indicated range, value, or structure.
[0039] Method for Detecting Modified DNA Nucleobases
[0040] The present disclosure describes methods and compositions for the detection of modified nucleobases in DNA samples, where the modified nucleobases reflect, for example, epigenetic modifications and DNA damage. The methods include the enzymatic excision of a modified nucleobase of interest in a DNA target fragment to create an abasic site at each position where the modified nucleobase of interest occurs in the nucleic acid sequence of the DNA target fragment. The positions of the abasic sites can be determined by DNA sequencing methods as described herein. The methods of the invention also include a workflow for generating a first complementary copy and a second complementary copy (i.e., a first strand and a second strand) of a DNA target fragment template. The first complementary copy is generated prior to the enzymatic excision of the modified nucleobase of interest, while the second complementary copy is generated after the enzymatic excision of the modified nucleobase of interest. Thus, the first complementary copy and the second complementary copy encode the genetic information and, for example, epigenetic information of the DNA target fragment, respectively. The sequence information obtained from the first complementary copy and the second complementary copy can be compared to determine the position of the modified nucleobase of interest in the nucleic acid sequence of the original DNA target fragment.
[0041] Overview
[0042] According to the methods described herein, modified nucleobases of interest can include, but are not necessarily limited to, one or more of the following: 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxylcytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (*-oxoG), uracil (U), 6-methyladenine (6-mA), 8-oxoadenine, O-6-methylguanine, 1-methyladenine, O-4-methylthymine, 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxymethyluracil, or thymine dimer. In certain cases, multiple any combinations of these exemplary and other modified nucleobases can be detected by the methods of the invention.
[0043] In one aspect, methods are provided for detecting modified nucleobases (e.g., modified nucleobases of interest) in a sample of nucleic acid. Figures 1A to 1D A schematic overview of an exemplary method is shown in Figures 1A to 1D . The method can include step A: obtaining a nucleic acid sample and fragmenting the nucleic acid to produce a sample comprising DNA target fragment 100. As used herein, the term "target fragment" means the corresponding nucleic acid fragment is derived from a biological sample and is the target of the methods described herein, which interrogate the presence of specific modified nucleobases in the nucleic acid sequence. In this non-limiting description, the modified nucleobase of interest is methylated cytosine (5-mC), and the DNA target fragment is a double-stranded nucleic acid fragment. Here, the strands of the DNA target fragment are depicted as "parental (+)" 100a (i.e., the sense strand) and "parental (-)" 100b (i.e., the antisense strand). For simplicity, each strand of the DNA target fragment in this example contains a single 5-mC residue.
[0044] In certain cases, the DNA target fragment can be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof obtained from a biological sample.
[0045] In certain embodiments, the method can include step B: ligating (i.e., joining) adapters 101 and 103 to the 5' and 3' ends of the DNA target fragment to produce an adapter-ligated DNA target fragment. The adapters can include regions of double-stranded DNA and regions of single-stranded DNA. In Figure 1A the example shown, the adapters are Y-shaped adapters (YADs) and comprise one double-stranded region and two single-stranded DNA regions. The adapters can also include sequences or other features that mediate downstream steps of the workflow. For example, in certain embodiments, the adapters can include sequences for immobilizing the adapter-ligated DNA target fragment on a solid support, sequences for oligonucleotide primer hybridization, sequences enabling bioinformatics analysis of DNA sequence information (e.g., unique molecular identifier barcodes [UMI], sample identifiers [SID]), chemical moieties for solid-phase immobilization, etc. In certain embodiments, the structures of adapters 101 and 103 can be the same or different, depending on the specific application.
[0046] The method can then include step C: denaturing the DNA target fragment to produce single-stranded parental (+) strand 105a and single-stranded parental (-) strand 105b. As used herein, the terms "target" and "parental" are used interchangeably when they refer to nucleic acid strands. In addition, as used herein, single-stranded DNA target fragments can be interchangeably referred to as "DNA templates", which refers to a polynucleotide strand to which a complementary polynucleotide can hybridize or be synthesized by a nucleic acid polymerase, e.g., in a primer extension reaction.
[0047] Then, the method can include step D: performing a first primer extension reaction. The first primer extension reaction is guided by an extension oligonucleotide (i.e., primer) that hybridizes with a DNA template using a first DNA polymerase. In some cases, the extension oligonucleotide can hybridize to a region in the adaptor sequence. The first primer extension reaction produces a sample of double-stranded DNA fragments, each fragment including a newly synthesized first complementary copy strand (i.e., first daughter strands 107b and 107b) that hybridizes (i.e., couples) with a target fragment template (i.e., parental strands 105a and 105b). In some cases, the first DNA polymerase is a high-fidelity DNA polymerase. In this step, the sample of double-stranded DNA fragments differs from the sample of DNA target fragments in step A in that it contains complementary copy strands synthesized in vitro. The primer extension reaction can be carried out under conditions where the resulting complementary copy strands are "native" strands because they do not include the modified nucleobases of interest present in the target strands. For example, in this description, the first complementary copy strands incorporate native cytosine residues at the positions of methylated cytosine residues in the corresponding target strands. As used herein, the term "native" refers to a nucleobase, nucleotide, or polynucleotide that is similar to the relevant modified nucleobase, nucleotide, or polynucleotide except for a specific modification to the modified nucleobase, nucleotide, or polynucleotide. Thus, in certain aspects, each modified nucleobase, nucleotide, or polynucleotide can have a similar native nucleobase, nucleotide, or polynucleotide, and vice versa.
[0048] In some cases, prior to performing step (D) of the first primer extension reaction, the target fragment template is immobilized on a solid support, such as Figure 2A shown. As shown herein, the newly synthesized complementary copy strands are not immobilized on the solid support and can be physically separated from the immobilized template strands when the double-stranded DNA fragments are denatured. In other cases, as Figure 2B shown, an oligonucleotide complementary to the template strand (e.g., to the adaptor sequence) is immobilized on the solid support and is capable of "capturing" the template strand via hybridization. After capturing the target fragment, the first primer extension reaction can be carried out using the hybridized oligonucleotide as a primer to produce first complementary copy strands that are also immobilized on the solid support. In this case, denaturation of the resulting double-stranded DNA fragments will release the template strands from the solid support while retaining the complementary copies.
[0049] Then, the method can include step E: treating a sample of double-stranded DNA fragments with a DNA glycosylase capable of excising a modified nucleobase of interest (e.g., 5-mC as described herein). As used herein, the term "excise" means to cleave the N-glycosidic bond between the sugar and the base of a nucleotide. Excising a modified nucleobase of interest creates an abasic site (e.g., apurinic or apyrimidinic, AP site) at each position of the modified nucleobase of interest in the DNA target fragment. In some cases, more than one DNA glycosylase or other enzyme can be used to generate abasic sites. DNA glycosylases can also be engineered to inactivate functions that are not suitable for the desired result. For example, the lyase activity of the enzyme can be selectively inactivated while maintaining the glycosylase activity. It is noted that the first complementary copy strand is resistant to DNA glycosylase treatment such that the sites of its native nucleobases are not converted to abasic sites.
[0050] As used herein, the term "converted", when used in reference to a DNA target fragment, refers to a DNA target fragment or a portion thereof that has been treated under conditions sufficient to excise a modified nucleobase of interest to generate an abasic site in an otherwise continuous polynucleotide strand. This process is also referred to herein as "conversion of a modified nucleobase to an abasic site". Compared to existing epigenetic detection methods that rely on chemical conversion of native nucleobases to distinguish native and modified bases (e.g., bisulfite conversion of native cytosine), the method of the present invention provides the advantage of selective enzymatic cleavage of modified nucleobases while leaving native nucleobases unchanged. Thus, compared to bisulfite conversion-based methods, the overall damage to the DNA target fragment is not as extensive and the complexity of the genetic code is not significantly reduced.
[0051] Then, the method can include step F: denaturing the sample of double-stranded DNA fragments to release the converted parental DNA template strands 105a and 105b. As described, in some cases, the DNA template is immobilized on a solid support prior to the first primer extension reaction to enable separation from the first complementary copy strand, which partitions into solution upon denaturation. In other cases, the first complementary copy strand remains on the solid support such that the DNA target fragment can partition into solution upon denaturation. After step F, the DNA template strand and the first complementary copy strand are no longer coupled. As used herein, the term "coupled" is well known to those of skill in the art and refers to the process by which two nucleic acid strands bind together. Coupling is achieved, for example, by forming hydrogen bonds between the DNA template strand and their complementary copy strands. Thus, in the context of the present disclosure, the terms "hybridized" and "hybridization" fall within the definitions of "coupled" and "couple", respectively. That is, for example, a complementary copy of a DNA template can be coupled to the template by hybridization.
[0052] Then, the method can include step G: performing a second primer extension reaction. The second primer extension reaction is directed by an extension oligonucleotide that hybridizes, using a second DNA polymerase, to a region in an adaptor sequence of, for example, a DNA template to produce second complementary copies 109a and 109b of the DNA target strand template. The second DNA polymerase is selected for its ability to synthesize a complementary copy strand past (e.g., through and beyond) the position of an abasic site in the target fragment template. A DNA polymerase that exhibits this property may be referred to as a "bypass polymerase" and can include a translesion DNA polymerase. As discussed with reference to step D, in certain embodiments, either the DNA template strand or the second complementary strand can be selectively immobilized on a solid support to enable purification of the second complementary strand from the template strand.
[0053] According to the present invention, under the extension conditions used in this step, the nucleobase incorporated into the daughter strand at the position opposite the abasic site in the parental template does not form a canonical Watson-Crick base pair with the original, unmodified nucleobase. In Figure 1C the example shown, the nucleotide incorporated opposite the abasic site in the template strand is identified as "non-G" because G would typically base pair with 5-mC (the modified nucleobase of interest in this case). Thus, "non-G" is any nucleobase other than G, such as any one of adenine (A), cytosine (C), or thymine (T).
[0054] In some cases, the second DNA polymerase can be selected based on its substrate specificity and preference for incorporating a preferred nucleotide at the position opposite the abasic site in the modified template strand. For example, a DNA polymerase with a known preference for incorporating opposite an abasic site in the template a dATP will be suitable for detecting modified cytosine in the target fragment because "A" does not typically base pair with "C". Several DNA polymerases are known in the art to exhibit specific preferences for nucleotide incorporation at abasic sites, as further discussed herein.
[0055] Then, the method can include step H: determining the nucleotide sequences of the first and second complementary copy strands. A variety of sequencing platforms and methods are suitable for the practice of the present invention. In one example, the sequencing method is nanopore-based "amplicon sequencing" See, for example, Applicant's U.S. Patent Nos. 7,939,259 and 10,301,345 and Published Application Nos. WO2020 / 172,479 and WO2020 / 236,526, which are incorporated herein by reference in their entirety.
[0056] Then, the method can include step I: comparing sequence reads of the first complementary copy strand and the second complementary copy strand to determine the position of the modified nucleobase of interest in the original DNA target fragment (e.g., using bioinformatics analysis tools well-known in the art). The first complementary strand serves as a reference sequence as it encodes the genetic information of the DNA target fragment. In contrast, the second complementary strand encodes the epigenetic information of the DNA target fragment. A difference (e.g., a base substitution) in the sequences of the first complementary copy strand and the second complementary copy strand at a specific position indicates the position of the modified nucleobase of interest in the DNA target fragment sequence. In Figure 1D the illustrated example, a "non-G" detected in the second complementary strand at the position identical to the "G" in the first complementary strand indicates that the DNA target fragment initially contained a 5-mC residue at this position in the opposite strand.
[0057] In some cases, the method of the present invention can include an additional step (step G): stabilizing the abasic sites generated in the transformed DNA template before generating the second complementary copy. As Figure 3A summarized in the illustration in, abasic sites in DNA are known in the art to exist as an equilibrium mixture of two structural forms: (I) a closed-ring hemiacetal 301 and (II) an open-ring aldehyde hydrate 303. The open-ring aldehyde 303 is a highly reactive compound. Thus, abasic residues in a DNA fragment are converted to strand breaks by a β-elimination reaction, in which the 3'-phosphodiester bond in the open-ring aldehyde form hydrolyzes to generate a 3'-terminal unsaturated sugar and a terminal 5'-phosphate. The presence of nucleophilic molecules in the environment (including thiols, amines, polyamines, and basic proteins) further favors this undesired reaction. As will be apparent to those skilled in the art, strand breaks are detrimental as they prevent replication of the target fragment and result in loss of information.
[0058] To overcome this problem, in certain embodiments, the methods disclosed herein can include using a stabilizer that prevents chemical degradation of the open-ring aldehyde 303 and subsequent strand breaks. As Figure 3B shown, in one case, the stabilizer can be a chemical that reacts covalently with the abasic site to form a stable adduct 305. As used herein, the term "adduct" refers to the product of the direct covalent addition of two or more different molecules (resulting in a single reaction product containing all the atoms of all the components) and is thus a different molecular species. In other cases, the stabilizer can be a soluble buffer additive or other physicochemical reaction conditions that do not react covalently with the abasic site.
[0059] Further details regarding the above methods and embodiments are provided below.
[0060] Unless otherwise indicated, the practice of the present invention will employ conventional techniques within the skill of the art, such as molecular biology, microbiology, recombinant DNA, etc. These techniques are well explained in the literature. See, for example, Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989), OLIGONUCLEOTIDE SYNTHESIS (M.J. Gait ed., 1984), the series METHODS IN ENZYMOLOGY (Academic Press, Inc.), CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F.M. Ausubel, R. Brent, R.E. Kingston, D.D. Moore, J.G. Siedman, J.A. Smith and K. Struhl, eds., 1987).
[0061] DNA Sample / DNA Target Fragment
[0062] In one aspect, DNA is obtained or provided from a biological sample. The DNA obtained or provided from a biological sample can be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof.
[0063] The DNA sample can be obtained from a patient or subject, from an environmental sample, or from an organism of interest. In embodiments, the DNA sample is extracted, purified, or derived from a cell or collection of cells, a body fluid, a tissue sample, an organ, and / or an organelle. In some embodiments, the sample DNA is genomic DNA.
[0064] In certain cases, genomic DNA and mitochondrial DNA can be obtained separately from the same biological sample or source. Many different methods and techniques can be used to isolate genomic DNA and mitochondrial DNA. Generally, such methods involve disruption and lysis of the starting material, followed by removal of proteins and other contaminants, and finally recovery of the DNA. Removal of proteins can be achieved, for example, by digestion with proteinase K followed by salting out, organic extraction, gradient separation, or binding the DNA to a solid support (anion exchange or silica technology). Mitochondrial DNA can be isolated similarly after initial isolation of the mitochondria. The DNA can be recovered by precipitation using ethanol or isopropanol. There are also commercial kits available for isolating nuclear DNA or mitochondrial DNA. The choice of method depends on many factors, such as the sample volume, the amount and molecular weight of DNA required, the purity required for downstream applications, and time and cost.
[0065] In certain embodiments, the methods of the present disclosure utilize mild enzymatic and chemical reactions, avoiding substantial degradation associated with methods such as bisulfite sequencing. Accordingly, the methods can be used for the analysis of low-input samples, such as cell-free circulating DNA, circulating tumor DNA, and single-cell analysis.
[0066] In some embodiments, the DNA sample is cell-free circulating DNA (cfDNA), i.e., DNA that is found in blood and is not present within cells. cfDNA can be isolated from blood or plasma using methods known in the art. Commercial kits are available for isolating cfDNA, including, for example, the Circulating DNA Kit (Qiagen). The DNA sample may be from an enrichment step, including but not limited to antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0067] In some cases, the isolated DNA is fragmented into multiple shorter double-stranded DNA target fragments. Generally, DNA fragmentation can be carried out physically or enzymatically.
[0068] For example, physical fragmentation can be carried out by sonication, ultrasonication, microwave radiation, or hydrodynamic shearing. Sonication and ultrasonication are the primary physical methods used to shear DNA. For example, the instrument (Woburn, MA) is an acoustic device used to break DNA into 100 bp - 5 kb. Covaris also produces tubes (gTubes) that will process samples of 6 - 20 kb for Mate-Pair libraries. Another example is (Denville, NJ), which is an ultrasonication device used to shear chromatin, DNA, and disrupt tissues. Small amounts of DNA can be sheared to lengths of 150 bp - 1 kb. The from Digilab (Marlborough, MA) is another example, and it utilizes hydrodynamic forces to shear DNA. Nebulizers (such as those manufactured by Life Technologies (Grand Island, NY)) can also be used to atomize a liquid using compressed air, shearing DNA into fragments of 100 bp - 3 kb within a few seconds. Since nebulization may result in sample loss, in some cases, it may not be an ideal fragmentation method for limited amounts of sample. For smaller sample volumes, ultrasonication and sonication may be better fragmentation methods because the entire amount of DNA from the sample can be more effectively retained. Other physical fragmentation devices and methods known or developed can also be used.
[0069] Various enzymatic methods can also be used to fragment DNA. For example, DNA can be treated with DNase I, or a combination of maltose-binding protein (MBP)-T7 Endo I and a non-specific nuclease such as Vibrio vulnificus nuclease (Vvn). The combination of the non-specific nuclease and T7 Endo acts synergistically to produce non-specific nicks and back nicks, generating fragments that are separated from the nick site by 8 or fewer nucleotides. In another example, DNA can be treated with dsDNA (NEB, Ipswich, MA). dsDNA Fragmentase generates dsDNA breaks in a time-dependent manner to produce DNA fragments of 50 - 1,000 bp depending on the reaction time. NEBNext dsDNA Fragmentase contains two enzymes, one that randomly generates nicks on dsDNA and another that recognizes the nick site and cleaves the DNA strand opposite the nick, producing dsDNA breaks. The resulting DNA fragments contain short overhangs, 5'-phosphate, and 3'-hydroxyl groups.
[0070] In some cases, a DNA sample is fragmented into target fragments within a specific size range. For example, a DNA sample can be fragmented into fragments in the range of about 25 - 100 bp, about 25 - 150 bp, about 50 - 200 bp, about 25 - 200 bp, about 50 - 250 bp, about 25 - 250 bp, about 50 - 300 bp, about 25 - 300 bp, about 50 - 500 bp, about 25 - 500 bp, about 150 - 250 bp, about 100 - 500 bp, about 200 - 800 bp, about 500 - 1300 bp, about 750 - 2500 bp, about 1000 - 2800 bp, about 500 - 3000 bp, about 800 - 5000 bp, or any other size range within these ranges. For example, a DNA sample can be fragmented into fragments of about 50 - 250 bp. In some cases, the fragments can be larger or smaller than about 25 bp.
[0071] A DNA target fragment can be any DNA fragment from a biological sample that has a sequence of interest, which may or may not include epigenetic modifications or DNA damage to one or more nucleobases. In some aspects, the DNA target fragment can include cytosine modifications (i.e., 5-mC, 5-hmC, 5-fC, and / or 5-caC). The DNA target fragment can be a single DNA molecule in the sample or an entire population (or subset) of DNA molecules in the sample that have, for example, cytosine modifications. The DNA target fragment can contain multiple DNA sequences such that the methods described herein can be used to generate a library of DNA target fragments that can be analyzed individually (e.g., by determining the sequence of a single target) or in groups (e.g., by multiplexed DNA sequencing methods).
[0072] In an embodiment, the methods described herein include the step of adding an adaptor DNA molecule to a double-stranded DNA target fragment. An adaptor DNA or DNA linker is a short, chemically synthesized single-stranded or double-stranded oligonucleotide that can be ligated to one or both ends of other DNA molecules. Double-stranded adaptors can be synthesized such that each end of the adaptor has blunt ends or 5' or 3' overhangs (i.e., sticky ends). The DNA adaptor is ligated to the DNA target fragment to provide sequences for, for example, primer extension reactions and sequencing reactions with complementary primers and / or for bioinformatics analysis (e.g., clustering related sequences into families based on shared unique molecular identifier barcodes, UMIs).
[0073] Prior to ligation of the adaptor, the ends of the DNA fragment can be made ligation-ready. For example, by end repair and generating blunt ends with 5' phosphate groups. Fragmented DNA can be made blunt-ended by many methods known to those skilled in the art. In a particular method, the ends of the fragmented DNA are "polished" with T4 DNA polymerase and Klenow polymerase, a procedure well-known to those skilled in the art, and then phosphorylated with polynucleotide kinase. Then, a single 'A' deoxynucleotide is added to the two 3' ends of the DNA molecule using Taq polymerase or Klenow exo minus polymerase, generating a one-base 3' overhang that is complementary to the one-base 3' 'T' overhang on the ends of the adaptor duplex.
[0074] In some cases, the linker can include two partially complementary oligonucleotides such that they hybridize to form a region of double-stranded sequence, but also retain regions of single-stranded, non-hybridized sequence. The regions of single-stranded sequence can include "universal" oligonucleotide binding sequences such that all target fragments in the library can bind the same oligonucleotide, which can be a capture oligonucleotide to localize the target fragment to a solid support, an oligonucleotide primer for primer extension reactions, a PCR primer, a sequencing primer, or a combination thereof. In certain cases, the linker can include two regions of single-stranded, non-hybridized sequence (i.e., a first 5' single-stranded region and a second 3' single-stranded region). This configuration is known in the art as a "Y" - shaped linker. The first and second single-stranded regions of the Y - shaped linker are not complementary and can contain different primer hybridization sequences and other features.
[0075] The portions of the two single-stranded regions of the linker typically include at least 10, or at least 15, or at least 20 consecutive nucleotides on each strand. The lower limit of the length of the single-stranded region is generally determined by function, such as the need to provide a suitable sequence for primer binding for primer extension, PCR, and / or sequencing. In theory, there is no upper limit to the length of the single-stranded region, except that it is generally advantageous to minimize the total length of the linker, e.g., to facilitate separation of unbound linker from the double-stranded DNA target fragments to which the linker is ligated. Thus, preferably, the length of the single-stranded region on each strand should be less than 50, or less than 40, or less than 30, or less than 25 consecutive nucleotides.
[0076] The double-stranded region of the linker is a short double-stranded region, typically containing 5 or more consecutive base pairs, formed by annealing two partially complementary polynucleotide strands. Generally, it is advantageous for the double-stranded region to be as short as possible without loss of function. "Function" herein refers to the double-stranded region forming a stable duplex under the standard reaction conditions of an enzyme-catalyzed nucleic acid ligation reaction.
[0077] The exact nucleotide sequence of the linker is generally not important for the present invention and can be chosen by the user such that the desired sequence elements are ultimately included in the common sequence of the library of double-stranded DNA target fragments ligated to the linker. Additional sequence elements can be included, e.g., to provide a binding site for a primer that will ultimately be used to sequence the complementary copy strand of the DNA target fragment. The linker can further include "tag" sequences, unique molecular identifiers (UMIs), and / or sample identifier sequences, which can be used to label, track, and distinguish target fragments and their complementary copies derived from a particular source. The general characteristics and uses of such sequences are well known in the art.
[0078] The ends of the single-stranded regions of the adaptor can be biotinylated or have another functional group that enables it to be captured or immobilized on a surface, such as a solid support. Alternative functional groups other than biotin are known in the art and are described, for example, in the patent application WO2020 / 172479 entitled "Solid Phase Synthesis Methods and Devices for Xpandomers for Single Molecule Sequencing" published by the applicant, which is incorporated herein by reference in its entirety.
[0079] The "ligation" of the adaptor to the 5' and 3' ends of each fragmented double-stranded nucleic acid target fragment involves ligating the two polynucleotide chains of the adaptor to the double-stranded target polynucleotide such that a covalent bond is formed between the two chains of the two double-stranded molecules. Preferably, this covalent ligation occurs by forming a phosphodiester bond between the two polynucleotide chains, but other means of covalent ligation (e.g., non-phosphodiester backbone bonds) can also be used. However, the basic requirement is that the covalent bond formed in the ligation reaction allows polymerase read-through such that the resulting construct can be replicated in a primer extension reaction using a primer that binds to a sequence in the region of the adaptor-target construct derived from the adaptor molecule.
[0080] In some cases, the adaptor and the DNA target fragment can be incubated with a ligase to covalently link the adaptor and the DNA target fragment. Ligases catalyze the formation of phosphodiester bonds between juxtaposed 5' phosphate and 3' hydroxyl termini in double-stranded DNA or RNA. The enzyme will ligate blunt ends and sticky ends and repair single-strand nicks in double-stranded DNA. An exemplary ligase is T4 ligase, which is the most commonly used enzyme in cloning. Another ligase that can be used is E. coli DNA ligase, which preferentially ligates sticky double-stranded DNA ends but is also active on blunt-ended DNA in the presence of Ficoll or polyethylene glycol. Another ligase that can be used is DNA ligase Ilia, which is known to function in mitochondria.
[0081] Prior to further processing of the adaptor-target construct, a purification step can be performed on the product of the ligation reaction to remove unbound adaptor molecules.
[0082] Ligation of the adaptor to the two free ends of the double-stranded DNA target fragment generates a library of adaptor-ligated double-stranded DNA target fragments, where the adaptor is located at the 5’ and 3’ ends of the target.
[0083] There are several standard methods for separating the strands of the adaptor-ligated double-stranded DNA target fragment by denaturation, including heat denaturation, or chemical denaturation in 100 mM sodium hydroxide solution or formamide solution. The pH of the single-stranded DNA fragment solution can be neutralized by adjusting with an appropriate acid solution or, preferably, by buffer exchange through a size exclusion chromatography column pre-equilibrated in a buffer solution.
[0084] First Complementary Copy Strand
[0085] In the embodiments disclosed herein, a single-stranded DNA target fragment (i.e., the parental strand) provides a template for synthesizing a first complementary copy (i.e., the first daughter strand) of the target fragment via a primer extension reaction. The term "primer extension reaction" may be used interchangeably herein with the term "nucleic acid polymerization reaction" and refers to an in vitro method of preparing a new nucleic acid strand or extending an existing nucleic acid in a template-dependent manner. The first complementary copy strand is synthesized by extending an oligonucleotide primer with a first DNA polymerase such that the first complementary copy of the template strand extends in the 3' direction of the oligonucleotide primer.
[0086] In embodiments where the DNA target fragment is double-stranded, one or both strands can serve as templates for the primer extension reaction. For example, when one strand ("sense" strand) is used as a template, a complementary copy that is complementary to the sense strand is generated. Similarly, when the antisense strand is used as a template, a complementary copy that is complementary to the antisense strand is generated. When both strands are used as templates, separate complementary copies are generated for each of the sense and antisense strands. In a preferred embodiment, each strand of the double-stranded DNA target fragment is a template nucleic acid.
[0087] As used herein, the term "complementary" refers to nucleic acid sequences that are capable of forming Watson-Crick base pairs. For example, the complementary sequence of a first sequence is a sequence that is capable of forming Watson-Crick base pairs with the first sequence. The term "complementary" does not necessarily mean that the sequence is fully complementary to its complementary strand, but the term can mean that the sequence is partially complementary to it. Thus, in some embodiments, complementarity encompasses sequences that are complementary along the entire length of the sequence or a portion thereof. For example, two sequences can be complementary to each other along at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the sequence length. As used herein, the term "sequence" encompasses, but is not limited to, nucleic acid sequences, polynucleotides, oligonucleotides, probes, primers, primer-specific regions, and target-specific regions. Despite any mismatches, two sequences should have the ability to selectively hybridize to each other under appropriate conditions.
[0088] Primer extension can be carried out by any method that allows polymerase-based extension of a primer that anneals (i.e., hybridizes) to a single-stranded DNA target fragment. In some embodiments, simple primer extension involves adding a primer and a first DNA polymerase to the target DNA fragment under conditions that allow primer hybridization and polymerase primer extension. Of course, such reactions include the necessary nucleotides, buffers, and other reagents known in the art for primer extension. Importantly, the nucleotides included in the primer extension reaction are "native", i.e., unmodified nucleotides, and thus, the first complementary copy strand will not include modifications to the nucleobases of interest. The first complementary copy strand is generated to encode and preserve the genetic sequence of the DNA target strand.
[0089] Many methods for detecting primer extension products are known. In some embodiments, the primer is detectably labeled (e.g., at its 5' end or otherwise positioned so as not to interfere with 3' extension of the primer), and after primer extension, the length and / or quantity of the labeled extension product is detected by detecting the label.
[0090] In certain embodiments, the primer used in the primer extension reaction anneals to a primer binding sequence (in one strand) in the single-stranded region of the adaptor. As used herein, the term "anneal" refers to the sequence-specific binding / hybridization of a primer to a primer binding sequence in the adaptor region of an adaptor-linked DNA target fragment under the conditions of the primer annealing step of an initial primer extension reaction. Primer annealing conditions are well known in the art (see, e.g., Sambrook et al., 2001, Molecular Cloning, A Laboratory Manual, 3rd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor Laboratory Press, NY; Current Protocols, edited by Ausubel et al.).
[0091] In preferred embodiments, the first DNA polymerase is a high-fidelity DNA polymerase. The fidelity of a DNA polymerase results in the accurate replication of the desired template. Specifically, this involves multiple steps, including the ability to read the template strand, select the appropriate nucleoside triphosphates, and insert the correct nucleotide at the 3' primer terminus such that Watson-Crick base pairing is maintained. In addition to effectively discriminating between correct and incorrect nucleotide incorporation, some DNA polymerases have 3'→5' exonuclease activity. This activity, known as "proofreading", is used to excise misincorporated single nucleotides and then replace them with the correct nucleotide.
[0092] In certain embodiments, suitable high-fidelity DNA polymerases for practicing the invention include the KAPA HiFi DNA polymerase commercially available from Roche Diagnostics Corp., the high-fidelity DNA polymerase and the engineered Pfu DNA polymerase, such as Pfu-X, commercially available from New England Biolabs, Inc. and Jena Biosciences, respectively.
[0093] Solid-Phase Synthesis
[0094] In certain embodiments, the first primer extension reaction can be carried out on a solid support. Thus, in a further aspect, the present invention provides a method for solid-phase nucleic acid synthesis using an adaptor-linked DNA target fragment having known sequences at its 5' and 3' ends (e.g., sequence features designed into the adaptor).
[0095] The terms "solid support", "solid state", "solid phase" and "substrate" are used interchangeably herein and refer to a material or group of materials having one or more rigid or semi-rigid surfaces. In many embodiments, at least one surface of the solid support will be substantially flat, such as the surface of a polymeric microfluidic card or chip. In some embodiments, it may be desirable to physically separate regions of the card or chip, e.g., by etching channels, trenches, cavities, raised areas, pins, etc., for different reactions. According to other embodiments, the solid support will take the form of insoluble beads, resins, gels, membranes, microspheres or other geometric configurations composed of, e.g., controlled pore glass (CPG) and / or polystyrene.
[0096] The present invention encompasses solid-phase synthesis methods in which a capture moiety is immobilized on a solid support. In certain cases, the capture moiety includes a first end covalently bound to the solid support and a second end providing a functional group capable of binding to the 5' end of a single-stranded adaptor-linked DNA target fragment. In such cases, the single-stranded DNA target fragment is immobilized on the solid support while the complementary copy strand is not immobilized on the solid support. In other cases, the capture moiety includes an extension oligonucleotide capable of hybridizing to the 3' end of the adaptor-linked target fragment. The single-stranded adaptor-linked DNA target fragment is hybridized to the extension oligonucleotide and a primer extension reaction is carried out. In such cases, only the complementary copy strand is immobilized on the solid support. These alternative solid-phase synthesis configurations are shown in Figure 2A and Figure 2B .
[0097] As used herein, the term "immobilized" refers to the association, attachment or binding of a molecule (e.g., a linker, adaptor or oligonucleotide) to a support in a manner that provides a stable association under the conditions of the extension, amplification, ligation and other processes described herein. Such binding can be covalent or non-covalent. Non-covalent binding includes electrostatic, hydrophilic and hydrophobic interactions. Covalent binding is the formation of a covalent bond characterized by the sharing of electron pairs between atoms. Such covalent binding can occur directly between the molecule and the support or can be formed through a crosslinker or by including specific reactive groups on the support or molecule or both. Covalent attachment of a molecule can be achieved using a binding partner immobilized on the support, such as avidin or streptavidin, and non-covalent binding of a biotinylated molecule to avidin or streptavidin. Immobilization can also involve a combination of covalent and non-covalent interactions.
[0098] Any suitable covalent linking means known in the art can be used for these purposes. The linking chemistry chosen will depend on the nature of the solid support and any derivatization or functionality imposed on it. The extended oligonucleotide can include moieties that can be non-nucleotide chemical modifications to facilitate attachment. Some exemplary embodiments of suitable surface chemistries include conventional streptavidin / biotin interaction chemistries and involve, for example, functionalizing the solid support with a linker moiety comprising a terminal biotin moiety. In this embodiment, the 5' end of the single-stranded DNA fragment (or oligonucleotide) binds to the linker moiety. Attachment is mediated by the streptavidin moiety provided by the 5’ end of the single-stranded DNA fragment. The linker moieties disclosed herein can have a length sufficient to link the single-stranded DNA fragment to the support such that the support does not significantly interfere with the primer extension reaction.
[0099] Alternatively, immobilization of the capture moiety or oligonucleotide (e.g., extended oligonucleotide) onto the solid support can be achieved by covalently linking the capture oligonucleotide to the solid support via a click reaction. In this embodiment, the covalent linkage can be mediated by a maleimide-PEG-alkyne linker crosslinked to the solid support. The alkyne moiety provided by the end of the linker distal to the substrate is capable of reacting with the azide group provided by the 5’ end of the capture oligonucleotide. Methods for functionalizing the solid support with maleimide-linker polymers are provided in applicant's published patent application No. WO 2020 / 172479, which is incorporated herein by reference in its entirety.
[0100] In some cases, the bond between the capture moiety and the solid support is cleavable such that the primer extension product can be released from the support after synthesis. Cleavable linkers and methods for cleaving such linkers are known and can be utilized by the knowledge of those skilled in the art in the methods provided. For example, a cleavable linker can be cleaved by an enzyme, catalyst, compound, temperature, electromagnetic radiation, or light. Optionally, the cleavable linker comprises a moiety cleavable by β-elimination hydrolysis, a moiety cleavable by acid hydrolysis, an enzymatically cleavable moiety, or a photocleavable moiety. In some embodiments, suitable cleavable moieties are photocleavable (PC) spacers or phosphoramidites available from Glen Research.
[0101] Glycosylase-Mediated Excision of Modified Nucleobases
[0102] In one aspect, the method of the present invention includes the step of treating the double-stranded DNA product of a first primer extension reaction with a DNA glycosylase to specifically excise a modified base of interest. Many DNA glycosylases are known in the art and target a wide range of specific modified nucleobases and DNA damage elements, including sequence mismatches and a broad range of epigenetic modifications. Exemplary epigenetic modifications detectable by the method include, but are not limited to, 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-carboxylcytosine (5-caC), 5-formylcytosine (5-fC), 8-oxo-7,8-dihydroguanine (oxoG), uracil, methyladenine (mA), and the like.
[0103] There are two main classes of DNA glycosylases: monofunctional enzymes and bifunctional enzymes. Monofunctional glycosylases only have glycosylase activity and cleave the N-glycosidic bond that links the damaged or modified nucleobase to the sugar-phosphate backbone of DNA. All DNA glycosylases cleave the glycosidic bond, but they differ in their specificity for base substrates and reaction mechanisms. Bifunctional glycosylases also have apurinic or apyrimidinic site (AP) lyase activity, which enables them to cleave the phosphodiester bond of DNA at the site of base damage, creating a single-strand break.
[0104] A non-limiting list of exemplary DNA glycosylases that can be used in the method of the present invention is listed in Table 1. In some cases, one or more of the DNA glycosylases listed in Table 1 can be used in the method to excise a modified base of interest from a DNA target fragment. Although the present disclosure specifically identifies the selected DNA glycosylases, it should be understood that any suitable DNA glycosylase can be used to perform the base excision step of the method.
[0105] Table 1 DNA Glycosylases
[0106]
[0107]
[0108] In one embodiment, the method utilizes a DNA glycosylase that acts directly on 5-mC, i.e., a glycosylase capable of hydrolyzing the glycosidic bond between the 5-mC residue and the sugar-phosphate backbone. For example, a suitable DNA glycosylase that directly excises 5-mC can be a member of the DEMETER (DME) family of DNA glycosylases, such as DME, ROS1, or DMEL. The DME gene of Arabidopsis thaliana encodes a 1,729-amino acid protein with a central DNA glycosylase domain (amino acids 1167-1368), which domain includes a helix-hairpin-helix (HhH) motif. The HhH motif in DME catalyzes the excision of 5-mC (see, e.g., Choi et al., 2002. Cell 110:33-42). In certain embodiments, the DME glycosylase can be a variant that includes amino acids 1167-1368 but lacks certain other regions of the protein.
[0109] In some cases, a suitable DNA glycosylase that acts directly on 5-mC can be an ortholog of DME. As used herein, the term "ortholog" means one of two or more homologous gene sequences found in different species. Table 2 lists an exemplary list of DME orthologs that can be used according to the present invention.
[0110] Table 2. DME orthologs
[0111]
[0112]
[0113]
[0114] In the case where the DNA glycosylase is a bifunctional enzyme, the glycosylase (e.g., DME or its ortholog) can be mutated to inactivate the lyase activity while still retaining the glycosylase activity, as Figure 4A shown. The reaction mechanism of bifunctional DNA glycosylases is well known in the art (see, e.g., Scharer and Jiricny. 2001. Bioessays 23:270-281). In some cases, a conserved aspartic acid obtains a proton from a conserved lysine residue, which lysine residue attacks the C1' carbon of the deoxyribose ring, generating a covalent DNA-enzyme intermediate. A β or γ elimination reaction releases the enzyme from the DNA and cleaves one of the phosphodiester bonds in the phosphodiester bond. Mutant forms of DME in which the invariant aspartic acid at position 1304 or the lysine at position 1286 has been altered (e.g., variant D1304N or K1286Q) have been shown to reduce DNA glycosylase activity while retaining enzyme structure and stability (see, e.g., Fromme et al., 2004. Nature 427:652-656).
[0115] The present invention also contemplates other mutations that inactivate or optimize suitable features of DNA glycosylases. For example, DNA glycosylases can be engineered to increase their stability and / or solubility. DNA glycosylases can also be engineered to optimize the desired substrate specificity.
[0116] In certain embodiments, thymine DNA glycosylase (TDG) can be used to excise its known targets 5-carboxylcytosine (5-caC) and 5-formylcytosine (5-fC). In further embodiments, as Figure 4B shown, TDG can be used to recognize 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC), which are modified bases that it does not specifically recognize. For example, prior to treatment with TDG, the DNA target fragment can also be treated with ten-eleven translocation (TET) enzymes. The TET family of proteins includes three human proteins (TET1, TET2, and TET3) and are cytosine oxygenases that catalyze the conversion of 5-methylcytosine (5-mC) to 5-hydroxymethylcytosine (5-hmC). 5-hmC can be further oxidized by TET proteins to 5-formylcytosine (5-fC) and 5-carboxylcytosine (5-caC) (see, e.g., Parker et al. 2019. Biochemistry 58:450-467). In another example, a suitable TET enzyme can be any TET ortholog isolated from Naegleria gruberi (see, e.g., Hashimoto et al. 2014. Nature 506(7488):391-395). Thus, in certain embodiments, TDG can be used to excise any existing 5-caC and 5-fC modified bases present in a DNA target fragment that has also been treated with a TET enzyme.
[0117] According to the present invention, other similar methods for altering the selective excision of modified bases are also possible. For example, thymine DNA glycosylase (TDG) and uracil DNA glycosylase (UDG) can be used in a similar method to detect the same bases discussed above.
[0118] The base excision processes discussed herein can be carried out using purified enzymes, which can be recombinant enzymes that contain heterologous tags to facilitate purification. Protein tags are well known in the art and include, for example, terminal polyhistidine tags that can be purified via immobilized metal affinity chromatography (IMAC). In some cases, it may be necessary to include more than one protein purification step. For example, the glycosylases used in the methods disclosed herein should preferably be free of contaminating nucleic acids. In some cases, the protein purification step can include one or more of size exclusion chromatography, ion exchange chromatography, affinity chromatography, heparin adsorption chromatography, etc.
[0119] Of course, the base excision reaction will include a suitable buffer, cofactors, additives, and an amount of purified DNA glycosylase sufficient to effect the desired base excision reaction such that the modified nucleobase of interest in the DNA target fragment is excised to generate an abasic site. Exemplary base excision reactions are described in Example 1.
[0120] After treatment with DNA glycosylase, the double-stranded DNA fragment will be asymmetrically altered. Notably, the DNA template strand will lack a nucleobase at the position of the original modified base of interest. In contrast, the first complementary copy strand remains unchanged (i.e., "unconverted") because the native nucleobases incorporated during the first primer extension reaction will resist glycosylation-mediated conversion to an abasic site.
[0121] Stabilization of Abasic Sites in DNA Target Fragments
[0122] Advantageously, in the method according to the invention, the abasic sites generated in the DNA target fragment can be protected from further degradation with a stabilizer. In certain embodiments, a suitable stabilizer can be a chemical that covalently binds to the abasic site to form a stable abasic adduct. As discussed with reference to Figure 3B it is known that certain aldehyde-reactive compounds react with the open-ring aldehyde form (II) of the abasic site to produce a stable open structure, which is referred to herein as an abasic adduct. Abasic adducts are refractory to enzymatic activity (e.g., lyase-mediated degradation) or chemical conditions that induce degradation (e.g., high pH). Some exemplary, non-limiting structural classes of aldehyde-reactive stabilizers are shown in Figure 5A and 5B and are described below. The reaction rates, stabilities, and sizes of the resulting protected adduct products vary for each class. The chemical nature of each abasic adduct product provides different chemoenzymatic properties in terms of stabilization duration and suitability as a template for DNA polymerase extension.
[0123] As Figure 5A shown, in one embodiment, a suitable stabilizer can be from a group of O-hydroxylamines (Compound IIIa), which are compounds known to react with the aldehyde group of the open-ring form of the abasic site (II) to produce a very stable class of oxime structures (Compound IVa) that are refractory to β-elimination by enzymatic activity (e.g., AP or dRp lyase) or by high pH.
[0124] In another embodiment, a suitable stabilizer can be from the group of hydrazides (Compound IIIb), which are a class of compounds that react with aldehyde (II) to form hydrazones (Compound IVb).
[0125] In another embodiment, suitable stabilizers may be from a group of tryptamines (compounds IIIc) which react with aldehydes (II) via a Pickett-Spengler annulation reaction to form tricyclic heterocycles (compounds IVc).
[0126] like Figure 5B As shown, in another embodiment, suitable stabilizers may be from a group of beta aminothiols (compound IIId) (eg, cysteine), which are a class of compounds that react with aldehydes (II) to form cyclic thiazolidines (compound IVd).
[0127] In another embodiment, suitable stabilizers may be from the group of alkyl hydrazines (Group IIIe), which are a class of compounds that react with aldehydes (II) to form alkyl hydrazones (compounds IVe).
[0128] In another embodiment, suitable stabilizers may be derived from a group of hydrazino-iso-Picot-Spengler indoles (Compound IIIf), which react with the abasic aldehyde (II) form to form a tricyclic structure (Compound IVf).
[0129] In another embodiment, suitable stabilizers may be from a group of methylaminooxy-iso-pict-Spengler indoles (group IIIg), which react with abasic aldehydes (II) to form tricyclic structures (compounds IVg).
[0130] In other cases, the stabilizer can be an agent that does not covalently react with the abasic site in the DNA target fragment, such as a reaction additive or other physicochemical reaction conditions. The following is a non-limiting list of exemplary stabilizers: 1. Aqueous buffers (e.g., water) without salt; 2. Alkaline buffers of various concentrations (e.g., buffers based on ammonia, NaOH, or other hydroxides); 3. Acidic buffers of various concentrations (e.g., buffers based on acetic acid, HCl, or nitric acid); 4. Urea; 5. Detergents (e.g., SDS, Tween, or Triton); 6. Solvents (e.g., acetonitrile, DMSO, formamide, DMF, or glycerol); 7. PEG and PEG variants; 8. Guanidine salts; and 9. Current or current pulses applied to the reaction. In some cases, any suitable combination of the aforementioned stabilizers can be used.
[0131] In certain embodiments, the chemistries described herein can be used to form stable abasic adducts during treatment of a DNA target fragment with one or more of a monofunctional DNA glycosylase, a bifunctional DNA glycosylase, or a bifunctional DNA glycosylase engineered to inactivate lytic enzyme activity.
[0132] In certain embodiments, the methods of the invention can utilize bifunctional DNA glycosylases to generate abasic sites that are stable and refractory to lyase-mediated backbone cleavage. In other words, the glycosylase activity can be uncoupled from the lyase activity of the bifunctional glycosylase by chemically "knocking out" the lyase activity. In some embodiments, this can be achieved by including one or more abasic stabilizers disclosed herein in the glycosylase reaction. As discussed, the stabilizer forms a stable adduct at the abasic site after modified base excision. Such abasic adducts are resistant to further lyase activity, such that strand excision does not occur at these sites. This phenomenon is referred to herein as the biochemical knockout or "hijacking" of DNA lyase activity.
[0133] The biochemical hijacking of DNA lyase activity is shown in simplified form in Figure 6A and 6B . Figure 6A depicts the native activity of an exemplary bifunctional DNA glycosylase (e.g., DEMETER) acting on 5-mC. After cleaving the N-glycosidic bond to release the methylated base, the enzyme forms a Schiff base intermediate (I) with the open-ring ribose moiety and proceeds to cleave the phosphodiester bond in the DNA backbone by a β-elimination reaction to generate a strand break (II). Figure 6B depicts knocking out the lyase activity with an aminooxyalkyl compound. As used herein, the term "aminooxyalkyl" is used to denote an O-alkylated derivative of hydroxylamine and has the general structure H 2 N-O-R, where R is an alkyl group. Here, during treatment of a DNA substrate with a DNA glycosylase, an exemplary aminooxyalkyl depicted as "H 2 N-O-R" is added. After enzyme-mediated cleavage of the N-glycosidic bond to release the modified base, the aminooxyalkyl reacts with the abasic site (I) to form a stable adduct (III), which prevents the enzyme from further interacting with the DNA substrate and, for example, cleaving the phosphodiester backbone.
[0134] Second Complementary Copy Strand
[0135] The methods described herein include the step of performing a second primer extension reaction to generate a second complementary copy (i.e., a second daughter strand) of the parental DNA template. This step is performed after cleavage of the modified base. Thus, the second complementary copy of the DNA template retains at least a portion of the epigenetic information encoded in the original DNA target fragment.
[0136] After glycosylase treatment, the asymmetrically altered DNA fragments are denatured using any suitable method recognized in the art, including acid-base denaturation (using, for example, acetic acid, HCL, or nitric acid), alkaline denaturation (using, for example, NaOH), solvent-based denaturation (using, for example, DMSO, formamide, guanidine, sodium salicylate, propylene glycol, or urea), or physical denaturation (using, for example, heating, beads, sonication, or radiation). The resulting single-stranded DNA template strand is then purified from the first complementary target strand. Purification of the transformed template strand population is facilitated by the solid-phase synthesis methods described herein, wherein one of the two parental strand populations and the daughter strand population is selectively immobilized on a solid support.
[0137] The second primer extension reaction is directed by an extension oligonucleotide that hybridizes to the DNA target template using a second DNA polymerase to produce a second double-stranded DNA fragment that includes a second complementary copy strand hybridized to the parental template strand. As described herein, the second primer extension reaction can be carried out on a solid support, wherein the parental template strand or the second daughter strand is selectively immobilized on the support.
[0138] The second DNA polymerase is selected because of its ability to synthesize the second complementary copy beyond the position of the abasic site in the transformed parental template. DNA polymerases that exhibit this property are known in the art and are referred to as, for example, "bypass" or "translesion" polymerases.
[0139] In some cases, the second DNA polymerase can be selected based on its activity to preferentially incorporate a specific nucleotide opposite the abasic site in the template. An object of the present invention is to generate a second complementary copy strand such that the nucleobase incorporated opposite the abasic site in the template does not form a Watson and Crick base pair with the modified nucleobase previously excised from the template. For example, in some cases, the modified base of interest is 5-mC. In this case, the second DNA polymerase is selected based on a preference for any nucleotide other than dGTP (i.e., "non-G") opposite the position where 5-mC has been converted to an abasic site. For example, the polymerase can preferentially incorporate dATP, dTTP, or dCTP at these sites.
[0140] It is known in the art that abasic sites represent the most common DNA damage in the genome and have mutagenic potential, leading to mutations common in human cancers. Although these lesions lack genetic information, it has been observed that adenine is the most efficiently inserted nucleobase during DNA polymerase bypass of abasic sites, a phenomenon known as the "A-rule". For DNA polymerases from families A (including human DNA polymerases γ and θ) and B (including human DNA polymerases α, ε, and δ), a strong preference for adenine (i.e., dATP) incorporation by DNA polymerases has been observed (see, e.g., Obeid et al. 2010. EMBO J. 29(10):1738 - 1747). In a preferred embodiment of the present invention, the second DNA polymerase will have a preference for incorporating opposite the abasic site in the template, particularly when the modified nucleobase of interest is a derivative of C (e.g., 5-mC).
[0141] A non-limiting list of exemplary second DNA polymerases is shown in Table 3.
[0142] Table 3 Exemplary Abasic Bypass DNA Polymerases
[0143] DNA Polymerase Species Family gp90(exo -) PaP1 Bacteriophage A Polα Homo sapiens A Klenow Fragment (exo -) Escherichia coli B Vent Polymerase (exo -) Hyperthermophile B Klentaq Chimera B RB69pol RB69 Bacteriophage B gp43 T4 Bacteriophage B Polδ Homo sapiens B Dpo4 Sulfolobus solfataricus Y Dpo4 Variant with Amino Acid Substitutions Sulfolobus solfataricus Y dinB (PolIV) Escherichia coli Y REV1 Homo sapiens Y umuC and umuD (PolV) Escherichia coli Y
[0144] In some cases, the second DNA polymerase can include a mixture of more than one DNA polymerase. For example, the mixture can contain a DNA polymerase that is capable of incorporating a nucleotide opposite the abasic site but cannot further extend the daughter strand, and another DNA polymerase that does have the ability to extend the daughter strand beyond the abasic site in the parental strand. In another example, the mixture can include a DNA polymerase having exonuclease activity. A combination of a bypass polymerase (e.g., DPO4 or its variants) and a polymerase having exonuclease activity (e.g., DPO1) can provide several advantages. For example, the exonuclease can provide proofreading activity, and the combination results in more efficient and accurate incorporation of the desired nucleotide by, for example, minimizing polymerase stalling and errors.
[0145] In some cases, the substrate preference of the bypass DNA polymerase at the abasic site can be optimized or directed by other methods of the present invention. For example, the DNA polymerase can be an engineered variant having mutations that increase its bypass activity or preference for incorporating a specific nucleotide opposite the abasic site.
[0146] In one embodiment, the engineered variant is a variant of DPO4 DNA polymerase (SEQ ID NO:1). DPO4 is a DNA polymerase (Y-family DNA polymerase) naturally expressed by the archaeon Sulfolobus solfataricus, which typically functions in the replication of damaged DNA through a process called translesion synthesis (TLS). Advantages of DPO4 include a monomeric structure, an open structure, the absence of an exonuclease domain, and the ability to bypass abasic sites. The crystal structure of DPO4 can be used to guide protein engineering, see, for example, Ling et al. (2001) "Crystal Structure of a Y-Family DNA Polymerase in Action: A Mechanism for Error-Prone and Lesion-Bypass Replication" Cell 107:91-102. As is known in the art, the inventors have engineered thousands of DPO4 variants that are optimized for, for example, the ability to utilize unconventional nucleotide analogs as substrates. A non-limiting list of DPO4 variants and screening methods that can be used in accordance with the present invention is disclosed in U.S. Patent Nos. 11,299,725, 11,530,392, and 11,708,566 issued to the applicant, the contents of which are incorporated herein by reference in their entirety.
[0147] The inventors had previously identified the region of the DPO4 polymerase corresponding to amino acids 76-86, which had already been a key target for modifying and optimizing the substrate specificity of the polymerase. Thus, in an additional wild-type background, many variants with mutations in this region were screened for abasic site bypass activity with dATP incorporation. A specific DPO4 polymerase variant was identified from the screening that exhibited robust abasic site bypass activity and is referred to herein as "C9110". Relative to the wild-type polymerase, this variant includes the following mutations: M76W_K78E_E79P_Q82W_Q83G_S86E and a deletion of amino acids 341-352 (SEQ ID NO:3).
[0148] Nucleotide Analogue
[0149] In some cases, the substrate preference of a bypass DNA polymerase can be modified or directed by utilizing alternative nucleotides (i.e., nucleotide analogs) in a second primer extension reaction. For example, when the modified nucleobase of interest is 5-mC, the primer extension reaction can include an analog of dATP, some examples of which are shown in Figure 7For example, dATP can be one or more of, for example, DAP (diaminopurine), 7-position substituents such as alkynyl C8, C10, phenyl, or analogs (a) on 7-deazadATP. Other exemplary dATP analogs include 7-deazas having an iodine group analog (B) or a bromine group analog (C) bonded to the C-7 atom, or a chlorine group analog (D) bonded to the C-2 atom. In other instances, as Figure 8 shown, dATP can be modified with 6-position substituents such as N6-methyldATP analog (A), N6-aminohexane analog (B), or 8-bromine group analog (C). In one embodiment, N6-methyldATP is used for the second primer extension reaction.
[0150] Designed Nucleotide Analogue
[0151] In a further aspect, the methods of the invention can include nucleotide analogs wherein the nucleobase is designed to introduce specific structural and / or chemical features that facilitate incorporation by bypass DNA polymerases. Exemplary nucleobase features include an overall geometry that is spatially compatible with the empty "pocket" left by nucleobase excision. For example, nucleotide analogs having the size and geometry of two bases (e.g., a base pair) may be advantageous. Other beneficial features can include an overall increase in hydrophobicity or the introduction of moieties known to enhance incorporation by bypass polymerases, such as spermine. In certain embodiments, the designed nucleotide analogs can include more than one such feature, e.g., they can include both polymerase-enhancing features and "bulky" hydrophobic features.
[0152] Certain exemplary designed nucleotide analogs include, but are not limited to Figure 9 the following depicted in: alkyl analogs, N6-ethyl-2'-dATP, analog (A), 2-methyl-2'-dATP, analog (B), 2-ethyl-2'-dATP, analog (C), and protected analogs, N6-benzoyl-2'-dATP, analog (D), and N6-phenoxyacetal-2'dATP, analog (E).
[0153] Other exemplary nucleotide analogs include, but are not limited to Figure 10 the following depicted in: 7-ethynylphenyl-7-deaza-2'-ATP, analog (A), N6-trifluoroacetamide-di-2'-dATP, analog (B), and N6-ethoxyacetyl-2'dATP, analog (C).
[0154] The design of nucleotide (e.g., dATP) analogs suitable for practicing the invention can be by Figure 11The general structure guidance shown includes the following: N6-(alkyl or acyl)-2'-dATP, compound (A), N6-(alkyl or acyl)-2-alkyl-2'-dATP, compound (B), N6,N6-(alkyl or acyl)-2-alkyl-2'-dATP, compound (C), N6,N6-(alkyl or acyl)-2-alkyl-7-deaza-2'-dATP, compound (D), N6,N6-(alkyl or acyl)-2-alkyl-7-alkynyl-7-deaza-2'-dATP, compound (E), N6,N6-(alkyl or acyl)-2-alkyl-7-alkynyl-3,7-dideaza-2'-dATP, compound (F), and γ-O-alkyl-N6,N6-(alkyl or acyl)-2-alkyl-7-alkynyl-3,7-dideaza-2'dATP, compound (G).
[0155] In another embodiment, a dGTP analog such as 7-deazadGTP is used, which is a less favorable polymerase substrate and may affect determination of which nucleotides are incorporated opposite an abasic site during an abasic site bypass primer extension reaction. In other embodiments, additional components of the primer extension reaction can be optimized to affect the substrate preference of the bypass DNA polymerase, such as buffer pH, solvent composition, relative ratios of dNTPs, etc. In some cases, the amount of polymerase protein may be limiting in the reaction, thereby minimizing synthesis of undesired primer extension by-products.
[0156] Aminooxyalkyl Nucleobase Analogue
[0157] As discussed herein and with reference to FIG. 5, certain chemical stabilizers react with abasic sites in DNA to form stable oxime adducts that prevent subsequent degradation of the phosphodiester backbone. As used herein, the term "oxime" refers to an organic compound belonging to the imine class and having the general formula RR'C═N-OH, where R is an organic side chain and R' can be hydrogen, forming an aldoxime, or another organic group, forming a ketoxime. O-substituted oximes form a closely related family of compounds. A particularly useful class of stabilizers for forming oxime adducts are those having the generalized aminooxyalkyl structure H 2 N-O-R stabilizers. Advantageously, the inventors have discovered that certain oximes have the ability to further biologically mimic the Watson-Crick base pairing activity of natural nucleobases. Thus, they not only stabilize abasic sites but also directly incorporate specific nucleotides opposite sites during daughter strand synthesis. Such aminooxyalkyl-based stabilizing reagents and their corresponding oxime adduct products may alternatively be referred to herein in certain embodiments as "nucleobase mimics", "aminooxyalkyl nucleobase mimics", or "nucleobase oxime mimics".
[0158] In one embodiment, the uracil mimetic 1-[2-(amino)ethyl]-uracil is used to stabilize abasic sites because the aminooxyalkyl moiety of the mimetic compound reacts with the abasic site to form a stable oxime adduct. Advantageously, the heterocyclic moiety of the compound is capable of forming Watson-Crick base pairs with adenine and will thus direct the incorporation of dATP during daughter strand synthesis.
[0159] Figure 12A An example of the conversion of 5-mC to a uracil oxime mimetic is shown. Here, as described previously, a DNA target molecule containing a 5-mC residue is treated with TET(I) to convert 5-mC to 5-caC and with TDG(II) to excise the 5-caC nucleobase and generate an abasic site. In this example, the DNA target is also treated with an aminooxyalkyl uracil mimetic (III), which reacts chemically with the abasic site to form a stable oxime mimetic adduct (IV). In this embodiment, the aminooxyalkyl uracil mimetic (III) is 1-[2-(aminooxy)ethyl]-uracil, available from Enamine, Ltd, Kyiv, Ukraine. Advantageously, the inventors have found that both the enzymatic conversion and excision of 5-mC with TET and TDG and the chemical conversion of the abasic nucleotide to a stable oxime adduct can be carried out in a single reaction, i.e., a "one-pot" reaction. This one-pot reaction is also referred to herein as a "chemoenzymatic nucleobase conversion reaction". Importantly, the oxime mimetic adduct (IV) is capable of base pairing with an adenine base and is thus read as uracil during daughter strand synthesis.
[0160] Figure 12B Shows how 5-mC can be chemoenzymatically converted to a uracil oxime mimetic for the detection of 5-mC in a DNA target fragment. Here, as discussed with reference to Figure 12A the parental DNA template is subjected to steps (I) through (IV) to chemoenzymatically convert 5-mC to a uracil oxime mimetic. Prior to this conversion, a first daughter strand copy (V) of the template is synthesized, as discussed with reference to Figure 1B This reaction is carried out with native nucleotides such that native G is incorporated into the daughter strand opposite the 5-mC in the parental template. After the chemoenzymatic conversion of the parental template, a second primer extension reaction generates a second daughter strand copy (VI). This reaction can also be carried out with native nucleotides such that native A is incorporated at the position opposite the uracil oxime mimetic. For sequence comparison analysis, both the first and second daughter strand copies are used as templates for amplicon sequencing A template of Scheme (VII), as further described herein. The sequencing reads of the resulting first daughter strand copy will indicate "C" at each position of 5-mC in the original parental template, while the sequencing reads of the second daughter strand copy will indicate "T" at each position of 5-mC in the parental template. Thus, the "C->T" substitution in the sequence of the Xpandomer copy of the second daughter strand reveals the positions of 5-mC in the target fragment.
[0161] Other exemplary aminooxyalkyl nucleobase analogs suitable for the methods of the present invention include 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, and are commercially available from, for example, Enamine Ltd. In other embodiments, the present invention contemplates new aminooxyalkyl nucleobase analogs in which certain chemical features are optimized for specific applications. For example, the analogs can include heterocycles other than uracil, such as thymine, cytosine, guanine, or adenine. In other embodiments, the analogs can include alternative atomic distances between the oxime and the heterocycle, such as from two carbons to three, four, or five carbons. Certain exemplary aminooxyalkyl nucleobase analogs are listed in Figure 13 and include the following: 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, Compound (A); 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, Compound (B); 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, Compound (C); 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, Compound (D); 1-[2-(aminooxy)ethyl]-thymine, Compound (E); and other predicted pseudouridine analogs, Compounds (F) and (G).
[0162] According to the present invention, in the case of incorporating a nucleotide at a position in the second complementary strand opposite an abasic site in the DNA target strand, it preferably cannot form a Watson-Crick base pair with the originally excised modified nucleobase under the primer extension conditions described herein. For example, if the modified nucleobase of interest is a derivative of cytosine, the nucleotide incorporated opposite the excised base will not be dGTP, but dATP, dCTP, or dTTP or a derivative thereof; if the modified nucleobase of interest is a derivative of guanine, the nucleotide incorporated opposite the excised base will not be dCTP, but dATP, dGTP, or dTTP or a derivative thereof; if the modified nucleobase of interest is a derivative of adenine, the nucleotide incorporated opposite the excised base will not be dTTP, but dATP, dCTP, or dGTP or a derivative thereof; if the modified nucleobase of interest is a derivative of thymine, the nucleotide incorporated opposite the excised base will not be dATP, but dCTP, dGTP, or dTTP or a derivative thereof. In a preferred embodiment, as discussed herein, natural dATP or a derivative thereof is the nucleotide incorporated opposite the abasic site generated by the excision (i.e., conversion) of a modified cytosine (e.g., 5-mC) in the original DNA target fragment.
[0163] In some cases, the yield of the desired incorporated nucleotide is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or nearly 100% of the total number of incorporation events for each second complementary copy strand produced. For example, the yield of the desired incorporated nucleotide can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or nearly 100% of the total events in each second primer extension reaction. In one instance, the yield of the desired incorporated product can be at least 80%. In one instance, the yield of the desired incorporated nucleotide can be at least 85%. In another instance, the yield of the desired incorporated nucleotide can be at least 90%. In another instance, the yield of the desired incorporated nucleotide can be at least 95%. In another instance, the yield of the desired incorporated nucleotide can be nearly 100%.
[0164] In certain cases, the second DNA polymerase can "skip" the abasic site during the second primer extension reaction and generate a deletion in the second complementary copy opposite the position of the abasic site. In still other cases, the second DNA polymerase can incorporate more than one nucleotide at a position in the target DNA polymerase opposite the abasic site, thereby generating an insertion in the second complementary copy. In either of these two cases, the sequence of the second daughter strand differs from the sequence of the first daughter strand, and these differences can provide information about the position of the modified nucleobase in the target fragment.
[0165] In some cases, once the first and second complementary copy strands of the DNA target fragment are generated as described above, they can be evaluated by many established and emerging nucleic acid sequencing techniques, including but not limited to deep sequencing, next-generation sequencing, and nanopore sequencing.
[0166] Chemo-Enzymatic Nucleobase Conversion Reaction Mixture
[0167] In certain aspects, the chemo-enzymatic nucleobase conversion reaction mixture according to the present invention can comprise at least one DNA glycosylase, a chemical stabilizer, and a suitable buffer.
[0168] Each DNA glycosylase may be specific for one or more different types of modified nucleobases or one or more types of nucleobase modifications. In some embodiments, the DNA glycosylase comprises one of the glycosylases listed in Table 1. In other embodiments, the chemo-enzymatic nucleobase conversion reaction mixture can comprise additional enzymes that chemically convert the modified nucleobases of interest without excising the nucleobases from the DNA fragment, such as TET enzymes. In some embodiments, the amount of DNA glycosylase in the nucleobase conversion reaction mixture will be an amount sufficient to completely excise most of the modified target nucleobases from the DNA target fragment. For example, the amount of DNA glycosylase can be about 0.1 μg of purified enzyme protein / pmol of DNA template, about 0.15 μg of purified enzyme protein / pmol of DNA template, about 0.2 μg of purified enzyme protein / pmol of DNA template, about 0.3 μg of purified enzyme protein / pmol of DNA template, about 0.5 μg of purified enzyme protein / pmol of DNA template, about 0.7 μg of purified enzyme protein / pmol of DNA template, about 1.0 μg of purified enzyme protein / pmol of DNA template, about 1.5 μg of purified enzyme protein / pmol of DNA template, about 2 μg of purified enzyme protein / pmol of DNA template, or more than 2 μg of purified enzyme protein / pmol of DNA template.
[0169] In some embodiments, the chemical stabilizer can be selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine. In some embodiments, the chemical stabilizer can be present in the nucleobase conversion reaction mixture at a final molar concentration of about 1 mM, about 5 mM, about 10 mM, about 15 mM, about 20 mM, about 25 mM, about 30 mM, up to 50 mM, up to 75 mM, up to 100 mM, or more than 100 mM.
[0170] In some embodiments, suitable buffers can be selected from the group consisting of MES, Tris-HCl, HEPES, etc. In further embodiments, the suitable buffer can include additional excipients such as salts (e.g., NaCl or NaOAc), DTT, MgCl 2 , DTT, PEG, etc. In other embodiments, the nucleobase conversion reaction can include cofactors suitable for a particular DNA glycosylase or other convertase, such as one or more of ammonium iron(II) sulfate, α-ketoglutaric acid, and sodium ascorbate. In some embodiments, the final pH of the nucleobase conversion reaction mixture can be about pH 4, about pH 5, about pH 6, about pH 7, or higher than pH 7. Of course, those skilled in the art will understand that the final pH will depend on the specific stabilizers, DNA glycosylases, and other enzymes present in the reaction mixture.
[0171] In certain embodiments, the chemo-enzymatic nucleobase conversion reaction mixture can be a liquid, frozen liquid, dried liquid, lyophilized liquid, or partially lyophilized liquid.
[0172] Kit
[0173] In another aspect, a kit is provided that contains reagents for performing the methods described herein. In certain embodiments, the kit can contain the chemo-enzymatic nucleobase conversion reaction mixture as described herein. The kit can contain various other enzymes. For example, the kit can contain one or more of a high-fidelity DNA polymerase, a base excision repair-deficient DNA polymerase, and a DNA polymerase with exonuclease activity. The kit can also contain a DNA ligase for library preparation, such as a ligase that ligates adapters to DNA target fragments to create a library of adapter-ligated DNA target fragments.
[0174] In some cases, the kit may contain one or more buffers and / or reaction components for performing the first primer extension reaction, base excision reaction, abasic site stabilization reaction, and second primer extension reaction steps of the method. For example, the kit may contain one or more of DNA polymerase buffer, DNA glycosylase buffer, DNA ligase buffer, or any combination thereof. The kit may also contain other reagents such as salts, cations, or detergents.
[0175] In some cases, the kit includes reagents and instructions for DNA sample fragmentation and adapter ligation. For example, the kit may contain one or more enzymes for DNA fragmentation and adapter ligation.
[0176] In some cases, the kit may further contain control DNA oligonucleotides containing one or more of the modified nucleobases of interest. The control oligonucleotides may be provided at a known concentration and have a known amount of modified nucleobases per DNA molecule or concentration. In some cases, the control DNA oligonucleotides may be within a specific size range. For example, the control DNA oligonucleotides may be in the range of 25 - 100 bp, 25 - 150 bp, 50 - 200 bp, 50 - 300 bp, 25 - 500 bp, etc. In some cases, the control DNA oligonucleotides may be within the same approximate size range as the DNA molecules to be analyzed using the kit.
[0177] In some cases, the kit may further include instructions. The instructions may specify how to perform one or more of the DNA isolation step, DNA fragmentation step, adapter ligation step, first primer extension reaction step, glycosylase treatment step, abasic site stabilization step, and second primer extension reaction step. Instructions for how to use the control DNA oligonucleotides may also be included in the kit.
[0178] By Amplification Sequencing
[0179] One nucleic acid sequencing method of the present invention is "sequencing by amplification" developed by Stratos Genomics (See, e.g., U.S. Patent No. 7,939,259 to Kokoris et al., "High-Throughput Nucleic Acid Sequencing by Amplification," which is incorporated herein by reference in its entirety). SBX is polymerization based on highly modified unnatural nucleotide analogs (referred to as "XNTPs"). Generally, SBX uses biochemical polymerization to transcribe the sequence of a DNA template (e.g., the first complementary copy and the second complementary copy of a DNA target fragment as described herein) onto a measurable polymer called an "Xpandomer." The transcribed sequence is encoded in a high signal-to-noise reporter spaced approximately 10 nm along the Xpandomer backbone, which is designed for high signal-to-noise, well-differentiated reactions. These differences provide a significant performance enhancement in terms of the sequence read efficiency and accuracy of the Xpandomer relative to native DNA.
[0180] XNTPs are deployable, 5'-triphosphate modified unnatural nucleotide analogs that are compatible with template-dependent enzymatic polymerization. XNTPs have two distinct functional regions; namely: a phosphoramidate bond that is selectively cleavable and that links the 5'-α-phosphate to the nucleobase; and a symmetrically synthesized reporter tether (SSRT) attached at certain positions within the nucleoside triaminophosphate, which allows for control of deployment by cleavage of the phosphoramidate bond. The SSRT includes linkers separated by selectively cleavable phosphoramidate bonds. Each linker is attached to one end of a reporter code. The XNTP substrate incorporated into the daughter strand product of template-dependent polymerization is in a "constrained" conformation. The constrained conformation of polymerized XNTPs is a precursor to the extended conformation, as seen in the Xpandomer product.
[0181] The transition from the constrained conformation to the extended conformation is caused by cleavage of the selectively cleavable phosphoramidate bond within the backbone of the daughter strand. In this embodiment, the SSRT contains one or more reporters or reporter codes that are specific to the nucleobase to which they are attached, thereby encoding the sequence information of the template. In this way, the SSRT provides a means to extend the length of the Xpandomer and reduce the linear density of the parental strand sequence information.
[0182] The SSRT (i.e., "tether") of XNTPs includes several different functional elements or features, such as a polymerase enhancement region, a reporter code, and a translation control element (TCE). Each of these features performs a unique function during Xpandomer translocation through a nanopore to generate a series of unique and reproducible electrical signals. The SSRT is designed to control the Xpandomer translocation rate of the TCE through a combination of steric and / or electrical repulsion, and the sizes of different reporter codes are designed to block the flow of ions through the nanopore at different measurable levels.
[0183] The specific SSRT polymerization sequences can be efficiently synthesized using phosphoramidite chemistry commonly used for oligonucleotide synthesis. The reporter codes and other features can be designed by selecting the sequences of specific phosphoramidites from commercial and / or proprietary libraries. Such libraries include, but are not limited to, polyethylene glycols with lengths of 1 to 12 or more ethylene glycol units and aliphatic polymers with lengths of 1 to 12 or more carbon units. In certain embodiments, the SSRT includes a feature called a "polymerase enhancement region" at the SSRT terminus near the nucleotide triphosphoramidite diester. The polymerase enhancement region can include a positively charged polyamine spacer (e.g., primary amine, secondary amine, tertiary amine, or quaternary amine) or a triamine spacer (three secondary amines, each separated by three carbons), which promotes incorporation through the XNTP structure by nucleic acid polymerase. In certain embodiments, the polymerase enhancement region includes two repeating units of spermine.
[0184] As used throughout this disclosure, the terms "linker A" and "linker B" refer to SSRT regions that each include a polymerase enhancement region and one or more translocation deceleration features or regions, and in certain embodiments, include a spacer containing, for example, a PEG6 polymer, which can be customized to adjust the length of the SSRT that traverses through the nanopore.
[0185] In certain embodiments, XNTP can be a compound having the following general structure:
[0186]
[0187] In one embodiment, R can be H, for example, when the compound is used for sequencing a DNA template.
[0188] In certain embodiments, the nucleobase is adenine, cytosine, guanine, thymine, uracil, or a nucleobase analog. Those skilled in the art will understand that adenine, cytosine, guanine, thymine, and uracil are naturally occurring nucleobases. As used herein, the term "nucleobase analog" refers to a non-naturally occurring nucleobase that is capable of forming Watson and Crick base pairs with a complementary nucleobase on an adjacent single-stranded nucleic acid template.
[0189] To obtain sequence information, the Xpandomer translocates from the cis reservoir to the trans reservoir through the nanopore. As the Xpandomer translocates, the reporter enters the stem until its translocation control element stops at the stem entrance. The reporter is held in the stem until the TCE can enter and pass through the stem, and then the translocation proceeds to the next reporter. After passing through the nanopore, each reporter code of the linearized Xpandomer generates a unique and reproducible electrical signal that is specific to the nucleobase to which it is attached.
[0190] In certain embodiments, nanopore-based sequencing chips can be used to analyze Xpandomers generated by SBX chemistry. The nanopore-based sequencing chips can include a large number of sensor units configured in an array. For example, the chip can include an array of one million units configured as 1000 rows by 1000 columns of units. Each unit in the array can include control circuitry integrated on a silicon substrate. Such nanopore-based sequencing chips, devices, and systems are described, for example, in the applicant's published patent application No. WO2021 / 219795, which is incorporated herein by reference in its entirety.
[0191] Proprietary in-house bioinformatics pipelines are typically used to process sequencing reads. The methods disclosed herein utilize UMIs to effect pairing of first and second complementary copy reads. The read pairs can be quality filtered and adapter and primer sequences trimmed. UMI sequences may cluster together, defining UMI families (all reads derived from a single DNA template).
[0192] Diagnostic and Prognostic Methods
[0193] In certain embodiments, the method can involve diagnosing an individual having a disorder characterized by a methylation level and / or methylation pattern at a particular locus in a test sample that is different from the methylation level and / or methylation pattern at the same locus in a sample considered normal or considered not to have the disorder. The method can also be used to predict an individual's susceptibility to a disorder characterized by a level and / or pattern of methylated loci that is different from the level and / or pattern of methylated loci exhibited in the absence of the disorder.
[0194] Particularly with regard to cancer, changes in DNA methylation have been recognized as one of the most common molecular alterations in human tumorigenesis. Hypermethylation of CpG islands located in the promoter regions of tumor suppressor genes is a well-recognized and common mechanism of gene inactivation in cancer (Esteller, Oncogene 21(35):5427-40(2002)). Conversely, global hypomethylation of genomic DNA has been observed in tumor cells; and a correlation between hypomethylation of many oncogenes and increased gene expression has been reported (Feinberg, Nature 301(5895):89-92(1983), Hanada et al. Blood 82(6):1820-8(1993)). Cancer diagnosis or prognosis assessment can be performed in the methods described herein based on the methylation status of a specific sequence region of a gene, which can include but is not limited to a coding sequence, a 5'-regulatory region, or other regulatory regions that affect transcriptional efficiency.
[0195] In a diagnostic or prognostic method, the reference genomic DNA (e.g., gDNA considered to be "normal") and the test genomic DNA to be compared can be obtained from different individuals, different tissues, and / or different cell types. In certain embodiments, the genomic DNA samples to be compared can be from the same individual but from different tissues or different cell types, or from tissues or cell types differentially affected by a disease or disorder. Similarly, the genomic DNA samples to be compared can be from the same tissue or the same cell type, where the cells or tissues are differentially affected by a disease or disorder.
[0196] Example
[0197] Example 1
[0198] One-pot chemoenzymatic conversion reaction
[0199] This example demonstrates glycosylase-mediated excision of 5-mC from a double-stranded DNA target fragment using an aminooxyalkyluracil mimic and chemical conversion of the resulting abasic site to a stable oxime adduct. Advantageously, the enzymatic and chemical conversion reactions are carried out simultaneously in a single reaction vessel (i.e., a "one-pot" reaction).
[0200] For this experiment, a single-stranded DNA target fragment (80mer) was designed to contain three spaced 5-mC residues. The 5' end of the target strand was covalently modified with biotin to facilitate physical manipulation of the strand with streptavidin-coated beads. The target strand was hybridized with a complementary oligonucleotide strand containing natural nucleotides at a molar ratio of 5:7.5 pmol to generate a double-stranded fragment. A 21mer oligonucleotide primer was designed to hybridize to the 3' end of the template.
[0201] The "one-pot" conversion reaction contained the following reagents: double-stranded DNA fragment, 3 μg of purified ngTET protein, 8 μg of purified TDG protein, 50 mM MES buffer (pH 6), 50 mM NaCl, 1 mM α-ketoglutarate (TET cofactor), 2 mM sodium ascorbate (TET cofactor), 1 mM DTT, 20% PEG, 0.1 mM ammonium iron(II) sulfate (Mohr's salt, TET cofactor), and 10 mM or 26 mM aminooxyalkyluracil mimic 1-[2-(aminooxy)ethyl]-4-hydroxy-1,2-dihydropyrimidin-2-one (C 6 H 9 N 3 O 3 )), which are commercially available from Enamine, Ltd., Kyiv, Ukraine. The final reaction mixture (50 μL) was incubated at 28 °C for 3 hours. Controls included similar one-pot reactions but without the uracil mimic and additionally without the TET and TDG enzymes.
[0202] To be able to detect the chemoenzymatic conversion of the target strand, the reaction products were subjected to mild alkaline conditions (100 mM NaOH, for 20’) to selectively cleave the target strand at the newly generated abasic sites. The reaction products were analyzed by gel electrophoresis and visualized by Sybr staining.
[0203] A representative gel is shown in Figure 14 . Lane 1 shows the products of the control reaction, lacking the TET and TDG proteins. The larger band corresponds to the longer target strand, and the smaller band corresponds to the shorter complementary strand. As expected, no degradation of the target strand was observed in the absence of DNA glycosylase. In contrast, lane 2 shows the degradation of the target strand in the presence of the TET and TDG proteins, indicating that the 5-mC residues were excised to generate labile abasic sites susceptible to base-mediated strand degradation. Notably, lanes 3 and 4 show that the inclusion of the aminooxyalkyluracil mimetic in the conversion reaction prevents target strand degradation. This observation is consistent with the mechanism of the mimetic forming stable oxime adducts at abasic sites generated by excision of nucleobases that are otherwise difficult to further degrade.
[0204] These results demonstrate the successful chemoenzymatic conversion of 5-mC residues to stable oxime adducts in DNA target fragments and provide proof-of-concept support that these discrete reactions can be carried out in a single one-pot reaction.
[0205] Example 2
[0206] DPO4 polymerase exhibits abasic bypass activity
[0207] This example demonstrates that DPO4 (a class Y DNA polymerase isolated from Sulfolobus solfataricus) is able to successfully synthesize full-length copies of a DNA template that includes several abasic challenges.
[0208] For this experiment, a single-stranded 80mer template was designed to contain three abasic (AP) sites. The 5' end of the template was covalently modified with a biotin moiety for immobilizing the template on streptavidin-coated beads. A 21mer extension oligonucleotide (EO) was designed to hybridize to the 3' end of the template. The 5’ end of the EO was covalently modified with a SIMA dye for fluorescence detection of the primer extension products.
[0209] Prior to the primer extension reaction, the template was prepared by incubating 75 pmol of the template with 100 pmol of the EO and 50 μl (10 mg / ml) of beads (Dynabeads TM MyOne TM Streptavidin C1, Thermofisher, Inc.) and incubating for 10 minutes at room temperature.
[0210] For a given primer extension reaction, 10 μl of the DNA-bead complex was used to provide the template and EO. The primer extension reaction contained the following reagents: 20 mM Tris-HCl (pH 8.8), 10 mM (NH 4 )2SO 4 , 10 mM KCl, 2 mM MgSO 4 , 0.1% Triton X-100, 200 μM dNTP, 1 mM MnCl, and 2 μg of purified DPO4 polymerase. The total reaction volume was 20 μl. The reaction was run at 37 °C for 1 h. After eluting the product from the beads with a buffer containing NaOH, the primer extension products were analyzed by gel electrophoresis.
[0211] A representative gel is shown in Figure 15 . Lane 1 shows the product of a primer extension reaction lacking DNA polymerase. As expected, no extension products were observed. Lanes 2-4 show the products of primer extension reactions that either did not contain additional additives (lane 2) or contained 50% 7-deaza-dGTP (lane 3) or 100% 7-deaza-dGTP (lane 4). As shown, DPO4 polymerase was able to efficiently synthesize full-length copies of an 80-mer template, indicating its surprising ability to bypass all three abasic sites in the DNA template.
[0212] These results indicate that DPO4 is able to synthesize a daughter strand through several abasic sites in the parental template and thus verify that the enzyme is a potentially suitable polymerase for practicing the methods disclosed herein.
[0213] Example 3
[0214] Improved bypass activity on a DNA template with stabilized abasic sites
[0215] This example demonstrates that a combination of an engineered DPO4 variant and wild-type DPO1 polymerase is able to synthesize full-length copies of a DNA template that contains three abasic sites stabilized as uracil oxime analogs. In addition, this example shows that the stabilization of the abasic sites as uracil oxime analogs directs the efficient incorporation of dATP at the relative sites in the newly synthesized daughter strand.
[0216] For this experiment, single-stranded DNA templates (80mer) were designed to contain three abasic (AP) sites that were relatively evenly spaced along the length of the template. The abasic oligonucleotides were synthesized using conventional phosphoramidite chemistry with abasic II phosphoramidite (5-O-dimethoxytrityl-1-O-tert-butyldimethylsilyl-2-deoxyribose-3-[(2-cyanoethyl)-(N,N-diisopropyl)]-phosphoramidite), according to the manufacturer's recommended protocol. The abasic II phosphoramidite is available from, for example, Glen Research, Sterling, VA. The abasic oligonucleotides were treated with 100 mM aminooxyalkyl at pH 4 - 5 to generate oxime adducts at the abasic sites and purified by gel electrophoresis. This experiment utilized the aminooxyalkyluracil mimics as described in Example 1.
[0217] The 5'-end of the template was conjugated with biotin to enable physical manipulation of the strand. A 21mer extension oligonucleotide (EO) was designed to hybridize to the 3'-end of the template. The 5'-end of the EO was covalently modified with a SIMA dye for fluorescence detection of the primer extension products.
[0218] The following primer extension reactions were performed using the abasic oligonucleotides as templates: A) extension with KAPA DNA polymerase, B) extension with wild-type DPO4 polymerase, C) extension with DPO4 polymerase variant C9110, and D) extension with a combination of DPO4 variant polymerase C9110 and DPO1 polymerase.
[0219] Primer extension reaction A contained the following reagents: 3 pmol of the abasic template, 2 pmol of the extension oligonucleotide primer, KAPA HiFi buffer, and polymerase, which are available from Roche Sequencing Solutions. The total reaction volume was 10 μl. The reaction was run at 55 °C for 30 minutes according to the manufacturer's instructions. As a control, the same primer extension reaction was performed using a native template without abasic sites. Primer extension reaction B contained the following reagents: 3 pmol of the abasic template, 2 pmol of the extension oligoprimer, 20 mM Tris-HCl (pH 8.8), 10 mM (NH 4 )2SO 4 、10 mM KCl、2 mM MgSO 4, 0.1% Triton X-100, 200 μM dNTP, and 2 μg of purified DPO4 polymerase. The total reaction volume was 10 μl. The reaction was run at 37 °C for 1 hour. Primer extension reaction C contained the following reagents: 3 pmol of abasic template, 2 pmol of extension oligonucleotide primer, 20 mM Tris-HCl (pH 8.8), 100 mM NaCl, 20 μM dNTPs / 1000 μM dATP, 1 μg of purified DPO4 polymerase variant C9110, 4 mM MgCl 2 , 10% PEG, 10% BHA NMP, 150 mM betaine, 1 mM spermine, 0.15 mM HMP, 1 mM PEM. The total reaction volume was 10 μl. The reaction was run at 55 °C for 14 hours. Primer extension reaction D contained the following reagents: 3 pmol of abasic template, 2 pmol of extension oligonucleotide primer, 20 mM Tris-HCl (pH 8.8), 100 mM NaCl, 20 μM dNTPs / 1000 μM dATP, 1 μg of purified DPO4 variant C9110, 25 nM Dpo1, 4 mM MgCl 2 , 10% PEG, 10% BHA NMP, 150 mM betaine, 1 mM spermine, 0.15 mM HMP, 1 mM PEM. The total reaction volume was 10 μl. The reaction was run at 55 °C for 14 hours. The primer extension products were analyzed by gel electrophoresis and visualized by exciting the SIMA (HEX) dye linked to the extension oligonucleotide.
[0220] Representative gels are shown in Figure 16As shown in gel (A), KAPA polymerase is able to synthesize full-length (FL) copies of the native 80mer template (lane “C”); however, this polymerase is unable to extend the extension oligonucleotide hybridized to the abasic template (lane “AP”), as the small fluorescent band observed in the gel indicates that the polymerase stalls at the first abasic site in the template. In contrast, as shown in gel (B), wild-type DPO4 polymerase is able to synthesize full-length copies of the abasic template, as evidenced by the large band corresponding to the full-length product in the gel. However, the wild-type polymerase also stalls at the abasic site in the template, as evidenced by the smear of incompletely extended products in the gel. As shown in gel (C), DPO4 variant C9110 exhibits improved extension activity relative to the wild-type polymerase, more efficiently synthesizing full-length copies of the abasic template. As shown in gel (D), the combination of DPO4 variant, C9110 and DPO1 polymerase demonstrates the most significant improvement in primer extension activity, as most of the extension products observed by the gel are full-length in size. Without being bound by theory, it is speculated that the exonuclease activity of DPO1 can act as a “proofreading factor”, for example, by reversing the misincorporation generated by DPO4 and allowing the polymerase to resume extension with higher fidelity.
[0221] These results indicate that, relative to wild-type DPO4 or DPO4 variant polymerase alone, the combination of DPO4 variant, C9110 and DPO1 polymerase is able to synthesize full-length daughter strands through several abasic adducts in the parental template with improved efficiency.
[0222] To identify the nucleotides incorporated by the DPO4 variant and DPO1 polymerase at the site opposite the abasic site in reaction D above, DNA sequence analysis was performed on the products of this primer extension reaction. The specific DNA sequencing method utilized was the nanopore-based amplification sequencing method developed by the inventors, which has been described in more detail above.
[0223] To synthesize Xpandomer copies of the primer extension product, an SBX reaction was performed, which included the following reagents: a 2:1 molar ratio of single-stranded DNA template to SBX extension oligonucleotide, 0.07 μg / μL DNA polymerase (DPO4 variant C7326, SEQ ID NO: 2), 15 mM AZ-43, 43PEM (i.e., compound 73 as disclosed in applicant's published PCT application No. WO 2019 / 135975, which is incorporated herein by reference in its entirety), 100 μM XNTPS (as disclosed in applicant's published PCT application No. WO 2020 / 236526, which is incorporated herein by reference in its entirety), 0.2 mM HMP, 0.6 mM MnCl 2, 50 mM TrisHCl, 175 mM NaCl, 200 mM imidazole, 350 mM betaine, 20% PEG, 7% NMP, 3% DMSO. The reaction was run at 37 °C for 2 hours. The resulting Xpandomer sample was treated with acid (7.5 M DCl) to cleave the phosphoramidate bond within the XNMP subunit and generate the extended form of Xpandomer. Xpandomer was sequenced using the Roche HTP high-throughput Nanopore sequencing platform as described, for example, in PCT application No. PCT / EP2019 / 084581 published by the applicant, which is incorporated herein by reference in its entirety.
[0224] For this experiment, over 10 6 individual full-length Xpandomer sequences were obtained and analyzed. The results of these analyses are presented in Figure 17A and 17B , which is a graph depicting the percentage of total sequences showing specific nucleotide incorporations at each of the three abasic sites in the parental DNA template. Notably, as Figure 17A shown (corresponding to primer extension reaction D), dATP was by far the most efficiently incorporated nucleotide opposite each abasic site in the template, with over 90% of the primer extension product sequences showing an A at each of these three positions. Additionally, incorporation of dGTP was observed to be a very rare event at any of these positions. In contrast, as Figure 17B shown (corresponding to primer extension reaction A), as expected, dGTP was by far the most efficiently incorporated nucleotide opposite each 5-mC residue in the native template.
[0225] In summary, these results demonstrate the chemoenzymatic conversion of 5-mC residues in the DNA template to uracil-mimicking adducts, and advantageously, the efficient incorporation of dATP opposite these sites by the combination of the engineered DPO4 variant and DPO1 polymerase. When comparing the sequence reads of the first and second strand copies of the untransformed and transformed DNA templates separately, this new conversion strategy enables the identification of G->A substitutions, thus providing an improved alternative method for identifying epigenetic information in DNA samples.
Claims
1. A method for identifying modified nucleobases in multiple nucleic acids, the method comprises: providing a sample comprising a plurality of DNA templates; generating a first complementary copy of the DNA templates, the generating being guided by an oligonucleotide primer in the presence of natural dNTPs using a first DNA polymerase, wherein the generating produces a complementary copy of each of the DNA templates such that each complementary copy comprises natural dNTPs, and wherein each complementary copy hybridizes to one of the DNA templates; subjecting the DNA templates and the first complementary copies to DNA glycosylase treatment, wherein the DNA glycosylase specifically excises the modified nucleobases in the DNA templates to convert the positions of the modified nucleobases into abasic sites, such that each DNA template converted by the glycosylase hybridizes to an unconverted complementary copy; generating a second complementary copy of the DNA templates converted by the glycosylase, the generating being guided by a second DNA polymerase, wherein the second DNA polymerase is capable of incorporating nucleotides opposite the abasic sites in the converted DNA templates, wherein the nucleotides do not Watson-Crick base pair with the modified nucleobases; determining the nucleotide sequences of the first complementary copies and the second complementary copies; and comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each of the DNA templates converted by the DNA glycosylase, thereby determining the positions of the modified nucleobases in the DNA templates prior to DNA glycosylase conversion.
2. The method according to claim 1, wherein the step of comparing the nucleotide sequence of the second complementary copy with the nucleotide sequence of the first complementary copy for each of the DNA templates converted by the DNA glycosylase identifies nucleotide substitutions in the sequence of the second complementary copy relative to the first complementary copy, wherein the positions of the nucleotide substitutions identify the positions of the modified bases in the DNA templates.
3. The method according to claim 1 or 2, wherein the modified nucleobases are selected from the group consisting of 5-mC, 5-hmC, 5-fC, and 5-caC.
4. The method according to any one of claims 1 to 3, wherein the DNA glycosylase is a monofunctional DNA glycosylase.
5. The method according to claim 4, wherein the monofunctional DNA glycosylase is thymine DNA glycosylase (TDG) or a variant thereof.
6. The method according to claim 5, wherein the step of subjecting the DNA templates and the first complementary copies to DNA glycosylase treatment further comprises subjecting the DNA templates and the first complementary copies to treatment with a ten-eleven translocation (TET) enzyme or a variant thereof.
7. The method according to claim 6, wherein the ten-eleven translocation (TET) enzyme or a variant thereof is ngTET.
8. The method according to any one of claims 1 to 3, wherein the DNA glycosylase is a bifunctional DNA glycosylase.
9. The method according to claim 8, wherein the bifunctional DNA glycosylase is a member of the DNA glycosylase DEMETER (DME) family or a variant thereof.
10. The method according to claim 9, wherein the member of the DNA glycosylase DEMETER (DME) family or a variant thereof is a variant engineered to inactivate the lyase activity.
11. The method according to any one of claims 1 to 10, wherein the second DNA polymerase is a base - bypass DNA polymerase.
12. The method according to claim 11, wherein the base - bypass DNA polymerase is DPO4 polymerase or a variant thereof.
13. The method according to claim 12, wherein the DPO4 polymerase or a variant thereof is a variant comprising the following mutations: M76W, K78E, E79P, Q82W, Q83G, and S86E (SEQ ID NO:3).
14. The method according to any one of claims 11 to 13, wherein the base - bypass DNA polymerase incorporates dATP in the second complementary copy at a position opposite the abasic site in the glycosylase - converted DNA template.
15. The method according to any one of claims 11 to 14, wherein the base - bypass DNA polymerase further comprises a third DNA polymerase, wherein the third DNA polymerase has exonuclease activity.
16. The method according to claim 15, wherein the third DNA polymerase is DPO1 polymerase.
17. The method according to any one of claims 1 to 16, wherein the first DNA polymerase is a high - fidelity DNA polymerase.
18. The method according to any one of claims 1 to 17, further comprising the step of treating the glycosylase - converted DNA template with a stabilizer before the step of generating the second complementary copy of the glycosylase - converted DNA template.
19. The method according to claim 18, wherein the stabilizer comprises an aldehyde - reactive compound that forms a stable adduct with the abasic site.
20. The method according to claim 19, wherein the stabilizer is selected from the group consisting of O - hydroxylamine, hydrazide, tryptamine, β - aminothiol, alkylhydrazine, hydrazino - iso - Pictet - Spengler indole, and methylaminooxy - iso - Pictet - Spengler indole.
21. The method according to claim 19, wherein the stabilizer comprises an aminooxyalkyl group capable of forming an oxime adduct with the abasic site.
22. The method according to claim 21, wherein the stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine.
23. The method according to claim 22, wherein the stabilizer is 1-[2-(amino)ethyl]-uracil.
24. The method according to claim 18, wherein the step of subjecting the DNA template and the first complementary copy to DNA glycosylase treatment and the step of treating the glycosylase-converted DNA template with a stabilizer before generating the second complementary copy occur in the same step.
25. The method according to any one of claims 1 to 24, wherein the DNA template is selected from the group consisting of genomic DNA, mitochondrial DNA, cell-free DNA, circulating tumor DNA, or a combination thereof.
26. The method according to any one of claims 1 to 25, wherein the DNA template is immobilized on a solid support.
27. The method according to any one of claims 1 to 26, wherein the first complementary copy or the second complementary copy is immobilized on a solid support.
28. The method according to any one of claims 1 to 27, wherein the step of determining the nucleotide sequences of the first complementary copy and the second complementary copy comprises synthesizing Xpandomer copies of the first complementary copy and the second complementary copy and passing the Xpandomer copies of the first complementary copy and the second complementary copy through a nanopore.
29. The method according to any one of claims 1 to 28, wherein the DNA template comprises a first adaptor ligated to the 5'-end of the DNA template and a second adaptor ligated to the 3'-end of the template.
30. The method according to claim 29, wherein the first adaptor or the second adaptor is a Y adaptor.
31. The method according to claim 29, wherein at least one of the first adaptor and the second adaptor comprises a unique molecular identifier barcode (UMI).
32. The method according to claim 31, wherein the step of comparing the sequences of the first complementary copy and the second complementary copy comprises bioinformatics pairing of sequences containing the same unique molecular identifier barcode (UMI).
33. A chemo-enzymatic nucleobase conversion reaction mixture comprising a DNA glycosylase, a chemical stabilizer, and a suitable buffer.
34. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 33, further comprising a DNA template strand hybridized to the first complementary copy strand, wherein the DNA template strand comprises a modified nucleobase and the first complementary copy strand comprises a natural nucleobase.
35. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 33 or 34, wherein the chemical stabilizer comprises an aminooxyalkyl group, and the aminooxyalkyl group is capable of reacting with an abasic nucleotide comprising an open-chain aldehyde moiety to form a stable oxime adduct.
36. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 35, wherein the chemical stabilizer is selected from the group consisting of 1-[2-(amino)ethyl]-uracil, 1-[3-(aminooxy)propyl]-uracil, 1-[4-(aminooxy)butyl]-uracil, 1-[5-(aminooxy)pentyl]-uracil, 1-[2-(aminooxy)ethyl]-2,4-diiodo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dibromo-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-dichloro-5-methylbenzene, 1-[2-(aminooxy)ethyl]-2,4-difluoro-5-methylbenzene, and 1-[2-(aminooxy)ethyl]-thymine.
37. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 35, wherein the chemical stabilizer is selected from the group consisting of O-hydroxylamine, hydrazide, tryptamine, β-aminothiol, alkyl hydrazide, hydrazino-iso-Pictet-Spengler indole, and methylaminooxy-iso-Pictet-Spengler indole.
38. The chemo-enzymatic nucleobase conversion reaction mixture according to any one of claims 33 to 37, wherein the DNA glycosylase is selected from the group consisting of N-methylpurine DNA glycosylase (MPG), MutY homolog (MUTYH), Nth-like DNA glycosylase 1 (NTHL1), Nei-like DNA glycosylase 1 (NEIL1), Nei-like DNA glycosylase 2 (NEIL2), Nei-like DNA glycosylase 3 (NEIL3), 8-oxoguanine DNA glycosylase (OGG1), uracil DNA glycosylase 1 (Ung1), uracil DNA glycosylase 2 (Ung2), single-strand selective monofunctional uracil glycosylase (SMUG1), thymine DNA glycosylase (TDG), methyl-binding domain 4 (MBD4), Fpg, Ung, Demeter (DME), and ROS1.
39. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 38, wherein the reaction mixture comprises more than one DNA glycosylase.
40. The chemo-enzymatic nucleobase conversion reaction mixture according to any one of claims 33 to 37, wherein the DNA glycosylase is TDG or a variant thereof.
41. The chemo-enzymatic nucleobase conversion reaction mixture according to claim 40, further comprising a TET enzyme.
42. A kit for detecting modified nucleobases in a DNA sample, the kit comprising: The chemo-enzymatic nucleobase conversion reaction mixture according to any one of claims 33 to 41; at least one enzyme selected from a high-fidelity DNA polymerase, a base excision bypass DNA polymerase, and a DNA polymerase having exonuclease activity; and a suitable mixture of dNTPs or analogs thereof.
43. The kit according to claim 42, further comprising one or more of buffers for the enzyme.
Citation Information
Patent Citations
Phosphoroamidate esters, and use and synthesis thereof
US10301345B2
DP04 polymerase variants
US11299725B2
DPO4 polymerase variants with improved accuracy
US11530392B2
DP04 polymerase variants
US11708566B2
High throughput nucleic acid sequencing by expansion
US7939259B2