Double-stranded DNA deaminase
Patent Information
- Application Number
- JP2024531071
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2022-11-22
- Publication Date
- 2025-12-02
AI Technical Summary
Current methods for identifying modified cytosines in DNA require a denaturation step, which is inefficient and cumbersome.
Development of double-stranded DNA deaminases that can deaminate cytosines without denaturing the DNA, allowing for direct analysis of double-stranded DNA substrates.
Enables efficient and direct identification of modified cytosines in DNA without the need for denaturation steps, enhancing the speed and efficiency of DNA analysis workflows.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] This application claims the benefit of U.S. Patent Application No. 18 / 058,115, filed November 22, 2022, which claims the benefit of U.S. Provisional Patent Application No. 63 / 264,513, filed November 24, 2021, which applications are incorporated by reference in their entireties herein.
[0002] The Sequence Listing is provided herewith as Sequence Listing XML, "NEB-451.xml," created on November 22, 2022, and having a size of 1.49 GB. The contents of the Sequence Listing XML are incorporated herein by reference in their entirety. [Background technology]
[0003] In many organisms, cytosines in the genome can be covalently modified, for example to 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC). These epigenetic changes are thought to play a role in a wide variety of phenomena, including gene expression. Global or localized changes in DNA methylation are among the early events known to occur in cancer. Identification of methylation profiles in humans is a key step for studying disease processes and is increasingly being used for diagnostic purposes.
[0004] Current methods for identifying modified cytosines include a deamination step that converts cytosines to uracils, with the modified cytosines remaining undeaminated. The uracils in these deaminated DNA molecules are copied to thymines during amplification, and after sequencing the amplification products, each modified cytosine in the starting sequence can be easily identified as a "C" in the sequenced amplification product, and each cytosine appears as a "T" in the sequenced amplification product.
[0005] DNA can be deaminated chemically (e.g., using bisulfite; see Frommer et al., PNAS 1992 89:1827-1831) or enzymatically using DNA deaminases (e.g., using APOBEC3A; see, e.g., Sun et al., Genome Res. 2021 31:291-300 and Vaisvila et al., Genome Res. 2021 31:1280-1289). However, both of these approaches require single-stranded substrates. Thus, current workflows for analyzing modified cytosines typically include a denaturation step. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Frommer et al., PNAS 1992, 89:1827-1831 [Non-Patent Document 2] Sun et al., Genome Res. 2021, 31:291-300 [Non-Patent Document 3] Vaisvila et al., Genome Res. 2021, 31:1280-1289 Summary of the Invention [Problem to be solved by the invention]
[0007] It would be desirable to eliminate the denaturation step from the current workflow. [Means for solving the problem]
[0008] Summary of the Invention The present disclosure relates in some embodiments to deaminases with one or more desirable properties, including, for example, cytosine deaminases that are active against double-stranded DNA substrates. These enzymes can deaminate cytosines in double-stranded DNA substrates (e.g., without denaturing the DNA). In addition to deaminating cytosines in double-stranded DNA, double-stranded DNA deaminases can deaminate cytosines in single-stranded DNA. Cytosines adjacent to guanines ("CG") may be deaminated as well as, less well, or better than cytosines in other sequence contexts ("CH", H=A, C, T) by the disclosed deaminases. Double-stranded DNA deaminase compositions may include a deaminase and, optionally, a buffer, one or more enzymes that alter the deamination susceptibility of one or more modified cytosines (e.g., TET methylcytosine dioxygenase and / or DNA beta-glucosyltransferase).
[0009] The present disclosure relates in some embodiments to a method for deaminating a double-stranded DNA substrate. For example, deaminating the double-stranded DNA may include, for example, contacting the double-stranded DNA substrate with a double-stranded DNA deaminase without denaturing the substrate to deaminate cytosines in the double-stranded substrate, or using any agent (e.g., gyrase or helicase) that unwinds or otherwise separates the strands of the substrate to generate a deaminated product. In some embodiments, the method may include sequencing at least one strand of the product of the deamination reaction (which is a deaminated double-stranded DNA molecule, referred to herein as a "deaminated product") to generate a sequence read. The method may include amplifying the deaminated product to generate an amplified product, and then sequencing the amplified product to generate a sequence read. The disclosed cytosine deaminases may deaminate cytosine without deaminating modified cytosines (e.g., 5mC, 5hmC, 5fC, 5caC, 5ghmC, N4mC) that are also present in the DNA substrate, or may also deaminate cytosine and one or more modified cytosines in the substrate. Thus, the location of modified cytosines (e.g., 5mC or 5hmC) in a double-stranded DNA substrate can be identified by analysis of sequence reads. Some double-stranded DNA deaminases do not deaminate N4mC but can deaminate other modified cytosines, others do not deaminate 5mC and 5hmC, some do not deaminate 5hmC but can deaminate 5mC, some do not deaminate 5ghmC but can deaminate 5mC and / or 5hmC, and others do not deaminate 5fc and 5caC but can deaminate 5mC and 5hmC. Thus, the location of one or more modified cytosines may be determined in a double-stranded substrate by contacting the substrate with a deaminase having a selected specificity and, optionally, pre-treating the substrate with one or more enzymes that alter the susceptibility of one or more modified cytosines to deamination.For example, the method may include pretreating a double-stranded DNA substrate with (a) TET methylcytosine dioxygenase and DNA beta-glucosyltransferase or (b) TET methylcytosine dioxygenase but not DNA beta-glucosyltransferase. These enzymes modify 5mC and / or 5hmC in double-stranded nucleic acids to make those residues resistant to certain double-stranded DNA deaminases. In some embodiments, the method may include contacting a double-stranded DNA deaminase with a double-stranded nucleic acid that has not been contacted (previously or simultaneously) with either TET methylcytosine dioxygenase or DNA beta-glucosyltransferase, e.g., if the double-stranded DNA deaminase does not deaminate 5mC and / or 5hmC.
[0010] In some embodiments, the double-stranded DNA substrate may include at least one N4mC or pyrrolo-dC. N4mC is found in prokaryotes and archaea. Thus, in some embodiments, the double-stranded DNA substrate may be prokaryotic or archaeal. In some embodiments, the double-stranded DNA substrate may be made by ligating a hairpin adaptor to a double-stranded fragment of DNA to create a ligation product, enzymatically creating a free 3' end at the double-stranded region of the hairpin adaptor in the ligation product, and extending the free 3' end with a dCTP-free reaction mixture that includes a strand-displacing or nick-translating polymerase, dGTP, dATP, dTTP, and modified dCTP. In this method, the modified dCTP is incorporated into the new strand to create a double-stranded nucleic acid with a modified C.
[0011] Enzymes and kits for carrying out the methods are also provided, including, for example, double-stranded DNA deaminase and reaction buffers.
[0012] The file of this patent contains at least one drawing executed in color. Copies of this patent with color drawing(s) will be provided by the Patent and Trademark Office upon request and payment of the necessary fee. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 shows the topology of a maximum likelihood phylogenetic tree of cytosine deaminases surrounded by descriptive activity data arranged in concentric circles, with each set of tree termini, enzyme names, and activity results aligned along the radial axis. The enzymatic activity results for various substrates shown in these circles were measured by in vitro screening assays using Illumina short-read sequencing-based detection (Example 3). The total area of the circle corresponds to the total activity, and the relative size of the colored regions indicates the relative activity for the indicated substrate. The innermost circle shows the relative deamination activity for unmodified cysteine in double-stranded DNA (blue region) compared to single-stranded DNA (red region). The middle circle shows activity for 5-methylated cytosine in double-stranded DNA. The outermost circle shows activity for 5-hydroxymethylated cytosine in double-stranded DNA. The enzyme names are colored according to their phylogenetic family. [Figure 2A] Figures 2A-C show enzyme activity for cytosine deaminases assayed according to the screening method of Example 3. Activity is expressed as the percentage of total cytosines in the sample that are deaminated. Figure 2A shows activity results for example deaminases on double-stranded versus single-stranded DNA. Figure 2B shows activity results for example deaminases on unmodified cytosine in a CG context versus a CH (a combination of CA, CC, and CT) context. Figure 2C shows activity results for example deaminases on cytosine versus 5-methylcytosine in all sequence contexts. [Figure 2B]Figures 2A-C show enzyme activity for cytosine deaminases assayed according to the screening method of Example 3. Activity is expressed as the percentage of total cytosines in the sample that are deaminated. Figure 2A shows activity results for example deaminases on double-stranded versus single-stranded DNA. Figure 2B shows activity results for example deaminases on unmodified cytosine in a CG context versus a CH (a combination of CA, CC, and CT) context. Figure 2C shows activity results for example deaminases on cytosine versus 5-methylcytosine in all sequence contexts. [Figure 2C] Figures 2A-C show enzyme activity for cytosine deaminases assayed according to the screening method of Example 3. Activity is expressed as the percentage of total cytosines in the sample that are deaminated. Figure 2A shows activity results for example deaminases on double-stranded versus single-stranded DNA. Figure 2B shows activity results for example deaminases on unmodified cytosine in a CG context versus a CH (a combination of CA, CC, and CT) context. Figure 2C shows activity results for example deaminases on cytosine versus 5-methylcytosine in all sequence contexts. [Figure 3A]3A-3D show example workflows for identifying the location of modified cytosines in DNA. FIG. 3A shows an example workflow for APOBEC3A deamination of ssDNA, and FIGS. 3B, 3C, and 3D show example workflows in which APOBEC3A is replaced by a cytosine deaminase that deaminates dsDNZA. FIG. 3B shows an example single-pot workflow in which the DNA denaturation step is eliminated by the use of a dsDNA deaminase that is active on ssDNA and dsDNA. As shown, the DNA deaminase can be added to the reaction mixture following reaction with TET and BGT without intermediate purification and denaturation steps, thereby enhancing detection and methylome mapping of target methylation sites on genomic DNA. FIG. 3C shows an example workflow in which a substrate is contacted with a deaminase that does not deaminate either 5fC or 5caC without requiring or including pretreatment with BGT. FIG. 3D shows an example methylome analysis workflow in which a substrate is contacted with a single enzyme, dsDNA deaminase. [Figure 3B] 3A-3D show example workflows for identifying the location of modified cytosines in DNA. FIG. 3A shows an example workflow for APOBEC3A deamination of ssDNA, and FIGS. 3B, 3C, and 3D show example workflows in which APOBEC3A is replaced by a cytosine deaminase that deaminates dsDNZA. FIG. 3B shows an example single-pot workflow in which the DNA denaturation step is eliminated by the use of a dsDNA deaminase that is active on ssDNA and dsDNA. As shown, the DNA deaminase can be added to the reaction mixture following reaction with TET and BGT without intermediate purification and denaturation steps, thereby enhancing detection and methylome mapping of target methylation sites on genomic DNA. FIG. 3C shows an example workflow in which a substrate is contacted with a deaminase that does not deaminate either 5fC or 5caC without requiring or including pretreatment with BGT. FIG. 3D shows an example methylome analysis workflow in which a substrate is contacted with a single enzyme, dsDNA deaminase. [Figure 3C]3A-3D show example workflows for identifying the location of modified cytosines in DNA. FIG. 3A shows an example workflow for APOBEC3A deamination of ssDNA, and FIGS. 3B, 3C, and 3D show example workflows in which APOBEC3A is replaced by a cytosine deaminase that deaminates dsDNZA. FIG. 3B shows an example single-pot workflow in which the DNA denaturation step is eliminated by the use of a dsDNA deaminase that is active on ssDNA and dsDNA. As shown, the DNA deaminase can be added to the reaction mixture following reaction with TET and BGT without intermediate purification and denaturation steps, thereby enhancing detection and methylome mapping of target methylation sites on genomic DNA. FIG. 3C shows an example workflow in which a substrate is contacted with a deaminase that does not deaminate either 5fC or 5caC without requiring or including pretreatment with BGT. FIG. 3D shows an example methylome analysis workflow in which a substrate is contacted with a single enzyme, dsDNA deaminase. [Figure 3D] 3A-3D show example workflows for identifying the location of modified cytosines in DNA. FIG. 3A shows an example workflow for APOBEC3A deamination of ssDNA, and FIGS. 3B, 3C, and 3D show example workflows in which APOBEC3A is replaced by a cytosine deaminase that deaminates dsDNZA. FIG. 3B shows an example single-pot workflow in which the DNA denaturation step is eliminated by the use of a dsDNA deaminase that is active on ssDNA and dsDNA. As shown, the DNA deaminase can be added to the reaction mixture following reaction with TET and BGT without intermediate purification and denaturation steps, thereby enhancing detection and methylome mapping of target methylation sites on genomic DNA. FIG. 3C shows an example workflow in which a substrate is contacted with a deaminase that does not deaminate either 5fC or 5caC without requiring or including pretreatment with BGT. FIG. 3D shows an example methylome analysis workflow in which a substrate is contacted with a single enzyme, dsDNA deaminase. [Figure 4A]Figures 4A-4C show example results of the workflow similar to Figure 3C, where the dsDNA deaminase used, CseDa01, detects 5mC and 5hmC but not 5caC and 5fC, without requiring or including BGT glycosyltransferase pretreatment. Figure 4A shows that CseDa01 DNA deaminase efficiently deaminates cytosine C, 5mC, 5hmC, and 5ghmC in both single-stranded and double-stranded substrates. Figure 4B shows that CseDa01 DNA deaminase showed no sequence bias, and deamination efficiency was greater than 95% in both ssDNA and dsDNA substrates for both CpG and CpH contexts in the E. coli genome. FIG. 4C shows that CseDa01 DNA deaminase does not deaminate 5caC and 5fC and may be useful for detecting 5mC and 5hmC without the BGT glycosylation step. [Figure 4B] Figures 4A-4C show example results of the workflow similar to Figure 3C, where the dsDNA deaminase used, CseDa01, detects 5mC and 5hmC but not 5caC and 5fC, without requiring or including BGT glycosyltransferase pretreatment. Figure 4A shows that CseDa01 DNA deaminase efficiently deaminates cytosine C, 5mC, 5hmC, and 5ghmC in both single-stranded and double-stranded substrates. Figure 4B shows that CseDa01 DNA deaminase showed no sequence bias, and deamination efficiency was greater than 95% in both ssDNA and dsDNA substrates for both CpG and CpH contexts in the E. coli genome. FIG. 4C shows that CseDa01 DNA deaminase does not deaminate 5caC and 5fC and may be useful for detecting 5mC and 5hmC without the BGT glycosylation step. [Figure 4C]Figures 4A-4C show example results of the workflow similar to Figure 3C, where the dsDNA deaminase used, CseDa01, detects 5mC and 5hmC but not 5caC and 5fC, without requiring or including BGT glycosyltransferase pretreatment. Figure 4A shows that CseDa01 DNA deaminase efficiently deaminates cytosine C, 5mC, 5hmC, and 5ghmC in both single-stranded and double-stranded substrates. Figure 4B shows that CseDa01 DNA deaminase showed no sequence bias, and deamination efficiency was greater than 95% in both ssDNA and dsDNA substrates for both CpG and CpH contexts in the E. coli genome. FIG. 4C shows that CseDa01 DNA deaminase does not deaminate 5caC and 5fC and may be useful for detecting 5mC and 5hmC without the BGT glycosylation step. [Figure 5A] Figures 5A-5B show example results of performing a single-tube oxidation of 5mC using CseDa01 and TET2. X-axis labels indicate serial dilutions of deaminase, with 1x being the most concentrated enzyme and 32x being a 32-fold dilution compared to 1x. Figure 5A shows results illustrating efficient deamination of a single-stranded substrate. Figure 5B shows results illustrating efficient deamination of a double-stranded substrate. [Figure 5B] Figures 5A-5B show example results of performing a single-tube oxidation of 5mC using CseDa01 and TET2. X-axis labels indicate serial dilutions of deaminase, with 1x being the most concentrated enzyme and 32x being a 32-fold dilution compared to 1x. Figure 5A shows results illustrating efficient deamination of a single-stranded substrate. Figure 5B shows results illustrating efficient deamination of a double-stranded substrate. [Figure 6A]Figures 6A-6B show example results using MGYPDa20, a modification-sensitive deaminase that efficiently deaminates cytosine to uracil. However, MGYPDa20 does not deaminate 5-methylcytosine and 5-hydroxymethylcytosine in dsDNA and ssDNA. This deaminase can be used to detect 5mC and 5hmC without protecting these modified bases. Figure 6A shows that MGYPDa20 DNA deaminase efficiently deaminates cytosine C, but does not deaminate 5mC, 5hmC, or 5ghmC. Figure 6B shows that MGYPDa20 DNA deaminase does not exhibit sequence bias. Sequence logos were created using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 6B] Figures 6A-6B show example results using MGYPDa20, a modification-sensitive deaminase that efficiently deaminates cytosine to uracil. However, MGYPDa20 does not deaminate 5-methylcytosine and 5-hydroxymethylcytosine in dsDNA and ssDNA. This deaminase can be used to detect 5mC and 5hmC without protecting these modified bases. Figure 6A shows that MGYPDa20 DNA deaminase efficiently deaminates cytosine C, but does not deaminate 5mC, 5hmC, or 5ghmC. Figure 6B shows that MGYPDa20 DNA deaminase does not exhibit sequence bias. Sequence logos were created using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 7A] Figures 7A-7B show example results using another modification-sensitive dsDNA deaminase, NsDa01, which can be used to detect 5mC and 5hmC without protecting the modified base. Figure 7A shows that NsDa01 DNA deaminase efficiently deaminates cytosine C but not 5mC, 5hmC, or 5ghmC. Figure 7B shows that NsDa01 DNA deaminase does not exhibit sequence bias. Sequence logos were created using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 7B] Figures 7A-7B show example results using another modification-sensitive dsDNA deaminase, NsDa01, which can be used to detect 5mC and 5hmC without protecting the modified base. Figure 7A shows that NsDa01 DNA deaminase efficiently deaminates cytosine C but not 5mC, 5hmC, or 5ghmC. Figure 7B shows that NsDa01 DNA deaminase does not exhibit sequence bias. Sequence logos were created using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 8A] Figures 8A-8B show example results using RhDa01, a CpG-specific modification-sensitive dsDNA deaminase, which can be used to detect 5mC and 5hmC in CpG contexts with or without protecting the modified base. Figure 8A shows that RhDa01 DNA deaminase efficiently deaminates cytosine C in CpG contexts, but not 5mC, 5hmC, or 5ghmC. Figure 8B shows that RhDa01 DNA deaminase displays CpG sequence specificity. Sequence logos were generated using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 8B] Figures 8A-8B show example results using RhDa01, a CpG-specific modification-sensitive dsDNA deaminase, which can be used to detect 5mC and 5hmC in CpG contexts with or without protecting the modified base. Figure 8A shows that RhDa01 DNA deaminase efficiently deaminates cytosine C in CpG contexts, but not 5mC, 5hmC, or 5ghmC. Figure 8B shows that RhDa01 DNA deaminase displays CpG sequence specificity. Sequence logos were generated using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 9A]Figures 9A-B show example results using MmgDa02, a CpG-specific modification-sensitive dsDNA deaminase, which can be used to detect 5mC and 5hmC in CpG contexts with or without protecting the modified base. Figure 9A shows that MmgDa02 DNA deaminase efficiently deaminates cytosine C in CpG contexts, but not 5mC, 5hmC, or 5ghmC. Figure 9B shows that MmgDa02 DNA deaminase displays CpG sequence specificity. Sequence logos were generated using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 9B] Figures 9A-B show example results using MmgDa02, a CpG-specific modification-sensitive dsDNA deaminase, which can be used to detect 5mC and 5hmC in CpG contexts with or without protecting the modified base. Figure 9A shows that MmgDa02 DNA deaminase efficiently deaminates cytosine C in CpG contexts, but not 5mC, 5hmC, or 5ghmC. Figure 9B shows that MmgDa02 DNA deaminase displays CpG sequence specificity. Sequence logos were generated using cytosine sites with >=90% deamination efficiency in the E. coli genome. [Figure 10] Figure 10 shows example results of using the one-tube-one-enzyme EM-seq method to locate 5mC in humans using MGYPDa20, a modification-sensitive dsDNA deaminase. This figure shows that MGYPDa20, a modification-sensitive DNA deaminase, can be used to accurately detect 5mC and 5hmC in the human GM12878 genome. In these experiments, two types of adapters were used, where all Cs were replaced by 5mC or pyrrolo-dC. In both cases, the overall methylation levels in the human GM12878 genome were accurately identified. [Figure 11A]Figure 11 shows example results using sequence logos of sites not deaminated by CseDa01 deaminase from N4mC-containing substrates of different genomes with different methyltransferase sequence specificities, i.e., Paenibacillus sp. JDR-2 (CCGG target sequence) and Salmonella enterica FDAARGOS_312 (CACCGT target sequence). The eukaryotic deaminase family of APOBEC3A deaminates N4mC, but bacterial deaminases do not, and thus the newly characterized bacterial deaminase may be used to detect N4mC modifications. Figure 11A shows that the detected N4mC motif matches the predicted CCGG methyltransferase motif in Paenibacillus sp. JDR-2. Figure 11B shows that the detected N4mC motif matches CACCGT from Salmonella enterica FDAARGOS_312. [Figure 11B] Figure 11 shows example results using sequence logos of sites not deaminated by CseDa01 deaminase from N4mC-containing substrates of different genomes with different methyltransferase sequence specificities, i.e., Paenibacillus sp. JDR-2 (CCGG target sequence) and Salmonella enterica FDAARGOS_312 (CACCGT target sequence). The eukaryotic deaminase family of APOBEC3A deaminates N4mC, but bacterial deaminases do not, and thus the newly characterized bacterial deaminase may be used to detect N4mC modifications. Figure 11A shows that the detected N4mC motif matches the predicted CCGG methyltransferase motif in Paenibacillus sp. JDR-2. Figure 11B shows that the detected N4mC motif matches CACCGT from Salmonella enterica FDAARGOS_312. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] The present disclosure provides double-stranded DNA deaminases, variants, progenitors, fusions, compositions, systems, instruments, methods, and workflows for deaminating double-stranded DNA (in duplex form, without denaturation). Applications of these deaminases include, for example, EM-seq, methyl-SNP-seq, and N4mC detection, among others.
[0015] Aspects of the disclosure can be understood in light of the descriptions, figures, sequences, embodiments, section headings, and examples provided, none of which should be construed as limiting the full scope of the disclosure in any way. Thus, the innovations presented herein should be interpreted in light of the full breadth and spirit of the disclosure.
[0016] Each of the individual embodiments described and illustrated herein has individual components and features which may be readily separated from or combined with the components and / or features of any of the other several embodiments without departing from the scope or spirit of the present teachings. Any recited method may be carried out in the order of events recited or in any other order which is logically possible. Unless otherwise expressly stated as required herein, each component, feature, and method step disclosed herein is optional, and the present disclosure contemplates embodiments in which each optional element may be explicitly excluded.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Nevertheless, certain terms are defined herein in connection with the embodiments of the present disclosure and for clarity and ease of reference.
[0018] Sources of commonly understood terms and symbols include Kornberg and Baker, DNA Replication, 2nd ed. (WH Freeman, New York, 1992); Lehninger, Biochemistry, 2nd ed. (Worth Publishers, New York, 1975); Strachan and Read, Human Molecular Genetics, 2nd ed. (Wiley-Liss, New York, 1999); Eckstein, editors, Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1991); Gait, editors, Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984); Singleton et al., Dictionary of Microbiology and Molecular biology, 2nd ed., John Wiley and Sons, New York (1994), and Hale & Markham, the Harper Collins Dictionary of Biology, Harper Perennial, NY (1991), and the like.
[0019] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a protein" refers to one or more proteins, i.e., a single protein as well as multiple proteins. Any element may be expressly excluded when exclusionary terminology such as "solely," "only," etc. is used in connection with a listing of any element or when a negative limitation is stated.
[0020] Numerical ranges are inclusive of the numbers defining the range. All numbers should be understood to include the midpoints of the integers above and below the integers, i.e., the number 2 includes 1.5 to 2.5. The number 2.5 includes 2.45 to 2.55, etc. Where sample numerical values are provided, each alone may represent an intermediate value in a range of values, and together may represent the ends of the range unless otherwise specified.
[0021] In the context of this disclosure, "buffer" and "buffering agent" refer to a chemical entity or composition that itself resists changes in pH and, when present in a solution, causes such a solution to resist changes in pH when such a solution is contacted with a chemical entity or composition having a higher or lower pH (e.g., an acid or an alkali). Examples of suitable non-naturally occurring buffers that may be used in the disclosed compositions, kits, and methods include HEPES, MES, MOPS, TAPS, tricine, and Tris. Additional examples of suitable buffers that may be used in the disclosed compositions, kits, and methods include ACES, ADA, BES, bicine, CAPS, carbonate / bicarbonate, CHES, citrate, DIPSO, EPPS, histidine, MOPSO, phosphate, PIPES, POPSO, TAPS, TAPSO, and triethanolamine.
[0022] In the context of this disclosure, a "deaminase substrate" refers to a polynucleotide (e.g., DNA) molecule that may optionally be entirely double-stranded, partially double-stranded and partially single-stranded, or entirely single-stranded. A deaminase substrate may include one or more cytosines, one or more modified cytosines, one or more adenines, one or more modified adenines, or combinations thereof. A DNA substrate may include one or more adaptors.
[0023] In the context of this disclosure, a "double-stranded DNA deaminase" is a hydro-lyase that deaminates cytosine in double-stranded DNA to uracil and / or adenine in double-stranded DNA to hypoxanthine. A double-stranded DNA deaminase may deaminate cytosine and / or adenine in double-stranded DNA as well as or better than it deaminates cytosine and / or adenine in single-stranded DNA, respectively. For example, a double-stranded DNA deaminase may deaminate cytosine double-stranded DNA but not cytosine in single-stranded DNA. The double-stranded DNA may be modification-sensitive. For example, a double-stranded DNA deaminase may deaminate unmodified cytosine or adenine in double-stranded DNA but not deaminate one or more corresponding modified cytosines or adenines.
[0024] In the context of this disclosure, "duplex" and "double stranded" refer to any three-dimensional conformation of a polynucleotide in which two polynucleotide strands (e.g., separate molecules or spatially separated portions of a single molecule) are arranged in an antiparallel helical fashion with complementary bases on each strand pairing with each other (e.g., Watson-Crick base pairing). The paired bases are stacked against each other, sharing the pi-electrons of the bases.
[0025] Duplex stability may be related, in part, to the ratio of complementary base pair mismatches (if any) in the two strands, the ratio of pairs with three hydrogen bonds (e.g., G:C) to pairs with two hydrogen bonds (e.g., A:T, A:U) in the duplex, and the length of the strands with higher ratios and longer strands generally associated with higher stability. Duplex stability may be related, in part, to the surrounding conditions, including, for example, temperature, pH, salinity, and / or the presence, concentration, and nature of any buffer(s), denaturant(s) (e.g., formamide), crowding agent(s) (e.g., PEG), detergent(s) (e.g., SDS), surfactant(s), polysaccharide(s) (e.g., dextran sulfate), chelating agent(s) (e.g., EDTA), and nucleic acid(s) (e.g., salmon sperm DNA). A double-stranded polynucleotide can contain one or more unpaired bases, including, for example, mismatched bases, hairpin loops, and single-stranded (5' and / or 3') ends.
[0026] A double-stranded polynucleotide (e.g., a double-stranded DNA deaminase substrate) can have any desired length. For example, a double-stranded polynucleotide can have a length of ≦50 nucleotides, 10-200 nucleotides, 80-400 nucleotides, 50-500 nucleotides, ≦500 nucleotides, ≦1 kb, ≦2 kb, ≦5 kb, or ≦10 kb. A double-stranded polynucleotide can have any desired number of mismatched or unpaired nucleotides, for example, ≦1 per 100 nucleotides, ≦2 per 100 nucleotides, ≦3 per 100 nucleotides, ≦5 per 100 nucleotides, or ≦10 per 100 nucleotides.
[0027] In the context of this disclosure, a "fusion protein" is a protein that is composed of two or more polypeptide components that are not associated in their natural state. A fusion protein may be a combination of two, three, or four or more different proteins. For example, a fusion protein may contain two naturally occurring polypeptides that are not associated in their respective natural states. A fusion protein may contain two polypeptides, one of which is naturally occurring and the other one of which is not naturally occurring. The term polypeptide is not intended to be limited to a fusion of two heterologous amino acid sequences. A fusion protein may have one or more heterologous domains added to the N-terminus, C-terminus, and / or middle portion of a protein. If the two portions of a fusion protein are "heterologous," the two portions are not part of the same protein in nature. Examples of fusion proteins include proteins comprising a double-stranded DNA deaminase fused to another enzyme (e.g., an endonuclease), an antibody, a binding domain suitable for immobilization such as maltose binding domain (MBP), a histidine tag ("His-tag"), a chitin binding domain, alpha mating factor or SNAP-tag® (New England Biolabs, Ipswich, MA (see, e.g., U.S. Pat. Nos. 7,939,284 and 7,888,090)), a DNA binding domain, and / or albumin, where the deaminase is optionally placed closer to the N-terminus or closer to the C-terminus than the other component(s). Binding peptides may be used to improve the solubility or yield of the deaminase during manufacture of a protein reagent. Other examples of fusion proteins include fusions of the deaminase with a heterologous targeting sequence, a linker, an epitope tag, a fluorescent protein, β-galactosidase, luciferase, and / or a detectable fusion partner such as a functionally similar peptide. The components of the fusion protein may be joined by one or more peptide bonds, disulfide bonds, and / or other covalent bonds.
[0028] In the context of this disclosure, a "modified cytosine" refers to any covalent modification of cytosine, including naturally occurring and non-naturally occurring modifications. Modified cytosines include, for example, 1-methylcytosine (1mC), 2-O-methylcytosine (m2C), 3-ethylcytosine (e3C), 3,N-methylcytosine (NMC), 2-O-methylcytosine (mC), 3-ethylcytosine (e3C), 2-O-methylcytosine (m ... 4 -ethylenocytosine (εC), 3-methylcytosine (3mC), 4-methylcytosine (4mC), 5-carboxylcytosine (5CaC), 5-formylcytosine (5fC), 5-hydroxymethylcytosine (5hmC), 5-methylcytosine (5mC), N 4 -methylcytosine (N4mC), and pyrrolo-cytosine (pyrrolo-C). Additional examples of modified nucleotides can be found at https: / / dnamod.hoffmanlab.org.
[0029] In the context of this disclosure, "non-naturally occurring" refers to a polynucleotide, polypeptide, carbohydrate, lipid, or composition that does not exist in nature. Such a polynucleotide, polypeptide, carbohydrate, lipid, or composition may differ from a naturally occurring polynucleotide, polypeptide, carbohydrate, lipid, or composition in one or more respects. For example, a polymer (e.g., a polynucleotide, polypeptide, or carbohydrate) may differ in the type and arrangement of its component components (e.g., nucleotide sequence, amino acid sequence, or sugar molecule). A polymer may differ from a naturally occurring polymer with respect to the molecule(s) to which it is linked. For example, a "non-naturally occurring" protein may differ in its secondary, tertiary, or quaternary structure from a naturally occurring protein by having chemical bonds (e.g., covalent bonds including peptide bonds, phosphate bonds, disulfide bonds, ester bonds, and ether bonds, etc.) to a polypeptide (e.g., a fusion protein), lipid, carbohydrate, or any other molecule. Similarly, a "non-naturally occurring" polynucleotide or nucleic acid may contain one or more other modifications (e.g., added labels or other moieties) at the 5' end, the 3' end, and / or between the 5' and 3' ends of the nucleic acid (e.g., methylation) or at least one other modification (e.g., added labels or other moieties). A "non-naturally occurring" composition may differ from a naturally occurring composition in one or more of the following ways: (a) having components that are not associated with nature; (b) having components at concentrations not found in nature; (c) excluding one or more components that would otherwise be found in the naturally occurring composition; (d) having a form not found in nature, e.g., dried, lyophilized, crystalline, aqueous; and (e) having one or more additional components beyond those found in nature (e.g., buffers, detergents, dyes, solvents, or preservatives).
[0030] With respect to amino acids, a "position" refers to the location that such amino acid occupies in a primary sequence of a peptide or polypeptide numbered from the amino terminus to the carboxy terminus. For example, a position in one primary sequence can match a position in a second primary sequence where the two positions face each other when the two primary sequences are aligned using an alignment algorithm (e.g., BLAST (Journal of Molecular Biology. 215(3):403-410) using default parameters (e.g., expectation threshold 0.05, word size 3, maximum matches in query range 0, matrix BLOSUM62, gap extent 11 extension 1, and conditional compositional score matrix adjustment) or custom parameters). An amino acid position in one sequence can match a position in a functionally equivalent motif or structural motif that can be identified in one or more other sequence(s) in a database by aligning the motifs. Analogously, with respect to nucleotides, a "position" refers to the location that such nucleotide occupies in a nucleotide sequence of an oligonucleotide or polynucleotide numbered from the 5' end to the 3' end.
[0031] All publications, patents, and patent applications mentioned herein are incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Reagents referenced in this disclosure may be made using available materials and techniques, obtained from the indicated sources, and / or obtained from New England Biolabs, Inc. (Ipswich, Mass.).
[0032] Double-stranded DNA deaminase The present disclosure relates to naturally occurring and non-naturally occurring double-stranded DNA deaminases. Non-naturally occurring double-stranded DNA deaminases are related to, but may differ from, naturally occurring proteins. Naturally occurring proteins often contain the deaminase as a single domain of a larger multi-domain structure in which the deaminase domain is located at the most C-terminal end. Non-naturally occurring double-stranded DNA deaminases may constitute truncated versions of naturally occurring proteins, in which case the non-naturally occurring double-stranded DNA deaminases have a high degree of identity to a portion of the naturally occurring sequence, but may, for example, lack structural and / or functional domains or subunits of the corresponding naturally occurring protein. Non-naturally occurring double-stranded DNA deaminases may have any number of insertions, deletions, or substitutions compared to naturally occurring enzymes. For example, a non-naturally occurring double-stranded DNA deaminase may have less than 100% identity, less than 99% identity, less than 98% identity, less than 90% identity, less than 85% identity, less than 80% identity, less than 70% identity, less than 60% identity, less than 50% identity, less than 40% identity, less than 30% identity, or less than 20% identity to a naturally occurring enzyme. A non-naturally occurring double-stranded DNA deaminase may include an expression and / or purification tag. A non-naturally occurring double-stranded DNA deaminase disclosed herein may have an amino acid sequence that is at least 80% identical (e.g., at least 90% identical, at least 95% identical, or at least 98% identical, or at least 99% identical) to a C-terminal deaminase domain of a naturally occurring protein, where the double-stranded DNA deaminase has double-stranded DNA deaminase activity but does not include the N-terminus (if any) of the corresponding naturally occurring protein. In some embodiments, the non-naturally occurring double-stranded DNA deaminase lacks at least 10, at least 20, at least 50, or at least 100 of the N-terminal amino acids of the corresponding naturally occurring protein. In some embodiments, the double-stranded DNA deaminase is 300 or less amino acids in length, e.g., 200 or less amino acids in length, or 150 or less amino acids in length.
[0033] According to some embodiments, the double-stranded DNA deaminase may comprise an amino acid sequence having at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 93%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of SEQ ID NOs: 1-152. In some embodiments, the double-stranded DNA deaminase may be encoded by a nucleic acid sequence that, when transcribed, translated, and / or processed, results in an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 93%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of SEQ ID NOs: 1-152. The double-stranded DNA deaminase may have an amino acid sequence that is at least 90% (e.g., at least 95%, at least 98%, at least 99%) identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, 99. In some embodiments, the non-naturally occurring double-stranded DNA deaminase lacks the N-terminus of the corresponding naturally occurring protein, e.g., at least 10, at least 20, at least 50, or at least 100 of the N-terminal amino acids. Sequence alignment and structural information can be used to design mutants. In some embodiments, the double-stranded DNA deaminase may contain a fragment of a wild-type protein, which contains the deaminase domain but lacks other domains of the wild-type protein that may be C-terminal and / or N-terminal to the deaminase domain. Examples of non-naturally occurring double-stranded DNA deaminases include SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, 99.
[0034] In some embodiments, the double-stranded DNA deaminase may be a fusion protein. For example, the double-stranded DNA deaminase may have a purification tag (e.g., His tag or the like) at either end. In some embodiments, the double-stranded DNA deaminase may be fused to a DNA-binding protein (e.g., the DNA-binding domain of a transcription factor) or a protein component of a nucleic acid-guided endonuclease (e.g., catalytically inactive Cas9 (dCas9) or Cas9 nickase (nCas9) or TALEN (transcription activator-like effector nuclease)) so that the fusion protein can affect site-specific C to T substitution in the genome. Exemplary methods of "base editing" are described in publications, for example, in Komor et al., among others (Nature 533:420-424).
[0035] The double-stranded DNA deaminase may optionally deaminate cytosine but not adenine ("dsDNA cytosine deaminase"), adenine but not cytosine ("dsDNA adenine deaminase"), or both adenine and cytosine (recognizing that one may be a better substrate than the other under otherwise equal conditions). The double-stranded DNA deaminase may be modification-sensitive. For example, the double-stranded DNA deaminase may deaminate cytosine but not one or more modified cytosines in the double-stranded DNA. For example, a double-stranded DNA deaminase may deaminate cytosine but not 5mC or N4mC, or a double-stranded DNA deaminase may deaminate C and 5mC but not 5hmC, 5ghmC, or N4mC.
[0036] Double-stranded DNA deaminase composition The disclosure provides, for example, a double-stranded DNA deaminase composition comprising a reaction mixture. According to some embodiments, the deaminase composition may comprise (a) a double-stranded DNA deaminase and (b) double-stranded DNA. The deaminase composition may comprise, for example, a deaminase variant (e.g., having an amino acid sequence at least 80% identical to one or more of SEQ ID NOs: 1-152). The double-stranded DNA deaminase composition may be free of one or more other catalytic activities. For example, the double-stranded DNA deaminase composition may be free of nucleases that cleave dsDNA, free of nucleases that cleave ssDNA, free of polymerase activity, free of DNA modifying activity, and / or free of protease activity, in each case, under desired test conditions (e.g., time, temperature, pH, salinity, model substrates, and / or other conditions), e.g., conditions intended to reproduce specific use conditions or to represent conditions for a range of use of the double-stranded DNA deaminase composition.
[0037] In some embodiments, the double-stranded DNA deaminase and compositions comprising one or more double-stranded DNA deaminases can have any desired form, including, for example, a liquid, gel, film, powder, cake, and / or any dry or lyophilized form. The double-stranded DNA deaminase composition can include, for example, a film, gel, fabric, or bead, comprising a double-stranded DNA deaminase and a support or matrix, for example, a magnetic material, agarose, polystyrene, polyacrylamide, and / or chitin.
[0038] In some embodiments, the reaction mixture may include a double-stranded DNA substrate containing cytosine and a double-stranded DNA deaminase. The double-stranded DNA substrate may include cytosine and at least one modified cytosine, such as 5fC, 5CaC, 5mC, 5hmC, N4mC, or pyrrolo-C. The double-stranded DNA substrate may be eukaryotic DNA (e.g., plant or animal) or bacterial. In some embodiments, the double-stranded DNA substrate may be from a mammal, such as a human. In some embodiments, the double-stranded DNA substrate may be human cfDNA. The reaction mixture may further include a TET methylcytosine dioxygenase (e.g., TET2) and a DNA beta-glucosyltransferase, and / or one or more of a ligase, a polymerase, a proteinase K, and / or a thermolabile proteinase K, as described herein. The reaction mixture may be free of an unwinding agent (eg, a gyrase, a topoisomerase, a single-stranded DNA binding protein, or a helicase) and / or free of a denaturing agent.
[0039] Double-stranded DNA deaminase method The present disclosure provides methods for using deaminases to, for example, identify the type and / or location of modified nucleotides in DNA. In some embodiments, the methods may include providing a double-stranded DNA substrate of any desired length. For example, the double-stranded DNA substrate may have a length of ≦50 nucleotides, 10-200 nucleotides, 80-400 nucleotides, 50-500 nucleotides, ≦500 nucleotides, ≦1 kb, ≦2 kb, ≦5 kb, or ≦10 kb. The double-stranded DNA substrate, in some embodiments, may be a fragment of genomic DNA, organelle DNA, cDNA, or other DNA of interest, and may be or originate from any desired source (e.g., human, non-human mammalian, plant, insect, microbial, viral, or synthetic DNA). The DNA substrate, in some embodiments, may be prepared by extracting (e.g., genomic DNA) from a biological sample and, optionally, fragmenting the genomic DNA. In some embodiments, fragmenting the DNA may include mechanically fragmenting the DNA (e.g., by sonication, nebulization, or shearing) or enzymatically fragmenting the DNA (e.g., using a double-stranded DNA "dsDNA" fragmentation mixture). Examples of enzymes for fragmentation include NEBNext® Fragmentase®, Ultrashear, and FS systems (New England Biolabs, Ipswich MA), among others. In some embodiments, DNA for deamination may already be fragmented (e.g., as in the case of FFPE samples and circulating cell-free DNA (cfDNA)).
[0040] According to some embodiments, the method may include polishing the DNA ends (e.g., ends of fragmented DNA). For example, the DNA ends may be contacted with (a) a proofreading polymerase that trims off 3' overhanging nucleotides, if any, (b) a proofreading and / or non-proofreading polymerase that fills in 5' overhangs, if any, and / or (c) a polynucleotide kinase (PNK) that phosphorylates unphosphorylated 5' ends, if any. In some embodiments, the method may include contacting the DNA ends (e.g., blunt ends) with a non-proofreading polymerase to add non-template A-tails (e.g., single-base overhangs that include adenine) to the 3' ends. According to some embodiments, the method may include ligating one or more adaptors to the DNA ends. The adaptors may include one or more sample tags, unique molecular identifiers (UMIs), modified nucleotides, primer sequences (e.g., for sequencing). In some embodiments, the adaptors may contain cytosines (or adenines) that are not substrates for the deaminase used. If necessary, the polishing and / or ligation products may be purified, e.g., to separate the polishing or ligation products from the enzyme, unreacted nucleotides and / or adaptors, if applicable.
[0041] In some embodiments, the method may include contacting (a) a deaminase substrate with (b) a glucosyltransferase (e.g., T4-BGT) and / or a 10-11 translocation (TET) dioxygenase to generate a modified deaminase substrate. BGT may glucosylate 5hmC to form 5ghmC. TET may oxidize 5mC to 5caC. Subsequent treatment with sodium bisulfite or apolipoprotein B mRNA editing enzyme subunit 3A (APOBEC3A) would result in deamination of all Cs in the modified deaminase substrate except for 5ghmC. The deaminases disclosed herein may eliminate the need to denature DNA prior to deamination (e.g., with APOBEC3A) and may provide methylation sensitivity.
[0042] The method may include contacting a double-stranded DNA substrate containing cytosine with a double-stranded DNA deaminase to generate a deamination product containing deaminated cytosine. The double-stranded DNA substrate may further include one or more modified cytosines, such as one or more modified cytosines selected from 5fC, 5CaC, 5mC, 5hmC, N4mC and pyrrolo-C, 4mC, εC, 3mC, e3C, m2C, and 1mC. The double-stranded DNA deaminase substrate does not need to be denatured before or during deamination. Thus, the method can be carried out without a denaturation step. In some embodiments, the deamination method may include contacting a double-stranded DNA substrate containing cytosine with a double-stranded DNA deaminase to generate a reaction mixture for generating a deamination product containing deaminated cytosine.
[0043] The deamination method may further include amplifying the deaminated product to generate an amplified product, thereby copying any deaminated C in the first strand to a T in the amplified product. The deamination method may further include ligating an asymmetric (or "Y") adapter, e.g., an Illumina P5 / P7 adapter, onto the deaminated product and amplifying the deaminated product using a primer complementary to a sequence in the adapter. In some embodiments, the method may include sequencing the deaminated product or amplifying the deaminated product to generate an amplified product and sequencing the amplified product, in each case to generate sequence reads. The deaminated product and / or the amplified product may be sequenced using any suitable system, including Illumina's reversible terminator method (see, e.g., Shendure et al., Science 2005 309:1728). In some embodiments, the deaminated product may be sequenced directly without amplification, e.g., by nanopore or PacBio sequencing. The sequencing step may generate at least 10,000, at least 100,000, at least 500,000, at least 1M, at least 10M, at least 100M, at least 1B, or at least 10B sequence reads per reaction. In some cases, the reads may be paired-end reads. The method may include analyzing the sequence reads to identify modified cytosines in the double-stranded DNA substrate, where the modified cytosines may be identified as "C" because they are deaminase resistant.
[0044] Double-stranded DNA deaminases that are "blocked" by or do not deaminate modified cytosines (e.g., 5mC, 5hmC, 5ghmC, N4mC) may be used in various "EM-seq"-like workflows for the analysis of modified cytosines. Current implementations of EM-seq use deaminases that prefer single-stranded substrates. Thus, current EM-seq workflows have a denaturation step (see FIG. 3A, Sun et al., Genome Res. 2021 31:291-300 and Vaisvila et al., Genome Res. 2021 31:1280-1289). In the present workflow, the denaturation step can be removed, thereby making the EM-seq workflow faster and more efficient.
[0045] The workflow of an example deamination method is shown in Figures 3B-D. As illustrated in Figure 3B, double-stranded DNA substrates may be prepared by pre-treating double-stranded DNA with a TET methylcytosine dioxygenase (e.g., TET2) and a DNA beta-glucosyltransferase to convert 5mC and 5hmC in the starting DNA to forms resistant to double-stranded DNA deaminases, such as MGYPDa829, MGYPDa06, CrDa01, AvDa02, CsDa01, LbsDa01, FlDa01, MGYPDa26, MGYPDa23, Chimera_10, and AncDa04. Double-stranded DNA deaminases useful in the illustrated workflow can have an amino acid sequence that is at least 90% identical to the amino acid sequence of any of MGYPDa829 (SEQ ID NO:96), MGYPDa06 (SEQ ID NO:4), CrDa01 (SEQ ID NO:12), AvDa02 (SEQ ID NO:21), CsDa01 (SEQ ID NO:9), LbsDa01 (SEQ ID NO:10), FlDa01 (SEQ ID NO:8), MGYPDa26 (SEQ ID NO:7), MGYPDa23 (SEQ ID NO:6), Chimera_10 (SEQ ID NO:97), and AncDa04 (SEQ ID NO:95) double-stranded DNA deaminases. As illustrated, the double-stranded DNA deaminase can be added to the reaction without any purification, denaturation, or addition of unwinding agents.
[0046] As shown in Figure 3C, double-stranded DNA substrates may be prepared by pre-treating double-stranded DNA with TET methylcytosine dioxygenase (e.g., TET2) without the use of DNA beta-glucosyltransferase to convert 5mC in the starting DNA to a form resistant to double-stranded DNA deaminases, e.g., CseDa01 and LbDa02. Double-stranded DNA deaminases useful in the illustrated workflow may have an amino acid sequence that is at least 90% identical to the amino acid sequence of any of the CseDa01 (SEQ ID NO:3) and LbDa02 (SEQ ID NO:1) double-stranded DNA deaminases. In this embodiment, the double-stranded DNA deaminase may be added to the reaction without the addition of any purification, denaturation, or unwinding agents.
[0047] As shown in Figure 3D, the double-stranded nucleic acid may not be contacted at any point in the workflow with TET methylcytosine dioxygenase or DNA beta-glucosyltransferase (or any other enzyme that converts modified cytosine to a form resistant to the selected double-stranded DNA deaminase). For example, the selected double-stranded DNA deaminase may be blocked by 5-hydroxymethylcytosine and 5-methylcytosine. A double-stranded DNA deaminase useful in the illustrated workflow may have an amino acid sequence that is at least 90% identical to the amino acid sequence of any of MGYPDa20 (SEQ ID NO: 11), NsDa01 (SEQ ID NO: 27), and AshDa01 (SEQ ID NO: 40) double-stranded DNA deaminases.
[0048] In some embodiments, the double-stranded DNA substrate may include at least one N4mC (N4-methyl-cytosine), which is a cytosine modification that is resistant to some double-stranded DNA deaminases. A double-stranded DNA deaminase that is useful for detecting N4mC may have an amino acid sequence that is at least 90% identical to any of the amino acid sequences of SEQ ID NOs: 1-28. For example, a double-stranded DNA deaminase that is useful for detecting N4mC may have an amino acid sequence that is at least 90% identical to any of the amino acid sequences of CseDa01 (SEQ ID NO: 3) and LbDa01 (SEQ ID NO: 19) double-stranded DNA deaminases. In these embodiments, the double-stranded DNA substrate may be or include prokaryotic or archaeal DNA.
[0049] In some embodiments, double-stranded DNA deaminases may be used in a "methyl-SNP-seq" workflow (see, e.g., Yan et al., Genome Res. 2022; gr. 277080.122). For example, a method may include (a) ligating a hairpin adapter to a double-stranded fragment of DNA to generate a ligation product, (b) enzymatically generating a free 3' end at the double-stranded region of the hairpin adapter in the ligation product, and (c) extending the free 3' end in a dCTP-free reaction mixture containing a strand-displacing or nick-translation polymerase, dGTP, dATP, dTTP, and modified dCTP, to generate a double-stranded DNA substrate, as described in U.S. Provisional Patent Application No. 63 / 399,970, filed August 22, 2022, which is incorporated herein by reference. Examples of modified dCTPs include 5mdCTP, pyrrolo-dCTP, and N4mdCTP, among others, modified dCTPs that can be incorporated by a polymerase. The deaminase can have an amino acid sequence that is at least 90% identical to the amino acid sequence of any of MGYPDa20 (SEQ ID NO: 11), NsDa01 (SEQ ID NO: 27), AshDa01 (SEQ ID NO: 40).
[0050] According to some embodiments, a double-stranded DNA deaminase composition may include a double-stranded DNA deaminase and, optionally, any of (including one or more of) a buffer (e.g., storage buffer, reaction buffer), an excipient, a salt (e.g., NaCl, MgCl2, CaCl2), a protein (e.g., albumin, enzymes), a stabilizer, a detergent (e.g., ionic, non-ionic, and / or zwitterionic detergent (e.g., octoxynol, polysorbate 20)), a polynucleotide, a cell (e.g., intact, digested, or any cell-free extract), a biological fluid or secretion (e.g., mucus, pus), an aptamer, a crowding agent, a sugar (e.g., monosaccharide, disaccharide, trisaccharide, tetrasaccharide, or further polysaccharide), a starch, a cellulose, a glass forming agent (e.g., for lyophilization), a lipid, an oil, an aqueous medium, a support (e.g., beads), and / or a (non-naturally occurring) combination thereof. A combination may include, for example, two or more of the listed components (e.g., a salt and a buffer) or multiple of a single listed component (e.g., two different salts or two different sugars). Examples of proteins that can be included in a double-stranded DNA deaminase composition include one or more enzymes that alter the deamination susceptibility of one or more modified cytosines (e.g., TET methylcytosine dioxygenase and / or DNA beta-glucosyltransferase).
[0051] Double-stranded DNA deaminase kit The present disclosure relates in some embodiments to a deaminase kit comprising a double-stranded DNA deaminase. The kit may comprise any of the components described herein. The double-stranded DNA deaminase composition or kit may comprise, for example, a double-stranded DNA deaminase and, optionally, a storage buffer (e.g., with a buffering agent and with or without glycerol) and / or a reaction buffer. The reaction buffer for the deaminase composition or deaminase kit may be in a concentrated form, and the buffer may comprise one or more additives (e.g., glycerol), one or more salts (e.g., KCl), one or more reducing agents, EDTA, one or more detergents, one or more non-ionic surfactants, one or more ionic (e.g., anionic or zwitterionic) surfactants, and / or crowding agents. The kit comprising dNTPs may comprise one, two, three of the total four dATP, dTTP, dGTP, and dCTP. The kit may further comprise one or more modified nucleotides.
[0052] One or more components of the kit may be included in one container for a single step reaction, or one or more components may be included in one container but separated from other components for sequential or parallel use. For example, the kit may include two components (e.g., deaminase and storage buffer) in a single tube and all other components in separate individual tubes, in each case the contents are provided in any desired form (e.g., liquid, dry, lyophilized). One tube in the kit may contain, for example, a master mix for receiving and amplifying DNA (e.g., deaminated DNA). For example, double-stranded DNA deaminase may be placed in the cap of the tube, and components for transcribing the template nucleic acid are placed in the body of the tube. As desired, for example, once the deamination reaction is completed, the tube may be tapped, rocked, inverted, rotated, or otherwise moved to contact the placed double-stranded DNA deaminase with the deamination reaction mixture. The kit may contain the double-stranded DNA deaminase and the reaction buffer in a single tube or in different tubes, and when contained in a single tube, the double-stranded DNA deaminase and the buffer may be in the same or separate locations in the tube. For example, the kit may contain the double-stranded DNA deaminase described above, and a reaction buffer (e.g., 5x or 10x buffer). The contents of the kit may be formulated for use in a desired method or process. In some embodiments, the kit may further contain (a) a TET methylcytosine dioxygenase (e.g., TET2) and a DNA beta glucosyltransferase, or (b) a TET methylcytosine dioxygenase and no DNA beta glucosyltransferase. In some embodiments, the kit does not contain either a TET methylcytosine dioxygenase or a DNA beta glucosyltransferase. In some embodiments, the kit further comprises a modified dCTP selected from 5hmdCTP, 5fdCTP, 5cadCTP, 5mdCTP, pyrrolo-dCTP, and N4mdCTP, and / or a strand displacement or nick translation polymerase. In some embodiments, the kit may further comprise a ligase, a polymerase, proteinase K, and / or a thermolabile proteinase K.The double-stranded DNA deaminase may be lyophilized or in a buffered storage solution containing glycerol.
[0053] As will be apparent to one having the benefit of this disclosure, double-stranded DNA deaminases may be used in a variety of genomic analysis methods, particularly those that aim to identify the location and / or type of one or more modified cytosines and / or determine the methylation status of cytosines. In other embodiments, double-stranded DNA deaminases can be components of fusion proteins for base editing, i.e., generating site-specific C to T substitutions in a genome.
[0054] Embodiment The present disclosure further relates to embodiments disclosed in U.S. Provisional Patent Application No. 63 / 264,513, including all of the following:
[0055] Embodiment 1. A polypeptide comprising at least 90% sequence identity to any of SEQ ID NOs:1-8, but not 100% identity to SEQ ID NO:3.
[0056] Embodiment 2. The polypeptide of embodiment 1, comprising at least 90% sequence identity to any of SEQ ID NOs: 1-3, but not 100% identity to SEQ ID NO: 3.
[0057] Embodiment 3. A polypeptide as described in embodiment 1, comprising at least 90% sequence identity with either SEQ ID NO: 1 or 2.
[0058] Embodiment 4. A polypeptide according to any one of embodiments 1 to 3, capable of deaminating cytosines in double-stranded DNA (dsDNA) without sequence bias.
[0059] Embodiment 5. A polypeptide according to any one of embodiments 1 to 3, capable of deaminating cytosines in single-stranded DNA (ssDNA) without sequence bias.
[0060] Embodiment 6. The polypeptide of any one of embodiments 1 to 5, comprising a fusion protein.
[0061] Embodiment 7. The polypeptide of any one of embodiments 1 to 6, wherein the polypeptide is lyophilized.
[0062] Embodiment 8 The polypeptide of any one of embodiments 1 to 7, wherein the polypeptide is immobilized on a substrate.
[0063] Embodiment 9. The polypeptide of any of embodiments 1-8, wherein the polypeptide is combined with one or more reagents in a mixture, the one or more reagents in the mixture comprising a second polypeptide.
[0064] Embodiment 10. The polypeptide of embodiment 9, wherein the second polypeptide is selected from the group consisting of a ligase, a polymerase, a methylcytosine (mC) dioxygenase, a DNA glucosyltransferase, a proteinase K, and a thermolabile proteinase K.
[0065] Embodiment 11. The polypeptide of any of embodiments 9 to 10, wherein one or more of the reagents in the mixture further comprises a reversible inhibitor of the deaminase.
[0066] Embodiment 12. The polypeptide of any one of embodiments 1 to 11, wherein the mixture further comprises DNA.
[0067]
[0036] Embodiment 13. (a) combining a reaction mixture containing genomic DNA with a sequence-bias-free double-stranded DNA (dsDNA) deaminase; (b) deaminating at least 50% of the cytosines in the genomic DNA to uracil without a denaturation step that converts the dsDNA to single-stranded DNA (ssDNA); 1. A method for methylome analysis comprising:
[0068] Embodiment 14. The method of embodiment 13, wherein prior to (a), a methylcytosine (mC) dioxygenase for genomic DNA is added to the reaction mixture to convert mC to hydroxymethylcytosine (hmC).
[0069] Embodiment 15. The method of any one of embodiments 13 to 14, wherein prior to (a), a hydroxymethylcytosine (hmC) modifying reagent is added to the reaction mixture.
[0070] The method of any one of embodiments 13 to 15, wherein (b) further comprises inactivating the DNA deaminase with proteinase K or thermolabile proteinase K.
[0071] The method of any one of embodiments 13 to 16, wherein (b) further comprises amplifying DNA containing the converted cytosine.
[0072] Embodiment 18. The method according to any one of embodiments 13 to 17, further comprising determining the sequence of the amplified DNA.
[0073] Embodiment 19. The method according to any one of embodiments 13 to 18, further comprising determining the location of methylcytosine (mC) in the genomic DNA.
[0074] Embodiment 20. A kit comprising a deaminase capable of deaminating cytosines in double-stranded DNA (dsDNA) and optionally single-stranded DNA (ssDNA) without sequence bias.
[0075] Embodiment 21. The kit of embodiment 20, further comprising a methyldioxygenase in a container separate from the dioxygenase.
[0076] Embodiment 22. The kit of embodiment 20 or 21, further comprising a hydroxymethylcytosine (hmC) modifying enzyme in the same container as the dioxygenase or in a different container. EXAMPLES
[0077] [Example 1] In vitro expression of DNA deaminase Candidate DNA deaminase genes were first codon-optimized, and then flanking sequences were added to each end, specifically a sequence containing a T7 promoter at the 5' end and a sequence containing a T7 terminator at the 3' end. These sequences were ordered as linear gBlocks from Integrated DNA Technologies (Coralville, IA, USA). Template DNA for in vitro protein synthesis was generated using Phusion® Hot Start Flex DNA polymerase with gBlocks as template and flanking primers. PCR products were purified using Monarch PCR and DNA Purification Kit (New England Biolabs, Inc. Ipswich, MA, USA). DNA concentration was quantified using a NanoDrop spectrophotometer (Thermo Fisher Scientific, Inc. Waltham, MA, USA). 100–400 ng of the PCR fragment was used as template DNA to synthesize analytical amounts of DNA deaminase using the PURExpress in vitro protein synthesis kit (New England Biolabs, Inc. Ipswich, MA, USA) according to the manufacturer's recommendations.
[0078] Example 2: Deamination assays on single-stranded and double-stranded substrates To test the activity of the in vitro expressed DNA deaminase, 2 μl aliquots of PURExpress samples were mixed with 300 ng of ΦX174 Virion DNA (ssDNA substrate) or ΦX174 RF I DNA (dsDNA substrate) in a buffer containing 50 mM Bis-Tris pH 6.0, 0.1% TritonX-100 and incubated at 37°C for 1 h. Deaminated ΦX174 DNA was purified using a Monarch PCR and DNA Purification Kit (New England Biolabs, Inc. Ipswich, MA, USA). DNA concentration was quantified using a NanoDrop spectrophotometer (Thermo Fisher Scientific, Inc. Waltham, MA, USA). 150 ng of deaminated DNA was digested to nucleosides using Nucleoside Digestion Mix (New England Biolabs, Inc. Ipswich, MA, USA) according to the manufacturer's recommendations. LC-MS / MS analysis was performed by injecting the digested DNA on an Agilent 1290 Infinity II UHPLC equipped with a G7117A diode array detector and a 6495C triple quadrupole mass detector operating in positive electrospray ionization mode (+ESI). UHPLC was performed on a Waters XSelect HSS T3 XP column (2.1 x 100 mm, 2.5 μm) with a gradient mobile phase consisting of methanol and 10 mM aqueous ammonium acetate (pH 4.5). MS data collection was performed with dynamic multiple reaction monitoring (DMRM). Each nucleoside was identified in the extraction chromatogram associated with its specific MS / MS transition: dC[M+H] + m / z 228.1→112.1; dU [M+H] + m / z229.1→113.1;d m C[M+H] + m / z 242.1→126.1; and dT[M+H] + At m / z 243.1→127.1, the ratios were calculated within the analyzed samples using external calibration curves with known amounts of nucleosides.
[0079] Example 3: NGS deamination assay 50 ng of E. coli C2566 genomic DNA was combined with control modified DNA:
[0080] [Table 1]
[0081] DNA preparation The DNA was then transferred to a Covaris MicroTUBE (Covaris, Woburn, MA, USA) and sheared to 300 bp using a Covaris S2 instrument. 50 μl of the sheared material was transferred to a PCR strip tube to begin library construction. NEBNext DNA Ultra II reagents (New England Biolabs, Inc. Ipswich, MA, USA) were used according to the manufacturer's instructions for end repair, A-tailing, and adapter ligation using Illumina compatible adapters. The ligated sample was mixed with 110 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. The library was eluted in 17 μl of water.
[0082] Deamination The DNA was then deaminated using 1 μl of dsDNA deaminase synthesized as described above in 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 at 37°C for 1 h incubation time. After the deamination reaction, 1 μl of heat-labile proteinase K (New England Biolabs, Inc. Ipswich, MA) was added and incubated for another 30 min at 37°C. 5 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (New England Biolabs, Inc. Ipswich, MA, USA) were added to the DNA and PCR amplified. The PCR reaction sample was mixed with 50 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. The library was eluted in 15 μl of water. The library was analyzed and quantified by high sensitivity DNA analysis using a chip inserted into an Agilent Bioanalyzer 2100. Whole-genomic libraries were sequenced using the Illumina NextSeq platform. 150 cycles (2 × 75 bp) paired-end sequencing was performed for every sequencing run. Base calling and demultiplexing were performed using the standard Illumina pipeline. Results for CseDa01 are shown in Figures 4A and 4B.
[0083] [Example 4] 1-Tube-3-Enzyme EM-seq (dsDNA deaminase MGYPDa829+ TET2+ BGT) 50 ng of NA12878 genomic DNA was combined with 0.1 ng of CpG methylated pUC19 and 1 ng of unmethylated lambda control DNA to make up to 50 μl in 5 mM Tris pH 8.0. DNA was prepared according to Example 3 and the library was eluted in 29 μl of water. DNA was oxidized in a 50 μl reaction volume containing 50 mM Tris HCl pH 8.0, 1 mM DTT, 5 mM sodium-L-ascorbate, 20 mM a-KG, 2 mM ATP, 50 mM ammonium iron(II) sulfate hexahydrate, 0.04 mM UDG-glucose (NEB, Ipswich, MA), 16 μg mTET2, 10 U T4-BGT (NEB, Ipswich, MA). The reaction was initiated by the addition of Fe(II) solution to a final reaction concentration of 40 μM and then incubated at 37° C. for 1 hour. The DNA was then deaminated using 1 μl of MGYPDa829 dsDNA deaminase at 37°C with an incubation time of 3 h. After the deamination reaction, 1 μl of heat-labile proteinase K (P8111S, New England Biolabs, Inc. Ipswich, MA) was added and incubated for another 30 min at 37°C and 15 min at 60°C. At the end of the incubation, the DNA was purified using 70 μl of resuspended NEBNext sample purification beads according to the manufacturer's protocol. The sample was eluted in 16 μl of water and 15 μl was transferred to a new tube. 1 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Inc. Ipswich, MA) were added to the DNA and PCR amplified. The libraries were analyzed and quantified on an Agilent Bioanalyzer 2100 DNA analyzer. The entire genomic library was sequenced and analyzed as described below.
[0084] Raw reads were first trimmed by Trim Galore software to remove adapter sequences and low-quality bases from the 3' end. Unpaired reads resulting from adapter / quality trimming were also removed during this step. Trimmed read sequences were C-to-T converted and then mapped to a composite reference sequence containing the complete sequence of the human genome (GRCh38) and lambda and pUC19 controls using the Bismark program (Langmead and Salzberg 2012) with default Bowtie2 settings. The aligned reads then underwent two post-processing QC steps: 1, alignment pairs sharing the same alignment start position (5' end) were considered as PCR duplicates and discarded; 2, reads that aligned to the human genome and contained excess cytosines in non-CpG contexts (e.g., more than 3 in 75 bp) were removed as they likely resulted from conversion errors. The number of T (converted and unmethylated) and C (unconverted and modified) at each covered cytosine position was then calculated from the remaining good alignments using the Bismark methylation extractor, and the methylation level was calculated as #C / (#C+#T). Figure 3C illustrates this workflow.
[0085] [Example 5] CseDa01 DNA deaminase does not deaminate 5caC and 5fC 1500 ng of oligonucleotides with one modified cytosine (5caC or 5fC) (ACACCCATCACATTTACAC(5caC)GGGAAAGAGTTGAATGTAGAGTTGG; SEQ ID NO: 157) or ACACCCATCACATTTACAC(5fC)GGGAAAAGAGTTGAATGTAGAGTTGG; SEQ ID NO: 158) were treated with CseDa01 DNA deaminase for 4 hours in a buffer containing 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 and incubated at 37°C for 1 hour. The deaminated oligonucleotides were purified using Monarch PCR and DNA purification kit (New England Biolabs, Inc. Ipswich, MA, USA). DNA concentration was quantified using a NanoDrop spectrophotometer (Thermo Fisher Scientific, Inc. Waltham, MA, USA). 1500 ng of deaminated DNA was digested to nucleosides using Nucleoside Digestion Mix (New England Biolabs, Inc. Ipswich, MA, USA) according to the manufacturer's recommendations. UHPLC-MS analysis was performed on a Waters XSelect HSS T3 XP column (2.1 × 100 mm, 2.5 μm) using a gradient mobile phase consisting of methanol and 10 mM ammonium acetate buffer (pH 4.5) using an Agilent 1290 Infinity II UHPLC equipped with a G7117A diode array detector and a 6135 XT MS detector. The identity of each peak was confirmed by MS. The relative amount of each nucleoside was determined by the integration of each peak at 260 nm or its respective UV absorption maximum. The results are shown in Figure 4C.
[0086] [Example 6] 1-Tube-2-Enzyme EM-seq using dsDNA deaminase CseDa01+TET2 50 ng of NA12878 genomic DNA was combined with 0.1 ng of CpG methylated pUC19 and 1 ng of unmethylated lambda control DNA to make up to 50 μl in 5 mM Tris pH 8.0. DNA was prepared according to Example 3 and the library was eluted in 29 μl of water. DNA was oxidized in a 50 μl reaction volume containing 50 mM Tris HCl pH 8.0, 1 mM DTT, 5 mM sodium-L-ascorbate, 20 mM a-KG, 2 mM ATP, 50 mM ammonium iron(II) sulfate hexahydrate, and 16 μg of mTET2. The reaction was initiated by adding Fe(II) solution to a final reaction concentration of 40 μM and then incubated at 37° C. for 1 hour. DNA was then deaminated using 1 μl of CseDa01 dsDNA deaminase at 37° C. with an incubation time of 3 hours. After the deamination reaction, 1 μl of heat-labile proteinase K (P8111S, New England Biolabs, Ipswich, MA) was added and incubated for an additional 30 min at 37°C and 15 min at 60°C. At the end of the incubation, DNA was purified using 70 μl of resuspended NEBNext sample purification beads according to the manufacturer's protocol. Samples were eluted in 16 μl of water and 15 μl was transferred to a new tube. 1 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Ipswich, MA) were added to the DNA and PCR amplified. Libraries were analyzed and quantified on an Agilent Bioanalyzer 2100 DNA analyzer. Whole genomic libraries were sequenced and analyzed as described below. Raw reads were first trimmed with Trim Galore software to remove adapter sequences and low quality bases from the 3' ends. Unpaired reads due to adapter / quality trimming were also removed during this step.Trimmed read sequences were converted from C to T and then mapped to a composite reference sequence containing the complete sequence of the human genome (GRCh38) as well as lambda and pUC19 controls using the Bismark program (Langmead and Salzberg 2012) with default Bowtie2 settings. The aligned reads then underwent two post-processing QC steps: 1, alignment pairs sharing the same alignment start position (5' end) were considered PCR duplicates and discarded; 2, reads that aligned to the human genome and contained excess cytosines in non-CpG contexts (e.g., more than 3 in 75 bp) were removed as they likely resulted from conversion errors. The number of T (converted and unmethylated) and C (unconverted and modified) at each covered cytosine position was then calculated from the remaining good quality alignments using the Bismark methylation extractor, and the methylation level was calculated as #C / (#C+#T). Figure 3C illustrates this workflow.
[0087] Example 7: DNA deaminase CseDa01 functions highly efficiently in TET2 buffer, allowing single-tube 5mC oxidation and DNA deamination reactions To test the activity of CseDa01 DNA deaminase in TET2 buffer, 2 μl of PURExpress sample was mixed with 300 ng of ΦX174 Virion DNA (ssDNA substrate) or ΦX174 RF I DNA (dsDNA substrate) in a buffer containing 50 mM Tris HCl pH 8.0, 1 mM DTT, 5 mM sodium-L-ascorbic acid, 20 mM a-KG, 2 mM ATP, 50 mM ammonium iron(II) sulfate hexahydrate, 0.04 mM, and incubated at 37° C. for 1 h. Deaminated ΦX174 DNA was purified using the Monarch PCR and DNA Purification Kit (New England Biolabs, Inc. Ipswich, MA, USA). DNA concentration was quantified using a NanoDrop spectrophotometer (Thermo Fisher Scientific, Inc. Waltham, MA, USA). 150 ng of deaminated DNA was digested to nucleosides using Nucleoside Digestion Mix (New England Biolabs, Inc. Ipswich, MA, USA) according to the manufacturer's recommendations. LC-MS / MS analysis was performed by injecting the digested DNA on an Agilent 1290 Infinity II UHPLC equipped with a G7117A diode array detector and a 6495C triple quadrupole mass detector operated in positive electrospray ionization mode (+ESI). UHPLC was performed on a Waters XSelect HSS T3 XP column (2.1 × 100 mm, 2.5 μm) with a gradient mobile phase consisting of methanol and 10 mM aqueous ammonium acetate (pH 4.5). MS data collection was performed with dynamic multiple reaction monitoring (DMRM) mode. Each nucleoside was identified in the extraction chromatogram associated with its specific MS / MS transition: dC[M+H] + m / z 228.1→112.1; dU [M+H] + m / z229.1→113.1;d m C[M+H] + m / z 242.1→126.1; and dT[M+H] +At m / z 243.1→127.1. The ratios within the analyzed samples were calculated using external calibration curves with known amounts of nucleosides. The results are shown in Figures 4A, 4B, 4C, 5A, and 5B.
[0088] Example 8 Modification-sensitive deaminases efficiently deaminate cytosine to uracil but do not deaminate 5-methylcytosine and 5-hydroxymethylcytosine in dsDNA and ssDNA 50 ng of E. coli C2566 genomic DNA was combined with 2 ng of unmethylated lambda, phage XP12 (all cytosines are 5-methylcytosines) and T4 phage DNA (all cytosines are 5-hydroxymethylcytosines) control DNA and made up to 50 μl in 10 mM Tris pH 8.0. The DNA was then prepared according to Example 3 with a shear size of 240-290 bp and a library elution volume of 15 μl of water. The DNA was then deaminated using 1 μl of a modification-sensitive dsDNA deaminase (e.g., MGYPDa20 or NsDa01) synthesized as described above in 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 at 37 °C with an incubation time of 1 h. After the deamination reaction, 1 μl of heat-labile proteinase K (P8111S, New England Biolabs, Ipswich, MA) was added and incubated for an additional 30 min at 37 °C. 1 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Ipswich, MA) were added to the DNA and PCR amplified. The PCR reaction sample was mixed with 50 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. The library was eluted in 15 μl of water. The library was analyzed and quantified by high-sensitivity DNA analysis using a chip inserted into an Agilent Bioanalyzer 2100. Whole genome libraries were sequenced using the Illmina NextSeq platform. 150 cycles (2 × 75 bp) paired-end sequencing was performed for all sequencing runs. Base calling and demultiplexing were performed using the standard Illumina pipeline. The raw reads were first trimmed by Trim Galore to remove adapter sequences and low-quality bases from the 3' end. Unpaired reads resulting from adapter / quality trimming were also removed during this step.Trimmed read sequences were C-to-T converted and then mapped to a composite reference sequence containing the complete sequences of the E. coli C2566 genome as well as lambda, phage XP12, and T4 controls using the Bismark program with default Bowtie2 settings.
[0089] The first 5 bp at the 5' end of R2 reads were removed to reduce end-repair errors, and aligned read pairs sharing the same alignment start position (5' end) were considered as PCR duplicates and discarded. The next deamination event (C->T) was called by comparing the remaining well-aligned sequences to the reference sequence using the Bismark methylation extractor program. Then, 20 bp flanking sequences (10 bp upstream and 10 bp downstream) of all covered cytosines from each individual genome were extracted, and cytosine sites were divided into different groups based on their deamination rate (>=90%, >=50%, >=25% or <=10%). The flanking sequences of each cytosine group were used to generate sequence logos using WebLogo3 to infer deamination sequence preference. The results are shown in Figures 6A and 6B for MGYPDa20, Figures 7A and 7B for NsDa01, Figures 8A and 8B for RhDa01_extN10, and Figures 9A and 9B for MmgDa02.
[0090] Example 9: Application of the 1-tube-1-enzyme EM-seq method to localize 5mC in humans using the modification-sensitive dsDNA deaminase MGYPDa20 50 ng of NA12878 genomic DNA was combined with 0.1 ng of CpG methylated pUC19 and 1 ng of unmethylated lambda control DNA and made up to 50 μl in 5 mM Tris pH 8.0. DNA was prepared according to Example 3 and the library was eluted in 17 μl of molecular grade water. The DNA was then deaminated using 1 μl of MGYPDa20 dsDNA deaminase in 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 at 37° C. for an incubation time of 3 hours. After the deamination reaction, 1 μl of heat labile proteinase K (P8111S, New England Biolabs, Ipswich, MA) was added and incubated for an additional 30 minutes at 37° C. PCR amplification was performed by combining 5 μM NEBNext Unique Dual Index primer, 20 μM deaminated DNA, and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Ipswich, MA). PCR reaction samples were mixed with 50 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. Libraries were eluted in 15 μl of water. Libraries were analyzed and quantified by high-sensitivity DNA analysis using a chip inserted into an Agilent Bioanalyzer 2100. Whole genome libraries were sequenced using the Illmina NextSeq platform and analyzed as described below. Raw reads were first trimmed with Trim Galore software to remove adapter sequences and low-quality bases from the 3' end. Unpaired reads resulting from adapter / quality trimming were also removed during this step. Trimmed read sequences were C-to-T converted and then aligned to a composite reference sequence containing the complete sequence of the human genome (GRCh38) and lambda and pUC19 controls using the Bismark program (Langmead and Salzberg 2012) with default Bowtie2 settings.The aligned reads then underwent two post-processing QC steps: 1, alignment pairs sharing the same alignment start position (5' end) were considered PCR duplicates and discarded; 2, reads that aligned to the human genome and contained excess cytosines in non-CpG contexts (e.g., more than 3 in 75 bp) were removed as they likely resulted from conversion errors. Next, the number of T (converted and unmethylated) and C (unconverted and modified) at each covered cytosine position was calculated from the remaining good quality alignments using the Bismark methylation extractor, and the methylation level was calculated as number of Cs / (number of Cs+number of Ts). Figure 3D illustrates this workflow. The results are shown in Figure 10.
[0091] Example 10: Preparation of methyl-SNP-seq libraries using MGYPDa20 DNA deaminase For whole human genome methyl-SNP-seq sequencing, 4 mg of NA12878 gDNA and 40 ng of unmethylated lambda DNA spiked in to monitor deamination efficiency were used. Genomic DNA was fragmented using a 250 bp sonication protocol using a Covaris S2 sonicator. Two technical replicates were set up. The fragmented gDNA was end-repaired, dA-tailed (NEB Ultra II E7546 module), and then ligated to custom hairpin adapters using NEB Ligase Master Mix (NEB, M0367). Incomplete ligation products (fragments with only one adapter ligated or no adapter ligated) were removed using two exonucleases (NEB exoIII and NEB exoVII). After treatment with UDG and EndoVIII, two nicks were created at the uracil positions in the hairpin adapters at both ends. The nick sites were translated towards the 3' end by DNA polymerase I in the presence of dATP, dGTP, dTGP and 5-methyl-dCTP. Nick translation causes a double-stranded DNA break when DNA polymerase I encounters the other nick on the opposite strand. The resulting fragments have one end ligated to a hairpin adaptor and a blunt end on the other side. The blunt end was dA-tailed and ligated with a methylated Illumina adaptor. The ligated products were deaminated with double-stranded DNA deaminase MGYPDa20 for 3 hours at 37°C. The deaminated DNA products were amplified using NEBNext Q5U Master Mix (NEB, M0597). The resulting indexed library was used for Illumina sequencing. The human methyl-SNP-seq library was sequenced using an Illumina Novaseq 6000 sequencer for 100 bp paired-end reads.
[0092] [Example 11] Detection of N4mC-modified DNA with CseDa01 dsDNA deaminase 50 ng of Paenibacillus sp. JDR-2 (CCGG target sequence) and Salmonella enterica FDAARGOS_312 (CACCGT target sequence) DNA were combined with 0.1 ng of CpG methylated pUC19 and 1 ng of unmethylated lambda control DNA and made up to 50 μl in 5 mM Tris pH 8.0. DNA was prepared according to Example 3 with a shear size of 240-290 bp and an elution volume of 15 μl of water. DNA was then deaminated using 1 μl of CseDa01 dsDNA deaminase synthesized as described above in 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 at 37°C with an incubation time of 1 h. After the deamination reaction, 1 μl of heat-labile proteinase K (P8111S, New England Biolabs, Ipswich, MA) was added and incubated for an additional 30 min at 37 °C. 1 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Ipswich, MA) were added to the DNA and PCR amplified. The PCR reaction sample was mixed with 50 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. The library was eluted in 15 μl of water. The library was analyzed and quantified by high-sensitivity DNA analysis using a chip inserted into an Agilent Bioanalyzer 2100. Whole genome libraries were sequenced using an Illmina NextSeq platform. 150 cycles (2 × 75 bp) paired-end sequencing was performed for all sequencing runs. The raw reads were first trimmed by Trim Galore to remove adapter sequences and low quality bases from the 3' end. Unpaired reads due to adapter / quality trimming were also removed during this step. Trimmed read sequences were converted from C to T and then aligned to the reference sequence and the complete sequences of lambda and pUC19 controls using the Bismark program with default Bowtie2 settings.The first 5 bp at the 5' end of R2 reads were removed to reduce end-repair errors, and aligned read pairs sharing the same alignment start position (5' end) were considered PCR duplicates and discarded. Subsequent deamination events (C->T) were called by comparing the remaining well-aligned sequences to the reference sequence using the Bismark methylation extractor program. N4mC modified sites were called if they were mostly deaminated (C->T conversion rate <=20%). The adjacent 20 bp sequences of all called N4mC sites were extracted and sequence logos were generated using WebLogo3. The results are shown in Figures 11A and 11B.
[0093] [Example 12] Detection of N4mC and 5mC modified DNA using CseDa01 dsDNA deaminase and MGYPDa20 dsDNA deaminase 50 ng of NEB1569 Thermus species M and NEB394 Acinetobacter species H genomic DNA were combined with 0.1 ng of CpG methylated pUC19 and 1 ng of unmethylated lambda control DNA and made up to 50 μl in 5 mM Tris pH 8.0. The DNA was then prepared according to Example 3 with a shear size of 240-290 bp and a library elution volume of 15 μl of water. The DNA was then deaminated using 1 μl of dsDNA deaminase synthesized as described above in 50 mM Bis-Tris pH 6.0, 0.1% Triton X-100 at 37° C. with an incubation time of 1 hour. After the deamination reaction, 1 μl of heat-labile proteinase K (P8111S, New England Biolabs, Ipswich, MA) was added and incubated for an additional 30 min at 37 °C. 1 μM of NEBNext Unique Dual Index primer and 25 μl of NEBNext Q5U Master Mix (M0597, New England Biolabs, Ipswich, MA) were added to the DNA and PCR amplified. The PCR reaction sample was mixed with 50 μl of resuspended NEBNext sample purification beads and purified according to the manufacturer's instructions. The library was eluted in 15 μl of water. The library was analyzed and quantified by high-sensitivity DNA analysis using a chip inserted into an Agilent Bioanalyzer 2100. Whole genome libraries were sequenced using the Illmina NextSeq platform. 150 cycles (2 × 75 bp) paired-end sequencing was performed for all sequencing runs. Base calling and demultiplexing were performed using the standard Illumina pipeline. The raw reads were first trimmed by Trim Galore to remove adapter sequences and low-quality bases from the 3' end. Unpaired reads resulting from adapter / quality trimming were also removed during this step.Trimmed read sequences were converted from C to T and then mapped to a composite reference sequence containing NEB1569 Thermus sp. M and NEB394 Acinetobacter sp. H, as well as the complete sequences of lambda and pUC19 controls, using the Bismark program with default Bowtie2 settings. The first 5 bp at the 5' end of R2 reads were removed to reduce end-repair errors, and aligned read pairs sharing the same alignment start position (5' end) were considered PCR duplicates and discarded. The next deamination event (C->T) was called by comparing the remaining well-aligned sequences to the reference sequence using the Bismark methylation extractor program. N4mC modifications are called from the CseDa01 deaminase-treated library. N4mC-modified sites are called if they are mostly not deaminated (C->T conversion rate <=20%). For 5mC modification detection, differential methylation analysis was performed between MGYPDa20 deaminase-treated libraries (detecting both N4mC and 5mC) and CseDa01 deaminase-treated libraries (detecting only N4mC) of the same sample to identify modified sites (i.e., 5mC) detected only in the MGYPDa20 library. Differentially methylated sites were called by logistic regression using the Methylkit program with a SLIM-corrected Q value <= 0.01 and methylation difference >= 80%. To identify methyltransferase recognition sequences, 9-bp flanking sequences including 4-bp upstream and 4-bp downstream of all modified sites were extracted, and unique 9-bp sequences were clustered using a hierarchical linkage method based on the differences between each pair of sequences. Sequence logos were created using WebLogo3 for each cluster representing different methyltransferase recognition motifs.
[0094] [Example 13] Candidate selection A list of HMMER3 (Eddy, SR Accelerated Profile HMM Searches. PLOS Comput. Biol. 7, e1002195 (2011)) cytosine deaminase sequence profiles was curated. Twenty-nine profiles were derived from a CDA clan (CL0109) from the Pfam (Mistry, J. et al., Pfam: The protein families database in 2021 Nucleic Acids Res. 49, D412-D419 (2021)) database (excluding TM1506, LpxI_C, FdhD-NarQ, and AICARFT_IMPCHas, which do not code for deaminases), 17 profiles were built from a multiple sequence alignment (MSA) of the deaminase family defined by Iyer et al. (Nucleic Acids Res. 39, 9473-9497, 2011), and one profile was built from a multiple sequence alignment found in Zhang et al., (Biol. Direct 7, 18, 2012).
[0095] Some candidate sequences were selected directly from the MSAs published in Iyer et al. (2011) and Zhang et al. (2012). Other candidate sequences were identified in six different databases: UniProt, Mgnify, IMG / VR, IMG / M, wastewater treatment plant metagenomes, and GenBank (respectively, The UniProt Consortium.UniProt: the universal protein knowledgebase in 2021 Nucleic Acids Res.49, D480-D489 (2021); Mitchell, A. L. et al., Mgnify: the microbiome analysis resource in 2020 Nucleic Acids Res.48, D570-D578 (2020); Paez-Espino, D. et al., IMG / VR: a database of cultured and uncultured DNA Viruses and retroviruses.Nucleic Acids Res.45, gkw1030 (2017); Chen, I. M. et al., The IMG / M data management and analysis system v.6.0: new tools and advanced capabilities.Nucleic Acids Res. 49, D751-D763 (2021); Singleton, C. M. et al., Connecting structure to function with the recovery of over 1000 high-quality metagenome-assembled genomes from activated sludge using long-read sequencing. Nat. Commun. 12, 2009 (2021); and Da, B. et al., GenBank. Nucleic Acids Res. 41, (2013)).
[0096] The majority of the deaminases tested were found as fusions with larger proteins, for example as part of polymorphic toxin systems. To define the boundaries of the deaminase domain, AlphaFold2 (Jumper, J. et al., Highly accurate protein structure prediction with AlphaFold. Nature 1-11 (2021) doi:10.1038 / s41586-021-03819-2) structural predictions were made and visualized. The N-terminal truncation site was generally chosen a few amino acids before helix 1 of the deaminase domain.
[0097] For convenience, each screened sequence was given a short name. The names are arbitrary but somehow related to the database or species of origin of the sequence. Da = deaminase, MGYP = Mgnify protein, Hm = hot metagenomics, VR = IMG / VR, WWTP = wastewater treatment plant, Chimera = chimeric sequence, Anc = ancestral sequence reconstruction. Other prefixes are mostly two or three letters derived from the name of the source organism or source environment of the metagenomic data. Some sequences also have prefixes or suffixes of the form extN#, extC#, d#, Cd#, which indicate an N-terminal extension, C-terminal extension, N-terminal deletion, and C-terminal deletion, respectively, of the indicated number of residues compared to the candidate without a name attached.
[0098] All amino acid sequence alignments were calculated using MAFFT (v7.490) (Katoh, K. & Standley, DMMAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability. Mol. Biol. Evol. 30, 772-780 (2013)) using global pair mode. Phylogenies were generated using raxml-ng (v.1.1) (Kozlov, A. M., Darriba, D. Flouri, T. Morel, B. & Stamatakis, A. RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics 35, 4453-4455 (2019)). Ancestral sequence reconstructions were constructed from the phylogenetic trees using raxml-ng (v.1.1).
[0099] [Example 14] Summary table The assay results for the 29 deaminases are shown in Table 1 below, where APOBEC3A (a single-stranded DNA deaminase) serves as a negative control. The other 28 deaminases in the table (double-stranded DNA deaminases) all have significant activity against double-stranded DNA substrates.
[0100] The double-stranded DNA deaminases disclosed herein may be used in numerous methods, processes, and workflows, including, for example, the applications shown in Table 2 below. The deamination products may contain one or more modified cytosines, for example, where the substrate dsDNA contained such modified cytosines and the operative deaminase does not deaminate or only poorly deaminates such modified cytosines. Each of the listed methods / applications may further include (a)(i) sequencing the deamination products, and / or (ii) amplifying the deamination products (e.g., by PCR) to generate amplification products and sequencing the amplification products to generate sequence reads in each of (a)(i) and (a)(ii), and (b) optionally determining the type and / or location of the modified cytosines in the dsDNA substrate from the sequence reads.
[0101] Screening results for over 100 deaminases are shown in Table 3 below, where APOBEC3A (single-stranded DNA deaminase) serves as a negative control. Many were observed to have double-stranded DNA deaminase activity under the conditions tested. The relevance of the enzymes tested is illustrated in Figure 1, in which it can be seen that deaminases that show limited or moderate activity under certain conditions tested may have higher activity under alternative or optimized conditions.
[0102] The names and sequence numbers of certain double-stranded DNA deaminases disclosed herein are set forth in Table 4, along with the corresponding names contained in U.S. Provisional Patent Application No. 63 / 264,513, filed November 24, 2022.
[0103] [Table 2]
[0104] [Table 3] TIFF2024543137000005.tif224168TIFF2024543137000006.tif220168TIFF2024543137000007.tif248168TIFF2024543137000008.tif209168TIFF2024543137000009.tif214168TIFF2024543137000010.tif255167
[0105]
Table 4
[0106]
Table 5
Claims
1. 1. A method for deaminating double-stranded nucleic acids, comprising: a double-stranded DNA substrate containing cytosine; contacting a double-stranded DNA deaminase having an amino acid sequence at least 80% identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, and 99; Producing a deaminated product comprising a deaminated cytosine. A method comprising:
2. The method of claim 1 , wherein the double-stranded DNA substrate further comprises a modified cytosine.
3. 3. The method of claim 2, wherein the modified cytosine is 5fC, 5CaC, 5mC, 5hmC, N4mC, 5ghmC, or pyrrolo-C.
4. sequencing the deaminated products, or amplifying the deaminated products to generate amplification products, and sequencing the amplification products to generate sequence reads, in each case. The method of claim 1 further comprising:
5. analyzing the sequence reads to identify modified cytosines in the double-stranded DNA substrate. The method of claim 4 further comprising:
6. 2. The method of claim 1, wherein the double-stranded DNA substrate is eukaryotic or bacterial DNA.
7. 2. The method of claim 1, wherein the double-stranded DNA substrate is human cfDNA.
8. 2. The method of claim 1, wherein the double-stranded DNA deaminase has an amino acid sequence that is at least 90% identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, and 99.
9. 2. The method of claim 1, wherein the double-stranded DNA substrate is pretreated with TET methylcytosine dioxygenase and DNA beta-glucosyltransferase.
10. 10. The method of claim 9, wherein the double-stranded DNA deaminase has an amino acid sequence at least 90% identical to any of the following SEQ ID NOs: MGYPDa829 (SEQ ID NO:96), MGYPDa06 (SEQ ID NO:4), CrDaOl (SEQ ID NO:12), AvDa02 (SEQ ID NO:2), CsDaOl (SEQ ID NO:9), LbsDaOl (SEQ ID NO:10), FlDaOl (SEQ ID NO:8), MGYPDa26 (SEQ ID NO:7), MGYPDa23 (SEQ ID NO:6), Chimera_10 (SEQ ID NO:97), and AncDa04 (SEQ ID NO:95).
11. 10. The method of claim 1, wherein the double-stranded DNA substrate is pretreated with TET methylcytosine dioxygenase but not with DNA beta-glucosyltransferase.
12. 12. The method of claim 11, wherein the double-stranded DNA deaminase has an amino acid sequence that is at least 90% identical to any of the SEQ ID NOs: CseDa01 (SEQ ID NO: 3) and LbDa02 (SEQ ID NO: 1).
13. 2. The method of claim 1, wherein the double-stranded DNA substrate is not pretreated with TET methylcytosine dioxygenase or DNA beta-glucosyltransferase.
14. The method of claim 1, wherein the double-stranded DNA substrate comprises at least one N4mC.
15. 15. The method of claim 14, wherein the double-stranded DNA substrate is bacterial DNA.
16. 14. The method of claim 13, wherein the double-stranded DNA deaminase has an amino acid sequence that is at least 90% identical to any of the following SEQ ID NOs: MGYPDa20 (SEQ ID NO: 11), NsDa01 (SEQ ID NO: 27), and AshDa01 (SEQ ID NO: 40).
17. To generate a double-stranded DNA substrate: (a) ligating hairpin adaptors to double-stranded fragments of DNA to create ligation products; (b) enzymatically generating a free 3′ end at the double-stranded region of the hairpin adapter in the ligation product; and 2. The method of claim 1, further comprising the step of (c) extending the free 3' end in a dCTP-free reaction mixture comprising a strand-displacing or nick-translation polymerase, dGTP, dATP, dTTP, and modified dCTP.
18. 18. The method of claim 17, wherein the modified dCTP is 5mdCTP, pyrrolo-dCTP, 5hmdCTP, or N4mdCTP.
19. 18. The method of claim 17, wherein the double-stranded DNA deaminase has an amino acid sequence that is at least 90% identical to any of the following SEQ ID NOs: MGYPDa20 (SEQ ID NO: 11), NsDa01 (SEQ ID NO: 27), AshDa01 (SEQ ID NO: 40).
20. 1. An enzyme comprising an amino acid sequence that is at least 80% identical to the C-terminal deaminase domain of a naturally occurring protein, (a) has double-stranded DNA deaminase activity; and (b) does not contain the N-terminus of a naturally occurring protein enzyme.
21. 21. The enzyme of claim 20, which is 300 amino acids or less in length.
22. 21. The enzyme of claim 20, which is at least 80% identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, and 99.
23. 21. The enzyme of claim 20, fused to a catalytically inactive Cas9 (dCas9) or a nicking Cas9 (nCas9) or a transcription activator-like effector nuclease (TALEN).
24. (a) the enzyme of claim 20; and (b) Reaction buffer Kit including:
25. further comprising TET methylcytosine dioxygenase and DNA beta-glucosyltransferase; or Further comprising a TET methylcytosine dioxygenase and no DNA beta-glucosyltransferase; 25. The kit of claim 24.
26. 25. The kit of claim 24, which is free of TET methylcytosine dioxygenase and DNA beta-glucosyltransferase.
27. 25. The kit of claim 24, further comprising a modified dCTP selected from 5mdCTP, pyrrolo-dCTP, 5hmdCTP, and N4mdCTP.
28. (a) a double-stranded DNA substrate containing cytosine; and (b) a double-stranded DNA deaminase having an amino acid sequence at least 80% identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, and 99; A reaction mixture comprising:
29. 30. The reaction mixture of claim 28, wherein the double-stranded DNA substrate comprises cytosine and at least one modified cytosine.
30. 30. The reaction mixture of claim 29, wherein the modified cytosine is 5fC, 5caC, 5mC, 5hmC, N4mC, or pyrrolo-C.
31. 30. The reaction mixture of claim 28, wherein the double-stranded DNA substrate comprises eukaryotic or bacterial DNA.
32. 29. The reaction mixture of claim 28, wherein the double-stranded DNA substrate is human cfDNA.
33. 29. The reaction mixture of claim 28, wherein the deaminase has an amino acid sequence at least 90% identical to any of SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 19, 24, 26, 27, 28, 33, 40, 49, 50, 63, 95, 96, 97, and 99.