Glycosyltransferases and related methods

The method of using glycosyltransferases and saccharides to modify and detect hydroxylated bases in nucleic acids addresses the limitations of current detection methods, providing a more effective means of characterizing these modifications and their biological significance.

WO2025129193A1PCT designated stage expired Publication Date: 2025-06-19NEW ENGLAND BIOLABS INC
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/060408
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-16
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods are limited in their ability to efficiently detect and characterize hydroxylated modified bases in nucleic acids across various organisms, which are crucial for understanding gene regulation, nucleic acid stability, and potential applications in vaccine and nucleic acid-based therapies.

Method used

A method involving the use of glycosyltransferases (GTs) and saccharides to add saccharides or oligosaccharides to hydroxylated modified bases in nucleic acids, allowing for detection and differentiation of these modified bases through labeling and sequencing techniques.

Benefits of technology

This method enables the accurate detection and characterization of hydroxylated modified bases, improving our understanding of their functions and roles in various biological processes, and potentially enhancing therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060408_19062025_PF_FP_ABST
    Figure US2024060408_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein is a method for making a nucleic acid that contains a saccharide- modified base. In some embodiments, the method may comprise contacting a nucleic acid with (i) a glycosyltransferase (GT) and (ii) an activated saccharide that is not UDP-glucose or a functionalized UDP-glucose, to add the saccharide of the activated saccharide onto a hydroxylated base in the nucleic acid. Kits for performing the method are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GLYCOSYLTRANSFERASES AND RELATED METHODS

[0002] CROSS-REFERENCING

[0003] This application claims the benefit of U.S. provisional application serial nos. 63 / 610,736, filed on December 15, 2023, and 63 / 610,706, filed on December 15, 2023, which applications are incorporated by reference herein.

[0004] SEQUENCE LISTING

[0005] This application is filed with a Sequence Listing in electronic form as a Sequence Listing XML, "NEB-499. xml" created on December 15, 2023, and having a size of 20,414 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in their entirety.

[0006] BACKGROUND

[0007] RNA and DNA modifications occur in eukaryotes and prokaryotes, as well as in their viruses, and serve a wide range of functions, from gene regulation to nucleic acid protection. Although the first nucleotide modification was discovered almost 100 years ago, new modifications are continuously being uncovered. A large diversity of nucleotide modifications is found among pyrimidines. Many of these modifications are hydroxylated and are found on bases at positions that do not interfere with Watson Crick bonding. These modifications may not only affect gene expression but also the stability of the nucleic acid through base-stacking possibly due to hydrophobicity, charge polarization, increased steric encumbrance, or blocking hydrogen bonds. Other functions of modified bases may include protection from nucleases and mutation. It would be desirable to develop tools to identify and characterize the presence and location of hydroxylated modified bases in nucleic acids throughout the nucleic acid universe from bacteriophages to eukaryotes. In addition to intrinsic biological information, identification of base modifications may provide means for stabilizing nucleic acids while permitting their replication, transcription and / or translation. Stabilization of nucleic acids may play a role in vaccine and nucleic acid-based therapies for preventing or treating human disease.

[0008] Additionally, because bacterial viruses are known to repurpose diverse bacterial base modifications, it is possible that such base modifications although rare may be found in the nucleic acid of eukaryotes. It would be desirable to identify enzymes that can recognize and label different modified nucleotides to analyze the presence and function of such modifications in viruses, bacteria, archaea and eukaryotic nucleic acids. Knowing when, where and how varied base modifications occur in eukaryotes may play a role in the understanding of variations in normal and disease phenotype in human populations.

[0009] Methylation of cytosine at the C5 position (5mC) in eukaryotic nucleic acid has been studied extensively in recent years. A few decades ago, an additional modification of 5mC was identified in bacteria and in phage, namely 5-hydroxymethylcytosine (5hmC). 5hmC has also been described in the human genome of various tissues particularly in the human brain. It was reported that 5hmC in nucleic acid samples can be modified by the attachment of glucose to yield 5-glucosylmethylcytosine (gmC) using phage T4 beta glucosyltransferase (BGT). Subsequently, it was reported that the enzyme T4 BGT could be used add keto-glucose or azide substituted glucose or glucosamine to 5hmC to enable the glucose to be further labelled with a synthetic label via azide using click chemistry or by reaction with the amine on glucosamine. Various methods for distinguishing between 5mC and 5hmC and for detecting and mapping 5hmC have been established (see for example US 10,227,646, US 10, 260,088, that utilize a Ten Eleven translocation dioxygenase (TET)), mutant, BGT and a deaminase (EM-Seq New England Biolabs, also see a review Wanget al in Mol Metab. 2022 Mar; 57: 101314). It would be desirable to have additional methods for characterizing the presence of 5hmC in nucleic acids.

[0010] SUMMARY

[0011] In general, a method is provided for detecting a hydroxylated base in a nucleic acid. The method includes combining a glycosyltransferase (GT) and a population of saccharides (e.g., activated saccharides) with a nucleic acid so that a saccharide (e.g., an oligosaccharide) is added to the hydroxylated modified base by the GT. The population of saccharides (e.g., activated saccharides) may include different types of saccharides. A single GT may have specificity for adding a single type of saccharide or optionally several types of saccharides so as to attach a saccharide to the hydroxylated modified base. For example, depending on the desired application, a selected GT may be specific for adding one type of monosaccharide, or nonspecific and perform one addition of different monosaccharides, or the selected GT may specifically or non-specifically add multiple monosaccharide units. The saccharide attachment provides a means to detect the hydroxylated modified base in the nucleic acid and optionally to distinguish it from other types of hydroxylated modified bases. In one embodiment, a plurality of different GTs having different saccharide specificities may be used to detect and optionally distinguish different hydroxylated modified bases.

[0012] If at least one of the saccharides added to the hydroxylated modified base is labelled with a characteristic label, the hydroxylated modified base can be readily detected, characterized, and optionally distinguished from other types of hydroxylated base.

[0013] Alternatively, the presence of an oligosaccharide on a hydroxylated modified base may generate a pronounced signal that is greater than one presently observed through glucosylation with a single glucose, which may detect, for example, by nanopore sequencing. In this context, an oligosaccharide containing 2 or more sugars, and in some embodiments, at least 3 sugars, may improve the accuracy of detection of rare hydroxylated modified bases in RNA or DNA.

[0014] In embodiments, GTs for use in the method are characterized by a sequence selected from sequences identified as SEQ ID NOs: 1-8 or having at least 90% sequence identity to sequences selected from SEQ ID NOs: 9-17.

[0015] In embodiments, examples of saccharides for use in the method are described in Figure 8A.

[0016] In embodiments the one or more types of saccharides used in the method have a label where the label enables detection and / or enrichment of nucleic acids containing the labeled hydroxylated modified base.

[0017] In embodiments, the hydroxylated base is selected from the group consisting of the bases described in Figure 9, for example, a hydroxylated modified pyrimidine in a DNA or RNA.

[0018] In one embodiment, the addition of a saccharide or oligosaccharide to the hydroxylated modified base enables identification of hydroxylated modified base by sequencing, for example, whole genome sequencing of the nucleic acid.

[0019] In general, a method is provided that includes the steps of combining: a nucleic acid comprising a hydroxylated base; a glycosyltransferase (GT); and a GT- specific saccharide substrate excluding a single glucose or glucosamine to produce a reaction mixture; and incubating the reaction mixture to modify the hydroxylated base in the nucleic acid with the saccharide substrate.

[0020] In embodiments, the hydroxylated modified base may be selected from the group consisting of bases described in Figure 9; the GT may comprise a sequence selected from SEQ ID NOs: 1-8 and / or may comprise a sequence having at least 90% sequence identity (e.g., at least 95%, at least 97, at least 98% or at least 99% identical) to a sequence selected from SEQ ID NOs: 9-15.

[0021] Examples of the saccharide used in the method are provided in Figure 8A and may be labelled to identify and / or enrich for the hydroxylated modified nucleotides. Identification of saccharide or oligosaccharide labelled hydroxylated modified nucleotides may be achieved without labelling by sequencing of the nucleic acid for example by whole genome sequencing.

[0022] In general, a method is provided for detecting hydroxylated modified bases in a nucleic acid comprising different modifications; the method comprising: (a) reacting the hydroxylated modified bases in the nucleic acid with a hydroxylated base modification specific GT in the presence of a plurality of different saccharides wherein at least one saccharide that is a substrate for the specific GT is labelled; (b) detecting the presence of the hydroxylated modified bases by identifying the specific label reacted with hydroxylated modified base. In an embodiment, the specific label enables the detection of the modified hydroxylated base in the nucleic acid or the separation of the nucleic acid containing the hydroxylated modified base from other nucleic acids that do not contain the hydroxylated modified base.

[0023] In general, a method is provided that includes: combining an enzymatically inactive glycosyl transferase (GT) mutant having specificity for a selected hydroxylated modified base in a nucleic acid containing hydroxylated modified bases, and a nucleic acid for detecting GT- linked hydroxylated modified bases in the nucleic acid. The use of saccharides is optional in this method. The inactive GT may be labelled to distinguish different base modification by means of different labelled GTs reacted thereto. In addition to labelling or alternative to labelling, sequencing of the nucleic acid by for example whole genome mapping allows to detection of hydroxylated modified bases.

[0024] In general, a method is provided for identifying hypermodified bases in nucleic acids; that includes: reacting a nucleic acid with a glycosyl transferase (GT) in the presence of a labelled sugar; and sequencing the labelled nucleic acid to identify the hypermodified base.

[0025] In general, a method is provided for in vivo labelling of modified bases in genomic DNA, chromatin or in RNA in a cell; comprising, reacting the genomic DNA, chromatin or RNA with a labelled mutant glycosyl transferase that binds to the modified genomic DNA or RNA and can be detected by electron microscopy or by other means via the label.

[0026] In general, a variant glycosyltransferase having a mutation in the catalytically active site in Figure 3 for inactivating the saccharide transferase activity of the glycosyl transferase while allowing binding to a targeted hydroxylated modified base in the nucleic acid, and optionally altering the transferase so as to enhance the specificity of binding of the glycosyl transferase to the target hydroxylated modified base.

[0027] In general, a kit is provided that includes one or more active or inactivated glycosyl transferases (GT) wherein each GT comprises a sequence selected from any of SEQ ID NOS: 1-8 or comprises a sequence having at least 90% identity (e.g., at least 90% identity, at least 95% identity, at least 98% identity, or at least 99% identity), to a sequence selected from SEQ ID NOS: 9-17 and optionally is labelled; one or more saccharides wherein optionally at least one saccharide is labelled; and instructions for use according to any of the methods described above. The GTS and saccharides may be in the same or different storage containers. The active or inactivated GTs and / or saccharides may be lyophilized or in liquid solution.

[0028] BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figures 1A-1C show a process for identifying glycosyl transferases (GTs). Figure 1A shows the workflow used to identify GTs: (1) Assemblies were designed from a database of metavirome pathway coding regions and genes of interest were amplified by PCR using positionspecific primers with the assistance of the Opentrons liquid handling system. (2) Amplicons were assembled into final expression vectors by Golden Gate Assembly. (3) Assembled plasmids were transformed into NEB T7 Express E. coli and individual colonies were picked and grown in 96- well plates for (4) pathway expression by IPTG induction and parallel replica plates for cPCR verification. (5) Modified E. coli gDNA were isolated using a cpce automation instrument and subjected to (6) nucleoside digestion and LC-MS analysis. (7) Data from individual assemblies are processed and masses for new products are determined. Figure IB shows a comparison against UHPLC traces from in vivo glucosylated DNA using co-expressed T4-bGT in same plasmid system. Nucleoside compounds eluting within each UHPLC peak were verified by MS and were labeled. Figure 1C shows a heat map visualization of peak areas for parallel gDNA samples from E. coli expressing empty vector pSL126 (pET28a backbone), MT14+TET43 alone (pSL126), and MT14+TET43+T4 0GT (pSL126+T4 0GT). Five replicates are shown for each condition. Heat map scale by grayscale intensity is shown below and corresponds to total normalized area for peaks across retention time intervals. The elution of a non-specific product peak of unknown mass is indicated with an asterisk. Figure 2A-2B shows cytosine sugar hypermodifications installed by phage DNA GTs. In Figure 2A, the left panel shows the position of ten hypermodifying GTs as they occur in pSL126 transformed into E.coli adjacent to a methyl transferase (MT) and a TET enzyme (i.e., a Ten- eleven translocation dioxygenase). The presence and position of the 10 GTs are shown i.e. GT87II, GT87I and II, GT32, GT57, GT69, GT91 GT92, GT53, GT81 and GT94. The UHPLC-MS analysis on the right of each GT shows the glycosylation of 5hmC (product of MT and TET). Masses of compounds eluting within new LC peaks were used for identification of sugar modifications on 5mC. Figure 2B shows classes of sugar products resulting from the biosynthetic activities of phage GTs based on gene architecture observed in metagenomic contigs. Chemical structures of the sugar-modified nucleobase are shown. Different GT-containing contigs are identified below the sugar products. These are class A and class B GTs. For example, GT32 is shown here to synthesize a single sugar and a disaccharide appended to 5hmC. The GTs from contig 87 can append a monosaccharide or a trisaccharide to 5hmC.

[0030] Figure 3 shows similar residues in the active site (catalytic domain) of Class A GTs (GT-A). The aligned active site of GT-As are characterized by a binding pocket for the key structural and catalytic residues highlighted as sticks having R...DDD...EDY / F conserved active site motifs. Below the structure, are multiple sequence alignments of the GT-A members tested herein with their cellular structural homologs (tarP and JGT). Aligned regions containing the conserved active site motifs highlighted above are indicated with asterisk (also see Figure 4A and 4C).

[0031] Figure 4A-4C show sequences that characterize GTs. Figure 4A shows alignment and conserved sequences for 7 GT-A enzymes. The black boxes highlight conserved active site residues for the GT-A family. Figure 4B shows alignment and conserved sequences for 2 GT-B enzymes. The black boxes highlight catalytic aspartate residue for two GT-B sequences. Figure 4C provides a table that shows the amino acid position of selected conserved amino acids in GT- A and GT-B proteins.

[0032] Figure 5A-5B shows a sample of new sugar modifications installed by successive GT-A GTs. Figure 5A depicts LC-MS analysis of nucleosides from gDNA isolated from E. coli expressing GT87-II which installs the first sugar on 5hmC precursors. LC trace (left) and MS spectra in positive (right, top) and negative (right, bottom) modes are shown. Expression vector assemblies are shown next to their corresponding LC traces. Empty white arrows represent positions in the pSL126 vector assembled with non-coding pET28a control amplicons to control for enzyme position and expression plasmid size. MS spectra are shown for the new peak at retention time ~4.58 min showing a mass of 419u. Figure 5B depicts LC-MS analysis of nucleosides from gDNA isolated from E. coli expressing GT87-I and GT87-II showing installation of subsequent sugars on 5hmC precursors. LC trace (left) and MS spectra in positive (right, top) and negative (right, bottom) modes are shown. Expression vector assemblies are shown next to their corresponding LC traces. Empty white arrows represent positions in the pSL126 vector assembled with noncoding pET28a control amplicons to control for enzyme position and expression plasmid size. MS spectra are shown for the new peak at retention time ~6.77 min showing a mass of 743u. Figure 5C is a heat map visualization of UHPLC results for contig 87 enzymes expressed in E. coli. Individual grayscale boxes represent peaks from the UHPLC traces. Heat map scale by grayscale intensity is shown to the right and corresponds to total normalized area for peaks across retention time intervals. Results from four biological replicates are shown for each condition. The elution of non-specific and low-abundance product peaks of unknown masses are indicated with asterisks.

[0033] Figure 6 shows an analysis of new sugar modifications installed by GTs that are dependent on neighboring enzymes (e.g., epimerases, HAD phosphomutase, CTP transferases). Heat map visualization of total nucleosides from E. coli expressing GTs with their corresponding sugar epimerases. All expression constructs encode MT14 and TET43 for biosynthesis of ShmC precursor residues. Empty white arrows represent positions in the pSL126 vector assembled with non-coding pET28a control amplicons to control for enzyme position and expression plasmid size. Masses for new products are labeled where appropriate. Results from 3-4 biological replicates are shown for each condition. The elution of non-specific and low- abundance product peaks of unknown masses are indicated with asterisks.

[0034] Figure 7A-7C shows orthogonal labelling of hydroxylated modified nucleotides. Although not shown, SI and S2 where applicable can be saccharides, oligosaccharides or polysaccharides or a mixture thereof for the different targeted modified nucleotides. Figure 7A shows different saccharides (SI and S2) carrying the same label (LI) with GT-1 and GT-2, (B) different saccharides (SI and S2 carrying different labels (LI and L2) with GT-1 and GT-2, and (C) different inactivated GTs (GT-1 and GT-2) with different or the same labels (LI and L2, or LI and LI) for binding to different hydroxylated modified bases. Examples of probes include fluorophores for imaging or affinity tags for enrichment.

[0035] Figure 8A-8B shows examples of sugars for reacting with a hydroxylated modified base in a nucleic acid and examples of labels inserted onto these sugars. Figure 8A shows examples of hydroxylated modified bases. Figure 8B shows examples of bioconjugation methods for labeling sugar molecule. A bioorthogonal functional label "L" (such an azide, alkyl, or other) is incorporated on a saccharide to produce an unnatural sugar. This is in turn incorporated onto the base by GT. Subsequent bioorthogonal reactions conjugate the glycosylated based with a label bearing a complementary bioorthogonal group. The labels include fluorophores for imaging or affinity tags for enrichment (see for example, Bioconjugate Techniques 3rd Edition by Greg T. Hermanson, Cheng, et al., Annual Review of Analytical Chemistry; Glycan Labeling and Analysis in Cells and In Vivo).

[0036] Figure 9 shows examples of hydroxylated bases C, G, A, and U that contain modifications and may be detected by the family of GT-A or GT-B enzymes characterized herein, (see for example, Boccaletto, P. et al. Nucleic Acids Research, Volume 50, Issue DI, 7 January 2022, Pages D231-D235; Parker, M. J. et al. (2020) Comprehensive Natural Products III: Chemistry and Biology, vol. 5, pp. 465-488.

[0037] Figure 10 shows consensus sequences that characterize GT-A.

[0038] Figure 11 shows consensus sequences that characterize GT-B.

[0039] Figure 12 provides full length sequences for examples of seven GT-A and two GT-B glycosyl transferases.

[0040] Figure 13 shows in vitro activity assay design for GT94 and biosynthesis of 5-KdomC.

[0041] Reaction components include the enzymes responsible for CMP-Kdo (KdsB) and 5-KdomC (GT94) biosynthesis, and the substrate precursors for each, including CTP, Kdo, and 5-hmC-containing DNA from the bacteriophage T4 gt_ / '. Modified DNAs from the reactions are digested to free nucleosides and analyzed by UHPLC and LC-MS.

[0042] Figure 14A-14B shows expression and purification of the enzymes responsible for 5- KdomC synthesis. GT94 (14A) and KdsB from E. coli (EcKdsB, 14B) were expressed in T7 Express Competent E. coli cells and purified by immobilized metal affinity chromatography. Fractions used for in vitro activity design and testing, notably the Wash and Elution fractions for GT94 and the elution fractions shown for EcKdsB are shown separated within Coomassie-stained SDS- polyacrylamide gels.

[0043] Figure 15 shows LC-MS analysis of digested nucleosides from GT94 in vitro reactions. The LC traces for each condition are shown with key nucleoside products highlighted, including canonical deoxyribonucleosides (dC, dG, dT, and dA), ribonucleosides (rC, rG, and rA), 5-hmC and background 5-a-gmdC present in T4 gt / _gDNA, and 5-|3-gmdC resulting from labeling with recombinant T4 0GT (NEB M0357). A peak specifically appearing in the presence of EcKdsB and GT94 is indicated with an asterisk and eluted at a very similar retention time to previously characterized 5-KdomC (Pyle, Lund, et al. 2024. Cell Reports.). This product appears with a concurrent slight reduction in the 5-hmC peak, which suggests that 5-hmC is being converted to the new product, as would be expected for synthesis of 5-KdomC. The top two traces are included as controls to show 5-hmC-containing DNA without (first trace) and with (second trace) addition of T4 0GT to demonstrate that 5-hmC is a suitable substrate for DNA GTs and can be converted to a new peak for 5-P-gmdC.

[0044] DETAILED DESCRIPTION

[0045] Described herein are a new family of glycosyl transferases that are capable of adding one or a multiple of glycosyl groups to a hydroxyl group on a nucleotide in a nucleic acid. These enzymes may achieve the addition of a plurality of glycosyl groups either singly or in a mixture of a plurality of glycosyl transferases that can work synergistically to add the desired sugar side chain with reactive groups to the nucleotide hydroxyl group.

[0046] All publications cited herein are incorporated by reference for all purposes.

[0047] The term "hypermodification" as used herein means complex modifications of bases in nucleic acid or nucleoside triphosphates. This may include for example, one or more saccharides added to a hydroxylated base or having a reactive sulfur, nitrogen or phosphorous on a carbon in the base.

[0048] The term "orthogonal" as used herein means the addition of a plurality of saccharides on a base using a glycosyl transferase.

[0049] The term "Saccharide" as used herein means natural or synthetic sugars. These include polyhydroxylated aldehydes or ketones with the chemical formula Cn(H2O)mwhere n and m may be different. This term includes acetylated sugars such as GalNAc "Saccharide" as used herein is intended to include monosaccharides, disaccharides, oligosaccharides and polysaccharides unless explicitly stated otherwise.

[0050] Examples of saccharide nucleotide donors used by mammals for glycosyltransferases include uridine diphospho-D-glucose (UDP-Glu), uridine diphospho-D-galactose (UDP-Gal), uridine diphospho-W-acetyl-D-glucosamine (UDP-GIcNAc), uridine diphospho-Af-acetyl-D- galactosamine (UDP-GalNAc), uridine diphospho-D-xylose (UDP-Xyl), uridine diphospho-D- glucuronic acid (UDP-GIcA), guanosine diphospho-D-mannose (GDP-Man), guanosine diphospho-L-fucose, and CMP-sialic acid, uridine diphospho-D-galactofuranose (UDP-Gal / ), guanosine diphospho-L-rhamnose (GDP-Rha), cytidine monophospho-W-acetylneuraminic acid (CMP-Neu5Ac), cytidine monophospho-2-keto-3-deoxy-D-mannooctanoic acid (CMP-Kdo), dolichol phosphomannose (DP-Man) decaprenolphosphoarabinose (PP-Ara), dolichol-PP- GlcaMangGIcNac.

[0051] Examples of saccharides containing two and three sugar molecules are shown in Figure 2A-2B.

[0052] As used herein, the term "buffering agent", refers to an agent that allows a solution to resist changes in pH when acid or alkali is added to the solution. Examples of suitable non- naturally occurring buffering agents that may be used in the compositions, kits, and methods of the invention include, for example, Tris, HEPES, TAPS, MOPS, tricine, or MES. The term "non- naturally occurring" refers to a composition that does not exist in nature.

[0053] Any protein described herein may be non-naturally occurring, where the term "non- naturally occurring" refers to a protein that has an amino acid sequence and / or a post- translational modification pattern that is different from the protein in its natural state. For example, a non-naturally occurring protein may have one or more amino acid substitutions, deletions or insertions at the N-terminus, the C-terminus and / or between the N- and C-termini of the protein. A "non-naturally occurring" protein may have an amino acid sequence that is different from a naturally occurring amino acid sequence (i.e., having less than 100% sequence identity to the amino acid sequence of a naturally occurring protein) but that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% identical to the naturally occurring amino acid sequence. In certain cases, a non-naturally occurring protein may contain an N-terminal methionine or may lack one or more post-translational modifications (e.g., glycosylation, phosphorylation, etc.) if it is produced by a different (e.g., bacterial) cell. A "mutant" protein may have one or more amino acid substitutions relative to a wild-type protein and may include a "fusion" protein. The term "fusion protein" refers to a protein composed of a plurality of polypeptide components that are unjoined in their native state. Fusion proteins may be a combination of two, three or even four or more different proteins. The term polypeptide includes fusion proteins, including, but not limited to, a fusion of two or more heterologous amino acid sequences, a fusion of a polypeptide with: a heterologous targeting sequence, a linker, an epitope tag, a detectable fusion partner, such as a fluorescent protein, 0- galactosidase, luciferase, etc., and the like. A fusion protein may have one or more heterologous domains added to the N-terminus, C-terminus, and or the middle portion of the protein. If two parts of a fusion protein are "heterologous", they are not part of the same protein in its natural state.

[0054] In the context of a nucleic acid, the term "non-naturally occurring" refers to a nucleic acid that contains: a) a sequence of nucleotides that is different from a nucleic acid in its natural state (i.e., having less than 100% sequence identity to a naturally occurring nucleic acid sequence), b) one or more non-naturally occurring nucleotide monomers (which may result in a non-natural backbone or sugar that is not G, A, T or C) and / or c) may contain one or more other modifications (e.g., an added label or other moiety) to the 5'- end, the 3' end, and / or between the 5'- and 3' -ends of the nucleic acid.

[0055] In the context of a preparation, the term "non-naturally occurring" refers to: a) a combination of components that are not combined by nature, e.g., because they are at different locations, in different cells or different cell compartments; b) a combination of components that have relative concentrations that are not found in nature; c) a combination that lacks something that is usually associated with one of the components in nature; d) a combination that is in a form that is not found in nature, e.g., dried, freeze dried, crystalline, aqueous; e) one or more components that are immobilized to a support and / or e) a combination that contains a component that is not found in nature. For example, a preparation may contain a "non-naturally occurring" buffering agent (e.g., Tris, HEPES, TAPS, MOPS, tricine or MES), a detergent, a dye, a reaction enhancer or inhibitor, an oxidizing agent, a reducing agent, a solvent or a preservative that is not found in nature.

[0056] Provided herein is a method for making a nucleic acid that contains a saccharide- modified base. In some embodiments, the method may comprise contacting a nucleic acid with (i) a glycosyltransferase (GT) and (ii) an activated saccharide that is not UDP-glucose or a functionalized UDP-glucose, to add the saccharide of the activated saccharide onto a hydroxylated base in the nucleic acid. In this context, an "activated" saccharide is a saccharide that contains an energy releasing group, typically at the C-l position, that allows it to act as a glycosyl donor in a glycosylation reaction. Saccharides are typically nucleotide-linked, i.e., they are a 'nucleotide sugar', examples of which comprise a UDP, GDP, CDP and others. As such, in some embodiments, the activated saccharide may be an NDP sugar (which term includes UDP sugars, etc.) or a NMP sugar. CMP-Kdo is an example of the latter. Also, the term "not UDP- glucose or a functionalized UDP-glucose" is intended to exclude UDP-glucose to which an amine, ketone or azide group has been added, typically at the 6-position. In any embodiment, the GT used in the method is not T4 p-glucosyltransferase (BGT).

[0057] In some embodiments, the nucleic acid comprises a hydroxylated base. However, in some cases the hydroxylated base may be provided by one or more other enzymes that are also combined with the nucleic acid. Specifically, in some embodiments, the nucleic acid may be additionally contacted with one or more enzymes that modify a non-hydroxylated nucleotide (which may be a pyrimidine) to a hydroxylated nucleotide in the nucleic acid (e.g., a ten-eleven translocation (TET) methylcytosine dioxygenases and, optionally, a cytosine methyltransferase) that modifies the nucleic acid to produce the hydroxylated base.

[0058] In some embodiments, the method may comprise detecting the saccharide-modified base in the nucleic acid. This can be done by sequencing, mass spectrometry or any of a variety of detection methods. As is described elsewhere in this disclosure, the saccharide can be detectably labeled, thereby allowing the base to be detected optically (e.g., via fluorimetry).

[0059] In some embodiments, the nucleic acid may contacted with a plurality of different types of GT and a plurality of different activated saccharides.

[0060] In any embodiment, the GT may be a type A GT (i.e., a "GT-A") and may comprise the following motif: XIXTXXRXXXQXTXXNJXXXXXXXXXLVVXXX (SEQ ID NO: 1) or any of SEQ ID NOs: 2- 4. GT-A enzymes have a "DDD" catalytic motif in the catalytic fold. In some embodiments, the GT is a type A GT (GT-A) and may comprise an amino acid sequence having at least 90% identity (e.g., at least 95% identity) to any of SEQ ID NOs: 9-15.

[0061] In alternative embodiments, the GT may a type B GT (GT-B) and may comprises any of SEQ ID NOs: 5-8 or the following motif: NXXXXXYXEJXEDXRXXXXXXXDXXNR (SEQ ID NO: 18). In some cases, the GT is a type B GT (GT-B) and may comprise a sequence having at least 90% identity (e.g., at least 95% identity) to any of SEQ ID NOs: 16 or 17.

[0062] The activated saccharide may be selected from the saccharides shown in Figure 8A. An activating group (e.g., an NDP) can be added to any of these saccharides.

[0063] In any embodiment, the activated saccharide may comprise a pentose sugar, a heptose sugar, a non-glucose hexose sugar, or 6-amino-2,6-dideoxy-a-Kdo (Kdo), for example.

[0064] In any embodiment, the activated saccharide may comprise a label. In these embodiments, the label may be transferred onto the hydroxylated base during the reaction. The label, for example, may comprise an affinity tag (e.g., a biotin or a click-chemistry group such as an azide) or a detectable moiety such as a fluorescent label. In some embodiments, the activated saccharide may a fluorescent label that is transferred onto the hydroxylated base.

[0065] The nucleic acid may be a DNA or an RNA and, in any embodiment, the method may further comprise sequencing the nucleic acid saccharide-modified nucleic acid. This may be done by nanopore sequencing, for example.

[0066] In any embodiment, the saccharide-modified nucleic acid may be conjugated to an antibody via the saccharide-modified base.

[0067] In some embodiments, the activated saccharide is an activated monosaccharide and the saccharide-modified hydroxylated base may comprises an oligosaccharide that comprises the monosaccharide.

[0068] The method may be done in vitro or in a cell. In some embodiments, the reaction is done in vitro. In these embodiments, the method may comprise: (a) combining: (i) the glycosyltransferase (GT), (ii) the activated saccharide, and (iii) the nucleic acid to produce a reaction mix; and (b) incubating the reaction mix to add the saccharide of the activated saccharide onto a hydroxylated base in the nucleic acid. In these embodiments, the nucleic acid may comprise a hydroxylated base or the reaction mix further comprises one or more other enzymes that modify the nucleic acid to produce the hydroxylated base.

[0069] In other embodiments, the reaction is done in a cell. In these embodiments, the method may comprise expressing the glycosyltransferase (GT) in a cell that comprises the nucleic acid; wherein the activated saccharide is either synthesized by the cell or exogenously added to the cell. These embodiments may include isolate the nucleic acid from the cell and, in some cases, sequencing the nucleic acid.

[0070] In an embodiment, a label is provided to readily detect and / or separate a molecule from its environment.

[0071] The label L of the substrate can be chosen by those skilled in the art from art described dependent on the application for which the is intended. Labels are such that the labeled hydroxylated modified nucleotide carrying label L is easily detected or separated from its environment. Other labels considered are those which are capable of sensing and inducing changes in the environment of the labeled hydroxylated modified nucleotide or labels which aid in manipulating the hydroxylated modified nucleotide by the physical and / or chemical properties of the saccharide substrate specifically introduced into the nucleotide. In one embodiment, the label L is designed to covalently react with the hydroxyl group of hydroxylated modified nucleotides as illustrated in Figure 8B. Examples of hydroxyl-reactive groups are isocyanates, isothiocyanates, active esters (e.g., succinimidyl esters, sulfosuccinimidyl esters, tetrafluorophenyl esters, sulfotetrafluorophenyl esters, sulfodicholorphenol esters), carboxylic acids, acid halides, anhydrides, acyl azides, dichlorotriazines, and sulfonyl chlorides. Other hydroxyl-reactive groups are aldehydes, dialdehydes, ketones, vinyl sulfones, vinyl esters, alkyl halides, peroxides, and epoxides.

[0072] In embodiments, a label L is a substituent that is different from hydrogen or from standard functional groups, in particular different from hydrogen, hydroxy, amino, halogen, carboxylate, carboxamide, carboxylic ester, nitrile, cyanate, isocyanate, sulfonate, sulfonamide, sulfonic ester, aldehyde, ketone, ether, and thioether substituent. Examples of a label include a spectroscopic probe such as a fluorophore or a chromophore, a magnetic probe or a contrast reagent; a radioactively labeled molecule; a molecule which is one part of a specific binding pair which is capable of specifically binding to a partner; a molecule that is suspected to interact with other biomolecules; a library of molecules that are suspected to interact with other biomolecules; a molecule which is capable of crosslinking to other molecules; a molecule which is capable of generating hydroxyl radicals upon exposure to H2O2 and ascorbate, such as a tethered metal-chelate; a molecule which is capable of generating reactive radicals upon irradiation with light, such as malachite green; a molecule covalently attached to a solid support, where the support may be a glass slide, a microtiter plate or any polymer known to those proficient in the art; a nucleic acid or a derivative thereof capable of undergoing base-pairing with its complementary strand; a lipid or other hydrophobic molecule with membrane-inserting properties; a biomolecule with desirable enzymatic, chemical or physical properties; or a molecule possessing a combination of any of the properties listed above.

[0073] Further labels L are positively charged linear or branched polymers which are known to facilitate the transfer of attached molecules over the plasma membrane of living cells. This is of particular importance for substances which otherwise have a low cell membrane permeability or are in effect impermeable for the cell membrane of living cells. A non-cell permeable hydroxylated-modified nucleotide and / or saccharide substrate will become cell membrane permeable upon conjugation to such a group L. Such cell membrane transport enhancer groups L comprise, for example, a linear poly(arginine) of D- and / or L-arginine with 6 - 15 arginine residues, linear polymers of 6 - 15 subunits which each carry a guanidinium group, oligomers or short-length polymers of from 6 to up to 50 subunits, a portion of which have attached guanidinium groups, and / or parts of the sequence of the HIV-tat protein, in particular the subunit Tat49-Tat57 (RKKRRQRRR in the one letter amino acid code).

[0074] Labels L may be spectroscopic probes and molecules which are one part of a specific binding pair that is capable of specifically binding to a partner (so-called affinity labels). Also, labels L are molecules covalently attached to a solid support. Spectroscopic probes may be fluorophores. When the label L is a fluorophore, a chromophore, a magnetic label, a radioactive label or the like, detection is by standard means adapted to the label and whether the method is used in vitro or in vivo. Particular examples of labels L are also radioactively labeled hexosamines.

[0075] Particular labels include those in which L of one hydroxylated modified nucleotide (LI) is one member and L of a another hydroxylated modified nucleotide or a differently labeled nucleotide (L2) is the other member of two interacting spectroscopic probes LI / L2, wherein energy can be transferred nonradiatively between the donor and acceptor (another fluorophore or a quencher) when they are in close proximity (less than 10 nanometer distance) through either dynamic or static quenching. An example of such a pair of labels LI / L2 is a FRET (Forster (Fluorescence) resonance energy transfer).

[0076] Particular fluorophores considered are: Alexa Fluor dyes, including Alexa Fluor 350, 405, 430, 488, 514, 532, 546, 555, 568,594, 610, 633, 647, 660, 680, 700, 750, and 790 (Life Technologies Corporation, Carlsbad, CA 92008, USA); coumarin dyes, including 3-cyano-7- hydroxycoumarin, 6,8-difluoro-7-hydroxy-4-methylcoumarin, 7-amino-4-methylcoumarin, 7- ethoxy-4-trifluormethylcoumarin, 7-hydroxy-4-methylcoumarin, 7-hydroxycoumarin-3- carboxylic acid, 7-dimethylamino-coumarin-4-acetic acid, 7-amino-4-methyl-coumarin-3-acetic acid, and 7-diethylamino-coumarin-3-carboxylic acid (Life Technologies Corporation, Carlsbad, CA 92008, USA); BODIPY dyes, including BODIPY 493 / 503, FL, R6G, 530 / 550, TMR, 558 / 568, 564 / 570, 576 / 589, 581 / 591, TR, 630 / 650, and 650 / 655 (Life Technologies Corporation, Carlsbad, CA 92008, USA); Quantum Dots, including Qdot® 545 ITK™, Qdot® 565 ITK™, Qdot® 585 ITK™ , Qdot® 605 ITK™ , Qdot® 655 ITK™ , Qdot® 705 ITK™ , and Qdot® 800 ITK™ (Life Technologies Corporation, Carlsbad, CA 92008, USA); Oregon Green dyes, including Oregon Green 488, 488, and 514 (Life Technologies Corporation, Carlsbad, CA 92008, USA); LanthaScreen™ Tb Chelates (Life Technologies Corporation, Carlsbad, CA 92008, USA); Rhodamine 110, Rhodamine Green, Rhodamine Red, Texas Red-X, Cascade Blue, Pacific Blue, Marina Blue, Pacific Orange, Dapoxyl Sulfonyl Chloride, Dapoxyl Carboxylic Acid, 1-Pyrenebutanoic Acid, 1-Pyrenesulfonyl Chloride, 2- (2,3-Naphthalimino)ethyl Trifluoromethanesulfonate, 2-Dimethylaminonaphthalene-5-Sulfonyl Chloride, 2-Dimethylaminonaphthalene-6-Sulfonyl Chloride, 3-Amino-3-Deoxydigoxigenin Hemisuccinamide, 4-Sulfo-2,3,5,6-Tetrafluorophenol, 5-(and-6)-Carboxyfluorescein, 5-(and-6)- Carboxynaphthofluorescein, 5-(and-6)-Carboxyrhodamine 6G, 5-(and-6)-

[0077] Carboxytetramethylrhodamine, 5-(and-6)-Carboxy-X-Rhodamine, 5-Dimethylaminonaphthalene- 1-Sulfonyl Chloride (Dansyl Chloride), 6-((5-Dimethylaminonaphthalene-l- Sulfonyl)amino)Hexanoic Acid, Succinimidyl Ester (Dansyl-X, SE), Lissamine Rhodamine B Sulfonyl Chloride, Malachite Green isothiocyanate, NBD Chloride, 4-Chloro-7-Nitrobenz-2-Oxa- 1,3-Diazole (4-Chloro-7-Nitrobenzofurazan), NBD Fluoride; 4-Fluoro-7-Nitrobenzofurazan, 4- Fluoro-7-Nitrobenz-2-Oxa-l,3-Diazole, NBD-X, and PyMPO (Life Technologies Corporation, Carlsbad, CA 92008, USA); CyDyes, including Cy 3, Cy 3B, Cy 3.5, Cy 5, Cy 5.5, Cy 7 (GE Healthcare, Little Chalfont, United Kingdom); ATTO dyes, including ATTO 390, 425, 465, 488, 495, 520, 532, 550, 565, 590, 594, 610, 611X, 620, 633, 635, 637, 647, 647N, 655, 680, 700, 729, and 740 (ATTO-TEC GmbH, D57076 Siegen, Germany); DY dyes, including DY 350, 405, 415, 490, 495, 505, 530, 547, 548, 549, 550, 554, 555, 556, 560, 590, 591, 594, 605, 610, 615, 630, 631,

[0078] 632, 633, 634, 635, 636, 647, 648, 649, 650, 651, 652, 654, 675, 676, 677, 678, 679, 680, 681,

[0079] 682, 700, 701, 703, 704, 730, 731, 732, 734, 749, 750, 751, 752, 754, 776, 777, 778, 780, 781,

[0080] 782, 800, 831, 480XL, 481XL, 485XL, 510XL, 520XL, and 521XL (Dyomics, Jena, Germany); CF dyes, including CF 350, 405, 485, 488, 532, 543, 555, 568, 594, 620R, 633, 640R, 647, 660C, 660R, 680, 680R, 750, 770, and 790 (Biotium, Inc. Hayward, CA 94545, USA); CAL Fluor® dyes, including CAL Fluor Gold 540, Orange 560, Red 590, Red 610, Red 635 (Biosearch Technologies, Novato, CA 94949, USA); Quasar® dyes, including Quasar 570, 670, 705 (Biosearch Technologies, Novato, CA 94949, USA); Biosearch Blue and Pulsar 650 (Biosearch Technologies, Novato, CA 94949, USA); DyLight Fluor dyes, including DYLight 350, 405, 488, 550, 594, 633, 650, 680, 755, 800 (Thermo Fisher Scientific, Waltham, Massachusetts, USA); FluoProbes dyes, including FluoProbes 390, 488, 532, 547H, 594, 647H, 682, 752, 782 (Interchim, Inc., 03100 Montlu?on Cedex, France); SeTau dyes, including SeTau 380, 425, 405, 404,655, 665, and 647 (SETA BioMedicals, Urbana, IL 61801, USA); Square dyes, including Square 635, 660, and 685 (SETA BioMedicals, Urbana, IL 61801, USA); Seta dyes, including, Seta 470, 555, 632, 633, 646, 650, 660, 670, 680, and 750 (SETA BioMedicals, Urbana, IL 61801, USA); SQ-565 and SQ-780 (SETA BioMedicals, Urbana, IL 61801, USA); Chromeo™ dyes, including Chromeo 488, 494, 505, 546, and 642 (Active Motif, Carlsbad, CA 92008, USA); Abberior STAR dyes, including STAR 440SX, 470SX, 488, 512, 580, 635, and 635P (Abberior GmbH, D-37077 Gottingen, Germany); Abberior CAGE dyes, including CAGE 500, 532, 552, and 590 (Abberior GmbH, D-37077 Gottingen, Germany); Abberior FLIP 565 (Abberior GmbH, D-37077 Gottingen, Germany); IRDye Infrared Dye, including IRDye® 650, 680LT, 680RD, 700DX, 750, 800CW, and 800RS (LI-COR Biosciences, Lincoln, NE 68504, USA); Tide Fluor™ dyes, including TF1, TF2, TF3, TF4, TF5, TF6, TF7, and TF8 (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); iFluor™ dyes, including i Fluor™ 350, 405, 488, 514, 532, 555, 594, 610, 633, 647, 680, 700, 750, and 790 (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); mFluor™ dyes, including mFluor™ Blue 570, Green 620, Red 700, Red780, Violet 450, Violet 510, Violet 540, and Yellow 630 (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); trFluor™ Eu and trFluor™ Tb (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); HiLyte Fluor™ dyes, including HiLyte Fluor™ 405, 488, 555, 594, 647, 680, and 750 (AnaSpec, Inc., Fremont, CA 94555, USA); Terbium Cryptate and Europium Cryptate (Cisbio Bioassays, 30200 Codolet, France); and other nucleotide classical dyes, including FAM, TET, JOE, VIC, HEX, NED, TAMRA, ROX, Texas Red (Biosearch Technologies, Novato, CA 94949, USA).

[0081] Particular quenchers considered are: QSY 35, QSY 9, QSY 7, and QSY 21 (Life Technologies Corporation, Carlsbad, CA 92008, USA); Black Hole Quencher™, including BHQ-0 , BHQ-1 , BHQ-2, and BHQ-3 (Biosearch Technologies, Inc., Novato, CA 94949, USA); ATTO 540Q, ATTO 580Q, and ATTO 612Q (ATTO-TEC GmbH, D57076 Siegen, Germany); 4-dimethylamino- azobenzene-4'-sulfonyl derivatives (Dabsyl), 4-dimethylaminoazobenzene-4'- carbonyl derivatives (Dabcyl), DNP and DNP-X [6-(2,4-Dinitrophenyl)aminohexanoic acid] (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); DYQ quenchers, including DYQ 425, 505, 1, 2, 660, 661, 3, 700, 4 (Dyomics, Jena, Germany); IRDye® QC-1 (LI-COR Biosciences, Lincoln, NE 68504, USA); Tide Quencher™, including TQ1, TQ2, TQ3, TQ4, TQ5, TQ6, and TQ7 (AAT Bioquest, Inc., Sunnyvale, CA 94085, USA); QXL™ quenchers, including QXL™ 490, 570, 610, 670, and 680 (AnaSpec, Inc., Fremont, CA 94555, USA); BlackBerry® Quenchers, including BBQ-650 (Berry & Associates, Inc., Dexter, Ml 48130, USA).

[0082] Other labels L are nonfluorescent but form fluorescent conjugates stoichiometrically with hydroxyl groups. These reagents are particularly useful for detecting and quantitating hydroxylated modified nucleotides. Examples of such labels are fluorescamine, aromatic dialdehydes (e.g., o-phthaldialdehyde (OPA) and naphthalene-2,3-dicarboxaldehyde (NDA)), ATTO-TAG CBQCA, ATTO-TAG FQ, 7-nitrobenz-2-Oxa-l,3-Diazole (NBD) derivatives, dansyl chloride, 1-pyrenesulfonyl chloride, dapoxyl sulfonyl chloride, coumarins, pyrenes, and N- methylisatoic anhydride (Life Technologies Corporation, Carlsbad, CA 92008, USA).

[0083] Depending on the properties of the label L, the hydroxylated modified nucleotide may be bound to a solid support on reaction with the hydroxyl moiety. The label L may already be attached to a solid support when entering into reaction with a saccharide, or may subsequently, i.e. after the saccharide is transferred to the nucleotide, be used to attach the labeled saccharide-containing nucleotide to a solid support. The label L may be one member of a specific binding pair, the other member of which is attached or attachable to the solid support, either covalently or by any other means. A specific binding pair considered is e.g., biotin and avidin or streptavidin. Either member of the binding pair may be the label L of the substrate, the other being attached to the solid support. Examples of specific binding pairs allowing covalent binding to a solid support are e.g. SNAP-tag / AGT and benzylguanine derivatives (WO 02 / 083937) or pyrimidine derivatives (WO 2006 / 1 14409), CLIP-tag / AGT and benzylcytosine derivatives (W02008 / 012296), Halotag and chloroalkane derivatives (Los G.V. et aL, Methods Mol BioL 356, 195-208, 2007), serine-beta-lactamases and beta-lactam derivatives (WO 2004 / 072232). Further examples of specific binding pairs allowing covalent binding to a solid support are acyl carrier proteins and modifications thereof (binder proteins), which are coupled to a phosphopantheteine subunit from Coenzyme A (binder substrate) by a synthase protein (WO 2004 / 104588). Examples of labels allowing convenient binding to a solid support are e.g. chitin binding domain (CBD), maltose binding protein (MBP), glycoproteins, transglutaminases, dihydrofolate reductases, glutathione-S-transferase al (GST), FLAG tags, His-tags, or reactive substituents allowing chemoselective reaction between such substituent with a complementary functional group on the surface of the solid support. Examples of such pairs of reactive substituents and complementary functional group are e.g. amine and activated carboxy group forming an amide, azide and a propiolic acid derivative undergoing a 1 ,3-dipolar cycloaddition reaction, amine and another amine functional group reacting with an added bifunctional linker reagent of the type of activated bis- dicarboxylic acid derivative giving rise to two amide bonds, or other combinations known in the art. Examples of a convenient solid support are e.g. glass surfaces such as glass slides, microtiter plates, and suitable sensor elements, in particular functionalized polymers (e.g. in the form of beads), chemically modified oxidic surfaces, e.g. silicon dioxide, tantalum pentoxide or titanium dioxide, or also chemically modified metal surfaces, e.g. noble metal surfaces such as gold or silver surfaces. When the label L is capable of generating reactive radicals, such as hydroxyl radicals, upon exposure to an external stimulus, the generated radicals can then inactivate proteins that are in close proximity of the hydroxylated modified nucleotide, allowing studying the role of these proteins. Examples of such labels are tethered metal-chelate complexes that produce hydroxyl radicals upon exposure to H2O2 and ascorbate, and chromophores such as malachite green that produce hydroxyl radicals upon laser irradiation. The use of chromophores and lasers to generate hydroxyl radicals is also known in the art as chromophore assisted laser induced inactivation (CALI). Furthermore, proteins which are in close proximity of the hydroxylated modified nucleotide can be identified as such by either detecting fragments of that protein by a specific antibody, by the disappearance of those proteins on a high-resolution 2D- electrophoresis gels or by identification of the cleaved protein fragments via separation and sequencing techniques such as mass spectrometry or protein sequencing by N-terminal degradation.

[0084] When the label L is a molecule that can cross-link to other nucleic acids or proteins, e.g. a molecule containing functional groups such as maleimides, active esters, or azides and others known to those proficient in the art, contacting such labeled hydroxylated modified nucleotide that interact with other nucleic acids or proteins leads to the covalent cross-linking of the hydroxylated modified nucleotide with its interacting nucleic acid or protein via the label. This allows the identification of the nucleic acid or protein interacting with the hydroxylated modified nucleotide. In a special aspect of cross-linking, the label L is a molecule which enables photo-reactive (light-activated) chemical crosslinking. Labels L for photo cross-linking are e.g. benzophenones, phenyl azides, and diazirines.

[0085] Other labels L considered are for example fullerenes, boranes for neutron capture treatment, nucleotides or oligonucleotides, e.g. for self-addressing chips, peptide nucleic acids, and metal chelates, e.g. platinum chelates that bind specifically to DNA. A particular biomolecule with desirable enzymatic, chemical or physical properties is methotrexate. Methotrexate is a tight-binding inhibitor of the enzyme dihydrofolate reductase (DHFR). Compounds of formula (I) wherein L is methotrexate belong to the well known class of so-called "chemical inducers of dimerization" (CIDs).

[0086] The term "base" as used herein means primary pyrimidines (cytosine, thymine, and uracil) and purines (adenine and guanine). The term "base" may also include damaged bases such as 8-oxoguanine etc. and artificial bases. Modified bases include any of the primary bases that have an additional chemical adduct. For example, a base may contain a reactive hydroxyl group that enables a glycosyl transferase to add a saccharide to the base.

[0087] Many different modified bases have been found in the nucleic acids of phage genomes and have also been described in RNA and DNA of eukaryotic and prokaryotic cells. Some examples of modified bases are provided in Figure 9.

[0088] The term "glycosyltransferase" as used herein means any naturally occurring or synthetic enzyme that establishes glycosidic linkages and variants thereof. These catalyze the transfer of saccharide moieties from an activated nucleotide sugar (also known as the "glycosyl donor") to a nucleophilic glycosyl acceptor molecule, the nucleophile of which can be oxygen- carbon-, nitrogen-, or sulfur-based.

[0089] GTs, unless otherwise specified, are identified as GT-A or GT-B and have characteristic amino acid residues within the active site for binding to the reactive group in specific modified bases for adding a saccharide. These have been recognized in the consensus sequences (SEQ ID. NOs 1-8) in Figure 10.

[0090] The term "nucleic acid" as used herein means DNA, RNA or a chimera of DNA and RNA where the nucleic acid may be single stranded or double stranded.

[0091] Hydroxyl groups on nucleotides may occur in a stable form in vivo or may be an intermediate and transitory product of an oxidation pathway such as observed with methyl cytosine that is converted to hydroxymethyl cytosine and then to formyl cytosine and carboxy cytosine in the presence of a methyl dioxygenase (mYOX). Other non-hydroxylated modifications of nucleotides may similarly be oxygenated and thereby susceptible to interactions with glycosylase transferase. In human DNA, DNA methyl transferase (DNMT) is known to add a methyl group to cytosine that then becomes a target for enzyme conversion to hydroxymethylcytosine. The availability of glycosyl transferases described herein alone or in conjunction with any of dioxygenases and methyltransferases enable detecting, mapping, and / or enriching for hereto undiscovered modified bases in nucleic acids.

[0092] Detecting modified bases in nucleic acids, mapping modified bases by sequencing and / or enriching for nucleic acids containing selected modified bases by immobilizing on a substrate can be further facilitated by using labelled sugar substrates of the glycosyl transferases described herein as GT-A and GT-B.

[0093] The inventors have used a high throughput (HT) biosynthetic pathway reassembly platform for co-expression of metavirome enzymes and isolation of their hypermodified DNA products. Using this workflow, thousands of new biosynthetic gene clusters (BGC) were discovered. Evolutionary connections between the newly identified phage DNA hypermodification pathways and endogenous bacterial lipopolysaccharide, wall teichoic acid, glycolipid and small molecule biosynthesis pathways in phage host organisms were demonstrated. This identification of structural homologs of the phage DNA GTs from bacterial and eukaryotic species is supported by the apparently ancient fold of GTs in evolutionary space The "uncharacterized protein" (UniProt Q4E4F0, putative JGT) from T. cruzi as a top hit for the GT-A phage GTs suggests that divergent eukaryotic DNA GT homologs occur in the biosphere. A blastp search using the putative base J glucosyltransferase from T. cruzi revealed many homologs primarily belonging to members of the Trypanosoma, Leishmania, and several other trypanosome genera, but also cyanobacteria (Spirulina sp.) and predatory bacteria (Bdellovibrio sp.).

[0094] Sugar hypermodifications on phage DNA can block endonucleolytic cleavage by host restriction enzymes. However, these sugars are typically installed at high sequence coverage across the phage genome resulting from pre-replicative ThyS-based modulation at the level of the cellular nucleotide pool (16, 25). By contrast, post-replicative 5mYOX-based sugar hypermodifications are determined by the sequence context of phage-encoded C5-MT enzymes. Most 5mYOX-coupled phage MTs have a GpC sequence motif preference (18), limiting the sitespecific installation of sugars onto 5hmC by downstream GTs. It therefore seems unlikely that this post-replicative strategy would broadly provide effective genome-wide protection from host endonucleases targeting cytosine in other sequence contexts. Additional motif specificity may be provided by the GT binding of the DNA polymer. Moreover, glycans can shield the phage DNA from GpC-specific endonuclease activity. Using the cell-based pathway reassembly platform described above, the inventors observed that moderate amounts of complex sugar hypermodifications were tolerated by E. coli when installed on the bacterial chromosome (Figures 2A-2B, 5A-5C and 6). Sequencing of phage and host genomes during infection and baseresolution mapping of installed modifications can reveal the distribution of sugars and position relative to functional genomic regions.

[0095] Many of the GTs from all subclasses share structural homology with bacterial GTs, specifically those involved in biosynthesis of cell surface molecules ( / .e., LPS, WTA) and glycolipids (Figures 3 and 4A-4C). Kdo, Hep, and additional LPS core sugars are receptors for phage infection and multiple host enzymes in LPS biosynthesis were identified in screens for phage receptors. The repurposing of a family of GTs by phage suggests a number of potential explanations. These include the following:

[0096] (a) Sequestration of LPS sugars through conjugation to DNA bases by phage- encoded GTs could indirectly affect receptor maturation or abundance at the host cell surface or serve to redirect sugar binding proteins from the host cell. Thus, phage or host DNA used as a template by the phage-encoded GTs would effectively act as a sugar and / or sugar-binding protein sponge during infection.

[0097] (b) Since Hep and Kdo are specialized and essential sugars for LPS maturation (83), sequestration of the activated sugar-nucleotide substrates could suppress receptor presence at the cell surface preventing attachment by extracellular phage in the surrounding environment as a new mechanism of superinfection exclusion (84).

[0098] (c) In addition to their role in DNA hypermodification, the LPS pathway-related phage GTs retain contacts with membrane structures or with membrane-associated LPS precursors. Directing phage DNA to membranous compartments within the cell could facilitate genome replication (85, 86) or other critical stages of the infection cycle, as observed recently by Armbruster and colleagues (87).

[0099] (d) Host sugar metabolites may be concentrated at the point of viral DNA replication and therefore used opportunistically in DNA hypermodification systems.

[0100] Kits

[0101] A kit is also provided. In some embodiments, the kit may comprise: (i) a glycosyltransferase (GT), (ii) an activated saccharide that is not UDP-glucose or a functionalized UDP-glucose; and (iii) a reaction buffer for the GT, where the reaction buffer may contain at least one non-natural buffering agent and provide optimal conditions for activity of the GT.

[0102] In some embodiments, the GT may a type A GT (GT-A) and may comprises the following motif: XIXTXXRXXXQXTXXNJXXXXXXXXXLVVXXX (SEQ ID NO: 1) or any of SEQ ID NOs: 2-4. In some embodiments, the GT is a type A GT (GT-A) and may comprise an amino acid sequence having at least 90% identity to any of SEQ ID NOs: 9-15.

[0103] In some embodiments, the GT may be a type B GT (GT-B) and may comprise any of SEQ ID NOs: 5-8 or the following motif: NXXXXXYXEJXEDXRXXXXXXXDXXNR (SEQ ID NO: 18). In some embodiments, the GT may be a type B GT (GT-B) and may comprise a sequence having at least 90% identity to any of SEQ ID NOs: 16 or 17. In some embodiments, the activated saccharide is selected from the saccharides shown in Figure 8A. For example, the activated saccharide may comprise a pentose sugar, a nonglucose hexose sugar, a heptose sugar, or 6-amino-2,6-dideoxy-a-Kdo (Kdo). The activated saccharide may comprise an affinity tag or detectable moiety, e.g., a fluorescent label.

[0104] The components of the kit may be in separate contains although, in some embodiments, certain components may be in the same container.

[0105] EMBODIMENTS

[0106] Embodiment 1. A method for detecting a hydroxylated modified base in a nucleic acid, comprising: (a) combining a glycosyltransferases (GT) and a population of saccharides with a nucleic acid; (b) permitting the GT to add an oligosaccharide or monosaccharide onto a hydroxylated base in the nucleic acid; and (c) detecting the oligosaccharide modified hydroxylated base in the nucleic acid.

[0107] Embodiment 2. The method of embodiment 1, wherein (b) further comprises: combining a plurality of different types of GT for attaching the saccharide onto the hydroxylated modified base and wherein the population of saccharide contains a plurality of different saccharides.

[0108] Embodiment 3. The method of any of embodiment 1 or 2, wherein at least one of the one or more GTs comprises a sequence corresponding to any of SEQ ID NOs: 1-4.

[0109] Embodiment 4. The method of any of embodiments 1 or 2, wherein at least one of the one or more GTs comprises a sequence corresponding to any of SEQ. ID NOs: 5-8.

[0110] Embodiment 5. The method of any of embodiments 1-3, wherein at least one of the one or more GTs comprises a sequence having at least 90% identity to any of SEQ ID NOs: 9-15.

[0111] Embodiment 6. The method of any of embodiments 1, 2 or 4, wherein at least one of the one or more GTs comprises a sequence having at least 90% identity to any of SEQ ID NOs: 16 or 17.

[0112] Embodiment 7. The method of any of embodiments 1-6, wherein the population of saccharides in (b) is selected from the group consisting of saccharides in Figure 8A.

[0113] Embodiment 8. The method of any of embodiments 1- 7, wherein the population of saccharides in (b) are labelled.

[0114] Embodiment 9. The method of embodiment 8, wherein the label enables the hydroxy modified base in the nucleic acid to be identified and / or enriched. Embodiment 10. The method of any of embodiments 1- 9, wherein the hydroxylated modified base is selected from the group consisting of the bases described in Figure 9.

[0115] Embodiment 11. The method of any of embodiments 1-10, wherein the hydroxylated modified base is a hydroxylated modified pyrimidine or purine.

[0116] Embodiment 12. The method of any of embodiments 1-11 wherein the nucleic acid is a DNA or an RNA.

[0117] Embodiment 13. The method of any of embodiments 1-12, further comprising: sequencing the nucleic acid.

[0118] Embodiment 14. The method of embodiment 13, wherein the step of sequencing further comprises whole genome sequencing.

[0119] Embodiment 15. A method, comprising:

[0120] (a) combining: i. a nucleic acid comprising a hydroxylated modified base; ii. a glycosyltransferase (GT); and iii. a GT specific saccharide substrate excluding a single glucose or glucosamine, to produce a reaction mixture; and

[0121] (b) incubating the reaction mixture to modify the hydroxylated modified base in the nucleic acid with the saccharide substrate.

[0122] Embodiment 16. The method of embodiment 15, wherein the hydroxylated modified base is selected from the group consisting of bases described in Figure 9.

[0123] Embodiment 17. The method of embodiment 15 or 16, wherein the GT comprises a sequence selected from SEQ ID NOs:l -8.

[0124] Embodiment 18. The method of any of embodiments 15-17, wherein the GT comprises a sequence having at least 90% sequence identity to a sequence selected from SEQ ID NOS: 9-15.

[0125] Embodiment 19. The method of embodiment 17, wherein the GT comprises a sequencing having at least 90% sequence identity to a sequence selected from SEQ ID NO: 16 or 17.

[0126] Embodiment 20. The method of any of embodiments 15-19, wherein the saccharide substrate is selected from the group consisting of saccharides provided in Figure 8A and wherein the saccharides are optionally labelled.

[0127] Embodiment 21. The method of any of embodiments 15-20, further comprising (c) whole genome sequencing the nucleic acid from (b) to detect and / or map the modified hydroxylated base in the nucleic acid. Embodiment 22. A method for detecting hydroxylated modified bases in a nucleic acid comprising different modifications; the method comprising:

[0128] (a) reacting the hydroxylated modified bases in the nucleic acid with a hydroxylated base modification specific glycosyl transferase in the presence of a plurality of different saccharides wherein at least one saccharide that is a substrate for the specific GT is labelled;

[0129] (b) detecting the presence of the hydroxylated modified bases by identifying the specific label reacted with hydroxylated modified base.

[0130] Embodiment 23. The method of embodiment 22, wherein the specific label enables the detection of the modified hydroxylated base in the nucleic acid or the separation of the nucleic acid containing the hydroxylated modified base from other nucleic acids that do not contain the hydroxylated modified base.

[0131] Embodiment 24. A method, comprising:

[0132] (a) Combining: (i) an enzymatically inactive glycosyl transferase (GT) mutant having specificity for a selected modified base in a nucleic acid containing modified bases, and (ii) a nucleic acid; and

[0133] (b) detecting GT- linked modified bases in the nucleic acid.

[0134] Embodiment 25. The method of embodiment 24, wherein (a) further comprises (iii) a saccharide.

[0135] Embodiment 26. The method of embodiment 24 or 25, wherein the GT is labelled.

[0136] Embodiment 27. The method of any of embodiments 24-26, further comprising distinguishing different hydroxylated modified bases by means of differently labelled GTs reacted thereto.

[0137] Embodiment 28. The method of any of embodiments 24-27, further comprising step (c) sequencing the nucleic acid to detect or map the modified bases in the nucleic acid.

[0138] Embodiment 29. A method for identifying hypermodified bases in nucleic acids; comprising: reacting a nucleic acid with a glycosyl transferase in the presence of a labelled sugar; and sequencing the labelled nucleic acid to identify the hypermodified base.

[0139] Embodiment 30. A method for in vivo labelling of modified bases in genomic DNA, chromatin or in RNA in a cell; comprising, reacting the genomic DNA, chromatin, or RNA with a labelled mutant glycosyl transferase that binds to the modified genomic DNA or RNA and can be detected by electron microscopy or by other means via the label. Embodiment 31. A variant glycosyltransferase having a mutation in the catalytically active site in Figure 3 for inactivating the saccharide transferase activity of the glycosyl transferase while allowing binding to a targeted hydroxylated modified base in the nucleic acid, and optionally altering the transferase to enhance the specificity of binding of the glycosyl transferase to the target hydroxylated modified base.

[0140] Embodiment 32. A variant glycosyl transferase having a mutation in the catalytically active site in Figure 3 for enhancing the transfer of a saccharide onto a modified base in a nucleic acid.

[0141] Embodiment 33. A kit comprising: one or more glycosyl transferases (GTs) wherein each GT comprises a sequence selected from any of SEQ ID NOS: 1-8 or comprises a sequence having at least 90% identity to a sequence selected from any of SEQ ID NOS: 9-17; one or more saccharides wherein at least one saccharide is labelled and instructions for use according to any of the methods embodimented herein.

[0142] Embodiment 34. The kit of embodiment 33, wherein the one or more GTs and the one or more saccharide are in different storage containers in liquid form or lyophilized.

[0143] Embodiment 35. The kit of embodiment 34, wherein the one or more GTs and the one or more saccharides are in the same storage container in liquid form or lyophilized.

[0144] EXAMPLES

[0145] All enzymes, buffers, plasmids, and strains were obtained from New England Biolabs (NEB, Ipswich, MA) unless otherwise noted. PCR setup, amplicon normalization, and Golden Gate Assembly (GGA) reaction setup steps were performed using custom python scripts operated on an OT-2 (Opentrons, Brooklyn, NY). Python packages used for data analysis include SciPy, Pandas, Seaborn, and Matplotlib.

[0146] Example 1: Detection of biosynthetic gene clusters using a high throughput in vivo discovery platform

[0147] A high-throughput platform for parallel analysis of biosynthetic gene clusters depicted in Figure 1A was used. As is shown, a co-expression platform for all enzymes in a pathway of interest was developed with each enzyme under separate transcription control by T7 RNA polymerase. Genes of interest with T7 promoter and terminator regions were amplified using position-specific primers, allowing for one-step assembly into a single plasmid expression vector using Golden Gate Assembly (New England Biolabs). Position of the enzymes and correct construct assembly was verified by colony PCR or whole-plasmid sequencing. Assembled constructs were transformed into E. coli with lac inducible expression of T7 RNA polymerase to coordinate transcription for the assembled genes of interest.

[0148] To detect modifications directly on nucleic acid bases, genomic DNA from E. coli was isolated and the DNA digested to individual nucleosides. The nucleoside pool was analyzed using reverse-phase liquid chromatography mass spectrometry (LC-MS), as described previously (E. J. Burke, et al.,. Proceedings of the National Academy of Sciences 118, e2026742118 (2021)). I products from total digested E. coli gDNA typically eluted within a 1.5-to-7-minute LC retention time window, among peaks for canonical dC, dG, dT, and dA nucleosides. Analytes were continuously monitored by mass spectrometry in positive and negative mode during LC separation and adduct ion masses were assigned to I peaks. For most sugar-modified nucleosides, [M+H]+, [M+Na]+, and [M+K]+adduct ions in positive mode and [M-H]' and [M+CHJCOJ-H]' adduct ions in negative mode were detected.

[0149] As a starting point for screening the activities of hypermodification enzymes on 5hmC, Methyl transferase 14 (MT14) and TET43 were coexpressed in each of the construct designs to maximize 5hmC substrate and increase the likelihood of modification detection. As a control, p- glucosyltransferase from T4 phage (T4-0GT) was coexpressed with MT and Tet in the same plasmid, which resulted in the observed additional peak for p-glucosyl-5-methylcytosine (5- GlcpmC). Samples were processed in 96-well plate format and the data visualized as heat maps displaying relative abundance of canonical and modified nucleosides (Figure 1A-1C and Figure 5A-5C). In this way, the high throughput discovery platform was validated as a robust system for parallel assessment of I hypermodification pathway enzyme functions.

[0150] Example 2: Phage glycosyltransferases install I sugars on cytosine

[0151] Systematic in vivo reconstitution of GT-containing phage pathways resulted in the production of original and highly diverse sugar-modified cytosine bases (Figure 2A-2B). The sugars appended to cytosine appeared to include six-, seven-, or eight-carbon sugars, mono-, di, or trisaccharide glycans, and ulosonic (e.g., Kdo) and uronic (HexA) sugar acids (Figure 2B).

[0152] For example, GTI alone from contig 87 produced a single hexose-5mC compound (419 amu) derived from 5hmC while expression of GTI alone did not result in any detectable modifications to 5mC and 5hmC (Figure 5A-5C). Co-expression of GTI and GTII produced a new hypermodified cytosine product trihexose-5mC (743 amu). The addition of a trihexose in contrast to a dihexose was unexpected. The natural products from contig 87 enzymes demonstrated that some glycosyltransferases work synergistically to add a plurality of saccharides to a hydroxyl group on a modified nucleotide.

[0153] Another example are the GTs from contigs 32 and 69, which are processive hexose (Hex)- and / V-acetylhexosamine (HexNAc)-transferases. In some cases, an epimerase partner modified the stereochemistry of the pool of available sugar nucleotides to enable the addition of a sugar onto the hydroxylated base. For example, GT32 was able to add 5- hexosemethylcytosine (5-HexmC, 419 amu) and 5-di-HexmC (581 amu) from 5hmC precursor bases preferably in conjunction with epimerase 32. GT69 produced 5-N- acetylhexosaminemethylcytosine (5-HexNAcmC, 460 amu) and 5-di-HexNAcmC (663 amu) preferably with epimerase 69. GT 92 added 5-di-HexNAcmC and GT 91 could produce a hexuronic acid (HexA)-modified cytosine (433 amu, Figure 2A).

[0154] GT- 53 and 94 yielded sugar-modified products corresponding to keto-deoxyoctonic (Kdo) sugar acid appendage.

[0155] Collectively, the screening of twelve candidate phage revealed the biosynthesis of ten new cytosine sugar hypermodifications stemming from the E. coli metabolomic substrate pool.

[0156] Example 3: Two families of Glycosyltransferase GT-A and GT-B were identifed

[0157] GT-A and GT-B were identified using the AlphaFold deep-learning network and revealed two general structural classes with GT-A and GT-B-structures, in agreement with our phylogenetic analyses. Using these predictive models, Protein Data Bank (PDB) and AlphaFold Database (AFDB) Proteome were probed for structural homologs using Foldseek (M. van Kempen, et al. Nat Biotechnol (2023).

[0158] Structure-guided searches of the AFDB using the phage GT-A structure predictions as query yielded consistent strong predicted matches (E-values = 1012to 1013) with an uncharacterized protein from Trypanosoma cruzi. Examination of proteins in UniProt (The UniProt Consortium, UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res 51, D523-D531 (2023).) with 90% sequence identity to this hit revealed that this enzyme is a close homolog of the base J-associated glucosyltransferase (JGT), providing further supporting evidence that the GT-A fold phage GTs directly transfer sugar moieties to DNA (Figure 3A,). A MSA of the phage enzymes with the T. cruzi JGT revealed limited (~9 % identity) sequence similarity, suggesting an ancient structure-function relationship between the phage and eukaryotic DNA GTs. The same structure-guided search against experimental structural models in the Protein Data Bank (PDB100) returned top hits (E-value = 10’5to 10'6) against tarP (Figure 3A,), a prophage-encoded GIcNAc-transferase for alternative biosynthesis of wall teichoic acid (WTA) in Staphylococcus aureus (D. Gerlach, et al., Nature 563, 705-709 (2018).). TarP-mediated 0-O-GlcNAcylation of ribitol-phosphate (RboP) subunits in WTA alters susceptibility of the bacterial cell to phage infection and modulates 5. aureus immune evasion in mammalian hosts (Gerlach et al; D. Gerlach, et al., Front Microbiol 13 (2022).46-48); X. Guoqing, et a!., J Bacterial 193, 4006-4009 (2011). Additionally, the chemistry of tarP-synthesized sugar conjugates matches those of the bacteriophage-encoded GTs - both transfer sugars to nucleophilic hydroxyl groups on RboP polymers or 5hmC acceptor substrates in DNA (Gerlach et al.2018).

[0159] Using GT-B fold enzymes to search the database, all three enzymes- GT53, GT66, GT94 shared some predicted structural folds, in agreement with the ability of these enzymes to act on DNA substrates. The predicted GT catalytic residues are observed within a charged groove. (Figure 3B). Two loop regions unique to GT94 and GT53 flanked the predicted sugar binding pocket adjacent to the putative catalytic aspartate, for providing flexibility and increased solvent-exposed cavity space for accommodating more bulky sugar moieties (e.g., Kdo, heptose). Other matches included BshA (N-acetyl-a-D-glucosaminyl L-malate synthase), MshA (D-inositol 3-phosphate glycosyltransferase), from multiple species and PgIH (GalNAc-a-(l,4)- GalNAc-a-(l, 3)-diNAcBac-PP-undecaprenol a-l,4-N-acetyl-D-galactosaminyltransferase) from Campylobacter jejuni (E-values = 10'5to 10-7,). Comparison with the top hits from AFDB Proteome revealed a pattern of structural similarities with bacterial and eukaryotic GT-B foldcontaining enzymes that function primarily on glycolipid or small molecule precursors with free hydroxyl groups (Figure 4A-4C).

[0160] Example 4: Enhanced discovery of DNA sugar hypermodification pathways in the microbial metagenome

[0161] Using the GT-A and GT-B fold-containing phage enzymes (Figure 3) hidden Markov model (HMM) profiles were generated to re-probe metavirome and metagenome databases for related enzymes and their neighboring gene clusters. The initial search was expanded to include diverse microbial metagenomes by ecosystem curated within the Joint Genome Institute (JGI) Integrated Microbial Genomes and Microbiomes (IMG / M) (Chen, et al., Nucleic Acids Res 51, D723-D732 (2023)), the gut phage database (GPD) . Cell 184, 1098-1109.e9 (2021)), RNA viruses in the metatranscriptome (Neri, et al., Cell 185, 4023-4037. el8 (2022).), and all viral, archaeal, and enterobacterial reference genomes available in NCBI. All hits from our GT HMM searches were retrieved with their surrounding gene neighborhoods, including ten predicted coding sequence regions upstream and downstream of the target gene. All coding regions in the retrieved contigs were annotated using HMMs from Pfam (Mistry, et al., Nucleic Acids Res 49, D412-D419 (2021). ), TIGRFAM REBASE (New England Biolabs), Carbohydrate Active Enzymes (CAZy) (Drula, et al., Nucleic Acids Res 50, D571-D577 (2022). ), and the prokaryotic antiviral defense locator (PADLOC) (Payne, et al., Nucleic Acids Res 50, W541-W550 (2022).) and DefenseFinder (Tesson, et al., Nat Commun 13, 2561 (2022)protein domain databases. This allowed was the basis for predictions about the roles of newly identified GTs in the context of their adjacent biosynthetic pathways.

[0162] Approximately 25,000 new GT-containing contigs were identified from all databases probed. The vast majority of these contigs (~95%) were identified within the latest release of the IMG / VR database confirming their viral origins. The remaining ~5% of assemblies were retrieved from IMG / M environment-specific metagenomes, the GPD, or from bacterial, archaeal, or phage genomes in NCBI. GT-A fold enzymes were more abundant in the metagenome (n=19,990 contigs, ~82% of total). Pathways containing a GT-B were associated with thymidylate synthase (ThyS) about 28% of the time whereas GT-A were found in similar amounts adjacent to ThyS and 5mYOX genes.

[0163] While not wishing to be limited by a hypothesis, present findings suggest that phage- encoded GTs have cellular homologs that have been repurposed throughout evolution to glycosylate diverse substrate acceptor molecules. Moreover, the functional and structural overlaps between phage GTs and those of endogenous eukaryotic and bacterial organisms suggest an evolutionary adaptation that has been shaped by the dynamic interplay between phage genome defense and bacterial immune systems.

[0164] Example 5. In vivo reconstitution of GT activity

[0165] GT94 is a metavirome DNA hypermodifying enzyme that was previously shown to synthesize 5-KdomC when expressed in E. coli (Pyle, Lund, et al. 2024). We attempted to reconstitute this activity in vitro using purified GT94, 5hmC-containing DNA, and the additional enzymatic and small molecule prerequisites for activated Kdo biosynthesis. The active form of Kdo in cells is present as the nucleotide CMP-Kdo, which is synthesized by the enzyme CMP-Kdo synthase (KdsB) in E. coli (EcKdsB). EcKdsB synthesizes CMP-Kdo from CTP and free Kdo sugar. CMP-Kdo is then available for subsequent reactions by cellular or phage-encoded GTs.

[0166] A gene fragment encoding GT94 with codon optimization for E. coli expression (Twist Bioscience, San Francisco, CA) was cloned into a pET28-based expression vector with a C- terminal histidine (6xHis) affinity tag. EcKdsB was amplified from purified T7 Express Competent E. coli (NEB C2566) genomic DNA and cloned in a similar manner with an N-terminal 6xHis tag. The expression constructs were transformed into NEB C2566 cells and protein expression was induced with the addition of IPTG (200 pM) to cells in mid-log growth. Induced cells were grown overnight (16-18 hours) at 18C in a shaker incubator set to 220 rpm.

[0167] EcKdsB was purified by IMAC on a HisTrap column (Cytiva Life Sciences, Marlborough, MA) and eluted with a gradient of imidazole from 0-500 mM. EcKdsB was dialyzed into storage buffer containing 20 mM potassium phosphate, 300 mM NaCI, 1 mM DTT, 20% glycerol, pH 7.0. Aliquots were flash frozen and stored at -80C.

[0168] Cell pellets expressing GT94 were lysed in E. coli lysis reagent (NEB P8116) containing T4 lysozyme (NEB P8115, 1 ug per mL of cells) at room temperature for 10 minutes. Clarified lysate was separated over equilibrated Ni-NTA resin (NEB S1428) in a column by gravity flow to bind 6xHis-tagged GT94. Settled resin was washed twice with ten column volumes of high-detergent wash buffer (50 mM Tris-HCI, 200 mM NaCI, 50 mM imidazole, 2.5 mM beta-mercaptoethanol, 50% glycerol, 1% Triton X-100, 1% DDM, pH 7.5) by gravity flow and GT94 was eluted in four 1 mL fractions containing elution buffer (50 mM Tris-HCI, 200 mM NaCI, 500 mM imidazole, 2.5 mM beta-mercaptoethanol, 50% glycerol, 1% Triton X-100, 1% DDM, pH 7.5). Eluted fractions were stored at -20C until further use in reconstituted in vitro reactions.

[0169] In vitro 5-KdomC synthesis reactions were assembled in a buffer containing 20 mM Trisacetate, 10 mM magnesium acetate, 50 mM potassium acetate, 100 pg / ml recombinant Albumin, pH 7.9, supplemented with CTP (1 mM), Kdo (1 mM), EcKdsB (100 pM), and GT94 (~0.25 pM, pooled wash fractions), where appropriate, to final reaction volumes of 100 pL. Each reaction included approximately 1 pg of fragmented T4 gf / _phage gDNA, which contains nearly 100% 5-hmC (equivalent to ~5.5 uM 5-hmC final, per reaction). Reactions were incubated at 37C for 4 hours and quenched with the addition of 8000U Proteinase K (NEB P8107) per reaction for 30 minutes at 37C. Fragmented gDNAs were purified using the NEB Monarch PCR & DNA Cleanup Kit (NEB T1030) and the total recovered gDNA was subjected to total nucleoside digest using the NEB Nucleoside Digestion Mix (NEB M0649) overnight at 37C. Digested nucleosides for each reaction were analyzed by LC-MS.

Claims

CLAIMSWhat is claimed:

1. A method for making a nucleic acid that contains a saccharide-modified base, comprising: contacting a nucleic acid with(i) a glycosyltransferase (GT) and(ii) an activated saccharide that is not UDP-glucose or a functionalized UDP- glucose, to add the saccharide of the activated saccharide onto a hydroxylated base in the nucleic acid.

2. The method of claim 1 or 2, wherein the nucleic acid comprises a hydroxylated base.

3. The method of any prior claim wherein the method further comprises contacting the nucleic acid with (iii) a one or more enzymes that modify the nucleic acid to produce the hydroxylated base.

4. The method of any prior claim, wherein further comprising: detecting the saccharide-modified base in the nucleic acid.

5. The method of any prior claim, wherein the nucleic acid is contacted with a plurality of different types of GT and a plurality of different activated saccharides.

6. The method of any prior claim, wherein the GT is a type A GT (GT-A) and comprises the following motif: XIXTXXRXXXQXTXXNJXXXXXXXXXLVVXXX (SEQ ID NO: 1) or any of SEQ ID NOs: 2-4.

7. The method of any prior claim, wherein the GT is a type A GT (GT-A) and comprises an amino acid sequence having at least 90% identity to any of SEQ ID NOs: 9-15.

8. The method of any of claims 1-5, wherein the GT is a type B GT (GT-B) and comprises any of SEQ ID NOs: 5-8 or the following motif: NXXXXXYXEJXEDXRXXXXXXXDXXNR (SEQ ID NO:18).

9. The method of any of claims 1-5 or 8, wherein the GT is a type B GT (GT-B) and comprises a sequence having at least 90% identity to any of SEQ ID NOs: 16 or 17.

10. The method of any of any prior claim, wherein the activated saccharide is selected from the saccharides shown in Figure 8A.

11. The method of any prior claim, wherein the activated saccharide comprises a pentose sugar, a non-glucose hexose sugar, a heptose sugar, or 6-amino-2,6-dideoxy-a-Kdo (Kdo).

12. The method of any prior claim, wherein activated saccharide further comprises a labelthat is transferred onto the hydroxylated base.

13. The method of claim 12, wherein the label comprises an affinity tag or a detectable moiety.

14. The method of claim 13, wherein activated saccharide comprises a fluorescent label that is transferred onto the hydroxylated base.

15. The method of any of any prior claim, wherein the nucleic acid is a DNA or an RNA.

16. The method of any prior claim, further comprising sequencing the saccharide-modified nucleic acid.

17. The method of claim 16, wherein the sequencing is done by nanopore sequencing.

18. The method of any prior claim, wherein the activated saccharide is an activated monosaccharide and the saccharide-modified hydroxylated base comprises an oligosaccharide that comprises the monosaccharide.

19. The method of any prior claim, wherein the reaction is done in vitro and the method comprises:(a) combining:(i) the glycosyltransferase (GT),(ii) the activated saccharide, and(iii) the nucleic acid; to produce a reaction mix; and(b) incubating the reaction mix to add the saccharide of the activated saccharide onto a hydroxylated base in the nucleic acid, wherein the nucleic acid comprises a hydroxylated base or the reaction mix further comprises one or more enzymes that modify the nucleic acid to produce the hydroxylated base.

20. The method of any of claims 1-18, wherein the reaction is done in a cell and the method comprises: expressing the glycosyltransferase (GT) in a cell that comprises the nucleic acid; wherein the activated saccharide is either synthesized by the cell or exogenously added to the cell, and the cell optionally contains one or more other enzymes that are capable of hydroxylating nucleic acid in the cell.

21. The method of claim 20, further comprising isolating the nucleic acid from the cell.

20. A kit comprising:(i) a glycosyltransferase (GT),(ii) an activated saccharide that is not UDP-glucose or a functionalized UDP-glucose; and(iii) a reaction buffer for the GT.

21. The kit of claim 20, wherein the GT is a type A GT (GT-A) and comprises the following motif: XIXTXXRXXXQXTXXNJXXXXXXXXXLVVXXX (SEQ ID NO: 1) or any of SEQ ID NOs: 2-4.

22. The kit of claim 20 or 21, wherein the GT is a type A GT (GT-A) and comprises an amino acid sequence having at least 90% identity to any of SEQ ID NOs: 9-15.

23. The kit of claim 20, wherein the GT is a type B GT (GT-B) and comprises any of SEQ ID NOs: 5-8 or the following motif: NXXXXXYXEJXEDXRXXXXXXXDXXNR (SEQ ID NO: 18).

24. The kit of claim 20 or 23, wherein the GT is a type B GT (GT-B) and comprises a sequence having at least 90% identity to any of SEQ ID NOs: 16 or 17.

25. The kit of any of claims 20-24, wherein the activated saccharide is selected from the saccharides shown in Figure 8A.

26. The kit of any of claims 20-25, wherein the activated saccharide comprises a pentose sugar, a non-glucose hexose sugar, heptose sugar, or 6-amino-2,6-dideoxy-a-Kdo (Kdo).

27. The kit of any of claims 20-24, wherein the activated saccharide further comprises an affinity tag or detectable moiety.

28. The kit of claim 27 , wherein activated saccharide comprises a fluorescent label.

29. The method of any prior method claim, further comprising conjugating the saccharide- modified nucleic acid to an antibody.

Citation Information

Patent Citations

  • Compositions and methods for analyzing modified nucleotides

    US10227646B2

  • Compositions and methods for analyzing modified nucleotides

    US10260088B2

  • Methods using o<6>-alkylguanine-DNA alkyltransferases

    WO2002083937A2

  • Covalent tethering of functional groups to proteins

    WO2004072232A2

  • Methods for protein labeling based on acyl carrier protein

    WO2004104588A1