Compositions and methods related to brig1 DNA glycosylase
Patent Information
- Application Number
- PCT/US2025/024986
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2025-04-16
- Publication Date
- 2026-01-15
AI Technical Summary
Existing methods for detecting and locating 5-hydroxymethylcytosine (5hmC) modifications in DNA are cumbersome and require numerous reagents and steps.
The use of Brigl DNA glycosylase, in combination with T4 alpha-glucosyltransferase (a-GT), to excise alpha-glucosyl-hydroxymethylcytosine nucleobases and generate abasic sites, facilitating improved detection and location of 5hmC modifications.
Enhances the efficiency and simplicity of detecting and locating 5hmC modifications by generating abasic sites, allowing for more effective analysis of these nucleobases.
Abstract
Description
[0001] COMPOSITIONS AND METHODS RELATED TO Brigl DNA GLYCOSYLASE
[0002] CROSS-REFERENCE TO RELATED APPLICATONS
[0003] This application claims priority to U.S. Provisional Application No. 63 / 634,593, filed April 16, 2024, the entire disclosure of which is hereby incorporated herein by reference.
[0004] SEQUENCE LISTING
[0005] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on April 15, 2025, is named “076091_00179_ST26.xml”, and is 22,085 bytes in size.
[0006] RELATED INFORMATION
[0007] Double stranded DNA, such as genomic DNA, often includes 5- hydroxymetnyl-cytosine (“5hmC”) nucleobases. Such modifications are common in the human genome, especially within neurons, as well as other cell types. While there are existing methods for detecting and locating 5hmC modifications, they are cumbersome and require numerous reagents and steps. There is thus a need for improved compositions and methods for analysis of 5hmC modifications, The disclosure is pertinent to this need.
[0008] BRIEF SUMMARY
[0009] The disclosure provides compositions and methods for improved detection and location of 5hmC modifications. In an example, the disclosure provides Brigl , a DNA glycosylase that excises alpha-glucosyl-hydroxymethylcytosine nucleobases to generate abasic sites. The Brigl is used in combination with a T4 alphaglucosyltransferase (a-GT). The described Brigl and a-GT proteins can be used in a variety of methods for determining the presence and / or location of 5-methylcytosine (5mC) by adapting and modifying existing sequencing reagents and methods by incorporating use of the Brigl protein in conjunction with a-GT. Fusion proteins comprising the Brigl and other proteins are included. Kits comprising the described protein are also included. BRIEF DESCRIPTION OF FIGURES
[0010] Figure 1. Screening of an environmental DNA (eDNA) library uncovers an unknown gene that protects E. coli against phage T4 infection by preventing viral replication, a, Ten-fold serial dilutions of phage T4 on lawns of E. coli EC100 that harbor the pWEB-TNC cosmid carrying different genes present in a three-gene operon isolated after the screening of an eDNA library with this phage. Gene c encodes Brig 1. Data are representative of three independent experiments, b, Quantitative PCR analysis of T4 DNA through amplification of the gp43 gene. Viral DNA was extracted from infected E. coli EC100 cells carrying pWEB-TNC or pBrigl at 2, 4 and 8 minutes after the addition of phage at an MOI of 1. Fold-change values were calculated relative to the pWEB-TNC 2-minute time point. Mean values are reported for three independent experiments. Error bars report the standard error of the mean (SEM). c, Normalized next-generation sequencing reads of T4 DNA, mapped to the viral genome. DNA for sequencing was obtained 8 minutes after infection of E. coli EC100 carrying pWEB-TNC or pBrig 1 at an MOI of 5.
[0011] Figure 2. The antiviral DNA glycosylase Brigl targets alpha-glucosyl- hydroxymethylcytosine nucleobases to restrict T4 viral replication, a, Ten-fold serial dilutions of different T4 phage stocks on lawns of E. coli EC100, each lawn carrying cosmid pWEB-TNC or pBrigl , and plasmids pEmpty, p(a-g ) or p(b-gt). Plaque images of one representative experiment from three independent experiments are shown, b, Normalized next-generation sequencing reads of T4 escaped DNA, mapped to the viral genome. DNA for sequencing was obtained 8 minutes after infection of E. coli EC100 carrying pWEB-TNC or pBrigl at an MOI of 5. c, Ten-fold serial dilutions of different T4 phage stocks on lawns of E. coli EC100, each lawn carrying cosmid pWEB-TNC or pBrigl , and plasmid pEmpty, p(a-gt) or p(gp42). Plaque images of one representative experiment from three independent experiments are shown.
[0012] Figure 3. Brigl excises alpha-glucosyl-hydroxymethylcytosine nucleobases. a, AlphaFold2 structure of Brigl , colored by position (N-terminal, blue; to C-terminal, red), with cavities shown in translucent grey. Inset; zoomed in view of putative glycosylase pocket, with uracil (pink-purple sticks) from PDB 4ZBY modeled into it, and showing the amino acid residues that could participate in substrate binding, b, Polyacrylamide gel electrophoresis of 60-nt single-stranded oligonucleotides containing a single modified base, incubated with either hSMUGI or Brig 1 at 37°C overnight, and then treated with an aldehyde-reactive Alexa 488 fluorescent probe to label and detect abasic sites. The same gel was stained with ethidium bromide to detect ssDNA. II, uracil; a-GIc-hmC, alpha-glucosyl- hydroxymethylcytosine; p-GIc-hmC, beta-glucosyl-hydroxymethylcytosine. Data are representative of one experiment, c, Urea-PAGE of dsDNA substrates harboring a- Glc-hmC in the top strand (see Figure 10a) treated either with hSMUGI or Brigl at 37°C overnight, with and without heating in the presence of NaOH for 30 minutes. Gels were stained with ethidium bromide. L, ssDNA size ladder. Data are representative of two independent experiments, d, Native PAGE of dsDNA substrates harboring a-GIc-hmC in both strands (see Figure 10a) treated either with hSMUGI or Brigl at 37°C overnight, with and without heating in the presence of NaOH for 30 minutes. Gels were stained with ethidium bromide. L, dsDNA size ladder. Data are representative of one experiment, e, Agarose gel electrophoresis of T4, T4 escaped or pWEB-TNC DNA (500 ng) treated with increasing concentrations (2, 20, 200, 400, 800 nM) of Brigl or with 10 units of NEB hSMUGI (Sm) for 30 minutes at 37°C. Data are representative of three independent experiments.
[0013] Figure 4. Brigl provides immunity against diverse phages that carry alpha-glucosyl-hmC nucleobases. a, Ten-fold serial dilutions of common coliphages spotted on lawns of E. coli EC100 carrying cosmid pWEB-TNC or pBrigl . T6 Dba-gt lacks the glucosyltransferase that adds the second glucose to alpha- glucosyl-hydroxymethylcytosine nucleobases in phage T6. Data are representative of three independent experiments, b, Agarose gel electrophoresis of 50 ng of T2, T4, T6, T6 Dba-gt or pWEB-TNC DNA (“p”) treated with 10 units of hSMUGI or 100 nM of Brigl for 30 minutes at 37°C. L, DNA size ladder. Data are representative of two independent experiments.
[0014] Figure 5. Brigl homologs provide anti-phage defense, a, Maximum likelihood tree of 42 Brigl homologs (noted by their NCBI protein accession numbers) found in different phyla: Firmicutes (purple background), Proteobacteria (green), Actinobacteria (pink), Cyanobacteria (light blue) and Planctomycetota (orange). Brigl homologs that provide anti-phage defense against T4 and T6 are indicated in red. Grey squares indicate the presence of putative defense genes in the immediate vicinity (within 10 genes upstream and downstream); brown squares, BREX / Pgl genes; yellow squares, transposases; orange squares, ADP-ribosyl glycohydrolases, b, Ten-fold serial dilutions of phage T4 or T6 on lawns of E. coli EC100 carrying the pAM39 vector to express the Brigl homologs shown in red in a using an arabinose-inducible promoter, in the presence (+ Ara) or absence (- Ara) of the inducer. Plaque images of one representative experiment from three independent experiments are shown.
[0015] Figure 6. Selection of soil metagenomic DNA fragments that provide immunity against phage T4 in E. coli. a, Schematic showing the experimental setup for the screening of soil DNA (environmental DNA, eDNA) libraries for the presence of clones resistant to T4 infection. eDNA is extracted from soil samples and cleaved into large fragments of ~40 kb that are cloned into E. coli EC100. The library is infected with phage T4 in top agar plates to isolate surviving colonies. The cosmid from the T4-resistant colonies is extracted and further analyzed to identify and confirm immunity genes, b, Ten-fold serial dilutions of v / ror T4 phage on lawns of E. coli EC100 cells generated from one of 16 colonies sampled randomly from the eDNA library population that survived T4 infection. Top row is a representative result for a surviving colony that did not contain a bona fide immunity gene in the eDNA (4 / 16 false positives). Bottom row is a representative result for a surviving colony that harbored a cosmid carrying immunity genes (12 / 16 true positives), c, Genes present within the 34.5 kb soil metagenomic DNA fragment that provided T4 immunity. Subcloned regions (Fragments C and D) are indicated. Transposases and defense genes are shown in yellow and grey, respectively, d, Ten-fold serial dilutions of phage T4 on lawns of E. coli EC100 carrying pWEB-TNC or cosmids containing Fragments C or D. Data are representative of two independent experiments, e, Genes present within Fragment D, showing the subclones generated for further analysis, Fragments D1-3; hyp, hypothetical protein, f, Ten-fold serial dilutions of phage T4 on lawns of E. coli EC100 carrying pWEB-TNC or cosmids containing Fragments D1-3. Data are representative of two independent experiments, g, Tenfold serial dilutions of phage T4 on lawns of bacteria expressing Brigl using an arabinose-inducible promoter, in the presence (+ Arabinose) or absence (- Arabinose) of the inducer. Data are representative of three independent experiments.
[0016] Figure 7. Brigl inhibits phage T4 DNA replication. a, Enumeration of T4 PFU in supernatants of infected cultures at different times after phage addition. Cultures of E. coli EC100 carrying pWEB-TNC or pBrigl were infected at MOI 0.01 . Addition of T4 phage to media lacking bacteria was used as a control, b, Quantitative PCR analysis of T4 DNA through amplification of the gp34 gene. Viral DNA was extracted from infected E. coli EC100 cells carrying pWEB- TNC or pBrigl at 2, 4 and 8 minutes after the addition of phage at an MOI of 1 . Foldchange values were calculated relative to the pWEB-TNC 2-minute time point. Mean values are reported for three independent experiments. Error bars report the standard error of the mean (SEM). c, Quantitative PCR results from Figure 1 b for DNA extracted from infected E. coli EC100 cells carrying pBrigl ; plotted using a Iog2 scale. Mean values are reported for three independent experiments. Error bars report the standard error of the mean (SEM). d, Same as c but for results shown in b. e, Same as b but using DNA extracted at 8- and 20-minutes post infection, amplifying the gp43 gene. Fold-change values were calculated relative to the pWEB- TNC 8-minute time point, f, Same as b but using DNA extracted at 8- and 20- minutes post infection. Fold-change values were calculated relative to the pWEB- TNC 8-minute post-infection time point. For all quantitative PCR graphs, mean values are reported for three independent experiments, with error bars reporting the standard error of the mean (SEM). g, Schematic of the cytosine modification pathway in phage T4, including a model for Brigl immunity through excision of alpha-glucosyl-hydroxymethylcytosine nucleobases from the viral genome to generate abasic sites. Gp42, dCMP hydroxymethylase; Gp56, dCTPase; a-GT, alpha-glucosyltransferase; b-GT, beta-glucosyltransferase. h, Model showing the activity of T4-encoded enzymes that affect DNA containing cytosine bases. Ale, dC- specific premature transcriptional terminator; DenB, dC-specific ssDNA endonuclease, i, Efficiency of plaquing of T4 and T4 phage passaged through E. coli EC100 / p(b-gt), [T4(+b-GT)], on lawns of E. coli EC100 carrying pBrigl . Mean are reported for three independent experiments. Error bars report the standard error of the mean (SEM); p-value is reported for an unpaired two-sided Student’s t-test. Figure 8. Oligonucleotides and nucleobases used in Brigl DNA glycosylase assays, a, AlphaFold2 structure of Brigl , colored by pLDDT (a per- residue model confidence score). Red to blue spectrum represents high to low confidence of secondary structure prediction, b, Crystal structure of a family 4 uracil DNA glycosylase from Sulfolobus tokodaii (PDB 4ZBY, a close structural homolog of Brigl) showing the uracil substrate (pink-purple sticks) in the glycosylase pocket and the Fe-S cluster (yellow and orange spheres). The structure is colored by position (N-terminal, blue; to C-terminal, red), with cavities shown in translucent grey. Inset, zoomed in view of the glycosylase pocket, showing the uracil substrate and the amino acid residues that participate in uracil binding, c, Sequence of the 60-nt single-stranded oligonucleotide used for testing the DNA glycosylase activity of Brigl . The nucleobase in red was synthesized as 5-hydroxymethylcytosine (hmC) and subsequently alpha- or beta-glucosylated. The Mfel restriction site is underlined. Mfel digestion was used to confirm glucosylation after annealing a complementary bottom strand oligonucleotide to create a dsDNA substrate for the Mfel enzyme, d, Chemical structures of alpha- and beta-glucosyl-hydroxymethylcytosine nucleobases. e, Mfel digestion of dsDNA substrates generated by annealing the oligonucleotide shown in panel c containing a 5-hydroxymethylcytosine nucleobase with a complementary bottom strand oligonucleotide. In each case, prior to Mfel digestion of the dsDNA oligonucleotides, the ssDNA (“ssDNA”) or dsDNA (“dsDNA”) was either untreated or treated with a low (+) or high (++) concentration of T4 alpha- or T4 beta-glucosyltransferase (a-GT or b-GT, respectively). Data are representative of one experiment, f, Sequence of the 60-nt oligonucleotide used for testing the DNA glycosylase activity of Brigl or hSMUGI in Fig. 3b and panel h below. The red X was replaced by the nucleobases shown in g. g, Chemical structures of different nucleobases used to test the specificity of Brigl . h, Polyacrylamide gel electrophoresis of 60-nt single-stranded oligonucleotides containing a single modified base, incubated with either hSMUGI or Brigl at 37°C overnight, and then treated with NaOH and heat for 30 minutes prior to gel electrophoresis. Gels were stained with ethidium bromide to detect ssDNA. U, uracil; T, thymine, mC, 5- methylcytosine; hmC, 5-hydroxymethylcytosine; 2-aminoA, 2-aminoadenine. Data are representative of two independent experiments, i, PAGE of the reaction products that resulted from the sequential treatment of the oligonucleotide shown in panel c, containing an alpha-glucosyl-hydroxymethylcytosine nucleobase, with Brigl at 37°C overnight followed by 50 units of endonuclease IV at 37°C for 4 hours or NaOH for
[0017] 30 minutes at 90°C. Gels were stained with ethidium bromide. L, ssDNA size ladder. Data are representative of two independent experiments.
[0018] The sequence in Fig. 8c is TAGACATTGCCCTCGAGGTA(7?mC)AATTGATCCGATTTCGACCTCAAACCTAGA CGAATTCCG (SEQ ID NOV). In this sequence hmC = 5-hydroxymethylcytosine. The Mfel restriction site is in italics.
[0019] The sequence in Fig. 8f is TAGACATTGCCCTCGAGGTAXCATGGATCCGATTTCGACCTCAAACCTAGACGA ATTCCG (SEQ ID NO:8), where X is any of II, T, 5-methylcytosine (mC) or 2- aminoadenine (2-aminodA).
[0020] Figure 9. Mass spectrometry of a Brigl -generated abasic site, a, Theoretical average masses, in daltons (Da), of the different oligonucleotides used for mass spectrometry. indicates an abasic site;hmC, 5-hydroxymethylcytosine;glc-hmC, alpha-glucosyl-hydroxymethylcytosine. c, Deconvoluted zero-charge mass spectra from high resolution mass spectrometry of the oligonucleotides shown on the top of each column incubated without any enzyme, with hSMUGI or with Brigl at 37°C overnight. Major peaks are indicated in red.
[0021] The sequences on Fig. 9C from left to right are TCGAGGTAUCATGGATCC (SEQ ID NO:9), TCGAGGTA(hmC)AATTGATCC (SEQ ID NQ:10), and TCGAGGTA(glc-hmC)AATTGATCC (SEQ ID NO: 1 1).
[0022] Figure 10. Brigl activity on dsDNA oligonucleotides. a, Sequence of the dsDNA oligonucleotide used for testing the DNA glycosylase activity of Brigl . The base marked as “X” in red was synthesized as 5- hydroxymethylcytosine (hmC) and subsequently alpha-glucosylated (a-GIc-hmC); “Y” was synthesized as cytosine (C) or as hmC that was subsequently alpha- glucosylated (a-GIc-hmC). The Mfel restriction site is underlined. Mfel digestion was used to confirm glucosylation. b, Sequence of the dsDNA oligonucleotide used for testing the uracil DNA glycosylase activity of hSMUGI and Brigl . The base marked as “Z” in red was synthesized as uracil (U) or thymine (T). c, PAGE of the reaction products obtained after Mfel digestion (40 units) of dsDNA oligonucleotides (500 ng) shown in panel a (the nucleotide modifications at X and Y positions are indicated). Data are representative for one experiment, d, PAGE of the reaction products obtained after treatment of dsDNA oligonucleotides shown in panels a (X in top strand is alpha-glucosyl-hmC; Y in bottom strand is cytosine) and b (Z in bottom strand is thymine), with hSMUGI or Brigi at 37°C overnight, with or without additional heating with NaOH for 30 minutes. L, dsDNA size ladder. Data are representative of two independent experiments, e, Urea-PAGE of the reaction products obtained after treatment of a dsDNA oligonucleotide shown in panel b (Z in bottom strand is thymine), with hSMUGI or Brigi at 37°C overnight, with or without additional heating with NaOH for 30 minutes. L, ssDNA size ladder. Data are representative of two independent experiments, f, PAGE of the reaction products obtained after treatment of a dsDNA oligonucleotide shown in panel a (X in top strand is hmC; Y in bottom strand is cytosine), with hSMUGI or Brigi at 37°C overnight, with or without additional heating with NaOH for 30 minutes. L, dsDNA size ladder. Data are representative of one experiment, g, PAGE of the reaction products obtained after treatment of ssDNA oligonucleotides shown in panel a (X and Y are alpha-glucosyl-hmC), with hSMUGI or Brigi at 37°C overnight, with or without additional heating with NaOH for 30 minutes. L, ssDNA size ladder. Data are representative of one experiment, h, PAGE of the reaction products obtained after treatment of a dsDNA oligonucleotide shown in panel b (Z in bottom strand is uracil), with hSMUGI or Brigi at 37°C overnight, with or without additional heating with NaOH for 30 minutes. L, dsDNA size ladder. Data are representative of one experiment, i, same as panel f, but X and Y are hmC. Data are representative of one experiment.
[0023] The sequences in Fig. 10a are TAGACATTGCCCTCGAGGTAXAATTGATCCGATTTCGACCTCAAACCTAGACGA ATTCCG (SEQ ID NO:12). X = 5-hydroxymethylcytosine (hmC) or alpha-glucosyl-5- hydroxymethylcytosine; and CGGAATTCGTCTAGGTTTGAGGTCGAAATCGGATYAATTGTACCTCGAGGGCAA TGTCTA (SEQ ID NO: 13). Y = C or hmC or alpha-glucosyl-5-hydroxymethylcytosine.
[0024] The sequences in Fig. 10b are: TAGACATTGCCCTCGAGGTAUAATTGATCCGATTTCGACCTCAAACCTAGACGA
[0025] ATTCCG (SEQ ID NO: 14) and
[0026] CGGAATTCGTCTAGGTTTGAGGTCGAAATCGGAZCAATTATACCTCGAGGGCAA TGTCTA (SEQ ID NO:15). Z = U or T
[0027] Figure 11. Generation of abasic sites in T4 genomic DNA by Brigl . a, Agarose gel electrophoresis of 125 ng of phage T4 genomic DNA treated with 10 units of hSMUGI (“Sm”) or increasing concentrations of Brigl (2, 20, 200, 400 nM) for 30 minutes at 37°C. Electrophoresis was performed under 40 V for 3 hours at 4°C. L, DNA size ladder. Data are representative for one experiment, b, Same gel shown in panel a run under higher voltage (85, 150 and 200 V for additional 25, 8 and 8 minutes, respectively) at room temperature. Data are representative for one experiment, c, Agarose gel electrophoresis of T4, T4 escaped or pWEB-TNC DNA (500 ng) treated with increasing concentrations (2, 20, 200, 400, 800 nM) of Brigl or with 10 units of hSMUGI (“Sm”) for 30 minutes at 37°C and followed by heat treatment (20 minutes at 65°C) prior to electrophoresis. L, DNA size ladder. Data are representative of three independent experiments, d, Agarose gel electrophoresis of T4 or pWEB-TNC DNA (500 ng) treated with increasing concentrations (20, 200 nM) of Brigl with or without treatment with SDS prior to electrophoresis. L, DNA size ladder. Data are representative of one experiment, e, Agarose gel electrophoresis of DNA from T4 or from T4 phage passaged through E. coli EC100 / p(b-gt), which overexpresses T4 beta-glucosyltransferase to increase the frequency of beta- glucosyl-hydroxymethylcytosine modifications within the T4 genome [T4(+b-GT)], after treatment with increasing concentrations (2, 20, 200, 400 nM) of Brigl for 30 minutes at 37°C. L, DNA size ladder. Data are representative of one experiment.
[0028] Figure 12. Brigl amino acid residues involved in base excision, a, Tenfold serial dilutions of phage T4 on lawns of E. coli EC100 carrying pWEB-TNC, pBrig 1 or pBrig 1 harboring substitutions in the amino acids thought to participate in catalysis shown in Fig. 3a. Plaque images of one representative experiment from three independent experiments are shown, b, PAGE of the reaction products obtained after treatment of the ssDNA oligonucleotide shown in Fig. 8c, harboring alpha-glucosyl-hmC, with hSMUGI (“Sm”; 5 units), Brigl or the Brigl Y121A, E147A mutant (50, 100, 200, 400, 800 and 1600 nM) at 37°C for 30 minutes, followed by heating with NaOH for 30 minutes. L, ssDNA size ladder. Data are representative of two independent experiments, c, Agarose gel electrophoresis of T4 or pWEB-TNC DNA (500 ng) treated with increasing concentrations (50, 500 nM) of Brigl , the Y121A, E147A mutant or with 10 units of hSMUGI (“Sm”) for 30 minutes at 37°C. L, DNA size ladder. Data are representative of one experiment.
[0029] Figure 13. E. coli AP endonucleases and DNA repair proteins are not specifically required for Brigl immunity, a, Ten-fold serial dilutions of phage T4 on lawns of different E. coli BW25113 mutants with deletions of genes involved in base excision repair, carrying the pAM38 vector to express Brigl using an arabinose-inducible promoter, in the presence (+ Arabinose) or absence (- Arabinose) of the inducer, b, Ten-fold serial dilutions of phage T4 on lawns of E. coli EC100, each lawn carrying cosmid pWEB-TNC or pBrigl , and plasmids pAM38(xf / ?A) or pAM38(nfo), which express the E. coli AP endonucleases XthA and Nfo, respectively, using an arabinose-inducible promoter, in the presence (+ Arabinose) or absence (- Arabinose) of the inducer, c, Same as a but using E. coli BW25113 mutants with deletions of genes involved in RecABCD or RecJQ DNA repair pathways. Plaque images of one representative experiment from two independent experiments are shown in a-c.
[0030] Figure 14. Brigl immunity against different T-even coliphages. a, Schematic of the cytosine modification pathway in phage T6, including a model for Brigl immunity in which the DNA glycosylase recognizes and excises alpha- glucosyl-hydroxymethylcytosine nucleobases from the viral genome, before the addition of the second glucosyl group to generate gentiobiosyl- hydroxymethylcytosine. T6 enzymes: Gp42, dCMP hydroxymethylase; Gp56, dCTPase; a-GT, alpha-glucosyltransferase; ba-GT, beta-alpha glucosyltransferase, b, Ten-fold serial dilutions of T6, T6 escaped and T6 escaper2 phages on lawns of E. coli EC100, each lawn carrying cosmid pWEB-TNC or pBrigl , and plasmids pEmpty or p(a-gt). c, Agarose gel electrophoresis of T2, T4, T6 or pWEB-TNC DNA (125 ng) treated with decreasing concentrations (200, 20 nM) of Brigl for 30 minutes at 37°C. L, DNA size ladder. Data are representative of one experiment, d, Efficiency of plaquing of T6 and T6 Aba-gt phages on lawns of E. coli EC100 carrying pBrigl . Mean values are reported for three independent experiments. Error bars report the standard error of the mean (SEM). N.D., no plaques detected; dotted line, limit of detection, e, Ten-fold serial dilutions of phages from the BASEL collection Bas35-47 spotted on lawns of E. coli EC100 carrying pWEB-TNC or pBrigl . Plaque images of one representative experiment from three independent experiments are shown.
[0031] Figure 15. Brigl homologs, a, Gene neighborhoods of Brigl homologs found in putative anti-phage defense islands. Brigl homologs and other DNA glycosylases are shown in magenta, Pgl / BREX genes in brown, and other putative defense genes in grey, b, Gene neighborhood of the Brigl homolog from Nocardioides zhouii showing ADP-ribosyl glycohydrolase, transposase and putative anti-phage defense genes in orange, yellow and grey, respectively, c, Same as b but for the Brigl homolog from Nocardioides anomalus. d, AlphaFold2 structure of the Brigl homolog from N. zhouii, colored by position (N-terminal, blue; to C-terminal, red), e, Same as d but colored by pLDDT (a per-residue model confidence score). Red to blue spectrum represents high to low confidence of secondary structure prediction, f, Same as d but showing the AlphaFold2 structure of the Brigl homolog from Nocardioides anomalus. g, Same as e but showing the AlphaFold2 structure of the Brigl homolog from Nocardioides anomalus.
[0032] DETAILED DESCRIPTION
[0033] Unless defined otherwise herein, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0034] Every numerical range given throughout this specification includes its upper and lower values, as well as every narrower numerical range that falls within it, as if such narrower numerical ranges were all expressly written herein.
[0035] As used in the specification and the appended claims, the singular forms “a” "and” and “the" include plural referents unless the context clearly dictates otherwise. Ranges and other values may be expressed herein as from “about” or “approximately” one particular value, and / or to “about” or “approximately” another particular value. When values are expressed as approximations by the use of the antecedent “about” or “approximately” it will be understood that the particular value forms another embodiment. The term “about” and “approximately” in relation to a numerical value encompasses variations of + / -10%, to + / - 1%.
[0036] The disclosure includes all steps and reagents such as proteins and nucleic acids, and all combinations of steps reagents, described herein, and as depicted on the accompanying figures. The described steps may be performed as described, including but not necessarily sequentially.
[0037] Amino acids of all protein sequences and all polynucleotide sequences encoding them are also included, including but not limited to sequences included by way of sequence alignments. Sequences of from 80.00%-99.99% identical to any sequence (amino acids and nucleotide sequences) of this disclosure are included. The disclosure includes any protein having at least 80% amino acid sequence identity with a specific amino acid sequence defined herein by way of a sequence identifier or database entry. Percent amino acid sequence identity with respect proteins means the percentage of amino acid residues in another sequence that are identical with the amino acid residues in the defined sequence, after aligning the sequences in the same reading frame and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and optionally not considering any conservative substitutions as part of the sequence identity.
[0038] The disclosure includes all polynucleotide and all amino acid sequences that are identified herein by way of a database entry. Such sequences are incorporated herein as they exist in the database on the filing date of this application or patent.
[0039] The disclosure includes all amino acid sequences that are defined by sequence identifier, but with one or more changed amino acids, relative to a native amino acid sequence. Amino acid changes include conservative changes, such as by changing an amino acid belonging to a grouping of amino acids having a particular size or characteristic to an amino acid belonging to the same grouping, and non-conservative changes, such as such as changing an amino acid belonging to a grouping of amino acids having a particular size or characteristic to an amino acid belonging to another grouping, including but not necessarily limited to modifications of the Brig 1 sequence to match the sequence of described Brig 1 homologues. In examples, based on the description described herein, a catalytically inactive Brig 1 may be provided, such as by modifying amino acids in a described nucleotide binding pocket.
[0040] In an example the disclosure provides an isolated described Brigl protein, having at least 90% sequence identity the Brigl amino acid sequence:
[0041] MSARERVADLWDEVVANWQPGQDVWPGPLDDWFKSYRGRGSGAVDLDQYPDP WVGDLRGRVREPRLVVLGLNPGIGYEELQGPNGTWTKRILDTGYSHCLHRSPPED PAGWIPVHKRPSAYWVRVEHFARRWLGDTSAGHRDILNFELYPWHSPKLTGGLAS PPAIVREFVLDPVAEVGVPQVFAFGAAWFKVAAGLALPILASWETNRGDWPSEFTG WRVGVFGLPSGQQLWSSQPGTGAPPSQAKVDALRVALQGLLG (SEQ ID NO:1)
[0042] In an example, the disclosure provides a polynucleotide encoding wild type Brigl that comprises or consists of the sequence:
[0043] ATGAGTGCACGCGAACGGGTGGCTGACCTCTGGGACGAGGTCGTCGCGAACT GGCAGCCCGGCCAGGACGTCTGGCCGGGTCCCCTCGATGACTGGTTTAAGAG TTATCGAGGCCGCGGCTCGGGTGCCGTCGACCTGGATCAATACCCCGACCCG TGGGTGGGAGATCTCCGCGGCCGCGTCCGCGAACCCCGGCTGGTGGTCCTCG GACTCAACCCCGGGATTGGCTACGAGGAACTCCAGGGGCCCAACGGGACCTG GACGAAGAGGATCCTGGACACGGGCTACAGCCACTGCCTCCACCGCAGCCCA CCCGAGGACCCTGCAGGCTGGATTCCCGTCCACAAGAGGCCGAGCGCCTACT GGGTGCGGGTCGAGCACTTTGCCCGGCGCTGGTTGGGTGACACATCCGCCGG CCATCGGGACATTCTGAACTTCGAGCTCTACCCTTGGCACTCGCCCAAGCTGA CGGGTGGGCTCGCCTCGCCTCCCGCGATCGTTCGCGAGTTCGTGCTGGATCC GGTCGCTGAGGTCGGCGTCCCGCAGGTTTTCGCCTTCGGTGCTGCATGGTTCA AAGTGGCGGCGGGATTGGCGCTGCCCATCCTTGCGTCGTGGGAGACCAATCG CGGTGACTGGCCCTCGGAGTTCACCGGCTGGCGCGTCGGCGTCTTCGGGTTA CCTTCGGGACAGCAACTCGTGGTGAGCAGTCAGCCTGGCACCGGAGCCCCTC CGTCCCAGGCGAAGGTCGACGCGCTGCGGGTGGCGCTGCAGGGACTACTGG GCTGA (SEQ ID NO:2) In an example the disclosure provides an isolated a-GT comprising an amino acid sequence at least 90% sequence identity the amino acid sequence:
[0044] MRICIFMARGLEGCGVTKFSLEQRDWFIKNGHEVTLVYAKDKSFTRTSSHDHKSFSI PVI LAKEYDKALKLVN DCDI LI I NSVPATSVQEATI N NYKKLLDN I KPSI RWVYQH DHS VLSLRRNLGLEETVRRADVIFSHSDNGDFNKVLMKEWYPETVSLFDDIEEAPTVYN FQPPMDIVKVRSTYWKDVSEINMNINRWIGRTTTWKGFYQMFDFHEKFLKPAGKST VMEGLERSPAFIAIKEKGIPYEYYGNREIDKMNLAPNQPAQILDCYINSEMLERMSK SGFGYQLSKLNQKYLQRSLEYTHLELGACGTIPVFWKSTGENLKFRVDNTPLTSHD SGIIWFDENDMESTFERIKELSSDRALYDREREKAYEFLYQHQDSSFCFKEQFDIITK (SEQ ID NO:3)
[0045] In an example the disclosure provides a DNA sequence encoding the a-GT comprising the sequence:
[0046] ATGCGTATTTGCATTTTTATGGCTCGAGGTCTTGAAGGTTGTGGTGTAACAAAAT TCTCACTCGAGCAACGTGATTGGTTTATTAAAAATGGTCATGAAGTAACTTTGGT TTATGCTAAAGATAAATCATTTACTCGTACAAGTTCTCATGACCACAAATCATTTT CAATTCCAGTTATTTTAGCTAAAGAATACGATAAAGCACTTAAGCTAGTAAATGA TTGTGATATTCTAATTATTAATTCTGTTCCTGCTACTTCCGTTCAAGAAGCTACGA TTAATAACTATAAAAAACTTTTAGATAATATTAAACCTTCTATTCGTGTTGTAGTTT ATCAGCATGATCATTCTGTTCTTTCTTTGCGTCGAAATTTGGGATTAGAAGAAAC TGTTCGTCGAGCTGATGTTATTTTTAGCCATTCTGATAATGGTGATTTTAATAAA GTTCTGATGAAAGAATGGTATCCAGAAACTGTTTCTCTGTTTGATGATATTGAAG AAGCACCGACAGTATATAATTTTCAGCCTCCTATGGATATTGTGAAGGTTCGGT CAACTTATTGGAAAGATGTTTCTGAAATTAACATGAATATCAACCGTTGGATTGG TCGTACGACTACATGGAAAGGTTTTTACCAGATGTTTGATTTTCATGAAAAATTC TTAAAACCTGCTGGTAAATCCACTGTAATGGAAGGTCTGGAACGTTCCCCTGCT TTTATTGCAATTAAGGAAAAAGGTATTCCGTATGAATATTACGGTAATCGTGAGA TTGATAAAATGAATCTCGCGCCGAATCAACCGGCACAAATCCTAGATTGTTATAT TAATAGTGAAATGCTTGAACGAATGAGTAAATCTGGCTTTGGATATCAGTTGAGT
[0047] AAACTTAACCAGAAATACTTACAACGCTCACTCGAATATACTCATCTCGAGCTTG GTGCATGTGGAACAATTCCGGTATTTTGGAAATCTACTGGCGAAAATTTAAAATT CCGTGTTGATAATACTCCTTTGACCTCGCATGATAGCGGTATCATTTGGTTTGAT GAAAATGATATGGAATCAACATTTGAACGTATTAAAGAACTGTCATCTGACCGAG CTCTTTATGACCGTGAGCGAGAAAAAGCATATGAATTTTTGTATCAGCATCAAGA TTCAAGCTTCTGCTTTAAAGAACAGTTTGACATTATTACAAAATAA (SEQ ID N0:4)
[0048] In examples, a described protein can be present in a fusion protein, i.e., a contiguous protein that contains two different protein segments. Thus, in embodiments, a fusion protein comprises additional amino acids that are added to a described protein. In examples, additional amino acids include any one or a combination of a protein purification tag, such as a Sumo or histidine tag, ribosomal skipping sequences, protease recognition sequences, and linker sequences. In an example, a described Brigl protein is modified to comprise a purification tag that comprises a poly-histidine tag. In examples, 4, 5, 6, or more histidine amino acids can be appended to the N- or C-terminus of the described Brigl protein. In an example, a Hise tag is used.
[0049] In examples, the disclosure is considered suitable for use in any eukaryotic or prokaryotic cells. For use with eukaryotic cells, a described protein may include a nuclear localization signal. In examples, a nuclear localization signal sequence comprises a nucleoplasm NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO:5) or and SV40 NLS having the sequence PKKKRKV (SEQ ID NO:6). Other NLS signals can be used and, in general, for eukaryotic purposes, a nuclear localization signal comprises one or more short sequences of positively charged lysines or arginines. In examples, a described protein or combination is introduced into cells using any suitable approach.
[0050] In examples, a described Brigl protein can be present in a fusion protein, i.e., a contiguous protein that contains two different protein segments. In examples, a fusion protein comprises an a-GT segment and a Brigl segment. The segments may be separated by a suitable linking amino acid sequence. In examples a described protein is present in a fusion protein with a segment comprising a Cas protein, including but not necessarily limited to Cas9 or dCas9. Complexes comprising the described protein(s) and DNA are provided. A described protein may also be detectably labeled to facilitate, for example, visualization of 5hmC or derivatives thereof that are bound by detectably labeled Brigl . In examples, the Brig 1 may be provided as a component of a fusion protein, wherein at least one other component of the fusion protein comprises a detectable protein, examples of which include GFP, RFP, BFP, mCherry, and the like.
[0051] In an exemplary method, a first step of the disclosure comprises treating the DNA with a T4 alpha-glucosyltransferase (a-GT) and a second step comprises treating the DNA with a described Brigl protein. Alternatively, a fusion protein comprising both the a-GT and the Brigl protein may be used to treat the DNA. “Treating” the DNA means exposing the DNA to the described protein or proteins such that detectable modifications to the DNA can be made. The disclosure includes using the described protein(s) to make detectable modifications that include but are not necessarily limited to production and / or modification of abasic sites, modifications of gentiobiosyl-nucleotides, and direct or indirect detection of the presence and / or location of 5-methylcytosine (5mC) within a particular DNA sequence. Thus, in an example, the disclosure provides a method for preparing DNA to undergo a described analysis. In examples, determining the presence and / or location of a 5mC site comprises determining an abasic site, as further described herein. In examples, the described protein are used to identify the presence and / or location of 5mC sites on ssDNA or dsDNA. Combinations of ssDNA and dsDNA may also be treated with the described proteins.
[0052] A described protein can be mixed with any pharmaceutically acceptable buffer, excipient, carrier and the like to form a pharmaceutical preparation. Suitable pharmaceutical compositions can be prepared by mixing one or more of the described proteins with a pharmaceutically-acceptable carrier, diluent or excipient, and suitable such components are well known in the art. Some examples of such carriers, diluents and excipients can be found in: Remington: The Science and Practice of Pharmacy (2023) 23rd Edition, Philadelphia, PA. Academic Press, the description of which is incorporated herein by reference. Any described protein or combination of proteins can be administered to an individual in need thereof.
[0053] In examples, a described Brigl protein of this disclosure is introduced into one or more prokaryotic or eukaryotic cells. In examples, the prokaryotic cells comprise or consist of gram positive, or gram-negative bacteria. The bacteria may be non- pathogenic, or pathogenic. In examples, a described protein is introduced into prokaryotic cells (e.g., bacterial or archaeal cells) in the context of a host, e.g., a human, animal, or plant host, e.g., or the bacteria are a component of a host’s microbiome or are an abnormal component of a microbiome, e.g., a pathogenic bacteria. In some examples, delivery of a protein described herein results in killing of the recipient cell. The protein may kill some or all of the cells, or render the cells non- pathogenic and / or sensitive to one or more antibiotics. In examples, the protein is used as a component of a food or beverage product, including but not limited to fermented food and beverages, and dairy products. In examples, selective delivery to a specific type of bacteria is used by way of a bacteriophage or packaged phagemids that can express a described Brig protein, but wherein the bacteriophage exhibits a specific tropism for a particular type of bacteria.
[0054] In examples, a protein of this disclosure is administered to an individual in a therapeutically effective amount. In examples, a therapeutically effective amount of a composition of this disclosure is used. The term “therapeutically effective amount” as used herein refers to an amount of an agent sufficient to achieve, in a single or multiple doses, the intended purpose of treatment. Appropriate effective amounts can be determined by one of ordinary skill in the art informed by the instant disclosure.
[0055] In examples, the disclosure provides for use of a described a-GT and Brig 1 for use with 5hmC determination, such uses including but not necessarily limited to the following:
[0056] ILLUMINA Next Generation Sequencing (NGS):
[0057] 1. snAP-seq:
[0058] In this example, genomic DNA is treated a-GT and Brig 1 to generate abasic sites at 5hmC nucleobases in a DNA sample. As a control, genomic (or another DNA sample) a -GT and Brig 1 treatment is used. The DNA is incubated with a Biotin- PEG3-azide probe in the presence of CuBr. This probe reacts with the aldehyde in the abasic site through a hydrazino- / so-Pictet-Spengler (HIPS) reaction, allowing for tagging and enrichment of abasic sites. A first adapter is ligated to the ends of the DNA, followed by enrichment of tagged DNA with Streptavidin beads. DNA is denatured and a single-strand break (SSB) is generated by alkaline-cleavage conditions at the abasic site. A second adapter is ligated at the abasic site position and the NGS library is generated by PCR and sequenced using for example an ILLUMINA MiSeq. Alignment of sequences against the reference genome reveals the location of the abasic site positions where enrichment is found. The original 5hmC position corresponds to the base located before the start of said enriched reads. A representative approach that can be adapted to be used with the described a-GT and Brigl protein as described in this example is described in Liu, Z.J., Martinez Cuesta, S., van Delft, P. et al. Sequencing abasic sites in DNA at singlenucleotide resolution. Nat. Chem. 11 , 629-637 (2019), the disclosure of which is incorporated herein by reference.
[0059] 2. SSiNGLe-AP:
[0060] In this example, genomic DNA is fragmented and denatured. Abasic sites are generated at the 5hmC nucleobase positions using a-GT and Brigl . As a control, genomic DNA with no Brigl treatment is used. The endogenous 3’OH termini of the fragments are blocked with terminal transferase and biotin-11-dCTP. The blocked DNA is captured using streptavidin-coated magnetic beads and treated with an enzyme that generates SSBs at the abasic sites. The newly created free 3’OH termini are tagged by polyA-tailing that is used for NGS library preparation. Paired- end sequencing is carried out with ILLUMINA NovaSeq. Abasic sites are located by shifting the position of the first aligned base, which corresponds to the position of the enzyme-generated 3’OH, by one base. A representative approach that can be adapted to be used with a described a-GT and Brigl as described in this example is described in Abasic site position determines original 5hmC position in the genome. Cai, Y., Cao, H., Wang, F. et al. Complex genomic patterns of abasic sites in mammalian DNA revealed by a high-resolution SSiNGLe-AP method. Nat Commun 13, 5868 (2022), the disclosure of which is incorporated herein by reference.
[0061] 3. TLS-APseq : In this example, genomic DNA is fragmented followed by denaturation. Enrichment of DNA with 5hmC nucleobases is performed using a-GT and Brigl and Ni-NTA magnetic beads. Different adapters are ligated to the 5’ and 3’ termini of the fragments followed by a-GT and Brigl treatment to generate abasic sites. A control DNA library is left untreated. Extension #1 of fragments is performed using a primer complementary to the 5’ adapter and the translesion synthesis (TLS) polymerase, Sulfolobus DNA polymerase IV, which is commercially available. This polymerase is able to replicate DNA past the abasic site by adding a nucleobase, typically adenine. Extension #2 is performed over the product of extension #1 using a primer complementary to the 3’ adapter and a high-fidelity DNA Polymerase. These extensions are expected to yield a DNA molecule with an A-T pairing at the abasic site position, if not a different “error”. Library preparation is carried out by PCR using primers complementary to sequences unique to primers 1 and 2, which are not found in the adapters used initially. Paired-end sequencing is carried out with ILLUMINA MiSeq or ILLUMINA NovaSeq. Sequencing reads are mapped back to the reference genome and only “errors” that are found in every copy are regarded as an abasic site, with a change from C to A on the leading strand and from G to T on the lagging strand. A representative approach that can be adapted to be used with a described a-GT and Brigl protein as described in this example is described in Thompson PS, Cortez D. New insights into abasic site repair and tolerance. DNA Repair (Amst). 2020 Jun;90: 102866. doi: 10.1016 / j.dnarep.2020.102866. Epub 2020 Apr 30; the disclosure of which is incorporated by reference
[0062] Nanopore Sequencing:
[0063] This example relates to a crown-ether electrolyte adduct. In this example, genomic DNA is extracted. Abasic sites are generated at the 5hmC nucleobase positions using a-GT and Brigl . As a control, genomic DNA with no a-GT and Brigl treatment is used. The DNA is incubated with 2-aminomethyl-18-crown-6 in the presence of NaBHsCN. 2-aminomethyl-18-crown-6 is functionalized to the abasic sites via reductive amination forming a stable adduct. This adduct generates unique amplitude current signatures when translocated through a nanopore. DNA fragments are sequenced using the MinlON from Oxford Nanopore. Base-calling software is modified to recognize the unique electrical signature generated by the adduct as an abasic site. A representative approach that can be adapted to be used with a described a-GT and Brigl as described in this example is described in An N, Fleming AM, White HS, Burrows CJ. Crown ether-electrolyte interactions permit nanopore detection of individual DNA abasic sites in single molecules. Proc Natl Acad Sci U S A. 2012 Jul 17; 109(29): 11504-9. doi: 10.1073 / pnas.1201669109. Epub 2012 Jun 1 , the disclosure of which is incorporated herein by reference.
[0064] In embodiments, the disclosure provides an article of manufacture, which may comprise a kit. In embodiments, the article of manufacture may comprise a described combination of a-GT and Brigl , or a fusion protein comprising a-GT and Brigl , and one or reagents for use in the described methods. An article of manufacture may include one or more sealed containers that contain any of the aforementioned components, and may further comprise packaging and / or printed material. The printed material may provide information on the contents of the article, and may provide instructions or other indication of how the contents of the article can be used, such as for isolating, immobilizing, detecting, or sequencing a DNA sample.
[0065] The following description is related to the present disclosure and describes the identification of Brigl and its interaction with a-GT.
[0066] Previous work has expanded understanding of the diversity of immune systems present in bacteria, but relied on the availability of sequenced genomes. Analyses of 16S ribosomal RNA sequences suggest that uncultivated and unsequenced microbes represent the majority of bacterial lineages. This “microbial dark matter”, which is beginning to be accessed through single-cell genomics may contain a vast assortment of unknown genetic pathways, including those involved in anti-phage defense. This disclosure describes analysis of this uncharted sequence space by screening a library of environmental DNA (eDNA) constructed in E. coli for clones showing resistance or immunity to phage T4 infection. This library was constructed by cloning microbial DNA isolated from an arid soil collected in Arizona into a cosmid vector. After challenging this library with the lytic coliphage T4, we isolated Brigl , which is, as described above, a DNA glycosylase from an unknown organism that provides immunity through the excision of alpha-glucosyl- hydroxymethylcytosine nucleobases present in the T4 genome, thus inhibiting phage replication after infection. Isolation of defense genes from eDNA
[0067] To uncover novel anti-phage defense systems present in unsequenced bacterial genomes, we screened an eDNA library generated in an earlier study, harboring large (~40 kb) DNA fragments extracted from a soil sample collected in Arizona that were cloned into pWEB-TNC cosmids, packaged into I phage and transfected into E. coli EC100 cells3 11. The library, containing at least 10 million clones, was infected with phage T4 to select resistant clones that enable colony formation (Fig. 6a). We picked sixteen random colonies and used a plaque assay to determine that twelve carried immunity to T4 phage infection, but did not provide immunity to an unrelated phage, \vir (Fig. 6b) for unedited images of plaque assays and gel electrophoresis). Sequencing of the selected cosmids revealed a 34.5 kb DNA insert (Fig. 6c), likely belonging to the phylum Actinobacteria. Further subcloning of cosmid fragments (Fig. 6c-f), which shows replicates for these and all subsequent plaque assays reported in this study) determined that gene “c", encoding an unknown protein belonging to the superfamily of uracil DNA glycosylases, was solely responsible for immunity (Fig. 1 a, Fig. 6g). Importantly, this gene lies within the vicinity of other putative defense systems (Fig. 6c, e), including Thoeris and Wadjet, a genetic context that suggests that gene c is part of a bacterial defense island present in the isolated eDNA.
[0068] To investigate how gene c affects T4 infection, we first determined that it does not affect phage adsorption (Fig. 7a). We also performed quantitative PCR (qPCR) to measure phage DNA accumulation (at two different loci, gp43 and gp34) in infected cells at 2, 4, 8 and 20 minutes post infection (Fig. 1 b and Fig. 7b-f). We found that, in contrast to susceptible hosts, which showed a steady increase in phage DNA over time, the T4 DNA content decreased in E. coli expressing gene c (Fig. 1 b and Fig. 7b-f). This result demonstrates that gene c not only inhibits T4 DNA replication but also causes a slight and gradual depletion of the phage DNA within the infected population (Fig. 7c, d). We performed next-generation sequencing (NGS) of E. coli cells infected with T4 for 8 minutes, which showed that phage DNA reads were severely depleted across the entire T4 genome in cells expressing gene c (Fig. 1 c). In light of these results we named this gene bacteriophage replication / nhibition DNA glycosylase 1, Brigl . The cosmid harboring only gene c, pFgmD3-4 (Fig. 1 a), was therefore renamed pBrigl . Brigl targets glucosylated DNA bases
[0069] To analyze how Brigl affects T4 replication, we isolated an “escaper” phage that completely bypassed Brigl defense (Fig. 2a) and displayed a very similar pattern of DNA reads during infection in the presence or absence of Brigl expression in E. coli hosts (Fig. 2b). Sequencing of the viral DNA revealed a single base pair deletion within the T4 alpha-glucosyltransferase (a-gf) gene, resulting in a frameshift (escaperl). Sequencing of PCR products obtained using DNA isolated from another 18 escaper phages showed additional mutations in a-gt, most of them frameshifts. We also generated an in-frame deletion of this gene, phage T4 Aa-gt, which phenocopied the escaped mutation (Fig. 2a). Finally, plasmid-borne expression of a-gt rescued E. coli from lysis by both T4 escaped and T4 Aa-gt (Fig. 2a), a result which demonstrates that this gene is required for Brigl defense.
[0070] Alpha- and beta-glucosyltransferases (a-GT and b-GT, respectively) add glucose in alpha- or beta-linkage to -70% and -30% of the 5-hydroxymethylcytosine (hmC) bases in the T4 genome, respectively (Fig. 7g-h). Brigl provided strong immunity upon infection with a mutant phage carrying an in-frame deletion of b-gt, T4 Ab-gt, (Fig. 2a). In addition, we constructed a cytosine-containing T4 mutant phage, T4(C), with the genotype Aalc AdenB Agp56 Agp42 (Fig. 7h), which has been shown to lack hmC in its genome. This phage was resistant to Brigl targeting, in contrast to the triple mutant Aalc AdenB Agp56 phage that harbors both gp42 to synthesize hmC nucleobases and a-gt to glycosylate the bases (Fig. 2c). Moreover, overexpression of gp42, but not a-gt alone, sensitized T4(C) to Brigl targeting (Fig. 2c). We passaged T4 on a strain that overexpressed b-GT to investigate the effect of an increase in the fraction of beta-glucosylated hmC nucleobases on Brigl immunity. We found that the resulting phage, T4(+b-GT), displayed a small but significant increase in propagation in the presence of the enzyme (Fig. 7i). Taken together, these demonstrate that Brigl targets alpha-glucosylated hmC nucleobases in the viral DNA to provide defense against T4.
[0071] Brigl excises a-glucosyl-hmC nucleobases
[0072] We performed an AlphaFold2 protein structure prediction of Brigl , which generated a high confidence structural model (Fig. 3a and Fig. 8a) that we used to find structural homologs. The top hits were all uracil DNA glycosylases, with the best match being a uracil DNA glycosylase from the archaeon Sulfolobus tokodaii (Dali Z score 9.2)20(Fig. 8b). Uracil DNA glycosylases recognize uracil bases in DNA (which may result from polymerase error or from cytosine deamination) and initiate base excision repair by hydrolyzing the N-glycosidic bond between the base and the deoxyribose sugar. We therefore analyzed whether Brig 1 removes alpha-, but not beta-glucosyl-hmC, from the T4 genome. To test this, we purified Brigl and determined its activity on a 60-nt ssDNA oligonucleotide substrate containing a single hmC residue within an Mfel restriction site (Fig. 8c), to which we introduced alpha- and beta-glucosyl-hmC modifications (Fig. 8d) using purified T4 a-GT and b- GT enzymes. Glucosylation was confirmed by annealing a complementary oligonucleotide to generate a dsDNA substrate for Mfel digestion, a restriction endonuclease that can cleave hmC- but not glucosyl-hmC-containing target sequences (Fig. 8e). The modified ssDNA oligonucleotides were incubated with Brigl to determine its DNA glycosylase activity using an aldehyde-reactive fluorescent probe that can detect abasic sites. The primary product of Brigl was a full-length oligonucleotide containing an abasic site, generated by excision of the alpha- but not the beta-glucosyl-hmC nucleobase (Fig. 3b). As a positive control, we treated an equivalent uracil-containing oligonucleotide (Fig. 8f, g) with hSMUGI , a previously characterized human uracil DNA glycosylase.
[0073] We further tested Brigl activity by treating reaction products with heat and NaOH, conditions that accelerate the cleavage of the DNA backbone at abasic sites via beta-elimination and thus enable the detection of these sites as DNA fragments. Using this method, Brigl exhibited robust base excision activity on the alpha- glucosyl-hmC-containing substrate but not on ssDNA substrates harboring beta- glucosyl-hmC, hmC, 5-methylcytosine, or 2-aminoadenine (Fig. 8g-h). We confirmed Brigl activity using a third method to detect abasic sites. We incubated the oligonucleotide product with endonuclease IV, an apurinic / apyrimidinic (AP) endonuclease that cleaves the sugar-phosphate backbone adjacent to an abasic site. This treatment resulted in the cleavage of the ssDNA substrate at the position where the glucosyl-hmC nucleobase is located (Fig. 8i). To obtain direct evidence of the removal of the glucosyl-hmC nucleobase, we performed high resolution mass spectrometry (MS). We treated an 18-nt ssDNA oligonucleotide containing either a single uracil, hmC or alpha-glucosyl-hmC nucleobase (Fig. 9a) with hSMUGI or Brigl . With hSMUGI and Brigl treatment of the uracil- and alpha-glucosyl-hmC-containing oligonucleotides, respectively, we recorded, in each case, strong primary peaks with mass values equivalent to the loss of the excised target nucleobase and the gain of a water molecule that were interpreted as the introduction of an abasic site (Figs. 9b, c). Together, these results demonstrate that Brigl is a DNA glycosylase that excises alpha-glucosyl-hmC nucleobases from ssDNA to generate abasic sites, with a high level of stereoisomeric specificity.
[0074] Brigl generates abasic sites in dsDNA
[0075] Although the T4 genome can have ssDNA intermediates during replication, most of the viral DNA is in a double-stranded form. We therefore tested whether Brigl can also introduce abasic sites in dsDNA oligonucleotide substrates (Fig. 10a, b) which were glucosylated at hmC sites and cleaved with Mfel to confirm the presence of the modification (Fig. 10c). We first treated with Brigl or hSMUGI a dsDNA substrate harboring either a single alpha-glucosyl-hmC or uracil in the top strand. When separated by non-denaturing polyacrylamide gel electrophoresis (PAGE), only the products treated with heat / NaOH showed a high molecular weight species (Fig. 10d). This is most likely due to the generation of nicked dsDNA that runs slower on a non-denaturing gel. To corroborate this, we separated the products by urea-PAGE, which revealed cleavage products generated by Brigl and subsequent heat / NaOH treatment, but not by treatment with the glycosylase alone (Fig. 3c and Fig. 10e). In contrast, dsDNA substrates containing hmC in the top strand, that were treated with Brigl and heat / NaOH, did not show cleavage products that are indicative of the introduction of abasic sites (Fig. 10f). Next, after confirming that both top and bottom ssDNA oligonucleotides were subject to glycosylase activity (Fig. 10g), we tested DNA duplex molecules containing modified bases in both strands. Brigl treatment led to the cleavage of both strands after beta-elimination, generating a dsDNA break and cleavage products that can be separated by nondenaturing, native PAGE (Fig. 3d). Similar results were obtained for hSMUGI experiments (Fig. 10h). Brigl did not display any activity on the dsDNA substrate harboring hmC in both strands (Fig. 10i). Altogether, these results indicate that Brigl is a monofunctional DNA glycosylase that generates abasic sites, with insignificant lyase activity.
[0076] Brigl degrades T4 phage DNA
[0077] We also tested the effect of Brigl on T4 phage DNA in vitro. We incubated wild-type and escaped viral DNA with Brigl for 30 minutes at 37°C and visualized the products via agarose gel electrophoresis. While the T4 DNA treated with Brigl was not distinguishable from an untreated control DNA when using low voltage and low temperature (4°C) to separate the reaction products, experiments at higher voltage and temperature resulted in DNA cleavage and degradation via betaelimination at abasic sites caused by the heat generated during electrophoresis (Fig. 11 a-b and Fig. 3e). In these conditions, increasing concentrations of Brigl only caused a mobility shift, but no degradation, in escaped DNA and a cosmid control DNA (pWEB-TNC) (Fig. 3e). Heating to 65°C or treatment with SDS of the reaction products before electrophoresis eliminated the mobility shifts (Fig. 11 c-d), results that suggest that Brigl can bind to non-target DNA. Brigl treatment of phage T4(+b- GT) DNA, which contains a higher proportion of beta-glucosyl-hmC nucleobases than wild-type T4 DNA (Fig. 7i) , resulted in the generation of a lower number of abasic sites, evidenced by the reduced DNA degradation via heat-promoted betaelimination during electrophoresis (Fig. 11e). Altogether these data indicate that heat-promoted beta-elimination at abasic sites generated by Brigl results in the degradation of wild-type, alpha-glucosylated, T4 phage DNA.
[0078] Brigl residues important for activity
[0079] Brigl ’s nucleotide binding pocket is predicted to be much larger (Fig. 3a) than that of the S. tokodaii uracil DNA glycosylase (Fig. 8b), with extra space adjacent to the C5 position of the pyrimidine where the additional alpha-glucosyl-hydroxymethyl group would protrude (Fig. 3a and Fig. 8d). To test if this putative binding pocket is important for Brigl activity, we mutated amino acids predicted to outline this area: Y121 , E147 and N 145 (Fig. 3a). Based on the structure of other related glycosylases, Y121 would stack against the flipped-out base (as is the case for F55 in S. tokodaii uracil DNA glycosylase, Fig. 8b), while E147 would form hydrogen bonds to its Watson:Crick face. Because this residue is often asparagine rather than glutamate (for example N82 in S. tokodaii uracil DNA glycosylase, Fig. 8b), we also considered the N145 residue (Fig. 3a). In vivo, the Y121A, E147A and E147Q, but not the N145A, substitutions affected Brigl -mediated immunity (Fig. 12a). In vitro, the Y121A / E147A double mutation abrogated base excision activity on ssDNA oligonucleotides as well as on T4 DNA (Fig. 12b-c). These results demonstrate that the putative DNA glycosylase catalytic pocket of Brigl is involved with base excision activity as well as defense against phage T4.
[0080] Involvement of host DNA repair pathways
[0081] Since endonuclease IV can cleave ssDNA oligonucleotides at the abasic sites generated by Brigl (Fig. 8g), we analyzed whether other enzymes that participate in base excision repair in E. coli could be involved with Brigl immunity in vivo. To explore this, we tested immunity in hosts lacking either one or both of the two major E. coli AP endonucleases, exonuclease III (XthA) and endonuclease IV (Nfo), the pyrimidine DNA glycosylase-lyase endonuclease III (Nth) or the abasic site sensor YedK. Deletion of any of the genes encoding these enzymes did not affect T4 PFU counts in the presence of Brigl (Fig. 13a), suggesting that they are not required for immunity. We also performed the opposite experiment, i.e., overexpressing XthA and Nfo to determine if they enhance Brigl immunity and found that neither of the AP endonucleases provided a further decrease in T4 PFUs (Fig. 13b). We determined that other nucleases, helicases and recombinases involved in recombinational DNA repair - RecBCD, RecQ, RecJ and RecA, which could process DNA ends generated by the sequential activity of Brigl and host- or phage-encoded AP endonucleases - did not affect immunity (Fig. 13c). Therefore, the described data indicate that the major E. coli DNA repair enzymes and AP endonucleases do not play a specialized role in Brigl -mediated anti-phage defense.
[0082] Brigl immunity against diverse phages
[0083] To test the range of phages restricted by Brigl , we infected E. coli with seven different coliphages and found that, in addition to T4, phages T2 and T6 were highly sensitive to Brigl targeting (Fig. 4a). These phages contain alpha-glucosylated (70% in T2; 3% in T6), but not beta-glucosylated hmC sites. In addition, both phage genomes carry beta-1 ,6-glucosyl-alpha-glucose (gentiobiose, Fig. 14a) adducts (T2, 5%; T6, 72%). Since the majority of the hmC nucleobases in the T2 genome are alpha-glucosylated, this phage is expectedly very sensitive to Brig 1 targeting (Fig. 4a). On the other hand, since only a small fraction of the T6 genome contains alpha- glucosyl-hmC, the high susceptibility of this phage to Brigl (Fig. 4a) is intriguing. To investigate this, we isolated two T6 phages that escaped targeting (Fig. 14b) and found that both carried inactivating mutations in the T6 a-gt gene which effect was reverted through expression of phage T4 a-gt (Fig. 14b). We also treated T2, T4 and T6 phage DNA with purified Brigl (Fig. 4c and Fig. 14c). We found that T4 and T2 DNA, but not T6 DNA, was partially degraded during electrophoresis, a result that, as opposed to the in vivo results, correlates with the low fraction of alpha-glucosyl- hmC nucleobases in the T6 genome. To test this, we deleted the ba-gt gene, which encodes beta-alpha glucosyltransferase (ba-GT), the enzyme required to add the second glucose in beta-linkage to alpha-glucosyl-hmC nucleobases and generate gentiobiosyl-hmC (Fig. 14a). This phage, T6 ba-gt, only carries alpha-glucosylated hmC nucleobases (presumably in 75% of the cytosines, Fig. 14a), and is more susceptible to Brigl immunity than wild-type T6 (Fig. 4a and Fig. 14d). In addition, treatment of T6 Aba-gt DNA with Brigl resulted in degradation after running the products on a gel (Fig. 4b). Altogether these results suggest that while gentiobiose modifications render T6 DNA resistant to Brigl in vitro, in vivo there is a window during the viral lytic cycle after the activity of a-GT on newly replicated hmC nucleobases but before the addition of the second glucose by ba-GT, in which a large proportion of the hmC nucleobases in T6 are modified only with alpha-glucose and therefore susceptible to Brigl restriction.
[0084] We tested 69 different E. coli phages from the BASEL collection and found that Brigl provides immunity against Bas35-45, all members of the T-even family that modify their genomes with alpha-glucosyl-hmC (Fig. 14e). In contrast, plaque formation by two other T-even phages within the collection, Bas46-47, predicted to carry arabinosyl-hmC nucleobases instead of glucosyl-hmC, was not affected by Brigl (Fig. 14e). Overall, these data demonstrate that Brigl restricts a large number of T-even phages that contain alpha-glucosylated hmC residues in their genomes. While additional modifications of these nucleobases prevent Brigl activity, their transient presence during the lytic cycle is sufficient for efficient immunity.
[0085] Homologs of Brigl also provide immunity We used PSI-BLAST to analyze the prevalence of Brig 1 in prokaryotic genomes and found 42 non-redundant homologs (annotated on NCBI as hypothetical proteins). Many of these are present within putative anti-phage defense islands (Fig. 5a and Fig. 15a), near other annotated anti-phage immunity genes. Most of the Brig 1 homologs currently available in genetic databases are found in Actinobacteria (Fig. 5a). We found that two related Brig 1 homologs, both present in Nocardioides, provided protection (Fig. 5a, b). These homologs are present in putative defense islands (Fig. 15b, c), with the one harbored by Nocardioides zhouii located in a similar genomic neighborhood as Brig 1 , i.e., adjacent to a predicted ADP-ribosyl glycohydrolase and near a ThsA-like SIR2-domain protein (Fig. 15b). Both share -50% amino acid identity with Brig 1 and a high level of predicted structural similarity (Figs. 15d-g).
[0086] METHODS
[0087] Bacterial strains and growth conditions
[0088] Cultivation of E. coli EC100 (Lucigen), E. coli K-12 MG1655, E. coli K-12 BW25113 and all other E. coli strains used in this disclosure were carried out in lysogeny broth (LB) at 37°C with shaking. Overnight cultures were inoculated from single bacterial colonies. Wherever applicable, media were supplemented with chloramphenicol at 12.5 pg / mL (for cosmids) or 25 pg / mL (for plasmids), spectinomycin at 50 pg / mL, kanamycin at 50 pg / mL, ampicillin or carbenicillin at 100 pg / mL, and / or tetracycline at 5 pg / mL to ensure cosmid or plasmid maintenance. E. coli Keio knockout strains were obtained from the Coli Genetic Stock Center at Yale University. Miniprepped plasmids (prepared by QIAprep Spin Miniprep Kit, QIAGEN, Cat# 27106) were cloned into chemically competent E. coli EC100 cells (Lucigen), electrocom petent E. coli EC100 cells (Lucigen) or rubidium chloride (RbC ) chemically competent E. coli K-12 MG1655 cells. For E. coli K-12 BW25113 and Keio knockout strains, protein purification strains, and strains with two plasmid combinations, existing strains were first made electrocom petent and then transformed with plasmid through electroporation (1 mm Bio-Rad Gene Pulser cuvette at 1.8 kV).
[0089] Gibson assembly For Gibson assemblies, 25-100 ng of the largest dsDNA fragment was combined with equimolar volumes of the smaller fragment(s) in a total volume of 5 pL in nuclease-free water. Reaction mixtures were prepared on ice and mixed with 15 pL of Gibson assembly master mix, pipette mixed and incubated at 50°C for 1 hour in a thermal cycler. Gibson reactions were transformed into chemically competent E. coli EC100 cells (Lucigen) or RbCh chemically competent E. coli K-12 MG1655 cells by mixing 5 pL of Gibson reaction with 50 pL cells and following a standard transformation protocol for chemically competent cells.
[0090] Oligo cloning
[0091] Oligo cloning was used to create a repeat-spacer-repeat CRISPR array with a desired spacer following a protocol previously described by this laboratory. Briefly, we used a Bsal restriction digest cloning approach. Parent type I l-A CRISPR arraycontaining plasmids with a repeat-spacer-repeat carried a 30 bp spacer sequence with two Bsal cut sites at either end (pCas9). To set up the Bsal plasmid digest, we mixed 42 pL of the parent CRISPR plasmid (40-60 ng / pL) with 6 pL Bsal-HF (NEB, Cat# R3535L), 6 pL NEB CutSmart buffer and 6 pL nuclease-free water. The restriction digest reaction was incubated at 37°C for approximately 6 hours. Two IDT oligonucleotides (oligos) comprised the type ll-A CRISPR spacer to be inserted into the Bsal cut plasmid CRISPR array: a “top” strand oligo with sequence 5’-AAAC-(30 bp spacer)-G-3’ and a “bottom” strand oligo with sequence 5’-AAAAC-(30 bp spacer reverse complement)-3’. For oligo cloning of type l-E spacers into pACYC184- TypelEspc / VT, the top strand oligo had sequence 5’-ACCG-(32 bp spacer)-3’ and the bottom strand oligo had sequence 5’-ACTC-(32 bp spacer reverse complement)-3’. The two oligos were phosphorylated with T4 polynucleotide kinase (NEB, Cat# M0201S) in a 50 pL reaction: 1.5 pL 100 pM top oligo, 1.5 pL 100 pM bottom oligo, 41 pL nuclease-free water, 5 pL T4 DNA ligase reaction buffer (NEB, Cat# B0202S), 1 pL T4 polynucleotide kinase (NEB, Cat# M0201S). The reaction was incubated at 37°C for 1 hour in a thermal cycler. After phosphorylation, oligos were annealed_by adding 2.5 pL of 1 M sodium chloride (Fisher Scientific, Cat# S271-3) solution to the 50 pL reaction and incubating for 5 minutes at 98°C and then allowing the reaction to gradually cool to room_temperature (approximately 2 hours). The annealed oligos were diluted 1 :10 in nuclease-free water and ligated into the Bsal-digested plasmid in a 20 pL reaction: 10 pL Bsal-digested plasmid, 6 pL nuclease-free water, 1 pL 1 :10 diluted annealed oligos, 5 pL T4 DNA ligase reaction buffer (NEB, Cat# B0202S), 1 pL T4 DNA ligase (NEB, Cat# M0202M). The ligation reaction was performed at room temperature overnight. The next day, 5 piL of the ligation reaction was transformed into 50 pL of chemically competent E. coli EC100 cells (Lucigen) and colonies were confirmed by PCR the next day.
[0092] Strain construction
[0093] A Red recombineering was used to generate the E. coli K-12 BW25113 DxthA Dnfo strain. An overnight culture of the E. coli K-12 BW25113 Dnfo Keio strain carrying the pAM38(red) plasmid with chloramphenicol resistance was diluted and grown to ODeoo~0.3 and then induced with 0.2% L-arabinose till ODeoo~1-1.2. Cells were made electrocom petent by washing twice with cold water and electroporated (1 mm Bio-Rad Gene Pulser cuvette at 1 .8 kV) with a PCR product carrying a xthA:tetR gene replacement matching the xthA:kanR gene replacement found in the E. coli K- 12 BW25113 DxthA Keio strain, with ~50 bp homology upstream and downstream of the xthA locus in the PCR product. After ~2 hours of recovery, cells were plated on LB agar plates with kanamycin at 50 pg / mL and tetracycline at 5 pg / mL to select for double mutants. Double knockouts were confirmed by PCR. After confirmation, strains were grown overnight in LB with kanamycin at 50 pg / mL and tetracycline at 5 pg / mL (but no chloramphenicol which selects for the plasmid) and with 0.2% L- arabinose induction. Without antibiotic selection, induced plasmid was rapidly lost due to toxicity from A Red overexpression. Strains were frozen at -80°C (900 pL culture + 100 pL DMSO) and struck out on appropriate antibiotic plates to confirm both double knockouts and loss of the recombineering plasmid.
[0094] Preparation of phage stocks
[0095] T2, T3, T5 and T6 phages were purchased from ATCC. Phages were first grown up in 10 mL cultures of exponentially growing E. coli K-12 MG1655 or EC100 cells at OD6oo~0.3. The phage-added cultures were incubated at 37°C with shaking overnight. Tubes were then spun down at 15,000 x g for 10 minutes at 4°C. Phagecontaining supernatants were filtered using Acrodisc 13 mm SUPOR 0.45 pm syringe filters (Pall, 4604) into 15 mL conical tubes and supernatants frozen down as phage stocks at -80°C (900 pL filtered supernatant + 100 pL DMSO). To grow up a phage stock for plaquing assays and other experiments, a pipette tip was used to scrape off a tiny portion of a frozen phage stock, which was then resuspended in 20 pL LB medium. Serial dilutions were prepared from the resuspended phage and spotted on a fresh LB top agar (LB broth Lennox base, 0.5% agar) lawn of E. coli EC100 in LB agar. The plate was incubated at 37°C overnight after drying at room temperature for 25 minutes. The next day a single phage plaque was picked from the top agar lawn using a P20 pipette set to 15 pL and resuspended in a 10 mL culture of exponentially growing E. coli EC100 at ODeoo~0.3. The phage-added culture was incubated at 37°C with shaking overnight. The tube was spun down the next day at 15,000 x g for 10 minutes at 4°C. The phage-containing supernatant was filtered using an Acrodisc 13 mm SUPOR 0.45 pm syringe filter (Pall, 4604) into a 15 mL conical tube. All final phage stocks were titered on top agar lawns of E. coli EC100 and stored at 4°C.
[0096] To grow phage stocks of Brig 1 escaper phages, single plaques formed by T4 or T6 phages on lawns of pBrig 1 -carrying EC100 cells were picked using a P20 pipette and resuspended in 20 pL LB medium. Serial dilutions were prepared from the resuspended phage and spotted on a fresh LB top agar lawn of E. coli EC100 carrying pBrig 1 to maintain selection of the escaper phage. The plate was incubated at 37°C overnight after drying at room temperature for 25 minutes. The next day a single phage plaque was picked from the top agar lawn using a P20 pipette set to 15 pL and resuspended in a 10 mL culture of exponentially growing ODeoo~0.3 E. coli EC100 carrying pBrigl for continued selection. The phage-added culture was incubated at 37°C with shaking overnight and filtered the next day as described earlier to generate the escaper phage stock. Final phage stocks were titered on top agar lawns of E. coli EC100 and stored at 4°C.
[0097] Generation of mutant phage stocks
[0098] T4 and T6 phage stocks were used to construct T4 Aa-g , T4 Ab-gt, T4 Aa / c AdenB Agp56, T4(C) and T6 ba-gt mutant phage stocks. In each case, a culture of E. coli EC100 cells carrying a recombinant pUT18C-based plasmid was grown overnight at 37°C with shaking in 10 mL LB supplemented with 100 pg / mL carbenicillin. The pUT18C plasmid contained a cloned segment of phage T4 or T6 DNA with the desired gene deleted and -750-1000 bp homology arms flanking the deleted genic region on either side. The overnight culture was diluted 1 :50 in 10 mL LB medium supplemented with 100 pg / mL carbenicillin. After approximately 1 hour of culture growth, ODeoo was measured for the culture and confirmed to be between 0.2-0.4. The 10 mL culture was then infected with 2 ja.L of T4 or T6 phage stock and grown overnight at 37°C with shaking to allow wild-type phages to recombine with the plasmid. The next day, the tube was spun down at 15,000 x g for 10 minutes at 4°C. The phage-containing supernatant was filtered using an Acrodisc 13 mm SUPOR 0.45 pm syringe filter (Pall, 4604) into a 15 mL conical tube.
[0099] Serial dilutions of recombinant phage were prepared and spotted on a fresh top agar lawn of E. coli EC100 containing a pCas9 plasmid in LB agar supplemented with 25 pg / mL chloramphenicol. The pCas9 plasmid carried a type I l-A CRISPR spacer targeting the phage gene that was deleted to select specifically for recombinant phage with the desired deletion. Top agar plates were incubated at 37°C overnight after drying at room temperature for 25 minutes. The next day multiple phage plaques were picked from the top agar lawn using a P20 pipette set to 15 pL and resuspended in 20 pL LB medium. 5 pL of the resuspend phage plaques were boiled in 15 pL colony lysis buffer48at 98°C for 15 minutes and then PCR checked to confirm that the desired gene was deleted, either with the deletion carried on the pUT18C recombinant plasmid or a de novo CRISPR-generated deletion that eliminated the appropriate gene. Serial dilutions were prepared for 1-2 correct phage plaques, which were then replaqued onto top agar lawns of pCas9 selection strains and incubated overnight at 37°C for stringent selection. The next day, a single phage plaque was picked from the top agar lawn using a P20 pipette set to 15 pL and pipetted directly into an OD6oo~0.2-0.4 exponentially growing culture that maintained the same selection for the mutant phage. The phage-infected culture was grown overnight at 37°C with shaking. The next day, the tube was spun down at 15,000 x g for 10 minutes at 4°C. The phage-containing supernatant was filtered using an Acrodisc 13 mm SUPOR 0.45 pm syringe filter (Pall, 4604) into a 15 mL conical tube. In some cases, an arabinose inducible type l-E CRISPR-Cas expressing E. coli MG1655 strain, ACT-01 , with a pACYCI 84-based plasmid expressing an arabinose-inducible type l-E CRISPR spacer was used to select for the recombinant phage. In these instances, 0.2% L-arabinose was included in all media for proper phage selection through type l-E CRISPR-Cas targeting. To make the T4 ,\a-gt phage, instead of CRISPR selection, E. coli EC100 / pBrig1 was used to select for the pUT18C-recombined phage. PCR and Sanger sequencing confirmed the desired in-frame deletion of a-gt in the mutant phage, matching the exact deletion carried on the pUT18C-da-g recombination plasmid.
[0100] To make the T4(+b-GT) phage, which is T4 phage carrying a higher-than- normal fraction of beta-glucosyl-hydroxymethylcytosine nucleobases, wild-type T4 was passaged through E. coli EC100 carrying the plasmid p(b-gt), which overexpresses T4 beta-glucosyltransferase (b-GT) under 1 mM isopropyl-p-D- thiogalactopyranoside (IPTG) induction. An overnight culture of E. coli EC100 / p(b-g ) was diluted 1 :50 in 10 mL LB medium supplemented with 50 g / mL spectinomycin and 1 mM IPTG. After approximately 1 hour 15 minutes of culture growth, ODeoo was measured for the culture and confirmed to be between 0.2-0.4. The 10 mL culture was then infected with 2 pL of wild-type T4 phage stock and grown overnight at 37°C with shaking. The next day, the tube was spun down at 15,000 x g for 10 minutes at 4°C. The phage-containing supernatant was filtered using an Acrodisc 13 mm SUPOR 0.45 pm syringe filter (Pall, 4604) into a 15 mL conical tube.
[0101] All final phage stocks were titered on top agar lawns of E. coli EC100 and stored at 4°C.
[0102] Plaque assays and efficiency of plaquing analysis
[0103] Overnight cultures were launched from single colonies in 3 mL of LB medium supplemented with appropriate antibiotic(s). Top agar lawns of E. coli were prepared by mixing 100 pL of overnight culture with 6 mL of LB top agar (LB broth Lennox base, 0.5% agar) supplemented with appropriate antibiotic(s). Top agar mixtures were poured over LB agar in 10 cm plates supplemented with appropriate antibiotic(s). Where necessary, 0.2% L-arabinose was included in the overnight media as well as in the LB top agar and the LB agar plate. Plates were dried at room temperature, partially open by a sterilizing flame, for 25 minutes for the top agar to solidify. Serial dilutions of phage stock were prepared and spotted on the top agar after drying. For imaging of plaque assays, 2.5 pL of each phage dilution was spotted on top agar using a multichannel pipette. For quantification of phage titers, efficiency of plaquing, and isolation of single phage plaques for phage DNA sequencing, 3-3.5 pL of each phage dilution was spotted on top agar using a multichannel pipette and the plate was tilted to allow phage spots to drip down the plate for easier quantification and isolation of single plaques. In all cases, plates were incubated at 37°C overnight after drying at room temperature for 25 minutes or until the plates were completely dry. Overnight plaque assays were imaged the next day (-16-24 hours after infection) using the FluorChem HD2 system (ProteinSimple). Plaque assay images were all auto-contrasted using Adobe Photoshop to give clearer images. In some cases, image brightness was enhanced further using Adobe Photoshop for better visualization of phage spots. Plaque assays with BASEL phages reported in Fig. 14e were performed in larger 15 cm plates of LB agar supplemented with 12.5 pg / mL chloramphenicol, to allow for plaquing of up to ten different phages on a single lawn. Here, the protocol was performed exactly as above, except with scaled up volumes: 300 uL of overnight culture was mixed with 15 mL of LB top agar supplemented with 12.5 pg / mL chloramphenicol. As before, 2.5 pL of each phage dilution was spotted on top agar using a multichannel pipette.
[0104] In Fig. 7i, efficiency of plaquing was quantified as the number of plaques formed by the phage on an E. coli EC100 / pBrig1 (targeting) lawn divided by the number of plaques formed by the same phage on an E. coli EC100 / pWEB-TNC (control) lawn. In Fig. 14c, to quantify phage plaques of T6 and T6 Aba-gt formed on E. coli EC100 / pBrig1 (targeting) lawns, infections were spread out across the entire top agar lawns to accurately count individual plaque-forming units (PFUs). To this end, 100 pL of phage stock normalized to -1 x 106PFU / pL (so -108PFUs total) was mixed with 100 uL of overnight culture and then mixed with 6 mL LB top agar (with 12.5 pg / mL chloramphenicol) and subsequently poured over an LB agar plate, supplemented with 12.5 pg / mL chloramphenicol. Top agar plates were incubated at 37°C overnight after drying at room temperature for 25 minutes. The next day, single plaques were counted across the entire top agar lawn. To accurately determine the total PFUs added of each phage, plaquing of serial dilutions of the -1 x 106PFU / pL normalized phage stocks was performed following the standard procedure of a plaque assay outlined above, using 3.5 pL drips of each phage dilution to facilitate more precise quantification of phage titers. Efficiency of plaquing was quantified as the total number of plaques formed by the phage across an entire E. coli EC100 / pBrig1 (targeting) lawn divided by the experimentally estimated total number of PFUs added.
[0105] Functional selection of a T4-resistant clone in the AZ52 soil DNA library in E. coli
[0106] The DNA library we used was generated in an earlier study using DNA extracted from an arid soil sample collected in Arizona. The library, AZ52, is comprised of large ~40 kb DNA fragments from soil microorganisms cloned into a pWEB-TNC cosmid. The insert-carrying cosmids were transformed into E. coli EC100 cells (Lucigen), generating a soil DNA library with approximately 20 million clones, divided into megapools carrying roughly 1.25 million clones each.
[0107] Each clone within the library houses a cosmid with a soil DNA insert, which carries genes from soil-derived microorganisms. Soil-derived genes can therefore be expressed heterologously in our library system. We performed our functional screen using the coliphage T4. To grow up libraries, we scraped frozen library stocks of E. coli EC100 carrying megapools 3-16 of the AZ52 DNA library into separate tubes with 10 mL LB supplemented with 12.5 pg / mL chloramphenicol and grew cultures overnight at 37°C with shaking. The next day, we infected E. coli EC100 overnight cultures with T4 at a multiplicity of infection (MOI) of 10, high enough to kill almost all clones without bona fide immunity. Infections were performed in 6 mL LB top agar with 500 pL of overnight stationary culture mixed with phage at MOI 10 on LB agar plates, supplemented with 12.5 pg / mL chloramphenicol. We incubated plates at 37°C for 36-48 hours and then inspected surviving colonies within top agar infections. We found that only megapool 4 showed an increased number of surviving colonies upon T4 infection compared to an infection of E. coli EC100 cells carrying an empty pWEB-TNC cosmid (control).
[0108] Since cells may survive T4 infection due to mutations within the E. coli host that prevent phage infection and not due to immunity genes carried within the soil DNA cosmids, we enriched for true immunity genes carried on cosmids. To eliminate false positive clones, we extracted pooled cosmid DNA from the surviving colonies on the enriched plate. To do this, we scraped top agar with surviving colonies into a 50 mL conical tube, melted the top agar in a 98°C heating block for 10-15 minutes until the top agar was completely melted, and then centrifuged the tube at -4000 x g for 5 minutes at room temperature to collect a cell pellet from which surviving cosmids were isolated using the QIAprep Spin Miniprep Kit (QIAGEN, Cat# 27106). The miniprepped cosmid pool was then transformed into 50 pL of electrocompetent E. coli EC100 cells (Lucigen) through electroporation (1 mm Bio-Rad Gene Pulser cuvette at 1 .8 kV) and recovered in 1 mL SOC medium. After 1 .5 hours of recovery, cells were assayed for transformation efficiency by pipetting ten-fold serial dilutions of the transformation culture on to an LB agar plate supplemented with 12.5 pg / mL chloramphenicol. While the plate was grown overnight at 37°C, the remaining transformation culture was stored overnight at 4°C. The next day, based on the calculated transformation efficiency, the transformation culture was spread onto ten 15 cm LB agar plates supplemented with 12.5 pg / mL chloramphenicol, plating for -30,000 colonies on each plate, for a total of -300,000 colonies. Plates were incubated overnight at 37°C and the next day colonies from all ten plates were scraped into 20 mL LB, vortexed and inverted to mix, and then diluted to ODeoo = 10. The ODeoo = 10 colony mixture was then mixed 1 : 1 with 50% glycerol to make a - 80°C freezer stock of a 1X phage-enriched DNA library for AZ52 megapool 4. This library was then grown up for re-infection with T4 and the steps described above were repeated two more times to generate a freezer stock of a 3X phage-enriched DNA library for AZ52 megapool 4.
[0109] We sampled colonies from the 3x-enriched library for anti-phage immunity by streaking the library to single colonies on an LB agar plate supplemented with 12.5 pg / mL chloramphenicol. Sixteen single colonies were grown overnight in LB supplemented with 12.5 pg / mL chloramphenicol at 37°C with shaking. Colonies were assayed for anti-phage immunity using plaque assays (described above) with phages Xv / r and T4. Of the sixteen colonies, twelve were found to carry immunity against T4 and none against ).vir. Cosmids were isolated from the twelve T4- resistant clones using the QIAprep Spin Miniprep Kit (QIAGEN, Cat# 27106) and sent for Sanger sequencing by Genewiz / Azenta using the universal primers T7 and M13F40, which flank the metagenomic DNA insert within the pWEB-TNC cosmid. Sequencing the T4-resistant cosmids revealed they all contained the same metagenomic DNA insert, suggesting that they all originated from the same T4- resistant library clone. One of the twelve T4-resistant clones was frozen at -80°C (900 pL culture + 100 pL DMSO) for use in future experiments.
[0110] Cosmid sequencing, assembly, and gene annotation
[0111] Cosmid DNA was extracted using the QIAprep Spin Miniprep Kit (QIAGEN, Cat# 27106). DNA was sequenced using the Nextera XT DNA Library Preparation Kit (Illumina, Cat# FC-131-1024). Paired-end 2 x 75 bp sequencing was conducted using the 150-cycle MiSeq Reagent Kit v3 (Illumina, Cat# MS-102-3001) on the Illumina MiSeq platform. Geneious Prime was used to assemble the cosmid genome, using the Geneious assembler (medium sensitivity / fast) on 100,000 paired- end DNA sequencing reads. The sequence of the cosmid harboring Brig 1 was deposited on GenBank, accession number OR880862. SnapGene was used to predict ORFs with ATG or GTG start codons (minimum length: 50 amino acids) within the metagenomic DNA insert of the assembled cosmid genome. Predicted ORFs were then run through NCBI PSI-BLAST blast. ncbi.nlm.nih.gov / Blast.cgi?PAGE=Proteins) and HHpred4950toolkit.tuebingen.mpg.de / tools / hhpred) to ascertain protein function where possible. Side-by-side genome annotation was also performed using the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) Genome Annotation Service (www.bv- brc.org / app / Annotation), with the annotation recipe for “Bacteria / Archaea” and the taxonomy name set to “Nocardiodes", taxonomy ID 1839. Defense genes and systems were identified using DefenseFinder ( / / defense-finder.mdmparis-lab.com / ) and the Prokaryotic Antiviral Defence LOCator (PADLOC) ( / / padloc. otago. ac.nz / padloc / ).
[0112] Subcloning of the T4-resistant cosmid to identify the T4 anti-phage system
[0113] To identify the immunity gene(s) in our cosmid, we subcloned four DNA fragments (A-D) that span the entire length of the metagenomic insert sequence. DNA fragments were amplified using 10 ng of cosmid DNA as template for PCR amplification using Phusion High-Fidelity DNA Polymerase (Thermo Scientific, Cat# F530L) with 1 M betaine (Sigma-Aldrich, Cat# B0300) and 1 pL DMSO in a 50 pL PCR reaction. Fragments were cloned into PCR-amplified pWEB-TNC cosmid backbones using NEBuilder® Hi Fi DNA Assembly Master Mix (NEB, Cat# E2621 L). NEBuilder® H i Fi DNA assembly was carried out at 50°C in a thermal cycler for 4 hours, and then 5 pL of the assembly reaction was transformed into 50 pL of chemically competent E. coli EC100 cells (Lucigen). Cells were incubated on ice for 30 minutes, heat shocked in a 42°C water bath for 30 seconds, placed back on ice for 2 minutes and then recovered in 250 pL SOC for 2 hours. Cells were then plated on LB agar supplemented with 12.5 pg / mL chloramphenicol and incubated overnight at 37°C. The next day, 8 colonies were picked, grown overnight in LB supplemented with 12.5 pg / mL chloramphenicol and their cosmids miniprepped the next day using the QIAprep Spin Miniprep Kit (QIAGEN, Cat# 27106). Miniprepped cosmids were sent for Sanger sequencing by Genewiz / Azenta using the universal primers T7 and M13F40, which flank the subcloned DNA fragment inserted into the pWEB-TNC cosmid backbone. Colonies that harbored cosmids with correct insert fragments were then assayed for immunity against phage T4 using plaque assays (see above). Plaque assays identified Fragment D as the fragment harboring anti-T4 immunity. Fragment D was further subdivided into Fragments D1 , D2 and D3, cloned and tested for immunity as described above. Fragment D3, containing a three-gene operon, was identified as the minimal DNA fragment carrying anti-T4 immunity. To determine the gene or genes responsible within the Fragment D3 operon, we generated six cosmid constructs (D3-1 to D3-6) containing different numbers and combinations of the three genes within the operon, each time being driven by the same promoter upstream of the first gene within the operon. These constructs were then tested using plaque assays to identify the gene within the operon that conveyed anti-T4 immunity.
[0114] NCBI blastn of the T4-resistant metagenomic DNA sequence
[0115] To identify possible organisms that our metagenomic DNA comes from, we performed a nucleotide BLAST on NCBI using the algorithm for somewhat similar sequences (blastn)
[0116] (blast. ncbi.nlm.nih.gov / Blast.cgi?PROGRAM=blastn&BLAST_SPEC=GeoBlast&PAG E_TYPE=BlastSearch. We performed blastn on the DNA sequences of Fragments C and D (see above and Fig. 6c).
[0117] T4 phage adsorption assay
[0118] Overnight cultures of E. coli EC100 cells carrying pWEB-TNC or pBrigl were diluted 1 :50 in 10 mL LB medium supplemented with 12.5 pg / mL chloramphenicol. After 1 hour 15 minutes of culture growth, ODeoo was measured for each culture and normalized to ODeoo = 0.3. Cultures were then infected with T4 at MOI 0.01 and incubated at 37°C with shaking for 50 minutes. A 10 mL bacteria-free, media-only control (LB + 12.5 |j.g / mL chloramphenicol) was mixed with the same volume of T4 and incubated alongside the cultures at 37°C with shaking for 50 minutes. Phage- infected cultures were sampled at the following time points: 0-, 10-, 20-, 30-, 40- and 50-minutes post infection. At each time point, 400 ptL was collected from each phage-infected culture and spun down in 1.5 mL Eppendorf tubes at 15,000 rpm for 2 minutes in a tabletop microcentrifuge at 4°C. Phage-containing supernatants were filtered using Acrodisc 13 mm SUPOR 0.45 pM syringe filters (Pall, 4604) into fresh 1.5 mL Eppendorf tubes. Filtered phage supernatants were used to prepare two sets of serial dilutions to estimate phage titers on fresh top agar lawns of E. coli K-12 MG1655. Top agar plates were incubated at 37°C. The next day, phage plaques were counted to determine phage titers at each timepoint. The experiment was repeated another two times in this manner for three independent biological replicates. qPCR of phage DNA replication
[0119] To quantify phage DNA replication within an infected E. co / / cell, overnight cultures of E. coli EC100 cells carrying pWEB-TNC or pBrigl were diluted 1 :50 in 50 mL of LB medium supplemented with 12.5 pg / mL chloramphenicol. After 1 hour 15 minutes of growth, ODeoo was measured, and the culture was normalized to ODeoo = 0.3. 700 pL of culture was dispensed between multiple 1 .5 mL Eppendorf tubes, corresponding to three replicates and multiple timepoints for each infection being monitored. These 700 pL cultures were infected with phage T4 at MOI 1 and incubated at 37°C with shaking for specified timepoints. At each timepoint, samples were removed from the incubator and tubes spun down at 15,000 rpm for 1 minute in a tabletop microcentrifuge at 4°C. Supernatants were removed and cell pellets immediately frozen down at -80°C for DNA extraction later. Additionally, 1-3 uninfected tubes for cells carrying pWEB-TNC or pBrig 1 were also prepared for DNA extraction as no-phage controls for qPCR.
[0120] Total DNA was extracted from frozen E. coli cell pellets using the Promega Wizard® Genomic DNA Purification Kit (Promega, Cat# A1125) following the protocol for Gram-negative bacteria. Extracted DNA was quantified using the Qubit™ dsDNA HS Assay Kit and each sample was normalized to 4 ng / pL. A total of 32 ng DNA was used as input for qPCR, performed using Fast SYBR Green Master Mix (Applied Biosystems, Cat# 4385612) and the QuantStudio® 3 Real-Time PCR System (Applied Biosystems) with primer pairs AA870 / AA871 (T4 gp43 target), AA872 / AA873 (T4 gp34 target) and AA387 / AA388 (E. coli K-12 MG 1655 dxs control). For qPCR data analysis, AACt values were calculated for the two T4 qPCR targets for each replicate at each timepoint. Fold-change values were then calculated for each replicate relative to the mean AACt value for cells carrying pWEB-TNC infected with T4 phage at the earliest timepoint post infection for a given experiment. The mean fold change of three biological replicates was plotted for each timepoint post infection.
[0121] Next-generation sequencing of phage DNA in T4-infected E. coli cells
[0122] Overnight cultures of E. coli EC100 cells carrying pWEB-TNC or pBrigl were diluted 1 :50 in 10 mL of LB medium supplemented with 12.5 pg / mL chloramphenicol. After 1 hour 15 minutes of growth, ODeoo was measured, and cultures were normalized to ODeoo - 0.3. Cultures were then infected at MOI 5 with T4 or T4 escaped for 8 minutes at 37°C with shaking, prior to centrifugation at 15,000 x g for 5 minutes at 4°C and subsequent freezing of cell pellets at -80°C. All cell pellets were stored at -80°C at least overnight, until ready for genomic DNA purification using the Promega Wizard Genomic DNA Purification Kit (Promega, Cat# A1125) following the protocol for Gram-negative bacteria. Purified genomic DNA was sheared using a pre-split snap-cap 6x16 mm Covaris microTUBE (Covaris, Cat# 520045) in a Covaris S220 focused-ultrasonicator and prepared for next generation sequencing using the Illumina TruSeq Nano DNA LT kit (Illumina, Cat# 20015964). Paired-end 2 x 75 bp sequencing was conducted using the 150-cycle MiSeq Reagent Kit v3 (Illumina, Cat# MS-102-3001 ) on the Illumina MiSeq platform. Illumina paired-end sequencing reads were aligned to phage genomes using a custom Python script, where the recorded number of phage-derived sequencing reads at a specific base pair position within the phage genome was normalized to the total sequencing reads for each sample.
[0123] Phage DNA extraction Phage genomic DNA was extracted from capsids using a previously described protocol55. Briefly, three tubes of 450 pL of a phage stock were first treated with DNase I (Invitrogen, Cat# 18068015) and RNase A (Promega, Cat# A7973) in DNase I buffer (20 mM Tris-HCI, pH 8, 2 mM MgCh), the reaction stopped with EDTA (Invitrogen, Cat# AM9260G), then capsids digested with Proteinase K (NEB, Cat# P8107S), and finally phage genomic DNA extracted using the DNeasy Blood & Tissue kit (QIAGEN, Cat# 69504). DNA was quantified using the Qubit™ dsDNA HS Assay Kit and assessed for quality using a nanodrop spectrophotometer.
[0124] T4 and T4 escaped genome sequencing and assembly
[0125] Phage genomic DNA was sequenced using the Nextera XT DNA Library Preparation Kit (Illumina, Cat# FC-131-1024). Paired-end 2 x 75 bp sequencing was conducted using the 150-cycle MiSeq Reagent Kit v3 (Illumina, Cat# MS-102-3001) on the Illumina MiSeq platform. Reads were quality-trimmed using Sickle (github.com / najoshi / sickle) and assembled into contigs using ABySS github.com / bcgsc / abyss). Finally, contigs were mapped to a reference phage T4 genome (GenBank: AF158101.6) using Medusa (combo.dbe.unifi.it / medusa). Automated genome annotation was performed using SnapGene and a reference phage T4 genome from NCBI (GenBank: AF158101 .6). Alignment of the T4 and T4 escaped genomes to the reference T4 genome revealed differential mutations between the two assembled phage genomes.
[0126] Sanger sequencing of bacteriophage escapers
[0127] T4 or T6 phage plaques on lawns of E. coli EC100 cells carrying pBrigl were isolated and resuspended in 20 pL of LB medium. Serial dilutions were prepared from the resuspended phage and spotted on a fresh LB top agar lawn of E. coli EC100 carrying pBrigl to maintain selection of the escaper phage. The plate was incubated at 37°C overnight. The next day a single phage plaque was picked from the top agar lawn using a P20 pipette set to 15 pL and resuspended in 20 pL of colony lysis buffer. Resuspended phage mixtures were boiled at 98°C for 15 minutes in a thermal cycler, and 1 pL of the boiled phage mixture was then used as template for PCR amplification using Phusion High-Fidelity DNA Polymerase (Thermo Scientific, Cat# F530L) with primers AA681 / AA682 to amplify T4 a-gt and primers AA1115 / AA1116 to amplify T6 a-gt. PCR products were submitted to Sanger sequencing by Genewiz / Azenta to identify mutations in a-gt. Wild-type T4 and T6 phage stocks were also PGR amplified at a-gt loci and sent for Sanger sequencing to provide reference sequences for comparison. Snapgene was used to align Sanger sequencing products of the escaper phages to wild-type a-gt sequences to identify escape mutations.
[0128] Brigl structural predictions using AlphaFold2
[0129] The structure of the intact (261 amino acid) Brigl protein was predicted using the colab implementation of AlphaFold2 (colab.research.google.com / github / sokrypton / ColabFold / blob / main / AlphaFold2.ipynb ) using default settings (except that the amber option was turned on to improve side chain rotamers). The highest ranked PDB structure produced by ColabFold (ptm = 0.86) was then visualized using PyMOL (The PyMOL Molecular Graphics System, Version 2.5.5, Schrodinger, LLC; www.pymol.org / pymol.html). Protein structure predictions of the Brigl homologs from Nocardiodes zhouii and Nocardiodes anomalus were performed in the same way. Cavities and pockets were visualized in PyMOL using default settings for surface calculation and the “cavities and pockets only” option for display.
[0130] Purification of Brigl
[0131] The brigl gene was recloned into the Nde1 and Xhol sites of pET21a using PCR primers that destroyed the Xhol site and added a Hiss tag immediately after the native C-terminal glycine of Brigl . The insert was verified by DNA sequencing. E. coli strain Rosetta(DE3)pLysS was used for protein expression. Cells were grown in LB medium with 100 pg / mL ampicillin at 37°C. 0.5 mM IPTG was added to induce protein expression when ODeoo~0.7, followed by further growth at 37°C for 2 hours. Cell pellets were resuspended in Ni column buffer A (50mM phosphate, 1 M NaCI, 5% glycerol, 1 mM DTT, pH 7.5) with complete mini protease inhibitor cocktail (Roche), one tablet / 1 L culture. After adding lysozyme to a final concentration of 200 mg / mL, the mixture was sonicated 3 times for 1 minute each, then centrifuged at 20,000 rpm in an SS-34 rotor for 1 hour. The supernatant was filtered and loaded onto a Ni column (Cytiva, HisTrap HP, Cat#17-5248-02), and eluted with a 30-minute gradient of O to 100% buffer B (Ni buffer A plus 500 mM imidazole, pH 7.5). Brigl - containing fractions were pooled and diluted with heparin column buffer A (25 mM MES, 0.5 mM EDTA, 5% glycerol, 1 mM DTT, pH 6), loaded on a heparin column (Cytiva, HiTrap Heparin HP, Cat# 17-0406-01 ) and eluted with a gradient from 10% to 70% heparin buffer B (heparin column buffer A + 2 M NaCI, pH 6) over 90 minutes. The purest fractions were pooled and concentrated, then dialyzed into storage buffer (20 mM Tris, 0.5 mM EDTA, 200 mM NaCI, 20% glycerol, 2 mM DTT, pH 8) and flash-frozen in small aliquots.
[0132] Purification of T4 alpha-glucosyltransferase
[0133] A pET15b derivative encoding N-terminally Hise tagged phage T4 alpha- glucosyltransferase (a-GT) was used for protein expression. Rosetta(DE3)pLysS cells harboring this plasmid were grown and induced as for Brigl , but after induction grown at 20°C overnight rather than for 2 hours at 37°C. The same purification protocol as for Brigl was followed except for a change in the pH of the heparin column buffers (A = 25 mM HEPES, 0.5 mM EDTA, 5% glycerol, 1 mM DTT, pH 7 and B = A + 2 M NaCI, pH 7). The purest fractions were pooled and concentrated, then dialyzed into storage buffer (20 mM Tris, 0.5 mM EDTA, 200 mM NaCI, 20% glycerol, 2 mM DTT, pH 8) and flash-frozen in small aliquots.
[0134] Purification of Brigl (Y121 A, E147A) mutant
[0135] The Brigl (Y121 A, E147A) mutant protein was purified according to a modified protocol. For consistency, wild-type Brigl was purified according to this same protocol, side-by-side, and this batch of purified Brigl protein was used only in experiments where Brigl (Y121 A, E147A) was used. Both Brigl and Brig1 (Y121A, E147A) were cloned into a pET21 a vector, with a Hise tag immediately after the native C-terminal glycine of Brigl . E. coli strain BL21 (DE3) was used for protein expression. Cells were grown in LB medium with 100 pg / mL ampicillin at 37°C overnight. The next day, a 1 : 100 dilution of the overnight was grown in 1 L of LB medium with 100 pg / mL ampicillin at 37°C for 3-4 hours. 0.5 mM IPTG was added to induce protein expression when ODeoo~0.7, followed by overnight growth (~16 hours) at 18°C. Cells were pelleted at 4500 rpm at 4°C (Eppendorf Centrifuge 5810 R) for 15 minutes. Cell pellets were resuspended in 20 mL of lysis buffer (50 mM HEPES, pH 7.7, 150 mM NaCI, 10% glycerol, 1 mM TCEP, 30 mM imidazole, 2 Roche mini protease inhibitor tabs EDTA free, 0.5 mg / mL lysozyme) and incubated on ice for 1 hour with shaking. The resuspended pellets were then sonicated using a Qsonica Q500 sonicator (70% amplitude with 10 seconds on, 30 seconds off for 2.5 minutes). The sonicated samples were spun down at 12,000 rpm at 4°C (Eppendorf Centrifuge 5810 R) for 30 minutes and the supernatant run through a gravity column loaded with 3 mL of HisPur Ni-NTA Resin (Thermo Scientific, Cat# 88222). Before passing supernatant, the column was equilibrated with equilibration buffer (50 mM HEPES, pH 7.7, 150 mM NaCI, 10% glycerol, 1 mM TCEP, 30 mM imidazole). Then, the ~20 mL of sonicated cell pellet supernatant was passed through the column. The column was washed twice with 25 mL wash buffer (50 mM HEPES, pH 7.7, 500 mM NaCI, 10% glycerol, 1 mM TCEP, 30 mM imidazole) and then eluted with 20 mL elution buffer (50 mM HEPES, pH 7.7, 150 mM NaCI, 10% glycerol, 1 mM TCEP, 300 mM imidazole). The eluted protein was concentrated to < 500 pL using an Amicon Ultra-4 Centrifugal Filter, 10 kDa MWCO (Millipore, Cat# UFC801024), with multiple rounds of centrifugation at 4300 x g for 10 minutes at 4°C (Eppendorf Centrifuge 5810 R), carefully resuspending the mixture between rounds of centrifugation via pipette mixing. The concentrated eluant was run on an AKTA pure™ chromatography system (Cytiva) fitted with a Superdex™ 75 Increase 10 / 300 GL column (Cytiva, Cat# 29148721) using storage buffer (50 mM HEPES, pH 7.7, 150 mM NaCI, 10% glycerol, 1 mM TCEP). Two peaks, corresponding to fractions 17-20 and 22-27, were collected and separately pooled. Pooled fractions were concentrated to < 500 pL using an Amicon Ultra-0.5 Centrifugal Filter, 10 kDa MWCO (Millipore, Cat# UFC501096), with multiple rounds of centrifugation at 13,000 x g for 5 minutes in a tabletop microcentrifuge at 4°C, resuspending the mixture between rounds of centrifugation via pipette mixing. For both Brigl and Brigl (Y121 A, E147A), the second peak (fractions 22-26 for Brigl and fractions 23-27 for the mutant) was determined to be free of nucleic acid contamination via nanodrop and found to contain pure protein (~29 kDa) by a Coomassie gel. Concentrated protein was flash- frozen in small aliquots and stored at -80°C for future use.
[0136] Annealing of ssDNA oligonucleotides
[0137] To generate dsDNA substrates for Mfel digestion and for DNA glycosylase assays, complementary ssDNA oligonucleotides were annealed. Briefly, 1 :1 molar ratios of top and bottom strand complementary ssDNA oligonucleotides (25-50 pM each) were mixed in a 60 pL reaction containing NaCI to a final concentration of 100 mM. The reaction was heated at 80°C for 20 minutes in a water bath or thermal cycler and then allowed to cool very slowly to room temperature. Annealed oligonucleotides were purified using an oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061) according to the manufacturer’s instructions.
[0138] Generation of glucosylated ssDNA and dsDNA oligonucleotides
[0139] We tested the activity of alpha-glucosyltransferase (a-GT) on both single- and double-stranded DNA. The ssDNA substrates were hmdC_18, hmdC_60_Mfel and hmdC_60_Mfel_Bot, which are 18mer and 60mer oligonucleotides, each containing a single hmC residue. The dsDNA substrate was hmdC_60_Mfel annealed to Bot_Mfel_60. Substrate DNAs (100 pM for ssDNA and 50 pM for dsDNA) were mixed at a 1 :1 molar ratio with a-GT in 1X NEBuffer 4 (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 1 mM DTT, pH 7.9) supplemented with 2 mM UDP-Glucose (NEB, supplied with NEB T4 beta-glucosyltransferase (b-GT)). All samples were incubated at 37°C overnight, then purified with an oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061) according to the manufacturer’s instructions. A subset of the ss- and dsDNA substrates were also treated with b-GT (NEB, Cat# M0357S) following the supplier’s instructions and purified as described above.
[0140] Modification by a-GT (or b-GT) was monitored by digestion with Mfel-HF (NEB, Cat# R3589S), which is blocked by the presence of glucosylated hmC but not by hmC (the modified C in hmdC_60_Mfel and in hmdC_60_Mfel_Bot is within an Mfel site. Before digestion, single-stranded a-GT- or b-GT-treated hmdC_60_Mfel oligonucleotides were annealed to Bot_Mfel_60 or to untreated or a-GT-treated hmdC_60_Mfel_Bot. Approximately 1.5 pg (Fig. 8e) or 500 ng (Fig. 10c) of each sample was digested with Mfel-HF (NEB, Cat# R3589S) for 1 hour, then electrophoresed on a 10% TBE gel (Invitrogen, Cat# EC6275BOX, 10-well or Cat# EC62755BOX, 15-well) at 140 V for 35 minutes. Gels were stained with 2 pg / mL ethidium bromide for 20 minutes, extensively rinsed with distilled water (3X for 10 minutes each), and then scanned using a ChemiDoc MP imager (BioRad) set to UV trans illumination and the machine’s 605 / 50 filter to detect ethidium bromide (Fig. 8e) or using the Amersham ImageQuant 800 set to UV fluorescence (Fig. 10c). DNA glycosylase assays with ssDNA oligonucleotides: detection of the abasic site with an aldehyde-reactive probe
[0141] We used an aldehyde-reactive fluorescent probe, AZDye 488 Hydroxylamine, (fluoroprobes.com) to detect removal of a base from the phosphodiester backbone in the absence of DNA cleavage. The dye was dissolved in distilled water to form a 10 pg / pL stock solution.
[0142] DNA glycosylase reactions were carried out in a reaction buffer containing 45 mM HEPES, pH 7.5, 0.4 mM EDTA, 2% glycerol, 1 mM DTT and 50 mM KCI, in a total reaction volume of 50 pL. The final DNA concentrations were 2 pM. Brig 1 was added to single-stranded a-GT- and b-GT-treated hmdC_60_Mfel to a final concentration of 35 pM, while 2 pL (10 units; 5 units / pL) of hSMUGI (NEB, Cat# M0336S) was added to dll_60 as a positive control. Reactions were incubated overnight at 37°C, after which 2 pL AZDye 488 dye was added, followed by incubation at 37°C for 30 minutes. 1 / 10 volume of 10% SDS was then added and incubated for another 30 minutes and purified by phenol / chloroform extraction. Samples were then treated with an oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061 ) according to the manufacturer’s instructions, eluted with 15 pL nuclease-free water, mixed with loading dye and electrophoresed for 45 minutes at 180 V on a 10% TBE gel (Invitrogen, Cat# EC62755BOX). The gel was stained with 2 pg / mL ethidium bromide for 20 minutes, extensively rinsed with distilled water (3X for 10 minutes each), then scanned using a ChemiDoc MP imager (BioRad) set to UV trans illumination and the machine’s 605 / 50 filter to detect ethidium bromide and then using blue epi illumination with the 530 / 28 filter for the AZDye 488 fluorescent probe.
[0143] DNA glycosylase assays with ssDNA and dsDNA oligonucleotides: detection by NaOH- or endonuclease IV-mediated cleavage of the abasic site
[0144] DNA glycosylase reactions were carried out in a reaction buffer containing 45 mM HEPES, pH 7.5, 0.4 mM EDTA, 2% glycerol, 1 mM DTT and 50 mM KCI, in a total reaction volume of 50 pL. The final ssDNA or dsDNA concentrations were 1 pM. Brigl or Brig1 (Y121A, E147A) was added to a final concentration of 1 pM (unless stated otherwise, e.g. range of 50-1600 nM in Fig. 12b), while 1 pL (5 units) of hSMUGI (NEB, Cat# M0336S) was added as a positive control. Reactions were incubated at 37°C overnight, unless stated otherwise (e.g. 30 minutes in Fig. 12b). Following enzymatic incubation, one set of samples was directly processed with an oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061) according to the manufacturer’s instructions. A second matched set of samples was treated with NaOH before cleanup: 25 pL of 0.5 M NaOH was added to each 50 pL sample and then heated at 90°C for 30 minutes before purification with the oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061) according to the manufacturer’s instructions. All samples were eluted from the cleanup columns in 15 pL nuclease-free water. 5 pL of each was mixed with loading dye and loaded onto a 10% TBE gel (Invitrogen, Cat# EC62755BOX) and electrophoresed at 140 V for 35 minutes. Gels were stained with 2 pg / mL ethidium bromide for 20 minutes, extensively rinsed with distilled water (3X for 10 minutes each), and then scanned using a ChemiDoc MP imager (BioRad) set to UV trans illumination and the machine’s 605 / 50 filter to detect ethidium bromide (Fig. 8h) or using the Amersham ImageQuant 800 set to UV fluorescence. For Urea-PAGE gels, eluted samples were first denatured by mixing 5 pL of purified sample with 5 pL of 2X TBE Urea Sample Buffer (Invitrogen, Cat# LC6876) and then heated at 70°C for 3 minutes. Denatured samples were loaded onto a 6% TBE-Urea gel (Invitrogen, Cat# EC68655BOX) and electrophoresed at 140 V for 35 minutes. Gels were soaked in ethidium bromide and rinsed with distilled water as described above, before imaging with the Amersham ImageQuant 800 set to UV fluorescence. For all gels, DNA ladders were made by mixing 20-, 40- and 60-bp ssDNA or dsDNA oligonucleotides and loading them onto their corresponding gels at -100 ng each oligonucleotide per load.
[0145] For abasic site detection by NEB endonuclease IV (Endo IV), DNA glycosylase reactions were set up as described above and incubated with Brig 1 overnight. Three matched sets of reactions were set up. After overnight incubation, one matched set of samples was treated with NaOH as described above and purified using the Zymo Research Oligo Clean & Concentrator Kit (Cat# D4061) according to the manufacturer’s instructions. The remaining two matched sets of samples were processed directly using the oligonucleotide cleanup kit. The purified samples were then incubated at 37°C for 4 hours in a 50 pL reaction with 1X NEBuffer 3 (100 mM NaCI, 50 mM Tris-HCI, 10 mM MgCh, 1 mM DTT, pH 7.9), with or without 50 units of NEB Endo IV (5 pL; 10 units / pL; Cat# M0304S). After 4 hours, reactions were purified using the Zymo Research Oligo Clean & Concentrator Kit according to the manufacturer’s instructions. All the purified samples were then loaded onto a 10% TBE gel (Invitrogen, Cat# EC62755BOX), electrophoresed at 140 V for 35 minutes, stained with ethidium bromide as described above and imaged with the Amersham ImageQuant 800 set to UV fluorescence.
[0146] High resolution mass spectrometry of hSMUGI- and Brigl -treated ssDNA oligonucleotides
[0147] DNA glycosylase reactions were carried out in a reaction buffer containing 45 mM HEPES, pH 7.5, 0.4 mM EDTA, 2% glycerol, 1 mM DTT and 50 mM KCI, in a total reaction volume of 50 pL. Reactions were performed with 18mer ssDNA oligonucleotides: dll_18, hmdC_18 and a-GT treated hmdC_18. The final ssDNA concentration in each reaction was 2 pM. Brigl was added to a final concentration of 2 pM, while 2 pL (10 units) of hSMUGI (NEB, Cat# M0336S) was added as a positive control. A no-enzyme reaction was used as a negative control. 2 x 50 pL reactions were set up for each reaction condition with the dU_18 oligonucleotide, while 8 x 50 pL reactions were set up for each reaction condition with hmdC_18 and a-GT treated hmdC_18. Reactions were incubated overnight at 37°C. After overnight incubation, all matched samples were pooled and processed with an oligonucleotide cleanup kit (Zymo Research, Oligo Clean & Concentrator Kit, Cat# D4061) according to the manufacturer’s instructions.
[0148] For mass spectrometry, purified oligonucleotide samples were dried using vacuum centrifugation and dissolved in 50 / 50 water / acetonitrile with 0.001 % triethylammonium bicarbonate. The pH of the solution was found to be comparable to that of deionized water. The samples were introduced to the mass spectrometer by manual injection using a Hamilton syringe applying pressure by hand at approximately 10 pL / min. Samples were analyzed using an orbitrap Ascend tribrid mass spectrometer (Thermo Scientific) operating in negative mode. Spectra were recorded in the mass range 600-1300 m / z at 120,000 resolution. A blank injection was introduced after each sample to eliminate carryover. Raw data was inspected using the Xcalibur Quality Browser (Thermo Scientific) and spectra were summed as necessary to provide representative spectra with a sufficient signal-to-noise ratio (S / N). Spectra were further processed using UniDec deconvolution software with the following parameters: sampling resolution and peak FWHM were both set to 0.1 , adduct mass was defined as -1.007276 Da, and charge states were defined 4-12 based on observations from the raw data. The m / z range was adjusted to fit the data and to exclude singly charged noise. Apart from the mass of the oligonucleotides, additional masses from metal adducts were also observed.
[0149] DNA glycosylase assays with phage and cosmid DNA
[0150] All reactions were performed in 50 pL reaction volumes in a reaction buffer containing 45 mM HEPES, pH 7.5, 0.4 mM EDTA, 2% glycerol, 1 mM DTT and 50 mM KCI. Assays were performed by incubating 50-500 ng of extracted phage genomic DNA from capsids or miniprepped pWEB-TNC cosmid DNA with varying concentrations (2-800 nM) of purified Brigl or Brigl (Y121 A, E147A) or with 10 units of NEB hSMUGI (NEB, Cat# M0336S) as a negative control. Reactions were incubated in a thermal cycler at 37°C for 30 minutes (or at 37°C for 30 minutes plus an additional 20 minutes at 65°C in Fig. 11 c, to cleave DNA at abasic sites and denature the glycosylase prior to gel electrophoresis). Reactions were then mixed with 10 pL of purple 6X loading dye with no SDS (NEB, Cat #B7025S) and the entire reaction volume was loaded onto a 1 % agarose gel containing ethidium bromide. Unless stated otherwise, the gel was run for 70 mins at 85 V at room temperature and then imaged using a UV gel imager (Amersham ImageQuant 800 set to UV fluorescence). Where SDS was used for protein denaturation (Fig. 11 d), all steps were carried out as described above except, before gel loading, a purple 6X gel loading dye containing a final 1X concentration of 0.08% SDS (NEB, Cat #B7024S) was used instead of loading dye without SDS.
[0151] For the gel in Fig. 11a, b, samples were loaded onto the 1% agarose gel in a cold room at 4°C and run at 40 V for 3 hours and then imaged using a UV gel imager (Amersham ImageQuant 800 set to UV fluorescence). The same gel was then allowed to equilibrate to room temperature for 30 minutes and then run longer by electrophoresis, this time under high voltage and at room temperature (in a room temperature gel box), first at 85 V for 25 minutes, then at 150 V for 8 minutes and finally at 200 V for another 8 minutes before final imaging using the Amersham ImageQuant 800 set to UV fluorescence.
[0152] Brigl multiple sequence alignment and phylogenetic tree construction
[0153] Brigl homologs were obtained using the NCBI PSI-BLAST protein homology search. Homologs were then subjected to a multiple sequence alignment (MSA) using MUSCLE v5 with 16 maximum iterations via the Geneious Prime software. A tree was built with the alignment output file via IQ-TREE 1.6.12 using the LG4M model with 1000 bootstrap alignments. The online tool ITOL was used for visualization of the resulting tree.
[0154] Brigl gene neighborhood analysis
[0155] Gene neighborhoods of the Brigl homologs from above (10 genes upstream and 10 genes downstream of each homolog) were constructed using a custom Python script. In brief, the script parses a blastp result XML file for accession numbers of each of the hits. For each hit accession, the script obtains the corresponding nucleotide accession from which the protein accession is derived. Finally, all annotated features within the nucleotide accession that are labelled as ‘CDS’ or ‘tRNA’ are built into a list, including their position within the nucleotide entry and their feature name. From this list, neighbors of the initial protein hit (10 genes upstream and 10 genes downstream) are extracted and built into a TSV file for subsequent analysis.
[0156] Statistical analysis
[0157] Statistical analyses were performed using GraphPad Prism version 10.1.0. Error bars and number of replicates for each phage experiment are defined in the figure legends. Statistical significance in Fig. 7i was determined using a two-tailed Student’s t-test (an unpaired parametric test assuming Gaussian distribution and that both populations have the same standard deviation).
[0158] The described eDNA library screen yielded a prokaryotic defense island containing a DNA glycosylase, Brigl , that provides immunity in E. coli against T- even bacteriophages by excising alpha-glucosyl-hmC nucleobases in the viral DNA. Brig 1 did not restrict the propagation of T4 Aa-gt nor Bas46-47 phages, whose genomes lack alpha-glucosyl-hmC and instead contain beta-glucosyl-hmC and arabinosyl-hmC nucleobases, respectively. The enzyme also did not degrade T6 phage DNA, which primarily harbors gentiobiosyl-hmC. Therefore, it is conceivable that T-even phages have diversified their hmC modification patterns to avoid restriction by DNA glycosylases involved in anti-phage defense such as Brig 1 . If so, the arms race between phages that modify their DNA and their hosts most likely resulted in the evolution of a larger family of Brig DNA glycosylases with activity against different hmC nucleobases, which probably includes some of the Brigl homologs that we found associated with other defense genes but failed to provide immunity against T4 and T6 (Fig. 15a).
[0159] The described assays with oligonucleotide substrates showed that the phosphate backbone of ssDNA and dsDNA substrates containing a single alpha- glucosyl-hmC nucleobase or dsDNA substrates containing two alpha-glucosyl-hmC nucleobases on opposite strands remained largely uncleaved despite overnight incubation with large amounts of Brigl protein (Fig. 3 and Fig. 10). The disclosure therefore provides a model that Brigl is primarily a monofunctional DNA glycosylase rather than a bifunctional glycosylase-lyase that would also nick the DNA phosphate backbone upon base excision. Without intending to be constrained by any particular theory, it is considered that Brigl activity disrupts the lytic cycle of T-even viruses by generating abasic sites throughout the viral genome that: (i) impede phage transcription and / or replication; (ii) lead to spontaneous hydrolysis of the phosphate backbone at the highly reactive abasic sites; and / or (iii) result in DNA interstrand crosslinks and DNA-protein crosslinks due to abasic site reactivity. In addition, given that Brigl can target ssDNA substrates, it would be possible for this enzyme to attack ssDNA intermediates that form during rolling-circle replication of T-even phages. We observed weak base excision activity after treating a uracil-containing ssDNA oligonucleotide with high concentrations of Brigl (Fig. 8h and Fig. 9c). This weak secondary activity on uracil suggests Brigl has likely evolved from members of the uracil DNA glycosylase superfamily. Given that there is no evidence for the misincorporation of alpha-glucosyl- hmC into bacterial DNA in the absence of phage infection, it is unlikely that Brig 1 would participate in a host base excision repair pathway dedicated to the removal of these nucleobases. On the contrary, and again without being constrained by any particular interpretation, it is considered that Brigl is a bona fide antiviral effector. Supporting this idea, the a-gt gene, responsible for the generation of alpha-glucosyl- hmC nucleobases, is widespread across phages infecting diverse hosts, and we found that Brigl provided immunity against 11 / 69 phages from the BASEL collection (Fig. 14e). Furthermore, many Brigl homologs are part of anti-phage defense islands, frequently associated with BREX / Pgl genes, but also with toxin-antitoxin cassettes, restriction endonucleases and CRISPR-Cas systems (Fig. 15a-c). In these genetic contexts, Brigl provides an additional layer of immunity against phages containing alpha-glucosyl-hmC that cannot be targeted by other systems present in the anti-phage defense islands, such as CRISPR-Cas and restriction endonucleases and possibly BREX immunity. Although the exact mechanism used by BREX systems to restrict phage infection is unknown, it was recently shown that wild-type T4, but not an a-gt / b-gt double mutant phage that lacks hmC glucosylation, evades BREX immunity in E. coli HS, suggesting that BREX systems are inhibited by glucosylated nucleobases. In hosts harboring defense islands with restriction enzymes, CRISPR or BREX in addition to Brigl , phages targeted by the latter will not be able to escape through mutations in a-gt that eliminate alpha-glucosylation of the viral DNA, since they will become susceptible to the defense mechanisms that target non-glucosylated DNA.
Claims
Claims:1 . A method comprising contacting an isolated DNA sample sequentially with T4 alpha-glucosyltransferase (a-GT) and subsequently with a Brig 1 protein, said Brig 1 protein comprising an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:1 .
2. The method of claim 1 , wherein abasic sites in the DNA are produced by the Brigl protein.
3. The method of claim 2, wherein the a-GT protein comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of SEQ ID NO:3.
4. The method of claim 3, further comprising determining the presence and / or location of the abasic sights within the isolated DNA sample, thereby determining the presence and / or the location of 5-hydroxymethylcytosine (5hmC) in the isolated DNA sample before being contacted with the a-GT and the Brigl protein.
5. The method of claim 4, wherein the isolated DNA sample comprises single-stranded DNA sample.
6. The method of claim 4, wherein the isolated DNA sample comprises double-stranded DNA.
7. The method of claim 6, wherein the double-stranded DNA comprises genomic DNA.
8. An isolated Brigl protein comprising an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 1 , and wherein optionally the isolated Brigl protein is a component of a fusion protein.
9. The isolated Brigl protein of claim 8, wherein the isolated Brigl protein is a component of the fusion protein, and wherein the fusion protein further comprises a T4 alpha-glucosyltransferase (a-GT) protein comprising an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:3.
10. A kit comprising a sealed container and an isolated Brigl protein comprising an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:1.11 . The kit of claim of 10, further comprising an isolated T4 alphaglucosyltransferase (a-GT) comprises an amino acid sequence that is at least 90% identical to the amino acid sequence of SEQ ID NO:3.
12. The kit of any one of claims 10-11 , further comprising one or more sealed containers that contain reagents for use in determining the presence and / or location of 5-hydroxymethylcytosine nucleobases in a DNA sample.
13. The kit of claim 12, further comprising printed material providing an indication that the Brigl protein is used for determining the presence and / or the location of 5-hydroxymethylcytosine (5hmC) within an isolated DNA sample.
Citation Information
Patent Citations
Methods and kits for detection of 5-hydroxymethylcytosine
US20130323728A1
Method for genomic profiling of DNA 5-methylcytosine and 5-hydroxymethylcytosine
US20170253924A1
SWI / SNF family chromatin remodeling complexes and uses thereof
US20230035235A1
Directed editing of cellular RNA via nuclear delivery of crispr / CAS9
US20230053915A1
Compositions and methods related to modification and detection of pseudouridine and 5-hydroxymethylcytosine
WO2022232795A1