Compositions and methods for characterizing RNA
The development of sequence-specific endoribonucleases with defined cleavage properties addresses the limitations of current RNA analysis methods, enabling controlled fragmentation and efficient analysis of RNA molecules, including therapeutic and functional RNA species.
Patent Information
- Application Number
- PCT/US2025/023483
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-05
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-09
AI Technical Summary
Current methods for analyzing RNA using ribonucleases result in uncontrollable and unpredictable digestion of RNA fragments that are too short for analytical tools, and there is a lack of robust methods to discover and produce highly pure, nucleotide- or sequence-specific endoribonucleases that can cleave double-stranded RNA, DNA/RNA hybrids, or structurally rigid or modified regions effectively.
Development of nucleotide- or sequence-specific endoribonucleases with desired recognition motifs and biophysical properties for controllable fragmentation of RNA molecules, including endoribonucleases with amino acid sequences similar to SEQ ID NOS: 1-139, 163, and 164, which can cleave single-stranded RNA selectively and produce defined cleavage products.
The endoribonucleases enable controlled fragmentation of RNA molecules, allowing for efficient analysis through methods such as sequencing, electrophoretic analysis, and spectroscopy, and can produce cleavage products suitable for further analysis and assembly into therapeutic or functional RNA species.
Smart Images

Figure US2025023483_09102025_PF_FP_ABST
Abstract
Description
[0001] COMPOSITIONS AND METHODS FOR CHARACTERIZING RNA
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application is a continuation in part of U.S. Application No. 18 / 969,910 filed December 5, 2024. This application also claims priority to U.S. Provisional Application No. 63 / 666,492 filed July 1, 2024 and U.S. Provisional Application No. 63 / 575,188 filed April 5, 2024. The contents of all of the above are hereby incorporated in their entirety by reference.
[0004] SEQUENCE LISTING STATEMENT
[0005] This disclosure includes a Sequence Listing submitted electronically in .xml format under the file name “NEB-483-484-489. xml” created on April 7, 2024, and having a size of 232 KB. This Sequence Listing is incorporated herein in its entirety by this reference.
[0006] BACKGROUND
[0007] Ribonucleases (RNases) are found in all domains of life and play an essential role in a wide range of physiological functions including RNA maturation, turnover, degradation, and anti-phage defense. Based on their enzymatic activity, they can be classified as exoribonucleases and endoribonucleases. Endoribonucleases can be sequence-specific (e.g., E. coli MazF cleaving at ACA), or non-specific (e.g., E. coli RNase I). Using ribonucleases that lack sequence and nucleotide specificity results in uncontrollable and / or unpredictable digestion of RNA into fragments that are too short for current analytical tools, such as mass spectrometry analysis, next generation sequencing (e.g., Illumina or Nanopore sequencing), and / or electrophoretic fragment analysis. Few endoribonucleases with discrete recognition motifs and defined cleavage specificities are known and a much smaller number of them are commercially available. A lack of robust methods to discover, fully characterize, and produce highly pure (e.g., contamination free) nucleotide and sequence specific endoribonucleases have contributed to their limited availability. Moreover, most known nucleotide- or sequencespecific endoribonucleases show limited cleavage of double-stranded RNA, DNA / RNA hybrids, structurally rigid or modified regions of RNA molecules. For example, the widely used RNase T1 enzyme has been shown to cleave any given tRNA only at bulges, loops, or other structured regions, leaving many of its predictive cleavage sites (‘GN’ sites) intact (i.e., uncleaved).
[0008] SUMMARY
[0009] Accordingly, needs have arisen for improved ribonuclease tools for analyzing RNA. The present disclosure relates, in some embodiments, to compositions and methods for characterizing RNA. For example, compositions and / or methods may include nucleotide-or sequence-specific endoribonucleases that have desired recognition motifs and / or biophysical properties (e.g., resistance to denaturing reagents and / or thermal stability) for controllable fragmentation of RNA molecules. The present disclosure relates, in some embodiments, to endoribonucleases targeting distinct sequence and / or structure motifs. Example endoribonucleases include endoribonucleases comprising an amino acid sequence according to one of SEQ ID NOS: 1-139, 163, and 164 and variants having > 95%, > 96%, > 97%, > 98%, or > 99% identity thereto.
[0010] The present disclosure relates, in some embodiments, to methods for cleaving a singlestranded RNA. SCleavage may be selective, for example, in that a single-stranded substrate having an identical sequence to a duplex substrate at the same enzyme to substrate ratio and like conditions will result in more cleavage product (more complete digestion) than the duplex substrate. In some embodiments, a method may comprise contacting a single-stranded RNA and an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 to produce cleavage products. A single-stranded RNA may comprise at least one endoribonuclease recognition sequence that corresponds to the endoribonuclease. In some embodiments, a method may comprise contacting a single-stranded RNA comprising at least one endoribonuclease recognition sequence that corresponds to the endoribonuclease and contacting further comprises contacting the endoribonuclease with the at least one recognition sequence to produce the cleavage products. Cleavage products may comprise one or more fragments of the single-stranded RNA. A single-stranded RNA may be a long RNA (e.g., at least 5000 nucleotides) and / or may comprise multiple (e.g., >10) sequences of interest separated by endoribonuclease recognition sequences (which may be referred to as polyci stronic).
[0011] According to some embodiments, a single-stranded RNA may comprise one or more modified nucleotides (e.g., anywhere in the nucleotide sequence including within the recognition sequence). Examples of modified nucleotides include N6-methyl-adenosine (m6A), 1-methyl-adenosine (nFA), 5-methyl-cytidine (m5C), 5-hydroxymethyl-cytidine (hm5C),N4- acetyl-cytidine (ac4C), 5 -methoxy cytidine (mo5C), 4-thiouridine (S4U), 2-thiouridine (S2U), pseudouridine ( ), NCmethyl-pseudouridine (nE ), 5-methyluridine (m5U), or 5- methoxyuridine (mo5U). Examples of single-stranded RNA include RNA molecules selected from (or RNA comprising RNA molecules selected from a messenger RNA (mRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), a small RNA (sRNA), a microRNA (miRNA), a long noncoding RNA (IncRNA), a circular RNA (circRNA), a transfer RNA (tRNA), an aptamer RNA, an antisense RNA, a silencing RNA (siRNA), a guide RNA (gRNA), or a therapeutic RNA. In some embodiments, a method may comprise contacting an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 and a polycistronic single-stranded RNA comprising at least ten endoribonuclease recognition sequences that each corresponds to the endoribonuclease to produce cleavage products. In some embodiments, contacting further comprises contacting the single-stranded RNA with a second endoribonuclease (e.g., an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 other than the amino acid sequence of the first endoribonuclease) and wherein the single-stranded RNA further comprises at least one endoribonuclease recognition sequence that corresponds to the second endoribonuclease. Cleavage products (e.g., cleavage products of a polycistronic template) may comprise at least one IVT template. In some embodiments, a method comprises contacting a single-stranded RNA with a first endoribonuclease and further comprises contacting the singlestranded RNA with a second endoribonuclease (e.g., concurrently) wherein the single-stranded RNA comprises at least one endoribonuclease recognition sequence that corresponds to each of the first and second endoribonucleases.
[0012] An endoribonuclease, according to some embodiments, comprises a sequence-specific endoribonuclease. In some embodiments, a method may further comprise analyzing at least one cleavage product. Analyzing a cleavage product may comprise analyzing the cleavage product by sequencing (e.g., NGS), electrophoretic analysis, photometric, NGS analysis, and / or spectroscopy (e.g., tandem mass spectrometry).
[0013] The present disclosure further relates to endoribonucleases and compositions thereof. For example, an isolated and purified endoribonuclease may have an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164. A composition may comprise a single-stranded RNA (e.g., a substrate RNA) and an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164. A composition may comprise, in some embodiments, a reaction buffer. A composition may comprise at least one additional endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 other than the amino acid sequence of the (first) endoribonuclease. The disclosure also relates to kits. In some embodiments, a kit may comprise a first endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164; and a second endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 other than the first endoribonuclease. A composition may comprise an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 and a buffering agent selected from HEPES, MES, MOPS, TAPS, tricine, Tris, ACES, ADA, BES, Bicine, CAPS, CHES, DIPSO, EPPS, MOPSO, PIPES, POPSO, TAPS, and TAPSO.
[0014] According to some embodiments, methods and compositions of the disclosure may include substrates and / or ligation products that include linear RNA and / or circular RNA.
[0015] The present disclosure provides, in some embodiments, methods for forming RNA cleavage products comprising a single-stranded therapeutic RNA. For example, a method may comprise contacting a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164; and a substrate RNA comprising, in a 5’ to 3’ direction, a therapeutic RNA, a 3’ untranslated region (3’ UTR), and a polyA tail, wherein at least one of the 3’ UTR and the polyA tail comprises at least one sequence-specific endoribonuclease recognition motif that corresponds to the endoribonuclease, to produce cleavage products comprising the single-stranded therapeutic RNA. The cleavage products may further include fragments of the single-stranded RNA (in addition to the single-stranded therapeutic RNA). In some embodiments, a substrate RNA may comprise a linear RNA, a circular RNA, or a mixture of linear RNA and circular RNA. In some embodiments, a singlestranded therapeutic RNA comprises at least 5000 nucleotides (e.g., at least 2000 nucleotides, at least 5000 nucleotides, at least 7500 nucleotides, at least 10000 nucleotides, or at least 15000 nucleotides), wherein the substrate RNA comprises at least one (e.g., at least 2, at least 3, at least 5) recognition sequence for the endonuclease. In some embodiments, a substrate RNA and / or a single-stranded therapeutic RNA may comprise one or more modified nucleotides (e.g., a naturally occurring modified nucleotide, an artificial modified nucleotide). Each modified nucleotide may be independently located at any position within the substrate RNA and / or therapeutic RNA. For example, each or the therapeutic RNA, the 3’ UTR, the polyA tail and / or the recognition motif may comprise one or more modified nucleotides. Examples of modified nucleotides include N6-methyl-adenosine (m6A), 1-methyl-adenosine (nfA), 5- methyl-cytidine (m5C), 5-hydroxymethyl-cytidine (hm5C),N4-acetyl-cytidine (ac4C), 5- methoxycytidine (mo5C), 4-thiouridine (S4U), 2-thiouridine (S2U), pseudouridine ( ), N1- methyl-pseudouridine (m1^), 5-methyluridine (m5U), or 5- methoxyuridine (mo5U). In some embodiments, a method may further include contacting the substrate RNA, the endoribonuclease, and at least one additional endoribonuclease, wherein each of the at least one additional endoribonucleases has an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1- 139, 163, and 164 other than the amino acid sequence of the first endoribonuclease.
[0016] The present disclosure provides, according to some embodiments, a method comprising contacting a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164; and a substrate RNA comprising a nucleotide sequence of interest heterologous to the endoribonuclease, a second nucleotide sequence, and at least one sequence-specific endoribonuclease recognition motif that corresponds to the sequencespecific endoribonuclease, wherein the recognition motif is disposed between the sequence of interest and the second nucleotide sequence, to produce cleavage products comprising the sequence of interest and the second sequence. A sequence of interest may comprise or consist of an in vitro transcription (IVT) template or a therapeutic RNA. In the case of an IVT template, the sequence of interest may be an RNA that is operable as an IVT template itself or it may be a precursor RNA from which an IVT template may be prepared (e.g., by copying the RNA into cDNA). In some embodiments, a substrate RNA may be, comprise or consist of a messenger RNA and the second sequence may be, comprise or consist of at least a portion of a 3’ untranslated region (3’ UTR) or at least a portion of a polyA tail.
[0017] The present disclosure further provides a method of analyzing a capped RNA. For example, a method may comprise contacting a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164; and a capped RNA, wherein the capped RNA comprises at least one sequence-specific endoribonuclease recognition motif cleavable by the endoribonuclease, to produce cleavage products, at least one of which comprises the cap; and analyzing the at least one product comprising the cap by gel fractionation, liquid chromatography, spectroscopy (e.g., tandem mass spectrometry), and / or nucleotide sequencing.
[0018] In some embodiments, the present disclosure provides methods comprising contacting an RNA polymerase and a polynucleotide template, the polynucleotide template comprising, in a 5’ to 3’ direction, an expression control sequence corresponding to the polymerase and a sequence of interest (e.g., a coding sequence) to produce transcription products, the transcription products comprising, in a 5’ to 3’ direction, a sequence complementary to the sequence of interest, a 3’ untranslated sequence, and a polyA tail, wherein at least one of the 3’ untranslated sequence and the polyA tail comprises at least one sequence-specific endoribonuclease recognition motif; and contacting the transcription products and a sequencespecific endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164, wherein the endoribonuclease is operable to cleave the transcription products comprising the endoribonuclease recognition motif to produce cleavage products, wherein the cleavage products comprising the sequence complementary to the sequence of interest are homogenously sized. In some embodiments, a cleavage product comprising the sequence complementary to the sequence of interest may be, comprise, or consist of a therapeutic RNA, a guide RNA, a tRNA, an mRNA, a rRNA, a miRNA, a IncRNA, a piRNA, a snRNA, a snoRNA, a chemically synthesized RNA. In some embodiments, a cleavage product comprising the sequence complementary to the sequence of interest may be, comprise, or consist of a cellular RNA. According to some embodiments, a method may further comprise contacting a cleavage product comprising the sequence complementary to the sequence of interest, a ligatable polynucleotide (which may be optional in some embodiments), and a ligase to produce a ligation product comprising the cleavage product comprising the sequence complementary to the sequence of interest and the ligatable polynucleotide (if included). For example, a method may comprise contacting a cleavage product comprising the sequence complementary to the sequence of interest and a ligase (without the ligatable polynucleotide) to produce a circular ligation product (comprising the cleavage product comprising the sequence complementary to the sequence of interest). A ligatable polynucleotide may comprise a 3’ OH, a 5’ phosphate, or both a 3’ OH and a 5’ phosphate. A ligatable polynucleotide may comprise RNA, DNA, or both RNA and DNA. A ligatable polynucleotide may comprise, for example, a label (e.g., fluorescent labels, affinity label, radiolabel, , glycan modification, deaminated base), a protein, an adapter (e.g., a sequencing adapter), a guide RNA, a linker, a bar code, and combinations thereof. A ligatable polynucleotide may comprise, for example, RNA, a sequencing adapter DNA, a sequencing adapter DNA-RNA chimera, a guide RNA, ssDNA, synthetic RNA, a portion of natural RNA or entire sequence of natural RNA (modified or unmodified), affinity tag attached RNA or RNA, artifificially or naturally modified DNA or RNA. In some embodiments, contacting the cleavage product comprising the sequence complementary to the sequence of interest, the ligatable polynucleotide (optionally referred to as a first ligatable polynucleotide), and the ligase further comprises contacting the cleavage product comprising the sequence complementary to the sequence of interest, the ligatable polynucleotide, the ligase, and a second ligatable polynucleotide. Methods may further include additional ligatable polynucleotides in like manner. Accordingly, it can be seen that a host of possible RNA molecules may be assembled. For example, first and second ligatable polynucleotides may be each of the two domains of a CRISPR nuclease guide. Extended (e.g., >1,000, >5,000, >10,000, >15,000, >20,000 nucleotide) RNAs may be assembled, according to some embodiments. Assembly reactions may include one or more splint DNAs to facilitate ordered assembly (e.g., ordered assembly in a single container and / or reaction). For example, a method may further include contacting the cleavage product comprising the sequence complementary to the sequence of interest and a splint DNA, the splint DNA comprising at least a portion of the sequence of interest and at least a portion of the sequence of the ligatable polynucleotide. Splint RNAs may be used in place of (or in addition to) DNA splints but without the advantage of removal by DNase I treatment. Assembly reactions may be performed in series to achieve ordered assembly of fragments without a splint.
[0019] The present disclosure further provides methods for cleaving a single-stranded RNA. For example, a method may comprise contacting a single-stranded RNA and an endoribonuclease having an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 to produce cleavage products, wherein the RNA comprises, in a 5’ to 3’ direction, (cs- ers)n(cs)m(A)x3 wherein cs encodes a sequence of interest (e.g., a coding sequence), ers encodes an endoribonuclease recognition sequence, A is a poly(A) tail, m=0 or 1, n=>l, >2, >3, >4, >5, >10, >15, >20, >25, or >50, x=0 or 1, and 3’ represents the 3’ end of the polycistronic RNA molecule, provided that if n=l, then mfO. According to some embodiments, a single-stranded RNA may comprise a plurality of sequences of interest (e.g., n is >2, n is >3, n is >4) wherein each sequence of interest may be the same as or different from the other sequence(s) of interest. In some embodiments, contacting may further comprise contacting the single-stranded RNA with a second endoribonuclease and wherein the polycistronic template further comprises at least one endoribonuclease recognition sequence that corresponds to the second endoribonuclease. At least one of the cleavage products produced may comprise a guide RNA, an IVT template or a therapeutic RNA. In some embodiments, contacting may include contacting the RNA and the endoribonuclease at high temperature (e.g., >43 °C, >65 °C, or >75 °C, or >85 °C, or >95 °C). For example, contacting the RNA and the endoribonuclease may comprise contacting the RNA and the endoribonuclease at >65°C. According to some embodiments, contacting may further comprise contacting the single-stranded RNA, the endoribonuclease, and at least one additional endoribonuclease, wherein each of the at least one additional endoribonucleases has an amino acid sequence that is at least 90% identical (e.g., >90%, >92%, >95%, >97%, or >99%, and optionally <100%) to any of SEQ ID NOS: 1-139, 163, and 164 other than the amino acid sequence of the endoribonuclease. A single— stranded RNA, in some embodiments, may be a polycistronic precursor RNA comprising at least one endoribonuclease recognition sequence that corresponds to the endoribonuclease and at least one endoribonuclease recognition sequence that corresponds to each of the at least one additional endoribonucleases. Cleavage products of such a precursor RNA may produce a plurality of copies of a single RNA species or a heterogeneous mixture of RNA species.
[0020] BRIEF DESCRIPTION OF THE FIGURES
[0021] FIGURE 1 shows an example embodiment of a ribonuclease expression construct, wherein “SS” represents a periplasmic signal sequence, “MBP” represents a maltose binding protein domain, “TCS” represents a thrombin cleavage sequence, “Insert” represents the coding sequence of an RNase gene, “6HIS” represents a polyhistidine (e.g., 6x His) tag, and “T7” represents a T7 RNA polymerase promoter.
[0022] FIGURE 2 illustrates an example embodiment of an endoribonuclease activity assay (EXAMPLE 2). Briefly, a substrate of a given ribonuclease (“pey43 RNA substrate”) is contacted with such ribonuclease (right) or subjected to a mock treatment (left) and reactions are resolved on a TBE-urea gel to assess whether cleavage has occurred.
[0023] FIGURES 3 A, 3B, and 3C show example cleavage activity results under the specific conditions tested. Activity of the subject endoribonucleases (e.g., ones displaying imperceptible or limited activity here) may be higher under other conditions. FIGURE 3 A shows example cleavage activity results for example endoribonucleases having the indicated amino acid sequences (SEQ ID NOS:79-81, 83-85, and 133). FIGURE 3B shows example cleavage activity results for example endoribonucleases having the indicated amino acid sequences (SEQ ID NOS:86-87, 134-139). FIGURE 3C shows example cleavage activity results for example endoribonucleases having the indicated amino acid sequences (SEQ ID NOS:88-94).
[0024] FIGURE 4 illustrates an example strategy for detecting ribonuclease recognition motifs for SEQ ID NOS: 1-139, 163, and 164. Briefly, a candidate substrate for an endoribonuclease is provided as shown. In a given candidate, each nucleotide represented as “N” may be provided as a specific nucleotide (e.g., A, G, C, or U) in a specific molecule. Populations of such constructs may be assembled covering a range of possible sequences. Such libraries may be subject to a mock treatment or a candidate endoribonuclease to form reaction products. Reaction products may be processed as shown. Figure legend: Compare control library to RNase treated library. Dropout library molecules contain RNase recognition motifs and underrepresented when compared to the control library.
[0025] FIGURE 5 shows example mass spectrometry characterization of cleavage sites within example motifs shown along the y-axis (SEQ ID NO: 152, {FAM}AAAAUAUU-p (SEQ ID NO: 153) (“-p” 3' phosphate), SEQ ID NO: 153 (“-p” 3' phosphate), SEQ ID NO: 154, 5’ FAM-labeled AAA>p (“>p” 2’, 3’ cyclic phosphate), SEQ ID NO: 155, GGUGAUA, SEQ ID NO: 156 (“-p” 3' phosphate), SEQ ID NO: 157 (“>p” 2’, 3' cyclic phosphate), SEQ ID NO: 158, SEQ ID NO: 159, (“>p” 2’, 3' cyclic phosphate), 5’ FAM- labeled AAAAU>p (“>p” 2’, 3' cyclic phosphate), 5’ FAM-labeled AAAAU>p (“>p” 2’, 3' cyclic phosphate), 5’ FAM-labeled AAAAUAU>p (“>p” 2’, 3' cyclic phosphate), [ FAM } AAAAUAUUOp (SEQ ID NO: 160) (“>p” 2’, 3' cyclic phosphate), 5’ FAM-labeled AAAAUAUU>p (“>p” 2’, 3' cyclic phosphate), R414 (SEQ ID NO: 141), and R415 (SEQ ID NO: 140)) using example endoribonucleases having amino acid sequences indicated along the x-axis.
[0026] FIGURE 6A shows an example method of assessing modification sensitivity of an endonuclease. As illustrated, a method may include comparing cleavage of a substrate RNA comprising no modified bases with cleavage of a substrate RNA having the same sequence but comprising one or more modified bases. The substrate RNA may have a sequence comprising all theoretical 5mers (SEQ ID NO: 145). FIGURE 6B shows example results of a modification sensitivity assay according to FIGURE 6A. Briefly, the assay comprised contacting E. coli MazF RNase and either unmodified RNA or m6A-modified RNA to form products and analyzing the products on a urea gel.
[0027] FIGURE 6C shows modification sensitivity of 22 selected example MazF enzymes using substrate RNAs (SEQ ID NO: 145) in which all As, Cs, or Us in the substrate were respectively replaced with the indicated modified base.
[0028] FIGURE 7A shows sensitivity results for the indicated example ribonucleases to a single modified nucleotide located in the recognition motif. FIGURE 7B shows an example modification sensitivity profile of endoribonuclease pEY586 having SEQ ID NO: 163 when contacted with substrate at one of the 4 indicated enzyme concentrations. FIGURE 7C shows an example modification sensitivity profile of endoribonuclease pEY546 having SEQ ID NO: 164 when contacted with substrate at one of the 4 indicated enzyme concentrations.
[0029] FIGURE 8A shows controlled fragmentation of RNA substrate (SEQ ID NO: 145) by example ribonucleases having the indicated amino acid sequences. FIGURE 8B shows the lengths of oligonucleotide reaction products arising from contacting pEY586 (SEQ ID NO: 163) with the indicated mRNA produced by in vitro transcription. Lengths were determined by UHPLC-MS / MS. FIGURE 8C shows the sequence coverage of pEY586 (SEQ ID NO: 163) in the IVT mRNA reactions of EXAMPLE 7.
[0030] FIGURES 9A - 9G show temperature dependent cleavage of an RNA substrate (SEQ ID NO: 151) by an example hyperthermophilic ribonuclease (SEQ ID NO:38) using the oligoribonucleotide substrate shown (SEQ ID NO: 151). FIGURE 9A shows an example reaction scheme in which an RNA substrate having the nucleotide sequence according to SEQ ID NO: 151 is contacted with an endoribonuclease having an amino acid sequence according to SEQ ID NO:38, incubated at a temperature of 37°C or 85°C to produce reaction products, and the reaction products are analyzed by liquid chromatography and tandem mass spectrometry (LC-MS / MS). FIGURE 9B shows the nucleotide sequence of the RNA substrate with the cleavage recognition sequences of the endoribonuclease shown in bold type and vertical bars marking the endoribonuclease cleavage site s(verti cal bars “ | ” ). Possible cleavage fragments of the RNA substrate are also shown ((l):nucleotides 1-20 of SEQ ID NO: 151; (2):nucleotides 21-29 of SEQ ID N0:151 and nucleotides 30-35 of SEQ ID NO: 151; (3):nucleotides 21-29 of SEQ ID N0:151 with a 3’ cyclic phosphate; and (4):nucleotides 30-35 of SEQ ID NO: 151). FIGURE 9C shows a possible secondary structure for the RNA substrate shown in FIGURE 9B. FIGURE 9D shows a LC-MS / MS chromatogram for products of a control reaction without enzyme incubated at 37°C.
[0031] FIGURE 9E shows a LC-MS / MS chromatogram for products of a reaction with enzyme incubated at 37°C. FIGURE 9F shows a LC-MS / MS chromatogram for products of a control reaction without enzyme incubated at 85°C. FIGURE 9G shows a LC-MS / MS chromatogram for products of a reaction with enzyme incubated at 85°C. FIGURE 9H shows example activity of an endoribonuclease having the amino acid sequence of SEQ ID NO:47 in the presence of increasing concentrations of urea on a substrate having the nucleotide sequence of SEQ ID NO: 141. FIGURE 91 shows example activity of an endoribonuclease having the amino acid sequence of SEQ ID NO:47 in the presence of increasing concentrations of formamide on a substrate having the nucleotide sequence of SEQ ID
[0032] N0:14L
[0033] FIGURE 10A shows an example method for denaturing an RNA substrate in the presence of an example hyperthermophilic ribonuclease (designated “peyl43” and having SEQ ID NO:38) at 85°C prior to nucleotide-specific ribonuclease (hRNase4) digestion.
[0034] FIGURE 10B shows a total ion chromatogram of example oligonucleotide cleavage products detected by LC-MS / MS following a no-enzyme control treatment at 85 °C (first panel), following cleavage of a yeast tRNAphesubstrate with the nucleotide-specific ribonuclease RNase 4 (second panel), following cleavage of a yeast tRNAphesubstrate with an example hyperthermophilic ribonuclease peyl43 (SEQ ID NO:38) at 85°C (third panel / Cont. 2), or following sequential digestion with peyl43 (SEQ ID NO:38) at 85°C to denature the tRNA substrate followed by RNase 4 digestion (fourth panel / Cont. 3).
[0035] FIGURE IOC shows the sequence coverage of pEY586 (SEQ ID NO: 163) in the E. coli rRNA reactions of EXAMPLE 8.
[0036] FIGURE 11 A and FIGURE 1 IB show an example application of sequence-specific endoribonucleases to the analysis of posttranscriptional 5’ capping reactions. Substrate RNAs were first capped and capping reaction products were cleaved with either a sequencespecific or a nucleotide-specific endoribonuclease. FIGURE 11 A shows the lengths of oligonucleotide products of cleavage reactions using an example sequence-specific RNase pey214 (SEQ ID NO: 132) and the nucleotide-specific ribonuclease RNase 4 (“R4”) to cleave RNA. As shown, pey214 resulted in a single fragment, while R4 cleaved at two positions. FIGURE 1 IB shows results of a mass spectrometry analysis of 5 ’-end capping reaction products utilizing an example sequence-specific RNase pey214 (SEQ ID NO: 132) and the nucleotide-specific ribonuclease RNase 4 to cleave RNA.
[0037] FIGURE 11C shows the results of a mass spectrometry analysis of 5 ’-end capping reaction products utilizing an example sequence-specific pEY586 (SEQ ID NO: 163).
[0038] FIGURE 12A shows an example method of generating homogenous 3’ ends of IVT products using a sequence-specific ribonuclease. Briefly, an mRNA substrate optionally may be produced by in vitro transcription (IVT) of a DNA template of interest comprising a terminal sequence encoding an RNase recognition motif. Under some conditions, some polymerases used in IVT reactions may add undesired extensions (templated or untemplated) to the 3’ end of an otherwise desired transcription product. Unwanted extensions may be removed by contacting the transcription products (the products comprising mRNA having an RNase recognition motif) and a ribonuclease operable to recognize the RNase recognition sequence and cleave the subject RNA, releasing the unwanted extensions. Optionally, desired mRNA cleavage products may be purified or otherwise separated from the unwanted extensions.
[0039] FIGURE 12B shows results of cleaving of 3’ ends of IVT products with example endoribonuclease pey214 (SEQ ID NO: 132) evaluated by capillary electrophoresis analysis. Briefly, a 5’ FAM-labeled 45-nt RNA oligonucleotide (SEQ ID NO: 142) was incubated at 37°C with no enzyme (upper trace), T7 RNA polymerase (upper middle trace), example ribonuclease pey214 (lower middle trace), or T7 RNAP followed by example ribonuclease pey214 (lower trace). Brackets at the top of the figure mark the substrate (S), (undesirably) extended oligonucleotides (E), and digested product (D).
[0040] FIGURE 12C shows results of the reactions of FIGURE 12B evaluated by mass spectrometry analysis.
[0041] FIGURE 12D shows a sketch of an example mRNA substrate comprising a pEY214 recognition sequence (AAA| ACAT, cleavage at the position marked with a vertical bar (“ | ”) in its poly A tail.
[0042] FIGURE 12E shows a gel fraction of reaction products of reactions illustrated in FIGURE 12D. Reactions were performed by contacting an example mRNA substrate (SEQ ID NO: 161) with no enzyme (control, “NE”) or endoribonuclease pEY214 at the indicated concentration. The uncleaved substrate in the NE lane is labeled “full length” while the expected 68 nt 3’ terminal cleavage product is labeled “68”. FIGURE 13 A illustrates an example cleavage ligation method comprising contacting a FAM-labeled substrate (SEQ ID NO: 143) having a stem loop structure and an example endoribonuclease recognition motif in the loop region (left image) and the corresponding endoribonuclease to form cleavage products (ntl-24 of SEQ ID NO: 143 and nt 25-47 of SEQ ID NO: 143; middle image) and contacting the cleavage products with an example RNA ligase operable to heal the cleavage, reforming the substrate (SEQ ID NO: 143).
[0043] FIGURE 13B shows separation of example reaction products on a denaturing urea gel. Substrate RNA (the substrate comprising a loop / single stranded region) was cleaved by an example endoribonuclease and the cleavage products were re-joined with RtcB ligase. Reaction products of a no-enzyme control appear in the lane on the left with the full-length RNA band (R498; SEQ ID NO: 143) marked. Reaction products of contacting the example endoribonuclease pey220 with the R498 substrate appear in the middle lane with the cleavage product (nt 1-24 of SEQ ID NO: 143) marked. Reaction products of contacting the example endoribonuclease pey220 with the R498 substrate followed by RtcB ligase appear in the right lane with the reformed substrate (SEQ ID NO: 143) and residual cleavage product (ntl-24 of SEQ ID NO: 143) comprising the FAM label is marked.
[0044] FIGURE 13C illustrates two example methods. The first, an example cleavage ligation method, comprises contacting a substrate RNA produced by IVT comprising a recognition / cleavage site and an example endoribonuclease operable to recognize the recognition site and cleave at such site to form cleavage products and contacting the cleavage products with an example RNA ligase to form a ligated RNA (e.g., reform the original substrate). The second, a control, comprises contacting unligatable synthetic RNA fragments with an example RNA ligase to form control reaction products. Synthetic substrates in the control reaction contain 5’ and 3’ hydroxyl groups thus they are not ligatable by RtcB ligase.
[0045] FIGURE 13D shows a gel fractionation of products of example methods illustrated in FIGURE 13C.
[0046] FIGURE 13E illustrates two example methods. The first, an example cleavage splinted ligation method, comprises contacting a substrate RNA comprising a recognition / cleavage site and an example endoribonuclease operable to recognize the recognition site and cleave at such site to form cleavage products, contacting cleavage products with a DNA splint complementary to at least a portion of each of the cleavage products (e.g., as pictured, fully complementary with both fragments) to form an RNA:DNA heteroduplex product comprising the RNA cleavage products and the DNA splint, contacting the heteroduplex and an example RNA ligase operable to ligate the RNA cleavage products within the splint heteroduplex (e.g., T4 RNA ligase 2) to form a heteroduplex ligation product, and optionally contacting the heteroduplex ligation product and a DNase operable to cleave the DNA strand of the heteroduplex ligation product (e.g., DNase I) to produce a single-stranded RNA ligation product. The second, a control, comprises contacting the same substrate with the same endoribonuclease to form cleavage products and contacting the cleavage products with the same RNA ligase (but without the splint) to produce control reaction products.
[0047] FIGURE 13F shows a gel fractionation of products of example methods illustrated in FIGURE 13E.
[0048] FIGURE 13G illustrates two example methods. The first, an example splinted ligation method, comprises contacting two (as illustrated) or more synthetic RNA molecules and a splint DNA complementary to at least a portion of each of the synthetic RNAs (e.g., as pictured, fully complementary with both RNAs) to form an RNA:DNA heteroduplex product comprising the synthetic RNAs and the DNA splint, contacting the heteroduplex and an example RNA ligase operable to ligate the synthetic RNAs within the splint heteroduplex (e.g., T4 RNA ligase 2) to form a heteroduplex ligation product, and optionally contacting the heteroduplex ligation product and a DNase operable to cleave the DNA strand of the heteroduplex ligation product (e.g., DNase I) to produce a single-stranded RNA ligation product. The second comprises contacting the same substrate with an example RNA ligase (optionally the same RNA ligase as the first method) but without the splint to produce control reaction products. These examples demonstrate that synthetic RNA can be converted by T4 PNK to RNA substrates that are amenable to ligation by T4 RNA ligase 1 and 2.
[0049] FIGURE 13H shows a gel fractionation of products of example methods illustrated in FIGURE 13G.
[0050] FIGURE 131 illustrates an example splinted RNA assembly method comprising contacting at least two synthetic RNA species (e.g., three pictured) wherein at least one of the species comprises a modified nucleotide and a splint DNA complementary to at least a portion of the synthetic RNA species (e.g., as pictured, fully complementary with all of the RNA species) to form an RNA:DNA heteroduplex product comprising the synthetic RNAs and the DNA splint, and contacting the heteroduplex product and an example RNA ligase operable to ligate the synthetic RNA species within the splint heteroduplex (e.g., T4 RNA ligase 2) to form a heteroduplex ligation product. A method may optionally further include contacting the heteroduplex ligation product and a DNase operable to cleave the DNA strand of the heteroduplex ligation product (e.g., DNase I) to produce a single-stranded RNA ligation product. A method may optionally further include contacting the single-stranded RNA ligation product and a ribonuclease operable to cleave the single-stranded RNA ligation product at the modified nucleotide (e.g., Endo V).
[0051] FIGURE 13 J shows a gel fractionation of products of the example method illustrated in FIGURE 131.
[0052] FIGURE 13K illustrates an example splinted RNA assembly method optionally comprising contacting a substrate RNA comprising a recognition / cleavage site and an example endoribonuclease operable to recognize the recognition site and cleave at such site to form cleavage products. A method may comprise contacting at least two RNA species (e.g., three pictured) wherein at least one of the species comprises a modified nucleotide and a splint DNA complementary to at least a portion of the RNA species (e.g., as pictured, fully complementary with the cleavage products and the RNA species comprising the modified nucleotide) to form an RNA:DNA heteroduplex product, the heteroduplex comprising the subject RNAs (e.g., as shown, the cleavage products separated by the RNA species comprising the modified nucleotide) and the DNA splint, and contacting the heteroduplex product and an example RNA ligase operable to ligate the subject RNA species within the splint heteroduplex (e.g., T4 RNA ligase 2) to form a heteroduplex ligation product. A method may optionally further include contacting the heteroduplex ligation product and a DNase operable to cleave the DNA strand of the heteroduplex ligation product (e.g., DNase I) to produce a single-stranded RNA ligation product. A method may optionally further include contacting the single-stranded RNA ligation product and a ribonuclease operable to cleave the single-stranded RNA ligation product at the modified nucleotide (e.g., Endo V).
[0053] FIGURE 13L shows a gel fractionation of products of the example method illustrated in FIGURE 13K.
[0054] FIGURE 14A illustrates an example method to generate circular RNA comprising contacting an example RNase and a synthetic or in vitro transcribed RNA to form a desired fragment and contacting the desired fragment and an RNA ligase operable to form an intramolecular (e.g., circular) ligation product.
[0055] FIGURE 14B shows an RNAfold webserver predicted structure of a RNA substrate (e.g., SEQ ID NO: 144) that may be used in a method of FIGURE 14A with the two cleavage sites indicated by black lines and arrows and the ligation site shown as the incomplete circle at the top of the predicted structure after RNase treatment.
[0056] FIGURE 14C shows results of a reaction with an example endoribonuclease of the disclosure, wherein cleavage products ran lower than the full-length starting substrate material (SEQ ID NO: 144) and the circularized products ran higher than the full-length material.
[0057] FIGURE 14D shows results of a reaction with an example endoribonuclease of the disclosure, wherein cleavage products ran lower than the full-length starting substrate material (syn97) and the circularized products ran higher than the full-length material. Lane 1 shows a no-enzyme control, lane 2 shows reaction products of the substrate and pEY214, lane 4 shows reaction products of the substrate, pEY214, and RtcB ligase, and lane 4 shows reaction products of the substrate, pEY214, RtcB ligase, and RNase R (to cleave any remaining single-stranded linear RNA).
[0058] FIGURE 15 illustrates an example method of using an endoribonuclease to cleave circular RNA molecules in preparation for sequencing.
[0059] FIGURE 16A and FIGURE 16B show methods of using an example endoribonuclease to cleave multiple mRNA molecules from a long transcript comprising multiple mRNA molecules. FIGURE 16A shows an example method using a substrate comprising multiple copies of a single mRNA. FIGURE 16B shows an example method using a substrate comprising multiple different RNAs (e.g., polycistronic).
[0060] FIGURE 16C illustrates a method of cleaving a substrate comprising multiple mRNA molecules comprising two example endoribonuclease recognition motifs and the corresponding endoribonuclease (as shown, pEY214)) to form cleavage products, wherein the cleavage products have 87 nucleotides, 200 nucleotides, and 323 nucleotides.
[0061] FIGURE 16D shows a gel fractionation of products of the example method illustrated in FIGURE 16C.
[0062] FIGURE 17 shows an example method for generating dsRNA comprising homogeneous ends, which dsRNA may be useful as an RNA size marker or in an RNA ladder.
[0063] FIGURE 18A shows an example method for generating point-modified RNA ^site- modified RNA) or recombinant RNA from fragments generated by endoribonucleases of the disclosure. FIGURE 18B shows a gel fractionation of products of the example programmable cleavage and ligation methods of EXAMPLE 16.
[0064] FIGURE 18C shows a gel fractionation of products of an example three-fragment assembly method of EXAMPLE 16 in which the middle fragment comprises inosine.
[0065] FIGURE 18D shows a gel fractionation of products of an example three-fragment assembly method of EXAMPLE 16 in which the middle fragment comprises a FAM label.
[0066] FIGURE 19 illustrates two example methods. The first method comprises contacting an RNA substrate and an example endoribonuclease (e.g., a modification insensitive endoribonuclease) to form cleavage products, wherein the RNA substrate comprises an endoribonuclease recognition motif comprising a modified nucleotide and an endoribonuclease recognition motif free of modified nucleotides. The second method comprises contacting the same substrate as above and a modification binding agent (e.g., a modification-specific antibody, a modification-binding protein) to form a binding product comprising the substrate and the binding agent bound to the modified nucleotide and contacting the binding product and the example endoribonuclease to form cleavage products.
[0067] FIGURE 20A illustrates four example methods. Method 20-1 comprises contacting an RNA substrate comprising a sequence-specific endoribonuclease cleavage motif and a corresponding example endoribonuclease to form cleavage products. Method 20-2 comprises contacting the same substrate and endoribonuclease in the presence of an DNA oligodeoxynucleotide that is complementary to a portion of the RNA substrate sequence (blocking the cleavage motif) and subsequent DNase treatment to remove the DNA oligo. Method 20-3 comprises contacting the same substrate and endoribonuclease in the presence of a DNA oligodeoxynucleotide that is complementary to the entire RNA substrate sequence and subsequent DNase treatment to remove the DNA oligo. Method 20-4 comprises contacting the same substrate and endoribonuclease in the presence of a DNA oligodeoxynucleotide that is complementary to a portion of the RNA substrate sequence and also comprises a central non-complementary sequence (operable to form a stem loop or bulge structure where the recognition motif is located) and subsequent DNase treatment to remove the DNA oligo.
[0068] FIGURE 20B shows gel fractionation of products of an example DNA oligonucleotide blocking method illustrated in FIGURE 20A (EXAMPLE 18).
[0069] FIGURE 21 A illustrates two example methods. The upper panel illustrates a method comprising contacting an RNA substrate with two cleavage sites and an example endoribonuclease to form cleavage products. The lower panel illustrates a method of programmable cleavage comprising contacting the same substrate and endoribonuclease in the presence of a DNA oligodeoxynucleotide that is complementary to one of the cleavage sites in the RNA substrate sequence and subsequent DNase treatment to remove the DNA inhibitor.
[0070] FIGURE 2 IB shows gel fractionation of products of an example DNA oligonucleotide blocking method illustrated in FIGURE 21A (EXAMPLE 18).
[0071] BRIEF DESCRIPTION OF THE SEQUENCES
[0072] Some embodiments of this disclosure relate to the following provided sequences of example polynucleotides and / or example polypeptides.
[0073] SEQ ID NOS: 1-139 are example endoribonucleases having names and cleavage recognition sequences as shown in TABLE 1.
[0074] SEQ ID NO: 140 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R415 and comprising recognition sequences for RNases designated pey216 and pey222 (nt 1=FAM label).
[0075] SEQ ID NO: 141 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R414 and comprising recognition sequences for RNases designated pey220 and pey222 (nt 1=FAM label).
[0076] SEQ ID NO: 142 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as LM001 and comprising a recognition sequence (nt 15-21) for an RNase designated pey214 (nt 1=FAM label).
[0077] SEQ ID NO: 143 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R498 and comprising a recognition sequence (nt 22-26) for an RNase designated pey220 (nt 1=FAM label).
[0078] SEQ ID NO: 144 is an example of a DNA sequence comprising a promoter (nt 6-23) and a template encoding a synthetic RNA referred to as Syn97. The template includes sequences encoding a recognition sequence (AAAACAT; nts 35-41 and 272-278) for an RNase designated pey214 and having the amino acid sequence of SEQ ID NO: 132.
[0079] SEQ ID NO: 145 is an example sequence of an RNA substrate. The sequence includes all possible (1024) distinct combinations of 5 nucleotides, 1534 of 4096 possible distinct combinations of 6 nucleotides (37.7%), and 1718 of 16384 possible distinct combinations of 7 nucleotides (10.4%). SEQ ID NO: 146 is an example sequence of an engineered and capped EPO mRNA, referred to as Syn90, that includes a pey214 recognition motif (AAAAACAU) with its cleavage product 61 nt away from the cap structure.
[0080] SEQ ID NO: 147 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R468 and comprising a recognition sequence (nt 11-13) for RNases designated pey77 and pey80 (nt 1=FAM label).
[0081] SEQ ID NO: 148 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R501 and comprising an m6A-modified recognition sequence (nl l=m6A; recognition seq nt 11-13) for RNases designated pey77 and pey80 (nt 1=FAM label).
[0082] SEQ ID NO: 149 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R502 and comprising an m6A-modified recognition sequence (nl3=m6A; recognition seq nt 11-13) for RNases designated pey77 and pey80 (nt 1=FAM label).
[0083] SEQ ID NO: 150 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R503 and comprising a 5mC-modified recognition sequence ((nl2=5mC; recognition seq nt 11-13) ) for RNases designated pey77 and pey80 (nt 1=FAM label).
[0084] SEQ ID NO: 151 is an example of a synthetic oligonucleotide comprising a U|GG recognition sequence (nt 20-22) and a secondary C|GG recognition sequence (nt 29-31) for the RNase designated peyl43. The cleavage sites are marked with a vertical bar (“ | ”).
[0085] SEQ ID NO: 152 is an example of a synthetic oligoribonucleotide.
[0086] SEQ ID NO: 153 is an example of a 5’ FAM-labeled synthetic oligoribonucleotide (nt 8=3’ phosphate).
[0087] SEQ ID NO: 154 is an example of a synthetic oligoribonucleotide.
[0088] SEQ ID NO: 155 is an example of a synthetic oligoribonucleotide.
[0089] SEQ ID NO: 156 is an example of a 5’ FAM-labeled synthetic oligoribonucleotide (nl3 = 3 '-phosphate).
[0090] SEQ ID NO: 157 is an example of a 5’ FAM-labeled synthetic oligoribonucleotide (nl3 = 2’, 3’ cyclic phosphate).
[0091] SEQ ID NO: 158 is an example of a synthetic oligoribonucleotide.
[0092] SEQ ID NO: 159 is an example of a 5’ FAM-labeled synthetic oligoribonucleotide (nlO = 2’, 3’ cyclic phosphate).
[0093] SEQ ID NO: 160 is an example of a 5’ FAM-labeled synthetic oligoribonucleotide (n9 = 2’, 3’ cyclic phosphate). SEQ ID NO: 161 is an example sequence of an engineered in vitro transcribed RNA that includes a pey214 recognition motif (AAAAACAU; nts 653-659) cleavage of which removes poly(A) tail from RNA.
[0094] SEQ ID NO: 162 is an example sequence of an engineered in vitro transcribed RNA that includes two pey214 recognition motifs (AAAAACAU; nts 85-91 and 285-291) cleavage of which generates three RNA fragments.
[0095] SEQ ID NO: 163 is an amino acid sequence of an example endoribonuclease designated pEY586 and shown in TABLE 1.
[0096] SEQ ID NO: 164 is an amino acid sequence of an example endoribonuclease designated pEY546 and shown in TABLE 1.
[0097] SEQ ID NO: 165 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R511 and comprising a 4-thiouridine-modified recognition sequence (underlined) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 4- thiouridine).
[0098] SEQ ID NO: 166 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R540 and comprising a 5,2’-O-dimethyluridine-modified recognition sequence (underlined) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 5,2’-O- dimethyluridine).
[0099] SEQ ID NO: 167 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R542 and comprising a 5,6-dihydrouridine-modified recognition sequence (underlined) for RNase designated pey586 (nt 1=FAM label; nt 12 = 5,6-dihydrouridine).
[0100] SEQ ID NO: 168 is an example sequence of EPO mRNA.
[0101] SEQ ID NO: 169 is an example sequence of Flue RNA.
[0102] SEQ ID NO: 170 is an example sequence of engineered 5kB-BGII mRNA.
[0103] SEQ ID NO: 171 is an example sequence of engineered 9kB-BGII mRNA.
[0104] SEQ ID NO: 172 is an example sequence of GFP Poly(A) mRNA.
[0105] SEQ ID NO: 173 is an example sequence of E. coli 5S RNA.
[0106] SEQ ID NO: 174 is an example sequence of E. coli 16S RNA.
[0107] SEQ ID NO: 175 is an example sequence of E. coli 23S RNA.
[0108] SEQ ID NO: 176 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R508 and comprising a recognition sequence (UGG; nts 20-22) for an RNases designated pey546 and pey586 (nt 1=FAM label). SEQ ID NO: 177 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R510 and comprising an pseudouridine-modified recognition sequence (UGG; nts 12-14) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = pseudouridine).
[0109] SEQ ID NO: 178 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R541 and comprising an 1-methylpseudouridine-modified recognition sequence (UGG; nts 12-14) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 1- methylpseudouridine).
[0110] SEQ ID NO: 179 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R543 and comprising an N3-methyluridine-modified recognition sequence (UGG; nts 12-14) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = N3- methyluridine).
[0111] SEQ ID NO: 180 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R544 and comprising an 5-bromouridine-modified recognition sequence (underlined) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 5- bromouridine).
[0112] SEQ ID NO: 181 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R545 and comprising an 5F-modified recognition sequence (UGG; nts 12-14) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 5-fluorouridine).
[0113] SEQ ID NO: 182 is an example of a FAM-labeled synthetic oligoribonucleotide referred to as R546 and comprising an 5-iodouridine-modified recognition sequence (underlined) for RNases designated pey546 and pey586 (nt 1=FAM label; nt 12 = 5- iodouridine).
[0114] SEQ ID NO: 183 is an example of a synthetic oligoribonucleotide referred to as R618 comprising recognition sequence (nts 38-44) for RNase designated pey214.
[0115] SEQ ID NO: 184 is an example of a synthetic oligoribonucleotide referred to as R619, used for ligation reactions.
[0116] SEQ ID NO: 185 is an example of a synthetic oligoribonucleotide referred to as R620 used for ligation reactions.
[0117] SEQ ID NO: 186 is an example of a synthetic oligoribonucleotide referred to as R625 comprising an Inosine-modification used for ligation reactions.
[0118] SEQ ID NO: 187 is an example of a synthetic oligo deoxyribonucleotide referred to as D850, used to block the cleavage of SEQ ID NO: 162. SEQ ID NO: 188 is an example synthetic DNA oligonucleotide referred to as D875, used as DNA splint and / or cleavage inhibitor for R618.
[0119] SEQ ID NO: 189 is an example of a synthetic oligodeoxyribonucleotide referred to as D876, DNA block for R618.
[0120] SEQ ID NO: 190 is an example of a synthetic oligodeoxyribonucleotide referred to as D877, DNA splint for R618 and R625.
[0121] SEQ ID NO: 191 is an example of a synthetic oligodeoxyribonucleotide referred to as D878, DNA block for R618.
[0122] SEQ ID NO: 192 is an example of a synthetic FAM labeled oligoribonucleotide referred to as R623, used for ligation (nt 17 = FAM labeled T).
[0123] SEQ ID NO: 193 is an example of a synthetic inosine containing oligoribonucleotide referred to as R627, used for ligation.
[0124] SEQ ID NO: 194 is an example of a 24 nt long DNA splint used in two-fragment ligation in FIGURE 18B.
[0125] SEQ ID NO: 195 is an example of an example of DNA disruptor binding to the left side of DNA split used in FIGURES 18B-18D.
[0126] SEQ ID NO: 196 is an example of an example of DNA disruptor binding to the right side of DNA split used in FIGURES 18B-18D.
[0127] SEQ ID NO: 197 is an example of a 55 nt long DNA splint used for 3-fragment ligation in FIGURE 18C and 18D.
[0128] SEQ ID NO: 198 is an example of an engineered human mRNA, PSMB2, referred to as PSMB2-600.
[0129] SEQ ID NO: 199 is an example of an engineered human mRNA, PSMB2, sequence, referred to as PSMB2-lkb.
[0130] SEQ ID NO:200 is an example of a DNA guide for Mucilaginibacter paludi AGO (Argonaute) protein.
[0131] DETAILED DESCRIPTION
[0132] The present disclosure relates, in some embodiments, to compositions and methods for characterizing and / or manipulating RNA. For example, compositions and / or methods may include nucleotide-or sequence-specific endoribonucleases having desired recognition motifs and / or biophysical properties (e.g., resistance to denaturing reagents and / or thermal stability) for controllable fragmentation of RNA molecules. RNA fragments having controlled and / or desired sizes may be analyzed by mass spectrometry, sequencing (e.g., next generation sequencing technologies including, for example, Illumina or Nanopore sequencers), and / or fragment analyzers. According to some embodiments, nucleotide-or sequence-specific endoribonucleases having desired recognition motifs and / or biophysical properties may be used to detect RNA modifications (enzymatic modification or chemical damages), cap status, capping efficiency, and / or polyadenylation for a subject RNA (e.g., a therapeutic RNA). Nucleotide-specific and / or sequence-specific endoribonucleases having desired recognition motifs and / or biophysical properties may be used in recombinant RNA methods, in some embodiments. For example, such enzymes may be used to linearize circular RNA molecules and / or may generate homogenous 5’ or 3’ RNA ends for accurate ligation of RNA substrates. In some embodiments, such enzymes may be used to make circular RNA or site-specifically modified RNA, to linearize naturally-occurring circular RNA for Nanopore sequencing or other applications, or to generate homogenous 5’ or 3’ RNA ends (removing undesired extension products after in vitro transcription). After cleavage by a sequence-specific endoribonuclease, byproducts (undesired extensions) could be removed using various methods. These methods include size selection, size exclusion chromatography, binding to cellulose matrix (dsRNA binding), HPLC, anion exchange chromatography, 5’-to-3’ or 3’-to-5’ exonuclease cleavage of byproducts specifically, and use of aptamer tags in RNA that can be removed after RNase cleavage.
[0133] Methods may include, in some embodiments, sequence mapping an RNA with one or more endoribonucleases. Nucleotide- and sequence-specific endoribonucleases may provide controlled cleavage of an RNA (e.g., a synthetic RNA or a cellular RNA). Controlled cleavage of an RNA substrate of interest can be used for characterization of the RNA sequence (e.g., the identity of the RNA), for assessment of the integrity and / or purity of the RNA. Controlled RNA cleavage may also provide detection of chemical modifications (e.g., chemical modifications incorporated synthetically in a synthetic RNA substrate during transcription using a modified nucleotide triphosphate and / or chemical modifications arising or induced post-transcriptionally using one or more of a chemical agent, an enzyme, temperature, and radiation). In some embodiments, endoribonucleases may be coupled with spectroscopy (e.g., mass spectrometry) or electrophoretic analysis for characterization and quality control of therapeutic RNA. Endoribonucleases may also be coupled with mass spectrometry or electrophoretic analysis for characterization of RNA samples extracted from cells, tissue preparations and bodily extracts (e.g., blood, saliva, nasal swabs, urine, etc.). The nucleotide-specificity or sequence-specificity of a ribonuclease may influence the ways in which a ribonuclease is used in biotechnology. For example, an RNA molecule comprising 25% A, 25% G, 25% C, and 25% U arranged in a random sequence would be cut (on average) once every 4 nt (e.g., yield fragments having an average size of 4 nt) by a nucleotide-specific (1-base) cutter, once every 16 nt (e.g., yield fragments having an average size of 16 nt) by a 2-base cutter, once every 64 nt by a 3-base cutter, once every 256 nt by a 4-base cutter, once every 512 nt by a 5-base cutter, and so on. However, in practice, the actual cleavage frequency may be affected by one or more of the following: (a) the nucleotide neighborhood (i.e., adjacent nucleotides next to the cleavage site) and / or (b) by the local RNA secondary structure at and around the cleavage site and / or (c) by chemical modifications present in RNA such as covalent epitranscriptomic modifications and / or (d) by the degeneracy of bases at the recognition motif. Known nucleotide- or sequence-specific endoribonucleases possess relative substrate specificity, not absolute specificity; thus, their cleavage activity and / or specificity may be affected by ratio of enzyme to substrate, and / or the nature and / or concentration of any reaction buffer(s). Appropriate (even preferred) reaction conditions may be determined for each endoribonuclease by identifying its cleavage motif / sequence, producing a long RNA substrate having a copy of this motif / sequence, contact 2-fold dilutions of the enzyme with a defined amount of the substrate (e.g., such that each dilution has a defined enzyme to substrate ratio), and examine the reaction products. In this way an enzyme to substrate ratio is identified at which the cleavage motif is cleave with little or no off-target cleavage. Reaction buffer, salt concentrations, and / or reaction temperature may be selected as desired. Many site-specific endoribonucleases either fail to cleave double-stranded regions of an RNA substrate, or cleave it inefficiently compared to the single-stranded RNA regions. Examples of cleavage reaction conditions include without limitation buffer composition and concentration, salts, pH, temperature, metal ions (e.g., kind and concentration), and additives (e.g., kind (such as detergents and denaturants) and concentration). The quality and purity of the enzyme preparations (including of their recombinant forms and mutants thereof) may also have a substantial impact on the RNA cleavage specificity.
[0134] Cut frequency and average fragment sizes may be influenced by endoribonuclease recognition sequences as well as the target substrate’s GC content, modified nucleotide content, and / or the presence and nature of any secondary and / or tertiary structure (e.g., dsRNA, helices, loops). Short, frequent cutters (e.g., endoribonucleases with single nucleotide or dinucleotide recognition sequences) may be used to generate short RNA fragments (average sizes of 4-150, 4-100, 4-80, 4-60, 4-40, or 4-20 nt), for sequence mapping by mass spectrometry. Long cutters (e.g., endoribonucleases with trinucleotide and longer recognition sequences) may be used to generate longer RNA fragments (average sizes of 5- 200, 20-200, 5-300, 5-500, or 5-1000 nt providing RNA signatures or fingerprints for very long transcripts (e.g., viral RNAs or self-amplifying RNAs, etc.) identification. The present disclosure also demonstrates that these enzymes enable analysis of modified and unmodified RNA. While enzymes sensitive to RNA modifications can be used for modification mapping and detection, the enzymes that can cleave modified RNA can be used for characterization of modified RNA therapeutics or vaccines. In some embodiments, endoribonucleases have surprisingly high resistance to denaturing conditions such as heat, formamide, and urea which are conditions that disrupt the secondary structure of ssRNA and dsRNA, thus facilitating the cleavage of the highly structured RNA substrates.
[0135] Some endoribonucleases may cut modified RNA and unmodified RNA, some endoribonucleases may cut only modified RNA, and some endoribonucleases may cut only unmodified RNA. In some embodiments, such endoribonucleases may be used alone or in combination, for example, to characterize the presence and / or location of a modified base in a subject RNA molecule. For example, engineered therapeutic RNA may comprise one or more nucleotide modifications (e.g., Nl-methyl-pseudouridine, 5-methoxyuridine, 5- methylcytidine, among many others). Controlled fragmentation of such RNA for downstream analysis (e.g., sequencing, spectroscopy, electrophoretic analysis, photometric analysis) may be achieved with an endoribonuclease that is site-specific, but not sensitive to (i.e., not blocked by) RNA modification (e.g., an enzyme that can cleave RNA whether or not it has a modification). Modification sensitive endoribonucleases (e.g., modification sensitive endoribonucleases that are completely inactive on substrates comprising a modified nucleotide in the recognition sequence) may be used to map and / or detect RNA modifications in both cellular RNA and chemically synthesized RNA. However, the site-specificity and activity of ribonucleases toward unmodified or modified RNA may defy reliable prediction based on amino acid sequence or protein structure alone. In such cases, site-specificity and / or activity may be determined empirically.
[0136] RNases, each targeting a distinct RNA sequence motif, may be used alone, in parallel (e.g., in separate reactions), or in combination (e.g., in a single reaction) to characterize RNA species of interest. RNA species of interest may have different sequence lengths (e.g., messenger RNA, circular RNA, self-replicating RNA or self-amplifying RNA, in each case, having any desired size, for example, in a range of X to Y nucleotides, where X is 100, 1,000, or 3,000, Y is 1,000, 9,000, or 50,000, and X<Y), structural motifs (e.g., hairpins, quadruplexes, i-motifs, etc.) (PMID: 11491284), and modification content (e.g., one or more modifications that are the same or different from each other, synthetic or naturally occurring, in any combination).
[0137] In some embodiments, the present disclosure provides endoribonucleases (and combinations thereof, compositions thereof, and kits thereof) each having a desired specificity and / or a desired RNA modification sensitivity for RNA analysis. The present disclosure provides, in some embodiments, nucleotide- or sequence-specific endoribonucleases that efficiently cleave single and / or double-stranded RNA and / or that are active in / tolerant of conditions that disrupt RNA secondary structure, such as high concentration of denaturing reagents (e.g., > 3 M formamide or urea) or high temperature (e.g., >43 °C, >65 °C, or >75 °C, or >85 °C, or >95 °C).
[0138] General Considerations
[0139] Aspects of the present disclosure can be understood in light of the provided descriptions, figures, sequences, embodiments, section headings, and examples, none of which should be construed as limiting the entire scope of the present disclosure in any way. Accordingly, the innovations set forth herein should be construed in view of the full breadth and spirit of the disclosure.
[0140] Each of the individual embodiments described and illustrated herein has discrete components and features which can be readily separated from or combined with the components and / or features of any of the other several embodiments without departing from the scope or spirit of the present teachings. Lists of example species within a particular genus may vary in length at different places throughout the disclosure. Species lists shortened for convenience shall not be construed to exclude example species listed elsewhere in the specification. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. Unless otherwise expressly stated to be required herein, each component, feature, and method step disclosed herein is optional and the disclosure contemplates embodiments in which each optional element may be expressly excluded. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements or use of a “negative” limitation. It is further intended to serve as antecedent basis for use of such elective terminology as “optionally” and the like in connection with the recitation of one or more claim elements. In addition, unless otherwise expressly stated to have only the disclosed property or function, each component, feature, and method step disclosed herein may have additional properties and / or functions. For example, a disclosure that an enzyme cleaves DNA would not preclude such enzyme from also having the property of cleaving RNA unless expressly disclosed to cleave only DNA.
[0141] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Still, certain terms are defined herein with respect to embodiments of the disclosure and for the sake of clarity and ease of reference.
[0142] Sources of commonly understood terms and symbols may include: standard treatises and texts such as Kornberg and Baker, DNA Replication, Second Edition (W.H. Freeman, New York, 1992); Lehninger, Biochemistry, Second Edition (Worth Publishers, New York, 1975); Strachan and Read, Human Molecular Genetics, Second Edition (Wiley-Liss, New York, 1999); Eckstein, editor, Oligonucleotides and Analogs: A Practical Approach (Oxford University Press, New York, 1991); Gait, editor, Oligonucleotide Synthesis: A Practical Approach (IRL Press, Oxford, 1984); Singleton, et al., Dictionary of Microbiology and Molecular biology, 2d ed., John Wiley and Sons, New York (1994), and Hale & Markham, the Harper Collins Dictionary of Biology, Harper Perennial, N.Y. (1991) and the like.
[0143] As used herein and in the appended claims, the singular forms “a” and “an” include plural referents unless the context clearly dictates otherwise. For example, the term “a protein” refers to one or more proteins, i.e., a single protein and multiple proteins.
[0144] Numeric ranges are inclusive of the numbers defining the range. All numbers should be understood to encompass the midpoint of the integer above and below the integer i.e., the number 2 encompasses 1.5-2.5. The number 2.5 encompasses 2.45-2.55 etc. When sample numerical values are provided, each alone may represent an intermediate value in a range of values and together may represent the extremes of a range unless specified. Percent ranges with only one end point (e.g., > 90% or < 10%) optionally include a second endpoint at the maximum or minimum percentage (e.g., > 90% includes a range of 90%-100% and < 10% includes a range of 0%-10%). Ranges (including percent ranges) with only one end point (e.g., > 90 or < 10) optionally include a second endpoint 10% higher or 10% lower than the provided endpoint (e.g., > 90 includes a range of 90-99 and < 10 includes a range of 1-10). Concentration percentages are w / v percentages unless otherwise indicated. In the context of the present disclosure, “3’ UTR” refers to a natural and / or artificial RNA sequence disposed 3’ of a coding sequence (e.g., 3’ of the stop codon) and 5’ of a polyA tail. A 3’ UTR sequence is not translated (except in rare or special cases, for example, where translation proceeds beyond the stop codon) and may be operable to regulate post- transcriptional expression of the RNA.
[0145] In the context of the present disclosure, “AGO guide” refers to a single-stranded oligonucleotide (a) capable of binding (e.g., hybridizing to) a polynucleotide having a complimentary sequence, (b) capable of binding an Argonaute, and (c) comprising (i) at least 12 nucleotides (e.g., 12-35 nucleotides), (ii) > 50% deoxyribonucleotides (e.g., > 60%, > 70%, > 80%, > 90%, or < 100% deoxyribonucleotides), (iii) < 50% ribonucleotides (e.g., < 10%, < 25%, < 35%, < 45%, < 50% ribonucleotides), (iv) optionally, a phosphorylated 5’ end, (v) optionally, a nucleotide sugar modification, and (vi) optionally, a nucleotide substitution. In some embodiments, an AGO guide may comprise a phosphorylated 5’ end or another chemical modification at its 5’ end. An AGO guide may be engineered or synthetic with a sequence selected to complement a desired target sequence. An AGO guide maybe capable of directing an Argonaute polypeptide:guide DNA complex to a target polynucleotide. A DNA guide may be an oligonucleotide or polynucleotide that is synthetic or from a natural source such as genomic DNA, cDNA, extrachromosomal DNA, microbial DNA or viral DNA (e.g., the natural source differing from the Argonaute such that the guide and Argonaute together form a non-naturally occurring combination). The guide DNA is generally single stranded when used with Argonaute although it may be derived from dsDNA. While RNA guides may be used, it will be recognized that RNA guides may be impractical where cost and / or stability considerations are important.
[0146] In some embodiments, an AGO guide length suitable for Argonaute cleavage of dsDNA (e.g., in the presence of a helicase or single strand binding protein) may comprise at least 12 nucleotides, for example, having a size range of 12-60 nucleotides, 14-50 nucleotides, 15-40 nucleotides, 16-35 nucleotides, 15-24 nucleotides, or 16-21 nucleotides. In some embodiments, an AGO guide DNA may be greater than 21 nucleotides or at least 24 nucleotides in length. In some embodiments, an AGO guide may be 16-21 nucleotides in length (e.g., 16, 17, 18, 19, 20 or 21 nucleotides).
[0147] In some embodiments, an AGO guide may comprise a nucleotide sugar modification or a nucleotide substitution. In some embodiments, a nucleotide sugar modification comprises a 2' sugar modification and maybe selected from the group consisting of a 2'-O— CH3, a 2'-F, and a 2'-M0E modification. In some embodiments, a nucleotide substitution comprises one selected from the group consisting of locked nucleic acid (LNA), an unlocked nucleic acid (UNA), deoxyuridine, pseudouridine, 5-methylcytosine, 2-aminopurine, 2,6- diaminopurine, deoxyinosine, 5-hydroxybutynl-2'-deoxyuridine, 8-aza-7-deazaguanosine, and 5-nitroindole. In some embodiments, an AGO guide molecule comprises a sugar modification and a nucleotide substitution.
[0148] The nucleotide sequence of an AGO guide may or may not be degenerate. For example, guides with degenerate sequences may be useful for targeting nucleic acid sequences that are not fully known and / or for targeting more than one variant in a population of polynucleotides.
[0149] In the context of the present disclosure, “buffer” and “buffering agent” refer to a chemical entity or composition that itself resists and, when present in a solution, allows such solution to resist changes in pH when such solution is contacted with a chemical entity or composition having a higher or lower pH (e.g., an acid or alkali). A buffering agent may better resist pH changes within a certain pH range than outside of such range. A buffering agent may comprise a proton-accepting component and / or a proton-donating component. Examples of suitable non-naturally occurring buffering agents that may be used in disclosed compositions, kits, and methods include HEPES, MES, MOPS, TAPS, tricine, and Tris. Additional examples of suitable buffering agents that may be used in disclosed compositions, kits, and methods include ACES, ADA, BES, Bicine, CAPS, carbonic acid / bicarbonic acid, CHES, citric acid, DIPSO, EPPS, histidine, MOPSO, phosphoric acid, PIPES, POPSO, TAPS, TAPSO, and triethanolamine. In some embodiments, a buffering agent may comprise a protein (e.g., a recombinant albumin).
[0150] As used herein, “catalytically active” refers to the property of a molecule (e.g., a proteinaceous molecule or macromolecule) to function as a catalyst of one or more chemical reactions relative to one or more substrates and products. A catalytically active sequence specific endoribonuclease or sequence specific endoribonuclease variant, for example, hydrolyzes one or more ribonuceic acids to yield (however briefly) oligonucleotide products comprising a 5’ phosphate, a 5’ OH, a 2-3’ cyclic phosphate, a 3’ phosphate, or a 3’ OH. Catalytic activity of a sequence specific endoribonuclease and / or sequence specific endoribonuclease variant may be assessed using existing techniques applied to one or more model substrates and / or one or more substrates of interest. For example, effective assays for catalytic activity of sequence-specific endoribonucleases may include size fractionation of products (e.g., on gels or other matrices), radioactive assays, spectroscopy (e.g., tandem mass spectrometry), liquid chromatography, and fluorometric assays. Catalytic activity may be assessed with respect to loss of a sequence-specific substrate or appearance of cleavage products (e.g., FAM-labeled products), and / or metrics that serve as a proxy thereof.
[0151] In the context of the present disclosure, “contacting” refers to any act of bringing one item (e.g., a molecule or group of molecules, object, material) in contact with one or more other items (e.g., a molecule or group of molecules, object, material, like or unlike the first molecule(s)). Contacting includes, for example, adding, amalgamating, blending, combining, connecting, emulsifying, joining, mixing, precipitating, stirring, and / or touching one item with one or more other items. In some embodiments, contacting includes contacting one item and another in or with a living cell. In some embodiments, contacting includes or consists of contacting under cell-free conditions (in vitro) and, accordingly, excludes contacting in vivo.
[0152] In the context of the present disclosure, “contact” refers to any physical, chemical, electrical, magnetic or other association between two or more like or unlike materials (e.g., between two molecules). Contacting includes any process, method, workflow or other means of bringing two (or more) materials into contact with one another. Contacting, for example, an enzyme and a substrate or a binding protein and its corresponding target, may include providing suitable conditions (e.g., concentration, pH, solvent, buffer, space (volume), temperature, time) and other parameters for the two materials to associate (e.g., for an enzyme to operatively interact with its substrate or a binding protein to bind its target). Contacting may be achieved by any method that brings two (or more) materials into operative association with one another including mixing (e.g., in solution), pouring, pipetting, flowing, injecting, vortexing, transferring, incubating, emulsifying, agitating, spraying, adhering, or coating one material with, in to, or on to another.
[0153] In the context of the present disclosure, “container” refers to a human-made container including containers made using human-programed machines. A container may comprise one or more walls (e.g., defining an interior volume) and optionally one or more openings. Containers comprising one or more openings may further comprise one or more closures (e.g., removable closures) for some or all such openings. A closure optionally may comprise an aperture or a septum, for example, to provide fluid communication with a volume of the container and a connected or inserted tube or syringe. Examples of containers include boxes, cartons, bottles, tubes (e.g., test tubes, microcentrifuge tubes), plates (e.g., 96-well, 384-well plates), vials, pipette tips, and ampules. Containers and / or closures may comprise any desired material including paper, plastics, glass, silicone, composites, metals, alloys, or combinations thereof. Containers and / or closures may comprise materials that are compostable, recyclable, and / or sustainable.
[0154] In the context of the present disclosure and with respect to an amino acid residue or a nucleotide base position, “corresponding to” refers to positions that lie across from one another when sequences are aligned, e.g., by the BLAST algorithm. An amino acid position in a functional or structural motif in one polymerase may correspond to a position within a functionally equivalent functional or structural motif in another polymerase.
[0155] In the context of the present disclosure and with respect to an endoribonuclease and an endoribonuclease recognition site of an RNA molecule, “corresponding to” refers to an enzyme-substrate relationship wherein the endoribonuclease is operable to recognize the recognition site and cleave the subject RNA in which the recognition site occurs.
[0156] In the context of the present disclosure, “CRISPR nuclease guide” refers to a singlestranded oligonucleotide (i) complementary to a CRISPR guide target site, (ii) operable to mediate binding of a CRISPR / Cas complex to the CRISPR guide target site, (iii) having at least 12 nucleotides (e.g., 12-95 nucleotides. 12-25 nucleotides, 20-40 nucleotides, 35-55 nucleotides, 50-75 nucleotides, 70-95 nucleotides, >95 nucleotides), (iv) comprising > 50% ribonucleotides (e.g., > 60%, > 70%, > 80%, > 90%, or < 100% ribonucleotides), (iv) optionally, having a phosphorylated 5’ end, (v) optionally, having a nucleotide sugar modification, and (vi) optionally, having a nucleotide substitution. A CRISPR nuclease guide may comprise a phosphorylated 5’ end or another chemical modification at its 5’ end. A CRISPR nuclease guide may be engineered or synthetic with a sequence selected to complement a desired target sequence. A CRISPR nuclease guide RNA may comprise two domains a complementary domain that is complementary to a desired target site (operable to bind to the target site) and a nuclease binding domain (operable to bind a CRISPR / Cas nuclease) which, together, are operable to guide the nuclease to the target site. Guide length may be as provided above or may be adjusted in view of the desired target site (e.g., Zetsche et al. Cell 163, 759-71 (2015)).
[0157] In the context of the present disclosure, “duplex” and “double stranded” refer to any conformation of a polynucleotide in which two polynucleotide strands (e.g., separate molecules or spatially separated portions of a single molecule) comprise secondary structure elements formed via intermolecular base complementation. For example, strands of a duplex may be arranged anti parallel to one another in a helix (e.g., A-form, B-form, Z-form) with complementary bases of each strand paired with one another (e.g., in Watson-Crick base pairs). Paired bases may be stacked relative to one another to permit pi electrons of the bases to be shared. A polynucleotide may have a homoduplex conformation (e.g., DNA:DNA or RNA:RNA) or a heteroduplex confirmation (e.g., DNA:RNA).
[0158] In the context of the present disclosure, “endoribonuclease” refers to an enzyme that hydrolyzes the phosphodiester backbone of one or more polyribonucleic acid substrates (e.g., single-stranded or double-stranded) to yield products comprising at least one 5’- phosphorylated polyribonucleotide, at least one 5 ’-hydroxylated polyribonucleotide, at least one 2’,3’-cyclicphosphorylated polyribonucleotide, at least on 3 ’-phosphorylated polyribonucleotide, and / or at least one 3 ’-hydroxylated polyribonucleotide. An endoribonuclease may recognize a motif in an RNA molecule comprising a single nucleotide or 2, 3, 4, 5, 6, 7 or more consecutive nucleotides. An endoribonuclease may cleave a substrate RNA comprising its recognition motif within or adjacent to or outside of such recognition motif.
[0159] Activity of an endoribonuclease may be expressed in terms of units where a unit of an endoribonuclease is the amount of enzyme that completely cleaves 5 pmol of substrate in 30 minutes at 37°C. Specific activity of an endoribonuclease may be expressed in terms of units per milligram of endoribonuclease protein.
[0160] Catalytic activity of an endoribonuclease may persist across a range of salt concentrations, temperatures and / or pH. For example, an endoribonuclease may display catalytic activity under a range of conditions and / or following removal from exposure to such conditions. An endoribonuclease may have catalytic activity at and / or following exposure to a pH from X'l to Y^, where X'l is any of pH 4, 4.5, 5, 5.5, 6, 6.5, 7, and Y^ is any of pH 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11 and X^Yt Reaction conditions including salt concentrations, temperatures and / or pH may impact off-target cleavage. Accordingly, it may be desirable to evaluate a range of conditions for use of a specific endoribonuclease with a specific substrate.
[0161] In the context of the present disclosure, “excipient” refers to any pharmaceutically acceptable material other than a therapeutic RNA that may be included with a therapeutic RNA contacted with a mammalian cell. Examples of excipients include diluents, emulsifiers, suspension aids, lipids (e.g., liposomes, lipidoids, lipid nanoparticles), dispersants, stabilizers, buffering agents, reducing agents, free radical scavengers, antioxidents, lubricants, sugars and other carbohydrates, solvents, microspheres, and combinations and mixtures thereof. In the context of the present disclosure, “fusion” refers to two or more polypeptides, subunits, or proteins covalently joined to one another (e.g., by a peptide bond). For example, a protein fusion may refer to a non-naturally occurring polypeptide comprising a protein of interest covalently joined to a second polypeptide. Examples of a second polypeptide include a reporter protein (e.g., a green fluorescent protein), a purification tag (e.g., a 6xHis or 8xHis tag), and expression tag, a polynucleotide binding protein, an enzyme, a conjugation tag (e.g., a SNAP® tag), and a peptide linker (e.g., a flexible linker, an inflexible linker, a cleavable linker). Unless otherwise disclosed, the protein of interest may be nearer to the N-terminal end or nearer to the C-terminal end than the second polypeptide to which it is joined. A fusion protein may have one or more heterologous domains added to the N-terminus, C-terminus, and or the middle portion of the protein. A fusion may comprise a non-naturally occurring combined polypeptide chain comprising two proteins or two protein domains joined directly to each other by a peptide bond or joined through a peptide linker. If two parts of a fusion protein are “heterologous”, they are not part of the same protein in its natural state. In some embodiments, a fusion may comprise an endoribonuclease (or an endoribonuclease variant) covalently joined to a second polypeptide. In some embodiments, a variant an endoribonuclease (or an endoribonuclease variant) may include a fusion to maltose binding protein and / or a histidine (e.g., 6xHis) tag, a heterologous targeting sequence, a linker, an epitope tag, a detectable fusion partner, such as a fluorescent protein, P-galactosidase, luciferase and / or functionally similar peptides. Components of a fusion protein may be joined by one more peptide bonds, disulfide linkages, and / or other covalent bonds.
[0162] In the context of the present disclosure, “guide” refers to and includes AGO guides and / or CRISPR nuclease guides.
[0163] In the context of the present disclosure, “heterologous” refers to materials from different sources. For example, a naturally-occurring enzyme (e.g., a bacterial sequencespecific endoribonuclease of Thermococcus) is heterologous to a substrate RNA where the RNA comprises a nucleotide sequence that does not occur in the native host (e.g., an artificial sequence that does not occur anywhere in nature, a mammalian sequence).
[0164] In the context of the present disclosure, “immobilized” refers to covalent attachment of an enzyme to a solid support with or without a linker. Examples of solid supports include beads (e.g., magnetic, agarose, polystyrene, polyacrylamide, chitin). Beads may include one or more surface modifications (e.g., O6-benzyleguanine, polyethylene glycol) that facilitate covalent attachment and / or activity of an enzyme of interest. For example, a support may comprise a ligand and an enzyme may have a receptor for such ligand or an enzyme may comprise a ligand and a support may comprise a receptor for such ligand. Receptor-ligand binding may be covalent or non-covalent. Non-covalent attachment (e.g., avidimbiotin, chitin:CBP) may be useful in some embodiments, for example, where the level of dissociation of the binding partner is deemed tolerable. A linker may be disposed between a support and an enzyme. For example, a linker disposed between a support and an enzyme may have a first covalent bond to the support and a second covalent bond to the enzyme. An immobilized enzyme comprising a ligand-receptor attachment may have a linker disposed between the support and the ligandreceptor attachment, a linker disposed between the enzyme and the ligand-receptor attachment, or both. An immobilized enzyme comprising a linker may also comprise an optional covalent bond directly between the enzyme and the support. A linker may be of any desired length and have any desired range of motion. A peptide linker may comprise one or more repeats (e.g., 1-10 repeats) of glycine-serine.
[0165] In the context of the present disclosure, “in vitro transcription” (IVT) refers to a cell- free reaction in which a DNA template is copied by a DNA-directed RNA polymerase (e.g., a cold-active RNA polymerase or a T7 RNA polymerase) to produce a product that comprises one or more RNA molecules having a sequence copied from the template. For clarity, IVT optionally may include co-transcriptional capping.
[0166] In the context of the present disclosure, “modification sensitive” refers to an endoribonuclease that displays differential activity depending on the modification status of its recognition sequence in a substrate. For example, a modification sensitive endoribonuclease may display its peak activity when combined with a substrate comprising its recognition sequence, all nucleotides of which are canonical nucleotides, but display less than its peak activity (e.g., any degree of partial activity down to no activity) when combined with an identical substrate except that one (or more) of the recognition sequence nucleotides is a modified nucleotide.
[0167] In the context of the present disclosure, “modified nucleotide” refers to nucleotides having a modification on the sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or in the phosphate groups (e.g., phosphorothioates and 5'-N- phosphoramidite linkages); and / or in the nucleotide base (e.g., as described in US 8,383,340; WO 2013 / 151666; US 9,428,535 B2; US 2016 / 0032316; Cappannini et al., MODOMICS: a database of RNA modifications and related information. 2023 update, Nucleic Acids Research, gkab!083, https: / / doi.org / 10.1093 / nar / gkadl083). Modifications include both naturally- occurring modifications and artificial or synthetic chemical modifications. Examples of modified nucleotides include N6-methyl-adenosine (m6A), 1-methyl-adenosine (nfA), 5- methyl-cytidine (m5C), 5-hydroxymethyl-cytidine (hm5C), N4-acetyl-cytidine (ac4C), 5- methoxycytidine (mo5C), 4-thiouridine (S4U), 2-thiouridine (S2U), pseudouridine ( ), N1- methyl-pseudouridine (nE ), 5 -methyluridine (m5U), and 5- methoxyuridine (mo5U).
[0168] In the context of the present disclosure, “non-naturally occurring” refers to a molecule (e.g., a polynucleotide, polypeptide, carbohydrate, or lipid) or composition that does not exist in nature. Such a molecule or composition may differ from naturally occurring molecules or compositions in one or more respects. For example, a polymer (e.g., a polynucleotide, polypeptide, or carbohydrate) may differ in the kind and arrangement of the component parts (e.g., nucleotide sequence, amino acid sequence, or sugar molecules). A polymer may differ from a naturally occurring polymer with respect to the molecule(s) to which it is linked. For example, a “non-naturally occurring” polypeptide (e.g., protein) may differ from naturally occurring polypeptides in its secondary, tertiary, or quaternary structure, by having (or lacking) a chemical bond (e.g., a covalent bond including a peptide bond, a phosphate bond, a disulfide bond, an ester bond, an ether bond, and others) to a lipid, a carbohydrate, a second polypeptide (e.g., a fusion protein), or any other molecule. Similarly, a “non-naturally occurring” polynucleotide or nucleic acid may comprise (or lack) one or more other modifications (e.g., an added label or other moiety) to the 5’- end, the 3’ end, and / or between the 5’- and 3 ’-ends (e.g., methylation) of the nucleic acid. A “non-naturally occurring” molecule or composition may differ from naturally occurring compositions in one or more of the following respects: (a) having components that are not combined in nature, (b) having components in ratios and / or concentrations not found in nature, (c) lacking one or more components otherwise found in naturally occurring molecules or compositions (e.g., a cell-free composition, a chromosome- free composition, a histone-free composition, a polymerase-free composition, a cell membrane- free composition), (d) having a form not found in nature (e.g., dried, freeze dried, lyophilized, crystalline, aqueous, immobilized), and (e) having one or more additional components beyond those found in nature (e.g., a buffering agent, a detergent, a dye, a solvent or a preservative).
[0169] In the context of the present disclosure, “nucleotide-specific” refers to an endoribonuclease that cleaves RNA adjacent to a recognition sequence that is no more than 2 bases long. Examples include RNase T1 (cleaves at G), RNase A (cleaves at C and U), MCI (cleaves at U), and RNase 4 (cleaves at UA or UG). With reference to an amino acid, “position” refers to the place such amino acid occupies in the primary sequence of a peptide or polypeptide numbered from its amino terminus to its carboxy terminus. A position in one primary sequence may correspond to a position in a second primary sequence, for example, where the two positions are opposite one another when the two primary sequences are aligned using an alignment algorithm (e.g., BLAST (Journal of Molecular Biology. 215 (3): 403-410) using default parameters (e.g., expect threshold 0.05, word size 3, max matches in a query range 0, matrix BLOSUM62, Gap existence 11 extension 1, and conditional compositional score matrix adjustment) or custom parameters). An amino acid position in one sequence may correspond to a position within a functionally equivalent motif or structural motif that can be identified within one or more other sequence(s) in a database by alignment of the motifs.
[0170] In the context of the present disclosure, an amino acid sequence having a percent identity to a reference sequence may be disclosed with or without specifying one or more of the variant positions. For example, if a 100-amino acid polypeptide is disclosed as having an amino acid sequence having at least 90% identity to a 100-amino acid reference sequence, the sequence may differ from the reference sequence at any positions up to 10. If a 100-amino acid polypeptide is disclosed as having an amino acid sequence having at least 90% identity to a 100-amino acid reference sequence and having a substitution at a given position, the sequence may differ from the reference sequence as noted at that position plus at any other positions up to 9.
[0171] In the context of the present disclosure, “RNA ligase” refers to any commercially available ligase operable to covalently join a 3’ RNA end and a 5’ RNA end. Examples of RNA ligases include CircLigase, RtcB ligase, SplintR® ligase, T4 RNA ligase 1, and / or T4 RNA ligase 2. An RNA ligase may or may not be operable to covalently join a 3’ RNA end and a 5’ DNA end and / or independently may or may not be operable to covalently join a 3’ DNA end and a 5’ RNA.
[0172] In the context of the present disclosure, “sequence-specific” refers to an endoribonuclease that recognizes a motif that is at least two (e.g., at least three) ribonucleotides long including, for example, 3-, 4-, 5-, 6-, or 7-base cutters or longer. Examples of sequencespecific endoribonucleases include E. coli MazF (a 3 -base cutter that recognizes the motif, ACA) and Mycobacterium MazF-mtl (a 3-base cutter that recognizes the motif, UCA).
[0173] In the context of the present disclosure, “single-stranded RNA” refers to any ribonucleic acid having a single-stranded conformation (in whole or substantial part) characterized by unpaired bases. Examples of single-stranded RNA include a messenger RNA (mRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), small RNA (sRNA), microRNA (miRNA), long noncoding RNA (IncRNA), circular RNA (circRNA), transfer RNA (tRNA), aptamer RNA, antisense RNA, silencing RNA (siRNA), guide RNA (gRNA), therapeutic RNA, other RNAs of interest or any combination thereof and can be or arise from any desired source (e.g., human, non-human mammal, plants, insects, microbial, viral, or synthetic DNA). A singlestranded RNA may be prepared, in some embodiments by extracting RNA (e.g., total RNA, mRNA, or other RNAs) from a biological sample. A single-stranded RNA may comprise a 5’ cap and / or a poly(A) tail. A single-stranded RNA, in some embodiments, may comprise one or more intramolecular hydrogen bonds (e.g., between non-consecutive bases). A singlestranded RNA, in some embodiments, may be free of secondary structure arising from interbase intermolulcular hydrogen bonds (hydrogen bonds between a base of such polynucleotide molecule and a base of any other (separate) polynucleotide molecule). For example, a single stranded polynucleotide (e.g., a ssRNA) may be free of sequence-specific hydrogen bonds, free of Watson-Crick hydrogen bonds, free of noncannonical hydrogen bonds or free of any two of up to all three of the foregoing hydrogen bonds. A ssRNA, in some embodiments, may be free of intermolecular and intramolecular hydrogen bonds at or near (within 20 nts of) an endonuclease recognition site. A single-stranded polynucleotide (e.g., a ssRNA) may be linear or circular. For clarity, a single-stranded polynucleotide molecule may exclude or include intramolecular hydrogen bonds. Intramolecular interbase hydrogen bonds, if present in a single stranded polynucleotide, may be associated with secondary structures that resemble separate polynucleotides annealed into a double-strand. Compositions having a single-stranded RNA may lack or substantially lack RNA species of the same or similar length having a complementary or substantially complementary sequence.
[0174] In the context of the present disclosure, “site-specific” refers to an endoribonuclease that cleaves RNA at a given nucleotide or at a specific sequence of nucleotides.
[0175] In the context of the present disclosure, “structured RNA” refers to any RNA having a secondary structure that includes hydrogen-bonded bases (e.g., canonical, Watson-Crick bonds and non-canonical bonds) resulting in one or more structural elements. Example structural elements include helixes, hairpins, bulges, junctions, internal loops, multi-branched loops, pseudoknots, quadruplexes, and i-motifs.
[0176] In the context of the present disclosure, “substitution” refers to an amino acid residue at a position in a comparator amino acid sequence that differs with respect to a corresponding position of a reference amino acid sequence, where the comparator and reference sequences are at least 60% identical to each other or at least 70% identical to each other or at least 80% identical to each other. A reference sequence and comparator sequence may have the same length or similar lengths (e.g., differing by < 12%, < 5%, < 1%). A substitute amino acid residue at a position, in addition to differing from the corresponding position of a reference amino acid sequence, may differ from the amino acid at the corresponding position of all naturally-occurring sequences that are at least 60% identical to each other or at least 70% identical to each other or at least 80% identical to the reference sequence. Optionally, a substitute amino acid may have different properties than the amino acid in the corresponding position of the reference sequence. Optionally, a substitute amino acid may have similar properties to the amino acid in the corresponding position of the reference sequence (a “conservative” substitution). For example, a non-polar amino acid (e.g., A, V, L, I, M, W, and F (and optionally C, G, and P) may substitute for another non-polar amino acid, a polar amino acid (e.g., N, Q, S, T, and Y) may substitute for another polar amino acid (e.g., C, D, E, H, K, N, P. Q, R, S, and T), a positively charged amino acid (H, K, and R) may substitute for another positively charged amino acid, and a negatively charged amino acid (e.g., D and E) may substitute for another negatively charged amino acid. A substitute amino acid may be a natural amino acid (e.g., replacing another natural amino acid or a non-natural amino acid). A substitute amino acid may be a non-natural amino acid (e.g., replacing a natural amino acid or another non-natural amino acid).
[0177] In the context of the present disclosure, “therapeutic RNA” refers to an RNA molecule operable in a mammal to provide a therapeutic modality. A therapeutic RNA may be free of coding sequences and / or comprise one or more coding sequences. A therapeutic modality may be provided by direct effect of the RNA or by translation of the RNA to produce a therapeutic protein. Examples of therapeutic proteins include immunogens for stimulating a desired immune response (e.g., a vaccine), proteins that takes the place of a defective protein (e.g., collagen, phenylalanine hydroxylase) or missing protein in the subject mammal, and enzymes that provide a catalytic activity not found in the subject mammal (e.g., an chemotherapeutic enzyme that catalyzes formation of a cytotoxic compound in a tumor). Examples of therapeutic RNA include, without limitation, antisense RNA (e.g., fomivirsen, nusinersen, eteplirsen, golodirsen), small interfering RNA (e.g., patisiran, inclisiran, lumasiran, nedosiran), RNA aptamers (e.g., pegaptanib), mRNA vaccine (e.g., Comirnaty and Spikevax). A therapeutic RNA may be a non-naturally occurring molecule in view of its structure and / or composition. For example, a therapeutic RNA may comprise a non-naturally occurring nucleotide sequence, non-naturally occurring nucleotides (e.g., modified nucleotides including naturally occurring modified nucleotides in non-naturally occurring sequence positions or abundance, synthetic modified nucleotides), and / or a non-naturally occurring structure (e.g., little or no duplex RNA relative to a naturally occurring RNA with the same nucleotide sequence). Therapeutic RNA may exclude natural RNA and / or exclude extracts of natural RNA (e.g., obtained from bacteria, yeast or plants) that are not suitable for administration to a mammalian subject. In this context, a preparation may not be suitable for administration to a mammal because of compromised safety and / or efficacy, whether actual or potential. Therapeutic RNA may be cell-free, non- naturally-occurring RNA.
[0178] In the context of the present disclosure, “uncapped” refers to an RNA (a) that does not have a cap and (b) that can be used as a substrate for a capping enzyme. Uncapped RNA typically has a tri- or di -phosphorylated 5’ end. RNAs transcribed in vitro have a triphosphate group at the 5’ end.
[0179] In the context of the present disclosure, “variant endoribonuclease” refers to a non- naturally occurring endoribonuclease. A variant endoribonuclease may comprise one or more amino acids in addition to any of SEQ ID NOS: 1-139, 163, and 164. For example, a variant endoribonuclease may comprise (e.g., at its amino terminal end or carboxy terminal end) 1-25 amino acids. Such additional amino acids may enable, facilitate and / or enhance translation, expression, cellular sorting, inactivation (e.g., by including a protease recognition and / or cleavage site), and / or purification. Such additional amino acids may constitute a linker, for example, to a support (e.g., a magnetic bead) or another protein.
[0180] A variant endoribonuclease may have an amino acid sequence sharing any desired degree of sequence identity with any of SEQ ID NOS: 1-139, 163, and 164 up to (but excluding) 100% identity. For example, a variant endoribonuclease may have an amino sequence having >85%, >90%, >95%, >96%, >97%, >98%, or >99% identity to any of SEQ ID NOS: 1-139, 163, and 164, wherein each substitution is a conservative substitution or wherein all substitutions are conservative substitutions except one or wherein all substitutions are conservative substitutions except two or wherein all substitutions are conservative substitutions except three or wherein all substitutions are conservative substitutions except four or wherein all substitutions are conservative substitutions except five.
[0181] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Reagents referenced in this disclosure may be made using available materials and techniques, obtained from the indicated source, and / or obtained from New England Biolabs, Inc. (Ipswich, MA). Endoribonucleases and Compositions
[0182] The present disclosure provides nucleotide-specific or sequence-specific endoribonucleases having desired recognition motifs and / or biophysical properties (e.g., resistance to denaturing reagents and / or thermal stability) for controllable fragmentation of RNA molecules, specific examples of which are shown in TABLE 1.
[0183] TABLE 1 Example Endonucleases
[0184] In some embodiments, endoribonucleases of TABLE 1 (e.g., SEQ ID NOS: 1-139,
[0185] 163, and 164) may be tagged as shown in FIGURE 1 or tagged with a 6xHIS and a linker.
[0186] The present disclosure further provides variants of these endonucleases having, for example, the same recognition motif and comprising an amino acid sequence having >85%, >90%, >95%, >96%, >97%, >98%, or >99% identity to any of SEQ ID NOS: 1-139, 163, and
[0187] 164. A variant nucleotide-specific and / or a variant sequence-specific endoribonucleases may comprise (e.g., at its amino terminal end or carboxy terminal end) 1-25 amino acids having no counterpart in the corresponding reference sequence (e.g., expression tags, purification tags).
[0188] Compositions, in some embodiments, may exclude whole cells and / or exclude cell extracts. For example, where a composition is configured to cleave a target RNA, it may be desirable or required to exclude cells and cell extracts that may interfere. Compositions may exclude, for example, enzymes or other materials that may nick a strand of (e.g., nicking enzymes) or cleave (e.g., single-stranded and double-stranded nucleases) polynucleotides or that ligate (e.g., ligases) polynucleotides cut by sequence-specific endoribonucleases. Compositions may lack, according to some embodiments, one or more of operable cell membranes, ribosomes, nucleus, cytoplasm, mitochondria, and / or a cell wall.
[0189] According to some embodiments, an endoribonuclease composition may comprise an endoribonuclease and, optionally, any of (including one or more of) a buffering agent (e.g., a storage buffer, a reaction buffer), an excipient, a salt (e.g., NaCl, MgCh, CaCh), a protein (e.g., albumin, polymerase), a stabilizer, a detergent (for example, ionic, non-ionic, and / or zwitterionic detergents (e.g., octoxinol, polysorbate 20), a polynucleotide, a cell (e.g., intact, digested, or any cell-free extract), a biological fluid or secretion (e.g., mucus, pus), an aptamer, a pH indicator (e.g., azolitimin, bromocresol purple, bromothymol blue, methylene blue, cresol red, neutral red, naphtholphthalein, phenol red), a crowding agent, a sugar (e.g., a mono, di, tri, tetra, or higher saccharide), a starch, cellulose, a glass-forming agent (e.g., glycerol, raffinose, stachyose, or trehalose for lyophilization), a lipid, an oil, aqueous media, a support (e.g., a bead) and / or (non-naturally occurring) combinations thereof. Combinations may include for example, two or more of the listed components (e.g., a salt and a buffer) or a plurality of species of a single listed component (e.g., two (or more) different salts or two (or more) different sugars). According to some embodiments, endoribonuclease compositions may comprise (a) an endoribonuclease, and (b) a polyribonucleotide (e.g., a known endoribonuclease substrate, a candidate endoribonuclease substrate, a test sample comprising or potentially comprising an endoribonuclease substrate), including, for example, a single or double stranded RNA, a DNA-RNA hybrid duplex, a fluorescent probe (e.g., TABLE 2)). A composition, in some embodiments, may comprise one or more modified nucleotides (e.g., base modified nucleotides, sugar modified nucleotides, and / or labeled nucleotides) and / or one or more modified nucleotide linkages (e.g., thiol linkages, azido linkages). An RNA strand may comprise, for example, one or modified nucleotides and / or nucleotide linkages.
[0190] Modified nucleotides and / or linkages may be included, for example, to increase or decrease the stability of a duplex molecule and / or increase or decrease susceptibility of a strand to cleavage by an endoribonuclease. Modified nucleotides include, for example, a 2'-O-CH3 modification, a 2'-F modification, a 2'-M0E modification, a locked nucleic acid (LNA), an unlocked nucleic acid (LINA), deoxyuridine, pseudouridine, 5-methylcytosine, 2- aminopurine, 2,6-diaminopurine, deoxyinosine, 5-hydroxybutynl-2'-deoxyuridine, 8-aza-7- deazaguanosine, 5 -nitroindole.
[0191] Kits
[0192] The present disclosure further relates to kits including an endoribonuclease (e.g., a nucleotide specific or a sequence specific endonuclease disclosed here) and, optionally, one or more additional components. For example, a kit may include an endoribonuclease and dNTPs, rNTPs, primers, other enzymes (e.g., polymerases, ligases, other nucleases), buffering agents, or combinations thereof. Enzymes may be included in a storage buffer. Any suitable storage buffer may be used, for example, buffers comprising one or more of a cryoprotectant (e.g., a polyol such as glycerol, an antifreeze protein), a salt, a detergent, a reducing agent, a sugar, a chelator, and an antimicrobial agent and having a pH tolerated by the enzyme to be stored, for example, between pH 6 and 9. A composition or kit may include a reaction buffer which may be in concentrated form, and the buffer may contain additives (e.g. glycerol), salt (e.g. NaCl, KC1), reducing agent, EDTA or detergents, among others. Detergents include nonionic detergents (e.g., t-octylphenoxypoly ethoxy ethanol), anionic detergents (e.g., alkylbenzene sulfonates), cationic detergents (e.g., alkylbenzene quaternary ammonium), and zwitterionic detergents. A composition or kit comprising dNTPs may include one, two, three of all four of dATP, dTTP, dGTP and dCTP. A kit comprising rNTPs may include one, two, three of all four of rATP, rUTP, rGTP and rCTP. A kit may further comprise one or more modified nucleotides. A kit may optionally comprise one or more primers (random primers, bump primers, exonuclease-resistant primers, chemically-modified primers, custom sequence primers, or combinations thereof). In some embodiments, a kit may include an endonuclease substrate, examples of which are shown in TABLE 2, for use as a control or other purposes.
[0193] TABLE 2 Example Endonuclease Substrates
[0194]
[0195]
[0196]
[0197]
[0198] A kit may be a non-natural collection of components configured, for example, for convenient storage, shipping, delivery, and / or use. One or more components of a kit may be included in one container for a single step reaction, or one or more components may be contained in one container, but separated from other components for sequential use or parallel use. The contents of a kit may be formulated for use in a desired method or process.
[0199] A kit is provided that contains: (i) an endoribonuclease; and (ii) a buffer. An endoribonuclease may have a lyophilized form and / or an immobilized form (each with or without a buffer composition) or may be included in a fluid (e.g., aqueous media) with buffer (e.g., a storage buffer or a reaction buffer in concentrated form). A kit may contain the endoribonuclease in a mastermix suitable for receiving and amplifying a template nucleic acid. An endoribonuclease may be a purified enzyme so as to contain substantially no DNA, no RNA, and / or no nucleases. A reaction buffer in (ii) and / or storage buffer containing the RNA polymerase in (i) may include non-ionic, ionic e.g. anionic or zwitterionic surfactants and crowding agents. A kit may include an endoribonuclease and a reaction buffer in a single tube or in different tubes. A kit may include, in some embodiments, an RNase inhibitor, for example, an RNase inhibitor selected to reduce / block activity of an undesired RNase (or class of RNases) without undue interference with the endoribonuclease of the kit. For example, a kit may include an RNase A inhibitor.
[0200] A subject kit may further include instructions for using the components of the kit to practice a desired method. The instructions may be recorded on a suitable recording medium. For example, instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or subpackaging) etc. Instructions may be present as an electronic storage data file residing on a suitable computer readable storage medium (e.g. a CD-ROM, a flash drive). Instructions may be provided remotely using, for example, cloud or internet resources with a link or other access instructions provided in or with a kit.
[0201] Methods
[0202] The present disclosure provides methods of using endoribonucleases. According to some embodiments, a method may include contacting an RNA substrate molecule and an endoribonuclease (e.g., any of the endoribonucleases disclosed herein). For example, an endoribonuclease may have an amino acid sequence according to any of SEQ ID NOS: 1-139, 163, and 164 or variants thereof. Variant endoribonucleases may comprise, for example, an amino acid sequence having >95%, >96%, >97%, >98%, or >99% identity to any of SEQ ID NOS: 1-139, 163, and 164. An RNA substrate molecule may have any desired length, for example, a length in a range from a to Z>, where a is 5, 10, 20, 30, 40, 50, 75, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, 500, 600, or 750 nucleotides, b is 50, 100, 250, 500, 750, 1,000, 2,500, 5,000, 7,500, 10,000, 12,500, 15,000, 17,500, or 20,000 nucleotides, and a < b. An RNA substrate molecule, in some embodiments, may comprise one or more (e.g., all) canonical bases and / or one or more modified nucleotides. For example, an RNA substrate molecule may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 modified nucleotides. An RNA substrate molecule may comprise >1, >2, >3, >4, >5, >6, >7, >8, >9, >10, >15, >20, >25, >30, >35, >40, >45, >50 or >55 mole percent modified nucleotides.
[0203] In some embodiments, all instances of a given nucleotide in an RNA substrate molecule may be modified nucleotides. For example, all A’s may be modified A’s, all C’s may be modified C’s, all G’s may be modified G’s, and / or all U’s may be modified U’s. As one example, an RNA substrate may be made according to a reference sequence consisting of the 4 canonical ribonucleobases except that every uridine in the reference sequence is occupied by a pseudouridine in the RNA substrate. Such an RNA substrate may be prepared, for example, in an IVT reaction comprising ATP, GTP, CTP and TTP (instead of UTP). In some embodiments, it may be desirable to prepare or provide a therapeutic RNA with a modified nucleotide in place of up to all positions of the corresponding unmodified nucleotide.
[0204] Contacting one or more endoribonucleases with an RNA substrate may support determination of one or more properties or characteristics of the RNA substrate. For example, methods may include sequence mapping an RNA substrate. In some embodiments, contacting an RNA substrate (e.g., a synthetic RNA, a cellular RNA) with nucleotide-specific and / or sequence-specific endoribonucleases may provide controlled cleavage of the RNA substrate. In this context, controlled cleavage refers to cleavage of the substrate RNA based on the sequence, modification status, cap status, and / or three-dimensional structure (e.g., single stranded, double stranded, hairpin, bulge, loop) to produce (a limited number of) cleavage products. The number of cleavage fragments produced by controlled cleavage, which will depend at least in part on the number of cleavage recognition sites in a given RNA substrate. Controlled cleavage may be achieved, in some embodiments, by choosing one or more sequence-specific endoribonucleases to target specific cleavage recognition sequences in a given RNA. If the size but not the sequence of a target RNA is not known, controlled cleavage may be achieved by selecting one or more endoribonucleases with known specificities (e.g., 2-base, 3-base, 4-base, 5-base cutter). Controlled cleavage (e.g., with one or more sequence-specific endoribonucleases) may produce cleavage products having an average length of >5, >10, >15, >20, >25, >30, >40, >50, >75, >100, >150, >200, >250, >300, >350, >400, >450, >500 or >1000 nucleotides. A non-specific ribonuclease, by contrast, may produce cleavage products having an average length of <5 nucleotides. Controlled cleavage of an RNA substrate of interest can be used for characterization of the RNA sequence (e.g., the identity of the RNA), for assessment of the integrity and / or purity of the RNA. In some embodiments, fewer and longer cleavage products may enable better characterization. For example, longer fragments can improve the accuracy of assembling sequencing reads into a complete RNA sequence. Controlled RNA cleavage may also provide detection of chemical modifications (e.g., chemical modifications incorporated synthetically in a synthetic RNA substrate during transcription using a modified nucleotide triphosphate and / or chemical modifications arising or induced post-transcriptionally using one or more of a chemical agent, an enzyme, temperature, and radiation). In some embodiments, endoribonucleases may be coupled with mass spectrometry or electrophoretic analysis for characterization and quality control of RNAs synthesized in vitro (e.g., a therapeutic RNA). Endoribonucleases may also be coupled with mass spectrometry or electrophoretic analysis for characterization of natural RNAs synthesized in vivo (e.g., RNA contained in or obtained from a microbe, a plant, or an animal including, for example, cells, tissue preparations and bodily fluids like blood, saliva, sinus fluids, urine, sweat). For example, a method may include contacting an RNA substrate (e.g., a synthetic RNA, a natural RNA) with an endoribonuclease to produce cleavage products, analyzing one or more of the cleavage products by mass spectrometry (e.g., LC-MS) or electrophoresis (e.g., capillary electrophoresis).
[0205] A method, in some embodiments, may be performed as a quality control check on a synthetic RNA. For example, a method may comprise contacting an endoribonuclease (e.g., a nucleotide specific endoribonuclease or a sequence specific endoribonuclease) having a cleavage recognition sequence comprising a modified nucleotide (e.g., pseudouridine) and an RNA substrate (e.g., a therapeutic RNA) to form cleavage products and analyzing at least one of the cleavage products to determine whether the RNA substrate had the desired pseudouridines in the right positions. For example, some therapeutic RNAs may comprise pseudouridine (e.g., to the exclusion of uridine). As a quality check, such an RNA may be contacted with a nucleotide specific endoribonuclease or a sequence specific nuclease and products may be analyzed for the presence of fragments that exceed predicted sizes (e.g., indicative of the omission of a pseudouridine where one should have been).
[0206] The present disclosure provides, in some embodiments, methods for cleaving singlestranded RNA and / or double-stranded RNA. For example, the present disclosure provides thermophilic and hyperthermophilic sequence specific endoribonucleases. A method may include contacting an endoribonuclease and a dsRNA substrate at a temperature and / or other conditions (e.g., urea, formamide) that melts or otherwise disrupts the duplex substrate (e.g., T >65 °C, or >75 °C, or >85 °C, or >95 °C).
[0207] In some embodiments, a method may comprise contacting an RNA substrate having a 5’ end (e.g., a linear RNA substrate) with an endoribonuclease that cleaves the substrate within 100 nucleotides of the 5’ end to produce cleavage products comprising a 5’ end fragment and analyzing the capping status of the 5’ end fragment. For example, a method may comprise contacting (a) a substrate RNA having a 5’ end and comprising a nucleotide sequence with at least one copy of the sequence AAAANAU near (e.g., within 100 nucleotides) the 5’ end and (b) a type III RNase (toxIN) of Methanosphaera stadtmanae (designated pey214 and having an amino acid sequence according to SEQ ID NO: 132), which recognizes the “AAAANAU” motif in RNA to produce cleavage products and analyzing at least one of the cleavage products. Such cleavage products may be like restriction enzyme like fragment(s) that are amenable to mass spectrometry analysis. Methods disclosed herein are more accurate than methods based on RNase H which requires a DNA-RNA chimeric probe and often cleaves at n-1, n. n+ positions (wherein n is the expected cleavage position) in an unpredictable way. Methods disclosed herein provide a more convenient option than methods based on shorter cutters (e.g., RNase Tl, RNase 4) in that no exogenous probe is required to produce 5’ cleavage products of a desired length. Often there will be multiple cleavage sites for such short cutters in any given non-engineered sequence thus resulting in cleavage products that are too short or not ideally suited for mass spectrometry or electrophoretic analysis.
[0208] The present disclosure also provides methods for producing RNA having one or more desired properties. For example, it may be desirable to manufacture single-stranded RNA that is with little to no double-stranded content along its length (e.g., for therapeutic applications). One approach to producing such RNA is to use an RNA polymerase that (intrinsically) produces little (e.g., nominal to no detectable) untemplated sequence at the 3’ end that is susceptible to forming duplex structures. An alternate or supplemental approach is to cleave undesired 3’ sequences. In some embodiments, a method may comprise contacting an RNA substrate with an endoribonuclease to generate a population of cleavage products having homogeneous 3’ ends. For example, a method may comprise producing (e.g., by in vitro transcription) an RNA substrate comprising, in a 5’ to 3’ direction, a sequence of interest and the cleavage recognition sequence of an endoribonuclease (e.g., AAAANAU) and contacting the RNA substrate with a type III RNase (toxIN) of Methanosphaera stadtmanae, which is active in IVT buffer and recognizes the “AAAANAU” motif in the RNA substrate to produce cleavage products comprising RNA with homogeneous 3’ ends and 3’ fragments, wherein the 3’ fragments optionally may be removed, for example, by size exclusion chromatography, affinity chromatography, capillary electrophoresis, enzymatic means (e.g., 5'-3' exonucleases that cannot cleave cap protected RNA or 5'-3' exonucleases that cannot cleave IVT RNA 5' ends comprising triphosphate) or other means.
[0209] According to some embodiments, a method may include contacting an endoribonuclease with a substrate comprising, in a 5’ to 3’ direction, an affinity capture sequence and a recognition sequence of the endoribonuclease to produce cleavage products (e.g., comprising upstream (of the cleavage site) products and downstream (of the cleavage site) products, the upstream cleavage products comprising the affinity capture sequence). Methods may include contacting the cleavage products with an immobilized support comprising or linked to an affinity partner of the affinity capture sequence to produce captured upstream cleavage products, the captured upstream cleavage products comprising the affinity capture sequence bound to the affinity partner. Example affinity capture sequences include, for example, aptamer sequences that bind to streptavidin or bacteriophage MS2 and MS2-like sequences that bind to MS2 coat protein (also referred as to MS2 tagging). These designed 3’ end can be encoded on the corresponding DNA templates that are used to generate the RNA of interest.
[0210] Methods include, according to some embodiments, contacting an endoribonuclease with a non-naturally occurring substrate. Where a specific cleavage or cleavage pattern is desired, a non-naturally occurring substrate may be engineered to add or to exclude one or more endoribonuclease recognition sites and / or one or more endoribonuclease recognition sites may be masked (e.g., by contacting the substrate and a DNA splint and / or a binding protein). If release of one or more fragment of a particular size and / or a particular sequence composition is desired, the substrate sequence may be modified to install one or more recognition sites (e.g., flanking the sequence of interest) or remove one or more recognition sites (e.g., where the sequence of interest has one or more internal recognition sequences that would confound creation of fragments comprising the full-length sequence).
[0211] The present disclosure further provides methods of cleaving and rejoining singlestranded RNAs. For example, a method may include contacting a substrate comprising one endoribonuclease recognition site and a corresponding endoribonuclease to produce cleavage products comprising a 5’ cleavage product and a 3’ cleavage product and contacting one or both products with an RNA ligase to form a ligation product. In some embodiments, a substrate may comprise one or more recognition sites for a given endoribonuclease and / or may comprise one or more recognition sites for one or more endoribonucleases. In some embodiments, a method may comprise contacting at least one cleavage product with an additional single- stranded polynucleotide (DNA, RNA) to form one or more chimeric ligation products. For example, a method may comprise contacting an RNA substrate of interest (e.g., as a simple composition having a single RNA substrate species or a complex composition having a one or more additional RNA substrate species beyond the species of interest) and a sequence-specific endoribonuclease to form cleavage products. A method may further include contacting at least one cleavage product (e.g., up to all cleavage products) with an additional single-stranded polynucleotide (e.g., an RNA adapter, a DNA adapter, a labeled polynucleotide) to form one or more ligation products.
[0212] EXAMPLES
[0213] Some specific example embodiments may be illustrated by one or more of the examples provided herein.
[0214] EXAMPLE 1 : Expression and purification of ribonucleases
[0215] Plasmids encoding endoribonucleases were codon optimized for E. coli and synthesized (Twist Bioscience, San Francisco, CA). E. coli cells (New England Biolabs, Inc., strain C2529, [can::CBD fliuA2 [Ion] ompT gal ( DE3) [dem] arnA::CBD slyD::CBD glmS6Ala \hsdS 2 DE3 = 2 sBamHIo \EcolU-B int::(lacE. :PlacUV5::T7 genel) i21 \nin5l) were transformed with periplasmic expression plasmids carrying a subject ribonuclease gene and grown overnight in LB-Kanamycin at 30°C. Secretion of ribonucleases into periplasmic space obviates the need for co-expression of neutralizing anti-toxin. The outline of an example construct is depicted in FIGURE 1. Individual colonies were inoculated in 5 mL of Dynamite media (12 g soy peptone, 24 g yeast extract, 6.3 mL glycerol, 3.8 g KH2PO4, 5 g glucose, 0.195 g MgSCh per liter) containing 80 pg / mL Kanamycin and 40 pM IPTG in culture tubes. Cultures were shaken at 300 RPM for 40 hours at 30°C. Cells were subsequently collected via centrifugation washed with PBS and frozen at -20°C. Frozen cell pellets were resuspended in 800 pL ice-chilled lysis buffer (500 mM NaCl, 20 mM Tris HC1, pH 7.5) containing protease inhibitors (1 mM PMSF, 0.5 nM Leupeptin, 275 mM benzamidine, 2 nM pepstatin). Cells were lysed via sonication and clarified by centrifugation. Clarified lysate was applied to 20 pL of amylose magnetic beads (New England Biolabs, Inc., E8035S) and mixed at 4°C for Ih in deep-well 96-well plates (Thermo Scientific, 95040450). After incubation the 96-deep well plates were transferred to the KingFisher Flex (Thermo Scientific) for automated purification. Magnetic beads were collected and washed twice with 1 mL wash buffer (1 M NaCl, 20 mM Tris-HCl pH 7.5) and once with 1 mL lysis buffer. Bound proteins were eluted using 100 pL of elution buffer (10 mM maltose, 500 mM NaCl, 20 mM Tris-HCl pH7.5). Purification of the target protein was confirmed via SDS-PAGE and Coomassie G-250 (Themo Fisher Scientific, LC6065). The eluate was mixed with an equal volume of 100% glycerol and stored at -20 °C.
[0216] EXAMPLE 2: Evaluating enzyme activity
[0217] To test the activity of purified ribonucleases, an RNA substrate containing all the theoretical combinations of 5 nucleotides (1,024 different 5-mers) occurring one or more times (SEQ ID NO: 139) was designed. This substrate includes 1,534 distinct 6-mer sequences, out of 4,066 theoretical 6-mer; and 1,718 distinct 7-mer sequences, out of 16,384 theoretical 7-mer sequences. Purified enzymes are expected to cleave the substrate RNA if the substrate includes the enzyme’s recognition sequence motif for the respective enzyme. The RNA substrate (named pey43, SEQ ID NO: 145) was generated using canonical ribonucleotide triphosphates by HiScribe T7 High Yield RNA synthesis kit (New England Biolabs, Inc., E2040S) per the manufacturer’s instructions. Synthesized RNA was treated with DNase I (New England Biolabs, Inc., M0303S) and column purified (New England Biolabs, Inc., T2050S). RNase activity was tested in the buffer containing 20 mM Tris, pH 7.5. 1 pL of 5 pM purified ribonuclease was used to cleave 400 nanogram (0.7 pmol) of RNA substrate (SEQ ID NO: 145) at 37°C for 30 minutes in total 10 pL of reaction volume in the presence of 10 Units of Murine RNase inhibitor (New England Biolabs, Inc., M0314L). Reactions were terminated by incubating with thermolabile proteinase K (New England Biolabs, Inc., P81 IS) at 37°C for 15 minutes followed by mixing with 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S). Afterwards, samples were denatured at 70°C for 5 minutes prior to loading onto a 6% Urea gel and run at 180Volt in IX TBE buffer. Gels were stained for 20 minutes in IX TBE buffer containing SYBR Gold stain (Themo Fisher, SI 1494) diluted 20,000-fold. Gels were rinsed to remove excess stain and imaged with 300 nm ultraviolet light or with laser scanner using Cy2 filter. A candidate protein was considered an active enzyme when the substrate RNA was fragmented relative to a control reaction in the absence of a ribonuclease (NE, no enzyme). Example results of such assays are shown in FIGURE 3A, B, C. The activity of the remaining candidate ribonucleases (not shown in FIGURE 3A, B, C) was assessed the same way.
[0218] Under the conditions tested, enzymes SEQ ID NOS: 83, 86, 89, and 133 displayed weak cleavage activity but cleavage was specific. Some enzymes exhibited very strong, highly sequence specific cleavage activity, e.g. SEQ ID NOS:79-82, 84, 85, 87, 88, 90-94. Some enzymes cleaved RNA into various fragment sizes suggesting that these are frequent cutters, e.g. SEQ ID NOS: 80 and 92. Some other enzymes produced fewer RNA fragments suggesting that these are less frequent cutters, e.g. SEQ ID NOS: 85, 88, and 89. Another subgroup of enzymes did not exhibit any convincing activity under the conditions tested. These enzymes may recognize motifs that are longer than 5-nucleotides or longer than 6- nucleotides or longer than 7-nucleotides.
[0219] EXAMPLE 3 : Determination of recognition motifs of ribonucleases
[0220] To determine the recognition motifs of ribonucleases, a library of synthetic RNA- DNA chimeras was chemically synthesized (Integrated DNA Technologies) each of which contained, in a 5’ to 3’ direction, an Illumina DNA adapter, a 3 nucleotide DNA flanking sequence, one of all possible RNA heptamers (7-mers), a second 3 nucleotide DNA flanking sequence, and a second Illumina DNA adapter. The oligonucleotide library containing all RNA heptamers (16,384 distinct sequences) was treated with ribonucleases to form cleavage products, which were then copied to cDNA using Luna Reverse Transcriptase (New England Biolabs, Inc.) and amplified by PCR using primers specific to Illumina adapters. The amplification products were then sequenced using Illumina sequencers. Sequencing confirmed the presence of all 16,384 RNA heptamers in the oligonucleotide library. Oligonucleotides cleaved by the ribonucleases were not expected to be fully copied by reverse transcriptase and not amplified by PCR (cleavage of a given oligonucleotide by a ribonuclease resulted in loss of that sequence or underrepresentation of it in Illumina sequencing library).
[0221] After Illumina sequencing and data analysis, motifs were inferred from underrepresented or absent 7-mer sequences in the sequenced library (compared to the input library). Underrepresented sequences were converted to sequence logos using The MEME suite software (meme-suite.org / meme / ). The motifs identified represented either a full recognition motif of ribonucleases, or a subset sequence of recognition motif. An outline of this approach is shown in FIGURE 4. Identified motifs for each enzyme are indicated in text form in TABLE 1. “N” refers to A, U, C, or G whose degeneracy level may vary for each RNase.
[0222] EXAMPLE 4: Determining cleavage position of ribonucleases within the recognition motif The exact cleavage position of ribonucleases was assessed utilizing a mass spectrometry-based cleavage assay. For each enzyme, a chemically synthesized 24 nucleotide long oligonucleotide comprising the predicted recognition motif was prepared (TABLE 2, SEQ ID NOS: 140 and 141). Buffered reactions (20 mM Tris, pH 7.5) were performed by combining 2 pL of 10 pM of oligonucleotide substrate and 1 pL of 5 pM endoribonuclease in a total volume of 20 pL for 30 minutes at 37°C. Each reaction was supplemented with 10 units of murine RNase inhibitor (New England Biolabs, Inc., M0314L). The masses of the resulting oligonucleotide cleavage products and undigested input oligonucleotides were measured by liquid chromatography mass spectrometry (LC-MS). Briefly, oligonucleotide analysis was performed on an Eclipse Fusion Orbitrap mass spectrometer (Thermo Fisher Scientific, Waltham, MA, USA) with a Vanquish ultra-high performance liquid chromatography system (UHPLC) (Thermo Fisher Scientific) equipped with an ACQUITY Premier Oligonucleotide C18 Column (Waters Corporation, Milford, MA, USA) (2.1 x 100 mm, 1.7 pm). A 23-min gradient of buffer A (1% hexafluoroisopropanol (HFIP), 0.1% N,N- diisopropylethylamine (DIEA), 1 pM EDTA) and 7 to 35% buffer B (90% Methanol, 10% water, 0.075% HFIP, 0.0375% DIEA, 1 pM EDTA) at a flow rate of 400 pL / min at 60°C. Intact oligonucleotide mass data were collected at a resolution of 60,000 (FWHM) at m / z 200. Cleavage products and input oligonucleotides with corresponding intact deconvoluted masses detected by LC-MS were annotated. The relative intensity of cleavage products was utilized to define the primary cleavage site of each ribonuclease in each target motif. The results are shown in FIGURE 5. Taken together, these data show that each of the tested ribonucleases cleaves RNA primarily at one specific position within each predicted recognition motif.
[0223] EXAMPLE 5: Assessing activity of ribonucleases on substrates containing RNA modifications.
[0224] To assess the modification sensitivity of endoribonucleases, the RNA substrate from EXAMPLE 2 (SEQ ID NO: 145) was generated by in vitro synthesis using a nucleotide triphosphate comprising one of the following nucleobase modifications m6A, nEA, m5C, hm5C, ac4C, s4U, s2U, , mlvP, or mo5U; in each case, to the exclusion of the corresponding canonical nucleotide triphosphates (ATP, GTP, CTP, UTP). For example, to generate m6A modified RNA, m6ATP was used instead of canonical ATP in the in vitro transcription reaction (together with GTP, CTP, and UTP), thus resulting RNA substrate containing m6A (i.e., all A positions according to the given template are occupied by m6A). Using this approach, 11 different modified RNA substrates were synthesized, each containing one of m6A, nEA, m5C, hm5C, ac4C, s4U, s2U, , mlT, or mo5U. As illustrated in FIGURE 6A, endoribonucleases were contacted with substrates and the cleavage products were evaluated on a denaturing urea gel. An example of such an assay is demonstrated in FIGURE 6B. In this control experiment, 1 pL of E. coli RNase MazF (Takara, 2415 A) was contacted with either 800 ng (1.39 pmol) of unmodified RNA substrate (SEQ ID NO: 145) or 800ng (1.39 pmol) of m6A-modified RNA substrate (SEQ ID NO: 145) in a total volume of 20 pL reaction at 37°C for 15 minutes; then cleavage efficiency of unmodified and modified RNA was evaluated. For all other RNases, 2 pL of 1 to 10 pM of an RNase was contacted with 800 ng (1.39 pmol) of RNA at 37°C for 15 minutes in a total volume of 20 pL reaction buffer (40mM Sodium phosphate, pH 7.5, 0.01% Tween-20).
[0225] Cleavage efficiency was classified as follows:
[0226] (i) insensitive (i.e., cleavage was observed with modified RNA as well as the corresponding unmodified RNA),
[0227] (ii) partially sensitive (i.e., detectably less cleavage was observed with modified RNA than with the corresponding unmodified RNA), and
[0228] (iii) fully sensitive (i.e., cleavage was not observed with modified RNA, but was observed with the corresponding unmodified RNA).
[0229] Partially sensitive (ii) and fully sensitive (iii) designations may also be referred to as partially blocked and fully blocked, respectively. The results for the tested enzymes are shown in FIGURE 6C. A ribonuclease that is fully sensitive or partially sensitive to one or more modifications can be used as an enzymatic tool to map the location of those modifications within an RNA sequence (e.g., through sequencing approaches similar to those described in EXAMPLE 3, wherein the motifs that contain modifications are overrepresented relative to the same sequence motifs containing only unmodified bases; the latter are dropped out of the library due to ribonuclease cleavage).
[0230] EXAMPLE 6: Assessing the effect of RNA modifications located within endoribonuclease recognition motifs.
[0231] Substrates produced in accordance with Example 5 may have a recognition motif comprising more than one modified nucleotide. For example, E. coli MazF ribonuclease recognizes an “ACA” motif. If a substrate is synthesized with m6ATP (to the exclusion of ATP), the resulting recognition motif would contain two modified nucleotides, namely {m6A}C{m6A}.
[0232] To assess the activity of endoribonucleases on a modification located on a single nucleotide on a specific position of a recognition motif, FAM labeled, short oligonucleotides were prepared, each containing a single modified nucleotide within the recognition motif. For example, to assess the activity of E. coli MazF on modified A and C, oligonucleotides containing single modification in the recognition motif {m6A}CA or A{m5C}A or AC{m6A} were synthesized. Cleavage efficiency of endoribonucleases was determined by comparing cleavage of an unmodified RNA relative to that of a modified RNA using the FAM signal as readout (FIGURES 7A-7C). Briefly, 1 pL of 5 pM oligonucleotide (one of SEQ ID NOS: 147-150 per reaction) was incubated with 1 pL of 1 pM ribonuclease pey77 (SEQ ID NO:7) or pey80 (SEQ ID NO:8), 1 pL of 10X Tris reaction buffer (200 mM) and 8 pL of nuclease-free water at 37 °C for 15 minutes. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 4 pL of each sample were then loaded on a pre-cast 20% polyacrylamide TBE- Urea gel that was run for 70 minutes at 180V before imaging on the Amersham Typhoon using the Cy2 setting. The results indicate that modifications at the cleavage sites may have affected some ribonucleases (SEQ ID NO: 148, enzymes SEQ ID NOS:7-8), while modifications in the vicinity of the cleavage site may affect only a few ribonucleases (SEQ ID NO: 149 [=oligo ID R502], enzyme SEQ ID NO:7 [=enzyme pey77]; FIGURE 7A). Certain modifications may not affect ribonucleases in any substantial way (SEQ ID NO: 150[=oligo ID R503], SEQ ID NOS:7-8 [=enzymes pey77 and pey80]; FIGURE 7A). Understanding the cleavage sensitivity of ribonucleases may allow their use for baseresolution mapping of individual modifications.
[0233] Oligonucleotides comprising single modification in the UGG recognition motif were also synthesized. These included pseudouridine, 4-thiouridine, 5-methyluridine, 1- methylpseudouridine, 5,6-dihydrouridine, N3-methyluridine, 5 -bromouridine, 5- fluorouridineand 5 -iodouridine. 1 pL of 10 pM of each of these oligonucleotides were incubated with 1 pL of either 0.05 pM, 0.1 pM, 0.2 pM or 1 pM of ribonuclease enzymes have the amino acid sequence of SEQ ID NO: 164 [=enzyme pEY546] or SEQ ID NO: 163 [=enzyme pEY586], 1 pL of 10X Tris reaction buffer (200 mM, pH 7.5) and 8 pL of nuclease-free water at 85 °C for 15 minutes. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 15% polyacrylamide TBE- Urea gel that was run for 70 minutes at 180V before imaging on the Amersham Typhoon using the Cy2 setting. Like with the ACA cleaving ribonucleases, the results indicate that cleavage activity of the two enzymes tested was not affected by the presence of certain modified nucleotides (SEQ ID NO: 166 [=OLIGO ID R540]) while activity of one enzyme was affected more than the other for other modifications (SEQ ID NO: 165 [=oligo ID R511 ], SEQ ID NO: 167 [=oligo ID R542]; FIGURES7B and 7C).
[0234] EXAMPLE 7: Controlled digestion of long RNA molecules synthesized by in vitro transcription.
[0235] Ribonucleases with little to no sequence or nucleotide-specificity produce RNA cleavage products having a large fraction of short RNA fragments. For example, cleavage products may have an average size of < 5 nucleotides. Fragments this short are not amenable to developing meaningful sequence information from mass spectrometry analysis and are not amenable to next generation sequencing such as Illumina and Nanopore sequencing or to electrophoretic fragment analysis. Attempts to control the length of RNA fragments by titrating down the ratio of enzyme-to-substrate, or reducing the incubation time, or physically removing the ribonuclease from the reaction vessel (e.g., using magnetic beads) are not always reliable, are difficult to control, and produce inconsistent results. For example, cleavage patterns and the species of fragments produced may not be reproducible (e.g., due to variations in ribonuclease activity from lot-to-lot, shelf stability or storage conditions; variations in the length and / or modification status of the RNA substrate; variations in the quality or self-stability of buffers used in the reaction; among many other factors). Cleavage products of long RNA substrates (e.g., > 1000 nt) using ribonucleases with little to no sequence or nucleotide-specificity may be highly complex in that many of the fragment species within the population of products may have identical sequences but trace their origin to multiple locations within the original substrate.
[0236] Endoribonuclease cleavage products produced in a more controlled, predictable, repeatable manner would be more compatible with mass spectrometry analysis and / or next generation sequencing and / or electrophoretic fragment analysis. For example, fragments of 4-40 nt may be conveniently sized for mass spec sequence mapping, fragments of 20-200 nt may be conveniently sized for RNA fingerprinting, and fragments of 50-10,000 nt may be conveniently sized for electrophoretic fragment analysis (e.g., gel electrophoresis or capillary electrophoresis). The present disclosure provides example ribonucleases that cleave RNA at distinct sequences allowing for controlled cleavage of RNA into fragments of a desired size (e.g., short, medium, and long fragments) to support different types of analysis. Results using examples of enzymes displaying such controllable cleavage are shown in FIGURE 8.
[0237] In these reactions, RNases were expressed and purified as described in EXAMPLE 1. Each RNase (1 pL of 1 pM to 5 pM enzyme) was incubated with 400 ng (0.7 pmol) of pey43 RNA substrate (SEQ ID NO: 145) in a total volume of 10 pL in 20 mM Tris, pH 7.5. Each reaction contained 10 Units Murine RNase inhibitor (New England Biolabs, Inc., M0314S). After incubation at 37°C for 30 minutes, reactions were terminated by incubating with thermolabile proteinase K (New England Biolabs, Inc., P81 IS) at 37°C for 15 minutes. Samples were mixed with 11 pL of 2X RNA loading dye (New England Biolabs, Inc.), denatured at 70°C for 5 minutes prior to loading onto a 6% TBE-Urea gel and ran at 180Volt in IX TBE buffer. Gels were stained in IX TBE buffer for 20 minutes which contained SYBR Gold stain at 20,000-fold dilution. Gels were rinsed to remove excess stain and imaged with laser scanner using Cy2 filter or UV light.
[0238] FIGURE 8 shows cleavage results with SEQ ID NO: 100 (pey333), SEQ ID NO: 104 (pey338), SEQ ID NO: 104 (pey339), SEQ ID NO: 105 (pey340), which yielded long fragments on average. In fact, these RNases seem to cleave 1800 nt long pey43 substrate in one to three sites as apparent from the pattern of the cleavage products. Pey338 and pey339 are the identical protein sequences but represent independently expressed and purified RNases, thus expected to give a similar cleavage pattern. FIGURE 8 also includes cleavage results with RNases having SEQ ID NO: 101-103 (pey335-337), which yielded medium size RNA fragments on average, under similar conditions.
[0239] FIGURE 8B and FIGURE 8C show the results of an additional set of cleavage reactions performed to produce a population of cleavage products for mass spectrometrybased sequence mapping. In these reactions, 5 pg of EPO (SEQ ID NO: 159), FLuc (SEQ ID NO: 160), BGII (5 kB) (SEQ ID NO: 170) or BGII (9kB) (SEQ ID NO: 162) IVT mRNA substrates were contacted with 400 nM of pEY586 (SEQ ID NO: 163) in a total volume of 10 pL in 10 mM Tris pH 8.0, 125 mM NaCl. Each reaction was incubated at 80°C for 15 minutes. Oligonucleotide cleavage products were purified by ethanol precipitation. The resultant cleavage products were characterized by ion pairing RP-LC-MS / MS and annotated with BioPharmaFinder 5.1 in Oligonucleotide mode. High-confidence oligonucleotide identifications were those detected within 5 ppm of the expected oligonucleotide monoisotopic with a confidence score greater than 90.
[0240] FIGURE 8B shows the length distribution of high-confidence oligonucleotide identifications mapping to EPO (SEQ ID NO: 168), FLuc (SEQ ID NO: 169), BGII (4.9 kB) (SEQ ID NO: 170) or BGII (8.9 kB) (SEQ ID NO: 162) IVT mRNA substrates following contact with pEY586 (SEQ ID NO: 163). A distribution of small sized RNA fragments ranging between 4-70 nucleotides in length was observed. FIGURE 8C shows the fractional sequence coverage of high confidence oligonucleotide identifications mapping to EPO (SEQ ID NO: 168), FLuc (SEQ ID NO: 169), BGII (4.9 kB) (SEQ ID NO: 170) or BGII (8.9 kB) (SEQ ID NO: 171) IVT mRNA substrates following contact with pEY586 (SEQ ID NO: 163). The small sized RNA fragments produced by pEY586 (SEQ ID NO: 163) yield high overall sequence coverages between 73-99% of each mRNA substrate sequence, suggesting that ribonucleases can be applied to produce populations of mass spectrometry amenable cleavage products from IVT mRNA.
[0241] EXAMPLE 8: Digestion of highly structured RNA with a hyperthermophilic MazF
[0242] RNase from Thermococcus profundus (SEQ ID NO:38 [=pey 1 3], TABLE 1) is disclosed here as an enzyme which surprisingly cleaves RNA at temperatures in a range of 4°C - 85 °C, although the optimum growth temperature of Thermococcus profundus is 80°C. The recognition motif was identified as UGG as described in EXAMPLE 3, based on a cleavage assay performed at 80°C. To test the specificity of peyl43 (SEQ ID NO:38) at multiple temperatures, 33 ng (2.4 pmol) of a peyl43 preparation was incubated with 20 pM of an RNA substrate comprising a single UGG motif (SEQ ID NO: 151) for 15 minutes (40 mM in sodium phosphate buffer pH 7.5) at either 37°C or 85°C. FIGURE 9A illustrates the protocol and FIGURES 9B and 9C illustrate the recognition and cleavage sites on the substrate (SEQ ID NO: 151). Each peyl43 (SEQ ID NO:38) digest was analyzed by UHPLC- MS / MS. FIGURES 9D - 9G show the resultant UHPLC-MS / MS chromatograms without (FIGURES 9D and 9F) and with (FIGURES 9E and 9G) enzyme at 37° C (FIGURES 9D - 9E) or 85°C (FIGURES 9F - 9G). Surprisingly, it was found that when reactions were performed at 37°C, the same enzyme preparation unexpectedly cleaves efficiently at an additional C|GG RNA recognition sequence in comparison to those cleaved at 85°C (FIGURES 9E and 9G). In other words, the cleavage preferences of this enzyme preparation (which included an MBP purification tag) may be changed or controlled by changing the temperature. This property of the recombinant protein makes it a multifunctional sequence specific RNase whose activity can be tuned by temperature based on the sequence composition of target RNA substrate if desired. Furthermore, the activity of this enzyme at >80°C enables digestion and / or denaturation of highly structured RNA substrates, such as tRNAs and ribosomal RNAs.
[0243] Structure of RNA can also be unfolded using denaturing reagents including, for example, urea or formamide (PMID: 23088364, PMID: 4611483). According to some embodiments, an endoribonuclease may retain catalytic activity in the presence of one or more denaturing reagents. An endoribonuclease that retains catalytic in the presence of one or more denaturing reagents, in some embodiments, may be contacted with an RNA substrate in the presence of the one or more denaturing reagents (e.g., to unfold or otherwise reduce the secondary and / or tertiary structure of the RNA substrate) to produce one or more RNA substrate cleavage products.
[0244] The activity of an endoribonuclease having the amino acid sequence of SEQ ID NO:47 was tested in using increasing concentrations of urea or formamide, and it was found that this endoribonuclease surprisingly cleaved RNA in the presence of up to 5 M urea and up to 35% final concentration of formamide (FIGURES 9H and 91). Briefly, 1 pL of 5 pM oligonucleotide (SEQ ID NO: 141) was incubated with 1 pL of 15 pM ribonuclease pey220 (SEQ ID NO:47), 1 pL of 10X Tris reaction buffer (200 mM), varying volumes of either urea (Millipore-Sigma U4883) or formamide (Millipore- Sigma F9037) to reach the desired final concentration for each reaction and nuclease-free water to bring the reaction volume up to 10 pL at 37 °C for 15 minutes. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 10 pL of each sample were then loaded on a pre-cast 20% polyacrylamide TBE-Urea gel that was run for 70 minutes at 180V before imaging on the Amersham Typhoon using the Cy2 setting.
[0245] FIGURE 10A and FIGURE 10B show another embodiment wherein a thermophilic or hyperthermophilic RNase is used to first cleave a structured RNA (e.g., yeast tRNAphe) prior to subsequent digestion with a non-thermophilic RNase such as human RNase 4 (hRNase 4). Briefly, 1 pg of yeast tRNAphewas initially denatured and cut by contacting it with 1 pL (330 ng) of peyl43 (SEQ ID NO:38) at 85°C for 15 minutes prior to incubation with hRNase 4 for 30 minutes at 37°C. The results demonstrate a general principle that pre-treatment of a highly structured RNA with a thermophilic or hyperthermophilic RNase, such as peyl43 (SEQ ID NO:38), improves subsequent cleavage of recognition motifs that may not be accessible by a non-thermophilic RNase, such as hRNase 4 (FIGURE 10B). This approach may be used to increase mapping coverage and / or modification analysis of structured RNAs.
[0246] Ribonucleases such as those capable of cleaving RNA at higher temperatures (>37°C; e.g., thermophilic or hyperthermophilic RNases) or in denaturing reagents (e.g., urea and formamide) may be used to cleave structured RNA since at higher temperatures or in denaturants the RNA structure may be at least partially disrupted. These types of enzymes can be used in biotechnology to sequence or characterize highly structured RNA (e.g., tRNA, rRNA) or RNA with structured regions (e.g., mRNA, IncRNA, sRNA). Alternatively, instead of higher temperatures, an endoribonuclease may be used in combination (e.g., as fusion constructs) with (a) a protein or domain, such as a DNA or RNA helicase or gyrase domains, that can at least partially unfold or unwind double-strand or quadruplex structures, thereby facilitating digestion of RNA substrates and / or (b) a denaturant (e.g., formamide or urea).
[0247] FIGURE IOC shows the results from an additional set of thermophilic cleavage reactions of highly structured RNA substrates to produce oligonucleotides amenable to mass spectrometry -based characterization. In these cleavage reactions, a mixture 5 pg of E. coli 5S rRNA (SEQ ID NO: 173), 16S rRNA (SEQ ID NO: 174) and 23S rRNA (SEQ ID NO: 175) extracted from purified E. coli 70S ribosomes (NEB Cat #P0763S) were contacted with 400 nM of pEY586 (SEQ ID NO: 154) in a total volume of 10 pL in 10 mM Tris pH 8.0, 125 mM NaCl. Each reaction was incubated at 80°C for 15 minutes. Oligonucleotide ethanol precipitation, UHPLC-MS / MS analysis, and high-confidence oligonucleotide identification were performed as described in EXAMPLE 7.
[0248] FIGURE 10C shows the fractional sequence coverage of high-confidence oligonucleotide identifications mapping to E. coli 5S rRNA (SEQ ID NO: 173), 16S rRNA (SEQ ID NO: 174), and 23S rRNA (SEQ ID NO: 175). High overall sequence coverages between 81-100% were observed for all three E. coli rRNA substrate sequences in the mixture, suggesting that thermophilic ribonucleases can be applied to mixtures of highly structured RNA substrates to generate mass spectrometry amenable cleavage products.
[0249] EXAMPLE 9: 5’ Cap analysis with sequence-specific RNases
[0250] Analytical characterization of mRNA 5 ’-end capping by UHPLC-MS / MS may be performed by selective cleavage of 5’ ends utilizing ribonucleases with aid of DNA / RNA probes; deoxyribozymes; or ribozymes. The present disclosure provides additional methods to characterize cap identity and capping status of an RNA of interest. To test whether the sequence-specific endoribonuclease pey214 (SEQ ID NO: 132) could be utilized for analysis of mRNA 5’ end capping, a 5’-engineered EPO mRNA (“Syn90”; SEQ ID NO: 146) was prepared by in vitro synthesis utilizing the CleanCap® Reagent AG. The engineered EPO mRNA (SEQ ID NO: 146) was designed to include a single pey214 cleavage site 61-nt away from the 5’ end. Specific cleavage at the pey214 motif will produce a 5’ end cleavage product that can be directly analyzed by UHPLC-MS / MS. This approach may be applicable to the cap analysis of therapeutic RNAs, both in terms of final product quality control as well as in- process mRNA optimization and development and may be compatible with high-throughput analysis. According to some embodiments, a specific cleavage motif (such as the one of pey214) may be designed and introduced in any synthetic RNA of interest. In other embodiments, a selected RNase may be chosen accordingly to cut a pre-existing cleavage motif in an RNA of interest. FIGURE 11 A and FIGURE 1 IB show a UHPLC-MS / MS analysis of the cleavage products produced upon cleavage of the engineered EPO mRNA (SEQ ID NO: 146) with ribonuclease pey214 (SEQ ID NO: 132). Specifically, 5 pg of EPO mRNA (SEQ ID NO: 146) was incubated with 0.8 pmol ribonuclease pey214 (SEQ ID NO: 132) in 20 mM Tris pH:7.0 in total volume of 10 pL reaction at 37 °C for 1 hour followed by T4 PNK treatment (New England Biolabs, Inc.). It was found that pey214 produces a set of cleavage products of a single and predicted length (61 nt) that are consistent with the position of its cleavage site within the engineered EPO mRNA sequence (FIGURE 11 A). Analysis of mRNA capping with pey214 produced comparable estimates of Capl incorporation in comparison to a validated capping assay utilizing hRNase4 (Wolf et al., 2023) (FIGURE 1 IB). Taken together, these data show that a sequence-specific endoribonuclease, pey214 can be utilized to selectively characterize mRNA 5’ end capping by UHPLC-MS / MS in a probe-independent manner. Such a probe-free sequence specific activity of RNases can be also used to analyze other regions of a given mRNA, such as the mRNA 3’ end. For example, in some embodiments, 3’ ends may be probed for length and / or sequence composition including, for example, the presence of desired (e.g., templated) or undesired (e.g., untemplated) extensions.
[0251] The following example experiment illustrates that a ribonuclease may be used to assess mRNA capping by UHPLC-MS / MS in an unmodified sequence in a probe-independent manner. In these cleavage reactions, 5 pg of FLuc mRNA (SEQ ID NO: 160) either reacted with faustovirus capping enzyme (FCE) or not reacted with FCE was contacted with 400 nM of pEY586 (SEQ ID NO: 163) in a total volume of 10 pL in 10 mM Tris pH 8.0, 125 mMNaCl. Each reaction was incubated at 80°C for 15 minutes. Oligonucleotide ethanol precipitation and UHPLC-MS / MS analysis were performed as described in EXAMPLE 7. Oligonucleotide modification status was determined by comparison of observed and expected monoisotopic masses within a threshold of 10 ppm.
[0252] FIGURE 11C shows a heatmap of the intensity of 5 ’-end oligonucleotides with a Capl modification, diphosphate, and triphosphate in the FLuc mRNAs (SEQ ID NO: 160) reacted with FCE or not reacted with FCE following contact with pEY586 (SEQ ID NO: 163). Primarily Capl modified 5 ’-end oligonucleotides were detected in the presence of FCE. Whereas, a mixture of diphosphate and triphosphate modified 5 ’-end oligonucleotides were detected in the absence FCE, consistent with no 5’-end cap addition. Taken together, these data show that a ribonuclease can be utilized to assess mRNA capping by UHPLC-MS / MS in an unmodified sequence in a probe independent manner.
[0253] EXAMPLE 10: Generating in vitro transcription products with 3’ homogeneous ends
[0254] Sequence-specific endoribonucleases with long recognition motifs (“long base cutters”) may be used to generate homogenous ends of in vitro transcribed or chemically synthesized mRNA by the removal of byproducts (e.g., undesired 3’ extensions). An example method of preparing RNA with 3 ’ homogenous ends is shown in FIGURE 12A in which an RNA substrate having or engineered to have an endoribonuclease cut site at or near (e.g., within 25 nt of) the desired 3’ end is contacted with the endoribonuclease. Undesired cleavage products (e.g., the 3’ end fragments) may be removed by available methods (e.g., by HPLC purification, by filtration using a membrane with the appropriate size cutoff, by size exclusion purification, by precipitation, by anion-exchange purification, by gel electrophoresis, by 5'-3 exonuclease recognizing the undesired cleavage product). Statistically, a 6-base cutter is expected to cleave an RNA having a random sequence on average every 4,096 nucleotides. This infrequency supports controlled cleavage of RNA substrates, for example, substrates that naturally comprise one or more 6-base cutter recognition motifs and substrates engineered to have one or more such motifs. In some embodiments, a synthetic RNA of interest may be engineered to display a unique recognition motif at a position where cleavage is desired. In other embodiments a synthetic RNA of interest may be further engineered to remove or alter one or more combinations of nucleotides that may coincide with the cleavage motif of a selected RNase thereby preventing the cleavage at undesired positions (this approach may enable the use of less frequent cutters such as 4- or 5-base cutters). In some other embodiments, a synthetic RNA of interest is designed so that the RNase cleavage motif precedes a nucleotide sequence that is known to bind to an appropriate target; in this way, once the RNA is cleaved, the resulting fragment containing a nucleotide sequence that is known to bind to an appropriate target may be captured by contacting it with such a target. Examples include one or more repeats of the RNA sequence of the bacterial phage MS2 that bind to the PP7 bacteriophage coat protein (PCP); the Pyrococcus juriosus CRISPR repeat RNA that binds to Cas6 protein; P. aeroginosa CRISPR RNA that binds to Csy4 endoribonuclease; streptomycin-binding RNA aptamers; Tobramycin-binding RNA aptamers; Streptavidin (SA)-binding RNA aptamers (e.g. SI, Sim, and others); and others. This approach enables specific removal of cleaved 3’ end fragment by affinity purification. Other examples, include the use of capture DNA and / or RNA probes that are at least partially complementary to the cleaved 3’ end fragment, thereby enabling their selective removal (e.g., if such probes may be covalently attached to solid surfaces, such as magnetic beads, or may be attachable to solid surfaces though affinity binding such as through biotin-streptavidin interactions; in some cases an additional protein is used to mediate capture by an appropriate probe, such as using the tombusvirus pl9 binding protein). In any embodiment, the resulting 5’ fragment containing the RNA sequence of interest is purified from the mixture that may contain the cleaved 3 ’fragment.
[0255] Methanosphaera stadtmaniae ToxIN protein (“pey214”; SEQ ID NO: 132) was discovered as a sequence-specific endoribonuclease that cleaves the sequence AAAAANAU, in which the caret symbol indicates the primary cleavage site within the seven-nucleotide long recognition motif in RNA. While this enzyme cleaves substrates having C, A, U, or G at position 5 of the recognition motif, under at least some conditions, substrates having a C, A, or U at position 5 of the recognition motif are cleaved better than substrates having a G at this position. For example, substrates having a C, A, or U at position 5 of the recognition motif are cleaved in a sequence-specific manner to homogeneity (FIGURE 12B, FIGURE 12C).
[0256] FIGURE 12B shows capillary electrophoresis and FIGURE 12C shows mass spectrometry of a 5’ FAM-labeled 45-nt RNA oligonucleotide (SEQ ID NO: 142) incubated at 37°C with either no enzyme, T7 RNA polymerase, ribonuclease (SEQ ID NO: 132), or T7 RNAP followed by SEQ ID NO: 132 ribonuclease digestion. The substrate (S), extended oligonucleotides (E), and digested product (D) were labeled.
[0257] To test if pEY214 could also specifically remove a section of the 3’ polyA tail of a full length mRNA (SEQ ID NO: 161; FIGURE 12D), 81 nM of an in vitro transcribed mRNA containing a single pEY214 (SEQ ID NO: 132) target site (AAAACAU) in the context of the polyA tail was incubated with various concentrations of pEY214 (0.02 - 2 pM) for 15 min at 37°C in 50 mM Tris / Cl pH 7.5. The reactions were stopped by adding 0.08 units of proteinase K and an additional incubation for 10 minutes at 24°C. Cleavage fragments were separated on a denaturing 6% TBE-urea gel (FIGURE. 12E). The full-length RNA substrate and the expected cleavage product (68 nt) are indicated on the side of the gel. When the enzyme was added in 25-fold excess (2 pM vs. 0.081 pM), most of the RNA substrate was digested non- specifically. Lower enzyme: substrate ratios showed the successful site-specific cleavage, resulting in a single 3’ cleavage fragment. Although disclosed in the context of trimming off such undesired 3’ ends, the same method may be adapted to remove the 5’ fragment and retain the 3’ end. This alternative may be useful where a 5’ end is capped, labeled, and / or tethered to a surface with the 3’ portion of the RNA being of interest. For example, an RNA that is immobilized by its 5’ end and comprising a recognition motif for a sequence-specific endoribonuclease between the tether point and the portion of the RNA of interest, may be contacted with a sequence-specific endoribonuclease to release fragments comprising the RNA of interest. In some embodiments, methods of trimming off undesired ends (3’ and / or 5’ ends) may be used to generate discrete homogenous preparations of single-strand RNA molecules to be used as authentic standards, for example, in electrophoretic, photometric or mass spectrometric analysis. Complementary single-stranded RNA molecules with homogenous ends may be used to generate homogenous double-stranded RNA molecules. These homogenous single- or double-stranded RNA molecules may find important biotechnological applications as molecular weight markers (or rulers) in electrophoretic gel ladders, as capillary electrophoresis standards, as chromatographic standards, as absorbance standards. Individual homogenous double-stranded RNA molecules of different lengths may be used to build calibration curves for the development of methods for detection of double-stranded RNA impurities that may be formed as by-products of during in vitro RNA synthesis (for example, using T7 RNA Polymerase). Many of the existing methods for detection of double-stranded RNA are based on the use of antibodies or other proteins (e.g. derived from double-stranded RNA cellular receptors) that bind to double-stranded RNA regions in a length-dependent manner. Thus, the ability to generate homogenous double-stranded RNA materials of defined lengths it is critical to obtain reliable measurement of double-stranded RNA levels in any RNA preparations.
[0258] EXAMPLE 11 : RNase cleaved RNA can be re-ligated with RtcB Ligase
[0259] FAM-labeled RNAs were designed to fold into a hairpin structure with a cleavage site located in the single-stranded loop region. These RNA substrates were cleaved into two fragments by a ribonuclease, leaving a 2', 3 '-cyclic phosphate and subsequently ligated back together by RtcB Ligase (FIGURE 13 A). Specifically, 10 pM stocks of these synthetic RNA (SEQ ID NO: 143) oligos from IDT (Integrated DNA Technologies) were heated in nuclease- free water to 90 °C for 5 minutes and then cooled to room temperature for 30 minutes to fold into hairpins. Aliquots (2 pL) of hairpin RNAs (10 pM) were then incubated with 8 pL of 15 pM endoribonuclease pey220 (SEQ ID NO:47), 4 pL of 10X Tris reaction buffer (200 mM) and 26 pL of nuclease-free water (a total reaction volume of 40 pL) at 37 °C for 15 minutes followed by a Monarch RNA cleanup column (NEB, Inc.) and 10 pL elution in nuclease-free water. The cleavage products were then incubated with 1 pL RtcB Ligase (at 15 pM concentration) (New England Biolabs, Inc., M0458) in 2 pL RtcB 10X reaction buffer (500 mM Tris-HCl, 750 mM KC1, 30 mM MgCh, 100 mM DTT, pH 8.3 at 25°C) supplemented with 2 pL of 1 mM GTP, 2 pL of 10 mM MnCh and 3 pL of nuclease-free water at 37 °C for 1 hour followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 15% polyacrylamide TBE- Urea gel that was run for 55 minutes at 180V before imaging on the Amersham Typhoon using the Cy2 setting. The cleaved products ran lower than both the full length starting material and ligated products (FIGURE 13B).
[0260] RNA may be ligated using RtcB ligase without a hairpin-like structure, according to some embodiments. For example, RNA oligonucleotides with a centralized recognition motif for ribonuclease pEY214 (SEQ ID NO: 132) were cleaved and then re-ligated using RtcB Ligase while two synthetic oligonucleotides with matching sequence of the two cleavage products were not ligated due to the absence of a 2', 3 '-cyclic phosphate (FIGURE 13C). Specifically, 1 pL of synthetic RNA (1 pM) was incubated with 1 pL of 10 pM endoribonuclease pey214 (SEQ ID NO: 132), 1 pL of 10X Tris reaction buffer (200 mM, pH7.5) and 7 pL of nuclease-free water (a total reaction volume of 10 pL) at 37 °C for 15 minutes followed by a Monarch RNA cleanup column (New England Biolabs, Inc., and 10 pL elution in nuclease-free water. The RNA was then incubated with 2 pL of RtcB Ligase (at 15 pM concentration) (New England Biolabs, Inc., M0458) in 2 pL of RtcB 10X reaction buffer (500 mM Tris-HCl, 750 mM KC1, 30 mM MgC12, 100 mM DTT, pH 8.3 at 25°C) supplemented with 2 pL of 1 mM GTP, 2 pL of 10 mM MnC12 6 pL of 50% PEG 8000 and 4 pL of nuclease-free water at 37 °C for 2 hours followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The cleaved products and synthetic fragments ran lower than both the full length starting material and ligated products (FIGURES 13C-13D). This experiment demonstrates that cleavage products of pEY214 (SEQ ID NO: 132) ribonuclease amenable to ligation by RtcB ligase. Ribonuclease-cleaved RNAs were also ligated by T4 RNA Ligase 1. Following a T4 PNK treatment, cleaved RNAs were ligated back to their full-length forms. Specifically, 6 pL of synthetic RNA (10 pM) was incubated with 6 pL of 10 pM endoribonuclease pey214 (SEQ ID NO: 132), 5 pL of 10X Tris reaction buffer (200 mM), 5 pL of Murine RNase Inhibitor (NEB, M0314) and 28 pL of nuclease-free water (a total reaction volume of 50 pL) at 37 °C for 15 minutes followed by a Monarch RNA cleanup column (NEB, Inc.) and 17 pL elution in nuclease-free water. 15 pL of cleaved RNA was then incubated with 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgC12, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB, M0314) and 28 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 2 pL of PNK treated RNA was then incubated with 1 pL of T4 RNA Ligase 1 (New England Biolabs, Inc., M0204) in 2 pL of T4 RNA Ligase 1 10X reaction buffer (500 mM Tris-HCl, 100 mM MgC12, 10 mM DTT, pH 7.5 at 25°C) supplemented with 2 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314), 2 pL of DMSO, 6 pL of 50% PEG 8000 and 4 pL of nuclease-free water at 25 °C for 2 hours followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The ligated product runs higher than the cleaved RNA and in line with the initial RNA (FIGURES 13E-13F).
[0261] Unlike RtcB Ligase, T4 RNA Ligase 1 also ligated two synthetic RNA fragments that have not been cleaved by a ribonuclease since T4 PNK treatment modifies synthetic RNA fragments in a way that they are amenable to ligation by T4 RNA ligase. 1 (FIGURE 13E).
[0262] Ribonuclease-cleaved RNAs were also ligated by T4 RNA Ligase 2 using a DNA oligonucleotide as a splint. Following a T4 PNK treatment, cleaved RNAs were ligated back to their full-length forms in both a splint-independent and splint-dependent manner. Specifically, 6 pL of synthetic RNA (10 pM) was incubated with 6 pL of 10 pM endoribonuclease pey214 (SEQ ID NO: 132), 5 pL of 10X Tris reaction buffer (200 mM, pH 7.5), 5 pL of Murine RNase Inhibitor (NEB M0314) and 28 pL of nuclease-free water (a total reaction volume of 50 pL) at 37 °C for 15 minutes followed by a Monarch RNA cleanup column (NEB, Inc.) and 17 pL elution in nuclease-free water. 15 pL of cleaved RNA was then incubated with 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgC12, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314) and 28 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. For the splint-independent reaction, 2 pL of PNK treated RNA was then incubated with 1 pL of T4 RNA Ligase 2 (New England Biolabs, Inc., M0239) in 1.5 pL of T4 RNA Ligase 2 10X reaction buffer (500 mM Tris-HCl, 20 mM MgC12, 20 mM DTT, 4 mM ATP, pH 7.5 at 25°C) supplemented with 3 pL of 200 mM Tris pH 7.5, 1.5 pL of 500 mM NaCl and 6 pL of nuclease- free water at 37 °C for 2 hours followed by a Monarch RNA cleanup. For the splint-dependent reaction, 6 pL of PNK treated RNA was mixed with 6 pL of 1 pM DNA splint, 1 pL of 200 mM Tris pH 7.5 and 2 pL of 500 mM NaCl and incubated at 65 °C for 3 minutes followed by cooling to room temperature for 15 minutes. The annealed RNA and splint were then incubated with 1 pL T4 RNA Ligase 2 (New England Biolabs, Inc., M0239) in 2 pL of T4 RNA Ligase 2 10X reaction buffer (500 mM Tris-HCl, 20 mM MgC12, 20 mM DTT, 4 mM ATP, pH 7.5 at 25°C) at 37 °C for 2 hours followed by a Monarch RNA cleanup. 10 pL of ligated RNA was then incubated with 1 pL of DNase I (New England Biolabs, Inc., M0303) in 5 pL of DNase I 10X reaction buffer (100 mM Tris-HCl, 25 mM MgC12, 5 mM CaC12, pH 7.6 at 25°C) and 35 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The ligated product runs higher than the cleaved RNA, in line with the initial RNA and remains after DNase I treatment (FIGURES 13E-13F). Unlike RtcB Ligase, T4 RNA Ligase 2, as expected, also ligated two synthetic RNA fragments that have not been cleaved by a ribonuclease but have been T4 PNK treated (FIGURES 13G-13H). This data demonstrates that RNA fragments generated by pey214 (SEQ ID NO: 132) enzyme can be converted by PNK / ATP treatment to molecules amenable to ligation by RNA Ligase 2.
[0263] T4 RNA Ligase 2 can also ligated three RNA fragments together in a desired order using a DNA oligonucleotide splint (FIGURES 131- 13 J). A central RNA fragment was designed to contain an inosine nucleotide that can be cleaved by Endo V to assess the identity of the ligation product. Specifically, 6 pL of synthetic central RNA (10 pM) was incubated with 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgC12, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314) and 32 pL of nuclease- free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 6 pL of each synthetic non-central RNA (10 pM) was incubated with 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgC12, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314) and 26 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 4 pL of both PNK treated RNAs were mixed with 2 pL of 1 pM DNA splint, 5 pL of 200 mM Tris pH 7.5 and 2 pL of 500 mM NaCl and incubated at 65 °C for 3 minutes followed by cooling to room temperature for 15 minutes. The annealed RNA and splint were then incubated with 3 pL T4 RNA Ligase 2 (New England Biolabs, Inc., M0239) in 2 pL of T4 RNA Ligase 2 10X reaction buffer (500 mM Tris-HCl, 20 mM MgC12, 20 mM DTT, 4 mM ATP, pH 7.5 at 25°C) at 37 °C for 2 hours followed by a Monarch RNA cleanup. 6 pL of ligated RNA was then incubated with 1 pL of DNase I (New England Biolabs, Inc., M0303) in 5 pL of DNase 1 10X reaction buffer (100 mM Tris-HCl, 25 mM MgC12, 5 mM CaC12, pH 7.6 at 25°C) and 35 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 6 pL of DNase I treated RNA was then incubated with 1 pL of Endo V (New England Biolabs, Inc., M0305) in 1 pL of NEB buffer 4 (500 mM Potassium Acetate, 200 mM Trisacetate, 100 mM Magnesium Acetate, 10 mM DTT, pH 7.9 at 25°C) and 2 pL of nuclease- free water at 37 °C for 1 hour followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1: 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The ligated product runs higher than the central and noncentral synthetic RNA, in line with the full-length splint and remains after DNase I treatment but not after Endo V treatment (FIGURES 131- 13 J).
[0264] Ribonuclease-cleaved RNAs were also ligated to a third central RNA fragment by T4 RNA Ligase 2 using a DNA oligonucleotide as a splint (FIGURE 13K, 13L). Following a T4 PNK treatment, cleaved RNAs were ligated to a central RNA fragment that was designed to contain an inosine nucleotide that can be cleaved by Endo V to assess the identity of the ligation product. Specifically, 6 pL of synthetic RNA (10 pM) was incubated with 6 pL of 10 pM endoribonuclease pey214 (SEQ ID NO: 132), 5 pL of 10X Tris reaction buffer (200 mM), 5 pL of Murine RNase Inhibitor (NEB M0314) and 28 pL of nuclease-free water (a total reaction volume of 50 pL) at 37 °C for 15 minutes followed by a Monarch RNA cleanup column (NEB, Inc.) and 17 pL elution in nuclease-free water. 15 pL of cleaved RNA was then incubated with
[0265] 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgCh, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314) and 28 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 6 pL of synthetic central RNA (10 pM) was incubated with 1 pL of T4 PNK (New England Biolabs, Inc., M0201) in 5 pL of T4 PNK 10X reaction buffer (700 mM Tris-HCl, 100 mM MgCh, 50 mM DTT, pH 7.6 at 25°C) supplemented with 5 pL of 10 mM ATP, 1 pL of Murine RNase Inhibitor (NEB M0314) and 32 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup.
[0266] 4 pL of both PNK treated RNAs were mixed with 2 pL of 1 pM DNA splint, 5 pL of 200 mM Tris pH 7.5 and 2 pL of 500 mM NaCl and incubated at 65 °C for 3 minutes followed by cooling to room temperature for 15 minutes. The annealed RNA and splint were then incubated with 3 pL T4 RNA Ligase 2 (New England Biolabs, Inc., M0239) in 2 pL of T4 RNA Ligase
[0267] 2 10X reaction buffer (500 mM Tris-HCl, 20 mM MgCh, 20 mM DTT, 4 mM ATP, pH 7.5 at 25°C) at 37 °C for 2 hours followed by a Monarch RNA cleanup. 6 pL ob ligated RNA was then incubated with 1 pL of DNase I (New England Biolabs, Inc., M0303) in 5 pL of DNase I 10X reaction buffer (100 mM Tris-HCl, 25 mM MgCh, 5 mM CaCh, pH 7.6 at 25°C) and 35 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 6 pL of DNase I treated RNA was then incubated with 1 pL of Endo V (New England Biolabs, Inc., M0305) in 1 pL of NEB buffer 4 (500 mM Potassium Acetate, 200 mM Tris-acetate, 100 mM Magnesium Acetate, 10 mM DTT, pH 7.9 at 25°C) and 2 pL of nuclease-free water at 37 °C for 1 hour followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for
[0268] 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The ligated product runs higher than the original substrate RNA, in line with the full-length splint and remains after DNase I treatment but not after Endo V treatment (FIGURE 13L).
[0269] Ligation strategies described above, for example two-fragment ligation or three- fragment-ligation, or multifragment ligation with and without splint DNA oligos splint RNA oligos can be used to generate recombinant RNA. Each ligated fragment may contain canonical nucleotides, chemically modified nucleotides or both. Chemical modification can be both naturally occurring such as epitranscriptomic modifications or synthetic non-natural modifications.
[0270] In two fragment ligation someone may use two different IVT products; one made with canonical nucleotides, second made with modified nucleotides. These fragments can be ligated to obtain partially modified recombinant RNA. In another application such as in 3 -fragment (3 -way ligation) or multiple fragment ligation, a point modified RNA or partially modified RNA can be generated. In another application, multiple site-specific modifications can be incorporated in as single molecule. RNA fragments containing desired modification or modifications can be synthesized chemically or enzymatically using RNA polymerases or installed by site-specific RNA modifying enzymes such pusl (pseudouridine synthase 1). Non -natural modifications on RNA fragments may include a fluorophore or any other type of chemical modification. Recombinant RNA can be circularized to obtain circular RNA.
[0271] EXAMPLE 12: Generation of circular RNA from RNase cleaved synthetic or in vitro transcribed RNA
[0272] To generate a circular RNA, an RNA substrate was designed to contain the cleavage site of a ribonuclease at both the 5’ and 3’ ends so that when RNA is cleaved by RNase, homogenous 5' and 3' ends are obtained. After ribonuclease cleavage, the two ends are then ligated together by RtcB ligase resulting in a circular RNA (FIGURE 14 A). To keep the 5’ and 3 ’ ends of the cleaved RNA in close proximity to each other for ligation, the RNA was designed to fold into a hairpin-like structure as predicted by RNAfold webserver (rna.tbi. univie. ac.at / cgi- bin / RNAWebSuite / RNAfold.cgi). The two cleavage sites are indicated by black lines and arrows on the predicted structure (FIGURE 14B, left). After RNase cleavage, the 5’ and 3’ ends of the RNA molecule are predicted to be in proximity, based on the RNAfold structure (FIGURE 14B, right). These ends can thus be ligated together at the ligation site shown as the incomplete circle at the top of the predicted structure (FIGURE 14B, right) after RNase treatment.
[0273] In vitro transcribed RNAs were generated using 1 pg of BspQI linearized plasmid per reaction at 37 °C for 6 hours with the HiScribe® T7 High Yield RNA Synthesis Kit followed by addition of 2 pL of DNase I and a further 30-minute incubation at 37 °C. After a Monarch RNA cleanup column, a 50 ng / pL stock was created for each RNA sample in nuclease-free water. 6 pL of this stock was then heated to 90 °C for 5 minutes and then cooled to room temperature for 30 minutes to fold before being combined with 12 pL of 10 pM ribonuclease pey214 (SEQ ID NO: 132), 6 pL of 10X Tris reaction buffer (200 mM) and 36 pL of nuclease- free water to a total reaction volume of 60 pL. The reactions were then incubated at 37°C for 15 minutes followed by a Monarch RNA cleanup column and the RNA sample was refolded by heating to 90 °C for 5 minutes and then cooling to room temperature for 30 minutes. 20 pL of the cleaved RNA was then incubated with 2 pL RtcB Ligase in 4 pL RtcB 10X reaction buffer supplemented with 4 pL of ImM GTP, 4 pL of 10 mM MnCh and 6 pL of nuclease-free water to a total reaction volume of 40 pL. The reactions were then incubated at 37 °C for 2 hours followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye was then added to each sample which were then incubated at 70 °C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 70 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. Ligation products were confirmed when cleaved products ran lower than the full length starting material and circularized products ran higher than the full length (FIGURE 14C). Cleavage of IVT RNA by RNase followed by ligation enables generation of a homogenous circular RNA population since RNase cleavage removes any undesired extension introduced by RNA polymerase during IVT reaction.
[0274] After ligation, to further confirm the presence of circular RNA, the RNA was incubated with RNase R to remove linear fragments. 7 pL of cleaned-up RNA sample were incubated with 1 pL of 1 :20 diluted RNase R (LGC Biosearch Technologies RNR07250), 1 pL of 10 X reaction buffer and 1 pL of nuclease-free water at 37 °C for 15 minutes followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye was then added to each sample which were then incubated at 70 °C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 70 minutes at 180V. The gel was stained using SYBR Gold stain (1: 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. Ligation products were confirmed when cleaved products decreased in intensity after RNase R treatment and circularized products remained after RNase R treatment (FIGURE 14D).
[0275] EXAMPLE 13: Linearization of circular mRNA with known or unknown sequences
[0276] Current direct RNA sequencing technologies (e.g., Oxford Nanopore direct RNA sequencing) require linear RNA molecules with open ends where adapters can be hybridized and / or ligated. Circular RNA molecules cannot be directly sequenced with these technologies unless the circular molecules are first linearized. The quality of sequencing results and even the ability to generate sequence information for circular RNAs can be compromised by uncontrolled or poorly controlled linearization reactions. In some embodiments, circular RNA (e.g., alone or as a mixed population with RNA in other conformations) may be contacted with one or more long-base cutter RNases selected to cleave the circular RNA (on average) once per molecule to produce linear RNA products, which may then be sequenced. As shown in FIGURE 15, total RNA may be subjected to ribosomal RNA depletion, poly(A) RNA depletion, or both ribosomal RNA depletion and poly(A) RNA depletion to produce an RNA pool depleted of linear RNA. The resulting RNA pool (enriched for circular RNA) optionally may be subject to any further steps desired to reduce the content of linear RNA including, for example, contacting the RNA pool with RNase R and / or a 5 ’-3’ exonuclease to form an RNA pool enriched in circular RNA. The circular RNA enriched pool may be contacted with one or more long-base cutter RNases to form linearization products (e.g., a pool of linearized (formerly) circular RNA molecules). Alternatively, multiple RNases can be used in parallel to increase the proportion of linearized circular RNA. Then, linearized RNA molecules can be used in standard ONT direct RNA sequencing or other sequencing platforms.
[0277] EXAMPLE 14: Modular synthesis and amplification of RNA molecules from single transcript
[0278] In the context of the present disclosure, “IVT template” refers to an RNA molecule that may be transcribed in vitro using any available in vitro transcription system including, for example, a cell lysate or a HiScribe® T7 high yield RNA synthesis kit (NEB, Inc.) An IVT template may be a single-stranded RNA comprising, in a 5’ to 3’ direction, a 5’ untranslated region (“UTR”), a coding sequence (e.g., defined by start and stop codons), a 3’ UTR, and / or a poly (A) tail.
[0279] In the context of the present disclosure, “polycistronic template” refers to a DNA (or RNA) molecule comprising, in a 5’ to 3’ direction, P(CS-ERS)n(CS)m(T)x3’, wherein P is a promoter, CS encodes a coding sequence, ERS encodes an endoribonuclease recognition sequence, «=>1, >2, >3, >4, >5, >10, >15, >20, >25, or >50, m=Q or 1, T encodes a poly(A) tail, x=0 or 1, and 3’ represents the 3’ end of the polycistronic template molecule. For clarity, if n=l, then mfO. Transcription of a polycistronic template may produce a polycistronic (single-stranded) RNA comprising, in a 5’ to 3’ direction, (cs-ers)n(cs)m(A)x3 ’ , wherein cs encodes a coding sequence, ers encodes an endoribonuclease recognition sequence, A is a poly(A) tail, and 3’ represents the 3’ end of the polycistronic RNA molecule (with m, n, and x as defined above). In some embodiments, coding sequences in a polycistronic RNA may be the same or different. In some embodiments, endoribonuclease recognition sequences may be the same or different.
[0280] Polycistronic RNAs may be produced efficiently from polycistronic templates, for example, by performing in vitro transcription with a viral RNA polymerase (e.g., T7 RNA polymerase, SP6 RNA polymerase, T3 RNA polymerase). In some embodiments, a method may include transcribing a polycistronic template to produce a polycistronic RNA and contacting the polycistronic RNA with an endoribonuclease (e.g., compatible with endoribonuclease recognition sequences in the polycistronic RNA) to produce endoribonuclease cleavage products comprising one or more (copies of) RNA molecules of interest (e.g., each comprising an IVT template). An example of such a method is illustrated in FIGURE 16 A.
[0281] In some embodiments, a polycistronic template may be configured to encode >1, >2, >3, >4, >5, >10, >15, >20, >25, or >50 different RNA species and a sequence specific endoribonuclease recognition site (optionally, wherein each recognition site may be the same or different) positioned between each of the different RNA species. In some embodiments, a method may comprise contacting a polycistronic template (the polycistronic template having sequences encoding two or more different coding sequences) with an RNA polymerase (e.g., a viral RNA polymerase or variant thereof) to form transcription products comprising polycistronic RNA (the polycistronic template having two or more different coding sequences) and contacting the transcription products with one or more endoribonucleases (the one or more endoribonucleases compatible with endoribonuclease recognition sequences in the polycistronic RNA) to produce endoribonuclease cleavage products comprising separate RNA molecules, each having one of the different species of coding sequences. An example of such a method is illustrated in FIGURE 16B.
[0282] 363 frnol (96.3 nM) of a 610 nucleotide long exemplary RNA sequence (SEQ ID: 162) that contained the sequence for pEY214 (AAA|ACAU, | indicating the cleavage site; FIGURE 16C) twice, were incubated for 15 min at 37°C with various concentrations of pEY214 (0.02 pM to 2 pM as indicated above the gel) in 50 mM Tris / Cl pH 7.5. The reactions were stopped by adding 0.08 units of proteinase K and an additional incubation for 10 minutes at 24°C. Cleavage fragments were separated on a denaturing 6% TBE-urea gel (FIGURES 16C-16D). pEY214 preferentially and sequence specifically cleaves at its recognition sites in the mRNA substrate, generating RNA fragments of the desired lengths. The lengths of the full-length RNA substrate (610 nt) and the expected cleavage products (87 / 200 / 323 nt) are indicated on the right side of the gel.
[0283] EXAMPLE 15: Generation of dsRNA using RNase cleaved sense and antisense RNA
[0284] In some embodiments, endoribonucleases may be used to produce high quality duplex RNA molecules. As illustrated in FIGURE 17 a method may include, for example, contacting a first template (DNA or RNA) comprising, in a 5’ to 3’ direction, a sequence encoding a sense strand coding sequence and a first endoribonuclease recognition site and an RNA polymerase to produce a first IVT product; contacting the first IVT product with the endoribonuclease corresponding to the endoribonuclease recognition site encoded by the first template to produce sense strand cleavage products (e.g., having homogenous 3’ ends); contacting a second template (DNA or RNA) comprising, in a 5’ to 3’ direction, a sequence encoding an anti-sense strand coding sequence (complementary to the sense strand coding sequence) and second endoribonuclease recognition site and an RNA polymerase to produce a second IVT product, contacting the second IVT product with the endoribonuclease corresponding to the endoribonuclease recognition site encoded by the second template (wherein the endoribonucleases sites may be the same or different) to produce antisense strand cleavage products (e.g., having homogenous 3’ ends); and contacting the sense strand cleavage products and anti-sense strand cleavage products to produce a high quality duplex RNA comprising the sense strand and the anti-sense strand. In some embodiments, the sense strand cleavage products and / or the anti-sense strand cleavage products may be purified to remove the respective 3’ fragments and / or enrich for the sense and anti-sense strands respectively. RNase cleavage may ensure that all the dsRNA molecules have uniform ends resulting in homogenous pool of dsRNA which could be used for different purposes including as dsRNA marker.
[0285] EXAMPLE 16: Generation of point-modified or site-specifically modified or recombinant RNA from RNase-cleaved fragments
[0286] Structural and functional analyses of noncoding RNA or coding mRNA may require preparation of highly pure and homogenous populations of RNA molecules. In some embodiments, it may be desirable to produce a recombinant RNA that mimics a native noncoding RNA or coding mRNA of interest with one or more site-specific alterations (examples of alterations include but are not limited to: a native base swap; the removal of a native RNA modification; the introduction of an epitranscriptomic modification; the introduction of a non-native chemical modification; the introduction of a spectroscopic probe, such as fluorophore label; and / or the introduction of an affinity probe, such as a biotin label) relative to the parent molecule. In the context of this disclosure, such recombinant RNA may be generated from the native (cellular) RNA species by selectively cutting out a fragment of choice from the native species by contacting the native RNA with one or more ribonucleases and then replacing that cut-out fragment with one harboring the desired modifications. In some embodiments, it may be desirable to produce a site-specifically modified RNA from an in vitro transcribed RNA (IVT RNA). Similarly, a site-specific modification may be introduced by cutting out a given sequence segment from an IVT RNA using one or more ribonucleases and then replacing it by ligating a fragment of choice (e.g., this fragment may be synthetic or derived from a native RNA) to the IVT RNA.
[0287] Producing highly pure and homogenous recombinant RNA (e.g., long RNAs) comprising site-specific chemical modifications (e.g., an epitranscriptomic modification or a fluorophore label) using prior tools can be challenging. In comparison, endoribonucleases disclosed herein (e.g., sequence-specific ribonucleases) may be used to selectively and precisely cut an RNA at predicted locations, thereby permitting that chemical modifications may be introduced to either native or synthetic RNA as described in this invention.
[0288] In some embodiments, endoribonucleases may be used to produce high quality recombinant RNA. As illustrated in FIGURE 18 A a method may include, for example, cleaving an in vitro produced or native RNA containing an RNase cleavage motif, wherein this motif may be either naturally occurring in the sequence or artificially installed. Between the cleavage sites a chemically synthesized RNA oligonucleotide 10 to 100 nt long, or 100 to 200 nt long, or 200 to 1000 nt long, containing a chemical modification may be introduced by ligation, for instance using an RNA ligase or using a splint ligation approach.
[0289] The present disclosure provides, according to some embodiments, methods of RNA assembly and / or RNA-DNA chimera assembly. For example, a method may comprise assembling >2, >3, >4, >5 RNA molecules (e.g., single-stranded RNA molecules; e.g., cleavage products of an endoribonuclease, synthetic RNA, native RNA, RNA produced by IVT). In some embodiments, a method may comprise contacting the >2, >3, >4, >5 RNA molecules of interest and an RNA ligase operable to join the fragments. Fragments may be assembled in a desired order. For example, fragments may be contacted and ligated in series (e.g., ligate the first two to each other, ligate a third, ligate a fourth and so on) to achieve a desired order of assembly. Fragments may be contacted with a splint DNA having domains complementary to at least some portions of at least some of the RNAs to be assembled. It may be generally desirable in this method and other disclosed methods for the splint DNA to be fully complementary to all of the RNAs to be assembled, but circumstances may arise where providing an alternate splint DNA is desired. For example, a splint DNA for a three-fragment assembly may be designed to comprise sequences fully complementary to desired 5’ and 3’ domains while allowing the central domain of the assembled RNA to be populated by a variable region.
[0290] In some embodiments, methods may include cleaving an RNA of interest. For example, a method may include contacting an RNA of interest and a programmable endoribonuclease (e.g., Mpa-AGO among other Argonautes; US Application 18 / 969,910 filed December 5, 2024 and incorporated herein by reference) to produce cleavage fragments. Programmable endoribonuclease cleavage fragments, according to some embodiments of this disclosure, may be assembled (e.g., reassembled or assembled in other desired combinations). For example, a method may comprise contacting an RNA of interest and a programmable endoribonuclease to produce cleavage fragments and contacting the programmable endoribonuclease cleavage fragments with an example RNA ligase. FIGURE 18B illustrates an example embodiment of a method in which a 600-nt region of the human PSMB2 mRNA (SEQ ID NO: 198) was in vitro transcribed using HiScribe T7 High Yield RNA Synthesis Kit (New England Biolabs) at 37°C for 2 hours. After column purification using the Monarch RNA Cleanup Kit (New England Biolabs), the IVT RNA was then cleaved into two fragments using a DNA-guided RNA-targeting Ago from Mucilaginibacter paludi (e.g., WP_008504757.1). The DNA guide (SEQ ID NO:200) used in the cleavage reaction was designed to produce two RNA cleavage fragments of sizes 359 nt and 241 nt, and the DNA guide was phosphorylated at the 5' end using T4 polynucleotide kinase (New England Biolabs), which improves cleavage efficiency. Ago cleavage reaction was performed at 50°C for 30 minutes, in NEB ThermolPol Reaction Buffer (New England Biolabs). 2 pg of RNA substrate was used per 100 pL reaction. The RNA substrate to Ago molar ratio was 1 :10, and the guide to Ago molar ratio was 1 : 1. After column purification, 700 ng of the cleavage products (3.64 pmol of the 5' fragment and 3.64 pmol of the 3' fragment) were annealed to a 24-nt long DNA oligo splint (SEQ ID NO: 194) (3.64 pmol), with or without two 60-nt long DNA oligo disruptors (SEQ ID NOS: 195-196) (87.36 pmol each), in 62.5 mM Tris, pH 7.5 and 50 mM NaCl, at 60°C for 5 minutes (12 pL reaction volume). The mixture was then cooled down 0.1 °C per second to 5°C. The first 12 nucleotides of the splint anneal to the first 12 nucleotides of the 3' cleavage fragment, and the last 12 nucleotides of the splint anneal to the last 12 nucleotides of the 5' cleavage fragment. One of the two disruptors anneals to the 5' cleavage fragment, one nucleotide away from the splint, while the other anneals to the 3' fragment, also one nucleotide away from the splint. These disruptors were expected to eliminate RNA secondary structures near the ligation site of the cleavage fragments. Following annealing, MgCh, DTT, ATP (final concentrations 2 mM, 2 mM, and 400 pM, respectively), and 4.95 pmol of T4 RNA Ligase 2 (New England Biolabs) were added to the reaction (15 pL final reaction volume), and the ligation reaction was performed at 37°C for 90 minutes. The DNA splint and disruptors were eliminated from the reaction by DNase I treatment (New England Biolabs) at 37°C for 15 minutes. Samples mixed with 2X RNA loading dye (New England Biolabs) and were analyzed by a 6% TBE-urea gel and stained with SYBR Gold (Thermo Fisher Scientific). As shown in FIGURE 18B, efficient splint-dependent ligation was achieved by including DNA disruptors in the reaction, and the DNA splint and disruptors can be removed from the reaction with DNase I. This ligation method may be applied to join one nuclease cleavage fragment to another cleavage fragment from a different cleavage reaction, or to a synthetic or natural RNA containing one or more chemical modifications.
[0291] FIGLTRE 18C illustrates the results of a three-RNA fragment assembly method. A 1- kb region of the human PSMB2 mRNA (SEQ ID NO: 199) was synthesized by IVT, and the resultant IVT RNA was cleaved using AGO, under the same reaction conditions as described for FIGURE 18B, to produce two cleavage fragments (590 nt and 410 nt). 1 pg of the cleavage fragments (3.12 pmol of the 5' fragment and 3.12 pmol of the 3' fragment) were then mixed with a synthetic 31-nt long middle RNA (SEQ ID NO: 193) (8 pmol) containing an inosine at position 16 (IDT) and were annealed to a 55-mer DNA oligo splint (SEQ ID NO: 197) (3.12 pmol), with or without two 60-nt long DNA oligo disruptors (SEQ ID NOS: 195-196) (75 pmol each), under the same reaction conditions as described for FIGURE 18B. The 55-ntDNA splint covers the entire middle RNA and extends 12 nucleotides into both the 5' and 3' cleavage fragments. Ligation reaction was performed with or without 4.95 pmol T4 RNA Ligase 2, under the same conditions as described for FIGURE 18B. To determine if the middle RNA was successfully ligated, ligation products were treated with endonuclease V (New England Biolabs), which cleaves 3' to an inosine in RNA / DNA. DNA splint and disruptors were removed by DNase I treatment. Samples were mixed with 2X RNA loading dye (New England Biolabs) and analyzed by a 6% TBE-urea gel and stained with SYBR Gold (Thermo Fisher Scientific). As demonstrated in FIGURE18C, ligation products of the expected size were detected in the samples with T4 RNA Ligase 2, and inclusion of the DNA disruptors slightly improved efficiency. Following endonuclease V treatment, these products were no longer detected, suggesting that they were digested by the inosine-targeting endonuclease V. This confirmed that they were the successful ligation products containing the inosine middle RNA.
[0292] FIGURE 18D illustrates the results of a three-RNA fragment assembly method in which the middle fragment was a 31-nt long synthetic RNA (SEQ ID NO: 192) comprising an internal FAM fluorescence label (on a dT) at position 17 (IDT). In this experiment, ligation was performed under the same reaction conditions as described for FIGURE 18C, except that the internally FAM-labelled middle RNA (SEQ ID NO: 192) (6.24 pmol) was used instead of the inosine-containing middle RNA (SEQ ID NO: 193). Ligation products were analyzed by 6% TBE-urea gel, and the gel was scanned directly for FAM fluorescence signals. As shown in FIGURE18D, successful ligation products containing the FAM-labeled middle RNA were detected in the samples with T4 RNA Ligase 2, and, consistent with all ligation experiments described in this EXAMPLE, inclusion of the DNA disruptors (SEQ ID NOS: 195-196) improved production of the desired ligation products.
[0293] EXAMPLE 17: Blocking RNase cleavage using RNA modification specific antibodies or modification binding natural or engineered proteins.
[0294] In some embodiments, a method may comprise contacting a substrate RNA having an endoribonuclease recognition motif and a binding protein (e.g., an antibody) operable to bind the endoribonuclease recognition motif. For example, if a substrate RNA of interest comprises an endoribonuclease recognition motif comprising a modified nucleotide, the binding protein (e.g., antibody) may be operable to bind the endoribonuclease recognition motif at the modified nucleotide (FIGURE 19).
[0295] As set forth in the instant specification, nucleotide- and / or sequence-specific RNases may exhibit partial sensitivity to RNA modifications. For example, in FIGURE 7, activity of RNases (SEQ ID NOS: 7 & 8) are impeded by the presence of methyl group modification in cytosine in the recognition motif "AC A" (modified version A{5mC}A). Cleavage activity of such RNases (e.g., SEQ ID NOS: 7 & 8) may be further impeded or completely blocked by contacting an RNA substrate with a modification-specific antibody (e.g., a commercially available 5mC-binding antibody, a m6A-binding antibody, and a pseudouridine-binding antibody). For example, modification-specific antibodies may bind to a recognition site comprising a modification, thereby limiting the ability of the endoribonuclease to gain access to its recognition site. In some embodiments, modification binding proteins (e.g., naturally occurring or engineered modification binding proteins) may also limit an endoribonuclease’s access to its recognition site (when such recognition site comprises a modified nucleotide). For example, naturally occurring m6A binding proteins including, for example, YTH domaincontaining proteins, IGF2BPs, and eIF3 (PMID:29103884, PMID:29476152, PMID:26593424 respectively) or engineered modification binding proteins may be used (e.g., instead of antibodies) to block RNase activity. By comparing cleavage pattern of antibody or protein treated and untreated RNA samples, modification sites can be inferred at single nucleotide resolution. Antibodies and modification-binding proteins can also be used for the enrichment of RNA molecules containing modifications followed by RNase treatment (FIGURE 19).
[0296] EXAMPLE 18: Inhibition of ribonuclease cleavage by annealing DNA probes to substrate RNA,
[0297] The present disclosure relates to methods of selectively cleaving an RNA molecule according to some embodiments. For example, a method may include blocking, binding, or otherwise shielding one or more endoribonuclease recognition motifs to limit or prevent cleavage by a corresponding sequence-specific endoribonuclease. A method may comprise, for example, contacting a substrate RNA having an endoribonuclease recognition motif and a DNA oligonucleotide complementary to the endoribonuclease recognition motif. If a substrate RNA of interest comprises more than one copy of the endoribonuclease recognition motif but only one copy or a subset of copies are to be blocked, the DNA oligonucleotide may be configured to have additional sequences (beyond the sequence complementary to the recognition motif) complementary to the sequence(s) proximal to the specific copies to be blocked. Without limiting any specific embodiment to any particular mechanism of action, hybridizing a recognition motif with a complementary DNA oligonucleotide may limit cleavage by sequence-specific endoribonucleases that cleave single-stranded substrates but do not cleave (or at least cleave less-efficiently) duplex substrates (e.g., DNA / RNA duplexes, RNA / RNA duplexes). Methods may include removing the DNA oligonucleotide, for example, by contacting the heteroduplex with a DNase (e.g., DNase I). In some embodiments, blocking endoribonuclease cleavage may be used to prevent off-target cleavage at degenerate sites in long RNA sequences. If an RNA substrate contains a degenerate site for cleavage by a ribonuclease, a DNA oligonucleotide can be annealed to the RNA at that off-target site to inhibit cleavage. While blocking oligonucleotides comprising DNA have several advantages as set forth herein, in some embodiments, it may be desirable to use a blocking oligonucleotide comprising or consisting of RNA. To further illustrate, 1 pL of synthetic RNA, R618, (1 pM) was incubated with 1 pL of 1 pM DNA oligo, 1 pL of 200 mM Tris pH 7.5, 1 pL of 500 mM NaCl and 4 pL of nuclease- free water and incubated at 65 °C for 3 minutes followed by cooling to room temperature for 15 minutes. The annealed RNA and DNA were then incubated with 1 pL of 10 pM ribonuclease pey214 (SEQ ID NO: 132) and 1 pL of Murine RNase Inhibitor (NEB M0314) at 37 °C for 15 minutes followed by a Monarch RNA cleanup. 10 pL of sample was then incubated with 1 pL of DNase I (New England Biolabs, Inc., M0303) in 5 pL of DNase I 10X reaction buffer (100 mM Tris-HCl, 25 mM MgC12, 5 mM CaC12, pH 7.6 at 25°C) and 34 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1: 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The cleaved product runs lower than the uncleaved RNA and ribonuclease cleavage was prevented by all three DNA oligonucleotides (FIGURE 20B). DNase I treatment removed the DNA oligonucleotides.
[0298] For a longer RNA substrate, IVT products were generated for a 610 nt RNA (SEQ ID NO: 162) substrate that contained two cleavage sites. 1 pL of 200 ng / pL IVT product was incubated with 1 pL of either 5, 10 or 15 pM DNA oligo, 1 pL of 200 mM Tris pH 7.5, 1 pL of 100 mM NaCl and 4 pL of nuclease-free water and incubated at 95 °C for 1 minute followed by cooling to room temperature for 15 minutes. The annealed RNA and DNA were then incubated with 1 pL of 10 pM ribonuclease pey214 (SEQ ID NO: 132) and 1 pL of Murine RNase Inhibitor (NEB M0314) at 37 °C for 15 minutes followed by a Monarch RNA cleanup. 10 pL of sample was then incubated with 1 pL of DNase I (New England Biolabs, Inc., M0303) in 5 pL of DNase I 10X reaction buffer (100 mM Tris-HCl, 25 mM MgC12, 5 mM CaC12, pH 7.6 at 25°C) and 34 pL of nuclease-free water at 37 °C for 30 minutes followed by a Monarch RNA cleanup. 10 pL of 2X RNA loading dye (New England Biolabs, Inc., B0363S) was then added to each sample which were then incubated at 70°C for 5 minutes. 18 pL of each sample were then loaded on a pre-cast 6% polyacrylamide TBE-Urea gel that was run for 40 minutes at 180V. The gel was stained using SYBR Gold stain (1 : 10,000) for 3 minutes before rinsing three times in distilled water and imaging on the Alphaimager HP. The three bands ran at 323 nt, 200 nt and 87 nt for the cleaved products. As the DNA to RNA ratio increased, ribonuclease cleavage was prevented and the intensity of the 200 nt and the 87 nt bands decreased (FIGURE 2 IB). DNase I treatment removed the DNA oligonucleotides. This data also demonstrates that (like in FIGURE 20B) certain cleavage of sites can be blocked by DNA oligos.
Claims
CLAIMSWhat is claimed is:
1. A method for forming RNA cleavage products comprising a single-stranded therapeutic RNA, the method comprising: contacting: a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164; and a substrate RNA comprising, in a 5’ to 3’ direction, a therapeutic RNA, a 3’ untranslated region (3’ UTR), and a polyA tail, wherein at least one of the 3’ UTR and the polyA tail comprises at least one sequence-specific endoribonuclease recognition motif that corresponds to the endorib onucl ease, to produce cleavage products comprising the single-stranded therapeutic RNA.
2. A method according to Claim 1, wherein cleavage products include fragments of the single-stranded RNA.
3. A method according to Claim 1 or Claim 2, wherein the single-stranded therapeutic RNA comprises at least 5000 nucleotides, wherein the single-stranded RNA has at least one recognition sequence for the endonuclease.
4. A method according to any preceding claim, wherein the single-stranded RNA further comprises one or more modified nucleotides.
5. A method according to any preceding claim, wherein at least one of the modified nucleotides is within the recognition motif.
6. A method according to any preceding claim, wherein at least one of the modified nucleotides is N6-methyl-adenosine (m6A), 1-methyl-adenosine (nEA), 5-methyl-cytidine (m5C), 5-hydroxymethyl-cytidine (hm5C),N4-acetyl-cytidine (ac4C), 5-methoxycytidine (mo5C), 4-thiouridine (S4U), 2-thiouridine (S2U), pseudouridine ( ), NCmethyl- pseudouridine (nE ), 5-methyluridine (m5U), or 5- methoxyuridine (mo5U).
7. A method according to any preceding claim, wherein the amino acid sequence is at least 95% identical to any of SEQ ID NOS: 1-139, 163, and 164.
8. A method according to any preceding claim, wherein contacting further comprises contacting the substrate RNA, the endoribonuclease, and at least one additional endoribonuclease, wherein each of the at least one additional endoribonucleases has an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164 other than the amino acid sequence of the endoribonuclease.
9. A method comprises contacting: a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164; and a substrate RNA comprising a nucleotide sequence of interest heterologous to the endoribonuclease, a second nucleotide sequence, and at least one sequence-specific endoribonuclease recognition motif that corresponds to the sequence-specific endoribonuclease, wherein the recognition motif is disposed between the sequence of interest and the second nucleotide sequence, to produce cleavage products comprising the sequence of interest and the second sequence.
10. A method according to Claim 9, wherein the sequence of interest is an IVT template or a therapeutic RNA.
11. A method according to Claim 9, wherein the substrate RNA is a messenger RNA and the second sequence is at least a portion of a 3’ untranslated region (3’ UTR) or at least a portion of a poly A tail.
12. A method of analyzing a capped RNA comprising: contacting: a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164; and the capped RNA, wherein the capped RNA comprises at least one sequencespecific endoribonuclease recognition motif cleavable by the endorib onucl ease, to produce cleavage products, at least one of which comprises the cap; analyzing the at least one product comprising the cap by gel fractionation, liquid chromatography, tandem mass spectrometry, and / or nucleotide sequencing.- I l l -13. A method compri sing : contacting an RNA polymerase and a polynucleotide template, the polynucleotide template comprising, in a 5’ to 3’ direction, an expression control sequence corresponding to the polymerase and a sequence of interest to produce transcription products, the transcription products comprising, in a 5’ to 3’ direction, a sequence complementary to the sequence of interest, a 3’ untranslated sequence, and a polyA tail, wherein at least one of the 3’ untranslated sequence and the polyA tail comprises at least one sequence-specific endoribonuclease recognition motif; and contacting the transcription products and a sequence-specific endoribonuclease having an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164, wherein the endoribonuclease is operable to cleave the transcription products comprising the endoribonuclease recognition motif to produce cleavage products, wherein the cleavage products comprising the sequence complementary to the sequence of interest are homogenously sized.
14. A method according to Claim 13, wherein the sequence of interest is a coding sequence.
15. A method according to Claim 14, wherein the cleavage product comprising the sequence complementary to the sequence of interest is a therapeutic RNA.
16. A method according to Claim 13, wherein the cleavage product comprising the sequence complementary to the sequence of interest is a tRNA, an mRNA, a rRNA, a miRNA, a IncRNA, a piRNA, a snRNA, a snoRNA, or a chemically synthesized RNA.
17. A method according to Claim 13 further comprising contacting the cleavage product comprising the sequence complementary to the sequence of interest, a ligatable polynucleotide, and a ligase to produce a ligation product comprising the cleavage product comprising the sequence complementary to the sequence of interest and the ligatable polynucleotide.
18. A method according to Claim 17, wherein the ligatable polynucleotide is a sequencing comprises DNA, comprises RNA or comprises both DNA and RNA.
19. A method according to Claim 17, wherein the ligatable polynucleotide comprises a label and / or a protein.
20. A method according to Claim 17, wherein the ligatable polynucleotide is a sequencing adapter, a guide RNA, a linker, or a bar code.
21. A method according to Claim 17, wherein contacting the cleavage product comprising the sequence complementary to the sequence of interest, the ligatable polynucleotide, and the ligase further comprises contacting the cleavage product comprising the sequence complementary to the sequence of interest, the ligatable polynucleotide, the ligase, and a second ligatable polynucleotide.
22. A method according to Claim 17 further comprising contacting the cleavage product comprising the sequence complementary to the sequence of interest and a splint DNA, the splint DNA comprising at least a portion of the sequence of interest and at least a portion of the sequence of the ligatable polynucleotide.
23. A method for cleaving a single-stranded RNA, the method comprising: contacting: the RNA; and an endoribonuclease having an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164; to produce cleavage products, wherein the RNA comprises, in a 5’ to 3’ direction, (cs-ers)n(cs)m(A)x3 wherein cs encodes a sequence of interest, ers encodes an endoribonuclease recognition sequence,A is a poly(A) tail, m=Q or 1, n=>\, >2, >3, >4, >5, >10, >15, >20, >25, or >50, x=0 or 1, and3’ represents the 3’ end of the polycistronic RNA molecule,provided that if n=l, then m 0.
24. A method according to Claim 23, wherein n=>2.
25. A method according to Claim 23 or Claim 24, wherein n=>3.
26. A method according to any of Claims 23-25, wherein contacting further comprises contacting the single-stranded RNA with a second endoribonuclease and wherein the polycistronic template further comprises at least one endoribonuclease recognition sequence that corresponds to the second endoribonuclease.
27. A method according to any of Claims 23-26, wherein at least one of the cleavage products comprises an IVT template or a therapeutic RNA.
28. A method according to any of Claims 23-27, wherein contacting the RNA and the endoribonuclease further comprises contacting the RNA and the endoribonuclease at >65°C.
29. A method according to any of Claims 23-28, wherein contacting further comprises contacting the single-stranded RNA, the endoribonuclease, and at least one additional endoribonuclease, wherein each of the at least one additional endoribonucleases has an amino acid sequence that is at least 90% identical to any of SEQ ID NOS: 1-139, 163, and 164 other than the amino acid sequence of the endoribonuclease.
30. A method according to Claim 29, wherein the single-stranded RNA is a polycistronic template comprising at least one endoribonuclease recognition sequence that corresponds to the endoribonuclease and at least one endoribonuclease recognition sequence that corresponds to each of the at least one additional endoribonucleases.
Citation Information
Patent Citations
Purification and Purity Assessment of RNA Molecules Synthesized with Modified Nucleosides
US20160032316A1
Method for purifying RNA on a preparative scale by means of HPLC
US8383340B2
Modified nucleosides, nucleotides, and nucleic acids, and uses thereof
US9428535B2
Modified polynucleotides for the production of biologics and proteins associated with human disease
WO2013151666A2
MazF mutant, recombinant vector, recombinant engineering bacterium and application thereof
CN114480345A