Method for high-resolution CRYO-electron microscopy structure determination of RNA
By fusing small RNAs to a group II intron scaffold and using cryo-EM, the challenges of determining RNA structures are overcome, achieving high-resolution structures and enabling precise modeling of RNA-ligand interactions.
Patent Information
- Application Number
- PCT/US2024/057848
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-05
AI Technical Summary
Current methods for determining the structure of RNA using cryo-EM are challenging due to RNA's inherent flexibility and heterogeneity, leading to low to moderate resolution and lack of nucleotide-level detail, especially for small RNAs.
A strategy involving the fusion of small RNAs to a group II intron scaffold, allowing for the formation of a construct that enables high-resolution structure determination through cryo-EM. This approach includes various methods for attaching the RNA of interest to the scaffold, such as embedding, non-covalent attachment, 3’ end labeling, and circular permutation.
Achieves high-resolution structures of small to moderate-sized RNAs, previously intractable for cryo-EM studies, with resolutions greater than 3 Å, allowing for visualization of individual nucleotides and precise modeling of ligand binding sites.
Smart Images

Figure US2024057848_05062025_PF_FP_ABST
Abstract
Description
METHOD FOR HIGH-RESOLUTION CRYO-ELECTRON MICROSCOPY STRUCTURE DETERMINATION OF RNACROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 63 / 603,582, filed November 28, 2023, the contents of which is incorporated herein by reference in their entireties.STATEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with government support under Grant No. R35GM141706 from the National Institutes of Health. The government has certain rights in the invention.TECHNICAL FIELD
[0003] The present technology is generally related to methods for determining the structure of RNA.BACKGROUND OF THE DISCLOSURE
[0004] Over the past decade, cryo-EM (electron microscopy) has seen a tremendous advancement in both hardware and software. The advents of direct electron detectors with high frame rates and microscopes with extremely stable beams have provided structural biologists with images of unparalleled quality. In addition, processing software has become increasingly more sophisticated and user friendly. These innovations have created a tool capable of determining the structures of a wide variety of biological molecules. However, even with this tremendous power available, most structural biologists focus their attention on proteins and how they function. As the understanding of biological processes deepens, it is clear that other types of molecules play significant functional roles. In particular, RNA has been observed to participate in chemistry at the heart of some of the most important active sites in biology. Structures of both the ribosome and spliceosome have revealed that these complex machines use RNA to catalyze their reactions. RNA structure also plays a key role in guiding telomerase and Cas proteins to specific DNA sequences. In addition, important RNA structures are being identified in the untranslated regions of expressed mRNAs. These structures have been found to participate as important regulatory elements for gene expression. With such a diverse set of biological roles, RNA structures have become the focus of multiple therapeutics approaches. mRNA based therapeutics are now an efficient way to provide genes to act as vaccines, gene therapies, and even fight cancer. There are alsopharmaceutical companies that are targeting RNA structures with small molecules. However, there is a lack of RNA-only structures that have been solved. This deficiency of structures is due to the fact that RNA alone does not maintain a rigid structure and only appears in ensembles of heterogeneous populations. This heterogeneity complicates the structural biology workflow and thus RNA has been avoided by most structural biologists. This disclosure addresses this unmet need and provides related advantages as well.SUMMARY OF THE DISCLOSURE
[0005] Cryo-EM structure determination of protein-free RNAs has remained difficult with most attempts yielding low to moderate resolution and lacking nucleotide-level detail. These difficulties are compounded for small RNAs as cryo-EM is inherently more difficult for lower molecular weight macromolecules. A strategy is provided herein for fusing small RNAs (i.e. target RNA) to a group II intron that yields high resolution structures of the appended RNA. The RNA scaffold and target RNA are attached to form a construct (see e.g. FIGS. 1-4, 21).
[0006] In one aspect, the embodiments of this disclosure have been demonstrated by determining the structures of the 86-nucleotide (nt) thiamine pyrophosphate (TPP) riboswitch and the recently described 210-nt raiA bacterial non-coding RNA involved in sporulation and biofilm formation. In the case of the TPP riboswitch, the scaffolding approach allowed visualization of the riboswitch ligand binding pocket at 2.5 A resolution. Additionally, the structure of the ligand-free apo state has been determined and observed that the aptamer domain of the riboswitch adopts an open Y-shaped conformation in the absence of ligand. Applicant has determined raiA structure at 2.5 A in the core.
[0007] The disclosed versatile scaffolding strategy enables efficient RNA structure determination for a broad range of small to moderate-sized RNAs, which were previously intractable for high-resolution cryo-EM studies. A preferred RNA size range of the target RNA is about 25 to about 400 nucleotides in length. In one aspect the target RNA is about 150 to about 250 nucleotides in length.
[0008] In one aspect, provided is a scientific technical method of imaging small molecules, particularly RNA. Preferably, the resolution of the target RNA in the images is >3 A. Using the methods provided, high-resolution around 2.5 angstroms or better may be achieved allowing for visualization of individual nucleotides within RNA. In one aspect high- resolution of around 1.8 angstroms may be achieved using the methods provided. Theexisting art is limited to low-resolution around 5 angstroms, which is not useful for optimizing small molecule therapeutics targeting RNA.
[0009] In another aspect, described are different methods of attaching the RNA of interest to the group II intron scaffold (see FIGS. 1-4 and SEQ ID NOs: ). Each method provides unique advantages that help to solve the diverse set of difficulties that arise in RNA structure determination.
[0010] Method 1 : RNA embedding method. In this method the RNA of interest contains a helix made up of the 5’ and 3’ ends of the RNA (called a Pl helix) (FIG. 1). The RNA is added with the correct phosphate backbone polarity at the position designated within the sequence of the scaffold.
[0011] Method 2: Non-covalent RNA attachment method. In this method, the scaffold has the Tecto RNA 1 stem loop added to the position designated within the sequence (FIG. 2). The RNA of interest then has a stem loop extend with the Tecto RNA 2 stem loop. The RNA of interest is then mixed with the scaffold to assemble a non-covalent complex. Additional information about Tecto RNA is found in Luc Jaeger, et al. TectoRNA: modular assembly units for the construction of RNA nano-objects, Nucleic Acids Research, Volume 29, Issue 2, 15 January 2001, Pages 455-463, DOI: 10.1093 / nar / 29.2.455, the contents of which are incorporated herein by reference.
[0012] Method 3: 3’ end labeling the scaffold. In this method, the RNA of interest is added to the 3’ end of the scaffold separated by a variable length poly adenosine (Poly A) linker (FIG. 3). The linker length can vary depending on the size of the RNA of interest.
[0013] Method 4: Circular permutation of RNA scaffold: In this method the scaffold RNA is circularly permuted as shown in FIG. 4. The circularly permuted scaffold is then embedded within a stem of the RNA of interest. This allows the RNA of interest to be embedded within the scaffold but maintains the endogenous folding of the 5’ and 3’ ends of the RNA of interest.
[0014] Thus, in one aspect, a method is provided for RNA imaging using an RNA scaffold. The method includes, or consists essentially of, or yet further consists of imaging a target RNA and a RNA scaffold construct using cryo-electron microscopy; wherein the target RNA is embedded within the RNA scaffold sequence.
[0015] In another aspect, a method for RNA imaging using an RNA scaffold, the method includes, or consisting essentially of, or yet further consisting of imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: a stem loop of the target RNA comprises a Tecto RNA 2 stem loop; the RNA scaffold comprises a Tecto RNA 1 stem loop; and the Tecto RNA 1 stem loop and the Tecto RNA 2 stem loop bind to each other or alternatively are bound to each other.
[0016] In a further aspect, a method for RNA imaging using an RNA scaffold is provided, the method comprising, or consisting essentially of, or yet further consisting of imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the target RNA is bound to a first end of a linker sequence; and the RNA scaffold is bound to a second end of a linker sequence.
[0017] In a yet further aspect, a method for RNA imaging using an RNA scaffold is provided, the method comprising, or consisting essentially of, or yet further consisting of imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the RNA scaffold is circularly permutated and is embedded in the target RNA; and the scaffold is embedded in a stem of the target RNA.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] FIG. 1 : is a secondary structure representation of an RNA scaffold using embedding technique, according to one embodiment. The sequence shown in FIG. 1 (SEQ ID NOs: ) is an alternative embodiment to the structure provided by Method 1 in Example 1 (SEQ ID NO. )•
[0019] FIG. 2 : is a secondary structure representation of an RNA scaffold using a Tecto RNA interaction, according to one embodiment. The sequences of FIG. 2 (SEQ ID NOs: ) is an alternative embodiment to the structure provided by Method 2 in Example 1.
[0020] FIG. 3 : is a secondary structure representation of an RNA scaffold using a 3 ’ end labeling technique, according to one embodiment. The sequences of FIG. 3 (SEQ ID NOs: ) is an alternative embodiment to the structure provided by Method 3 in Example 1.
[0021] FIG. 4 : is a secondary structure representation of an RNA scaffold using a circular permutation, according to one embodiment. The sequences of FIG. 4 (SEQ ID NOs: ) is an alternative embodiment to the structure provided by Method 4 in Example 1.
[0022] FIGS. 5A and 5B: illustrate an RNA scaffold application for determining the cryo- EM structure of the TPP riboswitch, according to various embodiments. FIGS. 5A and 5B describe the quality of the data and shows that the scaffold reaches 2.2 A in the core. The embedded riboswitch density, through focused classification, reaches 2.5 A in the core and allows for the unambiguous modeling of the bound TPP ligand and the metal ions associated with its binding.
[0023] FIGS. 6A, 6B, and 6C: Initial scaffold development using a group IIB intron RNP. (FIG. 6A) The maps of the group IIB RNP were generated using a tilted data collection strategy with the stage set to 30°. The map on the left was globally refined and shown at both a high and low threshold. The group II intron RNP displays clear RNA features at high threshold and density for the embedded RNA is observed at low threshold. Performing local refinement on the density corresponding to the embedded RNA yielded the map on the left which has been overlaid on the low threshold globally refined map for clarity. (FIG. 6B) A model for the embedded RNA was fit into the locally refined map. The fit shows that the overall fold of the embedded RNA was maintained. (FIG. 6C) The viewing direction distribution for the data set is shown. Even with a 30° stage tilt, there is still a preferred orientation for this specimen.
[0024] FIGS. 7A - 7D: Identification of a high-resolution group IIC intron scaffold. (FIG. 7A) The secondary structure of the Oceanobacillus iheyensis (O.i.) group IIC intron scaffold is shown, indicating domain III. The mutation to the G to A mutation to the active site is highlighted red. (FIG. 7B) Cryo-EM density is shown for the scaffold. Domain III forms a stable stem loop that is extruded away from the rest of the folded RNA. (FIG. 7C) A local resolution map shows that the core of the scaffold has a resolution of 2.4 A. (FIG. 7D) The viewing direction distribution plot generated from cryoSPARC is shown. The plot is evenly populated indicating that the group IIC intron scaffold exhibits no preferred orientation. The polynucleotide shown in FIG. 7A also is provided in SEQ ID NO: .
[0025] FIGS. 8A-8E: High-resolution cryo-EM structure determination of the TPP riboswitch using the group IIC intron scaffold. (FIG. 8A) A secondary structure of the TPP riboswitch attached to the scaffold through the stem of domain III of O.i. is shown. Nucleotides in the TPP riboswitch aptamer domain are indicated SHAPE reactivity. Experiments were performed in the presence of ligand; for full SHAPE data, see FIG. 17. . (FIG. 8B) 2D class averages of the scaffold alone and attached to the TPP riboswitch areshown. Signal corresponding to the TPP riboswitch (star) is clearly visible in the 2D classes. (FIG. 8C) A local resolution map of the scaffold (left) on was generated from focused refinement and overlaid on a low threshold globally refined map. The locally refined map of scaffold maintains high-resolution features after the TPP riboswitch is embedded. Density for the TPP riboswitch is visible in the globally refined map at low threshold. A locally refined map of the TPP riboswitch is shown on the right and overlaid on the same low threshold globally refined map for clarity. (FIG. 8D) A local resolution map for the TPP riboswitch. The resolution surrounding the TPP binding site is 2.5 A. (FIG. 8E) A viewing direction distribution plot is shown for the globally refined scaffold TPP data. The scaffold does not show any orientation preference with the added sequence of the riboswitch. FIG. 8A is the sequence also provided in SEQ ID NO: .
[0026] FIGS. 9A and 9B: High-resolution cryo-EM structure of the TPP riboswitch. (FIG. 9A) The unsharpened cryo-EM map as well as the build models are shown for the scaffolded TPP riboswitch. The density and model is colored to match the secondary structure in FIG. 8A. (FIG. 9B) The high-resolution density surrounding the thiamine pyrophosphate binding site allowed for the precise characterization of ligand binding by the riboswitch. The model overlaid with density is shown below.
[0027] FIGS. 10A and 10B: A model for translational regulation by the TPP riboswitch. (FIG. 10A) Density for the apo and TPP bound states of the TPP riboswitch are shown. The TPP bound map was low-pass filtered to 6 A. The P3-L5 interaction is disengaged in the apo state (left) allowing the pyrophosphate sensing stem to rotate causing the riboswitch to adopt an overall Y shaped conformation. In the apo state the base of the L5 stem loop is largely disordered with portions of Pl, J2-4, and P4 missing density completely (circled). In contrast, the thiamine sensing stem appears to have an alternate stable conformation in the apo state with the appearance of an additional minor groove. When TPP is bound (right), the P3-L5 interaction is engaged, and the riboswitch compacts leading to ordered Pl, J2-4, and P4 regions (circled) properly forming the three-way junction. (FIG. 10B) The above structural information provides insights into the mechanism of translational regulation by the TPP riboswitch. In the apo state, the base of the L5 stem loop is disordered allowing it to interact with the downstream regulatory element through complementary base pairing. This interaction causes a rearrangement of the regulatory element structure making the Shine- Dalgamo sequence (SD) accessible for the ribosome to initiate translation. The thiamine sensing stem having a stable conformation in the apo state likely promotes the formation ofthis L5 stem loop / regulatory element interaction. In the bound state, the equilibrium is shifted to the compacted state with an ordered L5 stem loop which prohibits the interaction of the L5 stem loop with the regulatory element. This ultimately leads to a loss in ribosome initiation and translation.
[0028] FIGS. 11A and 11B: cryo-EM structure of the raiA ncRNA motif. (FIG. 11 A) The secondary structure of the raiA motif from Clostridium acetobutylicum is shown. The RNA consists of seven helical stems (P1-P8). Each individual helix is indicated for clarity and relevant tertiary interactions are labeled. The pk-Loop (is made up of the junction nucleotides between P5 and P6. A secondary structure that more closely follows the tertiary structure is shown outlining the organization of the helices relative to each other in three- dimensional space. Symbols to denote noncanonical base pairs are shown in blue and use the Leontis and Westhof nomenclature54. (FIG. 11B) The unsharpened cryo-EM map of the scaffolded raiA non-coding RNA motif is shown. FIG. 11A has a sequence as provided in SEQ ID NO: .
[0029] FIGS. 12A - 12C: The pk-1 and pk-2 interactions form the core of the raiA motif. (FIG. 12A) The P3 loop (bottom, on left) of raiA forms an extended pseudoknot with the pk- Loop (top, on left) (pk-1). The pseudoknot consists of 7 consecutive cis-W:W base pairs. (FIG. 12B) pk-2 forms between the pk-Loop and the junction between P4 and Pic (J4- 1 c). This pseudoknot is not as extensive as pk-1 and only forms 2 cis-W:W. base pairs. (FIG. 12C) The sequence of pk-1 and pk-2 contained within the pk-Loop is only separated by a single nucleotide (G84). G84 base stacks with first nucleotide of the P3 Loop (A33) to further stabilize the pk-l / pk-2 structure.
[0030] FIGS. 13A and 13B: The Pl and P3 stems contain significant bends to their helical axis. (FIG. 13A) The Jlc-3a and J4-lc junctions turn the helical axis of Pl relative to P3 by about 180°. The helical axis of the Pl and P3 stems is represented by black arrows. (FIG. 13B) The I3a-3b internal loop forms an approximate 90° bend of the helical axis between the P3a and P3b stems.
[0031] FIGS. 14A - 14C: P4 and P5 form a continuous helix and stabilize pk-1 through an A-minor motif. (FIG. 14A) The structural organization of the P4 and P5 helices is shown relative to pk-1. The minor grooves of P5 and pk-1 point towards each other and form an extended interface of hydrogen bonds. (FIG. 14B) The base of P5 consists of several noncanonical base pairs the not only forms the continuous helix between P4 and P5, but alsoextrudes A57 and A58 from the helical axis. (FIG. 14C) A57 and A58 from the base of P5 extend into the minor groove of pk-1 forming an extended A-minor motif interaction.
[0032] FIGS. 15A and 15B: The P8 stem loop also interacts with pk-1 through an A-minor motif. (FIG. 15A) The organization of the P6, P7, and P8 stems are shown relative to pk-1. P8 consists of numerous noncanonical base pairs and as a result exhibits a distorted helical geometry. The minor grooved of P8 and pk-1 can be observed interacting. (FIG. 15B) The conserved E-Loop region of P8 forms through multiple consecutive noncanonical base pairs. This architecture results in a helical geometry capable of forming an A-minor motif with the opposite side of pk-1 relative to the P4 / P5 stem loop.
[0033] FIG. 16: Current state of cryo-EM on protein-free RNA samples. Unsharpened 3D reconstructions of a representative selection of protein-free RNAs from the RCSB Protein Data Bank are shown. The glycine riboswitch from V. cholera (PDB-6WLU) is a 231-nt RNA and 3D reconstruction was determined to a reported resolution of 5.7 A. A 3D reconstruction of a 171-nt tRNA-like structure from B. mosaic (PDB-7SAM) was determined to a reported 4.3 A. 3D reconstructions of the 52-nt fluoride riboswitch (PDB-8TJV) and 71- nt Zika virus xrRNA (PDB-8TJQ) were determined using a group I intron scaffold and were reported at 4.46 A and 5.05 A respectively. In all these examples, models were built into the determined maps and representative regions of the models fit to density are shown. In these examples, as shown in the zoomed-in insets, the sequence register is uncertain due to the low resolution. The bottom 3D reconstruction of the 86-nt thiamine pyrophosphate riboswitch (PDB-9C6K) was determined in this study using a group II intron as a scaffold to a resolution of 3.1 A. At this resolution, de novo modeling is possible with the separation between base pairs allowing for the accurate determination of sequence register. In addition, density for the bound thiamine pyrophosphate ligand is visible, which is not possible at the resolution seen in the prior examples.
[0034] FIGS. 17A - 17C: SHAPE-MaP probing of scaffold-TPP riboswitch construct in the absence and presence of TPP ligand. (FIG. 17A) SHAPE reactivity profiles in the absence or presence of TPP ligand. TPP riboswitch region is emphasized with green box. Riboswitch structural landmarks are highlighted. (FIG. 17B) SHAPE reactivities for the scaffold-riboswitch construct, measured in the presence of TPP, superimposed on a secondary structure model of the complete RNA. (FIG. 17C) Focused view of SHAPE reactivities for the TPP riboswitch aptamer domain, as appended to the scaffold RNA, in theabsence and presence of TPP ligand. Secondary structures were modeled using the AGSHAPE framework 41). For the ligand-free state, helices P2alt and P3alt overlap with, but are not identical to, P2 and P3 visualized in the TPP -bound state, consistent with prior studies (28). In panels B and C, nucleotides are shaded by SHAPE reactivity: shading from top to bottom on the SHAPE reactivity scale corresponds to high, medium, and low reactivities, respectively. FIG. 17A has a sequence according to SEQ ID NO: . FIG. 17B has a sequence according to SEQ ID NO: . FIG. 17C has a sequence according to SEQ ID NO: .
[0035] FIG. 18: Cryo-EM processing workflow of O.z.-TPP. Data processing workflow for the O.z.-TPP construct focused on O.i. and then focused on TPP. This represents the general workflow for the scaffold approach and can be applied to any RNA of interest. GSFSC resolution curves and viewing direction distribution plots are shown for both locally refined maps.
[0036] FIG. 19: Comparison of the scaffold structure with and without an attached RNA. Maps correspond to the scaffold alone and scaffold-TPP samples. The presence of the TPP riboswitch attached to Domain III does not alter the overall 3D reconstruction of the group II intron. A model for the O.i. intron (4DS6) was refined in real space into each of the maps and the resulting structure coordinates were aligned using LSQ superpose in COOT (right).
[0037] FIG. 20: Comparison of cryo-EM and crystal structures of the TPP riboswitch. The cryo-EM derived and crystallography derived (2GDI) structure coordinates were aligned using LSQ superpose in COOT. Full RMSD deviations were determined for the aligned maps in Chimera using Match -> Align. The results of this analysis were graphed (left) and annotated with the corresponding regions of the riboswitch.
[0038] FIG. 21: Scaffolded raiA construct. The RNA sequence for the O.i.-raiA construct is shown. FIG. 21 has a sequence according to SEQ ID NO: .
[0039] FIGS. 22A - 22D: Cryo-EM processing workflow of O.i.-raiA. (FIG. 22A) Data processing workflow for the O.i.-raiA construct is shown. (FIG. 22B) Local resolution map of the raiA reconstruction. (FIG. 22C) Viewing direction distribution plot. (FIG. 22D) Gold standard FSC curve.
[0040] FIG. 23: The Pl helix terminates in a discontinuous GNRA tetraloop. The three- way junction joining the Pl, P3, and P4 stems forms a discontinuous GNRA tetraloop.Within the discontinuous GNRA like motif, a trans-S:W base pair between G20:A183 closes the loop. A21, A22, and A183 then form a continuous base stack to complete the motif. Theresult of this discontinuous tetraloop is an approximate 180° bend of the helical axis between Pl and P3.
[0041] FIGS. 24A - 24C: P3 contains an internal loop that forms a 90° helical bend. (FIG. 24A) The I3a-3b internal loop adopts an approximate 90° bend in the helical axis between the P3a and P3b stems. Both A44 and A46 are extruded out of the internal loop. (FIG. 24B) The lib- 1 c internal loop contains multiple conserved nucleotides that form a hydrogen bonding network to stabilize the motif. (FIG. 24C) G10 forms a base triple with the G14:C193 cis- W:W base pair.
[0042] FIGS. 25A and 25B: SHAPE reactivity profile of the raiA RNA. (FIG. 25A) SHAPE reactivities of the raiA RNA as shown. RNA was inserted into the group II intron scaffold. Non-base-paired regions with low SHAPE reactivity are labeled below the nucleotide sequence. (FIG. 25B) SHAPE reactivities for the scaffold-ra / d construct, superimposed on a secondary structure model of the complete RNA as inferred from the cryo-EM structure. Pseudoknot pairings are indicated by dashed lines. Non-base-paired regions with low SHAPE reactivities, indicative of stable, non-canonical structure, are labeled with bold text. . FIG. 25B has a sequence according to SEQ ID NO: .DETAILED DESCRIPTION
[0043] Various embodiments are described hereinafter. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment(s).
[0044] As utilized herein with respect to numerical ranges, the terms “approximately,” “about,” “substantially,” and similar terms will be understood by persons of ordinary skill in the art and will vary to some extent depending upon the context in which it is used. If there are uses of the terms that are not clear to persons of ordinary skill in the art, given the context in which it is used, the terms will be plus or minus 10% of the disclosed values. When “approximately,” “about,” “substantially,” and similar terms are applied to a structural feature (e.g., to describe its shape, size, orientation, direction, etc.), these terms are meant to cover minor variations in structure that may result from, for example, the manufacturing orassembly process and are intended to have a broad meaning in harmony with the common and accepted usage by those of ordinary skill iFn the art to which the subject matter of this disclosure pertains. Accordingly, these terms should be interpreted as indicating that insubstantial or inconsequential modifications or alterations of the subject matter described and claimed are considered to be within the scope of the disclosure as recited in the appended claims.
[0045] The use of the terms “a” and “an” and “the” and similar referents in the context of describing the elements (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the claims unless otherwise stated. No language in the specification should be construed as indicating any non-claimed element as essential.
[0046] It is to be understood that this disclosure is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of this disclosure will be limited only by the appended claims.
[0047] The detailed description is divided into various sections only for the reader’s convenience and disclosure found in any section may be combined with that in another section. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0048] All numerical designations, e.g., pH, temperature, time, concentration, and molecular weight, including ranges, are approximations which are varied ( + ) or ( - ) by increments of 0.1 or 1.0, where appropriate. It is to be understood, although not always explicitly stated, that all numerical designations are preceded by the term “about.” It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art.
[0049] It must be noted that as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a cell” includes a plurality of cells.
[0050] As will be understood by one skilled in the art, for any and all purposes, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Furthermore, as will be understood by one skilled in the art, a range includes each individual member.
[0051] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of tissue culture, immunology, molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rd edition; the series Ausubel et al. eds. (2007) Current Protocols in Molecular Biology; the series Methods in Enzymology (Academic Press, Inc., N.Y.); MacPherson et al. (1991) PCR 1 : A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5th edition; Gait ed. (1984) Oligonucleotide Synthesis; U.S. Patent No. 4,683,195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins eds. (1984) Transcription and Translation; Immobilized Cells and Enzymes (IRL Press (1986)); Perbal (1984) A Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London);Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology; Manipulating the Mouse Embryo: A Laboratory Manual, 3rd edition (Cold Spring Harbor Laboratory Press(2002)); Sohail (ed.) (2004) Gene Silencing by RNA Interference: Technology and Application (CRC Press).
[0052] The term “about” when used before a numerical designation, e.g., temperature, time, amount, concentration, and such other, including a range, indicates approximations which may vary by ( + ) or ( - ) 10 %, 5 % or 1 %.
[0053] “Comprising” or “comprises” is intended to mean that the compositions, for example media, and methods include the recited elements, but not excluding others. “Consisting essentially of’ when used to define compositions and methods, shall mean excluding other elements of any essential significance to the combination for the stated purpose. Thus, a composition consisting essentially of the elements as defined herein would not exclude other materials or steps that do not materially affect the basic and novel character! stic(s) of the claimed invention. “Consisting of’ shall mean excluding more than trace elements of other ingredients and substantial method steps. Embodiments defined by each of these transition terms are within the scope of this invention.
[0054] As used herein, comparative terms as used herein, such as high, low, increase, decrease, reduce, or any grammatical variation thereof, can refer to certain variation from the reference. In some embodiments, such variation can refer to about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 80%, or about 90%, or about 1 fold, or about 2 folds, or about 3 folds, or about 4 folds, or about 5 folds, or about 6 folds, or about 7 folds, or about 8 folds, or about 9 folds, or about 10 folds, or about 20 folds, or about 30 folds, or about 40 folds, or about 50 folds, or about 60 folds, or about 70 folds, or about 80 folds, or about 90 folds, or about 100 folds or more higher than the reference. In some embodiments, such variation can refer to about 1%, or about 2%, or about 3%, or about 4%, or about 5%, or about 6%, or about 7%, or about 8%, or about 0%, or about 10%, or about 20%, or about 30%, or about 40%, or about 50%, or about 60%, or about 70%, or about 75%, or about 80%, or about 85%, or about 90%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% of the reference.
[0055] “Optional” or “optionally” means that the subsequently described circumstance may or may not occur, so that the description includes instances where the circumstance occurs and instances where it does not.
[0056] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).
[0057] “Substantially” or “essentially” means nearly totally or completely, for instance, 95% or greater of some given quantity. In some embodiments, “substantially” or “essentially” means 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0058] The terms or “acceptable,” “effective,” or “sufficient” when used to describe the selection of any components, ranges, dose forms, etc. disclosed herein intend that said component, range, dose form, etc. is suitable for the disclosed purpose.
[0059] The terms “polynucleotide”, “nucleotide”, “nucleotide sequence”, “nucleic acid” and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three dimensional structure, and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. The term also encompasses nucleic-acid-like structures with synthetic backbones, see, e.g., WO 97 / 03211 and WO 96 / 39154. A polynucleotide may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0060] As used herein the term “wild type” is a term of the art understood by skilled persons and means the typical form of an organism, strain, gene or characteristic as it occurs in nature as distinguished from mutant or variant forms.
[0061] As used herein, the terms “embedded” and “embeds” and “embedding” refer to the introduction of one nucleic acid sequence into a second nucleic acid sequence at a chosen site in the second nucleic acid sequence. In some embodiments, a target RNA is embedded within an RNA scaffold. In other embodiments, an RNA scaffold is embedded within atarget RNA. In some embodiments, the chosen site in the second nucleic acid sequence is within a structure of the second sequence, such as, for example, a stem or a loop structure.
[0062] As used herein, the terms “circularly permutated” and “circularly permute” refer to a method of macromolecular isomerization, wherein the endogenous termini of a nucleic acid sequence (e.g. the 5’ and 3’ ends of a RNA sequence) are covalently linked and new termini are introduced by breaking the sequence at another site. In some embodiments, the endogenous termini are covalently linked to each other directly. In some embodiments, the endogenous termini are covalently linked using a linker, such as a poly adenosine linker.
[0063] As used herein, the term “a polyadenylation linker (PolyA linker)” can be represented as “AA(A)„, - wherein n is an integer that is 0 or about from 1 to about 5, or about 10, or about 15, or about 20, or about 25, or about 30, or about 35, or about 40, or about 45, or about 50. In the RNA scaffold and target RNA construct, the PolyA linker is preferably less than 50 nucleotides, or preferably less than 40 nucleotides.
[0064] As used herein, “derived from” means isolated from, purified from, or engineered from, or any combination thereof.
[0065] As used herein, “tectoRNA” refers to RNA molecules capable of self-assembly to form nanoscale structures. TectoRNAs can be used as scaffolding to determine binding affinities of RNA-RNA interactions. See Luc Jaeger et al, TectoRNA: modular assembly units for the construction of RNA nano-objects, Nucleic Acids Research, Volume 29, Issue 2, 15 January 2001, Pages 455-463, DOI: 10.1093 / nar / 29.2.455.Modes For Carrying Out the Disclosure
[0066] Construct and Methods of Imaging
[0067] Provided herein is an RNA scaffold derived from or based on a group II intron that helps overcome the above obstacles. In some aspects, the RNA scaffold is derived from or based on a group IIC intron. In some aspects, the RNA scaffold is derived from or based on a group IIA or IIC intron. By providing the RNA of interest, also referred to as “target RNA,” with a stable platform to promote homogeneous folding, it is now possible to determine high resolution RNA-only structures. The RNA scaffold and target RNA are attached to form a construct (see e.g. FIGS. 1-4, 21).
[0068] In one aspect, described herein are four different methods of attaching the RNA of interest to the group II intron scaffold (FIGS. 1-4) to form the construct. Each methodprovides unique advantages that help to solve the diverse set of difficulties that arise in RNA structure determination.
[0069] In a first embodiment, an RNA embedding method is described. In the RNA embedding method, the RNA of interest contains a helix made up of the 5’ and 3’ ends of the RNA (called a Pl helix) (FIG. 1). The RNA is added with the correct phosphate backbone polarity at the position designated within the sequence of the scaffold. An exemplary structure of the RNA of interest embedded in the scaffold is provided in the Example Construct Sequences Method 1 as well as FIG. 1.
[0070] In a second embodiment, a non-covalent RNA attachment method is described. In this method, the scaffold has the Tecto RNA 1 stem loop added to the position designated within the sequence (FIG. 2). The RNA of interest then has a stem loop extend with the Tecto RNA 2 stem loop. The RNA of interest is then mixed with the scaffold to assemble a non-covalent complex. Exemplary structures of the non-covalent RNA attachment method is provided in the Example Construct Sequences Method 2 as well as FIG. 2.
[0071] In a third embodiment, a method is described for 3’ end labeling of the scaffold. In this method, the RNA of interest is added to the 3’ end of the scaffold separated by a variable length poly adenosine (Poly A) linker (FIG. 3). The linker length can vary depending on the size of the RNA of interest. In one aspect the linker length is about 2 to about 40 nucleotides in length. Exemplary structures of the 3’ end labeling of the scaffold method is provided in the Example Construct Sequences Method 3 as well as FIG. 3.
[0072] In a fourth embodiment, a method is described for circular permutation of an RNA scaffold. In this method the scaffold RNA is circularly permuted as shown in FIG. 4. The circularly permuted scaffold is then embedded within a stem of the RNA of interest. This allows the RNA of interest to be embedded within the scaffold but maintains the endogenous folding of the 5’ and 3’ ends of the RNA of interest. Exemplary structures of the circular permutation of an RNA scaffold method is provided in the Example Construct Sequences Method 4 as well as FIG. 4.
[0073] In some aspects, the group II intron is derived from at least one bacteria. For example the group II intron may be derived from Oceanobacillus iheyensis or Thermosynecoccus elongatus.
[0074] In some aspects, the group II intron RNA scaffold is about 80ka, about lOOkDa, about 120kDa, about 140kDa, about 160kDa, about 180kDa, about 200kDa, about 220kDa,about 240kDa, about 260kDa, about 280kDa, about 300kDa, about 320kDa, about 340kDa, about 360kDa, about 380kDa, about 400kDa, about 420kDa, about 440kDa, about 460kDa, or about 480kDa or about 500kDa. In one aspect, the group II intron RNA scaffold is about 80kDa to about 140kDa. In some aspects, the group II RNA scaffold is larger than 500kDa.
[0075] In some aspects, the group II intron RNA scaffold is about 350nt to about lOOOnt in length. In one aspect, the group II intron RNA scaffold is about 400nt to about 850nt in length. In some aspects, the group II intron RNA scaffold further includes an intron encoded protein (IEP). Group II introns with lEPs may be >1000nt.
[0076] Preferably, the RNA is about 25nt, 50nt, about 75nt, about lOOnt, about 125nt, about 150nt, about 175nt, about 200nt, about 250nt, about 250nt, about 275nt, or about 300nt in length. In some aspects the RNA is about 25nt to about 400nt or alternatively about 600nt in length. In some aspects, the RNA is about 25nt to about 400nt in length. In some aspects, the RNA is longer than 600nt. In some aspects the target RNA is a riboswitch. In some aspects the target RNA is a non-coding RNA. In some aspects, the target RNA is part of an RNA-protein complex. The target RNA can have a Y-shaped, compact, P-shaped, or candy conformation.
[0077] In some aspects, a ligand is bound to the target RNA. Non-limiting examples of ligands include thiamine pyrophosphate (TPP). In some aspects the methods are used to determine the structure of a drug in complex with the RNA it targets. A non-limiting example is risdiplam bound to its target site in SMN2 exon 7).
[0078] Images
[0079] According to another embodiment, provided herein are images taken using the methods described herein.
[0080] In one aspect, the resolution of the target RNA in the image is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A. Preferably, the resolution of the target RNA in the image is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, or about 3.0 A.
[0081] In one aspect, resolution of a core of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A. Preferably, the resolution of the coreof the target RNA in the image is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, or about 3.0 A. The core of the target RNA can be the center of mass of the RNA. The core of the target RNA is the focal point of the image alignment for cryo-electron microscopy.
[0082] The methods and images as described herein can be used, for example, in the visualization of binding of small molecule therapeutics to target RNAs for optimizing drug design.
[0083] The present aspects of this disclosure, thus generally described, will be understood more readily by reference to the following examples, which are provided by way of illustration and are not intended to be limiting.EXAMPLES
[0084] Example 1
[0085] Exemplary use of the scaffold: In this example, the thiamine pyrophosphate riboswitch was embedded within the scaffold using the RNA embedding method described above. FIG. 5 describes the quality of the data and shows that the scaffold reaches 2.2 A in the core. The embedded riboswitch density, through focused classification, reaches 2.5 A in the core and allows for the unambiguous modeling of the bound TPP ligand and the metal ions associated with its binding.
[0086] Exemplary constructs and sequences for each of Methods 1-4 are shown in FIGS. 1-4. Alternative embodiments of exemplary constructs are included herein in Example 1.Partial Sequence ListingConstruct SequencesMethod 1:General structure: scaffold fragment-Xi-scaffold fragment, wherein Xi is the target RNA (embedded RNA, noted in bolded lettering).5 ’ GTGTGCCCGGC ATGGGTGC AGTCT AT AGGGTGAGAGTCCCGAACT GTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGC GACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGA ACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATA ACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAG CAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCT - Embedded-RNA -AGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGT 3’ (SEQ ID NO: )Method 2:General structure:1) scaffold fragment (including Tecto RNA 1 sequence)2) Xi-Tecto RNA 2-X2,1) and 2) are non-covalently bound to each other wherein Xi and X2are fragments of the target RNA; and 1) and 2) are non-covalently bound to each otherScaffold: Group II intron (Black) Tecto RNA 1 (Bolded Text)5 ’ GTGTGCCCGGC ATGGGTGC AGTCT AT AGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGagctttcgagctCAGAAGTCAGCAGAAGTCATAGTACCCGGGAGAAGGGTAGTTCCGGGGAAACTTGGTTCTACCCCACGCTCCTGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCATGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGT 3’ (SEQ ID NO: )RNA of interest (Black) Tecto RNA 2 (Bolded Text)5 ’ - RNA-of-Interest -GGGATATGGAAGTTCCGGGGGAACTTGGTTCTTCCTAAGTCCT- RNA-of-Interest - 3 ’ (SEQ ID NO: )Method 3:General structure:Scaffold- AA(A)n-Xi, wherein Xi is the target RNA5 ’ GTGTGCCCGGC ATGGGTGC AGTCT AT AGGGTGAGAGTCCCGAAC TGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTG GCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGA GGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTG CATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAA TGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAG GTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTAGGGGAA GGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTA TACCGAACGGTACGTACGGTGGTGTG— PolyA-Linker -- RNA-of-Interest - 3’ (SEQ ID NO: )Method 4:General Structure:Xi-scaffold fragment- AA(A)n-scaffold fragment- X2, wherein Xi and X2 are fragments of the target RNA5 ’ - RNA-of-Interest -AGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGT - PolyA-Linker-GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCT -- RNA-of-Interest - 3’ (SEQ ID NO: )
[0087] Example 2
[0088] Experiment
[0089] The design of small molecule drugs to target proteins is notably accelerated by studying structure activity relationships (SAR) to optimize binding to protein targets, as informed by routine visualization of protein-ligand complexes. However, structure-informedSAR is currently difficult and rare for the development of small molecules targeting RNA structures. Most RNAs do not crystallize readily for x-ray crystallography and also present difficulties for cryo-electron microscopy (cryo-EM) structure determination. In Applicant’s experience, vitrification on cryo-EM grids often results in the denaturation and aggregation of RNA1and often only a small number of native particles are visualized, which is insufficient to obtain a high-resolution reconstruction. In addition, the few particles that are visualized often exhibit a preferred orientation such that the views of the particle are not sufficiently populated for high resolution 3D reconstruction. Therefore, there is a compelling need for technologies to overcome these obstacles and enable routine high-resolution cryo-EM analysis of protein-free RNAs. Obtaining high-resolution structures of RNA and RNA-ligand complexes is crucial for establishing molecular mechanisms of action in multiple biological contexts.
[0090] Cryo-EM has revolutionized structural biology with thousands of protein structures solved and shared in the Protein Data Bank (PDB). In contrast, there has only been a single protein-free RNA (Tetrahymena group I intron) determined at 3 A resolution or better using cryo-EM2. The remainder of current protein-free RNAs exhibit resolutions ranging from 5 to 10 A3, suggesting that RNA is not as amenable to cryo-EM structure determination as are proteins. At 5 A resolution, there are two major problems: de novo modeling is not possible, especially in complex helical junctions and sequence register cannot be determined due to the lack of nucleobase separation. It is possible to place computationally derived RNA structures into 5 A density, however these methods currently do not accurately model complex regions. For example, the glycine riboswitch3and a tRNA-like viral structure4were reported at 5.7 and 4.3 A, respectively. In each case, examination of the model fit to density shows an inability to accurately determine sequence register, with nucleotides modeled outside of density (FIG. 16).
[0091] Cryo-EM has revolutionized structural biology with thousands of protein structures solved and shared in the Protein Data Bank (PDB). In contrast, there has only been a single protein-free RNA (Tetrahymena group I intron) determined at 3 A resolution or better using cryo-EM2. The remainder of current protein-free RNAs exhibit resolutions ranging from 5 to 10 A3, suggesting that RNA is not as amenable to cryo-EM structure determination as are proteins. At 5 A resolution, there are two major problems: de novo modeling is not possible, especially in complex helical junctions and sequence register cannot be determined due to the lack of nucleobase separation. It is possible to place computationally derived RNA structuresinto 5 A density, however these methods currently do not accurately model complex regions. For example, the glycine riboswitch3and a tRNA-like viral structure4were reported at 5.7 and 4.3 A, respectively. In each case, examination of the model fit to density shows an inability to accurately determine sequence register, with nucleotides modeled outside of density (FIG. 16).
[0092] One approach for solving RNA structures via cryo-EM is to use a scaffold to which an RNA motif of interest is appended. The goal is that the favorable biophysical properties that result in high resolution for the scaffold would propagate to the target RNA. The larger mass of the overall RNA also facilitates particle picking and alignment for 3D reconstruction. This approach has been attempted with the Tetrahymena group I intron, in multiple adaptations5,6. This group I intron, in isolation, has been solved to about 3 A resolution2. However, to date, this approach has resulted in only about 5 A resolution for the target RNA5,6, which does not allow for the discrimination of individual nucleotides and makes precise modeling impossible (FIG. 16). For example, when the group I intron was used as a scaffold to determine the structures of the fluoride riboswitch and the Zika virus xrRNA, modeling of these RNA structures required the use of prior high-resolution crystal structures. At this resolution, it was not possible to perform de novo modeling (zoomed-in insets in FIG. 16), as noted previously by Langeberg et al. (2023). The scaffold approach has also been attempted using RNA origami7and the ribosome8, with similar results. In summary, current scaffolding approaches have not yet demonstrated high resolution structure determination of attached target RNAs to allow discrimination of individual nucleobases. Further development of technologies that would allow routine application of cryo-EM to protein-free RNAs would have a large impact on RNA structural biology and for development of small molecule drugs targeting RNA.
[0093] One approach for solving RNA structures via cryo-EM is to use a scaffold to which an RNA motif of interest is appended. The goal is that the favorable biophysical properties that result in high resolution for the scaffold would propagate to the target RNA. The larger mass of the overall RNA also facilitates particle picking and alignment for 3D reconstruction. This approach has been attempted with the Tetrahymena group I intron, in multiple adaptations.5,6This group I intron, in isolation, has been solved to about 3 A resolution.2However, to date, this approach has resulted in only about 5 A resolution for the target RNA5’6, which does not allow for the discrimination of individual nucleotides and makes precise modeling impossible (FIG. 16). For example, when the group I intron was used as ascaffold to determine the structures of the fluoride riboswitch and the Zika virus xrRNA, modeling of these RNA structures required the use of prior high-resolution crystal structures. At this resolution, it was not possible to perform de novo modeling (zoomed-in insets in FIG. 16), as noted previously by Langeberg et al. (2023). The scaffold approach has also been attempted using RNA origami7and the ribosome8, with similar results. In summary, current scaffolding approaches have not yet demonstrated high resolution structure determination of attached target RNAs to allow discrimination of individual nucleobases. Further development of technologies that would allow routine application of cryo-EM to protein-free RNAs would have a large impact on RNA structural biology and for development of small molecule drugs targeting RNA.
[0094] X-ray crystallography can be applied for the structure determination of RNA, but has limitations. It is challenging to form crystals with protein-free RNAs due to the relatively homogeneous negatively charged surface, making it difficult to form unique crystal contacts. RNA conformational dynamics are also not fully assessed in a crystal structure because the constraining crystal lattice packs macromolecules together at concentrations approaching hundreds of mg / ml, suppressing helical dynamics and favoring compact states. For example, in structural studies of group II introns, branch-site helix dynamics were characterized using crystallography and observed only small scale movements in the domain VI helix of a group IIB intron9. Subsequently, the cryo-EM structure of a homologous group IIB intron was determined, with a shared secondary structure, and found that the branch-site helix engages in a 90° swinging action between the two steps of splicing.10These outcomes emphasize that cryo-EM is better suited to providing insight into RNA dynamics.
[0095] It is believed that similar conformational dynamics could be observed for riboswitches if cryo-EM were feasible for these small RNAs. Riboswitches are RNA structures found in mRNAs that bind small molecule ligands to affect gene expression by altering either transcription or translation of a downstream gene. This regulation occurs through the sequestering or releasing of a regulatory RNA element.11Most riboswitches have two domains, an aptamer domain which binds the regulating ligand and forms an RNA motif with a specific higher-order tertiary structure, and an expression domain whose formation, or not, regulates either transcription or translation of the downstream gene. The riboswitch mechanism implies there is an equilibrium between different conformations, with the presence or absence of ligand shifting this equilibrium to govern the resulting effects upon gene expression.12'14However, crystal structures of most riboswitches typically show minordifferences between the bound and apo states.15'20As in the case of the group II intron, this similarly is likely due to the constraining environment of the crystal lattice.
[0096] In addition to riboswitches, there are large classes of other noncoding RNAs (ncRNAs) that have long been known to have important biological roles in both prokaryotes and eukaryotes.21These bacterial ncRNAs are predicted to have ordered RNA structures based on phylogenetically conserved secondary structures.21One recent example is the raiA non-coding RNA motif that is found in over 2500 bacterial species.22RaiA from Clostridium acetobutylicum is 210 nucleotides in length and more than 25% of its nucleotides exhibit 97% conservation, strongly indicating that the RNA plays an important biological role.22RaiA also contains highly-conserved nucleotides in single stranded regions whose conservation cannot be explained by the predicted secondary structure.22Knockouts of the gene encoding raiA results in defects in both bacterial sporulation and biofilm aggregation22The raiA ncRNA is also the fourth most abundant RNA when bacterial cells transition from exponential growth to the stationary phase and the resulting spores in raiA mutants exhibit 10% viability compared to wild-type (WT).22The cumulative evidence suggests that raiA has a conserved structure that is essential for its biological function in vivo. In contrast to riboswitches and catalytic introns, ncRNAs like raiA remain poorly characterized in terms of their 3D structure and function. Previous attempts at structural characterization of these ncRNAs were largely limited to chemical probing to confirm their secondary structures.23In the absence of known functional roles for ncRNAs, it would be useful to determine the high- resolution structure from which the first hints of possible function might be gleaned.
[0097] Here, a group II intron scaffold is provided that allows for cryo-EM structure determination of attached small RNAs at high resolution. This approach has been tested on RNA targets of both known and unknown structure. For the known structure candidate, cryo- EM structures of the 86-nucleotide (nt) thiamine pyrophosphate (TPP) riboswitch were determined, in both ligand-bound and ligand-free (apo) states. The ligand-binding pocket of the TPP riboswitch was visualized at 2.5 A resolution, enabling the precise modeling of the thiamine pyrophosphate ligand. Comparisons of the ligand-bound and apo structures reveal conformational dynamics that inform a mechanism for riboswitch function. For the target of unknown structure, the cryo-EM structure of the bacterial non-coding raiA at 3.0 A was determined, with the core at 2.5 A, allowing for de novo modeling of the entire RNA, as there are no homologues in the Protein Data Bank (PDB). With these advances, it is now possible to determine the structure of an ncRNA in the beginning stages of discovery to inform thedesign of subsequent biochemical and biological experiments, and to visualize bound ligands. Going forward, this technology allows structural determination at sufficiently high resolution to enable structure-based SAR for small molecules drugs that bind to RNA.
[0098] Results
[0099] Identification of a high-resolution RNA scaffold
[0100] An ideal RNA scaffold should have the following properties: 1) exhibit a minimal grid orientation preference, 2) have high solubility, 3) have a structure resistant to denaturation, 4) have a molecular mass >100 kDa, and 5) contain a bridging region that can accommodate target RNAs. An RNA with these properties should ideally result in cryo-EM density of 3 A or better resolution.
[0101] Multiple group II introns were screened for their suitability as scaffolds for cryo- EM. Group II introns are self-splicing catalytic RNAs that have six distinct domains and typically range in size from 400 to 850 nucleotides. The attachment was tested for a group IIC intron from Oceanobacillus iheyensis (O.i.)24(120 kDa) to a larger (scaffolding) group IIB intron ribonucleoprotein (RNP) complex from Thermosynecoccus elongatus (T.e )w(300 kDa) to form the T.el.-O.i. fusion construct. When local refinement was performed on the scaffold portion of the density, a 3.5 A map of the RNP was recovered (FIG. 6A). However, when local refinement was shifted to density corresponding to the embedded IIC intron, the resolution of the target RNA was limited to about 5 A. This resolution limit is consistent with prior experiences for scaffold-based approaches.5'8It was then observed that the overall fold of the embedded RNA was maintained even though the method resulted in only a moderate resolution map (FIG. 6B). This low resolution likely reflects that the RNP scaffold displays severe orientation bias, which is still present after employing a tilted data collection strategy at 30° (FIG. 6C). The ability of the group IIC intron was then tested in isolation to function as a scaffold. The O.i. intron (FIG. 7A) folds into a stable structure that crystallizes in a wide variety of precipitants.24However, this particular intron has the unwanted tendency to undergo intramolecular hydrolytic cleavage at multiple sites,25which may also cleave an attached RNA target. An inactive mutant was then used in which an active site residue is mutated to abolish catalytic activity.26A reconstruction at 2.4 A resolution was obtained for this group II intron (FIG. 7B, FIG. 7C, and Table 1). In addition, this intron exhibited a very favorable lack of orientation preference (FIG. 7D) and has a stable solvent-exposed stem that could be used to embed an RNA of interest (domain III). The orientationdistribution map is evenly populated, which is highly unusual for a protein-free RNA in Applicant’s experience. Based on these favorable properties, this group IIC intron was selected as a scaffold candidate for the attachment of target RNAs for high resolution cryo- EM.
[0102] Table 1. Cryo-EM data collection and refinement statistics
[0103] Construct design to attach target to the scaffold
[0104] It is necessary to attach the target RNA to the scaffold via a rigid helix to reduce flexibility and facilitate high-resolution reconstruction. The attachment site on the scaffold should have the following properties: 1) be solvent exposed so that the target RNA does not fold back onto the scaffold to disrupt the structure; 2) allow sequence changes and insertions that do not affect folding of the scaffold; and 3) exhibit limited flexibility. The domain III stem of the group II intron satisfies these requirements (FIG. 7A and FIG. 7B). In addition to the scaffold, there are also requirements for the target RNA to be amenable to attachment to the group II intron. The target RNA must contain a stem that can be modified / mutated without affecting its biochemical activity. This target stem is used to create a continuous helix with the attachment site on the scaffold. In some cases, a circular permutation of thetarget can enable use of a stem located deep within the RNA sequence. Ultimately, the target RNA is fused to the scaffold to form a single RNA that is then synthesized using in vitro transcription.
[0105] The fusion construct is directly purified from the in vitro transcription reaction using diafiltration. Before diafiltration, the transcription reaction is first treated with DNase to degrade the plasmid DNA template and then with proteinase K to cleave T7 RNA polymerase and DNase. This solution is then subjected to buffer exchange (molecular weight cut-off of 100 kDa). For some RNAs, this procedure results in the formation of hydrogels due to the high concentration gradient that builds up on the membrane. In these cases, size exclusion chromatography with a Superdex™ 200 column can be used to purify the transcription reaction without a concentration step. In contrast, denaturing purification requires the development of refolding protocols that can be difficult for larger RNAs, whereas native purification takes advantage of folding during transcription.
[0106] TPP riboswitch as a known RNA target
[0107] The aptamer domain of the thiamine pyrophosphate (TPP) riboswitch14(86-nt) was attached to the group II intron scaffold (( . / .-TPP). The TPP riboswitch binds to thiamine pyrophosphate and modulates gene expression, with a wide phylogenetic distribution in both prokaryotes and eukaryotes.14Multiple crystal structures of this riboswitch in the presence and absence of thiamine pyrophosphate revealed little difference in the overall RNA fold.27'29A crystal structure of a TPP riboswitch, as a crystallographic dimer, was recently solved in the absence of the ligand.30However, in this study, the density did not match the small-angle x-ray scattering reconstruction performed in solution, again emphasizing that crystal packing interferes with the sampling of conformations found natively in solution.
[0108] High resolution reconstruction of the TPP riboswitch
[0109] High-resolution cryo-EM typically could not be applied to the TPP riboswitch, given its small size of 86 nucleotides (27.5 kDa). The TPP riboswitch was covalently attached to the solvent exposed domain III stem of the O.i. group IIC intron through a rigid helix (FIG. 8A). SHAPE-MaP31confirmed that the secondary structure of the 493 -nt linked group II intron and riboswitch structure maintained the expected fold for both RNA domains (FIG. 8 A and FIG. 17). SHAPE-MaP data also confirmed binding of the TPP ligand to the embedded riboswitch and revealed expected local conformational changes in the riboswitch32. These ligand-induced changes included modest rearrangement of the P2 and P3helices, and reduced reactivities in the P3-L5 region and in the J3-2 and J2-4 elements. Cryo- EM data collection and subsequent processing were performed for the scaffold attached to the riboswitch (FIG. 18). Clear additional signal, corresponding to the attached riboswitch, was visible in all 2D class averages as compared to the 2D classes of the scaffold alone (FIG.8B) Local refinement, focused on the scaffold, yielded a 3D reconstruction with a resolution of 2.4 A for the core of the intron (FIG. 8C). A comparison of the resulting maps and models for the scaffold alone and with the TPP riboswitch embedded shows that the attachment process did not affect the overall fold of the scaffold (FIG. 19). Local refinement on the riboswitch component was then performed to yield a reconstruction with a global resolution of 3.1 A and 2.5 A for the ligand-binding pocket (FIG. 8D). The viewing distribution plot shows that embedding the riboswitch did not create a preferred orientation for the overall fusion construct (FIG. 8E). This resolution yielded clear density for the functional groups in this small molecule and enabled modeling of the TPP ligand with high confidence. Density for three magnesium ions, coordinated to the pyrophosphate moiety, are clearly visible (FIG. 9A and FIG. 9B). Only two metal ions were visualized binding to the pyrophosphate in previous crystal structures.28,29These magnesium ions are essential for binding of TPP to this riboswitch. The cryo-EM structure of this ligand-bound riboswitch exhibits the classic closed conformation observed in the crystal structures (FIG. 20).
[0110] In contrast, focused refinement of the ligand-free riboswitch revealed an open “Y” conformation at about 6 A resolution (FIG. 10). In the ligand-free (apo) structure, the thiamine-sensing and pyrophosphate-sensing helices form an approximate 90° angle. In contrast, upon ligand binding, the RNA adopts a compact tertiary structure in which these two stems are parallel to each other and the P3 and L5 motifs form a long-range tertiary interaction, stabilizing the binding pocket for TPP. The apo state is also more dynamic due to the absence of the ligand and absence of stabilizing tertiary interactions. The thiamine sensing-stem adopts an alternate stable conformation in the apo state with two major differences: the J2-4 element and the pyrophosphate-sensing stem are highly disordered, and the thiamine-sensing stem adopts a distinct conformation in which there is clear density for an additional helical groove that extends this helix compared to the bound state (FIG. 10A). These substantial conformational differences, visualized directly here, support a molecular model that rationalizes changes in the riboswitch structure as visualized by in-solution SHAPE probing of the ligand-free state (FIG. 17 and ref. 28).33This apo structure is thesame construct used for structure determination of the bound state, emphasizing that the ligand-free state is competent to bind TPP when scaffolded.[oni] The dynamic nature of the unbound riboswitch explains why this state has not previously been captured using x-ray crystallography. The resolution of the apo state is limited due to the inherent dynamics of this destabilized RNA, yielding fewer particles going into the 3D reconstruction. Fewer particles populating the holes compared to the bound state were observed. It was anticipated that collecting a significantly larger dataset would improve the resolution. Critically, the ligand-bound and apo structures support modeling the transition through which the riboswitch modulates gene expression upon ligand binding. In this model, the opening of the riboswitch allows it to interact with the downstream regulatory element and affect either translation or transcription through an alteration in RNA structure.12,13This study provides the first cryo-EM data showing large-scale conformational changes in a riboswitch upon ligand binding, and which supports a specific mechanistic model for changes in gene expression.
[0112] RaiA RNA as a target of unknown structure
[0113] The raiA RNA from Clostridium acetobutylicum was attached via its Pl helix to domain III of the group II intron scaffold. RaiA is larger than the TPP riboswitch and initially a small dataset of only 1,328 movies was collected, corresponding to about 3 hours of data collection on a Titan Krios G3i cryo-EM microscope with a Gatan K3 detector (UC Berkley Cryo-EM Facility). This small dataset, containing about 34,517 particles, was processed to obtain a reconstruction with an overall resolution of 4 A, with density approaching 3.5 A in the core of the RNA (FIG. 22). This initial map showed clear separation of individual nucleobases and was of sufficient quality to determine the sequence register, which allowed us to build a preliminary model of the RNA structure de novo. Subsequently 16,183 movies were collected, resulting in a reconstruction with an overall resolution of 3.0 A, and 2.5 A in the core (FIG. 22). This larger dataset improved the resolution and map quality to the point that clear delineation of the geometry of the backbone in regions with high distortion and assignment of accurate base geometries for the large number of non-canonical base pairs present in the raiA RNA was achieved.
[0114] Overall structure of raiA
[0115] The map quality allows for building a complete model for raiA and therefore define the overall architecture of this highly conserved bacterial ncRNA. Density is visible forevery single nucleotide of the raiA structure. RaiA contains 8 conserved Watson-Crick helices (P1-P8) separated by internal loops (I) and junctions (J) (FIG. 11 A). These internal loops and junctions consist of a large number of highly conserved non-canonical base pairs, which position the Watson-Crick helices (FIG. 11 A). These conserved Watson-Crick and non-canonical structural motifs yield a highly compacted RNA (FIG. 1 IB) with two pseudoknots (pk-1 and pk-2) at its core (FIG. 12). The P3 stem loop and J4-lc contain the nucleotides that form one half of pk-1 and pk-2 respectively. The other halves of both pk-1 and pk-2 reside in the junction between P5 and P6, which is referred to as the pk-Loop (pk- L).
[0116] For its size, raiA is one of the most complex RNA structures observed to date. The Pl and P3 helices have an intricate geometry that is organized by several internal loops and junctions that assist in positioning pk-1 in the core of raiA. The Pl stem terminates in a three-way junction formed by Jlc-3a and J4-lc. This junction forms a discontinuous GNRA tetraloop (G20-A21-A22-A183) with a trans-U19:hA23 closing base pair that results in a bend of the helical axis of nearly 180 degrees (FIG. 13A and FIG. 23). This extremely tight turn over a relatively short sequence allows the raiA RNA to maintain a very compact structure. The Pl and P3 stems then form a long-range tertiary interaction between Jlb-lc and I3a-3b. The I3a-3b internal loop contains five adenosines (A28 and A43-A46) (FIG. 24) that organize the majority of this structural feature. A trans-H:W pair between A28:A45 forms in the middle of I3a-3b with A43 base stacking underneath this non-canonical base pair. This motif extrudes both A44 and A45 which interact directly with lib- 1 c (FIG. 24). A trans-W:S pair forms between A44:G190 with A46 and G191 base stacking on both sides of this conserved non-canonical base pair. This motif stabilizes I3a-3b in a geometry that creates an approximately 90° bend between the P3a and P3b helices (FIG. 13B and FIG. 24) ultimately pointing the P3 stem loop in the correct geometry to form pk-1 (FIG. 12A).
[0117] The P4 and P5 stems from a single continuous helix through conserved non- canonical pairs at the base of P5 (FIG. 14A). These non-canonical pairs lead to the extrusion of A57 and A58 into the minor groove of pk-1 forming an A-minor motif (FIG. 14B and FIG. 14C). Immediately following P5 is the start of the pk-Loop which contains the first half of pk-2 and the second half of pk-1 (FIG. 11A). The pk-Loop then leads into the P6 stem, which consists of two variants. Variant l is a short stem loop and is found in 32% of the identified raiA sequences while variant 2 consists of a larger multi-stem domain (P6, P7, and P8) (FIG. 15A) in the remaining 68% of sequences. The raiA motif from Clostridiumacetobutylicum used in this study belongs to the variant 2 class and contains a conserved E- Loop (FIG. 11 A) in P8 that is found in 20% of identified sequences. The P8 stem loop in this construct contains 19 base pairs with 13 being non-Watson / Crick and extends parallel to the P4 / P5 stem loop. The most highly conserved region within P8 is the E-Loop motif, which further stabilizes pk-1 in the core of the RNA through a ribose zipper interaction (FIG. 15B).
[0118] Finally, pk-2 is formed between the pk-Loop and J4-lc. Two inner shell coordinated magnesium ions are observed stabilizing the nucleotides C80 and U82 that contribute the first half of pk-2 (FIG. 12B). Nucleotides G179 and Al 84 form a trans-S:W base pair that creates a tetraloop containing the second half of pk-2 (C181 / A182). The pk-1 and pk-2 pseudoknots are separated by a single nucleotide (G84) which base stacks with the first nucleotide (A33) of the P3 loop (FIG. 12C). Taken together, raiA forms a compact structure through short-range and long-range interactions between highly conserved nucleotides. SHAPE-MaP results in low SHAPE reactivities across the majority of the raiA RNA sequence indicating that most nucleotides are highly constrained (FIG. 25). A lack of nucleotide flexibility, even in single-stranded regions, supports the structural stability observed in the cryo-EM data.
[0119] Cryo-EM Data Processing of Scaffolded RNAs
[0120] The processing workflow for scaffolded RNAs requires special considerations. As a first step, refinement of the entire assembly is performed using CryoSPARC34to assess the quality of the cryo-EM data. At this stage, the scaffold serves as an internal control and should have a resolution of about 2.5 to 3 A. The overall fold of the group II intron should be maintained and density emanating from domain III for the target should be visible at a low map threshold. The signal corresponding to the scaffold is then subtracted from the particle images. In the case of smaller targets, such as the 86-nt TPP riboswitch, focused refinement is performed following signal subtraction. A focused refinement strategy compensates for the lower signal from the low-mass riboswitch, maintains alignment information from the initial global refinement, and results in a higher resolution reconstruction. In contrast, larger RNAs, like the 210-nt raiA molecule, do not require focused refinement because there is sufficient signal for subsequent non-uniform refinement.
[0121] Experimental Discussion
[0122] Technologies for high-resolution structure determination of small RNAs, through fusion to a group II intron scaffold have been developed. This scaffold strategy is the first todemonstrate cryo-EM structure determination of target RNAs to better than 3 A resolution, and yields two landmark results. For the TPP riboswitch, the resolution observed here is the highest ever achieved by cryo-EM for an RNA less than 100 nucleotides. For the raiA motif, an RNA with no structural homolog, Applicant’s scaffolding study was characterized by rapid progress from initial selection of the RNA target to determining its structure at sub-3 A resolution over a period of about 3 weeks.
[0123] The TPP riboswitch undergoes large-scale conformational changes when comparing the ligand-free to the bound state. The ligand-free state of the riboswitch reveals an initial Y- shaped fold that, when bound to ligand, closes into a compact fold. In addition to this long- range movement, the thiamine sensing stem in the ligand-free state has a more elongated helix with the appearance of an additional helical groove (FIG. 10A) compared with the bound state. This is consistent with the rearrangement of the secondary structure in the base of this stem to form additional pairing interactions, which agrees with SHAPE probing data (FIG. 17C). It is likely that a rearrangement in secondary structure is required for the interaction with the pyrophosphate group upon ligand binding.
[0124] The large conformational difference between the bound and apo forms supports a mechanism for riboswitch function, consistent with previous hypotheses.12'14The native TPP riboswitch features a downstream regulatory element consisting of a stem loop containing the start codon of the mRNA.14In this model, the ligand-free riboswitch initially exists in a relaxed, dynamic Y-shaped form and interacts with the regulatory stem loop to render the start codon accessible for translation initiation14FIG. 10B. Once the riboswitch binds the ligand, the riboswitch clamps down and forms a compact structure that disengages the riboswitch from the regulatory stem loop, creating new structures that sequester the start codon and inhibit translation. Large conformational changes, of similar magnitude upon ligand binding, are likely a general mechanism for riboswitch function, but have been difficult to visualize directly.
[0125] The core of the raiA RNA has a resolution of 2.6 A, which allows de novo modeling of the RNA structure and to distinguish between purines and pyrimidines. It was previously hypothesized that raiA is a ribozyme and not a riboswitch due to the lack of a downstream expression platform as observed with riboswitches.22The highly ordered nature of the raiA structure suggests that it is possible that this could be a novel class of structured RNA, including functioning as a ribozyme. RaiA is highly expressed in bacteria and a dominantcellular RNA, suggesting it plays an important role in bacteria.22CRISPR-based knockouts of the entire raiA RNA gene revealed its importance in sporulation and biofilm formation.22The high resolution structure of raiA RNA now allows detailed CRISPR-based mutagenesis of individual nucleotides to assess the structure with in vivo biological function. This approach flips the typical paradigm in RNA biochemistry, which usually involves extensive biochemical characterization before attempting structure determination. High resolution structure determination can now be carried out as part of the initial characterization of newly discovered RNA motifs. Knowledge of the structure could also allow for the detection of additional raiA -1 i ke motifs in the genomes of many more bacterial species, beyond the 2,500 seen to date.
[0126] Based on its importance in biofilm and spore formation, raiA may represent an antibiotic target. Given its high degree of conservation across many bacterial species, raiA appears to be a better drug target than most riboswitches. Multiple tertiary contacts and pockets have been identified that might serve as ligand binding sites to possibly disrupt its function in vivo. The scaffolding strategy enabled structure determination of raiA from only 1,200 micrographs, supporting efficient screening of both drug targets and ligands.
[0127] This work shows that it is possible to obtain high resolution structures of small RNAs and novel uncharacterized ncRNAs in the 86 to 210 nucleotide size range using cryo- EM. There should be no lower limit to RNA size using this approach. The upper limit may be about 400-600 nucleotides using this approach. The likely upper limit will occur when an attached RNA target overwhelms the favorable properties of the group II intron scaffold. The ability to visualize routinely a small molecule bound to RNAs with complex structures35using cryo-EM now creates a broad opportunity to use this approach in drug discovery efforts targeting RNA structures.
[0128] Materials and Methods
[0129] Plasmid cloning
[0130] The mutant O.i, O.i.-TPP, and T.el4h-O.i. genes were synthesized (Genscript) and cloned into a pUC57 vector using the EcoRV restriction site. The cloned plasmids were transformed into DH5a cells. The O.i. and O.i.-TPP genes contained a 4-nt 5' exon and a 3- nt 3' exon followed by a BamHI cut site. The T.el4h-O.i. gene contained a 19-nt 5' exon and a 9-nt 3' exon followed by a BamHI cut site. The T.el.4h maturase expression plasmid usedin this study was previously describedlO. The O.i.-raiA gene was synthesized by Genewiz using their PriorityGENE service and inserted into their pUC-GW-Amp vector. The O.i.- raiA gene contained a 4-nt 5' exon and a 3-nt 3 'exon followed by a BamHI cut site. The scaffold of the O.i.-raiA gene was modified from the one used in the TPP study. The following modifications were made: nucleotides 349 and 390-412 were deleted, the GAAA tetraloop (275-278) was mutated to a UUCG, U347 was mutated to an A, and A348 was mutated to a U. FIG. 7A and FIG. 21 have sequences according to SEQ ID NOs: .
[0131] In vitro RNA transcription and purification
[0132] The O.i, O.i-TPP, T.el4h-O.i., and O.i.-raiA plasmids were linearized using an engineered BamHI restriction site (NEB). 50 pg of template DNA was added to a total volume of 1 mL of in vitro transcription buffer (50 mM Tris-HCl pH 7.5, 25 mM MgC12, 5 mM DTT, 2 mM spermidine, 0.05% Triton X-100, and 5 mM of each NTP). T7 polymerase and thermophilic inorganic phosphatase was added to begin RNA synthesis. The reaction mixture was incubated at 37°C for 3 hrs. CaC12 was added to a final concentration of 1.2 mM along with Turbo DNase and placed at 37°C for 1 hour to fully digest the DNA template. Proteinase K was subsequently added and incubated at 37°C for an additional hour. The resulting solution was centrifuged to remove any precipitate and then filtered through a 0.2 pm filter. The filtered solution was buffer exchanged a total of 7 times, each time using 14 mL of filtration buffer (5 mM Na-cacodylate pH 6.5 and 4 mM MgC12) and a 100 kDa molecular weight cut-off filter. After the final buffer exchange step, the RNA was concentrated to approximately 10 mg / mL for use in downstream cryo-EM experiments.
[0133] SHAPE-MaP RNA structure probing
[0134] SHAPE-MaP assays were performed essentially as described previously36. Briefly, 10 pmol of either the TPP riboswitch or raiA RNA was diluted in 19 pL probing buffer (300 mM HEPES, pH 8.0, 2 mM MgC12, 40 mM NaCl). In the case of the ligand-bound TPP riboswitch, 10 pmol RNA was mixed with 1 uL of TPP (final concentration, 50 pM). After incubation for 15 min at 37 °C, 9 uL of RNA-ligand solution was added to 1 uL IM 2A337and mixed thoroughly. After incubation for 20 min, the reaction was quenched by addition of 5 uL of IM DTT. No ligand reactions were conducted in parallel. RNA was purified (G-25 spin column) and subjected to reverse transcription (SuperScript II, Invitrogen) using a heatcool protocol [25°C for 10 min, 42°C for 90 min, 10 cycles of (50°C for 2 min, 42°C for 2 min), 70°C for 10 min], DNA libraries were prepared from cDNA using a two-step PCRprocess (Q5 Hot Start High-Fidelity DNA Polymerase, New England Biolabs). PCR 1 used Step-1 Forward primer and Step-1 Reverse primer, and performed for 20 cycles; PCR 2 used Universal Forward primer and Universal Reverse prime, and performed for 10 cycles. All cDNA or dsDNA were purified (Mag-Bind TotalPure NGS beads, Omega), quantified (Qubit dsDNA Quantification Assay Kits, Life Technology), and visualized to confirm integrity (2100 Bioanalyzer Instrument, Agilent). The libraries were sequenced on a MiSeq system (Illumina). Mutations were aligned and parsed using ShapeMapper (v2.15)38using default parameters.
[0135] RT primer sequence: 5'-CTAATAGAGT AGAGCGAACT CCTCTC-3' (SEQ ID NO: ).
[0136] Step-1 Forward primer: 5'-CCCTACACGA CGCTCTTCCG ATCTNNNNNT TATGTGTGCC CGGCATG-3' (SEQ ID NO: ). Step-1 Reverse primer: 5'-GACTGGAGTT CAGACGTGTG CTCTTCCGAT CTNNNNNCTA ATAGAGTAGA GCGAACTCCT-3 (SEQ ID NO: )'.
[0137] raiA RT primer: 5’-ACACCACCGTACGTACC-3’ (SEQ ID NO: ). raiA Step-1 Forward primer: 5’- GACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNCTTTCGAGCTCAGAAGTC AG -3’ (SEQ ID NO: ). raiA Step-1 Reverse primer: 5’- CCCTACACGACGCTCTTCCGATCTNNNNN GTTCAAGCATGGCGC-3’ (SEQ ID NO: )•
[0138] Universal Forward primer: 5'-AATGATACGG CGACCACCGA GATCTACACT CTTTCCCTAC ACGACGCTCT TCCG-3' (SEQ ID NO: ). Universal Reverse primer: 5'- CAAGCAGAAG ACGGCATACG AGAT (SEQ ID NO: ) (8nt-barcode) GTGACTGGAG TTCAGAC-3' (SEQ ID NO: ).
[0139] Group IIB intron RNP assembly
[0140] The RNP used in this study was assembled and purified as previously described10.
[0141] EM sample preparation
[0142] For all grids prepared with TPP bound, 100 uL of 2 mg / mL RNA was incubated with 1 mM thiamine pyrophosphate at room temperature for 15 minutes prior to freezing. For all grids prepared with the riboswitch in the apo state as well as the O.i.-raiA motif, this binding step was omitted. To freeze the grids, 3.5 pL of freshly prepared RNA sample at 2.0mg / mL was applied to a glow discharged (40 mBar, 15 mA for 30 s using a PELCO easiGlow) copper R1.2 / 1.3 300-mesh grid (Quantifoil). The grid was blotted with a filter paper (Whatman No.l) at 4 °C in a cold room before plunging frozen into liquid ethane / propane (37.5 / 62.5) mix using a manual plunger. For the T.el4h-O.i RNP specimen, 3.5 pL of freshly prepared RNP at 1 mg / mL was applied to a glow discharged (40 mBar, 15 mA for 120 s using a PELCO easiGlow) UltrAuFoil Rl.2 / 1.3 300-mesh grid (Quantifoil). The grid was blotted with a filter paper (Whatman No.1) at 4 °C in a cold room before plunging frozen into liquid ethane / propane (37.5 / 62.5) mix using a manual plunger.
[0143] Cryo-EM data collection and processing
[0144] Movies were collected on the same Titan Krios microscope (Thermo Fisher) operating at 300 keV, equipped with a K3 Gatan direct electron detector, and were collected with a magnification of 105,000x for a physical pixel size of 0.811 Angstroms and camera operating in super-resolution mode (0.4055 A / pixel). O.i.-raiA movies were collected with a magnification of 105,000x for a physical pixel size of 0.848 A and camera operating in superresolution mode (0.424 Angstroms / pixel). Movies collected for O.i., O.z.-TPP, O. / .-TPP-Apo and O.i.-raiA were collected as 50 dose-fractionated frames with a dose rate of 8 e- / pixel / s for a total exposure time of 4.5 s and total dose of about 50 e- / A2. The O.i. and O.z.-TPP, O.i.- TPP-Apo datasets were collected with a defocus range of -0.5 pm to -1.5 pm, whereas the O.i.-raiA dataset was collected with a defocus range of -0.8 to -2.0 um. The T.el.-O.i. dataset was collected with the same parameters described previously (Ref) as reported in table SI, with the exception of a defocus range of -0.8 pm to -1.9 pm, and an exposure time of 5.4 s. All movies were semi-automatically collected using SerialEM.395724 (O.z.), 30,314 (O.i.- TPP), 15,876 (O.z.-TPP-Apo), 2596 (T.el.-O.i.), 1328 O.i.-raiA Small), and 16,183 (O.i.- raiA) movies were collected with these parameters.
[0145] For the O.i. dataset, 5,724 movies were collected and imported into Relion 3.1.40'42Super resolution movies were binned by two and motion-corrected using MotionCor2.43These motion-corrected micrographs were imported into CryoSPARC 3.3.2,34and CTF was estimated using Patch CTF. Micrographs were visually inspected and 343 micrographs were thrown out due to poor particle distribution and ice contamination. Particles were picked on the remaining 5,381 micrographs using a general model from Topaz 0.2.544as implemented in CryoSPARC. These about 1.6 million particles were extracted with a box size of 288 pixels downsampled to 64 pixels, for a final pixel size of 3.6 A / pixel. Particles were ranthrough 2D classification, and a small subset of particles was selected for ab initio model building. This model was then used for four subsequent rounds of 2-class heterogeneous refinement to classify the about 1.6 million particles into a 3D class that looked like the group II intron structure. The resulting 725,736 particles were extracted as 320 pixel boxed particles downsampled to 160 pixels for a pixel size of 1.6 A / pixel. The particles were then subjected to several additional rounds of 3D classification, refinement, and CTF refinement, before the final set of 607,776 particles were extracted with a box size of 360 pixels downsampled to 240 pixels, for a final pixel size of 1.22 A / pixel. These approximately 600,000 particles were put through a single round of 3D classification with alignment (562,162 particles selected), 2D classification (458,691 particles selected), and 9-class heterogeneous refinement (382,753 particles selected) for non-uniform refinement45This map refined to a final GSFSC resolution of 2.62 A as reported in cryoSPARC.
[0146] For the T.el.-O.i. dataset, 2596 movies were collected at a 30° stage tilt and imported into cryoSPARC 4.2.1. Super resolution movies were binned by two and motion- corrected and CTF was estimated using Patch Motion Correction and PatchCTF, respectively. 1,775 motion corrected micrographs were selected for further processing with good particle distribution and little ice contamination. 100 micrographs were randomly selected for initial particle picking using blob picker, and these approximately 60,000 particles were extracted for 2D classification. Good 2D references corresponding to approximately 22,000 particles were selected for ab initio model building and refinement. The 3D reference from this dataset was used for picking approximately 633,000 particles on the full 1,775 micrographs, 502,500 of which were then extracted at a box size of 512 pixels down sampled to 128 pixels, for a final pixel size of 3.2 A / pixel. The 3D volume generated by the initial refinement was used as the starting template for three rounds of consecutive 2-class heterogeneous refinement, and the resulting 331,439 particles were put through refinement and extracted with a box size of 512 pixels (0.811 A / pixel). The 322,683 particles were subjected to duplicate removal (20 Angstrom separation), another round of 2-class heterogeneous refinement, refinement, CTF refinement, and 2D classification, for a final set of approximately 283,000 particles. The GSFSC (0.143) global resolution of this particle set after non-uniform refinement was 2.88 A as reported by cryoSPARC. These particles were then subjected to hetero-refinement and 3D classification to select particles where the embedded O.i. insert was fully intact. The 34,822 particles selected from classification were used for masked local refinement around the O.i. insert, which had a GSFSC resolution of 4.02 A as reported in CryoSPARC.
[0147] For the O.z.-TPP-Apo dataset, 15876 super resolution movies were collected and imported into cryoSPARC 4.2.1. Movies were binned by two, and motion-corrected and CTF estimated using patch motion correction and PatchCTF, respectively. Exposures with CTF resolution fits above 10 A, relatively thick ice, and high values for the predicted stage tilt angle, were removed, leaving 11,485 micrographs for further processing. Due to difficulties with initial blob picking, a 1,800 micrograph subset was used for template picking using an O.i. map processed from a previous dataset. These approximately 800,000 particles were used for two rounds of 3-class hetero-refinement, with the final 93,000 particles used to generate a new template for particle picking against the full approximately 11,000 micrograph dataset. The approximately 8,000,000 particles picked were extracted with a 336 pixel box downsampled to 84 pixels (3.2 A / pixel). These particles were subjected to four consecutive rounds of hetero-refinement (4 classes, 2-class, 2-class, and 2-class). The resulting 1,212,955 particles were then refined and re-extracted to a box of 128 pixels (2.3 A / pixel). Duplicate particles (50 angstrom radius) were removed, and remaining particles were used for a 5-class hetero-refinement. The 1,056,211 particles selected were refined, and extracted with a 256 pixel box (1.27 A / pixel), and then put through another round of refinement and CTF refinement. These 1,032,440 particles were then ran through a 10-class hetero-refinement, and 2 classes that corresponded to the highest resolution densities were selected. These approximately 600,000 particles were then subjected to multiple rounds of refinement and CTF refinement to improve resolution, with one final 2D classification job used to remove particles with poor 2D classes. These 547,101 particles were re-extracted with a box size of 448 pixels downsampled to 288 pixels (1.26 A / pixels) and non-uniform refined to a final global GSFSC resolution of 2.78 A as reported by cryoSPARC
[0148] For the apo-TPP embedded structure, those about 547,000 particles were put through multiple rounds of 3D classification and hetero-refinement to remove conformations for the P2 and P3 stems that did not have resolvable densities, or corresponded to conformations that were not relevant to the TPP riboswitch. Classes that had intact density for an open conformation were pooled together, for a final particle set of 106,285 particles. In parallel, a mask around the O.i. scaffold was generated for local refinement of the scaffold, and these 3D coordinates were used for signal subtraction of the scaffold from the particle images. Then, the 3D coordinates of the approximately 106,000 particles after non-uniform refinement were used as a starting point for local refinement of the apo-TPP riboswitch in its open conformation. Gaussian priors of an 8 degree rotation and 4 A shift were used formasked local refinement of the apo state, which had a final GSFSC of 4.84 A as reported in CryoSPARC.
[0149] The basic workflow for the high resolution reconstruction of the TPP-riboswitch is outlined in FIG. 18. In brief, three separate data collections of the O.z.-TPP construct were collected (7,763 movies, 6303 movies, 16248 movies) with identical parameters, and pooled together as 30,314 super resolution movies. These movies were initially imported into cryoSPARC 4.2.1 for validation and a first reconstruction, and used a workflow similar to that implemented for the O.z.-TPP-apo construct. The first maps from cryoSPARC had a resolution close to 3.7 A / pix, and so the super resolution movies were imported into Relion 4.0 as separate exposure groups, and put through a standard workflow of particle picking, 3D classification, refinement, and particle polishing at the end. These particles were then exported into cryoSPARC, where they underwent the process of resolution improvement through non-uniform refinement and particle defocus refinement for a final GSFSC resolution of 2.53 A. Masked refinement of the scaffold was then used for signal subtraction from the polished particles, and the signal subtracted particles were recentered on the density for the TPP riboswitch. This TPP riboswitch density was locally refined with a Gaussian prior of a 3° rotation and 3 A shift, with a final reported GSFSC resolution of 2.96 A in CryoSPARC
[0150] The small O.i.-raiA dataset (1328 movies) was imported into cryoSPARC 4.5.1 using cryoSPARC Live, where super resolution movies were binned to 0.848 A / pixel and CTF was estimated using Patch Motion Correction and Patch CTF, respectively. 118 exposures were removed due to low resolution CTF fits, and 200,662 particles were then picked on these 1210 micrographs using a Topaz 0.2.5 general model. 171,146 particles were then extracted with a 108 pixel box (3.39 A / pixel) and subjected to a single round of 2D classification, ab initio model generation, and heterogeneous refinement (4 classes), with 34,293 particles being the final result. 33,096 particles were extracted (256 pixels 1.70 A / pixel) with a GSFSC resolution of 3.63 A after non-uniform refinement. A mask was then generated around the raiA insert, and this region of density had a final resolution of approximately 4.7 A. 2507 particles from 47 exposures were used to train a new Topaz model to pick O.i.-raiA particles on these micrographs.
[0151] Using the Topaz model, 263,878 particles were picked and then extracted from 1,194 micrographs (108 pixels, 3.4 A / pixel). These particles were subjected to a single roundof 2D classification to remove junk particles, and then two rounds of heterogeneous refinement (6 classes) to pick the highest resolution classes to pool together. These particles were refined, re-extracted (324 pixels, 1.34 A / pixel), and then subjected to non-uniform refinement with CTF defocus correction, for a final GSFSC resolution of 3.4 A. After attempting particle subtraction, heterogeneous refinement, 2D classification, 3D classification, and local refinement, the resolution and reconstruction of the raiA insert did not improve. After 2D classification to remove junk particles, 82,698 particles were re- centered on the raiA insert and re-extracted (192 pixels, 0.848 A / pixel). A homogenous reconstruction of the raiA density was then used as a reference for a single round of heterogeneous refinement (4-class), and density corresponding to an intact reconstruction of the raiA insert was used as selection criteria for a single 34,517 particle class, with the other three classes being representative of the O.i. scaffold or poor realignment. The particle sets utility was used to select these 34,517 particles from the original 324 pixel downsampled particles, and the raiA insert density was locally refined with a Gaussian prior of a 10° rotation and 10 A shift, for a final GSFSC resolution of 4.03 A before sharpening.
[0152] The large O.i. -raiA dataset processing is outlined in FIG. 22. In brief, 16,183 movies were imported into cryoSPARC 4.5.1 and motion-corrected and CTF estimated using cryoSPARC Live. After removing 3,749 movies, particles were picked using a general Topaz model, and extracted from 12,405 exposures for a total of 2,362,623 particles. These particles were subjected to three rounds of consecutive heterogeneous refinement (8 classes) using the global O.i. -raiA density from the smaller dataset as the initial reference volume. The approximately 1.3 million particles were then refined and re-extracted (216 pixels, 1.70 A / pixel) and subjected to another three rounds of heterogeneous refinement (5-classes) using the global reconstruction downsampled to five different resolutions. Particles assigned to classes corresponding to full intact O.i.-raiA were then refined, and re-extracted (512 pixels, 0.848 A / pixel) a final time. These particles were passed through non-uniform refinement and CTF correction for a final global resolution of 2.99 A.
[0153] After refinement, signal corresponding to O.i. was subtracted from the particle images, and then the particles for raiA were locally refined to a resolution of 3.0 A. The subtracted particle images were then re-centered on and cropped to a box size of 384 pixels around the raiA insert. These particles were then subjected to several rounds of heterogeneous refinement to remove 3D classes with poor density for the raiA insert, and thefinal set of 308,340 particles were put through non-uniform refinement for a final reconstruction with a GSFSC resolution of 3.04 A.
[0154] All data needed to evaluate the conclusions are available in the main text or the supplementary materials. Structure coordinates have been deposited in the Protein Data Bank (PDB) under accession numbers 9C6I (O.z.), 9C6J (O.z.-TPP focused on O.i. , 9C6K (O.i.- TPP focused on TPP) and 9CXF (O.i.-raiA). Cryo-EM density maps have been deposited in the Electron Microscopy Data Bank (EMDB) under accession numbers 4524745248(O.z.-TPP focused on O.i. 45249 (O.z.-TPP focused on TPP), 45250 (O.z.-TPP apo focused on TPP), 45251 (T.el.-O.i. focused on O.z. ), 45994 (O.i.-raiA small data set), and 45988 (O.i.- raiA large data set).
[0155] Model building and structure refinement
[0156] As a starting point for the model building of O.i., 4DS6 structure coordinates were refined in real space by PHENIX46,47The same process was repeated for the locally refined O.z.-TPP focused on the scaffold. For the locally refined map focused on the TPP riboswitch, 2GDI structure coordinates were real space refined by PHENIX. Any nucleotide that need to be remodeled was done in COOT48,49using the RCrane plugin.50,51For the raiA motif structure, no starting model was used. The RNA was built de novo in COOT using the secondary structure from Soares, L.W. et al. The final model was then real space refined in PHENIX. UCSF Chimera was used to make figures depicting the cryo-EM density.52All software was compiled by SBGrid.53
[0157] Quantification and Statistical Analysis
[0158] RNA concentrations were determined using a Nanodrop spectrophotometer (Thermo-Fisher). Maturase protein concentrations were determined using an SDS-PAGE gel with a titration of BSA (Thermo-Fisher). To calculate per-reside Full RMSD values, the two models being compared were superposed in COOT using LSQ using the entire sequence for alignment. The superposed models were opened in UCSF Chimera for evaluation using Match -> Align. All map and model validation and statistics were done in PHENIX (Table SI).
[0159] All publications, patent applications, issued patents, and other documents referred to in this specification are herein incorporated by reference as if each individual publication, patent application, issued patent, or other document was specifically and individually indicated to be incorporated by reference in its entirety. Definitions that are contained in textincorporated by reference are excluded to the extent that they contradict definitions in this disclosure.Partial Sequence Listing
[0160] Fig. 1 sequence (SEQ ID NO: )GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCC — Insert RNA--—GGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGTGAGAGGAGTTCGCTCTACTCTATTAG
[0161] Fig. 2 sequence-Scaffold with non-covalent attachment module (SEQ ID NO: )GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGagctttcgagctCAGAAGTCAGCAGAAGTCATAGTACCCGGGAGAAGGGTAGTTCCGGGGAAACTTGGTTCTACCCCACGCTCCTGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCATGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGT
[0162] Fig. 2 sequence-RNA of interest with non-covalent attachment module(SEQ ID NO: )RNA of interest 5’ end —GGGATATGGAAGTTCCGGGGGAACTTGGTTCTTCCTAAGTCCT — RNA of interest 3 ’end
[0163] Fig. 3 sequence-Scaffold 3’ end labeled with RNA of interest (SEQ IDNO: )GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTGTTCGCAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGT ACGGTGGTGT — RNA of Interest
[0164] Fig. 4 Sequence-Circular permutation of Scaffold 3’ with RNA of interest(SEQ ID NO: )RNA of Interest 5 ’end —GGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGAAACAAACAAATAAACTAAATTATGTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCC— RNA of Interest 3’ end
[0165] Fig. 7 sequence-Scaffold alone (SEQ ID NO: )GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTGTTCGCAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGTGAGAGGAGTTCGCTCTACTCTATTAG
[0166] Fig. 8 sequence-DIII plus TPP (SEQ ID NO: )AGAAGTCATAGTACCCTCAGTACTCGGGGTGCCCTTCTGCGTGAAGGCTGAGAAATACCCGTATCACCTGATCTGGATAATGCCAGCGTAGGGAAGTGCTGAGGGGAAGGACGGAA
[0167] Fig. 11 sequence-raiA (SEQ ID NO: )TTAAGTTAGGTTTGTGGTTGAAAGTCGATGCCAGTCGCAGGCAAAACGATCCACGTAAGTTAAACAAAGTTTTAATGAGCATGGTGCGGCTTAGAAGTAAGTCCTGCCGCTTTAGGCGAGAGTATTAGTAGTGAGAGGGTAATTCCGGGTAGCGAAACTTCCAG CAGGCGAGTGTGGGGTCAAAGACCAGGTCAACTAACTTAA
[0168] Fig. 17 sequence-Scaffold plus TPP (SEQ ID NO: )TTATGTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTCAGTACTCGGGGTGCCCTTCTGCGTGAAGGCTGAGAAATACCCGTATCACCTGATCTGGATAATGCCAGCGTAGGGAAGTGCTGAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGTGAGAGGAGTTCGCTCTAC TCTATTAG
[0169] Fig. 21 sequence-Scaffold plus raiA (SEQ ID NO: )GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGagctttcgagctCAGAAGTCAGCAGAAGTCATAGTACCCTTAAGTTAGGTTTGTGGTTGAAAGTCGATGCCAGTCGCAGGCAAAACGATCCACGTAAGTTAAACAAAGTTTTAATGAGCATGGTGCGGCTTAGAAGTAAGTCCTGCCGCTTTAGGCGAGAGTATTAGTAGTGAGAGGGTAATTCCGGGTAGCGAAACTTCCAGCAGGCGAGTGTGGGGTCAAAGACCAGGTCAACTAACTTAAGGGGAAGGACGGAACAAGTATGGCGTTCGC GCCATGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGT
[0170] Fig. 25 sequence-raiA shape data (SEQ ID NO: )TTAAGTTAGGTTTGTGGTTGAAAGTCGATGCCAGTCGCAGGCAAAACGATCCACG TAAGTTAAACAAAGTTTTAATGAGCATGGTGCGGCTTAGAAGTAAGTCCTGCCGC TTTAGGCGAGAGTATTAGTAGTGAGAGGGTAATTCCGGGTAGCGAAACTTCCAG CAGGCGAGTGTGGGGTCAAAGACCAGGTCAACTAACTTAAEmbodiments
[0171] 1. A method for RNA imaging using an RNA scaffold, the method comprising, consisting essentially of, or yet further consisting of imaging a target RNA and a RNA scaffold using cryo-electron microscopy.
[0172] 2. A method for RNA imaging using an RNA scaffold, the method comprising, consisting essentially of, or yet further consisting of: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: a stem loop of the target RNA comprises a Tecto RNA 2 stem loop; the RNA scaffold comprises a Tecto RNA 1 stem loop; and the Tecto RNA 1 stem loop and the Tecto RNA 2 stem loop bind to each other.
[0173] 3. A method for RNA imaging using an RNA scaffold, the method comprising, consisting essentially of, or yet further consisting of: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the target RNA is bound to a first end of a linker sequence; and the RNA scaffold is bound to a second end of a linker sequence.
[0174] 4. The method of embodiment 3, wherein the linker is a PolyA linker.
[0175] 5. A method for RNA imaging using an RNA scaffold, the method comprising, consisting essentially of, or yet further consisting of: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the RNA scaffold is circularly permutated and is embedded in the target RNA; and the scaffold is embedded in a stem of the target RNA.
[0176] 6. The method of any one of embodiments 1-5, wherein the RNA scaffold is a group II intron.
[0177] 7. The method of any one of embodiments 1-5, wherein the RNA scaffold is derived from a group II intron.
[0178] 8. The method of any one of embodiments 6-7, wherein the group II intron is a group IIC intron.
[0179] 9. The method of any one of embodiments 6-8, wherein the group II intronRNA scaffold is derived from at least one bacteria.
[0180] 10. The method of embodiment 9, wherein the at least one bacteria isOceanobacillus iheyensis or Thermosynecoccus elongatus.
[0181] 11. The method of any one of embodiments 1-10, wherein the target RNA is a riboswitch or non-coding RNA.
[0182] 12. The method of any one of embodiments 1-11, wherein the target RNA is bound to a ligand.
[0183] 13. The method of any one of embodiments 1-12, wherein the target RNA is about 25nt, 50nt, about 75nt, about lOOnt, about 125nt, about 150nt, about 175nt, about 200nt, about 250nt, about 250nt, about 275nt, or about 300nt in length.
[0184] 14. The method of any one of embodiments 1-13, wherein the group II intronRNA scaffold is about 80ka, about lOOkDa, about 120kDa, about 140kDa, about 160kDa, about 180kDa, about 200kDa, about 220kDa, about 240kDa, about 260kDa, about 280kDa, about 300kDa, about 320kDa, about 340kDa, about 360kDa, about 380kDa, about 400kDa, about 420kDa, about 440kDa, about 460kDa, or about 480kDa.
[0185] 15. The method of any one of embodiments 1-14, wherein the group II intronRNA scaffold is about 80kDa to about 140kDa.
[0186] 16. The method of any one of embodiments 1-15, wherein the target RNA has a Y-shaped, compact, P-shaped, or candy conformation.
[0187] 17. An image taken using the method of any one of embodiments 1-16.
[0188] 18. The image of embodiment 17, wherein the resolution of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A.
[0189] 19. The image of embodiment 17, wherein the resolution of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, or about 3.0 A.
[0190] 20. The image of embodiment 17, wherein the resolution of a core of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A.
[0191] 21. The image of embodiment 17, wherein the resolution of a core of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, or about 3.0 A.
[0192] 22. A construct comprising, consisting essentially of, or yet further consisting of a group II intron RNA scaffold and a target RNA.
[0193] 23. The construct of embodiment 22, further comprising a ligand bound to the target RNA.
[0194] 24. The construct of embodiment 22 or 23, wherein the intron is a group IIC intron.
[0195] 25. The construct of any one of embodiments 22-24, wherein the group II intron RNA scaffold is derived from at least one bacteria.
[0196] 26. The construct of embodiment 25, wherein the at least one bacteria isOceanobacillus iheyensis or Thermosynecoccus elongatus .
[0197] 27. The construct of any one of embodiments 22-26, wherein the target RNA attaches to the group II intron RNA scaffold at the group II intron domain III stem.
[0198] 28. The construct of any one of embodiments 22-27, wherein the group II intron is circularly permuted.
[0199] 29. The construct of any one of embodiments 22-28, wherein the target RNA is a riboswitch or non-coding RNA.
[0200] 30. The construct of any one of claims of embodiments 22-29, wherein the construct has a structure according to:GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCA GAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGG AAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGT ATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACA GAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGT CTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTXiAGGGGAAGGACGGAACA AGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGT GGTGT; wherein:Xi is the target RNA.
[0201] 31. A construct of any one of claims of embodiments 22-29, wherein the construct has a structure according to:1)GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGagctttcgagctCAGAAGTCAGCAGAAGT CATAGTACCCGGGAGAAGGGTAGTTCCGGGGAAACTTGGTTCTACCCCAC GCTCCTGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCATGCTTGAACC ACCGTATACCGAACGGTACGTACGGTGGTGT; and2) X1GGGATATGGAAGTTCCGGGGGAACTTGGTTCTTCCTAAGTCCTX2; wherein:Xi and X2 are fragments of the target RNA; and1) and 2) are non-covalently bound to each other.
[0202] 32. A construct comprising of any one of claims of embodiments 22-29, wherein the construct has a structure according to:GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACA GAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTAGGGGAAGGACGGAACAA GTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGTGAA(A)„Xi; wherein: n is an integer that is from 0 to 40; andXi is the target RNA.
[0203] 33. A construct comprising of any one of claims of embodiments 22-29, wherein the construct has a structure according to:XiAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCACCGTATACCGAACGGTACGTACGGTGGTGTAA(A)„GTGTGCCCGGCATGGGTGCAGTCTA TAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAA GGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTX2; wherein: n is an integer from 0 to 40; andXi and X2 are fragments of the target RNA.
[0204] 34. A method for RNA imaging using an RNA scaffold, the method comprising imaging the construct of any one of embodiments 17-33 using cryo-electron microscopy.Equivalents
[0205] While certain embodiments have been illustrated and described, it should be understood that changes and modifications can be made therein in accordance with ordinary skill in the art without departing from the technology in its broader aspects as defined in the following claims.
[0206] The embodiments, illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms “comprising,” “including,” “containing,” etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the claimed technology. Additionally, the phrase “consisting essentially of’ will be understood to include those elements specifically recited and those additional elements that do not materially affect the basic and novel characteristics of the claimed technology. The phrase “consisting of’ excludes any element not specified.
[0207] The present disclosure is not to be limited in terms of the particular embodiments described in this application. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art.Functionally equivalent methods and compositions within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims. The present disclosure is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled. It is to be understood that this disclosure is not limited to particular methods, reagents, compounds, compositions, or biological systems, which can of course vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0208] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0209] As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member.
[0210] Other embodiments are set forth in the following claims.References:1 Larsen, K. P. et al. Architecture of an HIV-1 reverse transcriptase initiation complex. Nature s, 118-122 (2018). DOI: 10.1038 / s41586-018-0055-92 Su, Z. et al. Cryo-EM structures of full-length Tetrahymena ribozyme at 3.1 A resolution. Nature 596, 603-607 (2021). DOI: 10.1038 / s41586-021-03803-w3 Kappel, K. et al. Accelerated cryo-EM-guided determination of three-dimensional RNA- only structures. Nat Methods 17, 699-707 (2020). DOI: 10.1038 / s41592-020-0878-94 Bonilla, S. L., Sherlock, M. E., MacFadden, A. & Kieft, J. S. A viral RNA hijacks host machinery using dynamic conformational changes of a tRNA-like structure. Science 374, 955-960 (2021). DOI: 10.1126 / science.abe85265 Liu, D., Thelot, F. A., Piccirilli, J. A., Liao, M. & Yin, P. Sub-3-A cryo-EM structure of RNA enabled by engineered homomeric self-assembly. Nat Methods 19, 576-585 (2022). DOI: 10.1038 / s41592-022-01455-w6 Langeberg, C. J. & Kieft, J. S. A generalizable scaffold-based approach for structure determination of RNAs by cryo-EM. Nucleic Acids Res 51, elOO (2023). DOI: 10.1093 / nar / gkad784Sampedro Vallina, N., McRae, E. K. S., Hansen, B. K., Boussebayle, A. & Andersen, E. S. RNA origami scaffolds facilitate cryo-EM characterization of a Broccoli-Pepper aptamer FRET pair. Nucleic Acids Res 51, 4613-4624 (2023). DOI: 10.1093 / nar / gkad224 Zhang, C. et al. Analysis of discrete local variability and structural covariance in macromolecular assemblies using Cryo-EM and focused classification. Ultramicroscopy 203, 170-180 (2019). DOI: 10.1016 / j.ultramic.2018.11.016 Robart, A. R., Chan, R. T., Peters, J. K., Rajashankar, K. R. & Toor, N. Crystal structure of a eukaryotic group II intron lariat. Nature 514, 193-197 (2014). DOI: 10.1038 / naturel3790 Haack, D. B. et al. Cryo-EM Structures of a Group II Intron Reverse Splicing into DNA. Cell 178, 612-623. e612 (2019). DOI: 10.1016 / j .cell.2019.06.035 Kavita, K. & Breaker, R. R. Discovering riboswitches: the past and the future. Trends Biochem Sci 48, 119-141 (2023). DOI: 10.1016 / j .tibs.2022.08.009 Serganov, A. & Nudler, E. A decade of riboswitches. Cell 152, 17-24 (2013). DOI: 10.1016 / j. cell.2012.12.024 Olenginski, L. T., Spradlin, S. F. & Batey, R. T. Flipping the script: Understanding riboswitches from an alternative perspective. J Biol Chem 300, 105730 (2024). DOI: 10.1016 / j jbc.2024.105730 Winkler, W ., Nahvi, A. & Breaker, R. R. Thiamine derivatives bind messenger RNAs directly to regulate bacterial gene expression. Nature 419, 952-956 (2002). DOI: 10.1038 / natureOl 145 Huang, L., Serganov, A. & Patel, D. J. Structural insights into ligand recognition by a sensing domain of the cooperative glycine riboswitch. Mol Cell 40, 774-786 (2010). DOI: 10.1016 / j .molcel.2010.11.026 Jenkins, J. L., Krucinska, J., McCarty, R. M., Bandarian, V. & Wedekind, J. E. Comparison of a preQi riboswitch aptamer in metabolite-bound and free states with implications for gene regulation. J Biol Chem 286, 24626-24637 (2011). DOI: 10.1074 / jbc.Ml 11.230375Schroeder, G. M. et al. Analysis of a preQi -I riboswitch in effector-free and bound states reveals a metabolite-programmed nucleobase-stacking spine that controls gene regulation. Nucleic Acids Res 48, 8146-8164 (2020). DOI: 10.1093 / nar / gkaa546Stagno, J. R., Bhandari, Y. R., Conrad, C. E., Liu, Y. & Wang, Y. X. Real-time crystallographic studies of the adenine riboswitch using an X-ray free-electron laser.FEBSJ2M, 3374-3380 (2017). DOI: 10.1111 / febs.14110Stagno, J. R. et al. Structures of riboswitch RNA reaction states by mix-and-inject XFEL serial crystallography. Nature 541, 242-246 (2017). DOI: 10.1038 / nature20599Stoddard, C. D. et al. Free state conformational sampling of the SAM-I riboswitch aptamer domain. Structure 18, 787-797 (2010). DOI: 10.1016 / j.str.2010.04.006Narunsky, A. et al. The discovery of novel noncoding RNAs in 50 bacterial genomes.Nucleic Acids Res 52, 5152-5165 (2024). DOI: 10.1093 / nar / gkae248Soares, L. W., King, C. G., Fernando, C. M., Roth, A. & Breaker, R. R. Genetic disruption of the bacterial raiA motif noncoding RNA causes defects in sporulation and aggregation. Proc Natl Acad Sci USA 121, e2318008121 (2024). DOI: 10.1073 / pnas.2318008121Somarowthu, S. et al. HOTAIR forms an intricate and modular secondary structure. Mol Cell 58, 353-361 (2015). DOI: 10.1016 / j.molcel.2015.03.006Toor, N., Keating, K. S., Taylor, S. D. & Pyle, A. M. in Science Vol. 320 77-82 (2008).Toor, N. et al. in RNA Vol. 16 57-69 (2010).Chan, R. T., Robart, A. R., Rajashankar, K. R., Pyle, A. M. & Toor, N. Crystal structure of a group II intron in the pre-catalytic state. Nat Struct Mol Biol 19, 555-557 (2012). DOI: 10.1038 / nsmb.2270Edwards, T. E. & Ferre-D'Amare, A. R. Crystal structures of the thi-box riboswitch bound to thiamine pyrophosphate analogs reveal adaptive RNA-small molecule recognition. Structure 14, 1459-1468 (2006). DOI: 10.1016 / j.str.2006.07.008Serganov, A., Polonskaia, A., Phan, A. T., Breaker, R. R. & Patel, D. J. Structural basis for gene regulation by a thiamine pyrophosphate-sensing riboswitch. Nature 441, 1167- 1171 (2006). DOI: 10.1038 / nature04740Thore, S., Leibundgut, M. & Ban, N. Structure of the eukaryotic thiamine pyrophosphate riboswitch with its regulatory ligand. Science 312, 1208-1211 (2006). DOI:10.1126 / science.1128451 Lee, H. K. el al. Crystal structure of Escherichia coli thiamine pyrophosphate-sensing riboswitch in the apo state. Structure 31, 848-859. e843 (2023). DOI: 10.1016 / j.str.2023.05.003 Weeks, K. M. SHAPE Directed Discovery of New Functions in Large RNAs. Acc Chem Res 54, 2502-2517 (2021). DOI: 10.1021 / acs.accounts.lc00118 Warner, K. D. et al. Validating fragment-based drug discovery for biological RNAs: lead fragments bind and remodel the TPP riboswitch specifically. Chem Biol 21, 591-595 (2014). DOI: 10.1016 / j.chembiol.2014.03.007 Steen, K. A., Rice, G. M. & Weeks, K. M. Fingerprinting noncanonical and tertiary RNA structures by differential SHAPE reactivity. J Am Chem Soc 134, 13160-13163 (2012). DOI: 10.1021 / ja304027m Punjani, A., Rubinstein, J. L., Fleet, D. J. & Brubaker, M. A. cryoSPARC: algorithms for rapid unsupervised cryo-EM structure determination. Nat Methods 14, 290-296 (2017). DOI: 10.1038 / nmeth.4169 Warner, K. D., Hajdin, C. E. & Weeks, K. M. Principles for targeting RNA with druglike small molecules. Nat Rev Drug Discov 17, 547-558 (2018). DOI: 10.1038 / nrd.2018.93 Smola, M. J., Rice, G. M., Busan, S., Siegfried, N. A. & Weeks, K. M. Selective 2'- hydroxyl acylation analyzed by primer extension and mutational profiling (SHAPE- MaP) for direct, versatile and accurate RNA structure analysis. NatProtoc 10, 1643- 1669 (2015). DOI: 10.1038 / nprot.2015.103 Marinus, T., Fessler, A. B., Ogle, C. A. & Incamato, D. A novel SHAPE reagent enables the analysis of RNA structure in living cells with unprecedented accuracy. Nucleic Acids Res 49, e34 (2021). DOI: 10.1093 / nar / gkaal255 Busan, S. & Weeks, K. M. Accurate detection of chemical modifications in RNA by mutational profiling (MaP) with ShapeMapper 2. RNA 24, 143-148 (2018). DOI: 10.1261 / rna.061945.117Mastronarde, D. N. Automated electron microscope tomography using robust prediction of specimen movements. J Struct Biol 152, 36-51 (2005). DOI: 10.1016 / j.jsb.2005.07.007Kimanius, D., Forsberg, B. O., Scheres, S. & Lindahl, E. Accelerated cryo-EM structure determination with parallelisation using GPUs in RELION-2. bioRxiv (2016).Scheres, S. H. RELION: implementation of a Bayesian approach to cryo-EM structure determination. J Struct Biol 180, 519-530 (2012). DOI: 10.1016 / j .j sb.2012.09.006Zivanov, J. et al. New tools for automated high-resolution cryo-EM structure determination in RELION-3. Elife 7 (2018). DOI: 10.7554 / eLife.42166Zheng, S. Q. et al. MotionCor2: anisotropic correction of beam-induced motion for improved cryo-electron microscopy. Nat Methods 14, 331-332 (2017). DOI: 10.1038 / nmeth.4193Bepler, T. et al. Positive-unlabeled convolutional neural networks for particle picking in cryo-electron micrographs. Nat Methods 16, 1153-1160 (2019). DOI: 10.1038 / s41592- 019-0575-8Punjani, A., Zhang, H. & Fleet, D. J. Non-uniform refinement: adaptive regularization improves single-particle cryo-EM reconstruction. Nat Methods 17, 1214-1221 (2020). DOI: 10.1038 / s41592-020-00990-8Adams, P. D. et al. PHENIX: a comprehensive Python-based system for macromolecular structure solution. Acta Crystallogr D Biol Crystallogr 66, 213-221 (2010).Afonine, P. V. et al. Real -space refinement in PHENIX for cryo-EM and crystallography. Acta Crystallogr D Struct Biol 74, 531-544 (2018). DOI: 10.1107 / S2059798318006551Emsley, P. & Cowtan, K. Coot: model-building tools for molecular graphics. Acta Crystallogr D Biol Crystallogr 60, 2126-2132 (2004). DOI: 10.1107 / s0907444904019158Emsley, P., Lohkamp, B., Scott, W. G. & Cowtan, K. Features and development of Coot. Acta Crystallogr D Biol Crystallogr 66, 486-501 (2010). DOI: S0907444910007493 [pii]10.1107 / S0907444910007493Keating, K. S. & Pyle, A. M. Semiautomated model building for RNA crystallography using a directed rotameric approach. Proc Natl Acad Sci U S A l 07, 8177-8182 (2010). DOI: 0911888107 [pii]10.1073 / pnas.0911888107 Keating, K. S. & Pyle, A. M. RCrane: semi-automated RNA model building. Acta Crystallogr D Biol Crystallogr 68, 985-995 (2012). DOI: 10.1107 / s0907444912018549 Pettersen, E. F. et al CCSF Chimera— a visualization system for exploratory research and analysis. J Comput Chem 25, 1605-1612 (2004). DOI: 10.1002 / jcc.20084 Morin, A. et al. Collaboration gets the most out of software. Elz e 2, e01456 (2013). DOI: 10.7554 / eLife.01456 Leontis, N. B. & Westhof, E. Geometric nomenclature and classification of RNA base pairs. RH4 7, 499-512 (2001). DOI: 10.1017 / sl 355838201002515
Claims
WHAT IS CLAIMED IS:
1. A method for RNA imaging using an RNA scaffold, the method comprising imaging a target RNA and a RNA scaffold using cryo-electron microscopy.
2. A method for RNA imaging using an RNA scaffold, the method comprising: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: a stem loop of the target RNA comprises a Tecto RNA 2 stem loop; the RNA scaffold comprises a Tecto RNA 1 stem loop; and the Tecto RNA 1 stem loop and the Tecto RNA 2 stem loop bind to each other.
3. A method for RNA imaging using an RNA scaffold, the method comprising: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the target RNA is bound to a first end of a linker sequence; and the RNA scaffold is bound to a second end of a linker sequence.
4. The method of claim 3, wherein the linker is a PolyA linker.
5. A method for RNA imaging using an RNA scaffold, the method comprising: imaging a target RNA and a RNA scaffold using cryo-electron microscopy; wherein: the RNA scaffold is circularly permutated and is embedded in the target RNA; and the scaffold is embedded in a stem of the target RNA.
6. The method of any one of claims 1-5, wherein the RNA scaffold is a group II intron.
7. The method of any one of claims 1-5, wherein the RNA scaffold is derived from a group II intron.
8. The method of any one of claims 6-7, wherein the group II intron is a group IIC intron.
9. The method of any one of claims 6-8, wherein the group II intron RNA scaffold is derived from at least one bacteria.
10. The method of claim 9, wherein the at least one bacteria is Oceanobacillus iheyensis or Thermosynecoccus elongatus.
11. The method of any one of claims 1-10, wherein the target RNA is a riboswitch or noncoding RNA.
12. The method of any one of claims 1-11, wherein the target RNA is bound to a ligand.
13. The method of any one of claims 1-12, wherein the target RNA is about 25nt, 50nt, about75nt, about lOOnt, about 125nt, about 150nt, about 175nt, about 200nt, about 250nt, about 250nt, about 275nt, or about 300nt in length.
14. The method of any one of claims 1-13, wherein the group II intron RNA scaffold is about80ka, about lOOkDa, about 120kDa, about 140kDa, about 160kDa, about 180kDa, about 200kDa, about 220kDa, about 240kDa, about 260kDa, about 280kDa, about 300kDa, about 320kDa, about 340kDa, about 360kDa, about 380kDa, about 400kDa, about 420kDa, about 440kDa, about 460kDa, or about 480kDa.
15. The method of any one of claims 1-14, wherein the group II intron RNA scaffold is about80kDa to about 140kDa.
16. The method of any one of claims 1-15, wherein the target RNA has a Y-shaped, compact, P-shaped, or candy conformation.
17. An image taken using the method of any one of claims 1-16.
18. The image of claim 17, wherein the resolution of the target RNA is about 2.0 A, about2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A.
19. The image of claim 17, wherein the resolution of the target RNA is about 2.0 A, about2.25 A about 2.5 A, about 2.75 A, or about 3.0 A.
20. The image of claim 17, wherein the resolution of a core of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, about 3.0 A, about 3.25 A about 3.5 A, about 3.75 A, about 4.0 A, about 4.25 A about 4.5 A, about 4.75 A, or about 5.0 A.
21. The image of claim 17, wherein the resolution of a core of the target RNA is about 2.0 A, about 2.25 A about 2.5 A, about 2.75 A, or about 3.0 A.
22. A construct comprising a group II intron RNA scaffold and a target RNA.
23. The construct of claim 22, further comprising a ligand bound to the target RNA.
24. The construct of claim 22 or 23, wherein the intron is a group IIC intron.
25. The construct of any one of claims 22-24, wherein the group II intron RNA scaffold is derived from at least one bacteria.
26. The construct of claim 25, wherein the at least one bacteria is Oceanobacillus iheyensis or Thermosynecoccus elongatus .
27. The construct of any one of claims 22-26, wherein the target RNA attaches to the groupII intron RNA scaffold at the group II intron domain III stem.
28. The construct of any one of claims 1-6, wherein the group II intron is circularly permuted.
29. The construct of any one of claims 1-7, wherein the target RNA is a riboswitch or noncoding RNA.
30. The construct of any one of claims of claims 22-29, wherein the construct has a structure according to:GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAA GGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGA ATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCA TATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTT TCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAG ACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCAT AGTACCCTXiAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCT TGAACCACCGTATACCGAACGGTACGTACGGTGGTGT; wherein:Xi is the target RNA.
31. The construct of any one of claims of claims 22-29, wherein the construct has a structure according to:1)GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCA TATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTT TCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAG ACTGCATTCTTACCCGGGGAGGTCTGagctttcgagctCAGAAGTCAGCAGAAGT CATAGTACCCGGGAGAAGGGTAGTTCCGGGGAAACTTGGTTCTACCCCAC GCTCCTGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCATGCTTGAACC ACCGTATACCGAACGGTACGTACGGTGGTGT; and2) X1GGGATATGGAAGTTCCGGGGGAACTTGGTTCTTCCTAAGTCCTX2; wherein:Xi and X2 are fragments of the target RNA; and1) and 2) are non-covalently bound to each other.
32. The construct of any one of claims of claims 22-29, wherein the construct has a structure according to:GTGTGCCCGGCATGGGTGCAGTCTATAGGGTGAGAGTCCCGAACTGTG AAGGCAGAAGTAACAGTTAGCCTAACGCAAGGGTGTCCGTGGCGACATG GAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTT CATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCC TTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAA GACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCA TAGTACCCTAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTT GAACCACCGTATACCGAACGGTACGTACGGTGGTGTGAA(A)„Xi; wherein: n is an integer that is from 0 to 40; andXi is the target RNA.
33. The construct of any one of claims of claims 22-29, wherein the construct has a structure according to:XiAGGGGAAGGACGGAACAAGTATGGCGTTCGCGCCTAAGCTTGAACCAC CGTATACCGAACGGTACGTACGGTGGTGTAA(A)„GTGTGCCCGGCATGGGT GCAGTCTATAGGGTGAGAGTCCCGAACTGTGAAGGCAGAAGTAACAGTTA GCCTAACGCAAGGGTGTCCGTGGCGACATGGAATCTGAAGGAAGCGGACGGCAAACCTTCGGTCTGAGGAACACGAACTTCATATGAGGCTAGGTATCAATGGATGAGTTTGCATAACAAAACAAAGTCCTTTCTGCCAAAGTTGGTACAGAGTAAATGAAGCAGATTGATGAAGGGAAAGACTGCATTCTTACCCGGGGAGGTCTGGAAACAGAAGTCAGCAGAAGTCATAGTACCCTX2; wherein: n is an integer from 0 to 40; andXi and X2 are fragments of the target RNA.
34. A method for RNA imaging using an RNA scaffold, the method comprising imaging the construct of any one of claims 17-33 using cryo-electron microscopy.
Citation Information
Patent Citations
RNA nanoparticles and nanotubes
US20120263648A1
Nucleic acid frameworks for structural determination
WO2017049136A1
Sub-3 Å CRYO-em structure of RNA enabled by engineered homomeric self-assembly
WO2023018878A2