Design of Structure-Based Therapeutics Targeting RNA Hairpin Loops
Through scaffold-oriented crystallography, combining scaffold RNA and RNA of interest, the three-dimensional structure of the RNA hairpin loop is determined, which solves the problem of difficult to obtain RNA hairpin loop structure information in the prior art, and realizes the design and development of targeted RNA therapeutic agents.
Patent Information
- Application Number
- CN202080093268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-19
- Filing Date
- 2020-11-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-11-19
AI Technical Summary
The prior art is difficult to effectively obtain the three-dimensional structural information of the RNA hairpin loop, which limits the design and development of targeted RNA therapeutic agents.
A scaffold-oriented crystallography method was developed to form fusion RNA crystals by combining specific scaffold RNA with RNA engineering of interest, and the three-dimensional structure of the RNA hairpin ring was determined using X-ray or electron crystallography techniques.
This method can quickly and simply determine the three-dimensional structure of the RNA hairpin loop and identify compounds with high affinity to interact with RNA, providing important information on the design of targeted RNA therapeutic agents.
Smart Images

Figure CN115038710B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application Ser. No. 62 / 937,657, filed Nov. 19, 2019, and entitled “Design of Structure - Based Therapeutics Targeting RNA Hairpin Loops,” which is co - pending and co - assigned, and which is incorporated herein by reference in its entirety under 35 U.S.C. § 119(e).
[0003] Statement of government support
[0004] This invention was made with government support under Grant No. 1616265, awarded by the National Science Foundation of the United States. The government has certain rights in this invention. Technical field
[0005] The present invention relates to methods and materials for determining the three - dimensional structure of RNA hairpin loops. Background art
[0006] RNA molecules are crucial for the development of many diseases, such as cancer and RNA virus infections. For this reason, RNA molecules are excellent therapeutic targets. In this context, almost all RNAs form hairpin secondary structures that are essential for their function. Therefore, it is necessary to understand these structures to facilitate the identification and design of therapeutic agents targeting these molecules. However, conventional methods for examining RNA, such as RNA interference and antisense oligonucleotides, are limited and avoid strong structures. Although conventional techniques can provide some information about RNA structure, the limitations of these techniques have made RNA hairpin loops an under - explored target for therapeutic inhibitor design.
[0007] There is a pressing need in the art for new methods and materials for obtaining information about the three - dimensional structure of RNA hairpin loops. Summary of the invention
[0008] As detailed below, we have developed a new scaffold - directed crystallography method that can be used to obtain information about the three - dimensional structure of RNA hairpin loops. The RNA crystallization scaffolds and related methods disclosed herein can be used to simply and rapidly determine the three - dimensional structure of RNA hairpin loops and their association with other reagents (e.g., inhibitors). The specific scaffold RNA used in the methods of the present invention is the YdaO - type c - di - AMP riboswitch from Thermoanaerobacter pseudethanolicus, which has been found to readily form a structure with a diameter exceeding Crystals of the large cavity. As discussed in detail below, we have determined that the RNA of interest can be engineered into the P2 stem of this scaffold RNA so that the hairpin can be accommodated in the cavity. The resulting fusion RNA can then be crystallized under conditions similar or unrelated to those that crystallize the individual scaffold. The three-dimensional structure of such molecules (e.g., these molecules alone and / or in association with other reagents) can then be determined using X-ray or electron crystallography techniques, etc.
[0009] The RNA crystallization scaffolds and related methods disclosed herein can be used to identify compounds that interact with target RNA molecules with high affinity and specificity, such as natural and chemically modified oligonucleotides, as well as small molecule drugs. This is important because the interaction between such compounds and RNA hairpin loops can affect the biological activity of these molecules in ways that can modulate their activity in vivo in pathologies such as cancer and RNA virus infections. In addition, since RNA is involved in almost every aspect of biology and disease, the methods disclosed herein are widely applicable protocols that can provide information on how to specifically modulate almost any target RNA. Thus, the methods disclosed herein allow the observation and evaluation of reagents that target specific RNAs, such as oligonucleotide analogs, including those that play a role in a variety of biological processes, such as processes involved in viral replication (e.g., replication of pathogens such as Severe Acute Respiratory Syndrome Coronavirus 2, Hepatitis C, and Zika), processes involved in pathological conditions such as cancer or neurodegenerative diseases, and processes involved in the production of microRNAs that regulate protein-coding genes, etc.
[0010] The present invention disclosed herein has multiple embodiments. One embodiment of the invention is a substance composition, which comprises a ribonucleic acid having at least 90% sequence identity with the following: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 1). Generally, the polynucleotide comprises the sequence of SEQ ID NO: 1. In this composition, residues 14-17 (GAAA) of SEQ ID NO: 1 of the ribonucleic acid are replaced with a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides (the above at least 90% sequence identity does not include the heterologous segment of nucleic acid that may be inserted at residues 14-17 into the ribonucleic acid). In these compositions, the heterologous segment of nucleic acid is generally a nucleic acid segment that forms a loop structure in a naturally occurring RNA molecule. In certain embodiments of the invention, the heterologous segment of nucleic acid comprises the complete loop structure in a naturally occurring RNA molecule, and optionally 0-5 base pairs of the stem structure. Optionally, these compositions may further comprise a reagent that binds to the ribonucleic acid, such as a polynucleotide that hybridizes to the ribonucleic acid.
[0011] Another embodiment of the invention is a system or kit for observing an RNA structure comprising a plasmid, the plasmid comprising a DNA sequence encoding a ribonucleic acid having at least 90% (and optionally less than 100%) identity with the following: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 1). In certain embodiments, the plasmid further comprises a promoter for expressing or transcribing the ribonucleic acid, and / or the system or kit further comprises an RNA polymerase. Optionally, the system or kit further comprises one or more primers that hybridize to a segment of nucleic acid in the plasmid.
[0012] Another embodiment of the invention is a method for obtaining information about the structure of ribonucleic acid. The method includes replacing residues 14-17 (GAAA) of SEQ ID NO: 1 (or ribonucleic acid having at least 90% identity with SEQ ID NO: 1) with a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides to form a fusion ribonucleic acid molecule, crystallizing the fusion RNA, subjecting the fusion ribonucleic acid molecule to X-ray or electron crystallography techniques in order to observe the results (e.g., the electron density map of the X-ray or electron crystallography technique) to obtain information about the three-dimensional structure of the heterologous segment of nucleic acid. In certain embodiments of these methods, the fusion ribonucleic acid molecule is combined with a ribonucleic acid-binding reagent (e.g., a polynucleotide that hybridizes to ribonucleic acid) prior to crystallographic analysis such that the structure of the RNA / reagent complex can be observed. Typically in these methods, the crystallographic analysis includes a comparison with a control sample lacking the ribonucleic acid-binding reagent. Optionally, in these methods, prior to the X-ray or electron crystallography technique, a plurality of fusion ribonucleic acid molecules are combined with a variety of ribonucleic acid-binding reagents (e.g., in a high-throughput screen). In some embodiments of the invention, at least two reagents are combined with the fusion ribonucleic acid molecule.
[0013] In an illustrative working embodiment of the invention, we examined nine structures of the pri-miRNA hairpin loop. These studies determined that loops of lengths 4-8 nucleotides are more structured than previously thought, making these and medium-length loops excellent targets for therapeutic agents. In embodiments of the invention, the target loop need not have a specific length and can be longer or shorter than the available examples. This realization and our new method for structure determination allow artisans to identify lead oligonucleotide compounds and perform structure-based iterative rounds quickly and cost-effectively. The methods of the invention have broad applications because they target processes that are important for combating infectious diseases and cancer, age-related pathologies and neurodegenerative diseases, and genetic disorders such as DiGeorge syndrome.
[0014] Other objects, features, and advantages of the invention will become apparent to those skilled in the art from the following detailed description. However, it should be understood that the detailed description and the specific examples, while indicating some embodiments of the invention, are given by way of illustration and not limitation. Many changes and modifications may be made within the scope of the invention without departing from the spirit of the invention, and the invention includes all such modifications. Brief Description of the Drawings
[0016] A brief description of the drawings is given below.
[0017] Figure 1 A-1E. Analyze the pri-miRNA terminal loop and search for potential crystallization scaffolds. Figure 1(a): Shows the distribution of pri-miRNA apical loop lengths. Figure 1 (b): Shows the comparison of the largest spherical cavity (with radius R max ) present in each RNA crystal structure against the diffraction resolution of that structure. Crystal forms with a single molecule in the asymmetric unit are shown as green crosses, and all other crystal forms are shown as black dots. Figure 1 (c): Shows the structure of RNA. Figure 1 (d): Shows the secondary structure of the YdaO-type ci-di-AMP riboswitch. The crystal packing of the riboswitch (PDB ID 4QK8) is shown. Molecules surrounding the large central channel (parallel to the c-axis) are gray, and blue spheres with radius are placed in the channel to illustrate its size. The L2 stem-loop that terminates within the channel is green. Figure 1 (e): Shows native gel analysis of W.T. YdaO and fusions with the pri-miR-9-1 terminal loop with 0 - 3 base pairs from the stem.
[0018] Figure 2 A-2F. Atomic structures of pri-miRNA terminal loops with lengths of 8 - 6 nt determined by scaffold-directed crystallography. Throughout the figure, the last base pair of the scaffold P2 stem is gray. Figure 2 A and 2D - 2F are shown in stereo view. The inset shows the secondary structure of the loop. The 2Fo - Fc electron density map is drawn at the level contour shown in each panel. Figure 2 (a) Shows pri-miR-378a (378a + 0bp). Figure 2 (b) Shows the pri-miR-378a loop with one base pair from the stem (378a + 1bp). Figure 2 (c) Shows the 378a + 1bp structure and electron density. Figure 2 (d) Shows pri-miR-340 (340 + 1bp). Figure 2 (e) Shows pri-miR-300 (300 + 0bp). The adjacent canonical pair in pri-miR-300 is C - G, the same as the pair in the scaffold. Thus, this structure is the required 300 + 1bp. Figure 2 (f) Shows pri-miR-202 (202 + 1bp).
[0019] Figure 3 A-3D. Structures of shorter (4 - 5 nt) pri-miRNA loops. The color scheme is the same as in Figure 2 . Figure 3(a) shows pri-miR-208a (208a + 1bp). Figure 3 (b) shows pri-miR-320b-2 (320b-2 + 1bp). Figure 3 (c) shows pri-miR-449c (449c + 1bp). Figure 3 (d) shows pri-miR-19b-2 (19b-2 + 1bp).
[0020] Figure 4 A-4E. Human pri-miRNA apical junctions and loop structural consensus, non-canonical pairs, and asymmetric flexibility. Figure 4 (a) shows Figure 2 and 3 structural alignments of all eight loops shown in. Positions that align well in most or all structures are marked. Figure 4 (b) shows plots of the folding ΔG values of eight pri-miRNA apical junctions and loops measured with 50 mM NaCl. Error bars represent the standard deviation obtained from 4 - 6 replicates. Each RNA contains the apical loop and the immediately adjacent base pairs from the stem, as well as five common base pairs (see Figure 8 the RNA secondary structure in a and the detailed thermodynamic parameters in Table 2). Figure 4 (c) shows the observed and expected counts of human pri-miRNAs with indicated apical closed-loop residue pairs. Expected counts were estimated based on the abundances of 5’ and 3’ loop residues. Figure 4 (d) shows the average atomic displacement parameter (ADP) for each residue, with all loops plotted on the same scale. The 5’ and 3’ termini represent the terminal base pairs of the pri-miRNA stem-loop. Structural plots illustrating the ADP distribution are presented in Figure 10 . Figure 4 (e) shows the root mean square fluctuation (RMSF, ) for each residue determined by molecular dynamics. Symbols and colors are the same as in Figure 4 (d).
[0021] Figure 5 A-5K. Association of the DGCR8 Rhed domain with the pri-miRNA apical junction. Figure 5 (a) - 5(h) Quantification of gel shift assays, with representative gel images shown in Figure 11 . Data points represent the mean fraction bound ± standard error (SE) from three replicate experiments. Data were fit with the Hill equation, and the dissociation constant (K d )(±SE) is shown. Figure 5(i) Shows the comparison of the free energy of Rhed binding (RTln(Kd)) with the terminal loop length, as predicted by mfold. Figure 5 (j) Shows the same as Figure 5 (i), except that the loop length was adjusted with the bases involved in the excluded non-canonical pairs.
[0022] Figure 6 A-6C. Results of the systematic mutagenesis of the U-U pairs observed in several crystal structures of the pri-miRNA apical junction (the U-U pairs are one of the best-processed pri-miRNA variants). The terminal residues in the pri-miRNA apical loop fine-tune miRNA production. Figure 6 (A) Shows a schematic diagram of the dual pri-miRNA construct used to measure the miRNA maturation efficiency in mammalian cells. Each pri-miRNA fragment contains a hairpin and flanking sequences of approximately 30 nt on each side, for a total of approximately 150 nt. The pri-miR-9-1 fragment was not altered and was used for normalization. The terminal loop residues of the 3' pri-miRNA fragment were mutagenized. The abundances of the two mature miRNAs were measured using quantitative RT-PCR. Figure 6 (B) Shows the maturation efficiency (miR-340 / miR-9 ratio) of the pri-miR-340 variants. Figure 6 (C) Shows the maturation efficiency of the pri-miR-193b variants. In these scatter plots, individual data points are shown as grey dots. Bars represent the mean and standard deviation.
[0023] Figure 7 A-7I. Simulated annealing composite omit maps calculated for all pri-miRNA loops. The color scheme is the same as Figure 2 and 3 . All map contours were drawn at 1.1σ. For details on calculating individual maps, see the Methods section. Figure 7 A shows 378a + 0 bp. Figure 7 B shows 378a + 1 bp. Figure 7 C shows 340 + 1 bp. Figure 7 D shows 300 + 0 bp. Figure 7 E shows 202 + 1 bp. Figure 7 F shows 208 + 1 bp. Figure 7 G shows 449c + 1 bp. Figure 7 H shows 320b-2 + 1 bp. Figure 7 I shows 19b-2 + 1 bp.
[0024] Figure 8 A-8I. RNA constructs for melting and binding assays. Figure 8(a) shows a short RNA oligonucleotide for optical melting assay. A common 5-bp helical segment serves as the stem for all hairpins (gray base pairs). The pri-miRNA apical junction and loop nucleotides are in black. Figure 8 (b)-8(i) shows the predicted secondary structures of all pri-miRNA fragments used in the Rhed binding assay. The additional G-C pairs added to the stem base to enhance transcription are highlighted in yellow. The boxes show the sequences of the apical loop and the terminal base pairs of the stem used for determining the crystal structure.
[0025] Figure 9 A-9B. Comparison of the pri-miRNA terminal loop structures with similar RNA folds found in the PDB. Figure 9 A shows a cartoon representation of the 8-nt loop of pri-miR-378a (378a+1, left), Figure 9 B shows similar loops from different structures of RNaseP (2), gua-I riboswitch (3), and tRNA Phe (4).
[0026] Figure 10 A-10H. Estimation of the flexibility of the apical loop using atomic displacement parameters (ADP). Figure 10 (a)- Figure 10 Each structure shown in (h) is colored from the lowest ADP in blue to the highest in red. The inset shows the range of the plotted ADPs.
[0027] Figure 11 A-11H. Example gel shift assays of each pri-miRNA fragment bound to Rhed. The pre-miRNA fragments are identified above each gel, and the free RNA and protein-bound species are labeled in the gels. The concentration (μM) of the Rhed dimer used in the binding reaction is shown below the gels.
[0028] Figure 12 A-12b. Analysis of the pri-miR-223 apical loop sequencing data from a previously reported high-throughput mutagenesis and processing assay (5). Figure 12 (a) shows the predicted secondary structure of the upper region of the pri-miR-223 hairpin. The base coloring reflects the level of evolutionary conservation in the Rfam entry of this RNA (Rfam accession number: RF0064). The major miRNA product from the 3p arm is highlighted in blue. The red letters show the mutations relative to the WT sequence. Compared with the alternative secondary structure shown in the inset, this model is likely the dominant conformation as it generates an optimal upper stem length of approximately 23 bp above the Drosha cleavage site and places the evolutionarily conserved residues within the stem and the less conserved residues in the bulges. Figure 12(b) shows a heat map that displays the frequency of C-A pairs in the 9-nt pri-miR-223 loop sequencing data. The lower left matrix shows the percentage frequency in the input library, and the upper right matrix shows the frequency in the processed RNA. For reference, the wild-type loop sequence is shown along the diagonal. The C-A pair is enriched to 69% in the processed fraction, compared to 22% in the input.
[0029] Figure 13 . The NMR ensemble of pri-miR-20b shows the U-G pair at the apical junction and the stacking of the adjacent 5′ G residue (6).
[0030] Figure 14A and 14B . RNA structure. Figure 14A shows the secondary structure of the HCV cis-acting replication element; and Figure 14B shows the HCV IRES domain IIIb (see, e.g., Quade et al., Nature Communications volume 6, Article number: 7646 (2015)). DETAILED DESCRIPTION OF THE INVENTION
[0032] Many of the techniques and procedures described or cited herein are well understood by those skilled in the art and are generally used using conventional methods. In the description of the preferred embodiments, reference may be made to the accompanying drawings, which form a part thereof, and in which are shown, by way of illustration, specific embodiments in which the invention may be practiced. It should be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the invention.
[0033] Unless otherwise defined, all technical terms, symbols, and other scientific terms or terminology used herein are intended to have the meaning commonly understood by those skilled in the art to which this invention pertains. In some instances, terms with commonly understood meanings are defined herein for clarity and / or for ease of reference, and such definitions included herein are not necessarily to be construed as representing a substantial difference from what is commonly understood in the art.
[0034] Metazoan pri-miRNAs fold into characteristic hairpin structures that are recognized by the Microprocessor complex during processing. For this recognition, the apical junction connecting the hairpin stem and loop guides the DGCR8 RNA-binding heme domain (Rhed) to the apex of the hairpin. Here, we describe a scaffold-guided crystallographic approach and report structures of the apical junctions and loops of a number of human pri-miRNAs. These structures reveal a consensus in which noncanonical base pairs and at least one 5′ loop residue stack on top of the hairpin stem. The noncanonical pairs contribute to thermodynamic stability in solution. U-U and G-A pairs are highly enriched at the apical junctions of human pri-miRNAs. We also find that Rhed binds more tightly to longer loops, biochemically explaining why pri-miRNAs with shorter loops are generally poorly processed. Our disclosure provides a structural basis for understanding the related molecular mechanisms of pri-miRNA and microRNA maturation.
[0035] As discussed below, we have developed methods and materials that can be used to determine the three-dimensional structures of pri-miRNA apical junctions and loops because of their important roles in miRNA maturation and regulation (7-10). These segments are present in both pri-miRNAs and pre-miRNAs, and thus their structures influence the Drosha and Dicer cleavage steps (8). The apical junctions and loops are also targets for drug discovery (11). To date, using NMR spectroscopy, the structures of only two pri-miRNA apical stem-loops have been characterized in the ligand-free state (6, 11, 12). The 13-nt apical loop of pre-miR-20b folds into a well-defined rigid structure (6), while weak signals suggest that the 14-nt loop of pri-miR-21 is unstructured (11, 12). The human genome encodes 1,881 pri-miRNA hairpins that vary widely from each other (13). To investigate the structures of a large number of pri-miRNAs, we have developed a scaffold-directed crystallization technique that can rapidly determine hairpin loop structures without interference from the lattice. We report nine apical junction and loop structures from eight pri-miRNAs and their biochemical characterization of interactions with Rhed.
[0036] Embodiments of the invention include a substance composition comprising a ribonucleic acid having at least 90% sequence identity with the following: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 1). Embodiments of the invention preferably exhibit at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity with the polynucleotide sequence of SEQ ID NO: 1. The percent identity can be readily determined by comparing the sequence of the polynucleotide variant with the corresponding portion of the full-length polynucleotide of SEQ ID NO: 1 (wherein the above sequence identity does not include a heterologous segment of nucleic acid that can be inserted into such ribonucleic acid in place of residues 14-17). Some techniques for sequence comparison include using computer algorithms well known to those of ordinary skill in the art, such as the Align or BLAST algorithms (Altschul, J. Mol. Biol. 219: 555-565, 1991; Henikoff and Henikoff, PNAS USA 89: 10915-10919, 1992)). Default parameters can be used.
[0037] Typically, the polynucleotide comprises the sequence of SEQ ID NO: 1. In this composition, residues 14-17 of the ribonucleic acid of SEQ ID NO: 1 (GAAA) are replaced by a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides (wherein at least 90% sequence identity as described above does not include the heterologous segment of nucleic acid that may be inserted at ribonucleic acid residues 14-17). In an illustrative embodiment, the polynucleotide comprises GGUUGCCGAAUCCXGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 29), wherein X comprises 4 to 33 heterologous nucleotides selected from A, U, G, and C (e.g., those that comprise a three-dimensional structure in naturally occurring RNA molecules such as human miRNAs). In these compositions, the heterologous segment of nucleic acid is typically a heterologous segment of nucleic acid that forms a three-dimensional structure (e.g., a loop structure) in a naturally occurring RNA molecule. In certain embodiments of the invention, the heterologous segment of nucleic acid comprises a complete loop structure in a naturally occurring RNA molecule, and optionally 0-5 base pairs of a stem structure. Optionally, these compositions may further comprise a reagent that binds to the ribonucleic acid, such as a polynucleotide that hybridizes to the ribonucleic acid.
[0038] Another embodiment of the present invention is a system or kit for observing RNA structure, which comprises one or more plasmids, and the one or more plasmids comprise a DNA sequence encoding a ribonucleic acid having at least 90% (and optionally less than 100%) identity with: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 1). In some embodiments of the present invention, the one or more plasmids comprise a polynucleotide sequence having at least 90% identity with the sequence GGTTGCCGAATCC (SEQ ID NO: 27) and / or a polynucleotide sequence having at least 90% identity with the sequence GGTACGGAGGAACCGCTTTTTGGGGTTAATCTGCAGTGAAGCTGCAGTAGGGATACCTTCTGTCCCGCACCCGACAGCTAACTCCGGAGGCAATAAAGGAAGGAG (SEQ ID NO: 28). In certain embodiments, the one or more plasmids further comprise a promoter for expressing or transcribing the ribonucleic acid, and / or the system or kit further comprises an RNA polymerase. Optionally, the system or kit further comprises one or more primers that hybridize to a nucleic acid segment in the plasmid.
[0039] Another embodiment of the invention is a method for obtaining information about the structure of ribonucleic acid. The method comprises replacing residues homologous to residues 14-17 (GAAA) of SEQ ID NO: 1 (or ribonucleic acid having at least 90% identity with SEQ ID NO: 1) with a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides (e.g., 4, 5, 6, or 7 nucleotides, etc., up to a 33-nucleotide heterologous segment) to form a fusion ribonucleic acid molecule, crystallizing the fusion RNA, and performing a structural analysis on the crystallized fusion ribonucleic acid molecule, such as a technique including X-ray or electron crystallography techniques, in order to observe the results to obtain information about the three-dimensional structure of the heterologous segment of the nucleic acid. In certain embodiments of these methods, the fusion ribonucleic acid molecule is combined with a ribonucleic acid-binding reagent (e.g., a polynucleotide or other reagent that binds to the heterologous segment of the ribonucleic acid) prior to crystallographic analysis, such that the structure of the RNA / reagent complex can be observed. Typically in these methods, the crystallographic analysis includes comparison with a control sample lacking the ribonucleic acid-binding reagent. Optionally, in these methods, multiple fusion ribonucleic acid molecules are combined with multiple ribonucleic acid-binding reagents (e.g., in a high-throughput screening procedure) prior to the structural analysis (e.g., X-ray or electron crystallography) technique. In some embodiments of the invention, at least two agents are combined with the fusion ribonucleic acid molecule.
[0040] Related embodiments of the present invention include methods for crystallographic analysis of polynucleotides. Generally, these methods include: selecting a first polynucleotide, wherein the first polynucleotide comprises the polynucleotide sequence of a first miRNA; identifying the polynucleotide segment that forms the first loop region in the first miRNA; selecting a second polynucleotide, wherein the second polynucleotide comprises the polynucleotide sequence of a second miRNA; identifying the polynucleotide segment that forms the first loop region in the second miRNA; forming a fusion polynucleotide, which is selected such that the polynucleotide segment containing the first loop region on the first polynucleotide is replaced or exchanged with the polynucleotide segment containing the first loop region on the second polynucleotide; and then performing crystallographic analysis on the fusion polynucleotide to observe the three-dimensional structure of the fusion polynucleotide; thereby performing crystallographic analysis of the polynucleotide. In certain embodiments of these methods, the first miRNA is a miRNA having at least 90% sequence identity with: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO: 1), wherein: the residues 14-17 (GAAA) of the ribonucleic acid segment are replaced with a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides and containing the first loop region on the second polynucleotide. In certain embodiments of the present invention, the first polynucleotide comprises the sequence of SEQ ID NO: 1; and / or the second miRNA comprises a human miRNA. Generally, in these methods, the crystallographic analysis is an X-ray or electron crystallography technique; and / or the crystallographic analysis is performed in the presence of a reagent that binds to the fusion polynucleotide (e.g., an antisense oligonucleotide homologous to the segment of nucleic acid containing the first loop region on the second polynucleotide).
[0041] In an illustrative working embodiment of the present invention, we examined nine structures of the pri-miRNA hairpin loop. These studies determined that loops of 4-8 nucleotides in length are more structured than previously thought, making these and medium-length loops excellent targets for therapeutic agents. In an embodiment of the present invention, the target loop need not have a specific length and can be longer or shorter than the available examples. This recognition and our new structure determination method allow artisans to identify lead oligonucleotide compounds and perform iterative rounds of structure-based design quickly and cost-effectively. The methods of the present invention have a wide range of applications because they target processes that are important in combating infectious diseases such as coronavirus disease 2019, as well as cancer, age-related pathologies, and neurodegenerative diseases, and genetic diseases such as Duchenne muscular dystrophy, DiGeorge syndrome, etc. In one illustration, embodiments of the present invention can be used to test and validate new antisense therapeutic agents that are designed to target genes associated with human cancer pathogenesis, particularly those cancers that are not amenable to small molecule or antibody inhibition.
[0042] As described below, we determined the three-dimensional structure of human microRNA primary transcripts (pri-miRNAs) (1). Briefly, pri-miRNAs are recognized and cleaved in the nucleus by the microprocessor complex containing the Drosha ribonuclease and its RNA-binding partner protein DGCR8. The apical junctions and loops of pri-miRNAs are also binding sites for other RNA-binding proteins and metabolites that regulate microRNA maturation. More importantly, when targeted by reagents such as polynucleotides, small molecules, etc., such pri-miRNA apical loops can then be observed. In this way, mature functional microRNAs and their structures can be observed when bound to or otherwise modulated by agents with therapeutic potential.
[0043] Other aspects and embodiments of the present invention are discussed in the following sections.
[0044] Investigation of pri-miRNA apical loop length
[0045] A previous study showed that pri-miRNAs with short (<10 nt) apical loops tend to be processed inefficiently by the microprocessor (7). Taking this into consideration and building on it, we compiled a list of human pri-miRNA apical loop sequences based on the predicted secondary structures we generated using mfold (14) and similar structures provided by miRBase (13). Most of them (1,314 out of 1,881, 70%) are less than 1 cnt in length, with the highest frequency in the 4-6 nt range ( Figure 1a). RNA secondary structure prediction programs tend to include base pairs in relatively long loops that are not necessarily stable (6, 11). We partially addressed this apparent bias by ignoring 1 or 2 base pairs separated from the hairpin stem. Although the list may still underestimate the number of longer loops, it still represents the best case we know. Thus, for most pri-miRNA recognition events, Rhed must interact with a relatively short apical loop to approach the apical junction.
[0046] Scaffold-directed crystallography
[0047] To determine the three-dimensional structures of pri-miRNA apical junctions and loops, we developed a scaffold-directed crystallization method. The concept is to fuse the target (unknown) sequence to a scaffold molecule that is known to crystallize well and has an available crystal structure. The fusion should crystallize under conditions similar to those of the individual scaffold. The lattice should be able to accommodate the target portion. The scaffold structure allows determination of the structure of the fusion by molecular replacement.
[0048] To identify suitable scaffolds, we mined the Protein Data Bank for RNA crystals that met four criteria. For each RNA structure entry, we first identified the largest sphere that could be accommodated in the lattice cavity, which is characterized by a radius R max Characterization ( Figure 1 b). We considered the reported diffraction resolution. To simplify the design, we restricted the search to entries with one molecule in the asymmetric unit. Finally, we manually reviewed the lattice to find stem-loops that pointed to the lattice cavity so that an NA hairpin could be fused to it. Among the hundreds of structures surveyed, we identified only one RNA that met these requirements, the YdaO-type c-di-AMP riboswitch from Thermoanaerobacter pseudethanolicus (hereafter simply referred to as YdaO) (15).
[0049] The YdaO lattice contains large solvent channels with a short P2 stem located within the channel and away from adjacent molecules ( Figure 1 c, d). The riboswitch has a complex pseudo-dyad symmetric "cloverleaf" fold ( Figure 1 d). We replaced the GAAA tetraloop on the YdaO P2 stem with the 14-nt apical loop of pri-miR-9-1 plus 0-3 additional base pairs from the stem. After annealing in the presence of the c-di-AMP ligand, all four fusion RNAs migrated as a single band on a native gel ( Figure 1 e), indicating that the engineered pri-mi-RNA sequence does not interfere with scaffold folding.
[0050] For our representative set of short pri-miRNA loops, we generated fusions with the YdaO scaffold containing the loop plus different numbers of base pairs from the stem and screened for crystallization. We successfully obtained crystals of constructs containing 0 or 1 base pair from the pri-miRNA stem. These crystals belong to the same space group P3 1 21, with similar unit cell dimensions (Table 1). We collected X-ray diffraction data and determined its structure, with a resolution range from 2.71 to (Table 1). For three pri-miRNAs, we also collected single-wavelength anomalous dispersion (SAD) data with redundancy in the range of 79 - 115. These SAD data helped with phasing and refinement. The refined native structures showed that the scaffold part is very similar to the structure of the wild type (WT), with C1' root mean square deviation (RMSD) values ranging from 0.22 to Below we describe the pri-miRNA part. Different from most RNA loop structures in the PDB, our structures have no crystal-ligand contacts and interactions and thus reflect their own folding propensities.
[0051] Structure of the pri-miRNA apical junction and loop
[0052] Our series of pri-miRNA loop structures covers the most common loop lengths in humans, ranging from 4 to 8 nt. The longest loop is 8 nt, from pri-miR-378a (referred to as 378a + 0bp, Figure 2 a and Figure 7 a). Since RNA loops can be flexible, they are usually not well resolved in electron density. To our surprise, the 2F o -F c map of 378a + 0bp reveals a highly structured conformation with clear density for all residues. The 378a + 0bp structure clearly shows that the outermost residues C1 and A8 of the loop form a non-canonical pair, which creates a platform on which the bases from the rest of the loop stack ( Figure 2 b). At the 5'-end, C2 and U3 stack above C1. Looking from the 3'-side, A4, G5, A6, and A7 stack above A8 in four layers. On the two-base stack, C2 O2 -A7 N6 U3 O4’ -A6 N6 U3 N3 -A6 OP2 U3 O2 -A6 N7 and U3 2'OH -G5 N7 The hydrogen bond between them further stabilizes the loop( Figure 2 b). Except for A4, each loop nucleotide of pri-miR-378a is coordinated by H-bonding.
[0053] We also solved the structure of the apical loop of pri-miR-378a with one base pair from the stem (378a+1bp, Figure 2 c and Figure 7 b). The models of the two loops are in good agreement (RMSD of all non-hydrogen atoms in the loop is Figure 2 c). The 378a+1bp structure confirms the non-canonical C1-A8 pair. Interestingly, the fact that 378a+0bp and 378a+1bp are almost identical indicates that the loop conformation is not strongly affected by the terminal A:U pair from the pri-miRNA stem.
[0054] The structures of pri-miR-340 (340+1bp) and pri-miR-300 (300+0bp) contain 7nt loops. The 340+1bp structure confirms the presence of a terminal A-U pair, which is covered by an unexpected U1-U7 pair ( Figure 2 d and Figure 7 c). The G2 and U3 bases from the 5’ end of the loop stack on top of the U-U pair. This leaves only three residues (C4, G5, and U6) with a more flexible conformation at the top of the loop. In the 300+0bp structure, the terminal C-G pair of the scaffold is the same as the last base pair of the pri-miR-300 stem, so this structure is actually 300+1bp. As in the cases of 378a and 340, we observed non-canonical pairing between U1 and U7 ( Figure 2 e and Figure 7 d). Similarly, base stacking interactions between U1, U2, U3, and A4 order the 5’ end of the loop. U6 is within hydrogen bonding distance of the U2 base and almost forms another non-canonical pair. C5 is outside the density and seems to be more flexible.
[0055] In the structure of pri-miR-202 (6nt loop), we did not observe non-canonical base pairs. However, similar to other structures, the A1 base at the 5’ end of the loop stacks onto the final G-C pair of the pri-miRNA stem ( Figure 2 f and Figure 7e). The remaining part of the loop shows continuous electron density at 1σ, but we were unable to determine the conformation with high confidence. Overall, the structures of the relatively long (6 - 8 nt) pri-miRNA loops reveal extensive base stacking and non-canonical base-pairing interactions, which may stabilize the loops more than previously expected. Thus, fewer loop residues are conformationally flexible.
[0056] Next, we studied the structures of shorter pri-miRNA terminal loops (4 - 5 nt, Figure 3 ). The structure of pri-miR-208a (208a + 1bp) with a 5-nt loop revealed an unpredicted A1-U5 Hoogsteen pair located above the final G-C pair of the stem ( Figure 3 a and Figure 7 f). The central 3 nt of the loop, U2, C4, and G3, stack together and are on the A1 base in the Hoogsteen pair. Additionally, a non-canonical U-U pair from 340 + 1bp is recapitulated between U1 and U5 in the pri-miR-449c structure ( Figure 3 b and Figure 7 g). Positions U1 and G2 stack together above the terminal base pair, leaving only A3 and U4 out of the density. These two five-nucleotide loops share a theme: the two outermost residues form non-canonical base pairs, while the middle three residues are unpaired and some of their bases are stacked.
[0057] Similar to the 202 + Ibp structure described above, for pri-miR-320b-2 (with a 5-nt loop), the A1 residue of the loop is located on top of the terminal A-U pair of the stem ( Figure 3 c and Figure 7 h). Finally, in the four-nucleotide loop structure of pri-miR-19b-2 (19b-2 + 1bp), the 5'-loop nucleotide U1 stacks above the terminal base pair, and there is a partial stacking interaction of A2 on top of U1 ( Figure 3 d and Figure 7 i). U3 and G4 are mainly out of the electron density, although there may be a contact between G4 N7 and the 2'-OH of A2 (approx. ). These structures confirm that the non-canonical pairing and base stacking of 5'-loop residues witnessed in the longer loop structures also dominate the folding of shorter loops.
[0058] Structural consensus of the pri-miRNA apical junction
[0059] Our pri-miRNA stem-loop structures point to a set of common structural features that define the terminal loops. To further illustrate these features, we generated a structural alignment of all eight pri-miRNA loops ( Figure 4a). First, we always observe the canonical base pairs (5'-1 paired with 3'-1) predicted by mfold at the apex of the pri-miRNA stem. Here, we use 5'-1 to denote the first residue at the 5'-end of the pri-miRNA sequence and 3'-1 to denote the first residue at the 3'-end, due to the different loop sizes. Second, in all structures, the first nucleotide at the 5'-end of the loop base-stacks with the terminal base pair (5'-2 stacks with 5'-1 / 3'-1). Third, in five of the eight loops (378a, 340, 300, 208a, 449c), this base-stacking is also accompanied by non-canonical base pairs (5'-2 paired with 3'-2), effectively making the apical loop two nucleotides shorter than predicted. Fourth, all eight structures reveal at least one additional level of base-stacking interaction on the 5'-side (5'-3 stacks on 5'-2). In contrast, only two structures indicate a second layer of stacking on the 3'-side. In addition to these common features, the other residues of the pri-miRNA loop appear to adopt completely different conformations or are flexible.
[0060] Non-canonical base pairs contribute to thermodynamic stability
[0061] To test whether the structures of the apical junctions and loops that we observed contribute to their stability in solution, we fused eight pri-miRNA sequences to a common 5-bp helical segment ( Figure 8 a) and used optical melting to measure their thermodynamic parameters. As in the crystal structures, each pri-miRNA sequence contains an apical loop and the adjacent canonical base pairs from the stem, thus including the minimal apical junction. We expected the canonical stem base pairs to contribute differently to the overall stability, with G-C or C-G pairs in three of the pri-miRNAs being more stable than A-U and U-A pairs in the others. However, this difference does not fully account for the changes in the free energy of folding (ΔG) that we measured (Table 2). A trend emerges when we consider the non-canonical pairs that we revealed in the three-dimensional structures. The two pri-miRNAs (pri-mir-300 and pri-mir-208a) that form non-canonical pairs and have G-C or C-G as the terminal stem pairs are the most stable, while those pri-miRNAs (pri-mir-320b-2 and pri-mir-19b-2) that do not form non-canonical base pairs and contain A-U or U-A canonical stem pairs are the least stable ( Figure 4b) Most other pri-miRNA sequences, which either contain non-canonical pairs but have A-U / U-A stem pairs (pri-mir-340 and pri-mir-449c), or do not form non-canonical pairs but have G-C / C-G stem pairs (pri-mir-202), have moderate stability. The apical junction / loop of pri-mir-378a contains a C-A non-canonical pair defined by a single hydrogen bond and thus shows a ΔG similar to those from the most unstable group. Collectively, these data suggest that non-canonical pairs at the pri-miRNA apical junction contribute to their structural stability in solution.
[0062] Human pri-miRNAs prefer U-U and G-A pairs at their apical junctions
[0063] We next estimated the abundance of non-canonical pairs at the pri-miRNA apical junction by analyzing all human pri-miRNA loop sequences. Among 1,881 such sequences, 340 contain U residues at the 5' and 3' termini, which are most likely to pair as in the pri-miR-340, pri-miR-300, and pri-miR-449c structures ( Figure 4 c). Among all possible combinations at these positions, the U-U pair is the most abundant, with an expected occurrence by chance of 181. This enrichment is highly significant because the probability of observing U-U 340 times is 3x10 -28 times lower than the probability of 181 times. The second most abundant combination is 5'-G and 3'-A, observed 245 times, with a likelihood by chance lower than the most likely count of 139 for an odd number by 1x10 -16 times. The loop sequence counts for other terminal combinations such as C-A (observed 122 times) differ significantly less from the chance expectation (109 times, P 122 / P 109 = 0.42). Thus, we conclude that human pri-miRNAs prefer U-U and G-A pairs adjacent to the hairpin stem.
[0064] Interestingly, U-U and G-A are known to stabilize the hairpin loop (16) when acting as closing pairs. Our pri-miRNA loop library was partially constructed based on secondary structure prediction, which has taken into account the stabilizing effect of U-U and G-A pairs. We do not think that this small additional energy term is responsible for the enrichment of U-U and G-A as closing pairs in the pri-miRNA apical loop, because for most pri-miRNAs, the loop sequence is defined by strong canonical base pairs as part of the pri-miRNA hairpin stem. In addition, other non-canonical pairs, such as G-G, C-A, and A-C, are also considered to be stable (although to a slightly lesser extent), but they are not enriched at the pri-miRNA apical junction. This result suggests that the U-U and G-A non-canonical pairs are preferred at the pri-miRNA apical junction, which may be due to their stabilizing effect and / or specific geometric features.
[0065] pri-miRNA loops share structural features with other RNAs
[0066] We asked whether the loop conformations we found were unique to pri-miRNAs or shared with other RNA stem-loops. To address this question, we threaded the RNA hairpin sequences from the PDB onto our pri-miRNA structures and then calculated the RMSD between the threaded poses and the original PDB conformations (see the Methods section). For pri-miR-378a, we identified three slightly shorter (6- or 7-nt) loops with different sequences but highly similar folds ( Figure 9 ). Comparison of these structures revealed a generalized loop motif that we call 3'-purine-rich stacking ( Figure 9 b). In 3'-purine-rich stacking, 4-5 mainly purine bases on the 3' side of the loop stack on top of each other at the top of the helical stem. One or two pyrimidines can be found at the position farthest from the stem. On the 5' side of the loop, two or three pyrimidine residues, most commonly uridine, act as connectors between the stacked residues and the stem. These connector pyrimidines form hydrogen bonds (sometimes non-canonical base pairs) with the stacked purines, thus further stabilizing the entire loop. More generally, in the pri-miR-320b-2 structure, three purines in the UGAA tetraloop stack on top of each other and are at the apex of an adjacent U-A pair, essentially forming a 3'-purine stack. Many pri-miRNAs and other hairpin loops contain sequences consistent with 3'-purine stacking. Overall, these observations suggest that the pri-miRNA loop structure is not necessarily unique to pri-miRNAs, which is also consistent with the previously reported interactions of DGCR8 and Drosha with many other cellular RNAs (17 - 21).
[0067] Asymmetric conformational flexibility of the pri-miRNA apical loop
[0068] Structural stability and dynamics are likely to be important for pri-miRNA junctions and loops for at least two reasons. First, common conformational features are expected to be stable. Second, dynamic regions make it easier to avoid steric hindrance when binding to processing proteins and to adopt conformations that are favorable for processing. To investigate this, we first reviewed the atomic displacement parameters (ADPs, also known as temperature or B-factors) refined during structure determination. Not surprisingly, residues at the top of the loop have large ADPs, indicating that they are highly dynamic; while residues near the stem, which are involved in common structural features such as non-canonical pairs and base stacking, tend to have lower ADPs( Figure 10 ). Importantly, with the exception of pri-miR-378a, most loops show a trend of higher stability in the 5’ region of the loop and greater flexibility in the 3’ region. Stacked 5’ residues are consistently more stable than 3’ nucleotides. To further compare the ADPs between structures, we calculated the average ADP for each residue and then plotted them at the same scale( Figure 4 d). In most structures, the peak of the ADP consistently lies near the middle to the 3’ end of the loop. Notably, the UGU motif (5, 10), previously identified as important for efficient processing, is located in the 5’ region of the loop.
[0069] To gain a more detailed understanding of loop dynamics, we performed molecular dynamics simulations of pri-miRNA junctions and loop nucleotides in explicit solvent. For simplicity, the simulations included only pri-miRNA residues plus two base pairs from the scaffold, and we constrained the positions of the scaffold nucleotides to prevent strand unwinding (see Methods for details). We ran the simulations for 1 μs at 300 K and analyzed the resulting trajectories by calculating the root mean square fluctuation (RMSF) for each residue( Figure 4 e). These statistics more clearly support the trend of a broader range of conformations for the central to 3’ loop residues sample.
[0070] Correlation of Rhed binding affinity with apical loop length
[0071] We wondered how Rhed recognizes all pri-miRNA apical junctions despite the different loop lengths. We addressed this question by measuring the affinity of Rhed for pri-miRNA fragments containing the apical loop plus approximately 20 bp from the stem( Figure 8 b-i). We used electrophoretic mobility shift assay (EMSA) to determine the Rhed dissociation constant (K d )( Figure 11 ). Rhed binds to all pri-miRNA fragments with K d values ranging from 1.9 to 9.2 μM( Figure 5a-h). This difference may be important for recognition, especially when pri-miRNAs compete for processing machinery. We plotted the binding ΔG versus the overall loop length ( Figure 5 i), and noted a trend for longer loops to bind more tightly. This trend became more pronounced when we corrected the loop lengths according to our 3D structures (length minus the number of residues involved in non-canonical pairs, Figure 5 j). Our results provide a biochemical explanation for pri-miRNA loop length preferences, although we cannot rule out the possibility that differences in the pri-miRNA stem also contribute to the range of Rhed affinities. We note that pri-miR-340, which contains a UGU motif on the 5’ side of the loop, binds Rhed with an affinity (K d = 3.5 μM) similar to that of other constructs lacking this sequence.
[0072] Discussion
[0073] We provide a working implementation that demonstrates a proof-of-concept that scaffold-directed crystallography can be a powerful tool for RNA structural biology. This approach is largely analogous to the popular fixed-arm MBP fusion technique, in which the target protein is attached to MBP in a fixed orientation via a continuous α-helical linker (22). However, our engineered approach specifically positions the target RNA in the lattice voids of the scaffold crystal. Such a design confers several additional advantages: (1) Since the target moiety does not disrupt existing lattice contacts, the fusion molecule can be crystallized under native conditions; (2) Since extensive conditions do not need to be re-screened, crystallization requires a minimal amount of purified fusion RNA; (3) The target does not interact with neighboring molecules in the lattice, thus enabling its structure to closely represent the conformation in solution.
[0074] Applying this technique to the pri-miRNA recognition problem provides an atomic-level survey of the apical junctions and loop structures of eight pri-miRNAs. These loops cover the most common loop lengths in human pri-miRNAs. Collectively, these structures reveal a structural consensus involving non-canonical base pairs that close the apical loop and further base stacking at the 5’ end. The previously reported NMR structure of pre-miR-20b supports this consensus (6). The pre-miR-20b stem terminates in a G-U pair, and the adjacent 5’ loop nucleotide (G) stacks on top of this pair ( Figure 13)。The comparison of the top 20 NMR solutions confirmed that these are stable features of the molecule. NMR studies of pre-miR-21 revealed weak signals corresponding to two tandem U-G / G-U pairs at the apical junction and indicated that the 14 nt apical loop is otherwise unstructured (11). Except for the apical junction, our apical loop and the one in the NMR structure differ in three-dimensional conformation, suggesting that their conformations are not direct specificity determinants. These conformations are related to their respective functions. For example, the pri-miR-125a loop can serve as an aptamer domain for binding folic acid (23).
[0075] The observation of non-canonical pairs at the pri-miRNA apical junction itself has important structural and functional implications. Our optical melting experiments showed that these pairings contribute to the thermodynamic stability of the RNA in solution ( Figure 4 b). In particular, U-U and G-A pairs are highly enriched at the apical junctions of human pri-miRNAs ( Figure 4 c). These pairs are generally conserved. For example, the U-U pair in pri-miR-340 is almost completely conserved, while nucleotide variations occur at all other positions. The only variation of the U-U pair is the substitution by a U-G pair in the central fruit bat (Pteropus alecto). Thus, these non-canonical pairs at the apical junction may be important for miRNA maturation, although their exact functions remain to be determined. The microprocessor recognizes the pri-miRNA hairpin by clamping its stem at both ends (24, 25). The optimal pri-miRNA hairpin stem length is estimated to be 35 ± 1 bp, counted by internal non-canonical pairs (10). Our study shows that the terminal non-canonical pairs at the apical junction must be considered. Previous high-throughput mutagenesis of pri-miR-16-1 showed that it is due to the stem length exceeding the optimal length, and the disruption of the canonical pairs at the stem tip increases the cleavage efficiency of the microprocessor (10). In pri-miR-16-1, a G-A pair is expected to form and stack at the end of the hairpin stem. In this case, the G-A pair would need to be disrupted together with the adjacent canonical pairs. This inhibitory effect makes it possible for RNA-binding proteins and RNA helicases to activate miRNA maturation (26). Conversely, we imagine that in the case where the pri-miRNA helical stem is shorter than optimal, non-canonical pairs will help the hairpin fit into the microprocessor complex.
[0076] The conformation of the apical junction can also be preferentially recognized by the microprocessor. In fact, the microprocessor prefers a U-G base pair over a Watson-Crick base pair at the 35th bp position of the pri-miR-30a stem (counted from the basal junction) (10). We re-analyzed another high-throughput mutagenesis data (5) and found that C-A pairs are highly enriched at the apical junctions of the microprocessor cleavage products. Figure 12)。In addition, the tendency of 5’-loop residue stacking and the more flexible tendency of the 3’-loop portion allow the UGU motif to be positioned and exposed for recognition by the processing machinery. Further studies are needed to validate this idea.
[0077] Our analysis of human pri-miRNA loop sequences showed that most of them are shorter than the optimal ≥10 nt. Among the eight pri-miRNAs with loop lengths between 4 and 8 nt, we observed a correlation between loop length and the change in free energy of binding to Rhed ( Figure 5 i). The correlation was improved when the residues involved in non-canonical pairs were excluded from the calculation of loop length ( Figure 5 j). Preferential binding to Rhed primes the pri-miRNA in a favorable position for processing, thus providing a biochemical explanation for the optimal loop length of ≥1 cnt (7). ΔG 结合 differences from Rhed are within 1 kcal / mol. We believe that such modest differences can have significant biological and pathological consequences, especially when the microprocessor becomes limiting (e.g., in many cancer cells). Preferential binding to the microprocessor, as represented by the apical junction interaction with Rhed shown here, may generate a processing hierarchy among pri-miRNAs and help determine miRNA expression profiles.
[0078] The apical junction and loop are also parts of pre-miRNA that are exported to the cytoplasm and cleaved by the Dicer ribonuclease in the miRNA maturation pathway. Previous studies have shown that the stem and loop lengths of pre-miRNA can affect the cleavage efficiency of both Drosha and Dicer (8). Further studies are needed to understand how the apical junction and loop structures contribute to the Dicer processing step. In addition, developing potential therapeutic agents targeting pri-miRNA, mRNA, and viral RNA hairpin loops is of great significance (11, 26, 27). Our structure shows that the pri-miRNA loop contains more structure than expected, which would reduce the entropy loss upon binding. Our crystallization method should allow structure-based inhibitor design.
[0079] Methods
[0080] Pri-miRNA apical loop analysis
[0081] To measure the approximate size of the apical loop, we downloaded all annotated human "hairpin" sequences and their genomic coordinates from miRBase (release 21). miRBase hairpins typically include the pre-miRNA portion as well as a variable number of additional base pairs from the basal stem. For each hairpin, we extended the genomic sequence by the same number of nucleotides at the 5' and 3' ends of the RNA until the total length equaled 150 nt. This 150-nt window contains the complete pri-miRNA hairpin, plus some single-stranded RNA on either side of the basal junction. We then used MFOLD (14) to generate predicted secondary structures for all pri-miRNA hairpins and generally retained the structure with the highest score (i.e., the lowest predicted folding free energy). We manually reviewed all predictions to ensure that they reflected the expected hairpin structure, where the mature miRNA sequence originated from one or both strands of the stem; in cases where mfold predicted alternative conformations, we selected the structure with the lowest free energy that contained a stem length of approximately three helical turns. We manually compared the secondary structures with those from miRBase and also removed 1-2 base pairs in the hairpin that were separated from the stem and thus considered unstable.
[0082] PDB Mining and Identification of the YdaO Crystallization Scaffold
[0083] We first filtered the PDB to obtain X-ray structures containing only RNA molecules (no proteins or DNA). To identify voids in the lattice, we wrote a PyMOL script that implemented a grid search algorithm in the following steps. (1) Generate a 3×3×3 unit cell block (i.e., 27 copies of the unit cell). The unit cell at the center of the block sees all possible lattice voids, whether inside or between unit cells. (2) Using three unit vectors along each unit cell axis (i.e., the a, b, and c vectors of length ), iteratively generate grid points of the form 5*i*a + 5*j*b + 5*k*c, where the integer values of i, j, and k are less than the respective unit cell edge lengths divided by 5. This gives grid points with spacing. (3) For each grid point, calculate the distance to all C1' atoms in the supercell and identify the shortest distance as R local . For each structure, identify the grid point with the largest R local as R max .
[0084] To find a suitable scaffold, we subsequently manually reviewed those with large R maxThe value and the structure of a single molecule in the asymmetric unit. We traced the chains to look for any stem-loops that project into the lattice cavity. Among hundreds of candidates examined, only the P2 stem-loop from the YdaO riboswitch (PDBID: 4QK8) met these criteria (15).
[0085] Preparation and native gel electrophoresis of YdaO WT and pri-miR-9-1 fusion RNAs
[0086] We initially designed the W.T. YdaO construct to contain a T7 promoter sequence at the 5’ end and an HDV ribozyme on the 3’ side, along with flanking EcoRI and BamHI restriction sites. This fragment was synthesized as a gene block (IDT), double digested, and cloned into the pUC19 plasmid. Clones were verified by Sanger sequencing. To replace the P2 loop nucleotides with the pri-miRNA stem-loop, we used a two-round PCR protocol. All reactions were carried out with Q5 high-fidelity DNA polymerase (New England Biolabs) according to the reaction settings and cycling conditions recommended by the manufacturer. All reaction mixtures contained the same reverse primer, which annealed to the 3’ end of the HDV and contained a BamHI site (5′-CGT GGATCC GGTCCCATTC-3′) (SEQ ID NO: 2). For the first PCR, the forward primer contained the pri-miRNA sequence plus approximately 20 nt upstream and downstream of the scaffold. The forward primer used for the pri-miR-9-1 fusion was 5′-CTATAGGTTGCCGAATCC GTGGTGTGGAGTCT GGTACGGAGGAACCGCTTTTTG-3′ (pri-miR-9-1 +0 bp) (SEQ ID NO: 3); 5′-CTATAGGTTGCCGAATCC AGTGGTGTGGAGTCTT GGTACGGAGGAACCGCTTTTTG-3′ (pri-miR-9-1 +1 bp) (SEQ ID NO: 4); 5′-CTATAGGTTGCCGAATCCGAGTGGTGTGGAGTCTTCGGTACGGAGGAACCGCTTTTTG-3′ (pri-miR-9-1 +2 bp) (SEQ ID NO: 5); 5′-CTATAGGTTGCCGAATCC AGAGT GGTGTGGAGTCTTCT GGTACGGAGGAACCGCTTTTTG-3′ (pri-miR-9-1 +3 bp) (SEQ ID NO: 6). This PCR product was gel purified and 1 μL was used as a template for the second round of PCR. All reaction mixtures contained the same reverse primer and forward primer (SEQ ID NO: 7), which anneals to the common scaffold residues (bold), with the addition of a T7 promoter (italic) and an EcoRI site (underlined). The second-round PCR product was gel purified, digested with EcoRI and BamHI, and ligated into pUC19. Clones containing the desired insert were sequence verified.
[0087] For the WTYdaO and pri-miR-9-1 fusion constructs, we prepared maxiprep plasmids and linearized them by overnight digestion with BamHI. The transcription reaction contained approximately 400 μg of linearized template, 40 mM Tris pH 7.5, 25 mM MgCl 2 , 4 mM DTT, 2 mM spermidine, 40 μg of inorganic pyrophosphatase (Sigma), 0.7 mg of T7 RNA polymerase, and 3 mM of each NTP in a total volume of 5 mL. After incubation at 37 °C for 4.5 hr, the final MgCl 2 concentration was adjusted to 40 mM, and the reaction was incubated for an additional 45 minutes. Despite the elevated Mg 2+ concentration, we only observed partial cleavage of the HDV ribozyme. The reaction was ethanol precipitated and purified using a denaturing 10% polyacrylamide slab gel. The desired product was visualized by UV shadowing and excised from the gel. The gel slice was crushed and extracted overnight at 4 °C in 30 mL of TEN buffer (150 mM NaCl, 20 mM Tris pH 7.5, 1 mM EDTA). Then, we spun down the gel slice and concentrated the RNA in an Amicon Ultra-15 centrifugal filter unit with a 10 kDa molecular weight cut-off (MWCO). The RNA buffer was exchanged three times into 10 mM HEPES pH 7.5 and concentrated to a final volume of approximately 50 μL.
[0088] For analysis on native gels, a 5 μM RNA stock solution was prepared by diluting the purified RNA into 5 mM Tris pH 7.0. Next, 2.5 μL of RNA was mixed with an equal volume of 2X annealing buffer, which contained 35 mM Tris pH 7.0, 100 mM KCl, 10 mM MgCl 2 and 20 μM c-di-AMP (Sigma). The mixture was heated at 90 °C for 1 minute, then rapidly cooled on ice, and then incubated at 37 °C for 15 minutes. The annealed RNA was mixed with a solution containing 40 mM Tris pH 7.0, 50 mM KCl, 5 mM MgCl 2, mixed with 20% (v / V) glycerol and xylene cyanol 2X loading dye, and analyzed on a 10% polyacrylamide gel using Tris-borate (TB) running buffer. The gel was stained with SybrGreen II and scanned on a Typhoon 9410 variable mode imager (GE Healthcare).
[0089] Preparation of the pri-miRNA-YdaO fusion for crystallization
[0090] Given the poor HDV self-cleavage efficiency of the pri-miR-9-1 fusion we observed, we chose to change strategies. Instead of using ribozymes to create homogeneous 3’ ends, we used PCR to generate transcription templates where the two 5’ residues on the antisense DNA strand were 2’-O-methylated. These modifications have been shown to reduce nontemplated nucleotides added by T7 RNA polymerase (28). We utilized a three-round PCR method to create the transcription template. All subsequent reactions contained the same reverse primer 5′-mCmUCCTTCCTTTATTGCCTCC-3′ (SEQ ID NO: 8), where “m” denotes 2’-O-methylation. For the first round of PCR, we set up a 50 μL reaction with Q5 polymerase to amplify the 3’ fragment of YdaO with the forward primer 5′-GGTACGGAGGAACCGCTTTTTG-3′ (SEQ ID NO: 9), and performed 30 cycles of amplification. The product was gel purified and 1 μL was used as the template for the next round. In the second round of PCR, we used unique forward primers for each construct containing the pri-miRNA loop and stem sequences that annealed to the 3’ YdaO fragment from the first stage. The primer sequences were
[0091] 5′-CTATAGGTTGCCGAATCC ATATGT GGTACGGAGGAACCGCTTTTTG-3′ (19b-2+1bp) (SEQ ID NO: 10);
[0092] 5′-CTATAGGTTGCCGAATCC GATCTGGC GGTACGGAGGAACCGCTTTTTG-3′ (202+1bp) (SEQ ID NO: 11);
[0093] 5′-CTATAGGTTGCCGAATCC GATGCTC GGTACGGAGGAACCGCTTTTTG-3′ (208a+1bp) (SEQ ID NO: 12);
[0094] 5′-CTATAGGTTGCCGAATCC GTTTACTTG GGTACGGAGGAACCGCTTTTTG-3′(300 + 1bp)(SEQ ID NO: 13);
[0095] 5′-CTATAGGTTGCCGAATCC AAAGTT GGTACGGAGGAACCGCTTTTTG-3′(320b - 2 + 1bp)(SEQ ID NO: 14);
[0096] 5′-CTATAGGTTGCCGAATCC ATGTCGTTT GGTACGGAGGAACCGCTTTTTG-3′(340 + 1bp)(SEQ ID NO: 15);
[0097] 5′-CTATAGGTTGCCGAATCC ACCTAGAAAT GGTACGGAGGAACCGCTTTTTG-3′(378a + 1bp)(SEQ ID NO: 16); and
[0098] 5′-CTATAGGTTGCCGAATCC ATGATTT GGTACGGAGGAACCGCTTTTTG-3′(449c + 1bp)(SEQ ID NO: 17).
[0099] The reactant was also 50 μL, and 30 cycles were carried out using Q5 polymerase. The product from the second round of PCR was analyzed by agarose gel electrophoresis to confirm amplification, and 40 μL of the reactant was used as the template for the third round of PCR without further purification. 2 mL of the PCR reactant was used with Phusion High-Fidelity DNA Polymerase (Thermo-Fisher) and the forward primer
[0100] 5′-GCAGAATTCTAATACGACTCACTATAGGTTGCCGAATCC-3′, (SEQ ID NO: 18) and 35 cycles were run.
[0101] The third-stage PCR products were purified on a HiTrap Q HP column (GE Healthcare). Buffer A contained 10 mM NaCl and 10 mM HEPES pH 7.5; buffer B was the same but contained 2 M NaCl. The column was equilibrated with 20% buffer B and the desired DNA product was eluted with a linear gradient to 50% B at 2 ml / min for 10 minutes. We analyzed the peak fractions on an agarose gel to confirm that they contained a single band of the correct size. The peak fractions were then pooled and concentrated in an Amicon filter unit (10 kDa MWCO) and washed with water to remove excess salt. The concentration of the DNA template (final volume of approximately 200 μL) was determined by UV absorbance.
[0102] The transcription reaction for the pri-miR-9-1 fusion was set up as described above but in a volume of 10 mL and contained 2.8 fmol of DNA template. The reaction was run at 37 °C for 4 hours and then subjected to phenol-chloroform extraction. The transcription was concentrated in an Amicon filter unit (10 kDa MWCO) and washed with 0.1 M triethylammonium acetate (TEAA) pH 7.0. The RNA (approximately 2 μL) was injected onto a Waters XTerra MS C18 reversed-phase HPLC column (3.5 μm particle size, dimensions 4.6 x 150 mm) maintained at 54 °C. TEAA and 100% acetonitrile were used as the mobile phases. The column was washed with 6% acetonitrile and the RNA was eluted with a gradient to 17% acetonitrile at 0.4 ml / min for 80 minutes. The peak fractions were analyzed on a denaturing 10% polyacrylamide gel. The pure fractions were pooled and buffer-exchanged into 10 mM HEPES pH 7.0 using an Amicon filter unit. The RNA was concentrated to a final volume of <50 μL and the concentration was determined by UV absorbance.
[0103] Crystallization, data collection, and structure determination
[0104] All RNA-c-diAMP complexes were prepared as described (15). Briefly, a solution containing 0.5 mM RNA, 1 mM c-di-AMP, 100 mM KCl, 10 mM MgCl 2 and 20 mM HEPES pH 7.0 was heated to 90 °C for 1 minute, rapidly cooled on ice, and equilibrated at 37 °C for 15 minutes immediately before crystallization. Screening was carried out in a 24-well plate containing 0.5 mL of well solution; the hanging drop consisted of 1 μL of RNA plus 1 μL of well solution. The plate was incubated at room temperature and crystals typically grew to full size (100 μm to over 200 μm) within a week. For 19b-2 + 1bp, the well solution contained 1.7 M (NH 4 ) 2 So 4 、0.2M Li2 SO 4 and 0.1 M HEPES pH 7.1. For 202+1 bp, 208a+1 bp, and 320b-2+1 bp, the wells contained 1.9 M (NH 4 ) 2 SO 4 , 0.2 M Li2SO 4 and 0.1 M HEPES pH 7.4. The well solution for 378a+0 bp contained 1.7 M (NH 4 ) 2 SO 4 , 0.2 M Li 2 SO 4 and 0.1 M HEPES pH 7.4. For the remaining constructs, crystallization was performed in 96-well plates, and the hanging drops consisted of 0.4 μL of RNA plus 0.4 μL of well solution. For 300+1 bp, the well solution contained 1.88 M (NH4) 2 SO 4 , 0.248 M Li2SO 4 and 0.1 M HEPES pH 7.4, and for 300+0 bp, it contained 1.90 M (NH 4 ) 2 So 4 , 0.158 M Li 2 So 4 and 0.1 M HEPES pH 7.4. Construct 340+1 bp crystallized from a well solution containing 1.89 M (NH 4 ) 2 SO 4 , 0.214 M Li 2 SO 4 and 0.1 M HEPES pH 7.4. Construct 378a+1 bp crystallized from 1.63 M (NH 4 ) 2 SO 4 , 0.272 M Li 2 SO 4 and 0.1 M HEPES pH 7.4. For construct 449c+1 bp, the wells contained 1.89 M (NH 4 ) 2 SO 4 , 0.128 M Li 2 SO 4 and 0.1 M HEPES pH 7.4.
[0105] All crystals were briefly soaked in a solution containing 20% (w / v) PEG 3350, 20% (v / v) glycerol, 0.2 M (NH 4 )2 SO 4 、 0.2 M Li 2 SO 4 and 0.1 M HEPES pH 7.3 cryoprotectant solution, and then flash-frozen in liquid nitrogen. Data were collected at 100 K at the Advanced Photon Source beamline 24-ID-C or the Advanced Light Source beamline 8.3.1. For all constructs, we collected native datasets at a wavelength of approximately . For 320b-2+1bp, 378a+0bp, and 449c+1bp, we measured phosphorus anomalous scattering by collecting approximately extra-high redundancy datasets from 1, 2, or 3 crystals, respectively. The data were indexed, integrated, and scaled using XDS (29).
[0106] When anomalous data were available, we used a combined molecular replacement / single anomalous dispersion method (MR-SAD) to generate partial experimental phases. The molecular replacement model consisted of the YdaOc-di-AMP riboswitch structure (PDB ID: 4QK8), and the GAAA tetraloop on the P2 stem had been removed from the model. Phases were obtained using the default settings in the Phaser-MR protocol in Phenix (30).
[0107] For all constructs, we obtained initial solutions by performing rigid-body fitting of the MR model (above figure) to the data using Phenix (including experimental phase restraints if available). This yielded excellent initial models with R 工作 < 30%. We then examined the electron density maps in the region of the P2 stem. For all RNAs, extra density for the missing base pairs and loops was clearly visible in the 2F o -F c and difference maps. We then modeled the missing residues in Coot (31). In cases where the density was unclear, we stopped modeling the incomplete loops and performed another round of refinement of coordinates, ADP, and TLS parameters using Phenix. This usually revealed extra density for the missing residues. Once the loops were fully modeled, we performed subsequent rounds of refinement and manual adjustment as described above until reasonable R factors and model geometries were obtained.
[0108] Simulated annealing composite omit maps were calculated in Phenix ( Figure 7)。In the cases of 19b-2+1bp, 202+1bp, 320b-2+1bp, 340+1bp, and 378a+1bp, the standard annealing temperature (5000 °C) and other default parameters produced reasonable plots. However, for 300+0bp, 300+1bp, and 378a+0bp, the default settings generated noisy plots with disrupted density regions. To improve the quality of the plots, we reduced the annealing temperature to 1000 °C and excluded the bulk solvent mask from the omitted regions. This type of composite-omission plot is called a Polder plot and prevents the solvent mask from obscuring weaker densities (32).
[0109] Comparison with known RNA loop structures in the PDB
[0110] To identify RNA loops in the PDB that are structurally similar to our pri-miRNA loop model, we first extracted the coordinates of the pri-miRNA apical junctions and loops. The search pool was the same set of RNA structures used to identify the above-mentioned crystallization scaffolds. For each structure in the PDB set, we used DSSR to identify all hairpin loops. We extracted the RNA sequences from each hairpin loop and eliminated loops shorter than the pri-miRNA sequence. For loops longer than the pri-miRNA, we used a sliding window to obtain all segments of loops of the same length. Then we threaded each loop sequence onto the pri-miRNA model using the "rna_thread" routine in Rosetta (33). Using a PyMOL script, we aligned the resulting threaded models with the original hairpin loops and calculated the RMSD between the two models. We aggregated and sorted the RMSD data for all PDB structures and manually inspected the loops with small RMSD to find hits with structural similarity.
[0111] Optical melting
[0112] RNA for the optical melting experiments was transcribed in vitro from synthetic DNA templates (IDT). The oligonucleotide template sequences used were
[0113] and The T7 promoter is shown in italics, while the pri-miRNA junction / loop segments are shown in bold. The template was annealed with a second strand complementary to the T7 promoter and added to the large-scale (10 mL) transcription reaction as described above. The reaction was precipitated with ethanol and purified on a 20% polyacrylamide denaturing gel. The desired band was recovered by UV shadowing. After gel extraction, the sample buffer was exchanged into water and concentrated in an Amicon centrifugal filter device.
[0114] For each RNA, a set of six dilutions was prepared in 50 mM NaCl and 10 mM sodium cacodylate pH 7.0 such that the initial absorbance ranged from approximately 1.0 to 0.1 AU. Samples were annealed by heating to 95 °C for 1 minute and rapidly cooling on ice, then equilibrated to 12 °C. Melting measurements were performed using a Cary Bio300 UV-visible spectrophotometer equipped with a Peltier-type temperature-controlled sample changer. Absorbance at 260 nm was recorded while the RNA was heated from 12 °C to 92 °C at a rate of 0.8 °C / min. Melting curves were analyzed using Prism (GraphPad, version 7) and fit with the equation where the absorbance (A) is approximated as a function of temperature (T). Changes in entropy (ΔS) and enthalpy (ΔH) as well as the slopes (m f and b f ) and y-intercepts (b u and b u ) of the linear regions of the double-stranded (m u and b u ) and single-stranded (m u and b u ) were all fit. The melting temperature and thermodynamic parameters at 37 °C were then derived from these parameters (Table 2).
[0115] Electrophoretic mobility shift assay
[0116] As previously described, human heme-bound Rhed protein was overexpressed in E. coli and purified using ion-exchange and size-exclusion chromatography (25). Radiolabeled pri-miRNA stem-loops ( Figure 8 b-i) were prepared by in vitro transcription. The DNA template consisted of an antisense oligonucleotide covering the desired sequence plus the T7 promoter, annealed to a sense oligonucleotide with the T7 promoter sequence (34). Each 20 μL transcription reaction contained 50 fmol template, 40 mM Tris pH 7.5, 25 mM MgCl 2 , 4 mM DTT, 2 mM spermidine, 2 μg T7 RNA polymerase, 0.5 mM each of ATP, UTP, CTP, and GTP, 3 mM each, and 3 nmol α- 32 P-ATP (10 μCi). Transcription was run at 37 °C for 2 hours and the RNA was purified on a denaturing 15% polyacrylamide gel. The RNA was extracted overnight in TEN buffer at 4 °C, precipitated with isopropanol, and resuspended in 40 μL of water.
[0117] We employed the recently reported EMSA protocol to examine the Rhed-pri-miRNA interaction (35). RNAs were diluted in 100 mM NaCl, 20 mM Tris pH 8.0 and heated at 90 °C for 1 minute, then rapidly cooled on ice. The annealed RNAs were added to binding reactions containing 10% (v / v) glycerol, c. 1 mg / ml yeast tRNA, 0.1 mg / ml BSA, 5 μg / ml heparin, 0.01% (v / v) octylphenoxypolyethoxylethanol (IGEPAL CA-630), 0.25 units RNase-OUT ribonuclease inhibitor, xylene cyanol, 20 mM Tris pH 8.0 and 0 - 20 μM Rhed protein. The final salt concentration of the solution was 150 mM NaCl. The binding reactions were incubated at room temperature for 30 minutes before loading onto a 10% polyacrylamide gel. Both the gel and the electrophoresis buffer contained 80 mM NaCl, 89.2 mM Tris base and 89.0 mM boric acid (final pH 8.2). The gel was run at 110 V for 45 minutes at 4 °C, then dried and exposed to a storage phosphor screen. Subsequently, the screen was scanned on a Typhoon scanner (GE Healthcare). Free and bound RNA bands were quantified using Quantity One software (BioRad) and fit to the Hill equation in Prism.
[0118] Molecular dynamics simulations
[0119] Coordinates corresponding to pri-miRNA residues plus two G-C pairs from the P2 stem of the scaffold were extracted from each crystal structure. Hydrogens were added to the models in GROMACS (36), and the RNAs were solvated in a truncated dodecahedron box with TIP3P water molecules. The box was large enough to space the RNA from any periodic copies of itself by at least 1 nm. Next, K + and Cl - ions were added to the system to neutralize the net charge and bring the final KCl concentration to 0.1 M. The CHARMM27 force field, Verlet cut-off scheme and particle-mesh Ewald electrostatics were used for all calculations. The system was energy minimized until the maximum force acting on any atom was less than 900 kJ / mol / nm. The final potential energy of the system was in the range of -1.3x10 5 kJ / mol.
[0120] Next, the system was initially equilibrated in two steps, first in the NVT ensemble and then in the NPT ensemble. Both equilibration simulations were run for 2 ns at 300 K with a time step of 2 fs. During NVT, the temperature was controlled by velocity rescaling. For NPT, the pressure was maintained at 1 bar using the Parrinello-Rahman barostat. For the production MD runs, positional constraints were applied to the G-C pairs from the scaffold, and all pri-miRNA nucleotides were unconstrained. All production simulations were run in NPT with a time step of 2 fs for a total of 1 μs. The rmsf and clustering functions in GROMACS were used to analyze the trajectories.
[0121] Reanalysis of the pri-miR-223 high-throughput processing assay
[0122] Sequencing data from the previously reported pri-miRNA-223 processing assay were downloaded from the Sequence Read Archive (accession number: SRA051323)(5). Reads corresponding to pri-miR-223 were aligned using Bowtie2(37). Any reads containing unknown nucleotides were eliminated. Reads from the input or selected libraries were separated by their corresponding barcodes and counted using Python.
[0123] Table 1. Data collection and refinement statistics for the pri-miRNA loop fusion structure.
[0124]
[0125]
[0126] Table 2. Thermodynamic parameters for pri-miRNA apical junctions and loop folding at 50 mM NaCl, reported as ± standard deviation.
[0127]
[0128] Public references
[0129] 1. Ha, M. and V. N. Kim, Regulation of microRNA biogenesis. Nat. Rev. Mol. Cell Biol., 2014. 15: 509-24.
[0130] 2. Krasilnikov, A. S., et al., Crystal structure of the specificity domain of ribonuclease P. Nature, 2003. 421: 760-4.
[0131] 3. Reiss, C.W., Y. Xiong, and S.A. Strobel, Structural Basis for Ligand Binding to the Guanidine-I Riboswitch. Structure, 2017. 25: 195 - 202.
[0132] 4. Byrne, R.T., et al., The crystal strucrute of unmodified tRNAPhe from Escherichia coli. Nucleic Acids Res. 2010. 38: 4154 - 62.
[0133] 5. Auyeung, V.C., et al., Beyond secondary structure: primary-sequence determinants license pri-miRNA hairpins for processing. Cell, 2013. 152: 844 - 58.
[0134] 6. Chen, Y., et al., Rbfox proteins regulate microRNA biogenesis by sequence-specific binding to their precursors and target downstream Dicer. Nucleic Acids Res., 2016. 44: 4381 - 95.
[0135] 7. Zeng, Y., R. Yi, and B.R. Cullen, Recogntiion and cleavage of primary microRNA precursors by the nuclear processing enzyme Drosha. EMBO J, 2005. 24: 138 - 148.
[0136] 8. Zhang, X. and Y. Zeng, The terminal loop region controls microRNA processing by Drosha and Dicer. Nucleic Acids Res, 2010. 38: 7689 - 97.
[0137] 9. Ma, H., et al., Lower and upper stem-single-stranded RNA junctions together determine the Drosha cleavage site. Proc Natl Acad Sci U S A, 2013. 110: 20687-92.
[0138] 10. Fang, W. and D. P. Bartel, The Menu of Features that Define Primary MicroRNAs and Enable De Novo Design of MicroRNA Genes. Mol. Cell, 2015. 60: 131-45.
[0139] 11. Shortridge, M. D., et al., A Macrocyclic Peptide Ligand Binds the Oncogenic MicroRNA-21 Precursor and Suppresses Dicer Processing. ACS Chem. Biol., 2017. 12: 1611-1620.
[0140] 12. Chirayil, S., et al., NMR characterization of an oligonucleotide model of the miR-21 pre-element. PLoS One, 2014. 9: e108231.
[0141] 13. Kozomara, A. and S. Griffiths-Jones, miRBase: integrating microRNA annotation and deep-sequencing data. Nucleic Acids Res., 2011. 39: D152-7.
[0142] 14. Zuker, M., Mfold web server for nucleic acid folding and hybridization prediction. Nucleic Acids Res, 2003. 31: 3406-3415.
[0143] 15. Gao, A. and A. Serganov, Structural insights into recognition of c-di-AMP by the ydaO riboswitch. Nat. Chem. Biol., 2014.10: 787-92.
[0144] 16. Serra, M.J., T.J. Axenson, and D.H. Turner, A model for the stabilities of RNA hairpins based on a study of the sequence dependence of stability for hairpins of six nucleotides. Biochemistry, 1994.33: 14289-96.
[0145] 17. Triboulet, R., et al., Post-transcriptional control of DGCR8 expression by the Microprocessor. RNA, 2009.15: 1005-11.
[0146] 18. Kadener, S., et al., Genome-wide identification of targets of the drosha-pasha / DGCR8 complex. RNA, 2009.15: 537-45.
[0147] 19. Macias, S., et al., DGCR8 HITS-CLIP reveals novel functions for the Microprocessor. Nat. Struct. Mol. Biol., 2012.19: 760-766.
[0148] 20. Heras, S.R., et al., The Microprocessor controls the activity of mammalian retrotransposons. Nat Struct Mol Biol, 2013.20: 1173-81.
[0149] 21. Han, J., et al., Posttranscriptional crossregulation between Drosha and DGCR8. Cell, 2009. 136: 75 - 84.
[0150] 22. Moon, A.F., et al., Asynergistic approach to protein crystallization: combination of a fixed - arm carrier with surface entropy reduction. Protein Sci., 2010. 19: 901 - 13.
[0151] 23. Terasaka, N., et al., A human microRNA precursor binding to folic acid discovered by small RNA transcriptomic SELEX. RNA, 2016. 22: 1918 - 1928.
[0152] 24. Nguyen, T.A., et al., Functional Anatomy of the Human Microprocessor. Cell, 2015. 161: 1374 - 87.
[0153] 25. Quick - Cleveland, J., et al., The DGCR8 RNA - binding heme domain recognizes primary microRNAs by clamping the Hairpin. Cell Rep., 2014. 7: 1994 - 2005.
[0154] 26. Michlewski, G., et al., Posttranscriptional regulation of miRNAs harboring conserved terminal loops. Mol Cell, 2008. 32: 383 - 93.
[0155] 27. Brakier-Gingras, L., J. Charbonneau, and S. E. Butcher, Targeting frameshifting in the human immunodeficiency virus. Expert Opin. Ther. Targets, 2012. 16: 249-58.
[0156] 28. Kao, C., M. Zheng, and S. Rudisser, A simple and efficient method roreduce nontemplated nucleotide addition at the 3 terminus of RNAs transcribed by T7 RNA polymerase. RNA, 1999. 5: 1268-72.
[0157] 29. Kabsch, W., XDS. Acta Crystallogr. D Biol. Crystallogr., 2010. 66: 125-32.
[0158] 30. Adams, P. D., et al., PHENIX: a comprehensive Python-based system for macromolecular structure solution. Acta Crystallogr D Biol Crystallogr, 2010. 66: 213-21.
[0159] 31. Emsley, P., et al., Features and development of Coor. Acta Crystallogr. D Biol. Crystallogr., 2010. 66: 486-501.
[0160] 32. Liebschner, D., et al., Polder maps: improving OMIT maps by excluding bulk solvent. Acta crystallographica. Section D, Structural biology, 2017. 73: 148-157.
[0161] 33. Cheng, C.Y., F.C. Chou, and R. Das, Modeling complex RNA tertiary folds with Rosetta. Methods Enzymol., 2015. 553: 35 - 64.
[0162] 34. Milligan, J.F., et al., Oligoribonucleotide synthesis using T7 RNA polymerase and synthetic DNA templates. Nucleic Acids Res., 1987. 15: 8783 - 98.
[0163] 35. Partin, A.C., et al., Heme enables proper positioning of Drosha and DGCR8 on primary microRNAs. Nat. Commun., 2017. 8: 1737.
[0164] 36. Abraham, M.J., et al., GROMACS: High performance molecular simulations through multi - level parallelism from laptops ro supercomputers. SoftwareX, 2015. 1 - 2: 19 - 25.
[0165] 37. Langmead, B. and S.L. Salzberg, Fast gapped - read alignment with Bowtie2. Nat. Methods, 2012. 9: 357 - 9.
[0166] Conclusion
[0167] The description of the preferred embodiments of the present invention ends here. For purposes of illustration and description, the foregoing description of one or more embodiments of the present invention has been presented. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. Given the above teachings, many modifications and variations are possible.
[0168] All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.
Claims
1. A substance composition comprising a ribonucleic acid having at least 90% sequence identity with the following: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGU UAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCG ACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO:1), wherein: Residues 14 - 17 (GAAA) of the ribonucleic acid are replaced by a heterologous segment of a nucleic acid having a length between 4 and 33 nucleotides.
2. The composition according to claim 1, further comprising a reagent that binds to the ribonucleic acid.
3. The composition according to claim 2, wherein the reagent is a polynucleotide that hybridizes to the ribonucleic acid.
4. The composition according to claim 1, wherein the heterologous segment of the nucleic acid forms a loop structure in a naturally occurring RNA molecule.
5. The composition according to claim 4, wherein the heterologous segment of the nucleic acid comprises the complete loop structure in the naturally occurring RNA molecule, and optionally 0 - 5 base pairs of the stem structure.
6. A method for obtaining information about the structure of a ribonucleic acid, which comprises: Obtaining a ribonucleic acid having at least 90% identity with SEQ ID NO:1; Replacing the residues corresponding to loop residues 14 - 17 (GAAA) in SEQ ID NO:1 with a heterologous segment of a nucleic acid having a length between 4 and 33 nucleotides to form a fusion ribonucleic acid molecule; Crystallizing the fusion ribonucleic acid molecule; Performing X - ray or electron crystallography techniques on the fusion ribonucleic acid molecule; and Observing the results of the X - ray or electron crystallography techniques to obtain information about the structure of the heterologous segment of the nucleic acid.
7. The method according to claim 6, wherein the fusion ribonucleic acid molecule is combined with a reagent that binds to the ribonucleic acid before the crystallographic analysis.
8. The method according to claim 7, wherein the reagent is a polynucleotide that hybridizes to the ribonucleic acid.
9. The method according to claim 7, wherein the crystallographic analysis includes comparison with a control sample lacking the reagent that binds to the ribonucleic acid.
10. The method according to claim 7, wherein a plurality of fusion ribonucleic acid molecules are combined with a plurality of reagents that bind to the ribonucleic acid before the X - ray or electron crystallography techniques.
11. The method according to claim 10, wherein at least two reagents are combined with the fusion ribonucleic acid molecule.
12. A method for performing crystallographic analysis on a polynucleotide, the method comprises: (a) Selecting a first polynucleotide, wherein the first polynucleotide comprises the polynucleotide sequence of a first miRNA; (b) Identifying the segment of the polynucleotide that forms a first loop region in the first miRNA; (c) Selecting a second polynucleotide, wherein the second polynucleotide comprises the polynucleotide sequence of a second miRNA; (d) Identify the polynucleotide segment that forms the first loop region in the second miRNA; (e) Form a fusion polynucleotide that is constructed such that the polynucleotide segment containing the first loop region on the first polynucleotide is replaced by the polynucleotide segment containing the first loop region on the second polynucleotide; and (f) Perform crystallographic analysis on the fusion polynucleotide to observe the three-dimensional structure of the fusion polynucleotide; Thereby performing crystallographic analysis of the polynucleotide, wherein the first miRNA is a miRNA having at least 90% sequence identity with: GGUUGCCGAAUCCGAAAGGUACGGAGGAACCGCUUUUUGGGGUUAAUCUGCAGUGAAGCUGCAGUAGGGAUACCUUCUGUCCCGCACCCGACAGCUAACUCCGGAGGCAAUAAAGGAAGGAG (SEQ ID NO:1), wherein: residues 14-17 (GAAA) of the ribonucleic acid are replaced by a heterologous segment of nucleic acid having a length between 4 and 33 nucleotides and containing the first loop region on the second polynucleotide.
13. The method according to claim 12, wherein: the first polynucleotide contains the sequence of SEQ ID NO:1; and / or the second miRNA contains a human miRNA.
14. The method according to claim 12, wherein the crystallographic analysis is an X-ray or electron crystallography technique.
15. The method according to claim 12, wherein the crystallographic analysis is performed in the presence of a reagent that binds to the fusion polynucleotide.
Citation Information
Patent Citations
High throughput ensemble-based docking and elucidation of three-dimensional structural confirmations of flexible biomolecular targets
US20110172981A1