Modified precursor of RNA interference-inducing nucleic acid molecule comprising structural motif promoting drosha cleavage

By inserting a GAG or GUG motif in the 3p strand of pri-miRNA precursors to form a bulge, the modified precursor enhances Drosha recognition, addressing inaccuracies in pri-miRNA structure prediction and improving RNA interference efficiency.

WO2025178406A1PCT designated stage Publication Date: 2025-08-28SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002495
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2025-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Current methods for predicting the secondary structures of pri-miRNAs are inaccurate, leading to uncertainties in understanding miRNA biogenesis, regulation, and evolution, as computational methods often misestimate stem lengths and apical loop structures, resulting in inconsistent processing efficiency and accuracy.

Method used

A modified precursor of RNA interference-inducing nucleic acid molecules is introduced, featuring a GAG or GUG motif inserted in the 3p strand near the Drosha cleavage site, forming a bulge, to enhance Drosha recognition and processing efficiency.

Benefits of technology

The modified precursor leads to higher expression levels of RNA interference-inducing nucleic acid molecules by optimizing Drosha cleavage, improving processing accuracy and homogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002495_28082025_PF_FP_ABST
    Figure KR2025002495_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a novel modified precursor of an RNA interference-inducing nucleic acid molecule, to a method for designing the modified precursor, and to a method for producing an RNA interference-inducing nucleic acid molecule. A modified precursor, according to one aspect, can be efficiently and homogeneously processed, and thus exhibits a higher expression level in cells, thereby being useful for design and regulation studies of RNA interference-inducing nucleic acid molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Modified precursor of RNA interference-inducing nucleic acid molecules containing structural motifs that promote DROSHA cleavage

[0001] It relates to a modified precursor of a novel RNA interference-inducing nucleic acid molecule.

[0002] MicroRNAs (miRNAs) are non-coding RNAs that are 20 to 24 nucleotides (nt) long. They form base pairs with target mRNAs through a seed sequence of 2 to 7 nt from the 5′ end, thereby inducing RNA interference (RNAi). RNA interference refers to a phenomenon in which double-stranded RNA specifically degrades mRNA at that base sequence, resulting in the inhibition of gene expression, or silencing of the corresponding gene. RNA interference can strongly suppress gene expression, and is widely used as a simple and effective gene expression suppression technique. The process of producing miRNA can also produce artificial small RNAs (siRNAs), which can be utilized in RNA interference technology.

[0003] The miRNA sequence is contained in a local hairpin of the pri-miRNA, and the pri-miRNA is processed by the Microprocessor complex. This complex consists of the RNase III enzyme Drosha and the cofactor DGCR8 (also called Pasha). Microprocessor cleaves the pri-miRNA hairpin approximately 13 bp from the base of the stem structure, releasing a short hairpin (pre-miRNA) of approximately 65 nt. After nuclear export, the RNase III enzyme Dicer cleaves the pre-miRNA approximately 22 bp from the end, forming a miRNA duplex. This miRNA duplex is loaded onto the Argonaute (AGO) protein, which aligns it in an orientation that favors either the 5′ strand (5p) or the 3′ strand (3p).

[0004] In this process, the Microprocessor acts as a key gateway, distinguishing genuine pri-miRNAs from thousands of transcripts. The Microprocessor examines various structural features of these transcripts to selectively recognize pri-miRNAs.

[0005] Pri-miRNA hairpins consist of an imperfect double-stranded RNA (dsRNA) stem containing an apical loop and a basal single-stranded RNA segment. Drosha interacts with the lower stem and the basal segment, while DGCR8 binds to the upper stem and the apical loop. Drosha and DGCR8 anchor at the basal and apical dsRNA-ssRNA junctions, respectively, accommodating a 35 ± 1 bp-long stem between the two junctions.

[0006] Microprocessors also recognize sequence motifs at specific locations relative to these structural elements. Drosha binds to the UG motif at the basal junction and the mismatched GHG (mGHG) motif of the lower stem, while DGCR8 interacts with the UGU motif at the apical junction. Serine / arginine-rich splicing factor 3 (SRSF3) acts as a cofactor to recognize the CNNC motif in the 3′ basal segment, thereby promoting processing. Negative regulators such as Lin28 have also been reported. Lin28 recognizes the UGAU and GGAG motifs in the apical loop of let-7 family members and inhibits processing by Drosha and Dicer. Because these sequence motifs are short and variable, it is believed that not only the primary sequence but also their relative positions to structural motifs play a crucial role in recognition.

[0007] Despite the crucial role of pri-miRNA structural elements, experimental studies on human pri-miRNAs and their variants are very limited. Consequently, the structures of most pri-miRNAs remain unknown. Currently, over 1,900 human miRNAs are registered in miRBase. Manual annotation of MirGeneDB (v1) yielded a list of 519 highly reliable miRNAs. However, high-throughput in vitro processing analysis revealed that only 758 of these human pri-miRNAs were significantly processed, while over 1,000 pri-miRNAs registered in miRBase were not processed, suggesting the possibility of misannotation. Significantly processed pri-miRNAs exhibited significant differences in cleavage efficiency and accuracy, which is not surprising given their diverse sequences and structures.

[0008] Although various computational methods have been developed to predict RNA structure, it remains uncertain whether these methods can accurately measure stem length and reliably determine the structures of apical loops and base regions. For example, previous predictions using Sfold and RNAfold suggested that more than 40% of human pri-miRNAs have stem lengths shorter than 30 nt or longer than 40 nt. However, this contradicts previous biochemical data showing a stem length of approximately 35 bp and structural studies of some individual pri-miRNAs. To further understand miRNA biogenesis, regulation, and evolution, it is crucial to elucidate pri-miRNA structures genomically and quantitatively.

[0009] In this study, we investigated the secondary structures of 476 high-fidelity human pri-miRNAs using SHAPE-MaP (selective 2′acylation analyzed by primer extension and mutational profiling), a high-throughput RNA structural analysis method. Combined with in vitro processing data, this study provides a powerful tool for clarifying the structural features of pri-miRNAs and elucidating the underlying mechanisms of miRNA processing regulation.

[0010] One aspect is to provide a modified precursor of an RNA interference-inducing nucleic acid molecule.

[0011] Another aspect is to provide RNA interference-inducing nucleic acid molecules generated from the above modified precursors.

[0012] Another aspect is to provide a nucleic acid molecule encoding the RNA interference-inducing nucleic acid molecule.

[0013] Another aspect is to provide an expression construct comprising a nucleic acid molecule encoding the RNA interference-inducing nucleic acid molecule.

[0014] Another aspect is to provide a vector comprising the above expression construct.

[0015] Another aspect provides a method for designing a modified precursor of an RNA interference-inducing nucleic acid molecule.

[0016] Another aspect provides a method for producing RNA interference-inducing nucleic acid molecules.

[0017] One aspect provides a modified precursor of an RNA interference-inducing nucleic acid molecule. The modified precursor can produce a nucleic acid molecule that induces RNA interference at a higher expression level.

[0018] Specifically, the modified precursor may be a precursor before modification having a stem-loop structure, in which a GAG or GUG motif is inserted. Here, the GAG ​​or GUG motif is inserted into the 3p strand of the region that interacts with Drosha within the stem of the precursor before modification, and the A or U base of the GAG ​​or GUG motif may form a bulge.

[0019] In one specific example, the RNA interference-inducing nucleic acid molecule may be, but is not limited to, miRNA, shRNA (short hairpin RNA), or siRNA (small interfering RNA), and may refer to any nucleic acid molecule capable of inducing RNA interference (RNAi), preferably a nucleic acid molecule including a region where the precursor interacts with Drosha.

[0020] The term "Drosha protein" in this specification refers to an RNase III family enzyme responsible for the first step of cleaving pri-miRNA into pre-miRNA in the nucleus during the miRNA biogenesis process, and acts by forming a Microprocessor complex together with DGCR8.

[0021] In one specific example, the "region that interacts with Drosha within the stem of the precursor before transformation" may mean a Drosha cleavage site.

[0022] The term "Drosha cleavage site" refers to the site where both RNA strands forming the stem structure of pri-miRNA are cleaved by the Drosha protein. Since the 5′-terminal cleavage site and the 3′-terminal cleavage site are separated by 2 bp, the relative distance from the cleavage site refers to the distance from the 5′-terminal cleavage site.

[0023] In one specific example, the modified precursor may comprise a GAG or GUG motif between two consecutive positions from -9 bp to -12 bp relative to the Drosha cleavage site of the 5p strand and two positions of the complementary 3p strand. For example, the modified precursor may comprise a GAG or GUG motif between any one of -9 bp and -10 bp, -10 bp and -11 bp, and -11 bp and -12 bp relative to the Drosha cleavage site of the 5p strand and a position of the complementary 3p strand. More specifically, the modified precursor may comprise a GAG or GUG motif between positions from -10 bp and -11 bp relative to the Drosha cleavage site of the 5p strand and a position of the complementary 3p strand.

[0024] In this specification, the positions of nucleotides are indicated by relative numbers. The position in the direction of the apical loop relative to the Drosha cleavage position of the 5p strand is indicated by +, and the position in the direction of each base segment is indicated by -.

[0025] That is, the modified precursor may contain a GAG or GUG motif between two consecutive positions 9 to 12 bp away from the Drosha cleavage site of the 5p strand in the 5'-terminal direction and two positions of the complementary 3p strand. More specifically, the modified precursor may contain a GAG or GUG motif between positions 10 and 11 bp away from the Drosha cleavage site of the 5p strand in the 5'-terminal direction and positions of the complementary 3p strand.

[0026] The above GAG ​​or GUG motif may be inserted between two base pairs in the 3p strand, so that the A or U base of the GAG ​​or GUG motif forms a bulge.

[0027] In one embodiment, the variant precursor may include a mismatched base pair at a position -6 bp, a position +1 bp, or both, relative to the Drosha cleavage site of the 5p strand. That is, the variant precursor may additionally include a mismatched base pair at a position 6 bp away from the Drosha cleavage site of the 5p strand in the direction of the base segment, 1 bp away from the Drosha cleavage site in the direction of the apical loop, or both.

[0028] In one specific example, the precursor of the RNA interference-inducing nucleic acid molecule may be a stem-loop structure and may include an apical loop and a basal segment.

[0029] As used herein, the term "apical loop" refers to a loop region located at the top of the stem structure in a precursor (pre-miRNA or pri-miRNA) of an RNA interference-inducing nucleic acid molecule. Since Dicer cleaves the double-stranded portion of a pre-miRNA from which the apical loop has been removed to produce a mature miRNA, this loop is involved in Dicer's processing efficiency.

[0030] As used herein, the term "basal segment" refers to the lower portion of the stem structure recognized by the Drosha / DGCR8 complex in a precursor (pre-miRNA or pri-miRNA) of an RNA interference-inducing nucleic acid molecule. The basal segment forms a ssRNA region of 5 to 7 bp in length, and Drosha recognizes the basal segment and cleaves it approximately 13 bp away from the junction of the basal segment and the stem structure in the direction of the apical loop.

[0031] In one embodiment, the loop may comprise 5 to 50 nucleotides. The loop may comprise 5 to 50, 10 to 50, 11 to 50, 12 to 50, 13 to 50, 13 to 20, or 13 to 15 nucleotides.

[0032] In one embodiment, the stem may comprise 20 to 50 base pairs. The stem may comprise 20 to 50, 30 to 40, 32 to 38, or 34 to 36 base pairs.

[0033] In one specific embodiment, the stem may comprise 10 or fewer mismatched base pairs. The stem may comprise 10 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer mismatched base pairs. A mismatched base pair refers to bases that are not completely complementary to each other pairing with each other.

[0034] In one embodiment, the base segment may comprise 1 to 15 nucleotides. The base segment may comprise 1 to 15, 1 to 10, 1 to 5, 2 to 10, or 2 to 4 nucleotides.

[0035] In one embodiment, the base segment may comprise 5 to 30 base pairs from the start base pair position of the 5p strand to the Drosha cleavage site of the 5p strand. The base segment may comprise 5 to 30, 7 to 20, or 10 to 15 base pairs from the start base pair position of the 5p strand to the Drosha cleavage site of the 5p strand.

[0036] Preferably, the modified precursor may comprise one or more of the following characteristics:

[0037] (a) the loop comprises 5 to 50 nucleotides;

[0038] (b) the stem contains 20 to 50 base pairs;

[0039] (c) the stem contains 10 or fewer mismatched base pairs;

[0040] (d) the base segment of the 5p strand or the 3p strand comprises 1 to 15 nucleotides; and

[0041] (e) 5 to 30 base pairs from the start base pair position of the 5p strand to the Drosha cleavage site of the 5p strand.

[0042] More preferably, the modified precursor may comprise one or more of the following characteristics:

[0043] (a) the loop comprises 13 to 15 nucleotides;

[0044] (b) the stem contains 34 to 36 base pairs;

[0045] (c) the stem contains no more than one mismatched base pair;

[0046] (d) the base segment of the 5p strand or the 3p strand comprises 2 to 4 nucleotides; and

[0047] (e) Contains 10 to 15 base pairs from the start base pair position of the 5p strand to the Drosha cleavage site of the 5p strand.

[0048] In one specific example, the modified precursor may further comprise at least one selected from the group consisting of a UG motif, an mGHG motif, a UGUG motif, and a CNNC motif.

[0049] The symbols used to express the sequence in the above motif follow the IUPAC code. For example, the H above means that it corresponds to one of A (adenine), C (cytosine), or U (uracil) and not G (guanine), and the N above means that it corresponds to one of A, C, G, or U.

[0050] In one embodiment, the modified precursor may additionally comprise a UG motif within the basal segment.

[0051] In one embodiment, the modified precursor may further comprise an mGHG motif within the stem.

[0052] In one embodiment, the modified precursor may additionally comprise a UGUG motif or a CNNC motif within the apical loop.

[0053] In one specific example, the modified precursor can promote cleavage of Drosha.

[0054] In one specific example, the precursor of the RNA interference-inducing nucleic acid molecule may be a pri-miRNA (primary miRNA). As used herein, the term "pri-miRNA" refers to a primary precursor produced during the biosynthesis of miRNA in eukaryotic cells, and includes a stem-loop structure with a hairpin-like structure.

[0055]

[0056]

[0057] Another aspect is to provide RNA interference-inducing nucleic acid molecules generated from the above modified precursors.

[0058] Another aspect is to provide a nucleic acid molecule encoding the RNA interference-inducing nucleic acid molecule.

[0059] Another aspect is to provide an expression construct comprising a nucleic acid molecule encoding the RNA interference-inducing nucleic acid molecule.

[0060] Another aspect is to provide a vector comprising the above expression construct.

[0061] As used herein, the term "expression construct" refers to a construct of genetic material containing a coding sequence and sufficient regulatory information to direct proper transcription of the coding sequence in a recipient cell, in vivo, and / or in vitro. The term "expression construct" may be used interchangeably with terms such as "expression cassette." The expression construct may be inserted into a vector for targeting to a desired host cell and / or subject.

[0062] The term "vector" as used herein refers to a genetic construct containing a nucleic acid molecule encoding a target RNA interference-inducing nucleic acid molecule operably linked to suitable regulatory sequences so as to enable expression of the target RNA interference-inducing nucleic acid molecule in a suitable host. The term "operably linked" as used herein means that the nucleic acid molecule sequence is functionally linked to a promoter sequence that initiates and mediates transcription.

[0063] The above vector, after being introduced into a suitable host cell, can replicate or function independently of the host genome, or can be integrated into the genome itself. The vector is not particularly limited as long as it can be expressed in the host cell, and any vector known in the art can be used to introduce the vector into the host cell. Examples of commonly used vectors include plasmids, cosmids, viruses, and bacteriophages, either in a natural or recombinant state.

[0064] What has been described in the above variant precursor also applies to the above nucleic acid molecule, the above expression construct and the above vector.

[0065]

[0066] Another aspect provides a method for designing a modified precursor of an RNA interference-inducing nucleic acid molecule.

[0067] Specifically, the method may comprise a step of inserting a GAG or GUG motif, wherein the A or U base forms a bulge, into the 3p strand of the region interacting with Drosha within the stem of the precursor before modification, as a method of designing a modified precursor capable of expressing more of the RNA interference-inducing nucleic acid molecule.

[0068] In one embodiment, the method may comprise inserting a GAG or GUG motif, wherein the A or U base forms a bulge, between two consecutive positions 9 to 12 bp in the 5'-end direction from the Drosha cleavage site of the 5p strand of the precursor before modification and two positions of the complementary 3p strand.

[0069] In one embodiment, the method may comprise inserting a GAG or GUG motif, wherein the A or U base forms a bulge, between two positions of the complementary 3p strand and positions 10 and 11 bp in the 5'-end direction from the Drosha cleavage site of the 5p strand of the precursor before modification.

[0070] The descriptions of the above variant precursor, the nucleic acid molecule, the expression construct and the vector also apply to the above design method.

[0071]

[0072] Another aspect provides a method for producing RNA interference-inducing nucleic acid molecules.

[0073] Specifically, the method may comprise the following steps as a method for producing more of the RNA interference-inducing nucleic acid molecule:

[0074] A step of designing a modified precursor of an RNA interference-inducing nucleic acid molecule;

[0075] producing a modified precursor of an RNA interference-inducing nucleic acid molecule; and

[0076] A step for producing an RNA interference-inducing nucleic acid molecule from a modified precursor of an RNA interference-inducing nucleic acid molecule.

[0077] In one specific example, the step of producing a modified precursor of an RNA interference-inducing nucleic acid molecule may be synthesizing RNA through biological synthesis, enzymatic synthesis, or chemical synthesis, and may include, without limitation, any RNA synthesis method known in the art, and may also be used in combination with other synthesis methods as needed.

[0078] The above biological synthesis refers to RNA synthesis through cellular expression, and specifically may include introducing a nucleic acid molecule encoding a precursor of an RNA interference-inducing nucleic acid molecule into a cell and producing a modified precursor of the RNA interference-inducing nucleic acid molecule through cellular expression.

[0079] The above enzymatic synthesis refers to synthesis through in vitro transcription (IVT), and specifically, may include preparing a nucleic acid molecule encoding a precursor of an RNA interference-inducing nucleic acid molecule as a template and enzymatically producing a modified precursor of the RNA interference-inducing nucleic acid molecule using RNA polymerase.

[0080] The above chemical synthesis is based on a synthetic method of sequentially linking and adding oligonucleotides, and may include sequentially synthesizing a desired sequence in the 5' to 3' direction of RNA to produce a modified precursor of an RNA interference-inducing nucleic acid molecule. For example, the above chemical synthesis may be solid-phase synthesis, etc.

[0081] The contents described in the above modified precursor, the nucleic acid molecule, the expression construct, the vector and the design method also apply to the above production method.

[0082]

[0083] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the following detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. In describing the present invention, detailed descriptions of related known technologies will be omitted if they are deemed to obscure the gist of the present invention.

[0084] Modified precursors of RNA interference-inducing nucleic acid molecules according to one aspect can be efficiently and homogeneously processed, resulting in higher expression levels of RNA interference-inducing nucleic acid molecules within cells, and thus are useful for research in designing and controlling RNA interference-inducing nucleic acid molecules.

[0085] Figure 1 is a schematic diagram illustrating the structural elements of a standard pri-miRNA. The bold line represents the miRNA duplex, and the arrow indicates the DROSHA cleavage site.

[0086] Figure 2 is a schematic diagram of the SHAPE-MaP procedure.

[0087] Figure 3 shows the accuracy of the SHAPE-based structure for a control RNA. Sensitivity is the proportion of known base pairs that are correctly predicted, and the positive predictive value (PPV) is the proportion of predicted base pairs that match the known structure.

[0088] Figure 4 is a heatmap showing the SHAPE reactivity of the 5p strand and the 3p strand in 476 pri-miRNAs. The numbers are assigned based on the 5′ end of the 5p miRNA or the 3′ end of the 3p miRNA, and the average SHAPE reactivity at each position is represented as a bar graph at the bottom.

[0089] Figure 5 is a diagram comparing the ratio of base pairs in pri-miRNA that are present in the sequence-based structure but not in the SHAPE-based structure.

[0090] Figure 6 illustrates the structural differences between the SHAPE-based and sequence-based structures. The base pair frequency of the SHAPE-based structure at each position is calculated by subtracting the base pair frequency of the sequence-based structure from the base pair frequency of the SHAPE-based structure. The sequence-based structure was predicted using RNAstructure4, RNAfold, mfold, CONTRAfold, and Eternafold.

[0091] Figure 7 shows the distribution of stem lengths derived from SHAPE-based and sequence-based structures. This was created by comparing pri-miRNAs (n = 414) analyzed in both studies.

[0092] Figure 8 is a schematic diagram showing the cross-validation of SHAPE-MaP and high-throughput in vitro processing results.

[0093] Figure 9 shows the average cleavage efficiency and cleavage homogeneity of each group as a result of in vitro DROSHA processing after grouping pri-miRNAs according to stem length measured from SHAPE-based or sequence-based structures. The number of pri-miRNAs in each group is indicated in parentheses.

[0094] Figure 10 is a diagram showing the length distribution of the upper and lower stems of pri-miRNA.

[0095] Figure 11 shows the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs according to the length of the lower or upper stem, as a result of in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0096] Figure 12 is a diagram showing the distribution of the number of mismatched base pairs in the stem.

[0097] Figure 13 shows the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs by the number of mismatched base pairs or the presence of mismatched base pairs (at the -6 or +1 position) as a result of in vitro DROSHA treatment. During the analysis, a mismatch at the -6 position, located in the middle of the mGHG motif, was excluded due to its association with accelerated processing. The number of pri-miRNAs in each group is indicated in parentheses.

[0098] Figure 14 is a diagram showing the locations of mismatched base pairs within the pri-miRNA stem, where the primary mature miRNA is derived from the 5p strand or the 3p strand. The numerical positions are indicated relative to the 5′ end (arrow) of the 5p miRNA.

[0099] Figure 15 is a diagram showing the distribution of the number of bulged bases from the stem, and the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs according to the number of bulged bases as a result of in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0100] Figure 16 shows the frequency of overhangs in the 5p and 3p strands predicted using SHAPE, and the results of in vitro DROSHA processing according to the presence or absence of overhangs at specific positions measured from the SHAPE-based structure. Arrows indicate the 5′ end of the 5p miRNA as the reference point and the 3p overhang between positions -10 and -11 (-10 / -11). Bars represent adjusted p-values ​​analyzed by the Mann-Whitney U test for differences in cleavage efficiency or cleavage homogeneity depending on the presence or absence of overhangs. Bonferroni correction was applied for multiple testing correction, and the dotted line indicates an adjusted p-value of 0.05.

[0101] Figure 17 is a diagram showing the frequency of protrusions in the 5p strand and 3p strand predicted using RNAstructure(Fold) and the results of in vitro DROSHA processing according to the presence or absence of protrusions at specific positions measured from the sequence-based structure.

[0102] Figure 18 is a schematic diagram showing the sequence preference for the 3p overhang at positions -10 / -11. The left schematic diagram shows the base sequence within the 3p overhang at positions -10 / -11, with the height of each letter being proportional to the nucleotide frequency, and the right schematic diagram shows the nucleotide pair frequency within the 3p overhang at positions -10 / -11, with the first letter in each pair representing the 5p nucleotide and the second letter representing the 3p nucleotide.

[0103] Figure 19 shows the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs into GUG bulge, GAG bulge, Else bulge, and No bulge bulge groups as a result of in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0104] Figures 20, 21, and 22 illustrate the quantification of pre-miRNA products following incubation with pri-miRNAs with or without a prominent GWG motif. When replacement processing occurred, pre-miRNA (Pre) and the replacement processing product (Alt) were measured separately. The amount of product was normalized by the input amount and then further normalized by the amount of pre-miRNA generated from the control group (pri-miRNA without a prominent GWG motif or wild-type pri-miRNA).

[0105] Figure 20 shows the results of a comparative experiment between artificial pri-miRNA with bulge (bGWG) and artificial pri-miRNA without bulge (No bulge). The bar graph shows the quantification of the product after 1 hour of incubation.

[0106] Figure 21 shows the results of a comparative experiment between Pri-mir-99a (WT) and a mutant with a protrusion removed. The bar graph shows the quantification of the product after 0.5 hours of incubation.

[0107] Figure 22 shows the results of a comparative experiment between Pri-mir-183 (WT) and a mutant with a protrusion removed (Mut). The bar graph shows the quantification of the product after 2 hours of incubation.

[0108] Figure 23 shows miRNA accumulation after transfection with wild-type (WT) and a prominent GWG-mutant (Mut) pri-miRNA. RNA was harvested 48 hours after pri-miRNA transfection into HEK293E cells, and TaqMan qPCR signals were first normalized to the miR-1 signal expressed by a different promoter within the plasmid and then further normalized to the WT signal. Biological replicates were performed three times.

[0109] Figure 24 shows the distribution of apical loop sizes and the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs according to apical loop size as a result of in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0110] Figure 25 shows the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs according to the distribution of base segment length and the distribution of base segment length as a result of in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0111] Figure 26 shows the structural context of the basal UG motif and its position frequency within the 5p strand, and the average cleavage efficiency and cleavage homogeneity of each group after grouping pri-miRNAs according to the basal UG structural context and performing in vitro DROSHA treatment. The number of pri-miRNAs in each group is indicated in parentheses.

[0112] Figure 27 is a diagram showing the processing results of artificial pri-miRNAs containing UG motifs in different structural contexts.

[0113] Figure 28 is a schematic diagram showing the general structure of pri-miRNA and the structural and sequence features of pri-miRNA optimized for DROSHA cleavage.

[0114] Figure 29 is a diagram showing the proportion of pri-miRNA groups sharing optimal structural combinations and the cleavage efficiency, cleavage homogeneity, cellular expression, and sequence conservation for each group.

[0115] Figure 30 is a schematic diagram of the design of conventional and improved TRIPZ shRNAs.

[0116] Figure 31 is a graph measuring miRNA accumulation and target luciferase activity after transfection of HEK293E cells with conventional and improved TRIPZ shRNAs. TaqMan qPCR signals were first normalized to U6 snRNA levels and then further normalized to unmodified TRIPZ signals. TaqMan qPCR was performed four times, and luciferase assays were performed three times in biological replicates.

[0117] The following examples are provided for more detailed description. However, these examples are provided solely to illustrate one or more specific examples, and the scope of the present invention is not limited to these examples.

[0118]

[0119] Example 1. Improved structural analysis of human pri-miRNA using SHAPE-MaP.

[0120] To analyze the structure of pri-miRNA (Fig. 1), SHAPE-MaP analysis was performed using 1-methyl-7-nitroisatoic anhydride (1M7). 1M7 selectively induces 2′-hydroxyl acylation in unstructured regions, and the acylated residues induce base substitutions or insertions / deletions (indels) during reverse transcription. The mutation rate of each nucleotide was measured using high-throughput sequencing, and the reactivity of 1M7, called "SHAPE reactivity," was quantified based on this. This reactivity was used for RNA secondary structure modeling.

[0121] SHAPE-MaP experiments were performed on 519 high-confidence human pri-miRNAs curated by MirGeneDB (v1). Pri-miRNAs were synthesized in vitro using 125 nt minigene constructs containing a central miRNA hairpin. These constructs have been used in previous high-throughput in vitro processing experiments to assess processing efficiency and homogeneity, allowing for seamless integration of structural and functional data (Fig. 2).

[0122]

[0123] Three well-defined RNAs were included as structural probe controls. The RNAs used as controls were: human U1 small nuclear RNA (snRNA), yeast tRNAAsp, and hepatitis C virus internal ribosome binding site (HCV IRES) domain II.

[0124] RNA pools, including control RNA, were incubated with 1M7. To ensure reliable quantification, 476 pri-miRNAs with more than 500 sequencing reads were selected. The resulting SHAPE reactivity was used as a constraint (Δ) in structural modeling, along with the underlying sequence information.

[0125] The control RNA showed high SHAPE reactivity (>0.4, Δ>0) in the known single-stranded RNA region and low reactivity in the structured region, and the SHAPE-based secondary structure model of the control RNA was very similar to the previously known structure (Fig. 3).

[0126]

[0127] Next, we analyzed the overall SHAPE reactivity pattern of the pri-miRNA. The reactivity data were aligned to the DROSHA cleavage site. A low-reactivity region of approximately 35 nt in length was observed on both the 5p and 3p strands, surrounded by a high-reactivity region. The average reactivity of the 5p strand showed marked changes at positions -13 and +22, which correspond to the basal and apical ss-dsRNA junctions, respectively. Similarly, the average reactivity of the 3p strand showed similar changes at the corresponding positions. Overall, these reactivity patterns were well consistent with the current understanding of the pri-miRNA structure (Fig. 4).

[0128]

[0129] 1.2 Comparison of SHAPE-based and sequence-based structures

[0130] Structures modeled using SHAPE reactivity (SHAPE-based structures) differed significantly from structures predicted solely from sequence information (sequence-based structures). The sequence-based structures showed more than 10% base-pair discrepancies compared to the SHAPE-based structures for most pri-miRNAs (61%) (Figure 5).

[0131] In general, applying SHAPE data revealed a tendency toward base-pairing, which was more pronounced in the apical loop and basal segments than in the stem. This pattern was consistently observed when comparing the results using various structure prediction software, including RNAfold, mfold, CONTRAfold, and Eternafold (Fig. 6).

[0132]

[0133] Furthermore, we compared the in vitro SHAPE-MaP data with previous in vivo SHAPE-MaP data for the mir-17-92 transcriptome (Ratnadiwakara M et al., EMBO Rep. 2023 Jul 5;24(7):e56021. doi: 10.15252 / embr.202256021. Epub 2023 Jun 12. PMID: 37306233; PMCID: PMC10328067.). The structural models based on both datasets showed high consistency (average sensitivity and positive predictive value of 96%), whereas the sequence-based structure showed lower consistency (89%) with the in vivo SHAPE model. This was due to mispredicted base pairs lowering the positive predictive value. This analysis reaffirms that current computational methods tend to overpredict RNA folding and underestimate unstructured regions. The differences in SHAPE reactivity between in vitro and in vivo at some locations may reflect protein interactions within cells. This was also highly consistent with models based on previously published chemical and enzymatic probing data (Krol, Jacek et al., Journal of Biological Chemistry, Volume 279, Issue 40, 42230-42239).

[0134] In summary, experimental data such as SHAPE-MaP can improve structure prediction models and are particularly useful for investigating the boundaries between structural elements and non-structural regions important for pri-miRNA processing.

[0135]

[0136] Example 2. Confirmation of the structural characteristics and functional effects of the stem.

[0137] The structural features of pri-miRNA were investigated as follows using the improved model of Example 1.

[0138]

[0139] 2.1 Confirmation of modal stem length and functional impact

[0140] First, the most frequent stem length was found to be 35 bp (SHAPE-based). Most pri-miRNAs contained stems of 35 ± 1 bp, accounting for 57%. This result contradicts previous sequence-only studies, which suggested that pri-miRNA stems are often very short (≤30 bp) or very long (≥40 bp) (Figure 7).

[0141] Noting these discrepancies, we re-evaluated the functional relevance of stem length using high-throughput in vitro processing data. After grouping pri-miRNAs by stem length, we identified a group of pri-miRNAs with significantly higher cleavage efficiency and homogeneity compared to other groups, thereby deriving the optimal stem length (Fig. 8). Pri-miRNAs with 35 ± 1 bp stems exhibited significantly higher processing efficiency and accuracy than pri-miRNAs with shorter or longer stems. In contrast, the same analysis using sequence-based structure revealed no significant differences between miRNA groups based on stem length (Fig. 9).

[0142] These results validate and extend previous mutagenesis studies on pri-mir-16-1, pri-mir-30a, and pri-mir-125a, and confirm that 35 ± 1 bp is the functionally optimal stem length.

[0143]

[0144] 2.2 Confirmation of the most frequent lengths and functional impact of the lower stem and upper stem

[0145] The stem of pri-miRNAs can be divided into lower and upper stems based on the DROSHA cleavage site. The most frequent lengths of the lower and upper stems were 13 bp and 22 bp, respectively (Fig. 10), which are consistent with the optimal values ​​previously reported for pri-mir-16-1 and pri-mir-30a. Some pri-miRNAs (e.g., pri-mir-455, pri-mir-490) have very short lower stems (≤6 bp) and may undergo non-canonical processing by the microprocessor, as recently proposed. However, in general, pri-miRNAs with lower stems as long as 13 bp exhibited the highest cleavage efficiency and homogeneity, whereas those with stems shorter than 11 bp were very poorly processed (Fig. 11, left).

[0146] Pri-miRNAs with 22-bp upper stems exhibited slightly higher processing quality than other cases, but there was no statistically significant difference compared to those with very long or very short stems (Fig. 11, right). Therefore, the apical element may be less important than the basal region, or may only affect a small group of pri-miRNAs. This observation is consistent with previous studies showing that the core region of the Microprocessor complex interacts closely with the lower stem and basal segments.

[0147]

[0148] 2.3 Identifying mismatches and bulges and their functional impact

[0149] Most stems are imperfect double-stranded structures containing multiple mismatched base pairs and / or bulges of varying locations and sizes. More than 98% of all pri-miRNAs have at least one mismatched base pair in their stem, the most common being 4 bp (Figure 12). Comparison with processing data revealed a negative correlation between the number of mismatched base pairs and processing efficiency, indicating that mismatched base pairs generally interfere with processing (Figure 13).

[0150] Nonetheless, we observed that mismatched base pairs occurred more frequently at certain positions than at others. The mismatched base pair at position -6 is located in the center of the mGHG motif, which is known to promote accurate processing. High-throughput data showed that pri-miRNAs with mismatched base pairs at position -6 exhibited significantly higher cleavage efficiency and homogeneity than those without. These results support and extend previous observations that many pri-miRNAs contain mismatched base pairs at position -6 (Figure 14, top).

[0151] In contrast, mismatched base pairs at the +1 position did not appear to correlate with pri-miRNA processing, and rather likely provide an advantage in later steps, such as strand selection of the 5p miRNA. Supporting this, mismatched base pairs at the +1 position were observed preferentially in pri-miRNAs where the 5p strand serves as a guide (Fig. 14, bottom).

[0152]

[0153] Example 3. Confirmation of the bulged GWG motif

[0154] 3.1 Cross-analysis with high-throughput data

[0155] Bulges are less common than mismatched base pairs, occurring approximately 3.7% of the time at individual sites. However, most pri-miRNAs (82%) contain at least one bulge within the stem (including both strands). Cross-analysis with high-throughput data revealed a negative correlation between the number of bulged bases and processing quality (Figure 15). Previous studies have suggested that bulges can impair processing yield and / or fidelity in some pri-miRNAs. The current analysis, utilizing hundreds of pri-miRNAs, confirmed that bulges generally impede processing.

[0156] Interestingly, bulges appeared to occur more frequently at specific locations. To investigate the significance of these specific bulges, pri-miRNAs were divided based on the presence or absence of bulges at individual locations on each strand, and cleavage efficiency and homogeneity were statistically compared. We found that bulges between positions -10 and -11 on the 3p strand were significantly associated with higher cleavage efficiency and homogeneity, suggesting a potential positive role in processing (Figure 16).

[0157] The abundance of this 3p bulge and its association with processing cannot be detected solely from sequence-based structures without experimental data (Fig. 17), as it is difficult to accurately predict the bulge.

[0158]

[0159] 3.2 Sequence enrichment analysis

[0160] Sequence abundance analysis revealed that the bulged base between positions -10 and -11 was primarily U (uracil), followed by A (adenine). The adjacent -10 position was strongly enriched in G (guanine) in the 3p strand and C (cytosine) in the 5p strand, forming a 5′′ pair. Position -11 showed a pattern of being slightly enriched in G in the 3p strand and C in the 5p strand (Fig. 18).

[0161] To investigate the importance of these sequence preferences, pri-miRNAs were grouped as follows:

[0162] - GUG: When the U bulge is surrounded by guanine on both sides

[0163] - GAG: When a bulge is surrounded by guanine on both sides

[0164] - Else: bulge that does not contain GUG or GAG, No bulge: if there is no bulge

[0165] The GUG and GAG groups exhibited high cleavage efficiency and homogeneity, while the remaining groups showed no significant differences compared to pri-miRNA without bulges (Fig. 19). This motif was designated "bulged GWG (bGWG)," where W represents U (uracil) or A (adenine).

[0166]

[0167] 3.3 In vitro treatment analysis

[0168] To confirm the causality between the bGWG motif and processing, we constructed panels of artificial or natural pri-miRNAs containing or lacking bGWG and performed in vitro processing assays (Figs. 20–22). Positions with bulges or deletions are indicated in bold.

[0169] - Pri-watson WT: Sequence number 1

[0170] AACUUGGCCCUUAAAAAAAAAAAAAAAA

[0171] - Pre-watson with no bulge: 서열번호 2

[0172] AACUUGGCCGUUAAAAAAAAAAAAAAAAAAAAAC

[0173] - First-99a WT: 서열번호 3

[0174] GGCCCAUGCAAGAUGUUGCCAUGGCAUAAACCCGUAGAUCCGAUCUUGUGGUGGUGGACCGCAAGCUCGCUUCUAUGGUCUGUGUGUGUGGUGGUGAUCUGACAAAAUGCUAUACAG

[0175] - First-99th Mutant(no bulge): 서열번호 4

[0176] GGCCCAUGCAAGAUGUUGCCAUGGCAUAAACCCGUAGAUCCGAUCUUGUGGUGGUGGACCGCCAAGCUCGCUUCUAUGGUCUGUGUGGUGGUGGUAAUCUGACAAAAUGCUAUACAG

[0177] - First-183 WT: 서열번호 5

[0178] CAGGCCGCAGUCUCCAUGAUGGCACUCCGGAAUUCCACUACAGAGAGCACUCCGAACAGGGCCUCCCGA

[0179] - First-183 Mutant(no bulge): 서열번호 6

[0180] CAGGCCGCAGAGUGUGACUCCUGUUCUGUGUAUGGCACUGGUAGAAUUCACUGUGAACAGUCUCAGUCAGUGAAUUACCGAAGGGCCAUAAACAGAGCAGGACAGAUCCACGAGGGCCUCCGGA

[0181]

[0182] A simple artificial hairpin, pri-watson, was processed more efficiently when it included the bGWG motif (Fig. 20). For the complex-structured natural pri-mir-99a, bulge removal reduced processing homogeneity and increased alternative products, but overall processing efficiency remained largely unchanged (Fig. 21). Another natural pri-miRNA, pri-mir-183, showed a significant decrease in both processing efficiency and homogeneity upon bulge removal (Fig. 22). Thus, although the effects of the bGWG motif vary depending on the structural context, we confirmed that the bGWG motif positively influences the efficiency and / or accuracy of in vitro processing.

[0183]

[0184] Additionally, ectopic expression experiments were performed to investigate the intracellular relevance of the bGWG motif. Pri-mir-99a, pri-mir-183, and their mutants containing or deleting bGWG were transfected into HEK293E cells, and the levels of mature miRNA were measured by reverse transcription-quantitative PCR (RT-qPCR). Consistent with the in vitro data, pri-miRNAs containing bGWG produced higher levels of mature miRNA (Figure 23), confirming that the bGWG motif is a positive factor in miRNA maturation.

[0185]

[0186]

[0187] Example 4. Confirmation of the apical loop structure

[0188] The apical loop plays a crucial role in miRNA regulation by providing specific binding sites for regulatory factors. SHAPE-based structural analysis identified points where local SHAPE reactivity contrasts sharply with adjacent sites, clearly defining the locations of the apical junction and the terminal loop.

[0189] Analysis results showed that loop sizes varied, but loops measuring 12–13 nt were the most common, accounting for 27% of the total. A positive correlation was observed between loop size and cleavage efficiency and homogeneity, but saturation occurred at 13 nt, indicating that a loop measuring 13 nt was sufficient to achieve maximum processing efficiency (Fig. 24).

[0190]

[0191] Some pri-miRNAs possessed exceptionally large loops (≥20 nt), suggesting additional functions beyond simply facilitating processing. Pri-miRNAs with these large loops include conserved members of the let-7 family, whose apical loops serve as binding sites for the Lin28 protein, a conserved protein that inhibits let-7 maturation. Lin28 contains two RNA-binding domains: a zinc-finger domain (ZFD) that specifically recognizes GGAG motifs and a cold-shock domain (CSD) that prefers UGAU motifs, although with less specificity.

[0192] Based on these sequence motifs, previous studies (Ustianenko et al., Lin28 Selectively Modulates a Subclass of Let-7 MicroRNAs, Mol Cell, 71 (2018), pp.271-283. e5) divided human let-7 miRNAs into two categories:

[0193] - CSD+ group (type I): Contains both motifs (GGAG, UGAU).

[0194] - CSD- group: A group that does not have the UGAU motif required for CSD binding.

[0195] SHAPE data also allowed analysis of the structural context of these sequence motifs, and the results showed that the UGAU motif is located in the loop portion of a small “sub-hairpin” within the apical loop, whereas the GGAG motif is always located in an unstructured region near the apical junction. This is consistent with the structural model of the mouse Lin28-let-7 precursor complex revealed by X-ray crystallography.

[0196]

[0197] Additionally, SHAPE data showed that the CSD group could be divided into two subgroups based on unique structural features:

[0198] - Unstructured loops (type II): let-7a-2, let-7c, let-7e.

[0199] - Structural loops (type III): let-7a-1, let-7a-3, let-7f-2.

[0200] The binding affinity of Lin28 measured in previous studies differed between these groups (type II > type III) and showed a negative correlation with the structural complexity of the loop. This suggests that the structure may affect Lin28 binding not only to the key sequences of the apical loop but also to the core sequences. In particular, because CSD binds RNA with low sequence specificity, it is possible that CSD may interact more easily with the flexible loop of type II pri-let-7 than with the structural loop of type III pri-let-7.

[0201]

[0202] Example 5. Identification of the basal segment and UG motif

[0203] The SHAPE-MaP model revealed that the single-stranded regions of the 5′ and 3′ base segments are typically short (≤4 nt), and that the 5′ base segment is generally not correlated with processing, but can interfere with processing when excessively long (≥20 nt). In contrast, the 3′ base segment appears to require a length of at least 5 nt to support efficient and homogeneous processing (Fig. 25). Therefore, many pri-miRNAs have suboptimal base segments and may require auxiliary factors such as SRSF3 to interact with the 3′ base CNNC motif.

[0204] The basal UG motif is a sequence element located at positions -14 to -13 relative to the DROSHA cleavage site, typically located at the basal junction. Using the SHAPE-MaP model, we identified pri-miRNAs with UG sequences located in either single- or double-stranded regions, and examined the possibility that the structural context of the UG motif regulates its activity. The UG motif structures were grouped into three categories:

[0205] - UGdd: Both U and G are double-stranded

[0206] - UGsd: U is single strand, G is double strand

[0207] - UGss: Both U and G are single stranded

[0208] Pri-miRNAs with UGdd or UGsd were processed more efficiently and homogeneously than those without UG motifs, but those with UGss did not show a significant difference from those without UG motifs. In human pri-miRNAs, UGdd and UGsd were relatively enriched at positions -14 to -13, whereas UGss was not (Fig. 26).

[0209] Additionally, since pri-miRNAs with UGdd or UGsd have longer sub-stems than those with UGss, we reanalyzed artificial pri-miRNA data with the same sub-stem length to control for bias due to sub-stem length. This dataset contains 20,049 artificial pri-miRNAs containing random sequences at the junctions, and the processing efficiency was measured by deep sequencing the products after a short incubation with the recombinant Microprocessor. The analysis results showed that artificial RNAs with UGdd and UGsd at positions -14 and -13 were processed more efficiently than RNAs without UG motifs, but RNAs with UGss did not show a significant difference from RNAs without UG motifs (Fig. 27).

[0210] In summary, we demonstrate that the UG sequence functions effectively in a specific structural context (where guanine forms base pairs).

[0211]

[0212]

[0213] Example 6. shRNA design

[0214] Based on the optimal structural and sequence features observed in this study, shRNA design was improved.

[0215]

[0216] 6.1 Analysis of the structure and function of pri-miRNA

[0217] The structural and functional relationship of human pri-miRNA was analyzed through the SHAPE experiment, and the functionally optimal features were determined through the analysis, as shown in Figure 28:

[0218] - Abortive loop of 13 nt or more

[0219] - 35 ± 1 bp stem and up to 1 mismatched base pair

[0220] - A 13 bp long lower stem containing a bulged GWG motif between positions -10 and -11.

[0221] - 3′ base segment of 5 nt or more.

[0222]

[0223] Figure 29 shows the proportion of pri-miRNAs possessing specific combinations of five representative structural features. While over 90% of all pri-miRNAs possess at least one optimal feature, only 1.3% (6 out of 476) possessed all five optimal features.

[0224] Pri-miRNAs with multiple optimal features tended to be processed more efficiently and homogeneously in vitro than those with fewer optimal features. Additionally, structurally optimal pri-miRNAs exhibited higher expression levels within cells and tended to be more deeply conserved than less optimal pri-miRNAs. These observations suggest that selective pressure exerted to maintain optimal structures in the primary transcriptome of physiologically important miRNAs (Figure 29).

[0225]

[0226] 6.2 Design of shRNA

[0227] Existing shRNA designs have primarily relied on the pri-mir-30a backbone. While pri-mir-30a possesses an optimal loop, its downstream elements are suboptimal, including a short lower stem (11 bp), a weak mGHG motif, and the absence of a bGWG motif (TRIPZ shRNA). To design an shRNA for RNA interference, we modified the lower stem and mGHG motif, substituting nine nucleotides and inserting one to create a bGWG motif (Improved TRIPZ shRNA) (Figure 30).

[0228] - TRIPZ shRNA: SEQ ID NO: 7

[0229] UGCUGUUGACAGUGAGCGAUACAUACUUCUUUAUAUUCCAUAGUGAAGCCACAGAUGUAUGGAAUGUAAAGAAGUAUGUAUUGCCUACUGCCUCGGACUUCA

[0230] - Improved TRIPZ shRNA: SEQ ID NO: 8

[0231] UGCUAUUGGCCGUCUCCGAUACAUACUUCUUUAUAUUCCAUAGUGAAGCCACAGAUGUAUGGAAUGUAAAGAAGUAUGUAUUGGCGCUGUGCCUCGGACUUCA

[0232] As a result of measuring the relative expression level of miRNA and target reporter activity, it was confirmed that the newly designed shRNA increased the miRNA expression level by approximately 2-fold and further improved the target reporter activity reduction rate by 16% to 18% (Fig. 31).

[0233]

[0234]

[0235] Reference example

[0236] Reference Example 1. Preparation of 1-methyl-7-nitroisatoic anhydride (1M7)

[0237] 1M7 was prepared according to a published protocol. Specifically, triphosgene (0.35 eq) was added to a solution of 2-amino-5-nitrobenzoic acid in dried tetrahydrofuran (THF, 100 ml), and the mixture was refluxed for 8 h. The solvent was then removed using a rotary evaporator to obtain 4-nitroisatoic anhydride as a pale yellow solid in theoretical yield. This was dissolved in dried dimethylformamide (DMF, 30 ml) together with N,N-diisopropylethylamine (DIEA, 3 eq) and methyl iodide (5 eq) without further purification, and the mixture was stirred at room temperature for 2 h. 8N hydrochloric acid (HCl, 50 ml) was then poured into the flask, and the mixture was stirred for 5 min. The yellow precipitate was isolated by vacuum filtration and three successive washes with 8N HCl and diethyl ether. Finally, the solvent was removed in vacuo to obtain 1M7 as a yellow solid in theoretical yield, which was confirmed to have sufficient purity by 1H NMR.

[0238]

[0239] Reference Example 2. Preparation of pri-miRNA

[0240] For SHAPE-MaP analysis, 531 miRNA loci annotated in MirGeneDB (v1) were selected. Pri-miRNAs were synthesized by in vitro transcription using a DNA construct generated in a previous study. This DNA construct contained a pri-miRNA hairpin (125 mer) with a 5′ consensus sequence (5′-GCCTATTCAGTTACAGCG-3′, SEQ ID NO: 9), and a 3′ consensus sequence (5′-CGTACTGAAGCTAGCAAC-3′, SEQ ID NO: 10). The pri-miRNA hairpin contained the pre-miRNA hairpin and conserved 5′ (minimum 20 nt) and 3′ (minimum 25 nt) sequences in vertebrates. The DNA constructs were pooled and purified using the QIAquick PCR Purification Kit (QIAGEN), and mutations were removed using the Surveyor® Mutation Detection Kit (IDT). Subsequently, DNA was amplified by PCR using KAPA HiFi HotStart ReadyMix (Kapa Biosystems) after gel purification, using primers Forward (5′-TAATACGACTCACTATAGGGCCTATTCAGTTACAGCG-3′, SEQ ID NO: 11) and Reverse (5′-m(2′-O-Me)GmUTGCTAGCTTCAGTACG-3′, SEQ ID NO: 12). Finally, pri-miRNA was prepared by in vitro transcription using MEGAscript™T7 Transcription Kit (Ambion).

[0241] DNA templates for control RNAs were prepared by PCR using plasmids containing each control RNA and 5′ and 3′ custom adapter sequences in the pTOP Blunt V2 vector. Primers used were SEQ ID NO: 11 (forward) and SEQ ID NO: 12 (reverse). Control RNAs were synthesized by in vitro transcription using the MEGAscript™T7 Transcription Kit (Ambion).

[0242]

[0243] Reference Example 3. RNA Folding and SHAPE Probe

[0244] RNA folding and SHAPE probing were performed according to the methods of a previous study. Pri-miRNA and control RNA (each pri-miRNA:each control RNA = 1:2) were mixed at a ratio of 5 pmol for each sample and refolded in a final volume of 10 μL in 50 mM Tris-HCl (pH 7.5), 100 mM NaCl, and 2 mM MgCl₂ at 37°C for 20 min. The folded RNA was denatured in the presence of 10 mM 1M7 and incubated at 37°C for 75 s. A no-reagent control was performed using DMSO instead of the SHAPE reagent. For denatured samples, RNA was treated with 1M7 (final 10 mM) in 50 mM HEPES (pH 8.0), 4 mM EDTA, and 50% formamide, and incubated at 95°C for 1 min. The modified RNA was recovered by ethanol precipitation.

[0245]

[0246] Reference Example 4. Reverse transcription, library preparation, and sequencing

[0247] Reverse transcription and library preparation under MaP conditions were performed with slight modifications to the method described in a previous study. Briefly, SHAPE-modified RNA was reverse transcribed at 42°C for 3 h using SuperScript II (Invitrogen). The reverse transcription primer (5′-CCTTGGCACCCGAGAATTCCANNNNNNGTTGCTAGCTTCAGTACG-3′, SEQ ID NO: 13) contained a sequence complementary to the 3′ custom adapter, six random nucleotides, and the RNA 5′ adapter sequence from the TruSeq Small RNA Library Preparation Kit (Illumina). The reverse transcription reaction buffer consisted of 0.5 mM mixed dNTPs, 50 mM Tris-HCl (pH 8.0), 75 mM KCl, 6 mM MnCl₂, and 10 mM DTT. After reverse transcription, cDNA was purified using Agencourt RNAClean XP (Beckman).

[0248] Sequencing libraries were generated using a one-step PCR approach. PCR was performed using Q5 Hot Start High-Fidelity DNA polymerase (NEB), and the forward primer (5′-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGACGATCNNNNNNGCCTATTCAGTTACAGCG-3′, SEQ ID NO: 14) contained a P5 sequence for Illumina paired-end sequencing (5′ end), six random nucleotides, and a sequence complementary to the 5′ consensus sequence. The RNA PCR Index Primer from the Truseq Small RNA Library Preparation Kit was used as the reverse primer. The one-step PCR was repeated three times to tag the cDNA and generate the final sequencing library. The PCR products were purified with AMPure XP beads (Beckman). Libraries were sequenced on an Illumina Hiseq X-10 (150x150 paired-end run) along with a 50% PhiX control library (Illumina).

[0249]

[0250] Reference Example 5. Sequence Processing and Sorting

[0251] Raw sequences were quality filtered using the FASTX Toolkit (v. 0.0.13.2) (fastq_quality_filter -q 25 -p 80). Filtered sequences were aligned to pri-miRNA structure sequences using bowtie2 (version 2.2.4) with reduced mismatch and gap extension penalties (--end-to-end --mp 5,1 --rdg 5,1 --rfg 5,1 --no-mixed --no-unal --norc).

[0252]

[0253]

[0254] Reference Example 6. SHAPE Reactivity Analysis

[0255] The number of mutations at each nucleotide position in the aligned sequences was calculated using the ShapeMapper package (version 1.0). Ambiguously mapped reads were removed, and consecutive mutations were merged during alignment parsing (parseAlignment -combine_strand -deletion_masking -remove_ambig_del -min_map_qual 2). To obtain reliable reactivity profiles, RNA samples with sufficient read counts were selected, and read count thresholds of 500, 100, and 500 were applied for the "1M7-treated", "1M7-untreated", and "RNA-denatured" samples, respectively. The thresholds were set to maximize the number of pri-miRNAs without compromising the accuracy of reference RNA structure prediction. If multiple libraries met the threshold for a pri-miRNA, the library with the largest number of reads was selected.

[0256] SHAPE reactivity was calculated as the difference in mutation rates between 1M7-treated and untreated samples divided by the mutation rate in the denatured RNA sample. SHAPE reactivity was normalized to the 90th percentile reactivity value within each RNA sample.

[0257]

[0258] Reference Example 7. Structural Prediction, Selection, and Illustration

[0259] Secondary structures were predicted using the RNAstructure package (Fold, version 5.835). When multiple structures were predicted, the following criteria were used to select a structure: 1 nt ≤ 3′ overhang length ≤ 4 nt, base pairs in the upper stem ≥ 14, base pairs in the lower stem ≥ 8, and nucleotide difference between the 5p and 3p strands ≤ 8. If no structure met the criteria, the criteria were gradually relaxed to select a structure. The selected secondary structures were visualized using PseudoViewer3 and ShapeMapper.

[0260] To prepare sequence-based structures, RNAfold (v. 2.2.5), mfold, RNAstructure (v. 5.835), CONTRAfold (v. 2.0240), and EternaFold were used with default settings. EternaFold was run with CONTRAfold using the EternaFold parameters (contrafold predict {fasta} --evidence --params EternaFoldParams_PLUS_POTENTIALS.v1 --numdatasources 1 --kappa 0.1). If multiple structures were predicted, the structure with the lowest free energy was selected.

[0261]

[0262] Reference Example 8. Prediction Evaluation for Control RNA Structures

[0263] Evaluation of the reference RNA structure predictions was performed using the "scorer" function in the RNAstructure package (version 5.835). Sensitivity was calculated as the proportion of correctly predicted base pairs among known base pairs, and the positive predictive value (PPV) was calculated as the proportion of predicted base pairs that are included in the known structure. Ct files for the known structures of the reference RNAs were manually generated.

[0264]

[0265] Reference Example 9. Cell SHAPE-MaP Analysis

[0266] The raw sequencing data were downloaded from PRJNA855586. To observe the structure under more physiologically relevant conditions, SHAPE data from native cells without SRSF3 depletion were used. The raw sequences were mapped using bowtie2 against two amplicons prepared by the authors.

[0267] - Amplification region 1: chr13: 91350853-91351411+

[0268] - Amplification region 2: chr13: 91350583-91351035+

[0269] The mapping command is: bowtie2 -q --local --mp 5,1 --rdg 5,1 --rfg 5,1 -N 1 --no-unal --norc -a

[0270] Mutation rate, reactivity, and structural models were derived according to the methods described in the “Reactivity Analysis” and “Structure Prediction, Selection, and Visualization” sections. Structural comparisons were performed using the “scorer” function in the RNAstructure package (version 5.835).

[0271]

[0272] Reference Example 10. Determination of basal and apical junctions

[0273] Characterization of the pri-miRNA structure began with determining the basal and apical junctions. Because multiple potential basal and apical junctions can exist in a pri-miRNA, a single basal and apical junction pair was selected using the following algorithm. Based on the notion that SHAPE reactivity represents the structural flexibility of each nucleotide, a "junction score" was formulated to quantify how well-defined a junction is. This score represents the difference in SHAPE reactivity between the single-stranded RNA (ssRNA) region of a potential junction and the adjacent double-stranded RNA (dsRNA) region. Based on the structural model of the microprocessor-pri-miRNA interaction, a 3-nt window was used for each of the ssRNA and dsRNA regions at the basal junction, and a 5-nt window was used at the apical junction.

[0274] The baseline joint score (BJS) was calculated as follows:

[0275] [Mathematical Formula 1]

[0276]

[0277]

[0278] The Apical Joint Score (AJS) was calculated as follows:

[0279] [Equation 2]

[0280]

[0281]

[0282] The positions of the basal and apical junctions (p, q, u, v) were determined as values ​​satisfying the following conditions:

[0283] [Equation 3]

[0284]

[0285]

[0286] However, the distance between (p, q) and (u, v) must be 31-39 bp. If no junction pair satisfied this range, the range was gradually expanded. After applying this algorithm, individual cases were manually reviewed and some junction locations were adjusted.

[0287]

[0288] Reference Example 11. Measurement of stem length, mismatched base pairs, and overhanging bases.

[0289] Stem length was measured by counting the symmetrical number of nucleotides in the 5p and 3p strands. Mismatched base pairs represent unpaired base pairs in both the 5p and 3p strands, and overhanging bases represent unpaired bases in one strand.

[0290] If the internal loop contains a different number of bases on both strands, the unpaired bases are counted as overhanging bases. For example, if the internal loop contains three nucleotides on the 5p strand and one nucleotide on the 3p strand, the stem length is calculated to increase by 1 bp and include one mismatched base pair and two overhanging bases.

[0291]

[0292] Reference Example 12. Measurement of the apical loop size and basal segment length.

[0293] When annotating single-stranded RNA (ssRNA) regions, we used an approach that approximates features associated with Microprocessor binding, in addition to the structural model. For the apical loop, the entire segment above the apical junction, rather than the terminal loop shown in the structural model, was considered the "apical loop." This is because the entire segment, rather than the terminal loop, in the structural model is expected to interact with the DGCR8 Rhed domain. This approach has been used in previous studies.

[0294] For the basal segment, we hypothesized that structurally dynamic bases near the basal junction would favor an unpaired state for DROSHA to anchor the junction. While the apical loop can be clearly demarcated by the apical junction, the basal segment required additional rules beyond the basal junction to determine its boundary. For this purpose, Shannon entropy, a measure of structural uncertainty, was used.

[0295] Shannon entropy calculations followed the procedures described in a previous study. Briefly, the partition function was obtained using the "partition" program from the RNAstructure package, which applied SHAPE reactivity (partition {fasta file} {pfs file} -sh {reactivity file}). Base-pairing probabilities were then calculated using the ProbabilityPlot program (ProbabilityPlot {fasta file} {pfs file} -t). Finally, the Shannon entropy at position iii (HiH_iHi) was calculated as follows:

[0296] [Equation 4]

[0297]

[0298]

[0299] Here, n represents the length of the RNA (125 nt). Using the Shannon entropy of the stem region as the null distribution, we classified structurally dynamic bases as those whose entropy exceeded the 95th percentile of the null distribution (p<0.05). Finally, the base segment was defined as a region containing unpaired or structurally dynamic bases.

[0300]

[0301] Reference Example 13. Cross-analysis between structural features and processed data.

[0302] Pri-miRNAs were divided into five to six groups of similar size based on structural characteristics. Cleavage efficiency and cleavage homogeneity were then compared between each pri-miRNA group. Cleavage efficiency and homogeneity were measured under SRSF3-deficient conditions. These values ​​were normalized to the cleavage efficiency and homogeneity of pri-mir-30a and analyzed.

[0303]

[0304] Reference Example 14. Purification of a Recombinant Microprocessor

[0305] To purify the recombinant full-length Microprocessor complex, transfections were performed using Drosha-HA (10 μg) and Flag-DGCR8 (2.5 μg) constructs, each cloned into the pCK expression vector. The plasmids were diluted in 1 mL OPTI-MEM and mixed with Fugene HD (Promega, 37.5 μL). The mixture was added to HEK293E cells (human fetal kidney 293 EBNA1; ATCC STR profiling-verified) at 50% confluence in 150 mm culture dishes. After 48 h of incubation at 37°C, the cells were harvested and lysed in 1 mL of lysis buffer (500 mM NaCl, 50 mM Tris-HCl pH 7.5, protease inhibitor cocktail (Calbiochem)).

[0306] The crude lysate was purified by sonication followed by centrifugation (13,000 rpm, 4°C, 10 min). During centrifugation, anti-Flag M2 affinity gel (Sigma-Aldrich) was prepared by washing three times with lysis buffer. 1 mL of the supernatant was incubated with 20 μL of pre-washed anti-Flag M2 affinity gel in a rotary incubator at 4°C for 1 h 30 min. The incubated affinity gel was washed twice with 1 mL of washing buffer (500 mM NaCl, 50 mM Tris-HCl, pH 7.5) containing 0.1% NP40, and then three times with the same washing buffer. Wash buffer was completely withdrawn using a syringe fitted with a 30G needle, and 100 μL of elution buffer (500 mM NaCl, 50 mM Tris-HCl pH 7.5, 1 mg / ml 3XFlag peptide (MilliporeSigma)) was added. The mixture was incubated at 4°C in a ThermoMixer (1300 rpm, 30 min). The eluate was collected using a 30G needle and stored at -80°C for further use.

[0307]

[0308] Reference Example 15. In vitro processing analysis

[0309] Pri-miRNA was internally labeled with α-32P-UTP using the MEGAscript T7 Transcription Kit (Ambion) according to the manufacturer's instructions. Radiolabeled RNA (RI) was analyzed and purified on a 6% urea polyacrylamide gel. Gel slices containing RI-labeled RNA were crushed in Gel Breaker Tubes (Istbiotech) and centrifuged (13,000 rpm, 4°C, 1 min). The crushed gel was incubated overnight in a 0.3 M NaCl solution (600 μL) in a rotary incubator. The RNA eluate was purified using a Corning Costar Spin-X centrifuge tube filter (MilliporeSigma) and centrifuged (13,000 rpm, 4°C, 5 min). The labeled RNA recovered from the flowthrough was precipitated at -80°C for 2 h by adding 100% EtOH (1 mL) and GlycoBlue Coprecipitant (Thermo Fisher Scientific, 1 μL).

[0310] In vitro pri-miRNA processing was performed by mixing labeled pri-miRNA with recombinant Microprocessor. The reaction was carried out at 37°C for 1 h in 10 μL buffer containing 100 mM NaCl, 50 mM Tris-HCl pH 7.5, 2 mM MgCl₂, 1 mM DTT, and 1 U / μL SUPERase inhibitor (Ambion). After the reaction, 2X RNA loading buffer (NEB, 9 μL) and 20 mg / ml Proteinase K (MilliporeSigma, 1 μL) were added to the mixture, and the mixture was incubated at 37°C for 30 min and at 50°C for 30 min, respectively. The samples were heated at 95°C for 5 min before gel loading. The processed pri-miRNA was analyzed on a 6% urea polyacrylamide gel using RNA decade marker (Ambion). Quantification of pri-miRNA processing was performed using Multi Gauge V3.0 software (Fujifilm).

[0311]

[0312] Reference Example 16. Ectopic pri-miRNA expression and miRNA qRT-PCR

[0313] 1 μg of pri-miRNA expression vector was cotransfected into 6-well plates (SPL) containing HEK293E cells at 50% cell density using Fugene HD (Promega). After 48 h, RNA was extracted using TRIzol. Mature miRNA levels were quantified using TaqMan MicroRNA Assays (Thermo Fisher Scientific; miR-1: #000385, miR-183: #002269, miR-99a: #000435) according to the manufacturer's instructions.

[0314]

[0315] Reference Example 17. Dissociation constant and free energy of loops

[0316] The dissociation constants were derived based on the Electrophoretic Mobility Shift Assay (EMSA) data from a previous study. The steady-state concentrations of LIN28A, RNA, and LIN28A-RNA complex were denoted as [L], [R], and [LR], respectively, and the initial concentrations of LIN28A and RNA were denoted as [L], [R], and [LR], respectively. i ] and [R i ] was defined as the initial concentration of LIN28A [L i ] is the initial concentration of RNA [R i ] is about 10,000 times larger than [L]~[L i ] was assumed. In this case, the dissociation constant (k d ) was calculated as follows:

[0317] [Equation 5]

[0318]

[0319]

[0320] [L i ] -1 x, [R i ] / [LR](the reciprocal of the bound fraction) is defined as y, then the following equation holds:

[0321] [Equation 6]

[0322]

[0323]

[0324] Dissociation constants were derived by plotting the x and y values ​​and estimating the slope. The x and y values ​​were manually extracted from the linear plots in the original paper. The free energy of the loop was calculated using the efn2 program from the RNAstructure package (efln2 {ct file} {output file} -sh {reactivity file} -w).

[0325]

[0326] Reference Example 18. Analysis of in vitro processing data of artificial pri-miRNA.

[0327] Libraries for the barcode dictionary, input, and product sequences were obtained from GEO (GSE117600). Raw sequences were processed according to the procedures described in the original study. The processing efficiency of each variant was calculated by dividing the number of products by the number of input sequences.

[0328]

[0329] Reference Example 19. Conservation and Cellular Expression of Pri-miRNA

[0330] The conservation of pri-miRNAs was quantified by the average phyloP score of the corresponding mature miRNAs. PhyloP scores for 100 vertebrate genomes were obtained from the UCSC Genome Browser. Intracellular miRNA expression was quantified by the average RPM in AQ-seq samples from HEK293T and HeLa cells. For each group, the first member listed among the seed families in MirGeneDB v2 was selected.

[0331]

[0332] Reference Example 20. Testing for a Stop Loop Dependency

[0333] We used small RNA-seq data from a previous study. Briefly, the DGCR8 P351A mutant, which weakly binds the heme cofactor that aids loop recognition, was overexpressed in DGCR8-deficient HCT116 cells, and miRNAs were analyzed by small RNA-seq. The top 250 miRNAs with the highest read counts in wild-type cells were selected, and miRNAs containing similar genes with identical mature sequences were excluded. Based on the apical loop size, the miRNAs were divided into optimal (≥ 12 nt) and non-optimal (≤ 9 nt) groups for analysis.

[0334]

[0335] Reference Example 21. Improved shRNA design

[0336] The original design was obtained by combining the pri-mir-30a loop and the base segment with the mir-1-1 duplex. Six nucleotide substitutions were made to increase the mGHG score. To create the bGWG motif, a 5'' pair was created at position -10, by substituting one nucleotide and inserting one. To extend the lower stem to 13 bp and maintain the base segment's length, two nucleotide substitutions were made.

[0337]

[0338] Reference Example 22. Dual luciferase assay

[0339] Reverse transfection was performed by trypsinizing HEK293E cells and counting the cell number. A total of 1.5E5 cells were seeded in 24-well plates (SPL) and transfected with pcDNA3-based shRNA plasmids (100 ng) and a dual luciferase reporter containing an incomplete 3XmiR-1 binding site at the 3′ end of Firefly (100 ng) using Fugene HD (Promega). After 24 h, cells were lysed and analyzed using the Dual Luciferase Reporter Assay System (Promega) according to the manufacturer's instructions.

[0340]

[0341] The foregoing description of the present invention is provided for illustrative purposes only. Those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.

Claims

1. As a precursor of modified nucleic acid molecules that induce RNA interference, A modified precursor of an RNA interference-inducing nucleic acid molecule having a stem-loop structure, wherein a GAG or GUG motif is inserted into a 3p strand of a region interacting with Drosha in the stem of the RNA interference-inducing nucleic acid molecule before modification, and wherein the A or U base of the GAG ​​or GUG motif forms a bulge.

2. In claim 1, A modified precursor of an RNA interference-inducing nucleic acid molecule comprising a GAG or GUG motif between two consecutive positions between -9 bp and -12 bp from the Drosha cleavage site of the 5p strand and two positions of the complementary 3p strand.

3. In claim 1, A modified precursor of an RNA interference-inducing nucleic acid molecule containing a mismatched base pair at a position -6 bp, +1 bp, or both relative to the Drosha cleavage site of the 5p strand.

4. In claim 1, A modified precursor of an RNA interference-inducing nucleic acid molecule comprising one or more of the following characteristics: (a) the loop comprises 5 to 50 nucleotides; (b) the stem contains 20 to 50 base pairs; (c) the stem contains 10 or fewer mismatched base pairs; (d) the base segment of the 5p strand or the 3p strand comprises 1 to 15 nucleotides; and (e) 5 to 30 base pairs from the start base pair position of the 5p strand to the Drosha cleavage site of the 5p strand.

5. In claim 1, A precursor of an RNA interference-inducing nucleic acid molecule, further comprising at least one selected from the group consisting of a basal UG motif, an mGHG motif, an apical UGUG motif, and a CNNC motif.

6. A precursor of an RNA interference-inducing nucleic acid molecule that promotes cleavage of Drosha in claim 1.

7. An RNA interference-inducing nucleic acid molecule produced from a precursor of any one of claims 1 to 6.

8. In claim 7, the RNA interference-inducing nucleic acid molecule is any one selected from the group consisting of miRNA, shRNA (short hairpin RNA), and siRNA (small interfering RNA).

9. A nucleic acid molecule encoding a precursor of an RNA interference-inducing nucleic acid molecule of any one of claims 1 to 6.

10. An expression construct comprising the nucleic acid molecule of claim 9.

11. A vector comprising the expression construct of claim 10.

12. A method for designing a modified precursor of an RNA interference-inducing nucleic acid molecule comprising: A step of designing a precursor of a modified RNA interference-inducing nucleic acid molecule by inserting a GAG or GUG motif into a precursor of a modified RNA interference-inducing nucleic acid molecule having a stem-loop structure, wherein the GAG ​​or GUG motif is inserted into the 3p strand of a region interacting with Drosha in the stem of the modified precursor, and the A or U base of the GAG ​​or GUG motif forms a bulge.

13. A method for producing an RNA interference-inducing nucleic acid molecule, comprising: A step of designing a modified precursor of an RNA interference-inducing nucleic acid molecule according to the method of claim 12; A step of producing a modified precursor of an RNA interference-inducing nucleic acid molecule through biological synthesis via cell expression, enzymatic synthesis (in vitro transcription; IVT), or chemical RNA synthesis; and A step for producing an RNA interference-inducing nucleic acid molecule from a modified precursor of an RNA interference-inducing nucleic acid molecule.

Citation Information

Patent Citations

  • VARIANT RNAi

    KR1020170110149A