Circular RNA and methods of making the same
By employing a multi-segment trans-splicing method using type I introns, the problems of numerous byproducts and low efficiency in circular RNA preparation have been solved, achieving efficient, byproduct-free circular RNA production, which is suitable for large-scale production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIGENE GUANGZHOU BIOLOGICS MFG CO LTD
- Filing Date
- 2024-09-05
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods for preparing circular RNA suffer from problems such as numerous byproducts, low efficiency, and difficulty in scaling up. In particular, the improved PIE system suffers from byproducts that impair RNA circularization efficiency during the circularization of long RNA.
By employing a trans-splicing method using multiple fragments of type I introns or their derivatives, and through the formation of the scaffold and catalytic domains of type I introns by RNA molecules arranged in a specific order, the efficient production of circular RNA is achieved.
It enables the production of circular RNA with virtually no byproducts, improves circularization efficiency, simplifies manufacturing and control, and is suitable for large-scale production.
Smart Images

Figure CN122139031A_ABST
Abstract
Description
Technical Field
[0001] This application relates to compositions and methods for efficiently constructing circular RNA via trans-splicing reactions.
[0002] References to sequence lists submitted electronically This application contains a sequence list submitted electronically. Information contained in the electronic sequence list (sequence list; size: 62 KB; and creation date: August 22, 2024) is incorporated herein by reference in its entirety. Background Technology
[0003] The clinical success of mRNA-based COVID-19 vaccines and, more recently, mRNA-based personalized cancer vaccines has fueled hope and enthusiasm for developing mRNA-based therapeutics. Any protein can be expressed via mRNA through cellular translation mechanisms. mRNA does not enter the cell nucleus. Therefore, mRNA does not pose the potential risk of altering DNA in the cell nucleus, leading to mutagenesis. Furthermore, mRNA production, including large-scale mRNA production, has become more efficient and faster. However, the instability of mRNA in vivo inhibits its translation efficiency, and its immunogenicity also hinders the development of mRNA-based therapeutics. Although the use of modified nucleotides (e.g., 1mψ) can reduce immunogenicity and increase mRNA stability and translation efficiency, other options for RNA-based therapies still need to be explored.
[0004] Circular RNAs (circRNAs) are a class of single-stranded RNAs distinguished by their covalently closed topology. The unique structure of circRNAs confers them greater stability, a longer half-life, and stronger resistance to RNase R than linear mRNAs. CircRNAs are ubiquitous in species ranging from viruses to mammals and possess numerous biological functions, including but not limited to acting as miRNA sponges and protein sponges, and interacting with many different RNA-binding proteins (RBPs). A subset of circRNAs can also undergo cap-independent translation. It has been shown that circRNAs with internal ribosome entry sites (IRES) can be translated both in vivo and in vitro. Furthermore, circRNAs have been reported to have lower immunogenicity than unmodified mRNAs and comparable immunogenicity to modified mRNAs. Molecular cells, 2019, 74, 508-520. ).
[0005] circRNAs with specific properties can be generated both in vivo and in vitro (Obi et al., Methods, 2021, 196: 85-103). In vitro circRNAs can be generated by synthesizing a linear RNA precursor followed by ligation to its ends to form a covalently closed circular structure. Several methods exist for the in vitro synthesis of circRNAs, including but not limited to chemical ligation, T4 RNA / DNA ligase-mediated enzymatic ligation, and ribozyme methods using self-splicing introns. Chemical ligation has relatively low ligation efficiency and requires the use of cyanogen bromide (BrCN), which may pose biosafety concerns. Enzymatic ligation is desirable for the ligation of large RNA molecules, but it may exhibit intermolecular end-joining side reactions, and it may be technically challenging to scale up. Ribozyme methods using self-splicing introns are currently the most popular and widely used method for synthesizing circRNAs.
[0006] Self-splicing introns act as ribozymes and catalyze the cleavage of themselves from pre-mRNA. Naturally occurring type I introns are characterized by a linear array of sequences (e.g., from the 5' end to the 3' end: 5' exon-5'-intron-3'-intron-3' exon) and structural features. Type I introns self-splice via a two-step transesterification mechanism to join the 5' and 3' exons to form a linked linear exon. Rearranging introns into the sequence 3'-intron-3' exon-5' exon-5'-intron allows the introns to undergo cleavage and ligation reactions to form circular intronic RNA. This ribozyme-catalyzed method (called the rearranged intron-exon self-splicing system or PIE system, for example, derived from...) Anabaena tRNA introns and the thymidine synthase (Td) gene of T4 phage have been reported for the formation of circular RNA in vitro with the aid of only GTP and Mg2+. A limitation of the PIE system is the difficulty in achieving circularization of long RNAs. PIE can be improved with the aid of strong homologous arms. Modified RNA molecules used for circRNA synthesis have the following components arranged from the 5' end to the 3' end: 5'-homologous arm-3'-intron-spacer-3'-exon-target RNA sequence-5'-exon-spacer-5'-intron-3'-homologous arm, where the 5'-homologous arm is perfectly paired with the 3'-homologous arm. Although improved PIE can efficiently circularize RNA molecules, byproducts of improved PIE (such as dimers or polymers generated by homologous arms) can impair RNA circularization efficiency and pose challenges to chemical, manufacturing, and controlled (CMC) production.
[0007] Therefore, there is still a need in this field for improved methods for preparing circular RNA. Summary of the Invention
[0008] The inventors of this application have surprisingly discovered that a trans-splicing method involving multiple fragments of type I introns or their derivatives can be used for the efficient production of circular RNA with virtually no byproducts (e.g., dimers or polymers).
[0009] In one general aspect, this application relates to RNA molecule pairs for preparing circular RNA, comprising: 1) A first RNA molecule comprising the following elements, said elements being operatively linked and arranged in order from the 5' end to the 3' end of said molecule: i) The 3' intron sequence of type I introns ii) The 3′ splice site of the type I intron, iii) Optionally, the 3' exon sequence of the type I intron, iv) Target RNA sequence, v) Optionally, the 5' exon sequence of the type I intron, vi) The 5′ splice site of the type I intron, and vii) The 5′ intron sequence of the type I intron, and (2) A second RNA molecule containing the intermediate intron sequence of the type I intron. When the first RNA molecule comes into contact with the second RNA molecule, the scaffold and catalytic domains of the type I intron are formed, thereby allowing the generation of the circular RNA containing the 3' exon sequence, the target RNA sequence, and the 5' exon sequence.
[0010] In some embodiments, the 3' intron sequence contains the R and S sequences of the type I intron, the 5' intron sequence contains the internal guide sequence (IGS) and P sequence of the type I intron, and the intermediate intron sequence contains the Q sequence of the type I intron.
[0011] In some implementations, the first RNA molecule comprises, from the 5' end to the 3' end of the molecule, a) Fragment 3, which includes the 3' intron sequence and the 3' splice site, which are operatively connected and arranged in order from the 5' end to the 3' end of fragment 3; b) Optionally, the 3' exon sequence; c) Target RNA sequence; d) Optional, 5' exon sequence; e) Fragment 1, comprising the 5' splice site and the 5' intron sequence operably connected and arranged in order from the 5' end to the 3' end of fragment 1; and The second RNA molecule contains segment 2 having the intermediate intron sequence.
[0012] Preferably, fragments 1-3 comprise nucleotide sequences or variants thereof of three parts of a type I intron obtained by dividing the intron using two interrupt sites, respectively. More preferably, the three parts of the type I intron are obtained by dividing the intron using an interrupt site located within the loop of the P6 stem-loop region of the type I intron and another interrupt site located within the loop of the P2, P5, P8, or P9 stem-loop regions of the type I intron.
[0013] In some embodiments, the target nucleotide sequence comprises at least one protein-coding sequence and an internal ribosome entry site (IRES) operably linked thereto. In some embodiments, the target nucleotide sequence comprises a non-protein-coding sequence.
[0014] In some embodiments, one or more of the 3' intron sequence, 3' splice site, 3' exon sequence, 5' exon sequence, 5' splice site, 5' intron sequence, and intermediate intron sequence used in the present invention may comprise or be derived from the corresponding natural sequence of a type I intron. In some embodiments, one or more sequences used in the components of the present invention have 100% identity with the corresponding natural sequence. In other embodiments, the sequences used in one or more components of the present invention are modified sequences having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the corresponding natural sequence. Compared to the natural sequence, the modified sequence may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions. The natural sequence may be derived from any suitable type I intron. Suitable examples of type I introns include, but are not limited to, type I introns selected from the following groups of organisms: cyanobacteria of the genus *Anabaena*, T4 bacteriophages, and species of the genus *Vibrio*. (Azoarcus sp.) BH72, Enterobacter T4 phage, Tervod phage, phage SpO1, phage S3b, Bacillus anthracis (Bacillus anthracis) Clostridium botulinum (Clostridium botulinum) Thermophilic Tetrahymena (Tetrahymena thermophila) Pavdulella (Dunaliella parva) Pneumocystis carinii (Pneumocystis carinii) Polycephalomycetes (Physarum polycephalum) , Anabaena Species PCC7120 Bifidobacterium hopsii (Scytonema hofmanni) Agrobacterium tumefaciens (Agrobacterium tumefaciens) Synechocystis (Synechocystis) PCC 6803 Slender Synechococcus (Synechococcus elongatus) PCC 6301, Neurospora crassa (Neurospora crassa) Candida albicans (Candida albicans) Seradiomycium selenophorum (Scytalidium cerradiumydiaces), Pediadis Esis Snowy Algae (Pediadiaces Chlamydomonas nivalis), Chlorella (Chlorella vulgaris) Parasitic Proteobacteria (Amoebidium parasiticum) Neurospora crassa, Necropoda (Emericella nidulans) brewing yeast (Saccharomyces cerevisiae) Saccharomyces cerevisiae (Schizosaccharomyces pombe), aquatic new green algae (Neochloris aquatica), Pavlovian algae, Negev simcania (Symkania negevensis) and Nesting naked cell shell. In some implementations, the natural sequence originates from... Anabaena Type I introns of cyanobacteria, for example Anabaena cyanobacteria Type I introns of the pre-tRNA-Leu gene of the species. In some other embodiments, the type I introns originate from T4 phage or... Nitrogenous Vibrio species BH72 Preferably, type I introns originate from... Anabaena cyanobacteria species pre-tRNA-Leu gene, T4 phage td gene or species of the genus *Vibrio* BH72 The pre-tRNA-Ile gene.
[0015] In some embodiments, the first RNA molecule contains either a 3' exon sequence or a 5' exon sequence. In another embodiment, the first RNA molecule contains both a 3' exon sequence and a 5' exon sequence. In some embodiments, the first RNA molecule does not contain either a 3' exon sequence or a 5' exon sequence. In some embodiments, the first RNA molecule does not contain both a 3' exon sequence and a 5' exon sequence.
[0016] In some embodiments, the first and second RNA molecules may include additional features to optimize circRNA production. For example, the first RNA molecule may further include a 3' end binding motif, and the second RNA molecule may further include a 5' end binding motif, wherein the 3' end binding motif and the 5' end binding motif are complementary to form a double-stranded region. Preferably, the 3' end binding motif and the 5' end binding motif are completely complementary to each other.
[0017] Another general aspect of this application relates to nucleic acid vectors for preparing circular RNA molecules. Specifically, nucleic acid vectors encoding a first RNA molecule and / or a second RNA molecule according to embodiments of this application. In a preferred embodiment, the vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of one or more RNA molecules. In some preferred embodiments, the promoter is selected from the T7 RNA polymerase promoter, the T6 viral RNA polymerase promoter, the SP6 viral RNA polymerase promoter, the T3 viral RNA polymerase promoter, or the T4 viral RNA polymerase promoter.
[0018] On the other hand, this application provides a method for preparing circular RNA, the method comprising: i) Provide or obtain RNA molecule pairs according to the embodiments of this application; ii) Add a buffer solution to the RNA molecule pair to obtain a reaction mixture to allow the formation of the scaffold and catalytic domains of the type I introns, and iii) Add GTP and divalent metal cations to the reaction mixture at a temperature that allows the formation of the circular RNA.
[0019] Optionally, the method further includes iv) harvesting the circular RNA formed in step iii).
[0020] In another aspect, this application provides circular RNA generated by the method of this application. This application provides a composition comprising an effective amount of the circular RNA of this application and a pharmaceutically acceptable carrier. Preferably, the pharmaceutically acceptable carrier comprises lipids, polymers, or lipid-polymer hybrids, such as lipid nanoparticles (LNPs), lipid microparticles, lipid suspensions, or liposomes.
[0021] Other aspects of this application include compositions comprising cells (e.g., eukaryotic cells) containing nucleic acids encoding circular RNA or circular RNA according to embodiments of this application. Compositions comprising cells and one or more pharmaceutically or physiologically acceptable carriers, excipients, or diluents are also provided.
[0022] Other aspects, features, and advantages of the invention will become apparent from the following disclosure, including the detailed description and preferred embodiments thereof and the appended claims. Attached Figure Description
[0023] The foregoing and other objects, aspects, features and advantages of the exemplary embodiments will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings.
[0024] Figure 1 From AnabaenaA schematic diagram of the secondary structure of type I introns of pre-tRNA.
[0025] Figure 2 A scheme for synthesizing classical rearrangements of cyclic exons-intron-exon (PIE).
[0026] Figure 3 A. A classic PIE strategy for generating circRNA from a target RNA sequence; Figure 3 B. Validate circRNA using 5% denatured TBE-PAGE.
[0027] Figure 4 A. A scheme for an improved PIE strategy to generate circRNA using the target RNA sequence; Figure 4 B. Cycloning with 10 / 15 / 19 bp homologous arms was confirmed by 5% denatured urea-TBE PAGE under two cyclization conditions (A and B); and Figure 4 C. Circulation of long RNAs with 19 bp homologous arms under B-circulation conditions was confirmed by 2% E-gel EX.
[0028] Figure 5 Multi-fragment activity Anabaena The structure of pre-tRNA type I introns.
[0029] Figure 6 The self-splicing activity of the multi-fragment type I introns was verified by 20% denatured urea-TBE PAGE. The first column is fragment 1 (F1) of the assembled intron-1, the second column is fragment 2 (F2) of the assembled intron-1, the third column is fragment 3 (F3) of the assembled intron-1, the fourth column is a mixture of F1, F2 and F3, and the multi-fragment intron-1 assembled from F1, F2 and F3 self-splicing to produce the linear exon (17 nt) indicated by the red arrow; the fifth column is fragment 1 (F1) of the assembled intron-4, the sixth column is fragment 2 (F2) of the assembled intron-4, the seventh column is fragment 3 (F3) of the assembled intron-4, and the eighth column is a mixture of F1, F2 and F3, and the multi-fragment intron-4 assembled from F1, F2 and F3 self-splicing to produce the linear exon (17 nt) indicated by the red arrow.
[0030] Figure 7 Multi-fragment activity Anabaena The splicing mechanism of pre-tRNA type I introns.
[0031] Figure 8 A-8D. Multi-fragment activity Anabaena Different cleavage sites on P6 of pre-tRNA type I introns.
[0032] Figure 9 A scheme for the trans-splicing strategy of synthesizing cyclic exons. Figure 10 The cyclic exons synthesized using the trans-splicing strategy were verified by 5% urea-TBE PAGE. Figure 10 A indicates the cyclic exon marked with a red arrow, generated by P-AI-1 with the help of its corresponding F2. Both P-AI-1 and its corresponding F2 are derived from self-assembled intron-1. Figure 10 B indicates the cyclic exon marked with a red arrow, generated by P-AI-2 with the help of its corresponding F2. Both P-AI-2 and its corresponding F2 are derived from self-assembled intron-2. Figure 10 C indicates the cyclic exon marked with a red arrow, generated by P-AI-4 with the help of its corresponding F2. Both P-AI-4 and its corresponding F2 are derived from self-assembled intron-4.
[0033] Figure 11 The effect of the P6 cleavage site on the circular exons generated by the trans-splicing strategy was verified using 5% urea-TBE PAGE. Circular exons are marked with red arrows.
[0034] Figure 12 A. A trans-splicing strategy and RNA splicing efficiency scheme for synthesizing circRNA using the target RNA sequence; Figure 12 B. Cycloning was confirmed by 5% denatured urea-TBE PAGE.
[0035] Figure 13 A. A scheme for generating circRNA using a trans-splicing strategy with binding motifs (BM) of different lengths. Figure 13B. The influence of binding motif (BM) length on the looping efficiency of the trans-splicing strategy was confirmed by 5% denatured urea-TBE PAGE. Column 1 (①) is I(5)-P-RNA-seq 1; Column 2 (②) is a mixture of I(5)-P-RNA-seq 1 and BM(5)-F2. Meanwhile, the circRNAs marked by red circles are synthesized from I(5)-P-RNA-seq 1 with the help of BM(5)-F2. Column 3 (③) is I(15)-P-RNA-seq 1; Column 4 (④) is a mixture of I(15)-P-RNA-seq 1 and BM(10)-F2. Meanwhile, the circRNAs marked by red circles are synthesized from I(15)-P-RNA-seq 1 with the help of BM(10)-F2. Column 5 (⑤) is a mixture of I(15)-P-RNA-seq 1 and BM(15)-F2. Meanwhile, the circRNAs marked by red circles are synthesized from I(15)-P-RNA-seq 1. 1 was synthesized with the help of BM(15)-F2; column 6 (⑥) is I(20)-P-RNA-seq 1, column 7 (⑦) is a mixture of I(20)-P-RNA-seq 1 and BM(20)-F2, and the circRNA marked by the red circle was synthesized by I(20)-P-RNA-seq 1 with the help of BM(20)-F2; column 8 (⑧) is I(50)-P-RNA-seq 1, column 9 (⑨) is a mixture of I(50)-P-RNA-seq 1 and BM(50)-F2, and the circRNA marked by the red circle was synthesized by I(20)-P-RNA-seq 1 with the help of BM(50)-F2; Figure 14 A. A strategy for improving trans-splicing of circRNA and a protocol for increasing RNA splicing efficiency; Figure 14 B. Cycloning was verified by 5% denatured urea-TBE PAGE; Figure 14 C. The circRNA was confirmed by RT-PCR using divergent primers spanning the circRNA linker site; and Figure 14 D. Sanger sequencing of the linker region confirmed the presence of circRNA.
[0036] Figure 15 An improved trans-splicing strategy for synthesizing circRNAs of the secretory protein Gaussian luciferase expressed in 293T cells.
[0037] Figure 16 An improved trans-splicing strategy for synthesizing circRNAs of the intercellular protein firefly luciferase in 293T cells.
[0038] Figure 17 An improved trans-splicing strategy for synthesizing circRNA expressing the nuclear protein eEGFP-NLS in 293T cells. Detailed Implementation
[0039] Various publications, articles, patents, and patent applications are cited or described in the background art and throughout the specification; each of these references is incorporated herein by reference in its entirety. Discussions of documents, actions, materials, devices, articles of manufacture, etc., already included in this specification are for the purpose of providing context for the invention. Such discussion does not acknowledge that any or all of these matters constitute part of the prior art with respect to any disclosed or claimed invention.
[0040] In this application, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. Furthermore, all terms and experimental procedures used herein related to protein and nucleotide chemistry, molecular biology, cell and tissue culture, microbiology, and immunology are terms and conventional methods commonly used in the art. For example, the standard DNA recombination and molecular cloning techniques used herein are well known to those skilled in the art and are described in detail in the following reference: Sambrook, J., Fritsch, Efland Maniatis, T, Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989. Meanwhile, to better understand this application, definitions and explanations of relevant terms are provided below.
[0041] For clarity and readability, the following definitions are provided. Any technical features mentioned in these definitions can be found in the various embodiments of this application. Further definitions and interpretations may be provided specifically in the context of these embodiments.
[0042] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” and “nucleic acid fragment” are used interchangeably and refer to a polymer of single-stranded or double-stranded RNA or DNA, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides (usually found in their 5'-monophosphate form) are designated by their single-letter names as follows: “A” represents adenosine or deoxyadenosine (for RNA or DNA, respectively), “C” represents cytidine or deoxycytidine, “G” represents guanosine or deoxyguanosine, “U” represents uridine, “T” represents deoxythymidine, “R” represents purine (A or G), “Y” represents pyrimidine (C or T), “K” represents G or T, “H” represents A, C, or T, “I” represents inosine, and “N” represents any nucleotide. Although nucleotide sequences in this document may be represented as DNA sequences (containing T(s)), when RNA is referred to, those skilled in the art can readily determine the corresponding RNA sequence (i.e., U instead of T).
[0043] The terms "circular RNA," "circular polynucleotide," and "circRNA" are used interchangeably and refer to polynucleotides that form a circular structure. Preferably, the circular structure is formed by covalent bonds.
[0044] As used throughout this application in the context of nucleic acid or amino acid sequences, the term "fragment" refers to a portion of the full-length sequence of a nucleic acid or amino acid sequence. Therefore, a fragment typically consists of a sequence identical to a corresponding segment within the full-length sequence. In the context of this application, a preferred fragment of a sequence consists of consecutive segments of an entity (e.g., a nucleotide or amino acid) corresponding to consecutive segments of an entity in the molecule from which the fragment is derived, said consecutive segments representing at least 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the total (i.e., full-length) molecule from which the fragment is derived. As used throughout this application, the term "fragment" typically includes sequences of self-splicing introns as defined herein, which, in terms of their nucleic acid sequence, are truncated at the 5' and / or 3' ends compared to the original intron's nucleic acid sequence. Therefore, such truncation can occur at the nucleic acid level. Thus, sequence identity relative to such a fragment as defined herein can preferably refer to the entire nucleic acid molecule.
[0045] The term "identity" has a generally accepted meaning in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using techniques known in the art. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example, Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, ed., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., ed., M Stockton Press, New York, 1991). Numerous methods exist for determining sequence identity. Examples of algorithms suitable for determining the percentage of sequence identity are those used in the Basic Local Alignment Search tool (hereinafter “BLAST”), see, for example, Altschul et al., J. Mol. Biol. 215:403-410, 1990 and Altschul et al., Nucleic Acids Res., 15:3389-3402, 1997. Software for performing BLAST analysis is publicly available from the National Center for Biotechnology Information (hereinafter “NCBI”). Default parameters for determining sequence identity using software available from NCB (e.g., BLASTN for nucleic acid sequences) are described in McGinnis et al., Nucleic Acids Res., 32: W20-W25, 2004.
[0046] As used herein, the terms “coding sequence,” “coding region,” and the corresponding abbreviation “cds” each refer to a sequence containing a triplet of multiple nucleotides that can be translated into a peptide or polypeptide. In the context of this application, a coding sequence may be an RNA sequence containing a triplet of multiple nucleotides, beginning with a start codon and preferably ending with a stop codon.
[0047] As used throughout the application in the context of nucleic acids, the term "derived from" means, for a nucleic acid "derived from" (another) nucleic acid, that is, a nucleic acid derived from (another) nucleic acid has, for example, at least 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the nucleic acid from which it is derived. Those skilled in the art will understand that sequence identity is typically calculated for the same type of nucleic acid, i.e., for DNA sequences or for RNA sequences. Therefore, it should be understood that if DNA is "derived from" RNA, or if RNA is "derived from" DNA, the RNA sequence is converted to the corresponding DNA sequence in the first step (particularly by replacing uracil (U) with thymidine (T) throughout the sequence), or vice versa, the DNA sequence is converted to the corresponding RNA sequence (particularly by replacing T with U throughout the sequence). Thereafter, the sequence identity of the DNA sequence or the sequence identity of the RNA sequence is determined. Preferably, "derived from" nucleic acid also refers to nucleic acid that has been modified compared to the nucleic acid from which it is derived, for example, to further increase RNA stability and / or elongate and / or increase protein production. In the context of amino acid sequences (e.g., antigenic peptides or proteins), the term "derived from" means that the amino acid sequence derived from (another) amino acid sequence has, for example, at least 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the amino acid sequence from which it is derived.
[0048] As used herein in the context of nucleic acid molecules, the term "rearrangement" refers to altering the sequence of a molecule. In this application, rearranged molecules are sometimes equivalent to the precursors of molecules (indicated by the prefix pre- in comparative examples, or by the prefix P- in embodiments of this application).
[0049] As used herein, the term “and / or” covers all combinations of items connected by the term, and each combination should be considered as listed separately herein. For example, “A and / or B” covers “A,” “A and B,” and “B.” For example, “A, B, and / or C” covers “A,” “B,” “C,” “A and B,” “A and C,” “B and C,” and “A and B and C.”
[0050] When referring to the “SEQ ID NO” of other patent applications or patents, the sequence (e.g., amino acid sequence or nucleic acid sequence) is expressly incorporated herein by reference. For the “SEQ ID NO” provided herein, the information <> under the identifier (in the sequence scheme) is also expressly included herein in its entirety. When “SEQ ID NO” is mentioned in the context of an RNA sequence, those skilled in the art will understand and be able to derive the RNA sequence from the mentioned SEQ ID NO, provided a DNA sequence is provided. When “SEQ ID NO” is mentioned in the context of a DNA sequence, those skilled in the art will understand and be able to derive the corresponding DNA sequence from the mentioned SEQ ID NO, provided an RNA sequence is provided.
[0051] As used herein, the term "type I intron" refers to a structured self-splicing intron that acts as a ribozyme and automatically catalyzes its removal (splicing) from primary transcripts. In a wide range of organisms, type I introns catalyze their own excision from mRNA, tRNA, and rRNA precursors. Type I introns comprise 14 subgroups, with the majority belonging to the IC3 subgroup. For example, a type I intron may belong to the IC3 subgroup. Anabaena Type I introns of cyanobacteria or type I introns of T4 bacteriophages belonging to the IA2 subgroup or type I introns belonging to the IC3 subgroup species of the genus *Vibrio* Type I introns of BH72. Other embodiments of self-splicing introns that can be used in this application include, but are not limited to, self-splicing introns derived from the following organisms: Enterobacterial phage T4, Tervod phage, phage SpO1, phage S3b, Anthrax bacteria , Clostridium botulinum , Thermophila tetrahymena , Bafdulella , Pneumocystis carinii , Polysaccharomyces multiceps , Anabaena species PCC7120 , Bifidus hominis , Agrobacterium tumefaciens , Synechocystis 6803 , Synechococcus slenderus PCC 6301, Neurospora crassa, Candida albicans, Ceratopsia Mildiasis Columnar Species, Pediasis Snowy Algae, Chlorella, Parasitic Proteobacteria, Neurospora crassa, and Navelnae Cell shell, Saccharomyces cerevisiae, Saccharomyces cerevisiae, aquatic new green algae, Dunaliella paff, Simcania negesii, naked celluloides shell.
[0052] Many type I introns can self-splice in vitro without the assistance of protein cofactors. Although type I introns are highly variable at the primary sequence level, they share characteristically conserved secondary and tertiary structures. Figure 1 As shown, type I introns contain pairing (P) elements (e.g., P1 to P9) and characteristic secondary structures of single-stranded loop regions, as defined by Waring and Davies (Davies, R W., Waring, R B., Ray, R A., Brown, TA and cao, C. (1982) Nature 300, 719-724.). When tracing the RNA strand from 5' to 3', they are numbered according to their occurrence. Note that P2 is the stem closest to the 3' side of P1, and P9 is the stem closest to the 3' side of P7.
[0053] Short, conserved sequences named P, Q, R, and S are involved in the formation of the core helical region. For example, the P sequence pairs with Q, which contributes to the P4 helix, and the R sequence pairs with S, which contributes to the P7 helix. The active core of a type I ribozyme contains the scaffold domain P4 / P6 (P4, P5, and P6) and the catalytic domain P3 / P9 (P3, P7, P8, and P9). The P3-P7-P9 helix contains a GTP-binding pocket. Base-pairing interactions between the 5' ends of introns and flanking exon sequences define the locations of the 5' and 3' splice sites. The 5' splice site is determined by the inner guide sequence (IGS) (a short intron sequence near the 5' end) pairing with the sequence of the upstream exon (5' exon) to form P1. The 3' splice site is determined by the short sequence of the downstream exon (3' exon) pairing with a portion of the IGS and mediating the interaction between the P9 and P3 / P8 helices that form the catalytic core.
[0054] In view of this disclosure, methods in the art can be used to determine the conserved P, Q, R, and S sequences, IGS, and P1 and P9 regions of group I introns, for example, by referring to the following references: Burke, JM, et al., (1987) Structural conventions for group I introns; Stahley, RM, et al., (2006) RNA splicing: group I intron crystal structures reveal the basis of splice site selection and metal ion catalysis; and / or Woodson, AS, (2005) Structure and assembly of group I introns.
[0055] Splicing of type I intron RNA is accomplished via a two-step transesterification reaction. The first reaction is initiated by the 3'-OH group of exogenous GTP, which docks in the G-binding pocket located in the p7 region, and the 3'-OH group attacks the 5' splice site. In the second reaction, the released 3'-OH of the 5' exon attacks the phosphodiester bond between the terminal G of the intron and the 3' exon, resulting in intron release and exon attachment. See, for example, Hausner et al. Mobile DNA, Volume 5, Article No. 8 (2014). Type I introns are also described in detail, for example, in: Nielsen H, Johansen SD (2009). "Group I introns: Moving in new directions". RNA Biol . 6(4): 375–83; Cate JH, GoodingAR, Podell E, et al. (September 1996). "Crystal structure of a group I ribozymedomain: principles of RNA packing". Science . 273 (5282): 1678–85; Cech TR(1990). "Self-splicing of group I introns". Annu. Rev. Biochem . 59: 543–68; Woodson SA (2005, June). "Structure and assembly of group I introns". Curr. Opin. Struct. Biol . 15 (3): 324–30.
[0056] In one general aspect, this application relates to RNA molecule pairs for preparing circular RNA, comprising: (1) A first RNA molecule comprising the following elements, said elements being operatively linked and arranged in order from the 5' end to the 3' end of said molecule: i) The 3' intron sequence of type I introns ii) The 3′ splice site of the type I intron, iii) Optionally, the 3' exon sequence of the type I intron, iv) Target RNA sequence, v) Optionally, the 5' exon sequence of the type I intron, vi) The 5′ splice site of the type I intron, and vii) The 5′ intron sequence of the type I intron, and (2) A second RNA molecule containing the intermediate intron sequence of the type I intron. When the first RNA molecule comes into contact with the second RNA molecule, the scaffold and catalytic domains of the type I intron are formed, thereby allowing the generation of the circular RNA containing the 3' exon sequence, the target RNA sequence, and the 5' exon sequence.
[0057] As used herein, "3' intron sequence of a type I intron" refers to an RNA sequence derived from the sequence adjacent to the 3' end of the type I intron. In some embodiments, "3' intron sequence of a type I intron" comprises the R and S sequences of the type I intron. In some embodiments, "3' intron sequence of a type I intron" has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the native sequence adjacent to the 3' end of the type I intron. For example, "3' intron sequence of a type I intron" may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions compared to the corresponding native sequence.
[0058] As used herein, "5' intron sequence of a type I intron" refers to the RNA sequence derived from the type I intron near its 5' end. In some embodiments, "5' intron sequence of a type I intron" includes the internal guide sequence (IGS) and the P sequence of the type I intron. In some embodiments, "5' intron sequence of a type I intron" has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the native sequence near the 5' end of the type I intron. For example, compared to the corresponding native sequence, "5' intron sequence of a type I intron" may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions.
[0059] As used herein, "intermediate intron sequence of a type I intron" refers to an RNA sequence derived from a type I intron, said RNA sequence being located between the natural 5' intron sequence and the natural 3' intron sequence of the type I intron. In some embodiments, the "intermediate intron sequence of a type I intron" comprises a type I self-splicing Q sequence. In some embodiments, the "intermediate intron sequence of a type I intron" has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with the natural sequence between the natural 5' intron sequence and the natural 3' intron sequence of the type I intron. For example, the "intermediate intron sequence of a type I intron" may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions compared to the corresponding natural sequence.
[0060] As used herein, a “splicing site” refers to a dinucleotide sequence in which the covalent bond between two nucleotides is cleaved during a splicing reaction. A “5’-splicing site” refers to a splicing site at the 5’ end of a type I intron, while a “3’-splicing site” refers to a splicing site at the 3’ end of a type I intron. Introns are removed from the primary transcript by cleavage at the 5’ and 3’ splicing sites. In some embodiments, the 5’-splicing site contains the sequence UA. In some embodiments, the 5’-splicing site contains the sequence UU. In some embodiments, the 3’-splicing site contains the sequence UG. In some embodiments, the removed RNA sequence begins with the dinucleotide UA at its 5’ end and ends with the dinucleotide UG at its 3’ end. In some embodiments, the removed RNA sequence begins with the dinucleotide UU at its 5’ end and ends with the dinucleotide UG at its 3’ end.
[0061] As used herein, “exon” or “exon sequence” means a sequence derived from a natural exon sequence located flanking a type I intron and capable of being recognized and / or spliced by the type I intron. In some embodiments, both the 3' exon sequence (i.e., the sequence of the exon adjacent to the 3' splice site or a fragment thereof) and the 5' exon sequence (i.e., the sequence of the exon adjacent to the 5' splice site or a fragment thereof) are necessary for the preparation of circular RNA using the methods according to this application. In other embodiments, the 3' exon sequence and / or the 5' exon sequence are not necessary for the preparation of circular RNA using the methods according to this application. See, for example, Rausch et al., Nucleic Acid Research, 2021, 49(6), e35, the contents of which are incorporated herein by reference in their entirety. In a preferred embodiment, the 3' exon sequence is derived from the natural 3' exon sequence of a type I intron, i.e., the exon sequence located flanking (downstream) of the 3' splice site, or a continuous segment thereof beginning with the 5' terminal nucleotide of the natural 3' exon sequence. In some embodiments, the 3' exon sequence comprises the complete natural sequence of the exon adjacent to the 3' splice site of the type I intron, or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the complete natural sequence beginning with the 5' terminal nucleotide of the natural 3' exon sequence or a continuous segment thereof. In some embodiments, the 3' exon sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions compared to the entire natural 3' exon sequence or a continuous segment thereof beginning with the 5' terminal nucleotide of the natural 3' exon sequence.
[0062] In a preferred embodiment, the 5' exon sequence is derived from the natural 5' exon sequence of a type I intron, i.e., the exon sequence located flanking (upstream) of the 5' splice site, or a continuous segment thereof starting from the 5' terminal nucleotide of the natural 5' exon. In some embodiments, the 5' exon sequence includes the complete natural sequence of the exon near the 5' splice site of the type I intron, or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with the complete natural sequence of the exon near the 5' splice site of the type I intron, or a continuous segment thereof starting from the 5' terminal nucleotide of the natural 5' exon. In some implementations, the 5' exon sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or more nucleotide substitutions, deletions or additions compared to the entire natural 5' exon of the self-splicing intron or a continuous segment starting from the 3' terminal nucleotide of the natural 5' exon.
[0063] In some embodiments, the 5' exon sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 59 (ACGGACUU). In some embodiments, the 3' exon sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 60 (AAAAUCCGU). In one embodiment, the first RNA molecule comprises a 3' exon sequence having a nucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% identity with SEQ ID NO: 60 or SEQ ID NO: 3. In another embodiment, the first RNA molecule comprises a 5' exon sequence having a nucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% identity with SEQ ID NO: 59 or SEQ ID NO: 4.
[0064] In some embodiments, when the first RNA molecule comes into contact with the second RNA molecule, the first intron fragment, the second intron fragment, and the third intron fragment interact with each other to form the secondary structure of the paired (P) elements necessary for the formation of the scaffold domain and the catalytic domain. The P elements are also described in this application as "P regions" or "P stem-loop regions".
[0065] In view of this disclosure, the first and second RNA molecules can be prepared by any suitable method. In some embodiments, the first and second RNA molecules can be obtained by rearranging type I introns. For example, the first RNA molecule may comprise, from its 5' end to its 3' end: (1) segment 3, which comprises a 3' intron sequence and a 3' splice site operably linked and arranged in order from the 5' end to the 3' end of segment 3; (2) a 3' exon sequence; (3) a target RNA sequence; (4) a 5' exon sequence; and (5) segment 1, which comprises a 5' splice site and a second intron segment operably linked and arranged in order from the 5' end to the 3' end of segment 1. The second RNA molecule may comprise segment 2 having an intermediate intron sequence. One or more of segments 1-3 may be derived from portions of a type I intron obtained by dividing the intron with two break sites.
[0066] As used herein, a “break site” or “splitting site” refers to a site in a type I intron where an RNA fragment is cleaved to obtain the intron. This cleavage generates an RNA fragment of a type I intron, which can be rearranged or arranged to facilitate the circularization of the target RNA sequence. According to embodiments of this application, two break sites are selected to allow the generation of three separate parts of a type I intron (e.g., fragment 1, fragment 2, and fragment 3), which together maintain the ribozyme activity required for self-splicing even after rearrangement. The three separate fragments maintaining the overall conformation of the type I intron maintain the ribozyme activity required for self-splicing.
[0067] In some embodiments, one interruption site is located within the loop of the P6 stem-loop region of the type I intron, and another interruption site is located within the loop of the P2, P5, P8, or P9 stem-loop regions of the type I intron. In one embodiment, the three parts of the type I intron can be obtained by dividing the intron with a first interruption site located within the loop of the P2 stem-loop region and a second interruption site located within the loop of the P6 stem-loop region. In another embodiment, the three parts of the type I intron are obtained by dividing the intron with a first interruption site located within the loop of the P5 stem-loop region and a second interruption site located within the loop of the P6 stem-loop region. In yet another embodiment, the three parts of the type I intron are obtained by dividing the intron with a first interruption site located within the loop of the P6 stem-loop region and a second interruption site located within the loop of the P8 stem-loop region. In a further embodiment, the three parts of the type I intron are obtained by dividing the intron with a first interruption site located within the loop of the P6 stem-loop region and a second interruption site located within the loop of the P9 stem-loop region.
[0068] In some embodiments, fragment 1 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence extending from the 5' end of the type I intron to the left of the first interruption site, such as at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. Fragment 2 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence extending from the right of the first interruption site to the left of the second interruption site, such as at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. Fragment 3 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence extending from the right of the second interruption site to the 3' end of the type I intron, such as at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. When tracing the RNA strand from 5' to 3', the first and second break sites are numbered according to their occurrence.
[0069] According to an embodiment of this application, fragment 1 may contain a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequences: SEQ ID NO: 11, 14, 17, or 20. Fragment 2 may contain a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequences: SEQ ID NO: 12, 15, 18, 21, 23, 25, 27, or 29. Fragment 3 may contain a nucleotide sequence that has at least 80% identity with the following nucleotide sequences, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity: SEQ ID NO: 13, 16, 19, 22, 24, 26, 28, or 30.
[0070] Preferably, the RNA molecule pair of this application comprises fragment 1, fragment 2, and fragment 3 having the following nucleotide sequences: a. These are SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO: 13, respectively; b. These are SEQ ID NO: 14, SEQ ID NO: 15, and SEQ ID NO: 16, respectively; c. These are SEQ ID NO: 17, SEQ ID NO: 18, and SEQ ID NO: 19, respectively; d. These are SEQ ID NO: 20, SEQ ID NO: 21, and SEQ ID NO: 22, respectively; e. These are SEQ ID NO: 14, SEQ ID NO: 23, and SEQ ID NO: 24, respectively; f. These are SEQ ID NO: 14, SEQ ID NO: 25, and SEQ ID NO: 26, respectively; g. SEQ ID NO: 14, SEQ ID NO: 27, and SEQ ID NO: 28 respectively; or h. are SEQ ID NO: 14, SEQ ID NO: 29 and SEQ ID NO: 30 respectively.
[0071] In a preferred embodiment, the type I intron originates from... Anabaena cyanobacteria T4 phage or Nitrogenous Vibrio species BH72 Preferably, type I introns originate from... Anabaena cyanobacteria species pre-tRNA-Leu gene, T4 phage td gene or Nitrogenous Vibrio species BH72 The pre-tRNA-Ile gene.
[0072] In some embodiments, three RNA fragments derived from type I introns can be used in the preparation of RNA molecule pairs that retain the self-splicing activity for the preparation of circular RNA as described in this application. For example, an RNA molecule pair according to one embodiment of this application may comprise: 1) a first RNA molecule comprising the following elements operatively linked and arranged in order from the 5' end to the 3' end of the molecule: i) fragment 3 according to this application, comprising a 3'-intron sequence, a 3'-splicing site of a type I intron, and a 3' exon sequence of a type I intron; ii) a target ribonucleotide sequence; iii) fragment 1 according to this application, comprising a 5' exon sequence, a 5'-splicing site of a set of introns, and a 5'-intron sequence; and 2) a second RNA molecule comprising fragment 2 according to this application, said fragment 2 comprising an intermediate intron sequence of a type I intron. The first RNA molecule is a rearranged RNA molecule that allows the generation of a circular RNA containing the target nucleotide sequence through self-splicing of the rearranged RNA molecule with the aid of fragment 2.
[0073] In some implementations, the target nucleotide sequence encodes a protein. Because circRNAs are covalently linked head-to-tail and lack a 5′ end, they lack the 7-methylguanosine (m7G) cap that is involved in the initiation of mRNA translation. circRNAs must rely on cap-independent mechanisms, such as IRES, to initiate translation.
[0074] In a preferred embodiment, the target nucleotide sequence comprises at least one protein-coding sequence and a translation initiation element operatively linked thereto, such as an internal ribosome entry site (IRES). Here, "operatively linked" means that the translation initiation element (e.g., IRES) can mediate the translation of the encoded protein.
[0075] As used herein, IRES sequences may be selected from, but are not limited to, the IRES sequences of the following organisms: Taura syndrome virus, schistosomiasis virus, Tyler's encephalomyelitis virus, simian virus 40, red imported fire ant virus 1, cereal constriction virus, reticulovirus endothelial proliferation virus, Forman poliovirus 1, soybean looper virus, Kashmir bee virus, human rhinovirus 2, glass leafhopper virus-1, human immunodeficiency virus type 1, glass leafhopper virus-1, lice P virus, hepatitis C virus, hepatitis A virus, hepatitis B virus, foot and mouth disease virus, human enterovirus 71, equine rhinovirus, tea looper-like virus, encephalocarditis virus (EMCV), fruit fly C virus, and cruciferous tobacco virus. Virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, hibiscus yellow ringspot virus, swine fever virus, Aichi virus, cristatavivirus, diechovirus, foot-and-mouth disease virus, enterovirus, bopivirus, diresapivirus, Aalivirus, Ailurivirus, Ampivirus, Anativirus, foot-and-mouth disease virus, Avihepatatovirus, Avisivirus, Boosepivirus, Caridiovirus, cosaviruvirus s), Crahelivirus, Crohivirus, Dicipivirus, Diresapivirus, Enteroviruses, Eribovirus, Felipivirus, Fipivirus, Galivirus, Grusopivirus, Harkavirus, Hemipivirus, Hepatoviruses, Hummivirus, Kunsagivirus, Limnipivirus, Livupivirus, LudopivirusMalagasyvirus, Megrivirus, Mischivirus, Mosavirus, Mupivirus, Orivirus, Parabovirus, Parebovirus, Pastivirus, Passerivirus, Pemapivirus, Potamipivirus, Rabovirus, Rafivirus, Rohelivirus, Rosavirus savirus, sakovuvirus, salivirus, sapelovirus, senecavirus, shanbavirus, sicinvirus, symapivirus, teschovirus, torchivirus, totottorivirus, tremovirus, tropivirus, human FGF2, human SFTPA1, human AML1 / RUNX1, drosophila Antennae), Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAPl, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIFla, Human n.myc, Mouse Gtx, Human p27kipl, Human PDGF2 / c-sis, Human p53, Human Pim-1, Mouse Rbm3, Drosophila reaper, Canine Scamper, Drosophila Ubx, Human UNR, Mouse UtrA, Human VEGF-A, Human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, Human c-src, Human FGF-1, Simian picomavirus, Turnip crepe disease virus, eIF4G aptamer, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVA1 / 2). Wild-type IRES sequences can also be modified and used in this application. Preferably, IRES is CVB3, BRAV-1L, PV1L, CAV2L, BRAV-1,PV1 or CAV2.
[0076] In some embodiments, the target nucleotide sequence may include additional sequences that enhance translation. For example, the target nucleotide sequence may further include 5'-UTRs and 3'-UTRs located flanking the IRES and the protein-coding sequence. The target nucleotide sequence may further include a spacer between the IRES and the 5' end of the protein-coding sequence. The target nucleotide sequence may also include a spacer between the 3' end of the coding sequence and the 3'-UTR.
[0077] The target nucleotide sequence may include the coding sequence of any target protein. The encoded protein may be any protein used for therapeutic or diagnostic purposes. For example, the encoded protein may be a secreted protein, intracellular protein, or nuclear protein in eukaryotic cells. In some embodiments, the encoded protein is a human protein, antigen, antibody, or gene-editing enzyme (e.g., CRISPR nuclease). In other embodiments, the encoded protein is a chimeric antigen receptor, an immunomodulatory protein, a transcription factor, etc.
[0078] In other embodiments, the target nucleotide sequence is a non-protein-coding sequence. For example, a non-protein-coding sequence can be an antisense RNA, aptamer, guide RNA, or other non-protein-coding RNA present in any organism. The non-protein-coding sequence may or may not contain specific secondary structures.
[0079] In some embodiments, the length of the target nucleotide sequence is at least 10, 20, 40, 60, 80, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 10000, or 20000 nucleotides. In some embodiments, the length of the target nucleotide sequence is from about 10 to about 20000 nucleotides.
[0080] According to embodiments of this application, the RNA molecules used in this invention can be (e.g., chemically) unmodified, partially modified, or fully modified. In some embodiments, the first RNA molecule and / or the second RNA molecule of this application contain at least one modified nucleotide. As used herein, a “modified nucleotide” is a nucleotide other than a ribonucleotide (2'-hydroxynucleotide). In some embodiments, at least 50% (e.g., nucleotides) of the first RNA molecule and / or the second RNA molecule contain nucleotides that are (e.g., nucleotides that are chemically) unmodified, partially modified, or fully modified. ,At least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% are modified nucleotides. All positions in a given RNA molecule do not need to be uniformly modified. More than one modification may be incorporated into an RNA molecule. A modification at one nucleotide is independent of a modification at another nucleotide. In view of this disclosure, RNA molecules can be synthesized and / or modified using methods known in the art.
[0081] Modified nucleotides can be modified nucleosides, modified phosphates, and / or modified nucleotide linkages. For example, modified nucleotides include, but are not limited to, 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), and N1-methylpseudouridine (ψ). The nucleotide modification may be an acyclic sugar, thiophosphate linker, or thioaminophosphate linker of 5-methoxyuridine (5mol), 2'-deoxy-2'-fluororibose (2'-F), 2'-O-methylribose (2'-O-Me), or UNA nucleotide. In some embodiments, at least one of the nucleotide modifications is a cytidine modification, a uridine modification, or an adenosine modification. In some embodiments, at least one of the modified nucleotides is selected from the group consisting of 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (ψ), and N1-methylpseudouridine (ψ). ) and 5-methoxyuridine (5 moU).
[0082] In some preferred embodiments, the RNA molecule contains less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, and less than 1%. As used herein, the percentage of a specific nucleotide modification refers to the ratio of nucleotides in the sequence that have undergone that specific modification to nucleotides that can undergo that specific modification.
[0083] In some preferred embodiments, the RNA molecule is unmodified. In some embodiments, the RNA molecule does not contain nucleotide chemical modifications.
[0084] In some embodiments, the first RNA molecule of this application further includes a 3' end binding motif, and the second RNA molecule of this application further includes a 5' end binding motif, wherein the 3' end binding motif and the 5' end binding motif are complementary to form a double-stranded region. Preferably, the 3' end binding motif and the 5' end binding motif are completely complementary to each other.
[0085] In some embodiments, the length of the binding motif can be, for example, about 5-50 nucleotides, such as about 5-50, about 10-50, about 20-50, about 30-50, about 40-50, about 5-40, about 5-30, about 5-20, or about 10-20 nucleotides. In some embodiments, the homology arm is 20 nucleotides long. In some embodiments, the homology arm is 40 nucleotides long. In some embodiments, the homology arm is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long. In some embodiments, the homology arm is no more than 50, 45, 40, 35, 30, 25, or 20 nucleotides long. In some implementations, the length of the homologous arm is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides.
[0086] In some implementations, the length of the double-stranded region formed by the 3' end binding motif and the 5' end binding motif is about 5-50, about 5-40, about 5-30, about 5-20, or about 10-20 nucleotides, for example, the length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides. Preferably, the length of the double-stranded region is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides.
[0087] Another general aspect of this application relates to nucleic acids encoding the first RNA molecule and / or the second RNA molecule of this application. The first RNA molecule and the second RNA molecule of this application may be encoded by one DNA molecule or two separate DNA molecules.
[0088] This application also relates to nucleic acid vectors that comprise or encode a first RNA molecule and / or a second RNA molecule. As used herein, a "vector" or "nucleic acid vector" is a nucleic acid molecule used to carry genetic material into another cell, said nucleic acid molecule being replicated and / or expressed in said cell. In view of this disclosure, any vector known to those skilled in the art can be used. Examples of vectors include, but are not limited to, plasmids, viral vectors (bacteriophages, animal viruses, and plant viruses), granules, and artificial chromosomes (e.g., YAC). The vector can be a DNA vector or an RNA vector. In view of this disclosure, those skilled in the art can construct the vectors of this application using standard recombination techniques.
[0089] The vector in this application may be an expression vector. As used herein, the term "expression vector" refers to any type of genetic construct containing nucleic acid encoding RNA capable of being transcribed. Expression vectors include, but are not limited to, vectors for recombinant expression (e.g., DNA plasmids or viral vectors), and vectors for delivery of nucleic acids to a subject for expression in the tissues of said subject (e.g., DNA plasmids or viral vectors). Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of host cells to be transformed and the desired protein expression level.
[0090] The vector used in this application can also be a viral vector. Typically, a viral vector is a genetically engineered virus carrying modified viral DNA or RNA, wherein the modified viral DNA or RNA has been rendered non-infectious but still contains a viral promoter and transgenes, thereby allowing the transgenes to be translated via the viral promoter. Because viral vectors often lack infectious sequences, they require helper viruses or packaging lines for large-scale transfection. Examples of viral vectors that can be used include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, poxvirus vectors, enterovirus vectors, Venezuelan equine encephalitis virus vectors, Semliki forest virus vectors, tobacco mosaic virus vectors, modified Ankara vaccinia virus (MVA) vectors, lentiviral vectors, etc.
[0091] In a preferred embodiment, the nucleic acid vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of an RNA molecule. The operatively linked promoter allows for in vivo and / or in vitro transcription of the coding sequence. The promoter may be, for example, a T7 RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter.
[0092] In some embodiments, a "nucleic acid vector" can be a virus, plasmid, or cellular DNA derived from higher organisms, into which a foreign DNA fragment may be or has been inserted for cloning and / or expression purposes. In some embodiments, the vector can be stably maintained in the organism. The vector may contain, for example, an origin of replication, optional markers or reporter genes, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). Terms include linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, granules, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc.
[0093] In another general aspect, this application provides a method for preparing circular RNA, wherein the method comprises: i) Provide or obtain RNA molecule pairs according to the embodiments of this application; (ii) Adding a buffer solution to the RNA molecule pair to obtain a reaction mixture to allow the formation of the scaffold and catalytic domains of the type I introns, and (iii) Add GTP and divalent metal cations to the reaction mixture at a temperature that allows the formation of the circular RNA. Optionally, the method further includes: (iv) Harvest the circular RNA formed in step iii).
[0094] In some embodiments, the first and second RNA molecules in the RNA molecule pair are provided in a molar ratio of 1:1, 1:2, 1:3, 1:4 or 1:5 (preferably in a molar ratio of 1:1).
[0095] In view of this disclosure, any suitable buffer solution can be used in the methods of this application. For example, the buffer solution may contain 10-200 mM Tris-HCl at pH 6-8.5 (e.g., pH 6, 6.5, 7, 7.5, 8, or 8.5), such as 50 mM Tris-HCl. In a preferred embodiment, the reaction mixture contains 50 mM Tris-HCl at pH 7.5.
[0096] In some embodiments, GTP is added to the reaction mixture at a concentration of 100 nM to 2 mM.
[0097] In some embodiments, a divalent metal cation is added to the reaction mixture at a concentration of at least about 5 mM, for example, about 5 mM to about 550 mM, for example, at least about 5 mM, about 10 mM, about 15 mM, at least about 20 mM, at least about 30 mM, at least about 40 mM, at least about 50 mM, at least about 60 mM, at least about 70 mM, at least about 80 mM, at least about 90 mM, at least about 100 mM, at least about 125 mM, at least about 150 mM, at least about 175 mM, at least about 200 mM, at least about 250 mM, at least about 300 mM, at least about 350 mM, at least about 400 mM, at least about 450 mM, at least about 500 mM, at least about 550 mM, or higher. Any suitable divalent metal cation, such as Mg2+, can be used.
[0098] In some embodiments, the RNA molecule pair is encoded by one or more non-naturally occurring DNA molecules and can be produced using recombinant techniques (methods described in detail below, e.g., in vitro derivatization using DNA plasmids) or chemical synthesis. In some other embodiments, the RNA molecule pair is chemically synthesized. Within the scope of this application, the DNA molecule used to produce circular RNA may comprise a DNA sequence of a naturally occurring original nucleic acid sequence, a modified version thereof, or a DNA sequence encoding a synthetic polypeptide not normally found in nature (e.g., a chimeric molecule or fusion protein). A variety of techniques can be used to modify DNA and RNA molecules, including but not limited to classical mutagenesis and recombinant techniques, such as site-directed mutagenesis, chemical treatment of nucleic acid molecules to induce mutations, restriction enzyme cleavage of nucleic acid fragments, ligation of nucleic acid fragments, polymerase chain reaction (PCR) amplification and / or mutagenesis of selected regions of nucleic acid sequences, synthesis of oligonucleotide mixtures, and ligation of mixture groups to “build” mixtures and combinations thereof of nucleic acid molecules.
[0099] The inventors have surprisingly discovered that, compared to conventional PIE methods, the method for preparing circular RNA provided in this application results in higher self-splicing efficiency and lower byproducts, such as dimers or polymers. Self-splicing efficiency depends on the correct folding and assembly of type I intron RNA. Circulation efficiency can be measured by the ratio of resulting circular RNA over a given time period. Dimers or polymers are byproducts of oligomers containing two or more monomers of precursor RNA linked by intermolecular non-covalent bonds, which can be strong or weak. Self-splicing efficiency, circularization efficiency, and the percentage of dimers or polymers can be determined using methods known in the art, such as those described in the embodiments herein.
[0100] In some implementations, the method of this application results in an increase in self-splitting efficiency of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 200%, at least about 300%, at least about 400%, at least about 500%, or more compared to conventional PIE methods.
[0101] In some implementations, the method of this application does not produce or substantially does not produce dimer or polymer byproducts.
[0102] Another general aspect of this application relates to circular RNAs produced by the methods described in this application, pharmaceutical compositions thereof, and their uses. Circular RNAs contain a target nucleotide sequence encoding a protein or a non-protein.
[0103] In some embodiments, this application provides pharmaceutical compositions comprising the nucleic acid vector of this application, a combination of three RNA fragments of the type I intron of this application, RNA molecule pairs and / or circular RNA, and a pharmaceutically acceptable carrier.
[0104] The specific use of the composition may depend on the target nucleotide sequence. In some embodiments, the pharmaceutical composition is used to treat a disease in a subject. The specific disease to be treated may depend on the specific target nucleotide sequence.
[0105] Pharmaceutically acceptable carriers may include, but are not limited to, buffers, excipients, stabilizers, or preservatives. Examples of pharmaceutically acceptable carriers include physiologically compatible solvents, dispersion media, coating agents, antibacterial and antifungal agents, isotonic agents, and absorption delay agents, such as salts, buffers, carbohydrates, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants, or emulsifiers, or combinations thereof. The amount of a pharmaceutically acceptable carrier in a pharmaceutical composition can be determined experimentally based on the activity of the carrier and the desired properties of the formulation (e.g., stability and / or minimal oxidation).
[0106] In further embodiments, the pharmaceutical composition may optionally comprise one or more additional active substances, such as therapeutic and / or preventative active substances. The pharmaceutical compositions of this application may be sterile and / or pyrogen-free. General considerations in the formulation and / or manufacture of pharmaceutical preparations can be found, for example, in Remington: The Science and Practice of Pharmacy, 21st edition, Lippincott Williams & Wilkins, 2005 (incorporated herein by reference).
[0107] Although the description of the pharmaceutical compositions provided herein primarily relates to pharmaceutical compositions suitable for administration to humans, those skilled in the art will understand that such compositions are generally suitable for administration to any other animal, such as non-human animals, or non-human mammals. It is well understood that pharmaceutical compositions suitable for administration to humans may be modified to make them suitable for administration to a variety of animals, and that ordinary veterinary pharmacologists can design and / or perform such modifications simply through ordinary (if any) experiments. Subjects intended to administer the pharmaceutical compositions include, but are not limited to, humans and / or other primates; mammals, including commercially relevant mammals such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats; and / or birds, including commercially relevant birds such as poultry, chickens, ducks, geese, and / or turkeys.
[0108] Formulations of the pharmaceutical compositions described herein can be prepared by any method known in or subsequently developed in the field of pharmacology. Typically, such preparation methods involve combining the active ingredient with an excipient and / or one or more other auxiliary ingredients, and then, if desired and / or as desired, separating, shaping, and / or packaging the product.
[0109] Circular RNA can be delivered to the recipient via any suitable delivery system. Delivery systems can include timed-release, delayed-release, and sustained-release delivery systems, such that delivery of the composition occurs prior to sensitization at the treatment site, allowing sufficient time for sensitization to occur. The composition can be used in combination with other therapeutic agents or therapies. Such systems avoid repeated administration of the composition, thereby increasing convenience for both the recipient and physician, and may be particularly suitable for certain formulations of this application.
[0110] Release and delivery systems include polymer-based systems such as poly(lactide-glycolic acid), copolyoxalate, polyesteramide, polyorthoester, polycaprolactone, polyhydroxybutyrate, and polyanhydride. Microcapsules containing the aforementioned polymers of a drug are described, for example, in U.S. Patent No. 5,075,109. Delivery systems also include non-polymer systems, which are lipids, including sterols (e.g., cholesterol, cholesterol esters) and fatty acids or neutral fats (e.g., monoglycerides, diglycerides, and triglycerides); silicone rubber systems; peptide-based systems; hydrogel release systems; wax coatings; compressed tablets using conventional adhesives and excipients; partially fused implants; and so on. In some embodiments, lipid nanoparticles or polymers are used as delivery mediators for the therapeutic circRNAs (including RNA delivery to tissues) described herein.
[0111] Lipid nanoparticles have shown significant potential for use as delivery mediators for therapeutic RNA, including mRNA delivery to tissues (Oberli, MA et al., Lipid Nanoparticle Assisted mRNA Delivery for Potent Cancer Immunotherapy. Nano Lett. 17, 1326-1335 (2017); YanezArteta, M. et al., Successful reprogramming of cellular protein production through mRNA delivered by functionalized lipid nanoparticles. Proc. Natl. Acad. Sci. USA 115, E3351-E3360 (2018); and Kaczmarek, JC, Kowalski, PS and Anderson, DG Advances in the delivery of RNA therapeutics: from concept to clinical reality. Genome Med. 9, 60 (2017)). To evaluate the efficacy of lipid nanoparticles for in vivo circRNA delivery, purified circRNA was formulated into nanoparticles with the ionizable lipid cKK-E12 (Dong, Y. et al. Lipopeptide nanoparticles for potent and selective siRNA delivery in rodents and nonhuman primates. Proc. Natl. Acad. Sci. USA 111, 3955-3960 (2014); and Kauffman, KJ et al. Optimization of Lipid Nanoparticle Formulations for mRNA Delivery in Vivo with Fractional Factorial and Definitive Screening Designs. Nano Lett. 15, 7300-7306 (2015)).
[0112] This application further relates to a method for protein expression, the method comprising translating at least one region of the cyclic polynucleotide provided herein.
[0113] In some embodiments, the method for protein expression includes translating at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the total length of the circular RNA into a polypeptide. In some embodiments, the method for protein expression includes translating the circular RNA into a polypeptide of at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 50 amino acids, at least 100 amino acids, at least 150 amino acids, at least 200 amino acids, at least 250 amino acids, at least 300 amino acids, at least 400 amino acids, at least 500 amino acids, at least 600 amino acids, at least 700 amino acids, at least 800 amino acids, at least 900 amino acids, or at least 1000 amino acids. In some embodiments, the method for protein expression includes translating circular RNA into a polypeptide of about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 50 amino acids, about 100 amino acids, about 150 amino acids, about 200 amino acids, about 250 amino acids, about 300 amino acids, about 400 amino acids, about 500 amino acids, about 600 amino acids, about 700 amino acids, about 800 amino acids, about 900 amino acids, or about 1000 amino acids. In some embodiments, the method includes translating circular RNA into a continuous polypeptide, a discrete polypeptide, or both as provided herein.
[0114] In some embodiments, translation of at least one region of the circular RNA occurs in vitro, for example, in rabbit reticulocyte lysates. In some embodiments, translation of at least one region of the circular RNA occurs in vivo, for example, after transfection of eukaryotic cells or transformation of prokaryotic cells such as bacteria.
[0115] In some embodiments, this disclosure provides a method for in vivo expression of one or more expressed sequences in a subject, the method comprising: administering a circular RNA to cells of the subject, wherein the circular RNA comprises one or more expressed sequences; and expressing one or more expressed sequences from the circular RNA in the cells. In some embodiments, the circular RNA is configured such that the expression of one or more expressed sequences in the cells at a later time point is equal to or greater than the expression in the cells at an earlier time point. In some embodiments, the circular RNA is configured such that the expression of one or more expressed sequences in the cells does not decrease by more than about 40% over a period of at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days. In some embodiments, the circular RNA is configured such that the expression of one or more expressed sequences in the cells is maintained at a level with a change not exceeding about 40% for at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days. In some embodiments, the administration of the circular RNA is performed using any of the delivery methods described herein. In some embodiments, the circular RNA is administered to the subject via intravenous injection. In some implementations, the administration of circular RNA includes, but is not limited to, prenatal administration, neonatal administration, postnatal administration, oral administration, administration by injection (e.g., intravenous, intra-arterial, intraperitoneal, intradermal, subcutaneous, and intramuscular), ocular administration, and intranasal administration.
[0116] In some embodiments, methods for protein expression include modification, folding, or other post-translational modifications of the translation product. In some embodiments, methods for protein expression include in vivo post-translational modifications, such as via cellular mechanisms.
[0117] All references and publications cited in this article are incorporated herein by reference.
[0118] The above-described implementation schemes can be combined to achieve the aforementioned functional characteristics. This is also illustrated by the following embodiments, which illustrate exemplary combinations and implemented functional characteristics.
[0119] Implementation Plan The following are exemplary numbered embodiments of the present invention.
[0120] Implementation Scheme 1. A combination of three short RNA fragments derived from a type I intron, wherein the three short RNA fragments are divided by two interruption sites located within the type I intron.
[0121] Implementation Scheme 2. A combination of three short RNA fragments as described in Implementation Scheme 1, wherein one interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region, and the other interruption site is located in the P2, P5, P8, or P9 region of the type I intron, for example, within the loop of the P2, P5, P8, or P9 region.
[0122] Implementation Scheme 3. A combination of three short RNA fragments as described in Implementation Scheme 2, wherein one interruption site is located in the P2 region of the type I intron, for example, within the loop of the P2 region, and the other interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region.
[0123] Implementation Scheme 4. A combination of three short RNA fragments as described in Implementation Scheme 2, wherein one interruption site is located in the P5 region of the type I intron, for example, within the loop of the P5 region, and the other interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region.
[0124] Implementation Scheme 5. A combination of three short RNA fragments as described in Implementation Scheme 2, wherein one interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region, and the other interruption site is located in the P8 region of the type I intron, for example, within the loop of the P8 region.
[0125] Implementation Scheme 6. A combination of three short RNA fragments as described in Implementation Scheme 2, wherein one interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region, and the other interruption site is located in the P9 region of the type I intron, for example, within the loop of the P9 region.
[0126] Implementation Scheme 7. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, the combination comprising: i) Fragment 1 (F1), which includes a 5'-splicing site extending from the 5' end of the type I intron to the left of the first interruption site located within the type I intron; ii) Fragment 2 (F2), which extends from the right side of the first interruption site to the left side of the second interruption site located within the type I intron; and iii) Fragment 3 (F3), which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron.
[0127] Implementation Scheme 8. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, the combination comprising: i) Fragment 1 (F1), which includes a 5'-splicing site, extending from the 5' end of the type I intron to the left of a first interruption site located within the loop of the P2 region, for example; ii) Fragment 2 (F2), which extends from the right side of the first interruption site to the left side of the second interruption site located within the loop of region P6, for example, region P6; and iii) Fragment 3 (F3), which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron.
[0128] Implementation Scheme 9. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, the combination comprising: i) Fragment 1 (F1), which includes a 5'-splicing site, extending from the 5' end of the type I intron to the left of a first interruption site located within the loop of the P5 region, for example; ii) Fragment 2 (F2), which extends from the right side of the first interruption site to the left side of the second interruption site located within the loop of region P6, for example, region P6; and iii) Fragment 3 (F3), which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron.
[0129] Implementation Scheme 10. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, the combination comprising: i) Fragment 1 (F1), which includes a 5'-splicing site, extending from the 5' end of the type I intron to the left of a first interruption site located within the loop of the P6 region, for example; ii) Fragment 2 (F2), which extends from the right side of the first interruption site to the left side of the second interruption site located within a ring in region P8, for example, region P8; and iii) Fragment 3 (F3), which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron.
[0130] Implementation Scheme 11. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, the combination comprising: i) Fragment 1 (F1), which includes a 5'-splicing site, extending from the 5' end of the type I intron to the left of a first interruption site located within a loop in region P6, for example, region P6; ii) Fragment 2 (F2), which extends from the right side of the first interruption site to the left side of the second interruption site located within a ring in region P9, for example, region P9; and iii) Fragment 3 (F3), which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron.
[0131] Implementation Scheme 12. A combination of three short RNA fragments as described in any one of Implementation Schemes 7-11, wherein the 5'-splicing site is located within the natural 5' exon region of the type I intron, and the 3'-splicing site is located within the natural 3' exon region of the type I intron.
[0132] Implementation Scheme 13. A combination of three short RNA fragments as described in any one of Implementation Schemes 1-12, wherein the type I introns are selected from the IC3 subgroup of type I introns.
[0133] Implementation Scheme 14. A combination of three short RNA fragments as described in any one of Implementation Schemes 1-13, wherein the type I introns are derived from cyanobacteria of the genus Anabaena. 、 T4 phage or species of the genus *Vibrio* The type I intron of BH72m; preferably, the type I intron is derived from the type I intron of cyanobacteria of the genus Anabaena.
[0134] Implementation Scheme 15. A combination of three short RNA fragments as described in any one of Implementation Schemes 1-14, wherein the type I introns are derived from cyanobacteria of the genus Anabaena. Species pre-tRNA-Leu gene, td gene of T4 phage or Nitrogenous Vibrio species kind The pre-tRNA-Ile gene of BH72.
[0135] Implementation Scheme 16. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, and 100% sequence identity with SEQ ID NO: 11, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, and 100% sequence identity with SEQ ID NO: 12, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, and 100% sequence identity with SEQ ID NO: 13.
[0136] Implementation Scheme 17. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 14, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 15, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 16.
[0137] Implementation Scheme 18. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 17, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 18, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 19.
[0138] Implementation Scheme 19. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 20, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 21, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 22.
[0139] Implementation Scheme 20. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 14, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 23, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 24.
[0140] Implementation Scheme 21. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 14, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 25, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 26.
[0141] Implementation Scheme 22. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 14, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 27, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 28.
[0142] Implementation Scheme 23. A combination of three short RNA fragments as described in Implementation Scheme 1 or 2, wherein fragment 1 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 14, fragment 2 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 29, and fragment 3 is a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100% sequence identity with the following: SEQ ID NO: 30.
[0143] Implementation Scheme 24. An RNA molecule pair for preparing circular RNA, said RNA molecule pair comprising: (1) A first RNA molecule comprising the following elements, said elements being operatively linked and arranged in order from the 5' end to the 3' end of said molecule: i) The 3' intron sequence of type I introns ii) The 3′ splice site of the type I intron, iii) Optionally, the 3' exon sequence of the type I intron, iv) Target RNA sequence, v) Optionally, the 5' exon sequence of the type I intron, vi) The 5′ splice site of the type I intron, and vii) The 5′ intron sequence of the type I intron, and (2) A second RNA molecule containing the intermediate intron sequence of the type I intron. When the first RNA molecule comes into contact with the second RNA molecule, the scaffold and catalytic domains of the type I intron are formed, thereby allowing the generation of the circular RNA containing the 3' exon sequence, the target RNA sequence, and the 5' exon sequence.
[0144] Implementation Scheme 24a. The RNA molecule pair as described in Implementation Scheme 24, wherein the 3' intron sequence comprises the R and S sequences of the type I intron, the 5' intron sequence comprises the internal guide sequence (IGS) and P sequence of the type I intron, and the intermediate intron sequence comprises the Q sequence of the type I intron.
[0145] Implementation Scheme 24b. An RNA molecule pair as described in Implementation Scheme 24 or 24a, wherein the first RNA molecule comprises from the 5' end to the 3' end of the molecule, a) Fragment 3, which includes the 3' intron sequence and the 3' splice site, which are operatively connected and arranged in order from the 5' end to the 3' end of fragment 3; b) Optionally, the 3' exon sequence; c) The target RNA sequence; d) Optionally, the 5' exon sequence; e) Fragment 1, comprising the 5' splice site and the 5' intron sequence operably connected and arranged in order from the 5' end to the 3' end of fragment 1; and The second RNA molecule contains segment 2 having the intermediate intron sequence.
[0146] Implementation scheme 24c. The RNA molecule pair as described in implementation scheme 24b, wherein fragments 1-3 each contain the nucleotide sequence or a variant thereof of the three parts of the type I intron obtained by dividing the type I intron with two interrupt sites.
[0147] Implementation scheme 24d. The RNA molecule pair as described in implementation scheme 24c, wherein the three parts of the type I intron are obtained by dividing the intron by an interruption site located in the loop of the P6 stem-loop region of the type I intron and another interruption site located in the loop of the P2, P5, P8 or P9 stem-loop regions of the type I intron.
[0148] Implementation Scheme 24e. An RNA molecule pair, the combination of which retains self-splicing activity to prepare circular RNA, wherein said RNA molecule pair comprises: a) A rearranged RNA molecule comprising elements operatively linked to each other and, in some embodiments, arranged in a 5' to 3' orientation in the following order: i) Fragment 3, which includes a 3'-splicing site extending from the right side of the second interruption site to the 3' end of the type I intron; ii) Optionally, the 3' exon sequence; iii) Target nucleotide sequence; iv) Optionally, the 5' exon sequence; and v) Fragment 1, which includes a 5'-splicing site extending from the 5' end of the type I intron to the left of a first interruption site located within the type I intron; and b) Segment 2, which extends from the right side of the first interruption site to the left side of the second interruption site located within the type I intron; Fragments 1-3 were obtained by separating them with two interruption sites, one of which was located in the P6 region of the type I intron, and the other interruption site was located in the P2, P5, P8 or P9 region of the type I intron. The rearranged RNA molecule allows for the generation of a circular RNA containing the target nucleotide sequence through self-splicing of the rearranged RNA molecule with the aid of fragment 2; and When tracing the RNA strand from 5' to 3', the first and second interrupt sites are numbered according to their occurrence.
[0149] Implementation Scheme 25. An RNA molecule pair as described in any one of Implementation Schemes 24-24e, wherein the 3' exon sequence is derived from the natural 3' exon of the type I intron.
[0150] Implementation Scheme 25a. An RNA molecule pair as described in any one of Implementation Schemes 24-25, wherein the 3' exon sequence has 100% identity with the natural 3' exon sequence of the type I intron, or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 99%, or at least 99% identity with the full-length sequence of the natural 3' exon sequence.
[0151] Implementation Scheme 25b. An RNA molecule pair as described in any one of Implementation Schemes 24-25a, wherein the 3' exon sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or more nucleotide substitutions, deletions or additions compared to the natural 3' exon sequence.
[0152] Implementation Scheme 25c. An RNA molecule pair as described in any one of Implementation Schemes 24-25b, wherein the 3' exon sequence comprises SEQ ID NO: 60 (AAAAUCCGU).
[0153] Implementation scheme 25d. An RNA molecule pair as described in any one of implementation schemes 24-25b, wherein the 3' exon sequence consists of: SEQ ID NO: 60 (AAAAUCCGU).
[0154] Implementation Scheme 25e. An RNA molecule pair as described in any one of Implementation Schemes 24-25b, wherein the 3' exon sequence comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% identity with the following: SEQ ID NO: 60.
[0155] Implementation scheme 25f. An RNA molecule pair as described in any one of implementation schemes 24-25b, wherein the 3' exon sequence comprises SEQ ID NO: 3.
[0156] Implementation Scheme 25g. An RNA molecule pair as described in any one of Implementation Schemes 24-25b, wherein the 3' exon sequence consists of: SEQ ID NO: 3.
[0157] Implementation scheme 25h. An RNA molecule pair as described in any one of implementation schemes 24-25b, wherein the 3' exon sequence comprises a nucleotide sequence having at least 80%, 85%, 90%, 95% or 100% identity with the following: SEQ ID NO: 3.
[0158] Implementation scheme 25i. An RNA molecule pair as described in any one of implementation schemes 24-25h, wherein the first RNA molecule does not contain the 3' exon sequence.
[0159] Implementation scheme 25j. An RNA molecule pair as described in any one of implementation schemes 24-25h, wherein the first RNA molecule contains the 3' exon sequence.
[0160] Implementation Scheme 26. An RNA molecule pair as described in any one of Implementation Schemes 24-25j, wherein the 5' exon sequence is derived from the natural 5' exon of the type I intron.
[0161] Implementation Scheme 26a. An RNA molecule pair as described in any one of Implementation Schemes 24-26, wherein the 5' exon sequence has 100% identity with the natural 5' exon sequence of the type I intron, or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the full-length sequence of the natural 5' exon sequence.
[0162] Implementation Scheme 26b. An RNA molecule pair as described in any one of Implementation Schemes 24-26a, wherein the 5' exon sequence has 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 or more nucleotide substitutions, deletions or additions compared to the natural 5' exon sequence.
[0163] Implementation Scheme 26c. An RNA molecule pair as described in any one of Implementation Schemes 24-26b, wherein the 5' exon sequence comprises SEQ ID NO: 59 (ACGGACUU).
[0164] Implementation Scheme 26d. An RNA molecule pair as described in any one of Implementation Schemes 24-26b, wherein the 5' exon sequence consists of: SEQ ID NO: 59 (ACGGACUU).
[0165] Implementation Scheme 26e. An RNA molecule pair as described in any one of Implementation Schemes 24-26b, wherein the 5' exon sequence comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% identity with the following: SEQ ID NO: 59.
[0166] Implementation scheme 26f. An RNA molecule pair as described in any one of implementation schemes 24-26b, wherein the 5' exon sequence comprises SEQ ID NO: 4.
[0167] Implementation scheme 26g. An RNA molecule pair as described in any one of implementation schemes 24-26b, wherein the 5' exon sequence consists of: SEQ ID NO: 4.
[0168] Implementation Scheme 26h. An RNA molecule pair as described in any one of Implementation Schemes 24-26b, wherein the 5' exon sequence comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, or 100% identity with the following: SEQ ID NO: 4. Implementation scheme 26i. An RNA molecule pair as described in any one of implementation schemes 24-26h, wherein the first RNA molecule does not contain the 5' exon sequence.
[0169] Implementation scheme 26j. The RNA molecule pair as described in implementation scheme 26i, wherein the first RNA molecule also does not contain the 3' exon sequence.
[0170] Implementation scheme 26k. An RNA molecule pair as described in any one of implementation schemes 24-26h, wherein the first RNA molecule contains the 5' exon sequence.
[0171] Implementation scheme 261. An RNA molecule pair as described in implementation scheme 26k, wherein the first RNA molecule further contains the 3' exon sequence.
[0172] Implementation Scheme 27. An RNA molecule pair as described in any one of Implementation Schemes 24-261, wherein the first interruption site is located in the P2 region of the type I intron, for example, within a loop of the P2 region, and the other interruption site is located in the P6 region of the type I intron, for example, within a loop of the P6 region.
[0173] Implementation Scheme 28. An RNA molecule pair as described in any one of Implementation Schemes 24-261, wherein the first interrupt site is located in the P5 region of the type I intron, for example, within a loop of the P5 region, and the second interrupt site is located in the P6 region of the type I intron, for example, within a loop of the P6 region.
[0174] Implementation Scheme 29. An RNA molecule pair as described in any one of Implementation Schemes 24-261, wherein the first interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region, and the other interruption site is located in the P8 region of the type I intron, for example, within the loop of the P8 region.
[0175] Implementation Scheme 30. An RNA molecule pair as described in any one of Implementation Schemes 24-261, wherein the first interruption site is located in the P6 region of the type I intron, for example, within the loop of the P6 region, and the other interruption site is located in the P9 region of the type I intron, for example, within the loop of the P9 region.
[0176] Implementation Scheme 31. An RNA molecule pair as described in any one of Implementation Schemes 24-30, wherein fragments 1-3 are as defined in any one of Implementation Schemes 16-23.
[0177] Implementation Scheme 32. An RNA molecule pair as described in any one of Implementation Schemes 24-31, wherein fragment 1 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence from the 5' end of the type I intron to the left of the first interrupt site, for example at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
[0178] Implementation Scheme 32a. An RNA molecule pair as described in any one of Implementation Schemes 24-32, wherein fragment 1 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequence: SEQ ID NO: 11, 14, 17, or 20.
[0179] Implementation Scheme 32b. An RNA molecule pair as described in any one of Implementation Schemes 24-32a, wherein fragment 2 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence from the right side of the first interrupt site to the left side of the second interrupt site, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
[0180] Implementation scheme 32c. An RNA molecule pair as described in any one of embodiments 24-32b, wherein fragment 2 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequence: SEQ ID NO: 12, 15, 18, 21, 23, 25, 27, or 29.
[0181] Implementation scheme 32d. An RNA molecule pair as described in any one of embodiments 24-32c, wherein fragment 3 comprises a nucleotide sequence having at least 80% identity with the nucleotide sequence from the right side of the second interrupt site to the 3' end of the type I intron, for example at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
[0182] Implementation scheme 32e. An RNA molecule pair as described in any one of embodiments 24-32d, wherein fragment 3 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequences: SEQ ID NO: 13, 16, 19, 22, 24, 26, 28, or 30.
[0183] Implementation Scheme 32f. The RNA molecule pair as described in any one of Implementation Schemes 24-32e, wherein fragment 1, fragment 2, and fragment 3 comprise the following nucleotide sequences: a. These are SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO: 13, respectively; b. These are SEQ ID NO: 14, SEQ ID NO: 15, and SEQ ID NO: 16, respectively; c. These are SEQ ID NO: 17, SEQ ID NO: 18, and SEQ ID NO: 19, respectively; d. These are SEQ ID NO: 20, SEQ ID NO: 21, and SEQ ID NO: 22, respectively; e. These are SEQ ID NO: 14, SEQ ID NO: 23, and SEQ ID NO: 24, respectively; f. These are SEQ ID NO: 14, SEQ ID NO: 25, and SEQ ID NO: 26, respectively; g. SEQ ID NO: 14, SEQ ID NO: 27, and SEQ ID NO: 28 respectively; or h. are SEQ ID NO: 14, SEQ ID NO: 29 and SEQ ID NO: 30 respectively.
[0184] Implementation Scheme 33. An RNA molecule pair as described in any one of Implementation Schemes 24-32f, wherein the target nucleotide sequence comprises at least one protein-coding sequence and an internal ribosome entry site (IRES) operatively linked thereto.
[0185] Implementation Scheme 34. The RNA molecule pair as described in Implementation Scheme 33, wherein the IRES is derived from any of the following: Taura syndrome virus, schistosomiasis virus, Tyler's encephalomyelitis virus, simian virus 40, red imported fire ant virus 1, cereal constriction virus, reticulovirus endothelial proliferation virus, Forman poliovirus 1, soybean looper virus, Kashmir bee virus, human rhinovirus 2, glass leafhopper virus-1, human immunodeficiency virus type 1, glass leafhopper virus-1, lice P virus, hepatitis C virus, hepatitis A virus, hepatitis B virus, foot and mouth disease virus, human enterovirus 71, equine rhinovirus, tea looper-like virus, encephalocarditis virus (EMCV), fruit fly C virus, cruciferous tobacco virus. Virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, hibiscus yellow ringspot virus, swine fever virus, Aichi virus, cristatovirus, diechovirus, apthovirus, enterovirus, Borna virus, diresapivirus, Aalivirus, Ailurivirus, Ampivirus, Anativirus, Aphthovirus, Avihepatatovirus, Avisivirus, Boosepivirus, Caridioviruvirus s), cosavirus, Crahelivirus, Crohivirus, Dicipivirus, Diresapivirus, Enteroviruses, Eribovirus, Felipivirus, Fipivirus, Galivirus, Grusopivirus, Harkavirus, Hemipivirus, Hepatoviruses, Hummivirus, Kunsagivirus, Limnipivirus, LivupivirusLudopivirus, Malagasyvirus, Megrivirus, Mischivirus, Mosavirus, Mupivirus, Orivirus, Parabovirus, Parebovirus, Pastivirus, Passerivirus, Pemapivirus, Potamipivirus, Rabovirus, Rafivirus, Rohelivirus s), Rosavirus, Sakovuvirus, Salivirus, Sapelovirus, Senecavirus, Shanbavirus, Sicinvirus, Symapivirus, Teschovirus, Torchivirus, Totottorivirus, Tremovirus, Tropivirus, Human FGF2, Human SFTPA1, Human AML1 / RUNX1, Drosophila antennae Antennae), Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAPl, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIFla, Human n.myc, Mouse Gtx, Human p27kipl, Human PDGF2 / c-sis, Human p53, Human Pim-1, Mouse Rbm3, Drosophila reaper, Canine Scamper, Drosophila Ubx, Human UNR, Mouse UtrA, Human VEGF-A, Human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, Human c-src, Human FGF-1, Simian picomavirus, Turnip crepe disease virus, eIF4G aptamer, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVA1 / 2). Wild-type IRES sequences can also be modified and used in this application. Preferably, the IRES is CVB3,BRAV-1L, PV1L, CAV2L, BRAV-1, PV1, or CAV2.
[0186] Implementation Scheme 35. The RNA molecule pair as described in Implementation Scheme 34, wherein the IRES comprises or consists of the nucleotide sequence of SEQ ID NO: 51.
[0187] Implementation Scheme 36. An RNA molecule pair as described in any one of Implementation Schemes 24-35, wherein the target nucleotide sequence is a protein-coding sequence.
[0188] Implementation Scheme 37. The RNA molecule pair as described in Implementation Scheme 36, wherein the protein coding region may encode secretory proteins, intracellular proteins, and nucleoproteins in eukaryotic cells.
[0189] Implementation Scheme 38. The RNA molecule pair as described in Implementation Scheme 37, wherein the protein coding region may encode human proteins, antigens, antibodies, gene editing enzymes such as CRISPR nucleases.
[0190] Implementation Scheme 39. The RNA molecule pair as described in Implementation Scheme 38, wherein the protein coding region may encode a chimeric antigen receptor, an immunomodulatory protein, or a transcription factor.
[0191] Implementation Scheme 40. An RNA molecule pair as described in any one of Implementation Schemes 24-35, wherein the target nucleotide sequence is a non-protein coding sequence.
[0192] Implementation Scheme 41. The RNA molecule pair as described in Implementation Scheme 40, wherein the non-protein coding sequence is selected from antisense RNA, aptamers, guide RNA or non-protein coding RNA present in any organism.
[0193] Implementation Scheme 42. An RNA molecule pair as described in any one of Implementation Schemes 24-41, wherein the length of the target nucleotide sequence is at least 10, 20, 40, 60, 80, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 10000, or 20000 nucleotides.
[0194] Implementation Scheme 43. An RNA molecule pair as described in any one of Implementation Schemes 24-42, wherein the first RNA molecule, the second RNA molecule, or the rearranged RNA molecule comprises modified nucleotides.
[0195] Implementation Scheme 43a. The RNA molecule pair as described in Implementation Scheme 43, wherein the modified nucleotide comprises a modified nucleoside.
[0196] Implementation Scheme 43b. An RNA molecule pair as described in Implementation Scheme 43 or 43a, wherein the modified nucleotide comprises a modified phosphate.
[0197] Implementation Scheme 43c. An RNA molecule pair as described in any one of Implementation Schemes 43-43b, wherein the modified nucleotide comprises a modified nucleotide linker.
[0198] Implementation scheme 43d. An RNA molecule pair as described in any one of implementation schemes 43-43c, wherein the modified nucleotide comprises cytidine modification, uridine modification, guanine modification or adenosine modification.
[0199] Implementation Scheme 43e. The RNA molecule pair as described in any one of Implementation Schemes 43-43d, wherein the modified nucleotide includes, but is not limited to, 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (ψ), etc. ), 5-methoxyuridine (5 moU), 2'-deoxy-2'-fluoro-ribose (2'-F), 2'-O-methylribose (2'-O-Me), acyclic sugars of UNA nucleotides, thiophosphate linkages, or thioaminophosphate linkages.
[0200] Implementation Scheme 43f. An RNA molecule pair as described in any of Implementation Schemes 43-43e, wherein the first RNA molecule, the second RNA molecule, or the rearranged RNA molecule is chemically modified, partially modified, or fully modified; preferably, the rearranged RNA molecule contains at least one nucleotide modification; more preferably, up to 100% of the nucleotides of the rearranged RNA molecule are modified.
[0201] Implementation Scheme 44. An RNA molecule pair as described in any one of Implementation Schemes 43-43f, wherein the at least one nucleotide modification is a cytidine modification, a uridine modification, or an adenosine modification. In some embodiments, the at least one nucleotide modification is selected from the group consisting of: 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (ψ), etc. ) and 5-methoxyuridine (5 moU).
[0202] Implementation Scheme 45. The RNA molecule pair as described in Implementation Scheme 44, wherein the rearranged RNA molecule contains less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, or less than 1%.
[0203] Implementation Scheme 46. An RNA molecule pair as described in any one of Implementation Schemes 24-42, wherein the first RNA molecule, the second RNA molecule, or the rearranged RNA molecule is unmodified.
[0204] Implementation Scheme 47. An RNA molecule pair as described in any one of Implementation Schemes 24-46, wherein the first RNA molecule or the rearranged RNA molecule further comprises a 3' end binding motif, and the second RNA molecule or fragment 2 (F2) further comprises a 5' end binding motif, wherein the 3' end binding motif and the 5' end binding motif are complementary to form a double-stranded region.
[0205] Implementation scheme 47a. The RNA molecule pair as described in implementation scheme 47, wherein the 3' end binding motif and the 5' end binding motif are completely complementary to each other.
[0206] Implementation Scheme 48. An RNA molecule pair as described in any one of Implementation Schemes 24-47a, wherein the length of the 3' end binding motif or the 5' end binding motif is about 5-50, about 10-50, about 20-50, about 30-50, about 40-50, about 5-40, about 5-30, about 5-20, or about 10-20 nucleotides.
[0207] Implementation Scheme 48a. The RNA molecule pair as described in Implementation Scheme 48, wherein the length of the 3' end binding motif or the 5' end binding motif is about 10, 15, or 19 nucleotides.
[0208] Implementation Scheme 49. A nucleic acid vector for preparing circular RNA molecules, wherein the vector encodes a first RNA molecule, a second RNA molecule, or a rearranged RNA molecule as defined in any one of Implementation Schemes 24-42.
[0209] Implementation Scheme 50. The nucleic acid vector of Implementation Scheme 49, wherein the vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of the first RNA molecule, the second RNA molecule, or the rearranged RNA molecule.
[0210] Implementation Scheme 51. The nucleic acid vector as described in Implementation Scheme 50, wherein the promoter is selected from the T7 RNA polymerase promoter, the T6 viral RNA polymerase promoter, the SP6 viral RNA polymerase promoter, the T3 viral RNA polymerase promoter, or the T4 viral RNA polymerase promoter.
[0211] Implementation Scheme 52. The nucleic acid vector of any one of Implementation Schemes 50-51, wherein the vector is selected from linear DNA fragments, plasmid vectors, viral vectors, granules, bacterial artificial chromosomes (BACs) or yeast artificial chromosomes (YACs).
[0212] Implementation Scheme 53. A method for preparing circular RNA, wherein the method comprises: i) Providing or obtaining rearranged RNA molecules as defined in any one of embodiments 24-48a by transcription from the nucleic acid vector described in this application; ii) Anneal the rearranged RNA molecules and fragment 2 in a buffer solution, then add GTP and divalent metal cations at the temperature at which RNA cyclization occurs; and iii) Harvest the circular RNA obtained in step ii).
[0213] Implementation Scheme 53a. A method for preparing circular RNA, wherein the method comprises: i) Provide or obtain an RNA molecule pair as described in any one of embodiments 24-48a; ii) Add a buffer solution to the RNA molecule pair to obtain a reaction mixture to allow the formation of the scaffold and catalytic domains of the type I introns, and iii) Add GTP and divalent metal cations to the reaction mixture at a temperature that allows the formation of the circular RNA. Optionally, the method further includes: iv) Harvest the circular RNA formed in step iii).
[0214] Implementation Scheme 54. The method for preparing circular RNA as described in Implementation Scheme 53, wherein the rearranged RNA molecule and fragment 2 are annealed in a buffer solution at a molar ratio of 1:1; 1:2; 1:3; 1:4; or 1:5, preferably at a molar ratio of 1:1.
[0215] Implementation Scheme 54a. A method for preparing circular RNA as described in Implementation Scheme 53a, wherein the first RNA molecule and the second molecule are provided in the reaction mixture in a molar ratio of 1:1, 1:2, 1:3, 1:4 or 1:5.
[0216] Implementation Scheme 55. A method for preparing circular RNA as described in any one of Implementation Schemes 53-54a, wherein the buffer solution comprises 10-200 mM Tris-HCl.
[0217] Implementation Scheme 55a. The method of Implementation Scheme 55, wherein the buffer solution comprises 50 mM Tris-HCl at pH 6-8.5, such as pH 6, 6.5, 7, 7.5, 8 or 8.5.
[0218] Implementation Scheme 55b. The method as described in Implementation Scheme 55a, wherein the buffer solution contains 50 mM Tris-HCl, pH 7.5.
[0219] Implementation Scheme 56. A method for preparing circular RNA as described in any one of Implementation Schemes 53-55b, wherein the concentration of GTP in the reaction mixture is in the range of 100 nM to 2 mM.
[0220] Implementation Scheme 57. A method for preparing circular RNA as described in any one of Implementation Schemes 53-56, wherein the divalent metal cation is Mg2+, and the concentration of the divalent metal cation in the reaction mixture is at least about 5 mM, for example about 5 mM to about 550 mM, for example at least about 5 mM, about 10 mM, about 15 mM, at least about 20 mM, at least about 30 mM, at least about 40 mM, at least about 50 mM, at least about 60 mM, at least about 70 mM, at least about 80 mM, at least about 90 mM, at least about 100 mM, at least about 125 mM, at least about 150 mM, at least about 175 mM, at least about 200 mM, at least about 250 mM, at least about 300 mM, at least about 350 mM, at least about 400 mM, at least about 450 mM, at least about 500 mM, at least about 550 mM or higher.
[0221] Implementation Scheme 58. A circular RNA, which is produced by any one of Implementation Schemes 53-57.
[0222] Implementation Scheme 59. A eukaryotic cell comprising circular RNA as described in Implementation Scheme 58.
[0223] Implementation Scheme 60. A cell population comprising circular RNA as described in Implementation Scheme 58.
[0224] Implementation Scheme 61. A composition comprising an effective amount of the circular RNA as described in Implementation Scheme 58 and a pharmaceutically acceptable carrier.
[0225] Embodiment 61a. The composition of embodiment 61, wherein the pharmaceutically acceptable carrier comprises lipids, polymers, or lipid-polymer hybrids, such as lipid nanoparticles (LNPs), lipid microparticles, lipid suspensions, or liposomes.
[0226] Implementation Scheme 62. A composition comprising: a) an effective amount of circular RNA as described in Implementation Scheme 58, and b) a delivery medium for delivering the circular RNA to cells, said delivery medium comprising a nanocarrier selected from the group consisting of lipids, polymers, and lipid-polymer hybrids.
[0227] Implementation Scheme 63. A composition comprising a cell expressing circular RNA and a combination of one or more pharmaceutically or physiologically acceptable carriers, excipients, or diluents, wherein the circular RNA is the circular RNA as described in Implementation Scheme 58.
[0228] Implementation Scheme 64. The composition as described in Implementation Scheme 63, comprising a plurality of cells expressing circular RNA.
[0229] Implementation Scheme 65. The composition as described in Implementation Scheme 63 or 64, wherein the composition is a nanoparticle formulation.
[0230] Example This application is further illustrated by the following embodiments, but is not limited thereto.
[0231] Comparative Example: Type I Intron-Mediated RNA Circulation Method Among the known type I introns tested, those known to originate from... Figure 1 The secondary structure shown Anabaena In vitro self-splicing of type I introns of pre-tRNA (SEQ ID: 58) is most efficient. It has been reported that rearranged type I introns can be used... Figure 2 The PIE (rearrangement-intron-exon) mechanism shown performs RNA circularization. This PIE splicing circularization method forms a 3' exon-5' exon sequence with a fusion of a 3'-intron and a 5'-intron, and the self-splicing of this sequence produces a circular exon.
[0232] 1.1 Experimental Methods 1) Using the PIE strategy as a starting point to generate circRNAs with the target RNA sequence. Figure 3 The 3'-intron, 3'-exon, 5'-exon, and 5'-intron sequences were used as a scaffold to construct a circRNA containing seq 1. The seq 1 sequence is shown in SEQ ID:6 (seq 1 is used here only as a representative example of the target RNA sequence).
[0233] 2) A DNA fragment including the coding sequence of the T7 promoter (SEQ ID: NO. 1), 3'-intron, 3' exon, seq 1, 5' exon, and 5'-intron was cloned into plasmid pUC57 by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized. An RNA consisting of a 3'-intron (SEQ ID: NO. 2), a 3'-intron (SEQ ID: NO. 3), seq 1, a 5'-intron (SEQ ID: NO. 4), and a 5'-intron (SEQ ID: NO. 5) from the 5' end to the 3' end (named preRNA-seq 1) was transcribed in vitro using the linearized plasmid DNA as a template. preRNA-seq 1 was purified by spin column chromatography.
[0234] 3) Circularization of purified preRNA-seq 1: Heat 1 μg RNA to 70°C for 5 min, then immediately place on ice for 3 min. Add GTP to a final concentration of 2 mM, along with a magnesium-containing buffer (50 mM Tris-HCl, 10 mM MgCl2, pH 7.5). Heat the RNA to 52°C for 10 min, and then check the circularization reaction using a 5% pre-denatured urea-TBE PAGE gel (Wshtbio).
[0235] 4) Improved PIE method: Placing homologous arms at the 5' and 3' ends of preRNA-seq 1 and inserting spacer sequences between the 3' exon and seq 1, and between seq 1 and the 5'-exon, can improve the self-splicing efficiency of PIE. Figure 4 An improved PIE strategy was reported in Wesselhoeft et al., Nature Communications, 2018, Vol. 9, Article No. 2629, the contents of which are incorporated herein by reference in their entirety. Therefore, the 5'-homologous arm, 3'-intron, 3'-exon, spacer, spacer, 5'-exon, 5'-intron, and 3'-homologous arm were used as a scaffold to construct seq 1 containing circRNA. The 5'-homologous arm and 3'-homologous arm are perfectly paired. The sequence of seq 1 is shown in SEQ ID: 6.
[0236] 5) A DNA fragment including the T7 promoter, a coding sequence of 5' arms of length 10, 15, or 19 nt (5'-arm-10 / 15 / 19), 3'-introns, 3' exons, spacers, seq 1, spacers, 5' exons, 5'-introns, and 3'-arms of length 10, 15, or 19 nt (3'-arm-10 / 15 / 19) was cloned into plasmid pUC57, which was performed by Suzhou Genweiz Biotechnology.
[0237] The plasmid DNA was linearized. RNA containing 5'-arm-10 (SEQ ID: 8) or arm-15 (SEQ ID: 9) or arm-19 (SEQ ID: 10), 3'-intron, 3' exon, seq 1, 5' exon, 5'-intron, and 3'-arm-10, 15, or 19 (named I-preRNA-seq 1-10, I-preRNA-seq 1-15, or I-preRNA-seq 1-19) from the 5' end to the 3' end was transcribed in vitro using the linearized plasmid DNA as a template. 5'-arm-10 perfectly pairs with 3'-arm-10. 5'-arm-15 perfectly pairs with 3'-arm-15. Similarly, 5'-arm-15 perfectly pairs with 3'-arm-15. I-preRNA-seq 1-10, I-preRNA-seq 1-15, or I-preRNA-seq 1-19 were purified by rotating column.
[0238] 6) A modified PIE strategy was used to circularize seq 2, an expression RNA of a larger size, similar to the circularization process of seq 1 for expressing circRNA (seq 2 is used here only as a representative example of a large target RNA sequence). The RNA contains 5'-arm-19, 3'-intron, 3' exon, spacer, seq 2, spacer, 5' exon, 5'-intron, and 3'-arm-19 from the 5' end to the 3' end, and is named I-preRNA-seq2-19. The seq 2 sequence is shown below: SEQ ID: 7.
[0239] 7) The purified I-preRNA-seq 1-10, I-preRNA-seq 1-15, I-preRNA-seq 1-19, and I-preRNA-seq 2-19 were circularized as follows: Method A: 1 μg of RNA was heated to 55°C for 8 min with 2 mM GTP and a buffer including magnesium (50 mM Tris-HCl, 10 mM MgCl2, 1 mM DTT, pH 7.5). To further improve splicing efficiency, the optimized Method B was followed: RNA was heated to 85°C for 5 min, then immediately placed on ice for 3 min, followed by adding GTP to a final concentration of 2 mM, along with a buffer including magnesium (10 mM), and then heating the RNA to 55°C for 8 min. The cyclization reaction (I-preRNA-seq1-10 / 15 / 19) was then examined using a 5% pre-denatured urea-TBE PAGE gel (Wshtbio), and I-preRNA seq 2-19 was examined using a 2% E-gel EX (Invitrogen).
[0240] result like Figure 3 B and Figure 4 As shown in B, 5% pre-denatured urea-TBE PAGE allows for a simple and efficient separation of circular splicing products from linear precursor molecules, nick RNA, and splicing intermediates. Since circular splicing consists of two steps of transesterification, the intermediate product here refers to the product of the first transesterification step, and the nick RNA product is a discontinuous circRNA molecule where there are no phosphodiester bonds between adjacent nucleotides, signifying a break in the circRNA. This break in circRNA is called nick RNA, and it shares the same nucleic acid sequence as the circRNA.
[0241] For the PIE strategy, the self-splicing efficiency was 27.2%, which was assessed by the disappearance of the preRNA band and quantified using Image J. Figure 3 B). The self-splitting efficiency equation is as follows: Self-splitting efficiency = Meanwhile, the RNA circularization efficiency of the PIE strategy was 19.1%, which was assessed by the appearance of circRNA bands and quantitatively analyzed using ImageJ. Figure 3 B). The cyclization efficiency equation is as follows: Circulation efficiency = For the improved PIE strategy, the results ( Figure 4B) shows that, compared with PIE ( Figure 3 Compared to method B, arm-10 and the spacer contribute slightly higher RNA splicing efficiency, from 27.2% to 29.7%. Simultaneously, RNA circularization efficiency increases from 19.1% to 22.3%. By optimizing splicing conditions (method B), the RNA circularization efficiency and self-splicing efficiency of arm-10 reach 29.4% and 38.8%, respectively. RNA circularization efficiency and self-splicing efficiency are calculated using the equations described above. However, for the longer homologous arms arm-15 or arm-19, the preRNA band is significantly reduced, but the overall circularization efficiency is severely impaired due to the formation of I-preRNA dimers. The dimer formation rate is quantified based on the following equation: For example, the dimer formation rate of arm-15 using method A reached 58.6%, or 54.0% using the optimized method B. For arm-19, the dimer formation rate using method A was 63.5%, or 55.8% using the optimized method B. The results show that the optimized method B can slightly reduce dimer formation. The improved PIE self-splicing efficiency and RNA circularization efficiency were also evaluated using the above equations: For arm-15, the self-splicing efficiency was 94.3% in method A or 97.3% in method B, while the RNA circularization efficiency was 28.8% in method A or 36.2% in method B. For arm-19, the self-splicing efficiency was 98.1% in method A or 97.7% in method B, while the RNA circularization efficiency was 23.5% in method A or 27.9% in method B. This indicates that even long arms contribute to high self-splicing efficiency as measured by the disappearance of the preRNA band, while dimers generated by long arms significantly impair circRNA generation.
[0242] Because a strong homologous arm of 19 nucleotides in length facilitates higher self-splicing efficiency, a circRNA with seq 2 (I-preRNA-seq 2-19) was synthesized using a 19-nt homologous arm. The length of I-preRNA-seq 2-19 is approximately 1.7 kb. Figure 4 As shown in C, the I-preRNA-dimer band is still clearly visible even after RNA circularization treatment (B-method), indicating that the 1.7 kb-RNA circularized by the improved PIE also has the problem of forming I-preRNA-dimers.
[0243] Therefore, although existing improved PIE methods can be used to improve splicing efficiency, they also lead to the formation of undesirable dimers.
[0244] Example 1. Multi-segment assembled active type I introns In this embodiment, the source Anabaena The type I introns of pre-tRNA were divided into three short RNA fragments. Self-splicing activity was measured to verify whether multiple fragments of type I introns could be assembled into active introns that underwent self-splicing.
[0245] 1.1 Experimental Methods We selected an interruption site located in the P6 region of the type I intron and a P2 region located in the type I intron. Figure 5 A), P5 ( Figure 5 B), P8 Figure 5 C) or P9 ( Figure 5 Another interruption point within region D) to receive from Anabaena The type I intron of pre-tRNA is divided into three short RNA fragments.
[0246] For assembled intron-1 ( Figure 5 A) The resulting fragment is referred to as fragment 1 (F1), which includes a 5'-splicing site and extends from the 5' end of the type I intron to the left of the first interruption site located in the P2 region of the type I intron; fragment 2 (F2), which extends from the right of the first interruption site in the P2 region to the left of the second interruption site located in the P6 region of the type I intron; and fragment 3 (F3), which includes a 3'-splicing site and extends from the right of the second interruption site located in the P6 region to the 3' end of the type I intron.
[0247] Similarly, for assembled intron-2 ( Figure 5 B), fragment 1 includes a 5'-splicing site and extends from the 5' end to the left of the first interruption site located in the P5 region; fragment 2 extends from the right of the first interruption site located in the P5 region to the left of the second interruption site located in the P6 region; and fragment 3 includes a 3'-splicing site and extends from the right of the second interruption site located in the P6 region to the 3' end of the type I intron.
[0248] For assembled intron-3 ( Figure 5 C) Segment 1 includes a 5'-splicing site and extends from the 5' end of the type I intron to the left of the first interruption site located in the P6 region; segment 2 extends from the right of the first interruption site located in the P6 region to the left of the second interruption site located in the P8 region; and segment 3 includes a 3'-splicing site and extends from the right of the second interruption site located in the P8 region to the 3' end of the type I intron.
[0249] For assembled intron-4 ( Figure 5D), fragment 1 includes a 5'-splicing site and extends from the 5' end of the type I intron to the left of the first interruption site located in region P6; fragment 2 extends from the right of the first interruption site located in region P6 to the left of the second interruption site located in region P9; and fragment 3 includes a 3'-splicing site and extends from the right of the second interruption site located in region P9 to the 3' end of the type I intron.
[0250] Fragments 1 (F1) and 3 (F3) can pair with fragment 2 (F2) to facilitate the assembly of active type I introns. Analysis of the assembled introns, such as assembled intron-1 (F2), is then performed. Figure 5 A) Assembled intron-2 ( Figure 5 B) Assembled intron-3 ( Figure 5 C) and assembled intron-4 to verify self-splicing activity. Multiple fragments of the intron (e.g., F1, F2, and F3 of the assembled intron-1 (AI-1)) were synthesized by Genscript solid-phase synthesis. The sequences AI-1-F1, AI-1-F2, and AI-1-F3 are shown in SEQ IDs: 11, 12, and 13; the sequences AI-2-F1, AI-2-F2, and AI-2-F3 of the assembled intron-2 (AI-2) were synthesized by Genscript solid-phase synthesis, and the sequences AI-2-F1, AI-2-F2, and AI-2-F3 are shown in SEQ IDs: 14, 15, and 16. The sequences AI-3-F1, AI-3-F2, and AI-3-F3 of the assembled intron-3 (AI-3) were synthesized by Genscript solid-phase synthesis, and the sequences AI-3-F1, AI-3-F2, and AI-3-F3 are shown in SEQ IDs: 17, 18, and 19. Furthermore, the F1, F2, and F3 of the assembled intron-4 (AI-4) were synthesized by Genscript solid-phase synthesis, and the AI-4-F1, AI-4-F2, and AI-4-F3 sequences are shown below: SEQ ID: 20, 21, 22.
[0251] Multi-segment introns were assembled using F1, F2, and F3, followed by self-splicing. Specifically, 50 pmol AI-1-F1 / AI-2-F1 / AI-3-F1 / AI-4-F1, 50 pmol AI-1-F2 / AI-2-F2 / AI-3-F2 / AI-4-F2, and 50 pmol AI-1-F3 / AI-2-F3 / AI-3-F3 / AI-4-F3 were incubated at 37 °C in buffer (50 mM pH 7.5 Tris-HCl) for 1 h. After incubation, GTP and MgCl2 were added to final concentrations of 2 mM and 34 mM, respectively, followed by further incubation at 37 °C for 30 min. The self-splicing reaction was analyzed using 20% pre-denatured urea-TBE PAGE gel (Wshtbio).
[0252] result The results showed that exon bands (17 nt) (the self-splicing products of type I introns) appeared after AI-1-F1, AI-1-F2, and AI-1-F3 assembled into multi-fragment active type I introns, and after AI-4-F1, AI-4-F2, and AI-4-F3 assembled into multi-fragment active type I introns, which verified the self-splicing activity of multi-fragment type I introns. The positions of the exons (the splicing products of multi-fragment type I introns) in the gel are indicated by red arrows. Figure 6 ).
[0253] In summary, a type I intron can be divided into three parts (fragment 1 (F1), fragment 2 (F2), and fragment 3 (F3)), which can then recombine to form a multi-fragment type I intron, which can self-splice to synthesize exons, such as... Figure 7 As shown.
[0254] Example 2. Trans-splicing strategy for synthesizing cyclic exons This is an exemplary illustration of generating circular RNA using a method involving multiple fragment ribozymes. Specifically, RNA containing fragment 1 (F1), exon 5, fragment 3 (F3), and exon 3' is rearranged to obtain a first RNA molecule (the rearranged RNA molecule) containing F3 - 3' exon - 5' exon - F1. The first RNA molecule self-splices with the aid of fragment 2 (F2) to generate circular exons. This method of generating circular exons is called the trans-splicing strategy. Figure 9 ).
[0255] 2.1 Experimental Methods: To generate circular exons, a DNA fragment containing a T7 promoter operatively linked to the coding sequence of F3-3' exon-5' exon-F1 was synthesized by PCR, with primers obtained from Suzhou Genweiz Biotechnology. The DNA fragment was transduced in vitro to produce a rearranged RNA molecule containing F3-3' exon-5' exon-F1. A second RNA molecule containing fragment 2 (F2) was chemically synthesized by Genscript. The rearranged RNA molecule and the second RNA molecule were derived from each of the assembled introns-1 to-4. The activity of generating circular exons from each rearranged RNA molecule with the aid of the corresponding F2 sequence was measured. For the assembled intron-1 (AI-1, Figure 5A), the sequence of fragment 3 (F3) is shown in SEQ ID NO: 13, the sequence of F2 is shown in SEQ ID NO: 12, and the sequence of fragment 1 (F1) is shown in SEQ ID NO: 11. The sequence of the rearranged RNA molecule of assembled intron-1 (P-AI-1) is shown in SEQ ID NO: 31. For assembled intron-2 (AI-2, Figure 5 B), the F3 sequence is shown in SEQ ID NO: 16, the F2 sequence is shown in SEQ ID NO: 15, and the F1 sequence is shown in SEQ ID NO: 14. The sequence of the rearranged RNA molecule of the assembled intron-2 (P-AI-2) is shown in SEQ ID NO: 32. For the assembled intron-3 (AI-3, Figure 5 C), the F3 sequence is shown in SEQ ID NO: 19, the F2 sequence is shown in SEQ ID NO: 18, and the F1 sequence is shown in SEQ ID NO: 17. The sequence of the rearranged RNA molecule of the assembled intron-3 (P-AI-3) is shown in SEQ ID NO: 33. For the assembled intron-4 (AI-4, Figure 5 D), the F3 sequence is shown in SEQ ID NO: 22, the F2 sequence in SEQ ID NO: 21, and the F1 sequence in SEQ ID NO: 20. The sequence of the rearranged RNA molecule with intron-4 (P-AI-3) assembled is shown in SEQ ID NO: 34. The effects of coupling different cleavage sites on P6 with the cleavage site at P5 on the generation of circular exons were also analyzed. Details of the cleavage are in Figure 8 As explained in AD. The segments are named as assembled introns 2-1 (AI-2-1, Figure 8 A) Assembled intron 2-2 (AI-2-2, Figure 8 B) Assembled intron 2-3 (AI-2-3, Figure 8 C) and assembled intron 2-4 (AI-2-4, Figure 8D). As described above, circular exons are generated using rearranged RNA molecules. For the assembled intron 2-1 (AI-2-1), the F2 sequence is shown in SEQ ID NO: 23, and the F3 sequence is shown in SEQ ID NO: 24. The sequence of the rearranged RNA molecule of AI-2-1 (P-AI-2-1) is shown in SEQ ID NO: 35. For the assembled intron 2-2 (AI-2-2), the F2 sequence is shown in SEQ ID NO: 25, and the F3 sequence is shown in SEQ ID NO: 26. The sequence of the rearranged RNA molecule of the assembled AI-2-2 (P-AI-2-2) is shown in SEQ ID NO: 36. For the assembled intron 2-3 (AI-2-3), the F2 sequence is shown in SEQ ID NO: 27, and the F3 sequence is shown in SEQ ID NO: 28. The sequence of the rearranged RNA molecule of AI-2-3 (P-AI-2-3) is shown in SEQ ID NO: 37. For the assembled intron 2-4 (AI-2-4), the F2 sequence is shown in SEQ ID NO: 29, and the F3 sequence is shown in SEQ ID NO: 30. The sequence of the rearranged RNA molecule of AI-2-4 (P-AI-2-4) is shown in SEQ ID NO: 38. The assembled introns 2-1 / 2 / 3 / 4 share F1 as shown in SEQ ID NO 14.
[0256] DNA fragments were synthesized by polymerase chain reaction (PCR) using primers obtained from Suzhou Genweiz Biotechnology. RNA containing F3, 3' exon, 5' exon, and F1 from the 5' end to the 3' end (named rearranged intron-1 (P-AI-1), rearranged intron-2 (P-AI-2), rearranged intron-3 (P-AI-3), rearranged intron-4 (P-AI-4), rearranged intron-2-1 (P-AI-2-1), rearranged intron-2-2 (P-AI-2-2), rearranged intron-2-3 (P-AI-2-3), and rearranged intron-2-4 (P-AI-2-4)) was transcribed in vitro using the DNA fragments obtained by PCR as templates. P-AI-1 / P-AI-2 / P-AI-3 / P-AI-4 / P-AI-2-1 / P-AI-2-2 / P-AI-2-3 / P-AI-2-4 RNA was purified by spin column. Circular exons (products of the trans-splicing strategy) are shown in SEQ ID NO: 39.
[0257] Column-purified P-AI-1 / P-AI-2 / P-AI-3 / P-AI-4 / P-AI-2-1 / P-AI-2-2 / P-AI-2-3 / P-AI-2-4 were cyclized with the aid of their respective F2. Specifically, P-AI-1 / P-AI-2 / P-AI-3 / P-AI-4 / P-AI-2-1 / P-AI-2-2 / P-AI-2-3 / P-AI-2- and their respective F2 were annealed in buffer (50 mM Tris-HCl, pH 7.5) at a 1:1 molar ratio. GTP and MgCl2 were added to the annealing system (final concentration of GTP was 2 mM, and final concentration of MgCl2 was 34 mM), and the system was then heated to 42 °C for 10 min. Trans-cyclization reaction system 1 contained the reagents shown in Table 1.
[0258] Table 1. Trans-cyclization reaction system 1
[0259] The cyclization products were examined using 5% pre-denatured urea-TBE PAGE gel (purchased from Wshtbio) at 180 V for 45 min.
[0260] result To investigate the possibility of generating circular exons (circ-exons) via a reverse splicing strategy, we used F2 of P-AI-1 and its corresponding derived self-assembled intron-1 to generate circular exons. (See data...) Figure 10 As shown in A), the cyclic exon bands are clearly present in the PAGE gel (marked with red arrows). Similarly, the cyclic exons are generated from P-AI-2 and its corresponding F2 (derived self-assembled intron-2), and P-AI-4 and its corresponding F2 (derived self-assembled intron-4), as... Figure 10 B and Figure 10 As shown in C. Furthermore, we need to mention that the purity of in vitro transcribed RNA, P-AI-1, P-AI-2, and P-AI-4 is relatively poor.
[0261] Furthermore, we tested the effect of coupling different cleavage sites on P6 with the cleavage site at P5 on the generation of circular exons, specifically the efficiency of P-AI-2-1, P-AI-2-2, P-AI-2-3, and P-AI-2-4 in generating circular exons with the aid of their respective F2 fractions. We found that P-AI-2-3 with the aid of its respective F2 fraction and P-AI-2-4 with the aid of their respective F2 fractions can generate circular exons. Figure 11 The circular exons are indicated by the red arrows in the middle.
[0262] In short, the data above indicates that the inverse splicing strategy can be used to generate cyclic exons.
[0263] Example 3. Trans-splicing strategy for synthesizing circRNA with target RNA sequence.
[0264] In this embodiment, using Figure 12 The trans-splicing strategy shown in A produces circRNA with the target RNA sequence.
[0265] 3.1 Experimental Methods To validate the trans-splicing strategy used to generate circRNA with the target sequence, a fragment (F3) including a 3' splice site, a 3' exon, and a 5' exon, and fragment 1 (F1) including a 5' splice site were used as scaffolds to assemble rearranged RNA molecules for the synthesis of circRNA with the target RNA sequence (seq 1 is used herein as a representative example of the target RNA sequence). The seq 1 sequence is shown in SEQ ID: 6, the same as in the comparative example. DNA was synthesized by Suzhou Genweiz Biotechnology. Fragment 2 (F2) was synthesized by Genscript. The F1 sequence is identical to that of Example 2 shown in SEQ ID NO: 14. The F2 sequence is identical to that of Example 2 shown in SEQ ID NO: 15, and the F3 sequence is identical to that of Example 2 shown in SEQ ID NO: 16.
[0266] DNA fragments including the coding sequences of the T7 promoter and F3, the 3' exon, seq 1, the 5' exon, and F1 were cloned into plasmid pUC57 by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized. RNA containing F3, the 3' exon, seq 1, the 5' exon, and F1 from the 5' end to the 3' end (named rearranged -RNA-seq 1 (P-RNA-seq 1)) was transcribed in vitro using the linearized plasmid DNA as a template. The P-RNA-seq 1 RNA was purified by spin column chromatography. The P-RNA-seq 1 sequence is shown in SEQ ID NO: 40.
[0267] The purified p-RNA-seq 1 was cyclized with the aid of F2. Specifically, p-RNA-seq 1 and F2 were annealed in a 1:1 molar ratio in buffer (50 mM Tris-HCl, pH 7.5). GTP and MgCl2 were added to the annealing system (final concentration of GTP was 2 mM, and final concentration of MgCl2 was 34 mM), and the system was then heated to 52 °C for 10 min. Trans-cyclization reaction system 2 contained the reagents shown in Table 2.
[0268] Table 2. Trans-cyclization reaction system 2
[0269] The cyclization products were examined using 5% pre-denatured urea-TBE PAGE gel (purchased from Wshtbio) at 180 V for 45 min.
[0270] 3.2 Results The results showed that after circularization with the help of F2, the circRNA bands appeared clearly. Figure 12 B) Demonstrates that the trans-splicing strategy can be used to generate circRNA. The splicing efficiency of the trans-splicing strategy is approximately 32.3%, which was assessed by the disappearance of the p-RNA-seq 1 band and quantified using ImageJ. The splicing efficiency was calculated according to the equations mentioned in the comparative examples above.
[0271] RNA circularization efficiency was 16.8%, measured by the appearance of circRNA bands and quantified using ImageJ. Figure 12 B). Calculate the cyclization efficiency equation based on the equations mentioned in the comparative embodiments above.
[0272] In summary, the trans-splicing strategy according to the embodiments of this application has splicing activity and can be used to synthesize circRNA with a target RNA sequence.
[0273] Example 4. An improved trans-splicing strategy for synthesizing circRNA using a target RNA sequence. A fully complementary binding motif was placed at the 3' end of p-RNA-seq 1 (named Improved Rearranged RNAseq 1 or IP-RNA-seq 1) and the 5' end of fragment 2 (named BM-F2). Assembly of IP-RNA-seq 1 and BM-F2 was analyzed. Results showed that, with the aid of BM-F2, the well-folded active complex IP-RNA-seq 1 assembled with BF-M2 exhibited higher self-splicing efficiency and higher efficiency in generating circRNA. Figure 13 A).
[0274] 4.1 Experimental Methods To test the possibility of an improved trans-splicing strategy, fragment (F3), exon 3, exon 5, fragment 1 (F1), and binding motif sequences of lengths 5, 10, 15, 20, and 50 nt were used as scaffolds to construct seq 1 expressing circRNA. The seq 1 sequence is shown in SEQ ID: 6. DNA was synthesized by Suzhou Genweiz Biotechnology. Fragment 2 (F2) with binding motifs of lengths 5, 10, 15, 20, and 50 nt at the 5' end (BM(5 / 10 / 15 / 20 / 50)-F2) was synthesized by Genscript chemistry. The sequences BM(5)-F2, BM(10)-F2, BM(15)-F2, BM(20)-F2, and BM(50)-F2 are shown in SEQ ID NO: 41, 42, 43, 44, and 45, respectively.
[0275] DNA fragments including the T7 promoter, F3, 3' exon, seq 1, 5' exon, F1, and binding motifs of lengths of 5, 10, 15, 20, and 50 nt (BM-5 / 10 / 15 / 20 / 50) were cloned into plasmid pUC57 by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized. RNA containing F3, 3' exon, seq 1, 5' exon, F1, and BM-5 / 10 / 15 / 20 / 50 from the 5' end to the 3' end (named I(5 / 10 / 15 / 20 / 50)-P-RNA-seq 1) was transcribed in vitro using the linearized plasmid DNA as a template. I(5 / 10 / 15 / 20 / 50)-P-RNA-seq 1 was purified by spin column chromatography. The sequences I(5)-P-RNA-seq 1, I(10)-P-RNA-seq 1, I(15)-P-RNA-seq 1, I(20)-P-RNA-seq 1 and I(50)-P-RNA-seq 1 are shown in SEQ ID NO: 46, 47, 48, 49 and 50, respectively.
[0276] The purified I(5 / 10 / 15 / 20 / 50)-p-RNA-seq 1 was cyclized with the aid of BM(5 / 10 / 15 / 20 / 50)-F2. Specifically, I(5 / 10 / 15 / 20 / 50)-p-RNA-seq 1 and BM(5 / 10 / 15 / 20 / 50)-F2 were annealed in a 1:1 molar ratio in buffer (50 mM Tris-HCl, pH 7.5). GTP and MgCl2 were added to the annealing system (final concentration of GTP was 2 mM, and final concentration of MgCl2 was 34 mM), and the system was then heated to 52 °C for 10 min. Transcyclization reaction system 3 contained the reagents shown in Table 3.
[0277] Table 3. Trans-cyclization reaction system 3
[0278] Meanwhile, we optimized the trans-cyclization reaction to improve the cyclization efficiency of the improved trans-splicing strategy. The cyclization products were then examined using a 5% pre-denatured urea-TBE PAGE gel (purchased from Wshtbio) at 180 V for 45 min.
[0279] To verify that the cyclization product was circRNA, the cyclization reaction was treated with RNase R. A typical 20 μL reaction mixture contained: 1 μg RNA, 20 mM Tris–HCl (pH 8.0), 10 mM NaCl, 0.1 mM MgCl2, and 1 U / 2 U / 3 U RNase R (Beyotime). The reaction mixture was incubated at 37 °C for 30 min, followed by heating to 70 °C for 10 min to inactivate RNase R. The RNA digested by RNase R was then analyzed by 5% pre-denatured urea-TBE PAGE gel.
[0280] According to the manufacturer's instructions, circularized IP-RNA-seq 1 was reverse transcribed (RT) into cDNA using the Hiscript First-Strand cDNA Synthesis Kit (vazyme, R111-02) with random primers. One-tenth of the RT product was used for PCR amplification in a real-time PCR system with divergent primers that amplify transcripts that span the splice junction of circRNAs. PCR products were analyzed by 1% agarose gel electrophoresis. The PCR products were also cloned into plasmids for sequencing using TA cloning (Shanghai Sangon).
[0281] 4.2 Results First, the influence of the binding motif length on the ringing efficiency of the inverse splicing strategy was investigated. For example... Figure 13As shown in B, all tested binding motifs (5 nt, 10 nt, 15 nt, and 20 nt) promoted circRNA generation, especially the 10 nt, 15 nt, and 20 nt lengths. Although the 50 nt binding motif also promoted trans-splicing strategy-mediated circRNA generation, partially assembled BM(50)-F2 and I(50)-P-RNA-seq 1 (in Figure 13 In B, the proportion of circRNA was reduced by the yellow star marker.
[0282] Furthermore, after optimizing the cyclization reaction conditions, the RNA cyclization efficiency of the improved trans-splicing strategy was examined. Results ( Figure 14 B) shows that adding binding motifs to F2 (BM(10)-F2) and p-RNA-seq 1 (IP-RNA-seq 1) reduced the splicing efficiency of the trans-splicing strategy from 32.3% ( Figure 12 B) increased to 77.9% Figure 14 Lane B, third lane Figure 14 The first lane of lane B is the RNA ladder, which was assessed by the disappearance of IP-RNA-seq 1 and quantified by Image J. Furthermore, this improved method increased the RNA circularization efficiency of the trans-splicing strategy from 16.8% ( Figure 12 B) increased to 57.2% Figure 14 B). Most importantly, no dimers were detected using this improved method. Splicing efficiency and RNA circularization efficiency were each calculated according to the corresponding equations cited in the comparative examples. Therefore, the improved trans-splicing strategy according to the embodiments of this application not only improves RNA circularization efficiency but also results in no detectable dimer formation, which is surprising compared to existing improved PIEs. Figure 4 B).
[0283] To ensure that the major circularization product of the improved trans-splicing strategy was indeed circRNA, the circularization reaction was treated with different units of RNase R (1 U, 2 U, or 3 U). Even after RNase R treatment, RNase R-resistant circRNA was confirmed by observing clear circRNA bands, while the linear product was partially or completely degraded by RNase R treatment. Figure 14 Lanes B: 4, 5, and 6.
[0284] As is well known, divergent primers designed to span the circRNA splice junction sequence can specifically amplify circRNA, but not the corresponding linear RNA. The 231-bp PCR product amplified by the divergent primers was analyzed using 1% agarose gel electrophoresis. Figure 14The results in C showed that the target band (231 bp) of the PCR product exhibited the expected size. Furthermore, sequencing results of the PCR product confirmed the correct splicing junction of the circRNA. Figure 14 D).
[0285] In summary, this embodiment demonstrates the successful generation of circRNA via the improved trans-splicing strategy according to the embodiments of this application, which has higher RNA circularization efficiency and no detectable dimer byproducts.
[0286] Example 5. circRNA synthesized using an improved trans-splicing method can express the secretory protein Gaussian luciferase in 293T cells. 5.1 Experimental Methods EMCV IRES, F3, 3' exon, 5' exon, F1, and the binding motif sequence were used as basic elements for constructing a circRNA expressing the secretory protein Gaussian luciferase (Gluc). The EMCV IRES sequence is shown in SEQ ID NO: 51. The Gluc sequence is shown in SEQ ID NO: 52. DNA was synthesized by Suzhou Genweiz Biotechnology. BM-F2 was chemically synthesized by Genscript. The BM-F2 sequence is the same as that used in Example 4.
[0287] A DNA fragment including the T7 promoter, F3, 3' exon, EMCV IRES, Gluc coding region, 5' exon, F1, and binding motif was cloned into plasmid pUC57, which was performed by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized, and RNA including F3, 3' exon, EMCV IRES, Gluc, 5' exon, F1, and binding motif (named IP-RNA-Glu) was transcribed in vitro using the linearized plasmid DNA as a template. IP-RNA-Glu (SEQ ID NO: 53) was purified, circularized with the aid of BM-F2, and then purified again using the procedure described in Example 4.
[0288] To enrich circRNA, 20 μg of RNA was diluted in water (86 μL), then heated at 65 °C for 3 min and cooled on ice for 3 min. 30 U RNase R and 10 μL of 10× RNase R buffer (Beyotime) were added, and the reaction was incubated at 37 °C for 15 min, then heated to 70 min and held for 10 min. The RNase R-digested RNA was then purified by column chromatography.
[0289] To evaluate the translation efficiency of enriched circRNA-Gluc in 293T cells, 293T cells were first cultured at a concentration of 1x10⁻⁶ cells. 4 Cells were seeded per well in 96-well plates and cultured at 37°C and 5% CO2. After cell adhesion, each well was transfected with 200 ng of RNase R-digested RNA using Lipofectamine MessengerMax (Invitrogen) transfection reagent, or with 200 ng of RNase R-digested RNA expressing eGFP-NLS (circRNA-eGFP-NLS) as a control.
[0290] Transfected cells were analyzed using a Gaussian luciferase assay. Specifically, transfected cells were incubated at 37°C with 5% CO2 for 24 hours, followed by the removal of 10 μL of culture medium from the transfected cells. The photometer was programmed. 10 μL of the collected culture medium was added to each well of a black 96-well plate, and then 50 μL of working solution (containing coelentrin) was added to each well. The light output was detected 10 minutes after signal stabilization.
[0291] 5.2 Results Gluc activity was determined using the Pierce Gluc luminescence assay kit from 293 T culture medium transfected with circRNA-Gluc. Gluc activity is shown in... Figure 15 In contrast to the control (circRNA-eGFP-NLS transfected in 293 T cells), circRNA-Gluc transfected in 293 T cells resulted in robust production of the secreted protein Gluc.
[0292] Example 6. circRNA synthesized using an improved trans-splicing method can express the intracellular protein firefly luciferase in 293T cells. 6.1 Experimental Methods EMCV IRES, F3, 3' exon, 5' exon, F1, and binding motif sequences were used as basic elements for constructing circRNA expressing the intracellular protein firefly luciferase (Fluc). The Fluc sequence is shown in SEQ ID NO: 54. DNA was synthesized by Suzhou Genweiz Biotechnology. BM-F2 was chemically synthesized by Genscript. BM-F2 is the same as that used in Example 4.
[0293] A DNA fragment including the T7 promoter, F3, 3' exon, EMCV IRES, Fluc coding region, 5' exon, F1, and binding motif was cloned into plasmid pUC57, which was performed by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized, and RNA including F3, 3' exon, EMCV IRES, Fluc, 5' exon, F1, and binding motif (named IP-RNA-Fluc) was transcribed in vitro using the linearized plasmid DNA as a template. IP-RNA-Fluc (SEQ ID: 55) was purified by cycling with the aid of BM-F2, and then purified again using the procedure described in Example 4.
[0294] circRNA-Fluc was enriched using the procedure described in Example 5.1 and transfected into 293T cells. Control cells were transfected with RNA expressing eGFP-NLS (circRNA-eGFP-NLS) digested with 200 ng RNase R.
[0295] Transfected cells were analyzed using a firefly luciferase assay. Specifically, transfected cells were incubated at 37°C and 5% CO2 for 24 hours. Culture medium was aspirated, and cells were washed with 100 μL / well of 1X DPBS buffer (Thermofish). After aspirating the DPBS, 100 μL / well of 1X cell lysis buffer was added to the cells. The plate was shaken at medium speed for 15 minutes on a plateau shaker. Cell lysis buffer was transferred to black 96-well plates, and an equal volume of working solution (containing the substrate D-luciferin) was added to each well. Light output was detected 10 minutes after signal stabilization.
[0296] 6.2 Results Fluc activity in cell lysis from 293 T cells transfected with circRNA-Fluc was determined using the Pierce Fluc luminescence assay kit. Fluc activity was shown in... Figure 16 In contrast to the control (circRNA-eGFP-NLS transfected in 293 T cells), circRNA-Fluc transfected in 293 T cells resulted in robust production of the intracellular protein Fluc.
[0297] Example 7. circRNA synthesized using an improved trans-splicing method can express nuclear proteins in 293T cells. 7.1 Experimental Methods EMCV IRES, F3, 3' exon, 5' exon, F1, and binding motif sequences were used as basic elements for constructing circRNAs expressing nuclear proteins. A fluorescent protein, eGFP, fused to the nuclear localization signal (NLS), was used to mimic nuclear proteins. The eGFP-NLS sequence is shown in SEQ ID NO: 56. DNA was synthesized by Suzhou Genweiz Biotechnology. BM-F2 was identical to that in Example 4.
[0298] A DNA fragment including the T7 promoter, F3, 3' exon, EMCV IRES, eGFP-NLS coding sequence, 5' exon, F1, and binding motif was cloned into plasmid pUC57 by Suzhou Genweiz Biotechnology. The plasmid DNA was linearized, and RNA including F3, 3' exon, EMCV IRES, eGFP-NLS, 5' exon, F1, and binding motif (named IP-RNA-eGFP-NLS) was transcribed in vitro using the linearized plasmid DNA as a template. IP-RNA-eGFP-NLS (SEQ ID: 57) was purified, circularized, and then purified again, as described in Example 4.
[0299] circRNA-eGFP-NLS was enriched using the procedure described in Example 5.1.
[0300] 293T cells were transfected with circRNA-eGFP-NLS using the procedure described in Example 5.1, except that 20,000 cells / well were seeded in 12-well plates.
[0301] Transfected cells were analyzed using eGFP-NLS assays. Specifically, Transfected cells were incubated at 37°C and 5% CO2 for 24 hours. Next, 10 μL of 100X Hoechst 33342 (Beyotime) viable cell dye was added to each transfected cell containing 1 mL of culture medium, followed by incubation at 37°C and 5% CO2 for 10 min. The culture medium was then aspirated from the cells. The cells were then washed twice with DPBS buffer. Subsequently, the cells were observed under a fluorescence microscope at 20x magnification.
[0302] 7.2 Results 293 T cells transfected with circRNA-eGFP-NLS were observed using fluorescence microscopy. To verify the presence of eGFP-NLS expressed by circRNA in the cell nucleus, Hoechst (a nuclear dye) was used to indicate nuclear localization. Figure 17As shown, eGFP-NLS colocalizes with Hoechst-stained cell nuclei, indicating that eGFP-NLS expressed by circRNA can be precisely relocalized to the cell nucleus.
[0303] Sequence information
Claims
1. An RNA molecule pair for preparing circular RNA, said RNA molecule pair comprising: (1) A first RNA molecule comprising the following elements, said elements being operatively linked and arranged in order from the 5' end to the 3' end of said molecule: i) The 3' intron sequence of type I introns ii) The 3' splice site of the type I intron. iii) Optionally, the 3' exon sequence of the type I intron, iv) Target RNA sequence, v) Optionally, the 5' exon sequence of the type I intron, vi) The 5′ splice site of the type I intron, and vii) The 5′ intron sequence of the type I intron, and (2) A second RNA molecule containing the intermediate intron sequence of the type I intron. When the first RNA molecule comes into contact with the second RNA molecule, the scaffold and catalytic domains of the type I intron are formed, thereby allowing the generation of the circular RNA containing the 3' exon sequence, the target RNA sequence, and the 5' exon sequence.
2. The RNA molecule pair of claim 1, wherein the 3' intron sequence comprises the R and S sequences of the type I intron, the 5' intron sequence comprises the internal guide sequence (IGS) and P sequence of the type I intron, and the intermediate intron sequence comprises the Q sequence of the type I intron.
3. The RNA molecule pair as claimed in claim 1 or 2, wherein the first RNA molecule comprises from the 5' end to the 3' end of the molecule. a) Fragment 3, which includes the 3' intron sequence and the 3' splice site, which are operatively connected and arranged in order from the 5' end to the 3' end of fragment 3; b) Optionally, the 3' exon sequence; c) The target RNA sequence; d) Optionally, the 5' exon sequence; e) Fragment 1, comprising the 5' splice site and the 5' intron sequence operably connected and arranged in order from the 5' end to the 3' end of fragment 1; and The second RNA molecule contains segment 2 having the intermediate intron sequence.
4. The RNA molecule pair of claim 3, wherein fragments 1-3 each comprise nucleotide sequences or variants thereof of the three parts of the type I intron obtained by dividing the type I intron with two interrupt sites.
5. The RNA molecule pair of claim 4, wherein the three parts of the type I intron are obtained by dividing the intron by an interruption site located within the loop of the P6 stem-loop region of the type I intron and another interruption site located within the loop of the P2, P5, P8 or P9 stem-loop regions of the type I intron.
6. The RNA molecule pair of claim 5, wherein the three parts of the type I intron are obtained by dividing the intron by a first interruption site located within the loop of the P2 stem-loop region and a second interruption site located within the loop of the P6 stem-loop region.
7. The RNA molecule pair of claim 5, wherein the three parts of the type I intron are obtained by dividing the intron by a first interruption site located within the loop of the P5 stem-loop region and a second interruption site located within the loop of the P6 stem-loop region.
8. The RNA molecule pair of claim 5, wherein the three parts of the type I intron are obtained by dividing the intron by a first interruption site located within the loop of the P6 stem-loop region and a second interruption site located within the loop of the P8 stem-loop region.
9. The RNA molecule pair of claim 5, wherein the three parts of the type I intron are obtained by dividing the intron by a first interruption site located within the loop of the P6 stem-loop region and a second interruption site located within the loop of the P9 stem-loop region.
10. The RNA molecule pair according to any one of claims 6-9, wherein: (b) Fragment 1 contains a nucleotide sequence that has at least 80% identity with the nucleotide sequence from the 5' end of the type I intron to the left of the first interruption site, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. (c) Fragment 2 contains a nucleotide sequence having at least 80% identity with the nucleotide sequence from the right side of the first interrupt site to the left side of the second interrupt site, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity; and (d) Fragment 3 contains a nucleotide sequence that has at least 80% identity with the nucleotide sequence from the right side of the second interruption site to the 3' end of the type I intron, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
11. The RNA molecule pair according to any one of claims 3-10, wherein fragment 1 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequence: SEQ ID NO: 11, 14, 17, or 20.
12. The RNA molecule pair according to any one of claims 3-11, wherein fragment 2 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequences: SEQ ID NO: 12, 15, 18, 21, 23, 25, 27, or 29.
13. The RNA molecule pair according to any one of claims 3-12, wherein fragment 3 comprises a nucleotide sequence having at least 80% identity with, for example, at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the following nucleotide sequences: SEQ ID NO: 13, 16, 19, 22, 24, 26, 28, or 30.
14. The RNA molecule pair of claim 13, wherein fragment 1, fragment 2, and fragment 3 comprise the following nucleotide sequence: a. These are SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO: 13, respectively; b. These are SEQ ID NO: 14, SEQ ID NO: 15, and SEQ ID NO: 16, respectively; c. These are SEQ ID NO: 17, SEQ ID NO: 18, and SEQ ID NO: 19, respectively; d. These are SEQ ID NO: 20, SEQ ID NO: 21, and SEQ ID NO: 22, respectively; e. These are SEQ ID NO: 14, SEQ ID NO: 23, and SEQ ID NO: 24, respectively; f. These are SEQ ID NO: 14, SEQ ID NO: 25, and SEQ ID NO: 26, respectively; g. SEQ ID NO: 14, SEQ ID NO: 27, and SEQ ID NO: 28 respectively; or h. are SEQ ID NO: 14, SEQ ID NO: 29 and SEQ ID NO: 30 respectively.
15. The RNA molecule pair as claimed in any of the preceding claims, wherein the first RNA molecule comprises the 3' exon sequence, for example, a 3' exon sequence derived from the natural 3' exon sequence of the type I intron.
16. The RNA molecule pair as claimed in any of the preceding claims, wherein the first RNA molecule comprises the 5' exon sequence, for example, a 5' exon sequence derived from the natural 5' exon sequence of the type I intron.
17. The RNA molecule pair of claim 15 or 16, wherein the 3' exon sequence is 100% identical to the natural 3' exon sequence of the type I intron, or is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 99%, or at least 99% identical to the full-length sequence of the natural 3' exon sequence, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions compared to the natural 3' exon sequence.
18. The RNA molecule pair according to any one of claims 15-17, wherein the 3' exon sequence comprises or consists of the following: SEQ ID NO: 60 (AAAAUCCGU), for example SEQ ID NO:
3.
19. The RNA molecule pair according to any one of claims 16-18, wherein the 5' exon sequence has 100% identity with the natural 5' exon sequence of the type I intron, or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a continuous fragment starting from the 3' terminal nucleotide of the natural 5' exon sequence, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions, or additions compared to the continuous fragment.
20. The RNA molecule pair according to any one of claims 16-19, wherein the 5' exon sequence comprises or consists of the following: SEQ ID NO: 59 (ACGGACUU), for example SEQ ID NO:
4.
21. The RNA molecule pair according to any one of claims 16-19, wherein the 3' exon sequence comprises: SEQ ID NO: 60 or SEQ ID NO: 3; and the 5' sequence consists of SEQ ID NO: 59 or SEQ ID NO:
4.
22. The RNA molecule pair as claimed in any of the preceding claims, wherein the target nucleotide sequence comprises at least one protein-coding sequence and an internal ribosome entry site (IRES) operatively linked thereto.
23. The RNA molecule pair of claim 22, wherein the IRES is selected from any one of the following: Taura syndrome virus, schistosomiasis virus, Theyl encephalomyelitis virus, simian virus 40, red imported fire ant virus 1, cereal constriction virus, reticulovirus endothelial proliferation virus, Forman poliovirus 1, soybean inchworm virus, Kashmir bee virus, human rhinovirus 2, glass leafhopper virus-1, human immunodeficiency virus type 1, glass leafhopper virus-1, lice P virus, hepatitis C virus, hepatitis A virus, hepatitis B virus, foot-and-mouth disease virus, human enterovirus 71. Equine rhinovirus, tea geometrid moth-like virus, encephalomyocarditis virus (EMCV), fruit fly C virus, cruciferous tobacco virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, hibiscus yellow ringspot virus, classical swine fever virus, Aichi virus, cristatoviruses, diechoviruses, foot-and-mouth disease viruses, enteroviruses, Borna virus, Diresapir virus, Ali virus, Arirure virus, Anpi virus, Anati virus, foot-and-mouth disease virus, Avihpata virus, Aves virus, Bosepiritis virus. Karidio virus, Cossa virus, Krahli virus, Kroshi virus, Dicipi virus, Diresapi virus, Enterovirus genus, Eribovirus, Filipi virus, Filipi virus, Gali virus, Grossopi virus, Haka virus, Hermipi virus, Hepatovirus genus, Hammuvirus, Kunsage virus, Linnipi virus, Liupi virus, Ludopi virus, Malagasy virus, Megari virus, Mishi virus, Mosa virus, Mupi virus, Ori virus, Parabovirus, Parabovirus, Parsi virus, Passeri virus, Pemapi virus, Potamipi virus, Labo virus, Rafi virus, Rochley virus, Rosa virus, Sakau virus, Sali virus, Sapellovirus genus, Senecavirus genus, Sumbavivirus, Sisin virus, Simapi virus, Simapi virus, Jeshen virus genus, Tolchi virus, Totori virus, Tremo virus, Tropi virus, Human FGF2, Human SFTPA1, Human AML1 / RUNX1, Drosophila antennae, Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAPl, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIFla, Human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila harvester protein, canine Scanper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila hairless protein, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, simian dermatologic virus, turnip crescendo virus, eIF4G aptamer, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVA1 / 2), or their modified IRES sequences, preferably, said IRES is an IRES of CVB3, BRAV-1L, PV1L, CAV2L, BRAV-1, PV1, or CAV2.
24. The RNA molecule pair of claim 23, wherein the IRES comprises or is composed of the following nucleotide sequences: SEQ ID NO:
51.
25. The RNA molecule pair as claimed in any of the preceding claims, wherein the target nucleotide sequence encodes a protein.
26. The RNA molecule pair of claim 25, wherein the protein is a secretory protein, intracellular protein, or nucleoprotein in a eukaryotic cell.
27. The RNA molecule pair of claim 25, wherein the protein is a human protein, an antigen, an antibody, or a gene-editing enzyme (e.g., a CRISPR nuclease).
28. The RNA molecule pair of claim 25, wherein the protein is a chimeric antigen receptor, an immunomodulatory protein, or a transcription factor.
29. The RNA molecule pair according to any one of claims 1-24, wherein the target nucleotide sequence is a non-protein coding sequence.
30. The RNA molecule pair of claim 29, wherein the non-protein coding sequence is selected from the group consisting of: antisense RNA, aptamers, guide RNA, and any other non-protein coding RNA present in any organism.
31. The RNA molecule pair as claimed in any of the preceding claims, wherein the length of the target nucleotide sequence is at least 10, 20, 40, 60, 80, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 10000, or 20000 nucleotides.
32. The RNA molecule pair as claimed in any of the preceding claims, wherein the first RNA molecule and / or the second RNA molecule comprises modified nucleotides.
33. The RNA molecule pair of claim 32, wherein the modified nucleotide comprises a modified nucleoside and / or a modified phosphate and / or a modified nucleotide linker, such as cytidine modification, uridine modification, guanine modification, or adenosine modification, such as 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (ψ), etc. ), 5-methoxyuridine (5 moU), 2'-deoxy-2'-fluoro-ribose (2'-F), 2'-O-methylribose (2'-O-Me) or UNA nucleotides, acyclic sugars, thiophosphate linkages or thioaminophosphate linkages.
34. The RNA molecule pair as claimed in claim 32 or 33, wherein the RNA molecule contains less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, or less than 1%.
35. The RNA molecule pair according to any one of claims 1-31, wherein the molecules are unmodified.
36. The RNA molecule pair as claimed in any of the preceding claims, wherein the first RNA molecule further comprises a 3' end binding motif, and the second RNA molecule further comprises a 5' end binding motif, wherein the 3' end binding motif and the 5' end binding motif are complementary to form a double-stranded region, preferably, the 3' end binding motif and the 5' end binding motif are completely complementary to each other.
37. The RNA molecule pair of claim 36, wherein the length of the 3' end binding motif or the 5' end binding motif is about 5-50, about 10-50, about 20-50, about 30-50, about 40-50, about 5-40, about 5-30, about 5-20, or about 10-20 nucleotides; preferably, the length is about 10, 15, or 19 nucleotides.
38. A nucleic acid vector encoding a first RNA molecule and / or a second RNA molecule as described in any of the preceding claims.
39. The nucleic acid vector of claim 38, further comprising an RNA polymerase promoter sequence operatively linked to the coding sequence of the first RNA molecule and / or the second RNA molecule.
40. The nucleic acid vector of claim 39, wherein the RNA promoter is selected from the T7 RNA polymerase promoter, the T6 viral RNA polymerase promoter, the SP6 viral RNA polymerase promoter, the T3 viral RNA polymerase promoter, or the T4 viral RNA polymerase promoter.
41. The nucleic acid vector according to any one of claims 38-40, wherein the vector is selected from linear DNA molecules, plasmids, viral vectors, granules, bacterial artificial chromosomes (BAC) or yeast artificial chromosomes (YAC).
42. A method for preparing circular RNA, wherein the method comprises: i) Providing or obtaining an RNA molecule pair as described in any one of claims 1-37; ii) Add a buffer solution to the RNA molecule pair to obtain a reaction mixture to allow the formation of the scaffold and catalytic domains of the type I introns, and iii) Add GTP and divalent metal cations to the reaction mixture at a temperature that allows the formation of the circular RNA. Optionally, the method further includes: iv) Harvest the circular RNA formed in step iii).
43. The method of claim 42, wherein the first RNA molecule and the second molecule are provided in the reaction mixture in a molar ratio of 1:1, 1:2, 1:3, 1:4, or 1:
5.
44. The method of claim 42 or 43, wherein the buffer solution comprises 10-200 mM Tris-HCl at pH 6-8.5, such as pH 6, 6.5, 7, 7.5, 8 or 8.5, for example 50 mM Tris-HCl.
45. The method of any one of claims 42-44, wherein the concentration of GTP in the reaction mixture is from 100 nM to 2 mM.
46. The method of any one of claims 42-45, wherein the divalent metal cation is Mg2+, and the concentration of the divalent metal cation in the reaction mixture is at least about 5 mM, for example about 5 mM to about 550 mM, for example at least about 5 mM, about 10 mM, about 15 mM, at least about 20 mM, at least about 30 mM, at least about 40 mM, at least about 50 mM, at least about 60 mM, at least about 70 mM, at least about 80 mM, at least about 90 mM, at least about 100 mM, at least about 125 mM, at least about 150 mM, at least about 175 mM, at least about 200 mM, at least about 250 mM, at least about 300 mM, at least about 350 mM, at least about 400 mM, at least about 450 mM, at least about 500 mM, at least about 550 mM or higher.
47. A circular RNA produced by the method described in any one of claims 42-46.
48. A composition comprising an effective amount of the circular RNA as described in claim 47 and a pharmaceutically acceptable carrier, preferably wherein the pharmaceutically acceptable carrier comprises a lipid, a polymer, or a lipid-polymer hybrid, such as lipid nanoparticles (LNPs), lipid microparticles, lipid suspensions, or liposomes.
49. A cell, such as a eukaryotic cell, comprising a nucleic acid vector as described in any one of claims 38-41 or a circular RNA as described in claim 47.
50. A composition comprising the cells as described in claim 49 and one or more pharmaceutically or physiologically acceptable carriers, excipients, or diluents.