Precursor RNA molecules and methods for producing circular rnas
The ICI system addresses limitations in circRNA production by integrating Group I/II introns with linkers and spacers for precise circularization, enhancing efficiency and stability while minimizing immune response.
Patent Information
- Application Number
- PCT/CN2025/109554
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-30
- Filing Date
- 2025-07-21
- Publication Date
- 2026-01-22
AI Technical Summary
Existing methods for producing circular RNA in vitro, such as chemical and enzymatic methods, are limited in size and efficiency, and ribozyme catalysis methods are not fully optimized for precise circularization without disrupting intron integrity.
A novel Integrated Catalytic Intron (ICI) system integrates Group I/II introns into the 5' end of RNA precursors, using linkers, spacers, and homology arms to achieve efficient and precise circularization of circRNA without disrupting the intron structure.
The ICI system enables high circularization efficiency and stability of circRNA, reducing immunogenicity and maintaining intron integrity, with improved size flexibility and reduced immune response.
Smart Images

Figure PCTCN2025109554-FTAPPB-I100001 
Figure PCTCN2025109554-FTAPPB-I100002 
Figure PCTCN2025109554-FTAPPB-I100003
Abstract
Description
Precursor RNA Molecules and Methods for Producing Circular RNAsTechnical field
[0001] The present invention relates to engineered precursor RNA molecules, vectors as well as methods for producing circular RNA molecules. The present invention also relates to circular RNA molecules thus produced and a composition comprising the same.Background
[0002] Circular RNA is a common type of RNA in eukaryotes. Naturally occurring circular RNAs are primarily produced by a molecular mechanism within cells called "back-splicing" . Eukaryotic circular RNAs have been found to have a variety of molecular and cellular regulatory functions. For example, circular RNAs can regulate the expression of target genes by binding to microRNAs. Circular RNAs can also regulate gene expression by directly binding to target proteins.
[0003] Due to its circular nature, circular RNA has a longer half-life than linear mRNA, so it is speculated that circular RNA synthesized in vitro may have higher stability. Methods of forming circular RNAs in vitro include the chemical method, the enzymatic catalysis method, and the ribozyme catalysis method. Chemical methods are expensive and the size of the circular RNA molecules that can be produced is limited. The enzymatic method mainly utilizes T4 RNA ligase to catalyze the circularization of linear RNA, and the size of RNA payload that can achieve circularization is also limited. Ribozyme catalysis (e.g., based on Group I introns) is a promising method for the preparation of circular RNAs.Summary of the invention
[0004] Circular RNA (circRNA) represents a unique class of single-stranded RNA molecules characterized by a covalently closed loop structure. Naturally occurring circRNAs are generally formed through a non-canonical splicing process known as ‘back-splicing’ . Unlike linear mRNAs, circRNAs do not require a 5’-cap or 3’-poly (A) tail for stability. Group I Intron-mediated circularization is one of the commonly used methods for in vitro preparation of circRNAs and Group II introns can also be used for such circularization. The specific circularization principle is as follows. The 5’ splice site, as marked by a conserved U-G pair, undergoes nucleophilic attack by the 3’-OH group of exogenous GTP bound at the G-binding site of the intron. Subsequently, a conformational change displaces the guanosine from the G-binding site by the 3’-terminal omega G (ωG) that signifies the 3’ splice site. The 3’-OH group of the terminal residue of the 5’ exon then participates in a reaction that mirrors the initial step. Ultimately, the 5’ and 3’ exons are joined, releasing the intron.
[0005] The permuted intron and exon (PIE) circularization system divides a Group I intron into two parts and positions them on both sides of the targeted RNA sequence. This arrangement requires that the two separated intron segments fold properly without interference from adjacent sequences. Additional non-structural segments, also known as spacers, are incorporated between the target RNA sequence and intron fragments to ensure correct folding and circularization. The classical PIE circularization system leverages modified Group I introns to achieve efficient RNA circularization by facilitating transesterification reactions at defined splice sites. In this study, we propose a novel integrated catalytic intron (ICI) system, which integrates the ribozymes (i.e., introns) as a whole into the 5’ end of the RNA precursors and achieves efficient and precise circularization of circRNA without disrupting the integrity of the intron.
[0006] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising:
[0007] a) a Group I / II intron,
[0008] b) a 3’ exon,
[0009] c) a sequence of interest (TOI) , and
[0010] d) a 5’ exon.
[0011] In embodiments of the above aspect, the circular RNA precursor further comprises linkers.
[0012] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0013] a) a Group I / II intron,
[0014] b) a 3’ exon,
[0015] c) a sequence of interest (TOI) , and
[0016] d) 5’ exon.
[0017] In another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0018] a) a Group I / II intron,
[0019] b) a 3’ exon,
[0020] c) a linker A,
[0021] d) a sequence of interest (TOI) ,
[0022] e) a linker B, and
[0023] f) a 5’ exon.
[0024] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0025] a) a linker I,
[0026] b) a Group I / II intron,
[0027] c) a 3’ exon,
[0028] d) a sequence of interest (TOI) ,
[0029] e) a 5’ exon, and
[0030] f) a linker II.
[0031] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0032] a) a linker I,
[0033] b) a spacer A,
[0034] c) a Group I / II intron,
[0035] d) a 3’ exon,
[0036] e) a sequence of interest (TOI) ,
[0037] f) a 5’ exon,
[0038] g) a spacer B, and
[0039] h) a linker II.
[0040] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0041] a) a linker I,
[0042] b) a spacer A,
[0043] c) a Group I / II intron,
[0044] d) a 3’ exon,
[0045] e) a linker A,
[0046] f) a sequence of interest (TOI) ,
[0047] g) a linker B,
[0048] h) a 5’ exon,
[0049] i) a spacer B, and
[0050] j) a linker II.
[0051] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0052] a) a Group I / II intron, and
[0053] b) a sequence of interest, comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end.
[0054] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0055] a) a linker I,
[0056] b) a Group I / II intron,
[0057] c) a sequence of interest (TOI) , comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end, and
[0058] d) a linker II.
[0059] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0060] a) a linker I,
[0061] b) a Group I / II intron,
[0062] c) a sequence of interest (TOI) , comprising: (i) an IRES fragment I which comprises a 3’ exon at its 5’ end, (ii) an ORF sequence, and (iii) an IRES fragment II comprising a 5’ exon at its 3’ end, and
[0063] d) a linker II.
[0064] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0065] a) a linker I,
[0066] b) a Group I / II intron,
[0067] c) a sequence of interest (TOI) , comprising: (i) an ORF sequence fragment I which comprises a 3’ exon at its 5’ end, (ii) an IRES sequence, and (iii) an ORF sequence fragment II comprising a 5’ exon at its 3’ end, and
[0068] d) a linker II.
[0069] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0070] a) a homology arm I,
[0071] b) a spacer A,
[0072] c) a Group I / II intron,
[0073] d) exon I,
[0074] e) a sequence of interest (TOI) , comprising: (i) an ORF sequence and (ii) an IRES sequence,
[0075] f) exon II,
[0076] g) a spacer B,
[0077] h) a homology arm II,
[0078] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0079] wherein said spacer A and B are non-complementary sequences.
[0080] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0081] a) a homology arm I,
[0082] b) a spacer A,
[0083] c) a Group I / II intron comprising a pairing G group and a 3’ end G group,
[0084] d) exon I,
[0085] e) a sequence of interest (TOI) , comprising: (i) an ORF sequence, and (ii) an IRES sequence,
[0086] f) exon II comprising a 3’ end U group,
[0087] g) a spacer B,
[0088] h) a homology arm II,
[0089] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0090] wherein said spacer A and B are non-complementary sequences, and
[0091] wherein said Group I / II intron forms a U-G base pair by its pairing G group with the 3’ end U of exon II.
[0092] In embodiments of the above aspect, the precursor further comprises a linker A located between the exon I and the TOI sequence, and a linker B located between the exon II and the TOI sequence, wherein the linker A and the linker B are partially complementary to each other and together are capable of forming a duplex.
[0093] In embodiments of the above aspect, said spacer A and B are of the same length. In embodiments of the above aspects, said spacer A and B are of different length.
[0094] In embodiments of the above aspects, the sequence of interest comprises a spacer C located between the IRES sequence and the ORF sequence.
[0095] In aspects and embodiments of the present invention, the Group I / II intron is a Group I self-splicing intron, which can be selected from the following group: cyanobacterium Anabaena Group I Intron, Azoarcus Group I Intron (Azo) , Scytalidium dimidiatum Group I Intron (Sd) , Staphylococcus phage Twort Group I Intron (Twort) , Scytonema-hofmani tRNA fMet group I intron (Sh) , and Agrobacterium-tumefaciens group I intron (At) .
[0096] In aspects and embodiments of the present invention, the 5’-exon and the 5’ end of the intron form a P1 structure containing a ribozyme recognition sequence I, and the 3’-exon and both the 5’ and 3’ ends of the intron form a P10 structure containing a ribozyme recognition sequence II. As used herein, the ribozyme recognition sequence refers to the specific sequence that the ribozyme recognizes in order to achieve self-cleavage / splicing, or the specific structure formed by this sequence.
[0097] In aspects of the present invention, the linker I sequence and the linker II sequence are functional sequences which can improve circularization. For example, linker I and linker II are homology arm sequences capable of complementarily pairing to each other to form a homology arm double-stranded region.
[0098] In some embodiments, linker I is upstream of the self-splicing intron and linker II is downstream of the 5’ exon.
[0099] In some embodiments, linker I or linker II is about 5-200 nucleotides in length, about 5-150 nucleotides in length, about 5-100 nucleotides in length, about 5-80 nucleotides in length, about 5-50 nucleotides in length, about 5-40 nucleotides in length, about 5-30 nucleotides in length, about 5-20 nucleotides in length, about 5-10 nucleotides in length, preferably, about 5, 10, 15, 20, 25, 30, 35 or 40 nucleotides in length.
[0100] In aspects of the present invention, linker A and linker B comprise homology arms and / or spacers. In embodiments, linker A and linker B are homology arm sequences capable of complementarily pairing to each other to form a homology arm double-stranded region. In embodiments, linker A and linker B are spacers. In embodiments, linker A and linker B both comprise a homology arm sequence and a spacer.
[0101] As used herein, a “spacer” refers to any contiguous nucleotide sequence that is predicted to be non-interfering with other nearby or adjacent structures in the circular RNA precursor, for example, the IRES, the sequence of interest, or intron. As used herein, the term "self-splicing intron" refers to an intron having self-splicing ribozyme activity and capable of excising itself and joining two flanking exons. In some embodiments, the splicing is autocatalytic splicing.
[0102] In some embodiments, the ORF sequence of the present invention is a protein-coding sequence. In some embodiments, the ORF sequence of the present invention is a non-coding sequence. It is to be understood that each of the method of prevention or treatment embodiments herein can also be formulated as corresponding use type embodiments.Brief description of the drawings
[0103] Figure 1. The influence of linkers located within the precursor (internal linkers) or external linkers located at both ends of the precursor on the circularization, in the ICI based circularization system.
[0104] Figure 2. The influence of Homology arms on circularization, in the PIE based circularization system.
[0105] Figure 3. Introduction of spacers between homology arm and intron / exon can significantly promote circularization efficiency on ICI system.
[0106] Figure 4. Impact of length of spacers between homology arm and intron / exon on ICI system circularization efficiency.
[0107] Figure 5. Impact of P1-disruptive spacers between homology arm and intron / exon on P1 structure compromises ICI circularization efficiency.
[0108] Figure 6. The influence of P1 sequence modification on the ICI system circularization efficiency.
[0109] Figure 7. The structure and circularization efficiency of ICI v3.0 system.
[0110] Figure 8. Pairing strength of homology arm affects ICI circularization efficiency.
[0111] Figure 9. RNA circularization efficiency of ICI v3.0 system with different Group I Introns.
[0112] Figure 10. Optimized ICI system has improved circularization efficiency.
[0113] Figure 11. Influence of the number of pairing bases and pairing strength on exon duplex on the circularization efficiency of the ICI system.
[0114] Figure 12. Circularization efficiency and expression level of seamless ICI system.Detailed description of the invention
[0115] In the present invention, unless indicated otherwise, the scientific and technological terminologies used herein refer to meanings commonly understood by a person skilled in the art. Also, the terminologies and experimental procedures used herein relating to protein and nucleotide chemistry, molecular biology, cell and tissue cultivation, microbiology, immunology, all belong to terminologies and conventional methods generally used in the art. For example, the standard DNA recombination and molecular cloning technology used herein are well known to a person skilled in the art, and are described in details in the following references: Sambrook, J., Fritsch, Efland Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989. In the meantime, in order to better understand the present invention, definitions and explanations for the relevant terminologies are provided below.
[0116] As used herein, the singular forms "a" , "an" , and "the" include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. For example, a circular RNA precursor refers to one or more circular RNA precursors. As such, the terms "a" , "an" , "one or more" and "at last one" can be used interchangeably. For example, the term “at least one” refers to one, two or more. This statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely, " "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation. Similarly the terms "comprising" , "including" and "having" can be used interchangeably.
[0117] As used herein, the term "and / or" encompasses all combinations of items connected by the term, and each combination should be regarded as individually listed herein. For example, "A and / or B" covers "A" , "A and B" , and "B" . For example, "A, B, and / or C" covers "A" , "B" , "C" , "A and B" , "A and C" , "B and C" , and "A and B and C" .
[0118] As used herein, "about" , "approximately" , "substantially" and "significantly" will be understood by persons of ordinary skill in the art and will vary to some extent on the context in which they are used. If there are uses of these terms which are not clear to persons of ordinary skill in the art given the context in which they are used, "about" and "approximately" will mean plus or minus <10%of the particular term and "substantially" and "significantly" will mean plus or minus >10%of the particular term.
[0119] In the present application, "optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes instances where said event or circumstance occurs and instances in which it does not.
[0120] "Polynucleotide" , "nucleic acid sequence" , "nucleotide sequence" , or "nucleic acid fragment" are used interchangeably to refer to a polymer of RNA or DNA that is single-or double-stranded, optionally containing synthetic, non-natural or altered nucleotide bases. Nucleotides (usually found in their 5’-monophosphate form) are referred to by their single letter designation as follows: "A" for adenylate or deoxyadenylate (for RNA or DNA, respectively) , "C" for cytidylate or deoxycytidylate, "G" for guanylate or deoxyguanylate, "U" for uridylate, "T" for deoxythymidylate, "R" for purines (A or G) , "Y" for pyrimidines (C or T) , "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide. Although the nucleotide sequences herein may be represented as DNA sequences (comprising T (s) ) , when referring to RNA, one skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U) .
[0121] For example, the nucleotide sequence of interest in the present invention can be a non-coding sequence, such as an antisense RNA, aptamer, guide RNA, or non-coding RNA existing in any organism.
[0122] Sequence "identity" has recognized meaning in the art, and the percentage of sequence identity between two nucleic acids or polypeptide molecules or regions can be calculated using the disclosed techniques. Sequence identity can be measured along the entire length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example, Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991) . There are many methods for determining sequence identity. An example of algorithms suitable for determining percent sequence identity is the algorithm used in the Basic Local Alignment Search Tool (hereinafter "BLAST" ) , see e.g., Altschul et al., J. Mol. Biol. 215: 403-410, 1990 and Altschul et al, Nucleic Acids Res., 15: 3389-3402, 1997. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (hereafter "NCBI" ) . Default parameters used to determine sequence identity using software available from NCBI (such as BLASTN for nucleic acid sequences) are described in McGinnis et al. Nucleic Acids Res., 32: W20-W25, 2004.
[0123] A polynucleotide sequence (e.g. an RNA polynucleotide sequence or a fragment thereof) "derived from" a designated polynucleotide sequence means the former polynucleotide sequence originates from the latter. In some embodiments, the polynucleotide sequence which is derived from a particular polynucleotide sequence has a polynucleotide sequence that is identical, essentially identical or homologous to that particular sequence or a fragment thereof. Polynucleotide sequences derived from a particular polynucleotide sequence may be variants of that particular sequence or a fragment thereof. For example, it will be understood by one of ordinary skill in the art that the circular RNA molecules suitable for use herein may be altered such that they vary in sequence from the sequences from which they were derived, while retaining the desirable activity thereof.
[0124] "Circular RNA precursor" herein refers to a linear RNA molecule capable of forming a covalently linked closed circular RNA molecule, e.g., by self-splicing. The circular RNA precursor may be produced by transcription from a nucleic acid vector comprising a coding sequence of the circular RNA precursor. Alternatively, the circular RNA precursor may also be obtained by chemical synthesis. As used herein, the terms “circular RNA precursor” and “linear RNA molecule for producing a circular RNA” have the same meaning and refer to the same thing, i.e., a linear RNA molecule capable of forming a covalently linked closed circular RNA molecule, e.g., by self-splicing under the action of the self-splicing intron, and thus they can be interchangeably used in the context of all aspects of the present invention described herein.
[0125] In some embodiments, the circular RNA precursor is capable of forming a covalently linked closed circular RNA molecule by self-splicing under the action of the self-splicing introns or intron fragments and the residual circularizing elements. As used herein, the term “residual circularizing elements” refer to exon I (i.e., 3’ exon) and exon II (i.e., 5’ exon) in the circular RNA precursor of the present invention which are elements necessary to effect self-splicing and circularization of the precursor and are retained in the circular RNA molecule thus generated.
[0126] As used herein, the term "self-splicing intron" refers to an intron having self-splicing ribozyme activity and capable of excising itself and joining two flanking exons. In some embodiments, the splicing is autocatalytic splicing. In the context of the present application, the first nucleotide of the 5’ end region of the Group I / II intron is always the nucleotide natually occurring at the naitve 5’ splice site of the Group I / II intron, and the last nucleotide of the 3’ end region of the Group I / II intron is the nucleotide natually occurring at the naitve 3’ splice site of the Group I / II intron. In preferred embodiments of the circular RNA precursor, the self-splicing intron is a naturally occurring Group I intron.
[0127] "Self-splicing introns" include, but are not limited to, Group I introns and Group II introns. Group I introns contain 14 subgroups, while most of the Group I introns belong to the IC3 subgroup. For example, the Group I intron may be a Group I intron of the cyanobacterium Anabaena belonging to the IC3 subgroup or a Group I intron from a T4 phage Group I intron belonging to the IA2 subgroup or a Group I intron from Azoarcus sp. BH72 belonging to the IC3 subgroup. Additional examples of self-splicing introns useful in the present invention include, but are not limited to, self-splicing introns derived from the following organisms: Enterobacteriophage T4, Bacteriophage Twort, Bacteriophage SPO1, Bacteriophage S3b, Bacillus anthracis, Clostridium botulinum, Tetrahymena thermophila , Dunaliella parva, Pneumocystis carinii, Physarum polycephalum, Anabaena sp. PCC7120, Scytonema hofmanni, Agrobacterium tumefaciens, Synechocystis PCC 6803, Synechococcus elongatus PCC 6301, Neurospora crassa, Candida albicans, Scytalidium cerradiumydiaces, Pediadiaces Chlamydomonas nivalis, Chlorella vulgaris, Amoebidium parasiticum, Neurospora crassa, Emericella nidulans, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Neochloris aquatica, Dunaliella parva, Symkania negevensis, Emericella nidulans. See e.g., Vicens, Q., et al., (2008) . Toward predicting self-splicing and protein-facilitated splicing of group I introns. RNA 14: 2013-2029; Tanner, A. M., et al., (1996) . Activity and thermostability of the small self-splicing group I intron in the pre-tRNAIIe of the purple bacterium Azoarcus. RNA 2: 74-83.
[0128] In some embodiments, the self-splicing intron is selected from the following group: cyanobacterium Anabaena Group I Intron such as AnaX, Azoarcus Group I Intron (Azo) , Scytalidium dimidiatum Group I Intron (Sd) , Staphylococcus phage Twort Group I Intron (Twort) , Scytonema-hofmani tRNA fMet Group I Intron (Sh) , or Agrobacterium-tumefaciens Group I Intron (At) . In some embodiments, the self-splicing intron has a sequence with 85%or more sequence identity to the above-mentioned self-splicing intron. In some embodiments, the self-splicing intron is a fragment of the above-mentioned self-splicing intron.
[0129] As used herein, “ORF sequence” refers to a sequence comprising a protein coding sequence or non-coding sequence.
[0130] “Circularization efficiency” as used herein may refer to the ratio of outcome circular RNA to input precursor in a given time period. Alternatively, “Circularization efficiency” as used herein may refer to the ratio of desired circular RNA to linear circRNA precursor or the ratio of circRNA to the sum of circRNA and linear circRNA precursor in the final product in a given time period. The circularizing efficiency can be determined by methods well known in the art, such as those described in Example 1 of the present application. It is expected that the circular RNA precursors of the present invention will have desired circularization efficiency.
[0131] “Reduced immunogenicity” as used herein may refer to that the circular RNA, upon contacted with cells, elicits a reduced immune response, i.e., an immune response at a level lower than a control circular RNA or control linear RNA. For example, reduced immune response refers to reduced expression of cytokines. The cytokines include, but are not limited to, IFNβ, TNFα, IL6 and / or RIG-I. In some embodiments, the immunogenicity of the circular RNA is reduced by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%or more. It is expected that the circular RNA precursors of the present invention will have reduced immunogenicity.
[0132] As used herein, "spacer" refers to any contiguous nucleotide sequence that at least does not negatively interfere with the function of the elements it connects. Generally, if it is desired to avoid the interaction of two near or adjacent elements, a spacer can be inserted between the two elements. The spacer sequences described herein can serve two functions: (1) to facilitate circularization and (2) to facilitate functionality by allowing correct folding of the residual circularizing element and the nucleotide sequence of interest (e.g., IRES) . In some embodiments, the spacer is no more than 150, no more than 100, no more than 50, no more than 30, no more than 10, no more than 5, or no more than 3 nucleotides in length. In some embodiments, the spacer is 5 nucleotides in length. In some embodiments, the spacer is 4 nucleotides in length. In some embodiments, the spacer is 3 nucleotides in length. In some embodiments, the first spacer may be absent. In some embodiments, the second spacer may be absent. In some embodiments, the first spacer and the second spacer may be absent.
[0133] As used herein, "ICI based circularization system” or “ICI system” , used interchangeably, refers to the Integrated Catalytic Intron (ICI) based circularization system of the present invention. For example, a precursor RNA molecule may contain the target of interesting (TOI) sequence, a Group I Intron, 3’-and 5’-exon sequences, and optionally homology arms and spacers which can facilitate RNA circularization. In one embodiment, the specific cleavage site conserved sequence of exon is at its 3’ end is cleaved by the nucleophilic attack of the free 3’ hydroxyl of guanylate, so that the 5’ exon produces a naked 3’ hydroxyl. Subsequently, the exposed 3’ hydroxyl of 5’ exon attacks the conserved sequence between the Group I Intron and 3’ exon (e.g., ωG) , and the Group I Intron is excised. Thus the two exons (exon 1 and exon 2) undergo a circularizing reaction to obtain the circular RNA.
[0134] In the context of the present application, a TOI sequence may also be referred to as “asequence of interest” . A TOI sequence, or a sequence of interest, may comprise a translation initiation element, such as an internal ribozyme entry site (IRES) , and an ORF sequence. In embodiments, the ORF sequence is a protein coding sequence. In embodiments, the TOI is a non-protein coding sequence. In one embodiment, the TOI sequence encodes an ornithine transcarbamylase (OTC) . In one embodiment, the TOI encodes a bone morphogenetic protein 2 (BMP2) . In one embodiment, the TOI encodes a collagen type III alpha 1 chain (COL3A1) .
[0135] As used herein, the terms “cleavage site” , “splice site” and “splicing site” can be used interchangeably.
[0136] Circular RNA precursor, vector, and method for preparation of circular RNA
[0137] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising:
[0138] a) a Group I / II intron,
[0139] b) a 3’ exon,
[0140] c) a sequence of interest, and
[0141] d) a 5’ exon.
[0142] In embodiments of the above aspect, the circular RNA precursor further comprises linkers.
[0143] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0144] a) a Group I / II intron,
[0145] b) a 3’ exon,
[0146] c) a sequence of interest, and
[0147] d) 5’ exon.
[0148] In another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0149] a) a Group I / II intron,
[0150] b) a 3’ exon,
[0151] c) a linker A,
[0152] d) a sequence of interest,
[0153] e) a linker B, and
[0154] f) a 5’ exon.
[0155] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0156] a) a linker I,
[0157] b) a Group I / II intron,
[0158] c) a 3’ exon,
[0159] d) a sequence of interest,
[0160] e) a 5’ exon, and
[0161] f) a linker II.
[0162] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0163] a) a linker I,
[0164] b) a spacer A,
[0165] c) a Group I / II intron,
[0166] d) a 3’ exon,
[0167] e) a sequence of interest,
[0168] f) a 5’ exon,
[0169] g) a spacer B, and
[0170] h) a linker II.
[0171] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0172] a) a linker I,
[0173] b) a spacer A,
[0174] c) a Group I / II intron,
[0175] d) a 3’ exon,
[0176] e) a linker A,
[0177] f) a sequence of interest,
[0178] g) a linker B,
[0179] h) a 5’ exon,
[0180] i) a spacer B, and
[0181] j) a linker II.
[0182] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0183] a) a Group I / II intron, and
[0184] b) a sequence of interest, comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end.
[0185] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0186] a) a linker I,
[0187] b) a Group I / II intron,
[0188] c) a sequence of interest, comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end, and
[0189] d) a linker II.
[0190] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0191] a) a linker I,
[0192] b) a Group I / II intron,
[0193] c) a sequence of interest, comprising: (i) an IRES fragment I which comprises a 3’ exon at its 5’ end, (ii) an ORF sequence, and (iii) an IRES fragment II comprising a 5’ exon at its 3’end, and
[0194] d) a linker II.
[0195] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0196] a) a linker I,
[0197] b) a Group I / II intron,
[0198] c) a sequence of interest, comprising: (i) an ORF sequence fragment I which comprises a 3’ exon at its 5’ end, (ii) an IRES sequence, and (iii) an ORF sequence fragment II comprising a 5’ exon at its 3’ end, and
[0199] d) a linker II.
[0200] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0201] a) a homology arm I,
[0202] b) a spacer A,
[0203] c) a Group I / II intron,
[0204] d) exon I,
[0205] e) a sequence of interest, comprising: (i) an ORF sequence, and (ii) an IRES sequence,
[0206] f) exon II,
[0207] g) a spacer B,
[0208] h) a homology arm II,
[0209] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0210] wherein said spacer A and B are non-complementary sequences.
[0211] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0212] a) a homology arm I,
[0213] b) a spacer A,
[0214] c) a Group I intron comprising a pairing G group and a 3’ end G group,
[0215] d) exon I,
[0216] e) a sequence of interest, comprising: (i) an ORF sequence, and (ii) an IRES sequence,
[0217] f) exon II comprising a 3’ end U group,
[0218] g) a spacer B,
[0219] h) a homology arm II,
[0220] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0221] wherein said spacer A and B are non-complementary sequences, and
[0222] wherein said Group I intron forms a U-G base pair with the 3’ end U of exon II.
[0223] In embodiments of the above aspect, the precursor further comprises a linker A located between the exon I and the TOI sequence, and a linker B located between the exon II and the TOI sequence, wherein the linker A and the linker B are partially complementary to each other and together are capable of forming a duplex.
[0224] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0225] a) optionally a linker I,
[0226] b) optionally a spacer A,
[0227] c) a Group I / II intron,
[0228] d) exon I,
[0229] e) optionally a linker A,
[0230] f) a sequence of interest,
[0231] g) optionally a linker B,
[0232] h) exon II,
[0233] i) optionally a spacer B,
[0234] j) optionally a linker II.
[0235] In embodiments of the above aspects, said spacer A and B are of the same length. In embodiments of the above aspects, said spacer A and B are of different length.
[0236] In the context of the present application, the first nucleotide of the 5’ end region of the Group I / II intron is always the nucleotide natually occurring at the naitve 5’ splice site of the Group I / II intron, and the last nucleotide of the 3’ end region of the Group I / II intron is the nucleotide natually occurring at the naitve 3’ splice site of the Group I / II intron. In preferred embodiments of the circular RNA precursor, the self-splicing intron is a naturally occurring Group I intron.
[0237] In embodiments of the above aspects, the Group I / II intron comprises a pairing G group and / or a 3’ end G group. In embodiments of the above aspects, said Group I / II intron forms a U-G base pair with the 3’ end U of exon II. In embodiments of the above aspects, the Group I / II intron comprises modification compared to the native Group I / II intron. In embodiments of the above aspects, the Group I / II intron does not comprise modification compared to the native Group I / II intron. In embodiments of the above aspects, the modification is not located on the sequence forming P1 structure compared to the native Group I / II intron.
[0238] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a native self-splicing introns of the following organisms: Enterobacteriophage T4, Bacteriophage Twort, Bacteriophage SPO1, Bacteriophage S3b, Bacillus anthracis, Clostridium botulinum, Tetrahymena thermophila, Dunaliella parva, Pneumocystis carinii, Physarum polycephalum, Anabaena sp. PCC7120, Scytonema hofmanni, Agrobacterium tumefaciens, Synechocystis PCC 6803, Synechococcus elongatus PCC 6301, Neurospora crassa, Candida albicans, Scytalidium cerradiumydiaces, Pediadiaces Chlamydomonas nivalis, Chlorella vulgaris, Amoebidium parasiticum, Neurospora crassa, Emericella nidulans, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Neochloris aquatica, Dunaliella parva, Symkania negevensis, Emericella nidulans. In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I self-splicing intron, which can be selected from the following group: cyanobacterium Anabaena Group I Intron such as AnaX, Azoarcus Group I Intron (Azo) , Scytalidium dimidiatum Group I Intron (Sd) , Staphylococcus phage Twort Group I Intron (Twort) , Scytonema-hofmani tRNA fMet group I intron (Sh) , or Agrobacterium-tumefaciens group I intron (At) . In some embodiments, the Group I / II intron has a sequence with at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to a native self-splicing intron (without any modification on the sequence forming P1 structure) . In some embodiments, the Group I / II intron retains the self-splicing activity of the self-splicing intron (a native self-splicing intron) . In some embodiments, the Group I / II intron retains more than 60%, more than 70%, more than 80%, more than 90%, more than 95%self-splicing activity compared to the (native) full length Group I / II intron. In some embodiments, the Group I / II intron is a native full length Group I / II intron or a Group I / II intron derived from a native Group I / II intron which retains more than 60%, more than 70%, more than 80%, more than 90%, more than 95%self-splicing activity compared to the native full length Group I / II intron. In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group II intron.
[0239] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron selected from Anabaena sp. YBS01, Staphylococcus phage Twort, Ncr. m. ND5, 1, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, Bfu. S1506, Mpl. L798, Pan. m. ND3, 1, Osp. S1199, Azoarcus olearius BH72, Scytalidium dimidiatum, Scytonema-hofmani tRNA fMet, and Agrobacterium-tumefaciens, preferably Anabaena sp. YBS01, Staphylococcus phage Twort, . Azoarcus olearius BH72, Scytonema-hofmani tRNA fMet, Agrobacterium-tumefaciens, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, and Pan. m. ND3, 1. In aspects and embodiments of the present invention, the Group I / II intron comprises or consists of a sequence selected from any one of SEQ ID NOs: 16-30 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NOs: 16-30 (preferably SEQ ID NO: 16, 17, 19, 20, 21, 22, 25, 27, 29, 30) , or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from any one of SEQ ID NOs: 16-30 (preferably SEQ ID NO: 16, 17, 19, 20, 21, 22, 25, 27, 29, 30) (without any modification on the sequence forming P1 structure) .
[0240] In aspects and embodiments of the present invention, the 5’-exon and the 5’ end of the intron form a P1 structure containing a ribozyme recognition sequence I, and the 3’-exon and both the 5’ and 3’ ends of the intron form a P10 structure containing a ribozyme recognition sequence II. As used herein, the ribozyme recognition sequence refers to the specific sequence that the ribozyme recognizes in order to achieve self-cleavage / splicing, or the specific structure formed by the participation of this sequence. As used herein, “exon I” and “exon 1” can also be used interchangeably with “3’-exon” , and “exon II” and “exon 2” can be used interchangeably with “5’-exon” . As used herein, in terms of a Group I intron that is used in the aspects and embodiments of the circular RNA precursor of the present invention, “5’-exon” refers to an RNA sequence which is the 3’ terminal portion of a 5’-flanking exon of the Group I intron (e.g., the exon naturally occurring 5’-upstream of the Group I intron) , and “3’-exon” refers to an RNA sequence which is the 5’ terminal portion of a 3’-flanking exon of the Group I intron (e.g., the exon naturally occurring 3’-downstream of the same Group I intron) . According to the present invention, the 5’-exon and 3’-exon are configured in such a way that the 3’ end region of the 5’-exon and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and the 5’ end region of the 3’-exon and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II. In such a P1 structure, the 5’ end region of the intron alone forms a part of the stem-loop structure of the P1 structure. That is, the 5’ end region of the Group I intron itself forms the loop and a part of the stem of the P1 structure, while the 3’ end region of the 5’-exon, by pairing with the 5’ end region of the Group I intron, forms the remaining part of the stem of the P1 structure with the 5’ end region of the Group I intron. The 5’ end region of the 5’-exon (exon II) and the 3’ end region of the 3’-exon (exon I) are partially complementary to each other and together are capable of forming a stem-loop-like structure. The 5’-flanking exon of the Group I intron and the 3’-flanking exon of the Group I intron can independently be either a naturally occurring exon sequence or a modified exon sequence.
[0241] As used herein, "ribozyme recognition sequence" refers to a segment of nucleotide sequence which belongs to a part of the exon. The ribozyme recognition sequence is involved in or required for circularization by the Group I / II intron and participates in circularization together with the Group I / II intron but is retained in the final circular RNA. In embodiments, the ribozyme recognition sequence I or II comprises 2 to 50 nucleotides, for example 2 to 40, 2 to 30, 2 to 25, 2 to 20, 2 to 15, 2 to 10 nucleotides, in particular 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more nucleotides. In embodiments, the ribozyme recognition sequence I is comprised of 1 to 10 nucleotides, in particular 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. In embodiments, the ribozyme recognition sequence II is comprised of 1 to 10 nucleotides, in particular 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. As used herein, the terms "ribozyme recognition sequence" and “ribozyme recognition site” can be used interchangeably.
[0242] In embodiments, the ribozyme recognition sequence I is capable of connecting with the 5’ end region of the Group I / II intron, particularly the first nucleotide of the 5’ end region of the Group I intron, and thus forms a P1 structure. In embodiments, the ribozyme recognition sequence II is connected with the 3’ end of the Group I / II intron. In embodiments, the ribozyme recognition sequence I is located at the 3’ end of the exon II. In embodiments, the ribozyme recognition sequence II is located at the 5’ end of the exon I.
[0243] During the circularization of the circular RNA precursor, the Group I / II intron and the sequence upstream of its 5’ end (if present) , and the sequence downstream of the 3’ end of the exon II (if present) are excised, and the 5’ end of the exon I and the 3’ end of the exon II are covalently linked to achieve circularization of the RNA. During the circularization of the circular RNA precursor, the ribozyme recognition sequence I is capable of connecting with ribozyme recognition sequence II to form the ribozyme recognition sequence. That is, during the circularization of the circular RNA precursor, the cleavage occurs between the 5’ splice site of the Group I intron and the 3’ end of ribozyme recognition sequence I and between the 3’ splice site of the Group I intron and the 5’ end of ribozyme recognition sequence II.
[0244] As mentioned above, “exon I” or “exon II” is a sequence derived from the native exon of the self-splicing intron (the exon flanking the self-splicing intron) and capable of being recognized and / or spliced by the self-splicing intron (the Group I / II intron) , and thus is required for circularization.
[0245] In aspects and embodiments of the present invention, the exon II comprises a 3’ end U group. In embodiments of the above aspects, the exon I and / or the exon II comprises modification compared to the native Group I / II exon. In embodiments of the above aspects, the exon I and / or the exon II does not comprise modification compared to the native Group I / II exon. In embodiments of the above aspects, the modification is not located on the sequence forming P1 structure compared to the native Group I / II exon.
[0246] In some embodiments, the 5’ end of the exon I is directly connected to the 3’ end of the Group I / II intron. In some embodiments, the 5’ end of the exon II is directly connected to the 3’ end of the sequence of interest or the linker B.
[0247] According to the present disclosure, the exon I and exon II each comprise an exon segment. In some embodiments, the exon I is derived from or contains a 3’ exon adjacent to 3’ end of the Group I / II intron (a native Group I / II intron) , and accordingly the exon II is derived from or contains a 5’ exon adjacent to 5’ end of the (same) Group I / II intron (a native Group I / II intron) (without any modification on the sequence forming P1 structure) . In some embodiments, the exon I, the exon II, and the Group I / II intron are derived from a same native Group I / II exon / intron (without any modification on the sequence forming P1 structure) . In some embodiments, the exon I, the exon II and the Group I / II intron in combination retain the activity of forming a P1 structure containing a ribozyme recognition sequence I and a P10 structure containing a ribozyme recognition sequence I.
[0248] In some embodiments, the exon I is derived from the native 3’ exon of the Group I / II intron (self-splicing intron) (the exon flanking (downstream of) the 3’ end of the self-splicing intron) or a contiguous fragment thereof starting from the 5’ terminal nucleotide. In some embodiments, the exon II is derived from the native 5’ exon of the Group I / II intron (self-splicing intron) (the exon flanking (downstream of) the 5’ end of the self-splicing intron) or a contiguous fragment thereof starting from the 3’ terminal nucleotide. In some embodiments, the exon I is the native 3’ exon of the Group I / II intron. In some embodiments, the exon II is the native 5’ exon of the Group I / II intron. In some embodiments, the exon I and / or exon II are a self-splicing exon segment. In some embodiments, the exon I and / or exon II comprise in part or in whole a naturally occurring exon sequence from a virus, bacterium or eukaryote. In other embodiments, the self-splicing exon segment comprises in part or in whole a non-naturally occurring sequence.
[0249] In some embodiments, the exon I is the entire native 3’ exon of the Group I / II intron (self-splicing intron) , or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%sequence identity with the entire native 3’ exon of the Group I / II intron (self-splicing intron) , or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to the entire native 3’ exon of the Group I / II intron (self-splicing intron) , (without any modification on the sequence forming P1 structure) .
[0250] In some embodiments, the exon I is derived from the native 3’ exon of the Group I intron. In some embodiments, the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon. In some embodiments, the exon I has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%sequence identity with a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon. In some embodiments, the exon I has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to a continuous fragment starting from the 5’ terminal nucleotide of the native 3’ exon, (without any modification on the sequence forming P1 structure) .
[0251] In some embodiments, the contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon comprises or consists of at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%nucleotides of the native 3’ exon. In some embodiments, the contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon is about 1-50, 2-30, 3-20, 4-18, 5-15 nucleotides in length. In some embodiments, the contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon is at least 1 nucleotide in length, such as at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15 nucleotides, at least 20, at least 25, at least 50 or more nucleotides in length, (without any modification on the sequence forming P1 structure) . In some embodiments, the contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon is 1 nucleotide in length or up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 15, up to 20, up to 25, up to 50 nucleotides in length or up to the total length of the native 3’ exon, (without any modification on the sequence forming P1 structure) .
[0252] In some embodiments, the exon I is a contiguous sequence having least 75%identity (e.g., at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 100%identity) with the entire native 3’ exon of the Group I / II intron (self-splicing intron) as described herein, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to the entire native 3’ exon of the Group I / II intron (self-splicing intron) , including the 3’ nucleotide of the splice site dinucleotide, (without any modification on the sequence forming P1 structure) .
[0253] In some embodiments, where the Group I / II intron (self-splicing intron) is a Group I intron, the 3’ exon region at least comprises a sequence (such as a sequence of about 1 to about 20 nucleotides) at it 5’ terminus which can pair with the P1 region of the corresponding Group I intron to form a P10 duplex region.
[0254] It is believed that for self-splicing of Group I introns, the consecutive one or more nucleotides (such as at least about 1 to about 7 nucleotides) from the 5’ end of the native 3’ exon can pair with the P1 region to form a P10 duplex region, and thus plays an important role in self-splicing. Definitions of the P1 and P10 regions of Group I introns are known in the art and can be determined, for example, with reference to the following documents: Burke, J. M., et al., (1987) Structural conventions for group I introns; Stahley, R. M., et al (2006) RNA splicing: group I intron crystal structures reveal the basis of splice site selection and metal ion catalysis; and / or Woodson, A. S., (2005) Structure and assembly of group I introns.
[0255] In some embodiments, the exon II is the entire native 5’ exon of the Group I / II intron (self-splicing intron) , or has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%sequence identity with the entire native 5’ exon of the Group I intron, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to the entire native 5’ exon of the Group I / II intron (self-splicing intron) , (without any modification on the sequence forming P1 structure) .
[0256] In some embodiments, the exon II is derived from the native 5’ exon of the Group I intron. In some embodiments, the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon. In some embodiments, the exon II has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%sequence identity with a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon, (without any modification on the sequence forming P1 structure) . In some embodiments, the exon II has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to a continuous fragment starting from the 3’ terminal nucleotide of the native 5’ exon, (without any modification on the sequence forming P1 structure) .
[0257] In some embodiments, the contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon comprises or consists of at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%of the nucleotides of the native 5’ exon, (without any modification on the sequence forming P1 structure) . In some embodiments, the contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon is at least 1 nucleotide in length, such as at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15 nucleotides, at least 20, at least 25, at least 50 or more nucleotides in length. In some embodiments, the contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon is 1 nucleotide in length or up to 2, up to 3, up to 4, up to 5, up to 6, up to 7, up to 8, up to 9, up to 10, up to 15, up to 20, up to 25, up to 50 nucleotides in length or up to the total length of the native 5’ exon, (without any modification on the sequence forming P1 structure) .
[0258] In some embodiments, the exon II is a contiguous sequence at least 75%identity (e.g., at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 100%identity) with the entire native 5’ exon of the Group I / II intron (self-splicing intron) as described herein, or has 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotide substitutions, deletions or additions compared to the entire native 5’ exon of the Group I / II intron (self-splicing intron) , including the 5’ nucleotide of the splice site dinucleotide, (without any modification on the sequence forming P1 structure) .
[0259] In some embodiments, where the Group I / II intron (self-splicing intron) is a Group I intron, the 5’ exon region comprises a sequence (for example, a sequence of about 3 to about 8 consecutive nucleotides) at its 3’ terminus which can pair with the internal guide sequence (IGS) of the corresponding Group I intron to form a P1 double-stranded region.
[0260] It is believed that for self-splicing of Group I introns, about 3 to about 8 consecutive nucleotides from the 3’ end of the native 5’ exon can pair with the internal guide sequence (IGS) of the intron to form the P1 double-stranded region, thus playing an important role in self-splicing. Definitions of IGS and / or P1 region of Group I introns are known in the art and can be determined, for example, with reference to the following documents: Burke, J. M., et al., (1987) Structural conventions for group I introns; Stahley, R. M., et al (2006) RNA splicing: group I intron crystal structures reveal the basis of splice site selection and metal ion catalysis; and / or Woodson, A. S., (2005) Structure and assembly of group I introns.
[0261] In embodiments, the exon I is derived from any one selected from 3’ exon of Group I intron of Anabaena sp. YBS01, Staphylococcus phage Twort, Ncr. m. ND5, 1, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, Bfu. S1506, Mpl. L798, Pan. m. ND3, 1, Osp. S1199, Azoarcus olearius BH72, Scytalidium dimidiatum, Scytonema-hofmani tRNA fMet, and Agrobacterium-tumefaciens. In embodiments, the exon II is derived from any one selected from 5’ exon of Group I intron of Anabaena sp. YBS01, Staphylococcus phage Twort, Ncr. m. ND5, 1, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, Bfu. S1506, Mpl. L798, Pan. m. ND3, 1, Osp. S1199, Azoarcus olearius BH72, Scytalidium dimidiatum, Scytonema-hofmani tRNA fMet, and Agrobacterium-tumefaciens. In embodiments, the exon I comprises or consists of a nucleotide sequence selected from any one of SEQ ID NOs: 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, and 59, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NOs: 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, and 59 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from any one of SEQ ID NOs: 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, and 59. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of any one of SEQ ID NOs: 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, and 59. In embodiments, the exon II comprises or consists of a nucleotide sequence selected from any one of SEQ ID NOs: 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, and 60, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NOs: 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, and 60 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from any one of SEQ ID NOs: 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, and 60 (without any modification on the sequence forming P1 structure) . In some embodiments, the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of any one of SEQ ID NOs: 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, and 60 (without any modification on the sequence forming P1 structure) .
[0262] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Anabaena sp. YBS01, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Anabaena sp. YBS01, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 31, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 31 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 31; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 32, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 32 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 32. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 31; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 32. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 16 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 16, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 16.
[0263] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Staphylococcus phage Twort, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Staphylococcus phage Twort, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 33, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 33 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 33; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 34, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 34 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 34. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 33;and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 34. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 17 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 17, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 17.
[0264] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Ncr. m. ND5, 1, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Ncr. m. ND5, 1, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 35, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 35 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 35; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 36, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 36 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 36. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 35; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 36. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 18 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 18, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 18.
[0265] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Aaz. b. trnL, and the exon I and the exon II are derived from 3’ exon and 5’exon of Group I intron of Aaz. b. trnL, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 37, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 37 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 37; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 38, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 38 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 38. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 37; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 38. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 19 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 19, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 19.
[0266] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Kap. S516, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Kap. S516, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 39, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 39 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 39; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 40, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 40 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 40. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 39; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 40. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 20 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 20, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 20.
[0267] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Sce. mL2449, and the exon I and the exon II are derived from 3’ exon and 5’exon of Group I intron of Sce. mL2449, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 41, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 41 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 41; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 42, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 42 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 42. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 41; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 42. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 21 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 21, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 21.
[0268] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of RB3. v. nrdB, 1, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of RB3. v. nrdB, 1, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 43, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 43 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 43; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 44, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 44 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 44. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 43; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 44. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 22 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 22, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 22.
[0269] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Bfu. S1506, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Bfu. S1506, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 45, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 45 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 45; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 46, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 46 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 46. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 45; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 46. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 23 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 23, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 23.
[0270] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Mpl. L798, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Mpl. L798, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 47, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 47 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 47; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 48, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 48 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 48. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 47; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 48. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 24 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 24, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 24.
[0271] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Pan. m. ND3, 1, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Pan. m. ND3, 1, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 49, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 49 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 49; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 50, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 50 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 50. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 49; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 50. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 25 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 25, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 25.
[0272] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Osp. S1199, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Osp. S1199, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 51, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 51 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 51; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 52, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 52 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 52. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 51; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 52. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 26 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 26, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 26.
[0273] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Azoarcus olearius BH72, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Azoarcus olearius BH72, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 53, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 53 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 53; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 54, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 54 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 54. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 53; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 54. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 27 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 27, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 27.
[0274] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Scytalidium dimidiatum, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Scytalidium dimidiatum, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 55, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 55 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 55; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 56, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 56 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 56. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 55; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 56. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 28 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 28, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 28.
[0275] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Scytonema-hofmani tRNA fMet, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Scytonema-hofmani tRNA fMet, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 57, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 57 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 57; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 58, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 58 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 58. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 57; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 58. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 29 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 29, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 29.
[0276] In aspects and embodiments of the present invention, the Group I / II intron is or is derived from a Group I intron of Agrobacterium-tumefaciens, and the exon I and the exon II are derived from 3’ exon and 5’ exon of Group I intron of Agrobacterium-tumefaciens, respectively. In embodiments, the exon I comprises or consists of a nucleotide sequence of SEQ ID NO: 59, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 59 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 59; and the exon II comprises or consists of a nucleotide sequence of SEQ ID NO: 60, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 60 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 60. In some embodiments, the exon I is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 59; and the exon II is 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, or 15 nucleotides of SEQ ID NO: 60. In embodiments, the Group I / II intron comprises or consists of a sequence of SEQ ID NO: 30 or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NO: 30, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 30.
[0277] In some embodiments, the exon I and the exon II are configured to be capable of forming a stem-loop-like structure. In some embodiments, the loop of the stem-loop-like structure comprises the splicing junction.
[0278] In some embodiments, the presence of the stem-loop-like structure can be predicted and / or determined by the nucleotide sequences of the Group I / II intron, the exon I and the exon II involved in circularization. In some embodiments, the presence of a stem-loop-like structure can be predicted and / or determined from the nucleotide sequence by RNA structure prediction tools such as RNAfold (http: / / rna. tbi. univie. ac. at / cgi-bin / RNAWebSuite / RNAfold. cgi) or RNAstructure (https: / / rna. urmc. rochester. edu / RNAstructureWeb / index. html) .
[0279] In some embodiments, the exon I comprises the sequence structure of the following formula: 5’-first loop sequence-first pairing sequence-first non-pairing sequence-3’ ; and the exon II comprises the sequence structure of the following formula: 5’-second non-pairing sequence-second pairing sequence-second loop sequence-3’ , wherein the first non-pairing sequence or the second non-pairing sequence may be independently present or absent, and the first pairing sequence and the second pairing sequence can complementarily pair to each other to form the stem of the stem-loop-like structure, which also called as “an exon duplex” , wherein the first loop sequence of exon I and the second loop sequence of exon II can form the loop of the stem-loop-like structure, e.g., through self-splicing for circularization.
[0280] Typically, the sequences forming the loop of the stem-loop-like structure are derived from the 3’ exon (the exon I) and / or the 5’ exon region (the exon II) .
[0281] In some embodiments, where the Group I / II intron is a Group I intron, the first loop sequence of exon I comprises or consists of one or more nucleotides (for example, about 1 to about 20 nucleotides) which can pair with the P1 region of the corresponding Group I intron (or the structure formed by the Group I / II intron) to form a P10 duplex region during the circularization.
[0282] In some embodiments, the first loop sequence of exon I may comprise or consist of a nucleotide sequence of (N) n, wherein N represents any nucleotide (A, G, U, or C) , n represents an integer from 1-20, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In some specific embodiments, n is 2, 4, or 5.
[0283] In some embodiments, the first loop sequence of exon I comprises or consists of the about 1 to about 7 consecutive nucleotides starting from the 5’ terminal nucleotide of the native 3’ exon of the Group I intron.
[0284] In some embodiments, the first loop sequence of exon I for example comprises or consists of AAAA.
[0285] In some embodiments, where the Group I / II intron is a Group I intron, the second loop sequence of exon II comprises or consists of one or more nucleotides (about 3 to about 8 nucleotides) which can pair with the internal guide sequence (IGS) of the corresponding Group I intron (or the structure formed by the Group I / II intron) to form a P1 duplex region during the circularization.
[0286] In some embodiments, the second loop sequence of exon II comprises or consists of the about 3 to about 8 consecutive nucleotides starting from the 3’ terminal nucleotide of the native 5’ exon of the Group I intron.
[0287] In some embodiments, the second loop sequence of exon II for example comprises or consists of CUU.
[0288] In some specific embodiments, a loop with a sequence of CUUAAAA can be formed after circularization.
[0289] In some embodiments, the first loop sequence of exon I comprises or consists of AAAA and the second loop sequence of exon II comprises or consists of CUU. In some specific embodiments, a loop with a sequence of CUUAAAA is formed after circularization.
[0290] The pairing sequences forming the stem of the stem-loop-like structure may be derived from the exon regions, however, it may also be derived from the spacer sequences. Alternatively, the pairing sequence may be derived from an exon region and a linker sequence, i.e., the pairing sequence comprises at least a portion of an exon region and at least a portion of the linker.
[0291] Without being bound by any theory, the RNA circularization efficiency based on intron self-splicing (e.g., Group I intron self-splicing) is related to the number of base pairs or the type or composition of base pairs in the stem portion of the stem-loop-like structure formed by the residual circularizing element. The stability of the stem-loop-like structure (e.g., as can be predicted from calculated free energies) may affect the circularization efficiency.
[0292] In some embodiments, the stem portion of the stem-loop-like structure (i.e. exon duplex) comprises at least 2 base pairs, such as at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15 or more base pairs, such as 2-30 base pairs, such as 2-25 base pairs, such as 2-20 base pairs, such as 2-15 base pairs, such as 2-10 base pairs, such as 5-30 base pairs, such as 5-25 base pairs, such as 5-20 base pairs, such as 5-15 base pairs, such as 5-10 base pairs, preferably consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 2-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 3-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 4-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 5-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 6-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 7-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 8-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 9-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 10-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 11-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 12-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 13-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 14-15 or more consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs, preferably consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 5 base pairs, preferably consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 6 base pairs, preferably consecutive matched base pairs. In some embodiments, the stem portion of the stem-loop-like structure comprises 7 base pairs, preferably consecutive matched base pairs.
[0293] In some embodiments, the stem portion in the stem-loop-like structure comprises up to 2 base mismatches, or up to 1 base mismatch, preferably, the stem portion comprises no base mismatches.
[0294] In some embodiments, the stem portion of the stem-loop-like structure comprises less than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%GC content. In some embodiments, the stem portion of the stem-loop-like structure comprises about 10%-60%, 20-50%, or 30-40%GC content.
[0295] In aspects and embodiments of the present invention, the sequence of interest comprises an ORF sequence. In aspects and embodiments of the present invention, the sequence of interest comprises a translation initiation element (TIE) and an ORF sequence. As used herein, the term “ORF” comprises a protein coding sequence or non-coding sequence. In embodiments, a sequence of interest comprises one or more ORF sequences. In embodiments, the sequence of interest comprises one or more TIE. In embodiments, the TIE is operably linked to the ORF. Here, “operably linked” refers to that the TIE such as IRES can mediate translation of the encoded protein. In embodiments, the TIE is located at the 5’ end of the ORF. In embodiments, the TIE is located at the 3’ end of the ORF. In embodiments, the ORF at the 5’ end of the sequence of interest comprises the 3’ exon or the linker A sequence at the 5’ end of the ORF. In embodiments, the ORF at the 3’ end of the sequence of interest comprises the 5’ exon or the linker B sequence at the 3’ end of the ORF. In embodiments, the TIE is selected from an IRES sequence, 5’ UTR sequence, Kozak sequence, a sequence containing m6A modification, a sequence complementary to ribosome 18S rRNA, or any combination thereof. In embodiments, the TIE is an Internal Ribosome Entry Site (IRES) .
[0296] “Translation initiation element (TIE) ” as used herein refers to a sequence in the RNA molecule that can initiate the translation of said RNA molecule.
[0297] In embodiments, the sequence of interest comprises one or more non-TIE functional element. As used herein, the term “non-TIE functional element” refers to a functional element that is not involved in translation initiation, such as a nucleotide sequence for regulating RNA translation, RNA circularization or RNA transcription. In embodiments, the non-TIE functional element is a 3’ UTR or a replicon.
[0298] In embodiments, the ORF encodes a protein. In embodiments, the ORF is a non-encoding sequence. For example, the non-encoding sequence can be antisense RNA, aptamer, guide RNA, or non-encoding RNA existing in any organism, and the like. The non-encoding sequence may or may not contain a specific secondary structure.
[0299] In embodiments, the ORF is a protein-coding sequence. Protein-coding sequences can encode proteins of eukaryotic, prokaryotic or viral origin. In certain embodiments, the protein can be any protein for therapeutic or diagnostic use. For example, the protein coding region can encode human proteins, antigens, antibodies, gene editing enzymes such as CRISPR nucleases, and the like. For example, the encoded protein can be a chimeric antigen receptor, an immunomodulatory protein, and / or a transcription factor, and the like. Some specific examples include, but are not limited to, EGF, FGF1, RBD, G6PC, PAH, HGF, and the like.
[0300] In embodiments, the ORF is a non-coding sequence. For example, the non-coding sequence can be an antisense RNA, an aptamer, a guide RNA, or non-protein-coding RNA existing in any organism, and the like. The non-coding sequence may or may not contain a specific secondary structure.
[0301] In some embodiments, the ORF is at least 10, 20, 40, 60, 80, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 10000, 20000 nucleotides in length. In some embodiments, the ORF is about 10-2000 nucleotides in length.
[0302] For example, the ORF encodes POLR2A (polymerase II subunit RPB1) , Luciferase, Spike, COL3A1 (collagen type III alpha 1 chain) , BMP2 (bone morphogenetic protein-2) , OTC (Ornithine Transcarbamylase) , EGFP (Enhanced Green Fluorescent Protein) , AQP1 (aquaporin-1) or SOD2 (Superoxide dismutase 2) protein or a protein having the same activity as the naturally occurring POLR2A, Luciferase, Spike, COL3A1, BMP2, OTC, EGFP, AQP1 or SOD2 protein and derived from the naturally occurring POLR2A, Luciferase, Spike, COL3A1, BMP2, OTC, EGFP, AQP1 or SOD2 protein through substitution, deletion or addition of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in the amino acid sequence of the naturally occurring POLR2A, Luciferase, Spike, COL3A1, BMP2, OTC, EGFP, AQP1 or SOD2 protein.. For example, the ORF may comprise or consist of a sequence selected from any one of SEQ ID NOs: 239-245, or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with any one of SEQ ID NOs: 239-245, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from any one of SEQ ID NOs: 239-245.
[0303] In some embodiments, the IRES is from Taura syndrome virus, Tiiatoma virus, Theiler’s encephalomyelitis virus, Simian Virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, Reticuloendotheliosis virus, Human poliovirus 1, Plautia stall intestine virus, Kashmir bee virus, Human rhinovirus 2, Homalodisca coagulata virus-1, Human Immunodeficiency Virus type 1, Homalodisca coagulata virus-1, Himetobi P virus, Hepatitis C virus, Hepatitis A virus, Hepatitis GB virus , Foot and mouth disease virus, Human enterovirus 71, Equine rhinitis virus, Ectropis obliqua picoma-like virus, Encephalomyocarditis virus, Drosophila C Virus, Human coxsackievirus B3, Crucifer tobamovirus, Cricket paralysis virus, Bovine viral diarrhea virus 1, Black Queen Cell Virus, Aphid lethal paralysis virus, Avian encephalomyelitis virus, Acute bee paralysis virus, Hibiscus chlorotic ringspot virus, Classical swine fever virus, Human FGF2, Human SFTPA1, Human AML1 / RUNX1, Drosophila antennapedia, Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAPl, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIFl alpha, Human n. myc, Mouse Gtx, Human p27kipl, Human PDGF2 / c-sis, Human p53, Human Pim-1, Mouse Rbm3, Drosophila reaper, Canine Scamper, Drosophila Ubx, Human UNR, Mouse UtrA, Human VEGF-A, Human XIAP, Drosophila hairless, S. cerevisiae TFIID, S. cerevisiae YAP1, tobacco etch virus, turnip crinkle virus, EMCV-A, EMCV-B, EMCV-Bf, EMCV-Cf, EMCV pEC9, Picobirnavirus, HCV QC64, Human Cosavirus E / D, Human Cosavirus F, Human Cosavirus JMY, Rhinovirus NAT001, HRV14, HRV89, HRVC-02, HRV-A21, Salivirus A SHI, Salivirus FHB, Salivirus NG-J1, Human Parechovirus 1, Crohivirus B, Yc-3, Rosavirus M-7, Shanbavirus A, Pasivirus A, Pasivirus A 2, Echovirus E14, Human Parechovirus 5, Aichi Virus, Hepatitis A Virus HA 16, Phopivirus, CVA10, Enterovirus C, Enterovirus D, Enterovirus J, Human Pegivirus 2, GBV-C GT110, GBV-C K1737, GBV-C Iowa, Pegivirus A 1220, Pasivirus A 3, Sapelovirus, Rosavirus B, Bakunsa Virus, Tremovirus A, Swine Pasivirus 1, PLV-CHN, Pasivirus A, Sicinivirus, Hepacivirus K, Hepacivirus A, BVDV1, Border Disease Virus, BVDV2, CSFV-PK15C, SF573 Dicistravirus, Hubei Picoma-like Virus, CRPV, Salivirus A BN5, Salivirus A BN2, Salivirus A 02394, Salivirus A GUT, Salivirus A CH, Salivirus A SZ1, Salivirus FHB, CVB3, CVB1, Echovirus 7, CVB5, EVA71, CVA3, CVA12, EV24 or an aptamer to eIF4G.
[0304] In some embodiments, the IRES comprises a SalivirusNG-J1_del34 IRES or a fragment or variant thereof. In certain embodiments, the IRES comprises a sequence according to SEQ ID NO: 250. In some embodiments, the IRES comprises a CVB3 IRES or a fragment or variant thereof, or the IRES comprises a nucleotide sequence according to SEQ ID NO: 251 or comprises a nucleotide sequence having at least 75%, e.g., at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 251. In some embodiments, the IRES comprises an Enterovirus A121 IRES or a fragment or variant thereof. In certain embodiments, the IRES comprises a sequence according to SEQ ID NO: 252 or comprise a nucleotide sequence having at least 75%, e.g., at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to SEQ ID NO: 252. In some embodiments, the IRES comprises a Canine-kobuvirus IRES or a fragment or variant thereof. In certain embodiments, the IRES comprises a sequence according to SEQ ID NO: 253 or comprise a nucleotide sequence having at least 75%, e.g., at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100% sequence identity to SEQ ID NO: 253.
[0305] In embodiments, the sequence of interest may comprise a spacer C. In embodiments, the spacer C is located between the TIE and the ORF. In some embodiments, the spacer C is no more than 50, no more than 40, no more than 30, no more than 20, no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, no more than 5, no more than 4, or no more than 3 nucleotides in length. In some embodiments, the spacer C is 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2 nucleotides in length. In some embodiments, the spacer C comprises or consists of a sequence selected from any one of SEQ ID NOs: 237-238.
[0306] In aspects of the present invention, the linker I sequence and the linker II sequence are functional sequences which can improve circularization. For example, linker I and linker II are homology arm sequences capable of complementary pairing to each other to form a homology arm double-stranded region. In some embodiments, the linker I sequence is a homology arm I. In some embodiments, the linker II sequence is a homology arm II.
[0307] As used herein, the terms “linker I (or linker I sequence) ” and “homology arm I” have the same meaning and refer to the same thing, and thus they can be interchangeably used in the context of all aspects of the present invention described herein. As used herein, the terms “linker II (or linker II sequence) ” and “homology arm II” have the same meaning and refer to the same thing, and thus they can be interchangeably used in the context of all aspects of the present invention described herein. The linker I (homology arm I) and the linker II (homology arm II) may be of the same or different length, as long as they are complementary to each other and together are capable of forming a double-stranded region. In some embodiments of the present invention, in the double-stranded region formed by the linker I (homology arm I) and the linker II (homology arm II) described herein, the 3’ end nucleotide of the linker I (homology arm I) forms a base pair with the 5’ end nucleotide of the linker II (homology arm II) . In embodiments, one base pair in the double-stranded region is formed by the 3’ end nucleotide of the linker I and the 5’ end nucleotide of the linker II.
[0308] In some embodiments, linker I is upstream of the self-splicing intron and linker II is downstream of the 5’ exon. In some embodiments, the linker I sequence is located at the 5’ end of the circular RNA precursor. In some embodiments, the linker II sequence is located at the 3’ end of the circular RNA precursor. In some embodiments, the linker I sequence is adjacent or very close to the group I / II intron.
[0309] In certain embodiments, the linker I sequence and the linker II sequence are completely or partially complementary sequences. Thus, in certain embodiments, at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%or 100%of the linker I sequence and the linker II sequence may be base paired with one another.
[0310] In some embodiments, the predicted mean minimum free energy (MFE) of the linker I sequence and the linker II sequence (the homology arms) is more than about-2.0 kal / mol, more than about-1.5 kal / mol, more than about -1.4 kal / mol, more than about -1.3 kal / mol, more than about -1.2 kal / mol, more than about -1.1 kal / mol, more than about -1.0 kal / mol, more than about -0.9 kal / mol, or more than about -0.8 kal / mol. In some embodiments, the predicted mean minimum free energy (MFE) of the linker I sequence and the linker II sequence (the homology arms) is from about -2.0 kal / mol to about -0.8 kal / mol. In some embodiments, the predicted mean minimum free energy (MFE) of the linker I sequence and the linker II sequence (the homology arms) is from about -1.5 kal / mol to about -1.0 kal / mol. In some embodiments, the predicted mean minimum free energy (MFE) of the linker I sequence and the linker II sequence (the homology arms) is from about -1.2 kal / mol to about -1.0 kal / mol. In some embodiments, the predicted mean minimum free energy (MFE) of the linker I sequence and the linker II sequence (the homology arms) is from about -1.1 kal / mol to about -1.0 kal / mol. The minimum free energy can be determined, for example, by RNAfold (http: / / rna. tbi. univie. ac. at / cgi-bin / RNAWebSuite / RNAfold. cgi) or RNAstructure (https: / / rna. urmc. rochester. edu / RNAstructureWeb / index. html) Structure Prediction Tool.
[0311] In some embodiments, linker I and / or linker II is about 4 to 400 nucleotides in length (e.g., 50-300 nucleotides in length, 100-200 nucleotides in length, 150-200 nucleotides in length) , about 5-200 nucleotides in length, about 5-150 nucleotides in length, about 5-100 nucleotides in length, about 5-80 nucleotides in length, about 5-50 nucleotides in length, about 5-40 nucleotides in length, about 5-30 nucleotides in length, about 5-20 nucleotides in length, about 5-10 nucleotides in length, preferably, about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250 or 300 nucleotides in length. In some embodiments, linker I and / or linker II is less than 400, less than 350, less than 300, less than 250, less than 200, less than 150, less than 100, less than 50, less than 40, less than 30 nucleotides in length.
[0312] In some embodiments, the linker I and the linker II are homology arm sequences capable of complementary pairing to each other to form a double-stranded region. In some embodiments, double-stranded region is about 4 to 400 base pairs (e.g., 50-300 base pairs, 100-200 base pairs, 150-200 base pairs) , about 5-200 base pairs, about 5-150 base pairs, about 5-100 base pairs, about 5-80 base pairs, about 5-50 base pairs, about 5-40 base pairs, about 5-30 base pairs, about 5-20 base pairs, about 5-10 base pairs, preferably, about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250 or 300 base pairs, preferably consecutive matched base pairs. In some embodiments, the double-stranded region is less than 400, less than 350, less than 300, less than 250, less than 200, less than 150, less than 100, less than 50, less than 40, less than 30 base pairs, preferably consecutive matched base pairs.
[0313] In certain embodiments, the linker I comprises a nucleotide sequence according to SEQ ID NO:227, 229, 231, or 233, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 227, 229, 231, or 233 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 227, 229, 231, or 233. In certain embodiments, the linker II comprises a nucleotide sequence according to SEQ ID NO: 228, 230, 232, or 234, a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 228, 230, 232, or 234 or a nucleotide sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 228, 230, 232, or 234. In certain embodiments, the linker I is 30 nucleotides, 40 nucleotides, 50 nucleotides, 100 nucleotides, 150 nucleotides, 200 nucleotides, 250 nucleotides, 300 nucleotides of SEQ ID NO: 227, 229, 231, or 233. In certain embodiments, the linker II is 30 nucleotides, 40 nucleotides, 50 nucleotides, 100 nucleotides, 150 nucleotides, 200 nucleotides, 250 nucleotides, 300 nucleotides of SEQ ID NO: 228, 230, 232, or 234. In certain embodiments, the linker I comprises a nucleotide sequence according to SEQ ID NO: 227, and the linker II comprises a nucleotide sequence according to SEQ ID NO: 228. In certain embodiments, the linker I comprises a nucleotide sequence according to SEQ ID NO: 229, and the linker II comprises a nucleotide sequence according to SEQ ID NO: 230. In certain embodiments, the linker I comprises a nucleotide sequence according to SEQ ID NO: 231, and the linker II comprises a nucleotide sequence according to SEQ ID NO: 232. In certain embodiments, the linker I comprises a nucleotide sequence according to SEQ ID NO: 233, and the linker II comprises a nucleotide sequence according to SEQ ID NO: 234.
[0314] In aspects of the present invention, linker A and linker B are any contiguous nucleotide sequence that at least does not negatively interfere with the function of the elements it connects. In aspects of the present invention, the linker A is located 3’ to the exon I. In aspects of the present invention, the linker B is located 5’ to the exon II. In aspects of the present invention, the linker A is located between the exon I and the sequence of interest. In aspects of the present invention, the linker B is located between the exon II and the sequence of interest. In aspects of the present invention, linker A and linker B comprise homology arms and / or spacers. In embodiments, linker A and linker B are homology arm sequences capable of complementary pairing to each other to form a homology arm double-stranded region. In embodiments, linker A and linker B are spacers. In embodiments, linker A and linker B both comprise a homology arm sequence and a spacer. In embodiments, the sequences of said linker A and B are identical. In embodiments, the sequences of said linker A and B are different. In embodiments, said linker A and B are of identical length. In embodiments, the sequences of said linker A and B are of different lengths. In some embodiments, linker A or linker B is about 4 to 400 nucleotides in length (e.g., 50-300 nucleotides in length, 100-200 nucleotides in length, 150-200 nucleotides in length) , about 5-200 nucleotides in length, about 5-150 nucleotides in length, about 5-100 nucleotides in length, about 5-80 nucleotides in length, about 5-50 nucleotides in length, about 5-40 nucleotides in length, about 5-30 nucleotides in length, about 5-20 nucleotides in length, about 5-10 nucleotides in length, preferably, about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250 or 300 nucleotides in length.
[0315] In certain embodiments, the linker A and the linker B are completely or partially complementary sequences and together are capable of forming a duplex. In certain embodiments, the linker A and the linker B are non-complementary sequences. In certain embodiments, about 60%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%of the linker A and / or the linker B may be base paired with one another. In certain embodiments, about 20 nucleotides, about 16 nucleotides, about 15 nucleotides, about 14 nucleotides, about 13 nucleotides, about 12 nucleotides, about 11 nucleotides, about 10 nucleotides, about 9 nucleotides, about 8 nucleotides, about 7 nucleotides, about 6 nucleotides, about 5 nucleotides, about 4 nucleotides, about 3 nucleotides, about 2 nucleotides, about 1 nucleotides of the linker A and / or the linker B may be base paired with one another. In certain embodiments, about 1-20 nucleotides, about 2-15 nucleotides, about 3-14 nucleotides, about 4-13 nucleotides, about 5-12 nucleotides, about 6-11 nucleotides, about 7-10 nucleotides, about 8-9 nucleotides of the linker A and / or the linker B may be base paired with one another. In certain embodiments, the linker A and the linker B are predicted to form a duplex of about 20 base pairs, about 16 base pairs, about 15 base pairs, about 14 base pairs, about 13 base pairs, about 12 base pairs, about 11 base pairs, about 10 base pairs, about 9 base pairs, about 8 base pairs, about 7 base pairs, about 6 base pairs, about 5 base pairs, about 4 base pairs, about 3 base pairs, about 2 base pairs, about 1 base pair in length. In certain embodiments, the linker A and the linker B are predicted to form a duplex of about 1-20 base pairs, about 2-15 base pairs, about 3-14 base pairs, about 4-13 base pairs, about 5-12 base pairs, about 6-11 base pairs, about 7-10 base pairs, about 8-9 base pairs in length. In certain embodiments, the duplex is formed by the 3’ end of the linker A and 5’ end of the linker B. In certain embodiments, the duplex is formed by the 5’ end of the linker A and 3’ end of the linker B.
[0316] In various embodiments, the linker A and the linker B are not predicted to form a duplex of more than 8 base pairs in length with any sequences within 250 nucleotides in either direction. In some embodiments, the linker A and the linker B are not predicted to form a duplex of more than 8 base pairs in length with any sequences within 1000 nucleotides in either direction.
[0317] In some embodiments, the linker A may be absent. In some embodiments, the linker B may be absent. In some embodiments, the linker A and linker B may be absent. In certain embodiments, the linker A comprises a sequence according to SEQ ID NO: 235. In certain embodiments, the linker B comprises a sequence according to SEQ ID NO: 236.
[0318] As used herein, “exon I” , “exon II” , “exon I / II” can be used interchangeably with “exon 1” , “exon 2” or “exon 1 / 2” . As used herein, “exon I” , “exon 1” , and “3’ exon” can be used interchangeably. As used herein, “exon II” , “exon 2” , and “5’ exon” can be used interchangeably. As used herein, “Group I / II intron” and “self-splicing intron” can also be used interchangeably with “ribozyme” . As used herein, the “target of interest” or “TOI” can be used interchangeably with the “sequence of interest” .
[0319] As used herein, "spacer" refers to any contiguous nucleotide sequence that at least does not negatively interfere with the function of the elements it connects. Generally, if it is desired to avoid the interaction of two near or adjacent elements, a spacer can be inserted between the two elements. The spacer sequences described herein can serve two functions: (1) to facilitate circularization and (2) to facilitate functionality by allowing correct folding of the residual circularizing element and the nucleotide sequence of interest (e.g., IRES) . In some embodiments, the spacer is no more than 150, no more than 100, no more than 50, no more than 30, no more than 10, no more than 5, or no more than 3 nucleotides in length. In some embodiments, the spacer is 5 nucleotides in length. In some embodiments, the spacer is 4 nucleotides in length. In some embodiments, the spacer is 3 nucleotides in length.
[0320] In all aspects of the present invention described herein, the spacer A and spacer B are non-complementary to each other. In embodiments of the above aspects, spacer A and spacer B do not form any base pair between each other and thus together are capable of forming a non-pairing region. Accordingly, in some embodiments of the present invention described herein, the double-stranded region formed by the linker I and the linker II and the non-pairing region formed by the spacer A and spacer B together are capable of forming a stem-loop-like structure.
[0321] In embodiments, said spacer A and B have same length. In embodiments, said spacer A and B have different length. In embodiments, the sequences of said spacer A and B are identical. In embodiments, the sequences of said spacer A and B are different. In some embodiments, the spacer A is located at the 3’ end of the linker A. In some embodiments, the spacer B is located at the 5’ end of the linker B. In some embodiments, the spacer A and the spacer B are each 2 to 150 nucleotides in length (e.g., 4-50 nucleotides in length, 6-30 nucleotides in length, 8-20 nucleotides in length, 10-15 nucleotides in length, 2-15 nucleotides in length, 4-15 nucleotides in length) . In some embodiments, the spacer A and the spacer B are each at least 2, 4, 6, 8, 10, 15, 20, 30, 40, 50, 70, 90, 100, 130, or 150 nucleotides in length. In some embodiments, the spacer A and the spacer B are each about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides in length. In some embodiments, the spacer A and the spacer B are each no more than 100, 90, 80, 70, 60, 50, 45, 40, 35, or 30 nucleotides in length. In some embodiments, the spacer A and the spacer B are each between 4 and 50, 10 and 50, 20 and 50, 20 and 40, and / or 25 and 35 nucleotides in length. In certain embodiments, the spacer A and the spacer B are each 4, 6, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides in length. In some embodiments, the spacer sequence is at least 4 nucleotides in length, and / or about 4 to about 60 nucleotides in length.
[0322] The inventors have surprisingly found that when self-splicing introns, self-slicing exons, linker I and II are used for RNA circularization, the introduction of an additional spacer A and B (and linker A and B) into the circular RNA precursor may result in increased circularizing efficiency.
[0323] In some embodiments, the precursor is in linear form and is capable of undergoing autocatalytic circularization which generates a circular RNA. In some embodiments, the circularizing efficiency of the circular RNA precursor is at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 200%, at least about 300%, at least about 400%, at least about 500%or more. In some embodiments, the circularizing efficiency of the circular RNA precursor is increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 200%, at least about 300%, at least about 400%, at least about 500%or more as compared to a control circular RNA precursor, for example, a corresponding circular RNA precursor without the spacer A and B (and linker A and B) .
[0324] In some embodiments, the spacer A and the spacer B do not interfere, inhibit or disrupt the formation of the functional P1 structure formed by the exon II and the 5’ end of the intron. In some embodiments, the spacer A and / or the spacer B comprises polyA, polyU, polyAC, polyC or polyG. In some embodiments, the spacer A and / or the spacer B is a polyA sequence. In some embodiments, the spacer A and / or the spacer B is a polyAC sequence. In some embodiments, the spacer A and / or the spacer B are unstructured sequences or non-interfering P1 sequences.
[0325] As used herein, “structured” with regard to RNA refers to an RNA sequence that is predicted by the RNAFold software or similar predictive tools to form a structure (e.g., a hairpin loop) with itself or other sequences in the same RNA molecule. As used herein, “unstructured” with regard to RNA refers to an RNA sequence that is not predicted by RNA structure predictive tools to form a structure (e.g., a hairpin loop) with itself or other sequences in the same RNA molecule. In some embodiments, unstructured RNA can be functionally characterized using nuclease protection assays.
[0326] In some embodiments, the spacer A and / or the spacer B comprises or consists of a sequence selected from any one of SEQ ID NOs: 209-222. In some embodiments, the spacer A comprises or consists of a sequence selected from any one of SEQ ID NOs: 209-215, 217, 219, and 221. In some embodiments, the spacer A comprises or consists of a sequence selected from any one of SEQ ID NOs: 209-214, 216, 218, 220, and 222. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 209. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 210. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 211. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 212. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 213. In some embodiments, the spacer A and the spacer B comprises or consists of a sequence of SEQ ID NO: 214. In some embodiments, the spacer A comprises or consists of a sequence of SEQ ID NO: 215, and the spacer B comprises or consists of a sequence of SEQ ID NO: 216. In some embodiments, the spacer A comprises or consists of a sequence of SEQ ID NO: 217, and the spacer B comprises or consists of a sequence of SEQ ID NO: 218. In some embodiments, the spacer A comprises or consists of a sequence of SEQ ID NO: 219, and the spacer B comprises or consists of a sequence of SEQ ID NO: 220. In some embodiments, the spacer A comprises or consists of a sequence of SEQ ID NO: 221, and the spacer B comprises or consists of a sequence of SEQ ID NO: 222.
[0327] In embodiments for “seamless ICI” system, the circular RNA precursor comprising in the 5’ to 3’ direction:
[0328] a) a linker I,
[0329] b) a spacer A,
[0330] c) a Group I / II intron,
[0331] d) a TOI sequence,
[0332] e) a spacer B, and
[0333] f) a linker II,
[0334] wherein the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region,
[0335] wherein the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region,
[0336] wherein the 5’ end region of the Group I / II intron is capable of forming the loop and a part of the stem of the P1 structure, the 3’ end region of the TOI sequence is capable of forming the remaining part of the stem of the P1 structure, and the 5’ terminal nucleotide of the Group I / II intron and the 3’ terminal nucleotide of TOI sequence together form the splice site in said P1 structure. In embodiments, the 3’ end region of the TOI sequence and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and wherein the 5’ end region of the TOI and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II.
[0337] In embodiments, the sequence of interest comprises a first fragment of the sequence of interest at its 3’ end and a second fragment of the sequence of interest at its 5’ end,
[0338] In embodiments, the first fragment of the sequence of interest and the second fragment of the sequence of interest are respectively derived from a 5’ terminal portion and a 3’ terminal portion of an ORF sequence, a TIE or a non-TIE functional element,
[0339] In embodiments, the first fragment of the sequence of interest comprises a ribozyme recognition sequence I located at its 3’ end, the second fragment of the sequence of interest comprises a ribozyme recognition sequence II located at its 5’ end,
[0340] In embodiments, in the circular RNA generated upon self-cleavage and circularization of the circular RNA precursor, the 3’ end of the ribozyme recognition sequence I is connected with the 5’ end of the ribozyme recognition sequence II,
[0341] In embodiments, the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region,
[0342] In embodiments, the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region,
[0343] In embodiments, the 5’ end region of the Group I / II intron is capable of forming the loop and a part of the stem of the P1 structure, wherein the 3’ end region of the ribozyme recognition sequence I is capable of forming the remaining part of the stem of the P1 structure with the 5’ end region of the Group I / II intron, and
[0344] In embodiments, the 5’ end region of the ribozyme recognition sequence I and the 3’ end region of the ribozyme recognition sequence II are partially complementary to each other and together are capable of forming an exon duplex or a first stem-loop-like structure.
[0345] In embodiments, in the circular RNA generated upon self-cleavage and circularization of the circular RNA precursor, the 3’ end of the first fragment of the sequence of interest is connected with the 5’ end of the second fragment of the sequence of interest to form the complete ORF sequence, TIE or non-TIE functional element.
[0346] The linker I, linker II, spacer A, spacer B, Group I intron, ORF, TIE, non-TIE functional element, ribozyme recognition sequence I, ribozyme recognition sequence II may have the definitions mentioned above.
[0347] In embodiments, a ribozyme recognition sequence I is located at the 3’ end of the sequence of interest, and a ribozyme recognition sequence II is located at the 5’ end of the sequence of interest.
[0348] For the seamless ICI system, “ribozyme recognition sequence II” can be used interchangeably with the term “exon I” defined herein, while “ribozyme recognition sequence I” can be used interchangeably with the term “exon II” defined herein.
[0349] For the seamless ICI system, the sequence of interest is scanned for sequences that are homologous to ribozyme recognition sequence II and / or ribozyme recognition sequence I, thereby allowing splicing to occur without introducing any exon sequence heterogenous to the sequence of interest in the produced circular RNA.
[0350] In embodiments, the ORF in the sequence of interest encodes AQP1 or SOD2 protein or a protein having the same activity as the naturally occurring AQP1 or SOD2 protein and derived from the naturally occurring AQP1 or SOD2 protein through substitution, deletion or addition of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in the amino acid sequence of the naturally occurring AQP1 or SOD2 protein.
[0351] In embodiments, the nucleotide sequence of the first fragment of the sequence of interest SEQ ID NO: 247, or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 247, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 247. In embodiments, the nucleotide sequence of the second fragment of the sequence of interest SEQ ID NO: 246, or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 246, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 246. In embodiments, the nucleotide sequence of the first fragment of the sequence of interest SEQ ID NO: 249, or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 249, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 249. In embodiments, the nucleotide sequence of the second fragment of the sequence of interest SEQ ID NO: 248, or a sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity with SEQ ID NO: 248, or a sequence having less than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides different from SEQ ID NO: 248.
[0352] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising:
[0353] a) a Group I / II intron,
[0354] b) a 3’ exon,
[0355] c) a sequence of interest, and
[0356] d) a 5’ exon.
[0357] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0358] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0359] a) a Group I / II intron,
[0360] b) a 3’ exon,
[0361] c) a sequence of interest, and
[0362] d) 5’ exon.
[0363] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0364] In another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0365] a) a Group I / II intron,
[0366] b) a 3’ exon,
[0367] c) a linker A,
[0368] d) a sequence of interest,
[0369] e) a linker B, and
[0370] f) a 5’ exon.
[0371] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0372] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0373] a) a linker I,
[0374] b) a Group I / II intron,
[0375] c) a 3’ exon,
[0376] d) a sequence of interest,
[0377] e) a 5’ exon, and
[0378] f) a linker II.
[0379] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0380] In yet another aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0381] a) a linker I,
[0382] b) a Group I / II intron,
[0383] c) a 3’ exon,
[0384] d) a linker A,
[0385] e) a sequence of interest,
[0386] f) a linker B,
[0387] g) a 5’ exon, and
[0388] h) a linker II.
[0389] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0390] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0391] a) a Group I / II intron, and
[0392] b) a sequence of interest, comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end.
[0393] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0394] a) a linker I,
[0395] b) a Group I / II intron,
[0396] c) a sequence of interest, comprising a 3’ exon at its 5’ end and a 5’ exon at its 3’ end, and
[0397] d) a linker II.
[0398] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0399] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0400] a) a linker I,
[0401] b) a Group I / II intron,
[0402] c) a sequence of interest, comprising: (i) an IRES fragment I which comprises a 3’ exon at its 5’ end, (ii) an ORF sequence, and (iii) an IRES fragment II comprising a 5’ exon at its 3’ end, and
[0403] d) a linker II.
[0404] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction:
[0405] a) a linker I,
[0406] b) a Group I / II intron,
[0407] c) a sequence of interest, comprising: (i) an ORF sequence fragment I which comprises a 3’ exon at its 5’ end, (ii) an IRES sequence, and (iii) an ORF sequence fragment II comprising a 5’ exon at its 3’ end, and
[0408] d) a linker II.
[0409] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0410] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0411] a) a homology arm I,
[0412] b) a spacer A
[0413] c) a Group I / II intron,
[0414] d) exon I,
[0415] e) a sequence of interest, comprising: (i) an ORF sequence, and (ii) an IRES sequence,
[0416] f) exon II,
[0417] g) a spacer B,
[0418] h) a homology arm II,
[0419] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0420] wherein said spacer A and B are non-complementary sequences.
[0421] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0422] In one aspect, it is an object of the present invention to provide a circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0423] a) a homology arm I,
[0424] b) a spacer A,
[0425] c) a Group I intron comprising a pairing G group and a 3’ end G group,
[0426] d) exon I,
[0427] e) a sequence of interest, comprising: (i) an ORF sequence, and (ii) an IRES sequence,
[0428] f) exon II comprising a 3’ end U group,
[0429] g) a spacer B,
[0430] h) a homology arm II,
[0431] wherein said homology arms I and II are complementary to each other and form a double stranded homology arm,
[0432] wherein said spacer A and B are non-complementary sequences, and
[0433] wherein said Group I intron forms a U-G base pair with the 3’ end U of exon II.
[0434] In embodiments of the above aspect, the sequence of interest comprises an IRES sequence, an ORF, and a spacer C located between the IRES sequence and the ORF sequence.
[0435] In embodiments of the above aspect, the circular RNA precursor of the invention allows generation of a circular RNA comprising the exon I, the sequence of interest, and the exon II through the self-splicing of the circular RNA precursor.
[0436] The circular RNA precursor may be (e.g., chemically) unmodified, partially modified or fully modified. In some embodiments, the circular RNA precursor comprises at least one nucleotide modification. In some embodiments, up to 100%of the nucleotides of the circular RNA precursor are modified. In some embodiments, the at least one nucleotide modification is a cytidine modification, an uridine modification, or an adenosine modification. In some embodiments, the at least one nucleoside modification is selected from the group consisting of 5-methylcytosine (m5C) , N6-methyladenosine (m6A) , pseudouridine (ψ) , N1-methylpseudouridine (m1ψ) and 5-methoxyuridine (5moU) . In some embodiments, the circular RNA precursor comprises less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 1%of a specific nucleotide modification. As used herein, the percentage of a particular nucleotide modification refers to the ratio of nucleotides in the sequence that have undergone that particular modification to nucleotides that can undergo that particular modification.
[0437] In some preferred embodiments, the circular RNA precursor is unmodified. In some embodiments, the circular RNA precursor does not contain nucleotide chemical modification.
[0438] In some preferred embodiments, the circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 61-147, preferably SEQ ID NOs: 127-132, or a nucleotide sequence with at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to any one of SEQ ID NOs: 61-147, preferably SEQ ID NOs: 127-132.
[0439] In another aspect, the present invention also provides a nucleic acid vector for generating a circular RNA molecule, said vector comprising a coding sequence of the circular RNA precursor of the invention. In another aspect, the present invention also provides a nucleic acid vector for generating a circular RNA molecule, comprising a DNA sequence encoding the circular RNA precursor of the invention.
[0440] As used herein, "vector" refers to a DNA derived from a virus, plasmid or cell of a higher organism into which a foreign DNA fragment can be or has been inserted for cloning and / or expression purposes. In certain embodiments, the vector can be stably maintained in the organism. The vector may contain, for example, an origin of replication, a selectable marker or a reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS) . The term includes linear DNA fragments (e.g., PCR products, linear plasmid fragments) , plasmid vectors, viral vectors, cosmids, bacterial artificial chromosomes (BACs) , yeast artificial chromosomes (YACs) , and the like.
[0441] In embodiments of the above aspect, the nucleic acid vector further comprises a promoter sequence operably linked to the coding sequence of the circular RNA precursor. The operably linked promoter allows in vivo and / or in vitro transcription of the circular RNA precursor. The promoter is, for example, a T7 RNA polymerase promoter, a T6 viral RNA polymerase promoter, a SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter or a T4 viral RNA polymerase promoter.
[0442] In embodiments of the above aspect, the nucleic acid vector further comprises an RNA polymerase.
[0443] In another aspect, the present invention also provides a circular RNA, which is prepared from the circular RNA precursor of the invention or the nucleic acid vector of the invention.
[0444] In another aspect, the present invention provides a circular RNA, which comprises an exon I, a sequence of interest, and an exon II.
[0445] In some embodiments, the exon I, the sequence of interest, and the exon II may have the definitions mentioned above.
[0446] In some embodiments, the exon I and the exon II (or the ribozyme recognition sequence I and the ribozyme recognition sequence II) are involved in or required for RNA circularization by a Group I / II intron (self-splicing intron) . The exon I and the exon II (or the ribozyme recognition sequence I and the ribozyme recognition sequence II) participate in circularization together with the Group I / II intron (self-splicing intron) but are retained in the final circular RNA.
[0447] In some embodiments, the exon I and the exon II (or the ribozyme recognition sequence I and the ribozyme recognition sequence II) are covalently linked. In some embodiments, 5’ end of the exon I is covalently linked to 3’ end of the exon II.
[0448] In some embodiments, the circular RNA is at least 10, 20, 40, 60, 80, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 10000, 20000 nucleotides in length. In some embodiments, the circular RNA is at least about 10 nucleotides in length. In some embodiments, the circular RNA is about 500 nt or less. In some embodiments, the circular RNA is at least about 1 knt.
[0449] The circular RNA may be unmodified, partially modified or fully modified. In some embodiments, the circular RNA comprises at least one nucleotide modification. In some embodiments, up to 100%of the nucleotides of the circular RNA are modified. In some embodiments, the at least one nucleotide modification is a cytidine modification, a uridine modification, or an adenosine modification. In some embodiments, the at least one nucleotide modification is selected from the group consisting of 5-methylcytosine (m5C) , N6-methyladenosine (m6A) , pseudouridine (ψ) , N1-methylpseudouridine (m1ψ) and 5-methoxyuridine (5moU) . In one embodiment, the circular RNA comprises less than 100%, less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 15%, less than 10%, less than 5%, less than 1%of a specific nucleotide modification. As used herein, the percentage of a particular nucleotide modification refers to the ratio of nucleotides in the sequence that have undergone that particular modification to nucleotides that can undergo that particular modification.
[0450] In some preferred embodiments, the circular RNA is unmodified. In some embodiments, the circular RNA does not contain nucleotide modification.
[0451] In some preferred embodiments, the circular RNA comprises a nucleotide sequence of any one of SEQ ID NOs: 148-208, preferably SEQ ID NOs: 188-193, or a nucleotide sequence with at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%or 100%sequence identity to any one of SEQ ID NOs: 148-208, preferably SEQ ID NOs: 188-193.
[0452] In particular embodiments, polynucleotides of the nucleic acid vector, the circular RNA precursor, or the circular RNA may be codon-optimized. A codon-optimized sequence may be one in which codons in a polynucleotide encoding a polypeptide have been substituted in order to increase the expression, stability and / or activity of the polypeptide. Factors that influence codon optimization include, but are not limited to one or more of: (i) variation of codon biases between two or more organisms or genes or synthetically constructed bias tables, (ii) variation in the degree of codon bias within an organism, gene, or set of genes, (iii) systematic variation of codons including context, (iv) variation of codons according to their decoding tRNAs, (v) variation of codons according to GC%, either overall or in one position of the triplet, (vi) variation in degree of similarity to a reference sequence for example a naturally occurring sequence, (vii) variation in the codon frequency cutoff, (viii) structural properties of mRNAs transcribed from the DNA sequence, (ix) prior knowledge about the function of the DNA sequences upon which design of the codon substitution set is to be based, and / or (x) systematic variation of codon sets for each amino acid. In some embodiments, a codon optimized polynucleotide may minimize ribozyme collisions and / or limit structural interference between the expression sequence and the core functional element. Codon optimization can be performed by methods known in the art.
[0453] In another aspect, the present invention provides a method for preparing a circular RNA, the method comprises:
[0454] 1) providing a circular RNA precursor of the invention or obtaining a circular RNA precursor by transcribing from the nucleic acid vector of the invention;
[0455] 2) incubating the circular RNA precursor under the conditions allowing autocatalytic circularization to generate a circular RNA; and
[0456] 3) harvesting the circular RNA obtained in step 2) .
[0457] In another aspect, the present invention provides a method for preparing a circular RNA molecule, comprising:
[0458] (a) providing a circular RNA precursor of the invention,
[0459] (b) allowing the RNA precursor to undergo intramolecular self-cleavage, thereby generating the circular RNA molecule, and
[0460] (c) recovering or enriching the circular RNA molecule generated from step (b) .
[0461] In one aspect, the present invention provides a circular RNA molecule prepared by the method of the invention.
[0462] Composition, cell and product
[0463] The invention also relates to a composition, which comprises the nucleic acid vector of the present invention and / or the circular RNA precursor of the present invention and / or the circular RNA of the present invention. In certain embodiments, the composition is a pharmaceutical composition. In certain embodiments, the composition is a cosmetic composition. In certain embodiments, the composition is a composition with a non-medical application. The specific use of the composition may depend on the sequence of interest.
[0464] In some embodiments, the composition is for use in preventing or treating a disease or condition in a subject. The specific disease or condition to be prevented or treated may depend on the specific sequence of interest.
[0465] Compositions of the invention may further comprise a carrier or excipient. In certain embodiments, the carrier is a pharmaceutically or cosmetically acceptable carrier or excipient.
[0466] In an embodiment, pharmaceutically or cosmetically acceptable carriers may include, but are not limited to, buffers, stabilizers, or preservatives. Examples of pharmaceutically or cosmetically acceptable carriers are physiologically compatible solvents, dispersion media, coating, antibacterial and antifungal agents, isotonic and absorption delay agents, etc., such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants or emulsifiers, or combinations thereof. The amount of pharmaceutically or cosmetically acceptable carrier in a composition may be experimentally determined based on the activity of the carrier and the desired properties of the preparation, such as stability and / or minimum oxidation. The carrier or excipient is "acceptable" in the sense that it is compatible with other components of the composition and is not harmful to its recipient.
[0467] The composition of the present invention may be formulated for any administration route. Preferably, the composition may be formulated in a form for intrarticularis (in joints) , parenteral, intravenous, intramuscular, dermal, buccal, subelingual, transnasal, intraperitoneal, subcutaneous, oral, topical, intrathecal, inhaled, transrectal, patch, pump, percutaneous, transrectal, muscular, body surface, mucosal, or intracranial administration.
[0468] The invention also provides a composition, which comprises a delivery means (e.g. lipid nanoparticles) comprising the nucleic acid vector of the present invention and / or the circular RNA precursor of the present invention and / or the circular RNA of the present invention or a delivery vector encoding the nucleic acid vector of the present invention and / or the circular RNA precursor of the present invention and / or the circular RNA of the present invention.
[0469] In one embodiment, the delivery means is selected from one or more of the following: macromolecular complexes, liposomes, nanocapsules, nanoparticles, exosomes, exosome-lipid conjugates, microspheres, beads, oil-in-water emulsions, lipid-nanoparticle conjugates, micelles, mixed micelles, and peptide-based polymeric complexes. In some embodiments, a suitable liposome, nanoparticle, or lipid-nanoparticle conjugate contains one or more non-cationic lipids, one or more cholesterol-based lipids, and / or one or more PEG-modified lipids. In another embodiment, a suitable liposome, nanoparticle, or lipid-nanoparticle conjugate contains lipids including (but not limited to) monoglycerides, diglycerides, thiolipids, lysolecithin, phospholipids, saponins, cholic acids, etc. In one embodiment, the compositions described herein comprise one or more liposomes or lipid nanoparticles.
[0470] In one embodiment, the delivery means comprises at least one targeting moiety. In another embodiment, the targeting moiety is a binding ligand, a mouse antibody, a human or humanized antibody, or a fragment thereof.
[0471] In some embodiments, the nucleic acid vector and / or the circular RNA precursor and / or the circular RNA described herein can be combined with the delivery means. In some embodiments, the combination may refer to incorporation into the delivery means, encapsulation within the liquid of the delivery means, dispersion within the delivery means, linking to the delivery means by linking molecules, embedding in the delivery means, compounding with the delivery means, dispersion in a solution containing the delivery means, mixing with the delivery means, binding with the delivery means, inclusion into the delivery means as a suspension, inclusion into or combination with micelles, attaching to a colloidal dispersion system, or combination with the delivery means in other ways.
[0472] In one embodiment, the delivery vector includes a non-viral, viral, plasmid, and non-plasmid vector. In one embodiment, examples of the virus delivery vectors include, but are not limited to, adenovirus vectors, adeno-associated virus (also known as adeno-associated virus, AAV) vectors, poxvirus vectors, herpes simplex virus I vectors, retrovirus vectors, lentiviral vectors, etc. In one embodiment, the delivery carrier is an AAV carrier. In one embodiment, the coding sequence of the nucleic acid vector and / or the circular RNA precursor and / or the circular RNA is operably linked to an expression element on the delivery vector.
[0473] Compositions described herein can be prepared by any method known in the art. In another aspect, the present invention provides a method for preparing a pharmaceutical or cosmetic composition for delivering to a subject, comprising:
[0474] (a) providing a circular RNA precursor of the invention,
[0475] (b) allowing the RNA precursor to undergo intramolecular self-cleavage, thereby generating the circular RNA molecule,
[0476] (c) recovering or enriching the circular RNA molecule generated from step (b) , and
[0477] (d) formulating the circular RNA molecule recovered or enriched from step (c) into a pharmaceutical or cosmetic composition.
[0478] In one aspect, the present invention provides a pharmaceutical or cosmetic composition prepared by the method of the invention.
[0479] In one aspect, the present invention provides a cell comprising the nucleic acid vector of the invention. The cell herein refers to a host cell. The nucleic acid vector of the invention can then be transcribed in the cell to produce a circular RNA molecule in the cell.
[0480] In one aspect, the present invention provides a product obtained by expressing the circular RNA of the present invention.
[0481] Use and method
[0482] In one aspect, the present invention provides use of the circular RNA precursor of the present invention and / or the circular RNA of the present invention, or the composition of the present invention as an expression vector.
[0483] In one aspect, the present invention provides the engineered RNA molecule of the present invention, the precursor RNA molecule of the present invention or the composition of the present invention, for use in preventing or treating a disease or condition in a subject.
[0484] In one aspect, the present invention provides a method of preventing or treating a disease or condition in a subject, comprising administering an effective amount of the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention to the subject.
[0485] In one aspect, the present invention provides use of the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention in preparation of a medicament for preventing or treating a disease or condition in a subject.
[0486] In one aspect, the present invention provides a method of delivering a circular RNA molecule to a subject, comprising administering to the subject a circular RNA molecule or composition of the invention. In embodiments, the TOI sequence comprises (i) a coding sequence of the polypeptide of interest, for example a therapeutic protein, and (ii) a translation initiation element, for example an IRES.
[0487] In one aspect, the present invention provides a method of providing a polypeptide of interest to a subject, comprising delivering to the subject a circular RNA molecule or composition of the invention. In embodiments, the TOI sequence comprises (i) a coding sequence of the polypeptide of interest, for example a therapeutic protein, and (ii) a translation initiation element, for example an IRES, so that the polypeptide of interest is expressed in the subject.
[0488] The specific disease or condition to be prevented or treated may depend on the specific sequence of interest.
[0489] As used herein, "subject" is intended to include living organisms in which an immune response can be elicited, and may include, but is not limited to, mammals, such as human or non-human mammals, such as domesticated, agricultural or wild animals, as well as birds and aquatic animals.
[0490] As used herein, the terms “administer” , “administering” , “administration” , “dose” , “dosing” , and the like refer to a method that can be used to enable the delivery of the composition to the desired site of biological action. These methods include, but are not limited to, intra-articular (in joints) , parenteral, intravenous, intramuscular, intradermal, buccal, sublingual, transnasal, intraperitoneal, subcutaneous, oral, topical, intrathecal, inhalation, transrectal, patch, pump, percutaneous, transrectal, muscular, body surface, mucosal, intracranial, etc administration. Possible administration techniques for the agents and methods described herein can be found, for example, in Goodman and Gilman, The Pharmacological Basis of Therapeutics, current edition; Pergamon and Remington, Pharmaceutical Sciences (current edition) , Mack Publishing Co., Easton, Pa.
[0491] As used herein, “treat” or "treating" an individual suffering from a disease means that the individual’s symptoms are partially or completely alleviated, or remain unchanged after treatment.
[0492] As used herein, “prevent” or "preventing" refers to prevention of a potential disease and / or prevention of worsening of symptoms or disease progression.
[0493] As used herein, "effective amount" refers to an amount when administered to the subject which results in beneficial or desired results, including clinical results, e.g., inhibits, suppresses or reduces the symptoms of the condition being treated in the subject as compared to a control. The precise amount of the component that provides an "effective amount" to an individual will depend on the mode of administration, the type and severity of the disease or condition, and on individual characteristics such as general health, age, sex, weight, and drug tolerance. One skilled in the art will be able to determine the appropriate dosage based on these and other factors.
[0494] As one skilled in the art will understand, the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention may be administered to patients in a variety of forms depending on the selected route of administration. The specific mode of administration and administration protocol will be selected by the attending clinician taking into account specific conditions (e.g. individual, disease, condition of disease involved, specific treatment) . In some embodiments, the method comprises administrating the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention by a mode selected from intraarticular (in joint) , parenteral, intravenous, intramuscular, dermal, buccal, sublingual, nasal, intraperitoneal, subcutaneous, oral, topical, intrathecal, inhaled, rectal, patch, pump, percutaneous, transrectal, muscular, body surface, mucous membrane, intracranial, etc. administration. In some embodiments, parenteral administration may be performed by continuous infusion over a selected period of time. In some embodiments, the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention is formulated in a form to be administered by a mode selected from intraarticular (in the joint) , parenteral, intravenous, intramuscular, dermal, buccal, sublingual, nasal, intraperitoneal, subcutaneous, oral, topical, intrathecal, inhaled, rectal, patch, pump, percutaneous, transrectal, muscular, body surface, mucous membrane, intracranial, etc. administration.
[0495] In some embodiments, the method may comprise the administration of the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention once or multiple times a day or less than once a day (e.g., once a week or once a month, etc. ) over a period of days to months or even years. In one embodiment, the method may comprise administrating to the subject the effective amount of the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention once a day, once every other day, twice a week, once a week, once every two weeks, once a month, once every other month, once every 3 months, or once every 6 months. In one embodiment, for the use described herein, the nucleic acid vector of the present invention, the circular RNA precursor of the present invention, the circular RNA of the present invention or the composition of the present invention is formulated in the form to be administered once a day, once every other day, twice weekly, once weekly, once every two weeks, once monthly, once every other month, once every three months, or once every six months.
[0496] The appropriate dose will depend, for example, on the specific nucleic acid vector of the present invention, the specific circular RNA precursor of the present invention, the specific circular RNA of the present invention or the specific composition of the present invention, the subject, the mode of administration and the nature and severity of the condition to be treated and the property of prior treatment the patient has undergone. Finally, the attending physician will determine the amount of the circular RNA, nucleic acid vector, or circular RNA precursor for an individual. In some embodiments, the attending physician may administer a low-dose of the circular RNA, the nucleic acid vector, or the circular RNA precursor of the present invention and observe the individual’s response. In other embodiments, the initial dose of the circular RNA, the nucleic acid vector, or the circular RNA precursor administered to an individual is higher, and the dose is then adjusted downward until symptoms of recurrence occur. Larger doses of the circular RNA, the nucleic acid vector, or the circular RNA precursor of the present invention may be administered until the individual has achieved optimal therapeutic effect, and at this point, the dose is generally not increased further. In some embodiments, the composition of the present invention is formulated to contain about 0.006mg to about 600mg, about 0.06mg to about 600mg, about 0.3mg to about 600mg, about 0.6mg to about 600mg, about 6mg to about 600mg, about 60mg to about 600mg, about 120mg to about 600mg, about 300mg to about 600mg, about 0.006mg to about 300mg, about 0.06mg to about 300mg, about 0.6mg to about 300mg, about 6mg to about 60mg, about 60mg to about 300mg, about 120mg to about 300mg, about 0.006 mg to about 60mg, about 0.06mg to about 60mg, about 0.3mg to about 60mg, about 0.6mg to about 60mg, about 6mg to about 60mg the circular RNA, the nucleic acid vector, or the circular RNA precursor of the present invention (per dose unit) . In some embodiments, a single administration of the composition is about 0.0001mg / kg to about 10mg / kg, about 0.001mg / kg to about 10mg / kg, about 0.005mg / kg to about 10mg / kg, about 0.01mg / kg to about 10mg / kg, about 0.1mg / kg to about 10mg / kg, about 1mg / kg to about 10mg / kg, about 2mg / kg to about 10mg / kg, about 5mg / kg to about 10mg / kg, about 0.0001mg / kg to about 5mg / kg, about 0.001mg / kg to about 5mg / kg, about 0.01m g / kg to about 5mg / kg, about 0.1mg / kg to about 10mg / kg, about 1mg / kg to about 5mg / kg, about 2mg / kg to about 5mg / kg, about 0.0001mg / kg to about 1mg / kg, about 0.001mg / kg to about 1mg / kg, about 0.005mg / kg to about 1m g / kg, about 0.01mg / kg to about 1mg / kg or about 0.1mg / kg to about 1mg / kg of the circular RNA, the nucleic acid vector, or the circular RNA precursor of the present invention.
[0497] In another aspect, the present invention also provides use of the nucleic acid vector of the invention, the circular RNA precursor of the present invention, the circular RNA of the invention or the composition of the invention in the manufacture of a cosmetic product.
[0498] In another aspect, the present invention also provides a method of expressing the circular RNA of the present invention in vivo, comprising
[0499] (a) delivering the circular RNA precursor of the present invention to a cell, and
[0500] (b) expressing the circular RNA of the present invention in vivo.
[0501] Clauses
[0502] 1. A circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0503] a) a linker I,
[0504] b) a spacer A,
[0505] c) a Group I / II intron,
[0506] d) an exon I,
[0507] e) a linker A,
[0508] f) a sequence of interest,
[0509] g) a linker B,
[0510] h) an exon II,
[0511] i) a spacer B,
[0512] j) a linker II,
[0513] wherein the linker I and the linker II are homology arm sequences capable of complementary pairing to each other to form a double-stranded region,
[0514] wherein said spacer A and B are non-complementary sequences,
[0515] wherein the Group I / II intron comprises a pairing G group and a 3’ end G group, and said Group I / II intron forms a U-G base pair with the 3’ end U of exon II,
[0516] wherein the exon I and the exon II are configured to be capable of forming a stem-loop-like structure, and
[0517] wherein the linker A and the linker B are partially complementary sequences and together are capable of forming a duplex.
[0518] 2. The circular RNA precursor of clause 1, wherein the Group I / II intron is or is derived from a Group I intron selected from Anabaena sp. YBS01, Staphylococcus phage Twort, Ncr. m. ND5, 1, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, Bfu. S1506, Mpl. L798, Pan. m. ND3, 1, Osp. S1199, Azoarcus olearius BH72, Scytalidium dimidiatum, Scytonema-hofmani tRNA fMet, and Agrobacterium-tumefaciens, preferably Anabaena sp. YBS01, Staphylococcus phage Twort, Azoarcus olearius BH72, Scytonema-hofmani tRNA fMet, Agrobacterium-tumefaciens, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, and Pan. m. ND3, 1.
[0519] 3. The circular RNA precursor of clause 1 or 2, wherein the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of the Group I / II intron and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of the Group I / II intron.
[0520] 4. The circular RNA precursor of any one of the preceding clauses, wherein the stem portion of the stem-loop-like structure comprises at least 2 base pairs, such as at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15 or more base pairs, such as 2-30 base pairs, such as 2-25 base pairs, such as 2-20 base pairs, such as 2-15 base pairs, such as 2-10 base pairs, such as 5-30 base pairs, such as 5-25 base pairs, such as 5-20 base pairs, such as 5-15 base pairs, such as 5-10 base pairs, preferably consecutive matched base pairs.
[0521] 5. The circular RNA precursor of any one of the preceding clauses, wherein the linker I and the linker II are of the same or different lengths.
[0522] 6. The circular RNA precursor of any one of the preceding clauses, wherein the double-stranded region comprises 50-300 base pairs, preferably consecutive matched base pairs.
[0523] 7. The circular RNA precursor of any one of the preceding clauses, wherein the predicted mean minimum free energy (MFE) of the linker I and the linker II is more than about-2.0 kal / mol, more than about-1.5 kal / mol, more than about -1.4 kal / mol, more than about -1.3 kal / mol, more than about -1.2 kal / mol, more than about -1.1 kal / mol, more than about -1.0 kal / mol, more than about -0.9 kal / mol, or more than about -0.8 kal / mol.
[0524] 8. The circular RNA precursor of any one of the preceding clauses, wherein said spacer A and B are non-complementary sequences.
[0525] 9. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and spacer B are each about 4-20 nucleotides in length, and the spacer A and spacer B have the same or different length.
[0526] 10. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and the spacer B do not interfere, inhibit or disrupt the formation of the functional P1 structure formed by the exon II and the 5’ end of the intron.
[0527] 11. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and / or the spacer B comprises polyA, polyU, polyAC, polyC or polyG.
[0528] 12. The circular RNA precursor of any one of the preceding clauses, wherein the linker A and the linker B are predicted to form a duplex of about 20 base pairs, about 16 base pairs, about 15 base pairs, about 14 base pairs, about 13 base pairs, about 12 base pairs, about 11 base pairs, about 10 base pairs, about 9 base pairs, about 8 base pairs, about 7 base pairs, about 6 base pairs, about 5 base pairs, about 4 base pairs, about 3 base pairs, about 2 base pairs, about 1 base pair in length.
[0529] 13. The circular RNA precursor of any one of the preceding clauses, wherein the sequence of interest comprises an ORF, optionally an IRES sequence operably linked to the ORF, and optionally a spacer C located between the IRES sequence and the ORF sequence.
[0530] 14. The circular RNA precursor of clause 13, wherein the ORF encodes POLR2A, Luciferase, Spike, COL3A1, BMP2, OTC, EGFP, AQP1 or SOD2 protein, or comprises a sequence selected from any one of SEQ ID NOs: 239-245.
[0531] 15. The circular RNA precursor of any one of the preceding clauses, wherein the exon I and / or exon II is a part of a sequence of interest.
[0532] 16. The circular RNA precursor of any one of the preceding clauses, wherein,
[0533] (i) the Group I / II intron is a Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 16, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 31, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 32, without any modification on the sequence forming P1 structure;
[0534] (ii) the Group I / II intron is a Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 17, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 33, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 34, without any modification on the sequence forming P1 structure;
[0535] (iii) the Group I / II intron is a Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 27, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 53, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 54, without any modification on the sequence forming P1 structure;
[0536] (iv) the Group I / II intron is a Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 29, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 57, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 58, without any modification on the sequence forming P1 structure;
[0537] (v) the Group I / II intron is a Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 30, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 59, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 60, without any modification on the sequence forming P1 structure;
[0538] (vi) the Group I / II intron is a Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 19, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 37, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 38, without any modification on the sequence forming P1 structure;
[0539] (vii) the Group I / II intron is a Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 20, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 39, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 40, without any modification on the sequence forming P1 structure;
[0540] (viii) the Group I / II intron is a Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 21, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 41, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 42, without any modification on the sequence forming P1 structure;
[0541] (ix) the Group I / II intron is a Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 22, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 43, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 44, without any modification on the sequence forming P1 structure; or
[0542] (x) the Group I / II intron is a Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 25, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 49, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 50, without any modification on the sequence forming P1 structure.
[0543] 17. The circular RNA precursor of any one of the preceding clauses, wherein
[0544] (i) the linker I comprises a nucleotide sequence of SEQ ID NO: 227, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 228 (corresponding to ICI 420 and 510) ;
[0545] (ii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 210, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 210, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 587) ;
[0546] (iii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 211, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 211, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 588) ;
[0547] (iv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 212, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 212, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 589) ;
[0548] (v) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 585) ;
[0549] (vi) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 214, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 214, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 700) ;
[0550] (vii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 215, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 216, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 666) ;
[0551] (viii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 217, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 218, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 667) ;
[0552] (ix) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 219, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 220, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 668) ;
[0553] (x) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 221, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 222, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 669) ;
[0554] (xi) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 590 and 607) ;
[0555] (xii) the linker I comprises a nucleotide sequence of SEQ ID NO: 231, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 232 (corresponding to ICI 591) ;
[0556] (xiii) the linker I comprises a nucleotide sequence of SEQ ID NO: 233, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 234 (corresponding to ICI 592) ;
[0557] (xiv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 22, the exon I comprises a nucleotide sequence of SEQ ID NO: 43, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 44, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230, and the sequence of interest comprises an ORF encoding OTC, BMP2 or COL3A1 (corresponding to ICI 1700, 1698, 1702) ; or
[0558] (xv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 17, the exon I comprises a nucleotide sequence of SEQ ID NO: 33, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 34, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230, and the sequence of interest comprises an ORF encoding OTC, BMP2 or COL3A1 (corresponding to ICI 1701, 1699, 1703) .
[0559] 18. The circular RNA precursor of any one of the preceding clauses, wherein the precursor is in linear form and is capable of undergoing autocatalytic circularization which generates a circular RNA with a circularization efficiency of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, or at least about 80%.
[0560] 19. The circular RNA precursor of any one of the preceding clauses, wherein the circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 127-132.
[0561] 20. A circular RNA precursor comprising from 5’ to 3’ direction, the following contiguously linked elements:
[0562] a) optionally a linker I,
[0563] b) optionally a spacer A,
[0564] c) a Group I / II intron,
[0565] d) an exon I,
[0566] e) optionally a linker A,
[0567] f) a sequence of interest,
[0568] g) optionally a linker B,
[0569] h) an exon II,
[0570] i) optionally a spacer B,
[0571] j) optionally a linker II,
[0572] wherein the exon I and the exon II are configured to be capable of forming a stem-loop-like structure.
[0573] 21. The circular RNA precursor of any one of the preceding clauses, wherein the Group I / II intron comprises a pairing G group and a 3’ end G group, and said Group I / II intron forms a U-G base pair with the 3’ end U of exon II.
[0574] 22. The circular RNA precursor of any one of the preceding clauses, wherein the Group I / II intron is or is derived from a Group I intron selected from Anabaena sp. YBS01, Staphylococcus phage Twort, Ncr. m. ND5, 1, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, Bfu. S1506, Mpl. L798, Pan. m. ND3, 1, Osp. S1199, Azoarcus olearius BH72, Scytalidium dimidiatum, Scytonema-hofmani tRNA fMet, and Agrobacterium-tumefaciens, preferably Anabaena sp. YBS01, Staphylococcus phage Twort, Azoarcus olearius BH72, Scytonema-hofmani tRNA fMet, Agrobacterium-tumefaciens, Aaz. b. trnL, Kap. S516, Sce. mL2449, RB3. v. nrdB, 1, and Pan. m. ND3, 1.
[0575] 23. The circular RNA precursor of any one of the preceding clauses, wherein the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of the Group I / II intron and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of the Group I / II intron.
[0576] 24. The circular RNA precursor of any one of the preceding clauses, wherein the stem portion of the stem-loop-like structure comprises at least 2 base pairs, such as at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15 or more base pairs, such as 2-30 base pairs, such as 2-25 base pairs, such as 2-20 base pairs, such as 2-15 base pairs, such as 2-10 base pairs, such as 5-30 base pairs, such as 5-25 base pairs, such as 5-20 base pairs, such as 5-15 base pairs, such as 5-10 base pairs, preferably consecutive matched base pairs.
[0577] 25. The circular RNA precursor of any one of the preceding clauses, wherein the linker I and the linker II are homology arm sequences capable of complementary pairing to each other to form a double-stranded region.
[0578] 26. The circular RNA precursor of any one of the preceding clauses, wherein the linker I and the linker II are of the same or different lengths.
[0579] 27. The circular RNA precursor of clause 25, wherein the double-stranded region comprises 50-300 base pairs, preferably consecutive matched base pairs.
[0580] 28. The circular RNA precursor of any one of the preceding clauses, wherein the predicted mean minimum free energy (MFE) of the linker I and the linker II is more than about-2.0 kal / mol, more than about-1.5 kal / mol, more than about -1.4 kal / mol, more than about -1.3 kal / mol, more than about -1.2 kal / mol, more than about -1.1 kal / mol, more than about -1.0 kal / mol, more than about -0.9 kal / mol, or more than about -0.8 kal / mol.
[0581] 29. The circular RNA precursor of any one of the preceding clauses, wherein said spacer A and B are non-complementary sequences.
[0582] 30. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and spacer B are each about 4-20 nucleotides in length, and spacer A and spacer B have the same or different length.
[0583] 31. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and the spacer B do not interfere, inhibit or disrupt the formation of the functional P1 structure formed by the exon II and the 5’ end of the intron.
[0584] 32. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and / or the spacer B comprises polyA, polyU, polyAC, polyC or polyG.
[0585] 33. The circular RNA precursor of any one of the preceding clauses, wherein the linker A and the linker B are partially complementary sequences, and the linker A and the linker B are predicted to form a duplex of about 20 base pairs, about 15 base pairs, about 15 base pairs, about 14 base pairs, about 13 base pairs, about 12 base pairs, about 11 base pairs, about 10 base pairs, about 9 base pairs, about 8 base pairs, about 7 base pairs, about 6 base pairs, about 5 base pairs, about 4 base pairs, about 3 base pairs, about 2 base pairs, about 1 base pair in length.
[0586] 34. The circular RNA precursor of any one of the preceding clauses, wherein the sequence of interest comprises an ORF, optionally an IRES sequence operably linked to the ORF, and optionally a spacer C located between the IRES sequence and the ORF sequence.
[0587] 35. The circular RNA precursor of clause 34, wherein the ORF encodes POLR2A, Luciferase, Spike, COL3A1, BMP2, OTC, EGFP, AQP1 or SOD2 protein, or comprises a sequence selected from any one of SEQ ID NOs: 239-245.
[0588] 36. The circular RNA precursor of any one of the preceding clauses, wherein the exon I and / or exon II is a part of a sequence of interest.
[0589] 37. The circular RNA precursor of any one of the preceding clauses, wherein,
[0590] (i) the Group I / II intron is a Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 16, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 31, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 32, without any modification on the sequence forming P1 structure;
[0591] (ii) the Group I / II intron is a Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 17, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 33, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 34, without any modification on the sequence forming P1 structure;
[0592] (iii) the Group I / II intron is a Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 27, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 53, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 54, without any modification on the sequence forming P1 structure;
[0593] (iv) the Group I / II intron is a Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 29, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 57, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 58, without any modification on the sequence forming P1 structure;
[0594] (v) the Group I / II intron is a Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 30, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 59, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 60, without any modification on the sequence forming P1 structure;
[0595] (vi) the Group I / II intron is a Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 19, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 37, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 38, without any modification on the sequence forming P1 structure;
[0596] (vii) the Group I / II intron is a Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 20, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 39, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 40, without any modification on the sequence forming P1 structure;
[0597] (viii) the Group I / II intron is a Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 21, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 41, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 42, without any modification on the sequence forming P1 structure;
[0598] (ix) the Group I / II intron is a Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 22, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 43, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 44, without any modification on the sequence forming P1 structure; or
[0599] (x) the Group I / II intron is a Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 25, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 49, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 50, without any modification on the sequence forming P1 structure.
[0600] 38. The circular RNA precursor of any one of the preceding clauses, wherein
[0601] (i) the linker I comprises a nucleotide sequence of SEQ ID NO: 227, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 228 (corresponding to ICI 420 and 510) ;
[0602] (ii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 210, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 210, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 587) ;
[0603] (iii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 211, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 211, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 588) ;
[0604] (iv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 212, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 212, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 589) ;
[0605] (v) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 585) ;
[0606] (vi) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 214, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 214, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 700) ;
[0607] (vii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 215, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 216, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 666) ;
[0608] (viii) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 217, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 218, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 667) ;
[0609] (ix) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 219, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 220, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 668) ;
[0610] (x) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 221, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 222, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 669) ;
[0611] (xi) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230 (corresponding to ICI 590 and 607) ;
[0612] (xii) the linker I comprises a nucleotide sequence of SEQ ID NO: 231, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 232 (corresponding to ICI 591) ;
[0613] (xiii) the linker I comprises a nucleotide sequence of SEQ ID NO: 233, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 16, the exon I comprises a nucleotide sequence of SEQ ID NO: 31, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 32, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 234 (corresponding to ICI 592) ;
[0614] (xiv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 22, the exon I comprises a nucleotide sequence of SEQ ID NO: 43, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 44, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230, and the sequence of interest comprises an ORF encoding OTC, BMP2 or COL3A1 (corresponding to ICI 1700, 1698, 1702) ; or
[0615] (xv) the linker I comprises a nucleotide sequence of SEQ ID NO: 229, the spacer A comprises a nucleotide sequence of SEQ ID NO: 213, the Group I / II intron comprises a nucleotide sequence of SEQ ID NO: 17, the exon I comprises a nucleotide sequence of SEQ ID NO: 33, the linker A comprises a nucleotide sequence of SEQ ID NO: 235, the linker B comprises a nucleotide sequence of SEQ ID NO: 236, the exon II comprises a nucleotide sequence of SEQ ID NO: 34, the spacer B comprises a nucleotide sequence of SEQ ID NO: 213, and the linker II comprises a nucleotide sequence of SEQ ID NO: 230, and the sequence of interest comprises an ORF encoding OTC, BMP2 or COL3A1 (corresponding to ICI 1701, 1699, 1703) .
[0616] 39. The circular RNA precursor of any one of the preceding clauses, wherein the precursor is in linear form and is capable of undergoing autocatalytic circularization which generates a circular RNA with a circularization efficiency of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, or at least about 80%.
[0617] 40. The circular RNA precursor of any one of the preceding clauses, wherein the circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 127-132.
[0618] 41. A circular RNA precursor, comprising in the 5’ to 3’ direction:
[0619] a) a linker I,
[0620] b) a spacer A,
[0621] c) a Group I intron,
[0622] d) an exon I,
[0623] e) a linker A,
[0624] f) a sequence of interest,
[0625] g) a linker B,
[0626] h) an exon II,
[0627] i) a spacer B, and
[0628] j) a linker II.
[0629] 42. The circular RNA precursor of any one of the preceding clauses, wherein the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region.
[0630] 43. The circular RNA precursor of any one of the preceding clauses, wherein the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region.
[0631] 44. The circular RNA precursor of any one of the preceding clauses, wherein the 3’ end region of the exon II (5’-exon) and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and the 5’ end region of the exon I (3’-exon) and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II.
[0632] 45. The circular RNA precursor of any one of the preceding clauses, wherein the 5’ end region of the exon II and the 3’ end region of the exon I are partially complementary to each other and together are capable of forming an exon duplex or a first stem-loop-like structure.
[0633] 46. The circular RNA precursor of any one of the preceding clauses, wherein the linker A and the linker B are partially complementary to each other and together are capable of forming a duplex.
[0634] 47. The circular RNA precursor of any one of the preceding clauses, wherein one base pair in the double-stranded region is formed by the 3’ end nucleotide of the linker I and the 5’ end nucleotide of the linker II.
[0635] 48. The circular RNA precursor of any one of the preceding clauses, wherein the double-stranded region formed by the linker I and the linker II and the non-pairing region formed by the spacer A and spacer B together are capable of forming a second stem-loop-like structure.
[0636] 49. The circular RNA precursor of any one of the preceding clauses, wherein the first nucleotide of the 5’ end region of the Group I intron is the splice site of the Group I intron.
[0637] 50. The circular RNA precursor of any one of the preceding clauses, wherein the 5’ end region of the Group I intron is capable of forming a part of the stem-loop structure of the P1 structure, and wherein the 5’ end region of the Group I intron is capable of forming the loop and a part of the stem of the P1 structure, wherein the 3’ end region of the exon II is capable of forming the remaining part of the stem of the P1 structure with the 5’ end region of the Group I / II intron.
[0638] 51. The circular RNA precursor of any one of the preceding clauses, wherein the exon II is the 3’ terminal portion of a 5’-flanking exon of the Group I intron, and the exon I is the 5’ terminal portion of a 3’-flanking exon of the Group I intron.
[0639] 52. The circular RNA precursor of any one of the preceding clauses, wherein the 5’ terminal nucleotide of the Group I / II intron and the 3’ terminal nucleotide of exon II together form the splice site in said P1 structure.
[0640] 53. A circular RNA precursor comprising in the 5’ to 3’ direction:
[0641] a) optionally a linker I,
[0642] b) optionally a spacer A,
[0643] c) a Group I intron,
[0644] d) a sequence of interest,
[0645] e) optionally a spacer B, and
[0646] f) optionally a linker II,
[0647] wherein the sequence of interest comprises a first fragment of the sequence of interest at its 3’ end and a second fragment of the sequence of interest at its 5’ end,
[0648] wherein the first fragment of the sequence of interest and the second fragment of the sequence of interest are respectively derived from a 5’ terminal portion and a 3’ terminal portion of an ORF sequence, a TIE or a non-TIE functional element,
[0649] wherein the first fragment of the sequence of interest comprises a ribozyme recognition sequence I located at its 3’ end, the second fragment of the sequence of interest comprises a ribozyme recognition sequence II located at its 5’ end, and
[0650] wherein in the circular RNA generated upon self-cleavage and circularization of the circular RNA precursor, the 3’ end of the ribozyme recognition sequence I is connected with the 5’ end of the ribozyme recognition sequence II.
[0651] 54. The circular RNA precursor of any one of the preceding clauses, wherein the ORF sequence is a protein-coding sequence or a non-coding sequence.
[0652] 55. The circular RNA precursor of any one of the preceding clauses, wherein the TIE is an IRES.
[0653] 56. The circular RNA precursor of any one of the preceding clauses, a ribozyme recognition sequence I is located at the 3’ end of the sequence of interest, and a ribozyme recognition sequence II is located at the 5’ end of the sequence of interest.
[0654] 57. A circular RNA precursor comprising in the 5’ to 3’ direction:
[0655] a) a linker I,
[0656] b) a spacer A,
[0657] c) a Group I intron,
[0658] d) a TOI sequence,
[0659] e) a spacer B, and
[0660] f) a linker II,
[0661] wherein the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region,
[0662] wherein the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region,
[0663] wherein the 5’ end region of the Group I / II intron is capable of forming the loop and a part of the stem of the P1 structure, the 3’ end region of the TOI sequence is capable of forming the remaining part of the stem of the P1 structure with the 5’ end region of the Group I / II intron, and the 5’ terminal nucleotide of the Group I / II intron and the 3’ terminal nucleotide of TOI sequence together form the splice site in said P1 structure.
[0664] 58. The circular RNA precursor of clause 57, wherein the 3’ end region of the TOI sequence and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and wherein the 5’ end region of the TOI and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II.
[0665] 59. The circular RNA precursor of clause 57 or 58, wherein the sequence of interest comprises a first fragment of the sequence of interest at its 3’ end and a second fragment of the sequence of interest at its 5’ end,
[0666] wherein the first fragment of the sequence of interest and the second fragment of the sequence of interest are respectively derived from a 5’ terminal portion and a 3’ terminal portion of an ORF sequence, a TIE or a non-TIE functional element,
[0667] wherein the first fragment of the sequence of interest comprises a ribozyme recognition sequence I located at its 3’ end, the second fragment of the sequence of interest comprises a ribozyme recognition sequence II located at its 5’ end, and
[0668] wherein in the circular RNA generated upon self-cleavage and circularization of the circular RNA precursor, the 3’ end of the ribozyme recognition sequence I is connected with the 5’ end of the ribozyme recognition sequence II.
[0669] 60. A nucleic acid vector for generating a circular RNA molecule, wherein said vector comprises a coding sequence of the circular RNA precursor of any one of clauses 1-59.
[0670] 61. The nucleic acid vector of clause 60, which further comprises a promoter sequence operably linked to the coding sequence of the circular RNA precursor.
[0671] 62. The nucleic acid vector of clause 60 or 61, which further comprises an RNA polymerase.
[0672] 63. A method for preparation of circular RNA, the method comprises:
[0673] 1) providing a circular RNA precursor of any one of clauses 1-59 or obtaining a circular RNA precursor by transcribing from the nucleic acid vector of any one of clauses 60-62;
[0674] 2) incubating the circular RNA precursor under the conditions allowing autocatalytic circularization to generate a circular RNA; and
[0675] 3) harvesting the circular RNA obtained in step 2) .
[0676] 64. A circular RNA, which is prepared from the circular RNA precursor of any one of clauses 1-59 or the vector of any one of clauses 60-62 or prepared from the method of clause 63.
[0677] 65. The circular RNA of clause 64, wherein the circular RNA comprises a nucleotide sequence of any one of SEQ ID NOs: 188-193.
[0678] 66. A composition comprising the circular RNA of clause 64 or 65.
[0679] 67. Use of the circular RNA of clause 64 or 65 or the composition of clause 66 in the manufacture of a medicament for preventing or treating a disease or condition in a subject.
[0680] 68. Use of the circular RNA of clause 64 or 65 or the composition of clause 66 in the manufacture of a cosmetic product.
[0681] 69. A cell comprising the nucleic acid vector of any one of clauses 60-62.
[0682] 70. A product obtained by expressing the circular RNA of clause 64 or 65.
[0683] 71. A method of expressing the circular RNA of clause 64 or 65 in vivo, comprising
[0684] (a) delivering the circular RNA precursor of any of clauses 1-59 to a cell, and
[0685] (b) expressing the circular RNA of clause 64 or 65 in vivo.
[0686] Examples
[0687] The following examples have been included to provide guidance to one of ordinary skill in the art for practicing representative embodiments of the presently disclosed subject matter. In light of the present disclosure and the general level of skill in the art, those of skill can appreciate that the following examples are intended to be exemplary only and that numerous changes, modifications, and alterations can be employed without departing from the scope of the presently disclosed subject matter. The synthetic descriptions and specific examples that follow are only intended for the purposes of illustration, and are not to be construed as limiting in any manner the present invention.
[0688] Example 1. CircRNA can be generated through integrated catalytic intron (ICI)
[0689] This embodiment demonstrates that the integrated catalytic intron (ICI) can effectively achieve RNA circularization. We designed a system in which the entire ribozyme is positioned on the 5’ end of the RNA precursor for circularization. This design mirrors the self-splicing Group I intron’s natural process that connects the exon regions on both sides of the intron through integrated intron, thus named the integrated catalytic intron (ICI) circularization system. We initially selected the cyanobacterium Anabaena group I intron (referred to as Ana below) for preliminary validation. The integrated Ana intron was placed at the 5’ end of the target of interest (TOI) , and the ribozyme recognition site Exon1 and ribozyme recognition site Exon2 were placed at both ends of the TOI (Figure 1A; ICI v0) . We understand that the Exons at both ends of the precursor can facilitate the second transesterification by coming into proximity and forming a duplex in space. Ultimately, after two complete transesterification reactions, the ends of the TOI region are connected to form a closed circular RNA structure. To assess the RNA circularization efficiency of this system, we performed in vitro transcription (IVT) and circularization reactions using a DNA template containing the POLR2A (which is a non-coding RNA) or SalivirusNG-J1_del34 IRES and a gene encoding Luciferase as the TOI, along with the previously mentioned circularization elements, demonstrating the system’s effectiveness in facilitating RNA circularization. According to the circularization products, the circularization efficiency is calculated as circ / (circ+pre) %by using Sciex (PA800 plus) . Subsequent enrichment of the circular RNA by RNase R resulted in a purer circular RNA product and to confirm bona fide circRNAs’ localization.
[0690] We have shown that RNA self-splicing circularization can be achieved by placing integrated intron at the 5’ end of the precursor RNA, but with limited circularization efficiency (Figure 1A; ICI v0) . Therefore, we optimized the ICI system by adding internal linkers (Figure 1A; ICI v1) (SEQ NO: 235, 236) and homology arms (HA) (Figure 1A; ICI v2) (SEQ NO: 227, 228) to the RNA precursors separately.
[0691] Firstly, internal linkers are added at both ends at the TOI of the precursor (Figure 1A, ICI v1) (SEQ NO: 63, 64) , assuming that the internal linker will function as homology arms and spacer sequences in this case. The results showed a slight improvement in the circularization efficiency of the precursor RNA (Figure 1B, D: POLR2A, ICI313→ ICI314, 35.04%→42.95%; Figure 1C, D: Salivirus NG-J1_del34-Luc, ICI412→ICI414, 20.14%→34.85%) .
[0692] However, the circularization efficiency of the precursor RNA was significantly inhibited when we seamlessly interacted with the homology arm at the 5’ and 3’ ends of the precursor (Figure 1B, D: POLR2A, ICI313→ ICI311, 35.04%→2.28%; Figure 1, C, D: Salivirus NG-J1_del34-Luc, ICI412→ICI310, 20.14%→0.18%) (SEQ NO: 65, 66) . The results suggest that the homology arm may be strongly complementary, which hinders the formation of intron P1 structure, thereby inhibiting the occurrence of the first transesterification reaction to significantly reduce the circularization efficiency.
[0693] Example 2. Homology arms can significantly improve circularization on P1 site of PIE
[0694] Similarly, in the PIE system, the design divides Ana intron into two parts in the P1 domain, which are placed at each end of the TOI. At the same time, homology arms were added at the 5’ and 3’ ends of the precursor to function as homology arms (Figure 2A) (SEQ NO: 227, 228) . The results showed that the homology arm could significantly promote the circularization efficiency of precursor RNA in the PIE system (Figure 2B, PIE293: 77.56%; PIE309: 84.24%) (SEQ NO: 67, 68) . These results suggest that the homology arms can act as homology arms to promote the formation of spatial duplexes of precursor RNA in the PIE system, thereby facilitating the occurrence of the first transesterification reaction. The impact of this design on circularization is significantly different between the ICI system and the PIE system, suggesting that the two systems are unique in design.
[0695] Example 3. Spacers addition between homology arm and intron / exon can significantly promote circularization efficiency on ICI
[0696] Furthermore, in order to analyze whether seamless homology arms may affect the formation of P1 structure in the ICI system and affect ribozyme splicing, we added spacer sequences of non-complementary pairs to the ends of the homology arms (the 3’ end of the 5’ homology arm and the 5’ end of the 3’ homology arm) (SEQ NO: 213) . This modification aims to create a looser and more flexible local space for the spatial duplex of the paired homology arm and intron, facilitating intron self-splicing (Figure 3A, ICI v2.1) . The results show that the combination of the spacer and homology arm design significantly enhanced the circularization efficiency of the ICI system in both POLR2A and Salivirus NG-J1_del34-Luciferae (Figure 3B, ICI420: 86.36%, ICI510: 78.31%) (SEQ NO: 69, 70) .
[0697] Example 4. Impact of additional spacer length between homology arm and intron / exon on ICI system circularization efficiency
[0698] This embodiment demonstrates that increasing the length of additional spacers can enhance the RNA circularization efficiency in the ICI system. Spacer sequences of defined lengths were inserted between the homology arm and intron, designating the SalivirusNG-J1_del34 IRES-Luciferase sequence as TOI (Figure 4A; SEQ ID NOs: 71-76) . After a unified in vitro transcription and circularization reaction, capillary gel electrophoresis (CGE) quantitative analysis revealed a significant correlation between spacer length and circularization efficiency. Specifically, incorporation of poly (A) tracts of increasing lengths (SEQ ID NOs: 209-214) resulted in progressively elevated circularization efficiencies (3.22%, 59.59%, 71.31%, 74.73%, 80.85%to 81.19%) as demonstrated by quantitative CGE data (Figure 4B) .
[0699] Example 5. Impact of additional spacers between homology arm and intron / exon on P1 structure compromises ICI circularization efficiency
[0700] In order to further explore the impact of spacer sequences on ICI, spacer pairs were randomly generated with a length of 10 nt and categorized into three types: unstructured base pairs (SEQ ID NOs: 215-218) , non-interfering P1-forming base pairs (SEQ ID NOs: 219-222) , and P1-disruptive base pairs (SEQ ID NOs: 223-226) (Figure 5A) . After a unified in vitro transcription and circularization reaction, CGE analysis showed that compared with the poly (A) spacer of equivalent length (Figure 5B; ICI585: 80.85%; SEQ ID NO: 75) , non-interfering P1-forming spacers did not alter circularization efficiency (ICI668: 79.67%, ICI669: 75.59%; SEQ ID NOs: 79, 80) . Unstructured spacer pairs moderately reduced circularization efficiency (ICI666: 64.56%, ICI667: 53.56%; SEQ ID NOs: 77, 78) , whereas P1-disruptive base pairs exhibited profound inhibition of this process (ICI670: 0.82%, ICI671: 0.57%; SEQ ID NOs: 81, 82) .
[0701] Example 6. P1 modification will affect the ICI system’s circularization efficiency
[0702] To explore the impact of P1 structure on the circularization efficiency of the ICI system, structural modifications were implemented targeting the P1 loop and duplex domains (SEQ ID NOs: 83-90) (Figure 6A) .
[0703] First, base-pair insertions were introduced into the P1 duplex region (Figure 6B) : 1) 5 base pairs were inserted upstream of CUU (between site: 0 and site: +1) , resulting in a marginal reduction in circularization efficiency (ICI595: 66.85%) ; 2) 5 base pairs were inserted downstream of CUU (site: -2) , significantly suppressing circularization efficiency (ICI674: 0.20%) ; 3) Insertions at both upstream and downstream positions of CUU caused a marked inhibition of circularization efficiency (ICI675: 0.27%) .
[0704] Next, base-pair disruptions were induced via nucleotide substitutions in the P1 duplex region (Figure 3B) : 1) Disrupting the base pairs by replacing the base pairs at upstream of CUU from AU to AC (site: +1 and site: +2) resulted in nearly a two-fold decrease in circularization efficiency (ICI644: 45.73%) ; 2) Disrupting the base pair by replacing the nucleotides at last two base pairs of CUU from UA to UC (site: -1) and from CG to CA (site: -2) significantly inhibited circularization efficiency (ICI645: 9.72%) ; 3) Combining the previous two modifications disrupted the base pairs at upstream and downstream of CUU, further suppressed circularization efficiency (ICI646: 12.30%) .
[0705] Then, the sequence replacement was made at the splicing site in the P1 duplex region. The CUU was replaced with CUC, resulting in more than a two-fold decrease in circularization efficiency (Figure 6B; ICI648: 39.60%) .
[0706] Finally, the substitution of AUAA with CCCC within the P1 loop domain induced a modest reduction in circularization efficiency (Figure 6B; ICI647: 59.93%) .
[0707] Example 7. Internal linker combines with homology arm and spacer further improve circularization
[0708] By combining internal linkers and homology arms while maintaining the P1 structure through spacer sequences, we developed the ICI v3.0 version (Figure 7A-B) . Performance validation using defined TOIs demonstrated that ICI v3.0 version (ICI590 / 607; Figure 7C; SEQ ID NOs: 91, 93) achieved significantly higher circularization efficiency compared to the previous ICI v2.1 version (ICI585 / 606; Figure 7C; SEQ ID NOs: 75, 92) . Notable improvements included: SalivirusNG-J1_del34-Luciferase constructs: 87.17%vs 80.85%; CAV2-Luciferase constructs: 69.22%vs 8.35%
[0709] Next, to assess system versatility, we evaluated circularization efficiency across TOIs with diverse length ranges and targets, like Spike, COL3A1, BMP2 and OTC (SEQ ID NOs: 241-244) . Results confirmed that the optimized ICI v3.0 system maintained high performance across different targets (Figure 7D; ICI639: 82.02%, ICI1560: 75.33%, ICI1561: 87.85%, ICI1562: 90.31%) (SEQ ID NOs: 94-97) .
[0710] Example 8. Pairing strength of homology arm affects ICI circularization efficiency
[0711] The influence of homology arm pairing strength on the circularization efficiency of the ICI system was investigated by designing sequence variants with varying lengths and minimum free energy (MFE) values (SEQ ID NOs: 227-234) targeting different ICI versions. Results demonstrated that for SalivirusNG-J1_del34-Luciferase construct, the circularization efficiency was positively correlated with the average MFE value (Figure 8A-B; SEQ ID NOs: 70, 75, 91, 98, 99) .
[0712] Example 9. Optimized ICI system markedly enhances RNA circularization efficiency of most Group I Introns
[0713] To further analyze whether the optimized ICI system has universality, we extended the optimized ICI system to different Group I Introns for further validation (SEQ ID NOs: 100-126) . After a unified in vitro transcription and circularization reaction, using CGE for quantitative analysis of circularization efficiency, the results showed that compared to the ICI v0 version, the optimized ICI system (ICI v3.0) significantly increased the circularization efficiency of most Group I Introns (Figure 9A) . Among them, RB3. v. nrdB, 1 and Staphylococcus phage Twort showed relatively higher circulation efficiency in ICI v3.0 version. Next, to assess their versatility, we evaluated circularization efficiency across TOIs with diverse targets, like COL3A1, OTC and BMP2 (SEQ ID NOs: 127-132) . Results confirmed that the two new group I introns in optimized ICI v3.0 system maintained high performance across different targets with different lengths (Figure 9B, ICI1700: 41.31%; ICI1701: 70.58%; ICI1698: 36.66%; ICI1699: 67.78%; ICI1702: 65.92%; ICI1703: 79.68%) .
[0714] Example 10. Optimized ICI system markedly better efficiency
[0715] Recent studies have reported that RNA circularization can be achieved through a Trans-Ribozyme-based circularization (TRIC) method (Du, Y., Zuber, P. K., Xiao, H. et al. Efficient circular RNA synthesis for potent rolling circle translation. Nat. Biomed. Eng, 2024. ) , which demonstrates comparable or better circularization efficiency compared to traditional PIE methods. Therefore, we compared the circularization efficiency of the optimized ICI system (ICI v3.0) with the TRIC system (TRIC-V2) utilizing the same Anabaena Group I Intron, and a TOI sequence derived from the reported literature (Du, Y., Zuber, P. K., Xiao, H. et al. Efficient circular RNA synthesis for potent rolling circle translation. Nat. Biomed. Eng, 2024. ) . The results showed that our optimized ICI v3.0 version exhibited significantly higher circularization efficiency than the best circularization system reported in the literature, TRIC-V2. Specifically, TRIC-V2 achieved 77.72%circularization efficiency (TRIC-V2 636; SEQ ID NO: 133) , while the ICI system reached 93.62% (ICI638; SEQ ID NO: 134) (Figure 10) .
[0716] Example 11. The number of pairing bases and pairing strength on exon duplex significantly affect the circularization efficiency of the ICI system
[0717] To further analyze the impact of exon duplex on the circularization efficiency of the ICI system, different lengths and pairing strengths of sequences were designed targeting the exon duplex (Figure 11A) .
[0718] First, we performed extension / truncation design on the base pairs of the exon duplex. Based on the natural 5 base pairs of the Anabaena Group I Intron: 1) We inserted additional 5 base pairs away from CUU. The results showed that the 10 base pairs exon sequence enhanced the circularization efficiency of ICI compared to the natural 5 base pairs (Figure 11B; ICI679: 90.95, ICI585: 78.60%; SEQ ID NOs: 135, 75) . 2) We deleted the GC, CG and AU base pairs in the exon-duplex region (namely, leaving only 2 base pairs close to the CUU) . The results indicated that the circularization efficiency of ICI with the 2 base pairs exon sequences was suppressed compared to the natural 5 base pairs (Figure 11B; ICI681: 56.79%, ICI585: 78.60%; SEQ ID NOs: 137, 75) .
[0719] Next, we enhanced the base pairing strength on the exon duplex: based on a 10-base-paired exon sequence that promotes circularization efficiency (ICI679) , we increased the GC content to 100% (ICI680; SEQ ID NO: 136) . The results are shown in Figure 11B.
[0720] Example 12. Seamless ICI has considerable circularization efficiency and expression level
[0721] For Group I Introns, the three-nucleotide motif (CUU or CAU) in Exon 2 serves as ribozyme recognition sites, pairing with the Internal Guiding Sequence (IGS) of the intron to form the P1 double-stranded region of the ribozyme, which can be used to determine the first splicing site. Exon 1 participates in forming the P10 double-stranded region of the ribozyme and determines the second splicing site. During the splicing process, Exon1 and Exon2 are retained in the circular RNA. Published studies indicated that the retained Exon sequences may induce potential immunogenicity (Liu CX, Guo SK, Nan F, Xu YF, Yang L, Chen LL. RNA circles with minimized immunogenicity as potent PKR inhibitors. Mol Cell. 2022 Jan 20; 82 (2) : 420-434. e6. ) . Therefore, we developed a seamless ICI system to generate the circular RNA by concealing the sequences of Exon 1 and Exon 2 within the TOI sequence (Figure 12A) . The design strategy was as follows: we identified the motifs CUU or CAU in the TOI as ribozyme recognition sites by using a degenerate codon mutation method. The upstream and downstream sequences flanking the ribozyme recognition site were positioned at the two ends of TOI position. Precursor sequences were specifically constructed in the 5’ to 3’ direction as follows: Linker1, 5’ Spacer, Ribozyme, sequence from downstream of the motif CUU or CAU (excluding the motif) to the 3’ end of the TOI, sequence from the 5’ end of the TOI to the motif CUU / CAU, 3’ Spacer, and Linker2.
[0722] Based on Anabaena Group I Intron and Twort Group I Intron, the seamless ICI systems were designed for the two targets of interest, AQP1 and SOD2. These systems were evaluated for circularization efficiency (SEQ ID NOs: 138, 139) . Experimental data demonstrated that both seamless ICI systems achieved effective RNA circularization: the Anabaena-based system (ICI593) achieved 48.44%circularization efficiency, while the Twort-based system (ICI594) exhibited higher efficacy at 78.89% (Figure 12B) . The expression levels in A253 cells confirmed successful intracellular expression of circRNAs generated by the seamless ICI systems (Figure 12C) .
[0723] Sequences mentioned in Examples are listed as below:
[0724] While the invention is described in conjunction with the enumerated embodiments, it will be understood that they are not intended to limit the invention to those embodiments. The invention is intended to cover all alternatives, modifications, and equivalents that may be included within the scope of the present invention. One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention. The present invention is in no way limited to the methods and materials described. In the event that one or more of the incorporated literature, patents, and similar materials differs from or contradicts this application, including but not limited to defined terms, term usage, described techniques, or the like, this application controls.
[0725] All references including patents, patent applications and publications cited in the present application are incorporated herein by reference in their entirety, as if each of them is individually incorporated. Further, it would be appreciated that one skilled in the art could make various changes or modifications to the invention without departing from the scope of the invention defined by the appended claims below. Accordingly, the present invention is not intended to be limited to the disclosed embodiments. Rather the present invention is intended to cover the disclosed embodiments as well as others falling within the scope and spirit of the invention to the fullest extent permitted in view of this disclosure and the inventions defined by the claims appended herein below.
Claims
A circular RNA precursor, comprising in the 5’ to 3’ direction:a) a linker I,b) a spacer A,c) a Group I intron,d) an exon I,e) a TOI sequence,f) an exon II,g) a spacer B, andh) a linker II,wherein the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region,wherein the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region,wherein the 5’ end region of the Group I intron is capable of forming the loop and a part of the stem of the P1 structure, the 3’ end region of the exon II is capable of forming the remaining part of the stem of the P1 structure with the 5’ end region of the Group I intron, and the 5’ terminal nucleotide of the Group I intron and the 3’ terminal nucleotide of exon II together form the splice site in said P1 structure, andwherein the 5’ end region of the exon II and the 3’ end region of the exon I are partially complementary to each other and together are capable of forming an exon duplex.The circular RNA precursor of claim 1, wherein the precursor further comprises a linker A located between the exon I and the TOI sequence, and a linker B located between the exon II and the TOI sequence, wherein the linker A and the linker B are partially complementary to each other and together are capable of forming a duplex.The circular RNA precursor of any one of the preceding claims, wherein the 3’ end region of the exon II and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and the 5’ end region of the exon I and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II.The circular RNA precursor of any one of the preceding claims, wherein the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of the Group I intron and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of the Group I intron.The circular RNA precursor of any one of the preceding claims, wherein the exon duplex comprises at least 2 base pairs, such as 2-30 base pairs, such as 2-25 base pairs, such as 2-20 base pairs, such as 2-15 base pairs, such as 2-10 base pairs, such as 5-30 base pairs, such as 5-25 base pairs, such as 5-20 base pairs, such as 5-15 base pairs, such as 5-10 base pairs, preferably consecutive matched base pairs.A circular RNA precursor comprising in the 5’ to 3’ direction:a) a linker I,b) a spacer A,c) a Group I intron,d) a TOI sequence,e) a spacer B, andf) a linker II,wherein the linker I and the linker II are complementary to each other and together are capable of forming a double-stranded region,wherein the spacer A and spacer B are non-complementary sequences and together are capable of forming a non-pairing region,wherein the 5’ end region of the Group I intron is capable of forming the loop and a part of the stem of the P1 structure, the 3’ end region of the TOI sequence is capable of forming the remaining part of the stem of the P1 structure with the 5’ end region of the Group I intron, and the 5’ terminal nucleotide of the Group I intron and the 3’ terminal nucleotide of TOI sequence together form the splice site in said P1 structure.The circular RNA precursor of claim 6, wherein the 3’ end region of the TOI sequence and the 5’ end region of the Group I intron are capable of forming a P1 structure containing a ribozyme recognition sequence I, and the 5’ end region of the TOI and both the 5’ and 3’ end regions of the Group I intron are capable of forming a P10 structure containing a ribozyme recognition sequence II.The circular RNA precursor of claim 6, wherein the TOI sequence comprises a first fragment of the TOI sequence at its 3’ end and a second fragment of the TOI sequence at its 5’ end,wherein the first fragment of the TOI sequence and the second fragment of the TOI sequence are respectively derived from a 5’ terminal portion and a 3’ terminal portion of an ORF sequence, a TIE or a non-TIE functional element,wherein the first fragment of the TOI sequence comprises a ribozyme recognition sequence I located at its 3’ end, the second fragment of the TOI sequence comprises a ribozyme recognition sequence II located at its 5’ end, andwherein in the circular RNA generated upon self-cleavage and circularization of the circular RNA precursor, the 3’ end of the ribozyme recognition sequence I is connected with the 5’ end of the ribozyme recognition sequence II.The circular RNA precursor of any one of the preceding claims, wherein the double-stranded region formed by linker I and linker II comprises 50-300 base pairs, preferably consecutive matched base pairs.The circular RNA precursor of any one of the preceding claims, wherein the predicted mean minimum free energy (MFE) of the linker I and the linker II is more than about-2.0 kal / mol, more than about-1.5 kal / mol, more than about -1.4 kal / mol, more than about -1.3 kal / mol, more than about -1.2 kal / mol, more than about -1.1 kal / mol, more than about -1.0 kal / mol, more than about -0.9 kal / mol, or more than about -0.8 kal / mol.The circular RNA precursor of any one of the preceding claims, wherein the spacer A and spacer B are each about 4-20 nucleotides in length, and spacer A and spacer B have the same or different length.The circular RNA precursor of any one of the preceding claims, wherein the spacer A and / or the spacer B comprises polyA, polyU, polyAC, polyC or polyG.The circular RNA precursor of any one of the preceding claims, wherein the double-stranded region formed by the linker I and the linker II and the non-pairing region formed by the spacer A and spacer B together are capable of forming a stem-loop-like structure.The circular RNA precursor of any one of the preceding claims, wherein,(i) the Group I intron is a Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 16, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 31, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Anabaena sp. YBS01 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 32, without any modification on the sequence forming P1 structure;(ii) the Group I intron is a Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 17, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 33, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Staphylococcus phage Twort or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 34, without any modification on the sequence forming P1 structure;(iii) the Group I intron is a Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 27, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 53, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Azoarcus olearius BH72 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 54, without any modification on the sequence forming P1 structure;(iv) the Group I intron is a Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 29, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 57, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Scytonema-hofmani tRNA fMet or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 58, without any modification on the sequence forming P1 structure;(v) the Group I intron is a Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 30, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 59, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Agrobacterium-tumefaciens or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 60, without any modification on the sequence forming P1 structure;(vi) the Group I intron is a Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 19, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 37, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Aaz. b. trnL or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 38, without any modification on the sequence forming P1 structure;(vii) the Group I intron is a Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 20, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 39, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Kap. S516 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 40, without any modification on the sequence forming P1 structure;(viii) the Group I intron is a Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 21, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 41, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Sce. mL2449 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 42, without any modification on the sequence forming P1 structure;(ix) the Group I intron is a Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 22, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 43, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of RB3. v. nrdB, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 44, without any modification on the sequence forming P1 structure; or(x) the Group I intron is a Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 25, and the exon I is a contiguous fragment starting from the 5’ terminal nucleotide of the native 3’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 49, and the exon II is a contiguous fragment starting from the 3’ terminal nucleotide of the native 5’ exon of Group I intron of Pan. m. ND3, 1 or comprises a sequence having at least 85%sequence identity with SEQ ID NO: 50, without any modification on the sequence forming P1 structure.The circular RNA precursor of any one of the preceding claims, wherein said TOI sequence encodes an OTC polypeptide.The circular RNA precursor of claim 15, wherein said circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 129 and 130.The circular RNA precursor of any one of the preceding claims, wherein said TOI sequence encodes a BMP2 polypeptide.The circular RNA precursor of claim 17, wherein said circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 127 and 128.The circular RNA precursor of any one of the preceding claims, wherein said TOI sequence encodes a COL3A1 polypeptide.The circular RNA precursor of claim 19, wherein said circular RNA precursor comprises a nucleotide sequence of any one of SEQ ID NOs: 131 and 132.A nucleic acid vector for generating a circular RNA molecule, comprising a DNA sequence encoding the circular RNA precursor of any one of claims 1-20.A cell comprising the nucleic acid vector of claim 21.A method for preparing a circular RNA molecule, comprising:(a) providing a circular RNA precursor of any one of claims 1-20,(b) allowing the RNA precursor to undergo intramolecular self-cleavage, thereby generating the circular RNA molecule, and(c) recovering or enriching the circular RNA molecule generated from step (b) .A method for preparing a pharmaceutical or cosmetic composition for delivering to a subject, comprising:(a) providing a circular RNA precursor of any one of claims 1-20,(b) allowing the RNA precursor to undergo intramolecular self-cleavage, thereby generating the circular RNA molecule,(c) recovering or enriching the circular RNA molecule generated from step (b) , and(d) formulating the circular RNA molecule recovered or enriched from step (c) into a pharmaceutical or cosmetic composition.The method of claim 24, wherein the composition further comprises a pharmaceutically or cosmetically acceptable carrier or excipient.The method of claim 24 or 25, wherein the composition further comprises a delivery means.The method of claim 26, wherein the delivery means is a lipid nanoparticle.A circular RNA molecule prepared by the method of claim 23 or a pharmaceutical or cosmetic composition prepared by the method of any one of claims 24-27.Use of the circular RNA precursor of any of claims 1-20, or the vector of claim 21, or the cell of claim 22, or the circular RNA molecule or the composition of claim 28 in the manufacture of a medicament for preventing or treating a disease or condition in a subject or in the manufacture of a cosmetic product.A method of delivering a circular RNA molecule to a subject, comprising administering to the subject a circular RNA molecule or composition of claim 28, wherein the TOI sequence comprises (i) a coding sequence of the polypeptide of interest, for example a therapeutic protein, and (ii) a translation initiation element, for example an IRES.A method of providing a polypeptide of interest to a subject, comprising delivering to the subject a circular RNA molecule or composition of claim 28, wherein the TOI sequence comprises (i) a coding sequence of the polypeptide of interest, for example a therapeutic protein, and (ii) a translation initiation element, for example an IRES, so that the polypeptide of interest is expressed in the subject.
Citation Information
Patent Citations
Circular RNA for translation in eukaryotic cells
CN112399860A
Constructs and methods for preparing circular RNA
CN117795073A
RNA circularization
TW202426640A
Circular rnas and their use in immunomodulation
US20190345503A1
Circularized engineered RNA and methods
US20210277393A1