Novel internal ribosome entry sites and uses thereof

Synthetic IRES sequences in nucleic acids address the limitations of naturally occurring IRES by enabling efficient and controlled protein expression in eukaryotic cells, suitable for therapeutic applications.

WO2025218812A1PCT designated stage Publication Date: 2025-10-23SHANGHAI CIRCODE BIOMED CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/090211
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2025-04-21
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Naturally occurring internal ribosome entry site (IRES) elements are too large, complex, and prone to host cell immune rejection, making them unsuitable for efficient protein expression in vivo applications such as protein replacement therapies and gene therapy.

Method used

Development of non-naturally occurring nucleic acids with synthetic IRES sequences that are at least 85-100% identical to specific nucleotide sequences, which can be used to express therapeutic proteins efficiently and immunologically inertly in eukaryotic cells.

Benefits of technology

The synthetic IRES sequences enable efficient and prolonged protein expression, suitable for therapeutic applications, providing precise control over protein expression and reducing immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025090211-FTAPPB-I100001
    Figure PCTCN2025090211-FTAPPB-I100001
  • Figure PCTCN2025090211-FTAPPB-I100002
    Figure PCTCN2025090211-FTAPPB-I100002
  • Figure PCTCN2025090211-FTAPPB-I100003
    Figure PCTCN2025090211-FTAPPB-I100003
Patent Text Reader

Abstract

Provided are novel Internal Ribosome Entry Site (IRES) sequences that improve cap-independent translation in eukaryotic cells. Provides herein are nucleic acids incorporating such sequences, vectors and host cells comprising these sequences, methods for their use in producing proteins of interest, and a method of improving expression, functional stability, immunogenicity, ease of manufacturing and / or half-life of a therapeutic protein encoded by the circular RNA.
Need to check novelty before this filing date? Find Prior Art

Description

NOVEL INTERNAL RIBOSOME ENTRY SITES AND USES THEREOFCross reference to related applicationsThe present invention claims the benefit of the priority of International Application No. PCT / CN2024 / 088899, filed April 19, 2024; the disclosure of each of which is incorporated hererin by reference in its entirety. REFERENCE TO A SEQUENCE LISTINGThe present invention incorporates herein by reference a Sequence Listing submitted with this application as an XML file entitled “TFI01225PCT-Sequence listing. xml” created on April 18, 2025 and having a size of 893,008 bytes.1. Field

[0001] The present invention relates to the field of molecular biology, in particular to constructs and methods for recombinantly expressing a protein of interest in a eukaryotic cell.2. Background

[0002] Internal ribosome entry site ( “IRES” ) elements are useful for cap-independent gene expression eukaryotic cells. However, naturally found IRES elements may 1) have a large length of nucleic acid residues, 2) contain complex secondary structures, and / or 3) be prone to host cell immune rejection, each of which is not preferable for in vivo applications, such as protein replacement therapies. Therefore, there exists a need in the art for the development of synthetic IRES elements that can result in efficient protein expression. There also exists a need in the art for providing nucleic acids (e.g., circular RNA) having synthetic IRES elements that can be suitable for use as a medicament or vaccine, such as for application in gene therapy and / or genetic vaccination. The compositions, methods and systems provided herein address these unmet needs and provide related advantages.3. Summary

[0003] Provided herein are non-naturally occurring nucleic acids comprising a translation initiation sequence (TI) comprising an internal ribosome entry site (IRES) that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.

[0004] In some embodiments, the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.

[0005] In some embodiments of the nucleic acids disclosed herein, the TI comprises at least two IRESs. In some embodiments, the TI further comprises a natural IRES sequence or a fragment thereof.

[0006] In some embodiments, nucleic acids provided herein are DNA. In some embodiments, nucleic acids provided herein are RNA. In some embodiments, nucleic acids provided herein are double stranded. In some embodiments, nucleic acids provided herein are single stranded. In some embodiments, nucleic acids provided herein are circular. In some embodiments, nucleic acids provided herein are linear. In some embodiments, nucleic acids provided herein are mRNA.

[0007] In some embodiments, nucleic acids provided herein further comprise a therapeutic protein-coding sequence (Z1) operatively linked to the TI. In some embodiments, Z1 is linked to TI via a linker. In some embodiments, Z1 encodes a therapeutic protein. In some embodiments, nucleic acids provided herein are express the protein for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid is administered to a human. In some embodiments, the expression of the therapeutic protein is immunologically inert after the nucleic acid is administered to a human.

[0008] Provided herein are also proteins expressed by the nucleic acids described herein.

[0009] Provided herein are also vectors comprising the nucleic acids described herein.

[0010] Provided herein are also cells comprising the nucleic acids described herein.

[0011] In some embodiments of the cells provided herein, the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI. Provided herein are also methods of expressing a protein comprising culturing the cell described herein under conditions and for a sufficient time for the expression of the protein.

[0012] Provided herein are also pharmaceutical compositions comprising the nucleic acid described herein, the vector described herein, or the cell described herein, and a pharmaceutically acceptable carrier, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI, wherein Z1 encodes a therapeutic protein.

[0013] Provided herein are also methods of expressing a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition described herein.

[0014] Provided herein are also methods of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition described herein. Provided herein are also uses of the pharmaceutical composition described herein for treating or preventing a disease or disorder.

[0015] Provided herein are also non-naturally occurring nucleic acids encoding an RNA comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circular RNA (circRNA) that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment; and wherein the circRNA comprises a TI that comprises an IRES that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.

[0016] In some embodiments, the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.

[0017] In some embodiments of the nucleic acid disclosed herein, the TI comprises at least two IRESs. In some embodiments, the TI further comprises a natural IRES sequence or a fragment thereof.

[0018] In some embodiments of the nucleic acid disclosed herein, the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.

[0019] In some embodiments, Z1 encodes a therapeutic protein. In some embodiments, the therapeutic protein is expressed for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid or the circRNA is administered to a human. In some embodiments, the expression of the therapeutic protein is immunologically inert after the nucleic acid or the circRNA is administered to a human.

[0020] In some embodiments of the nucleic acid disclosed herein have a structure selected from the group consisting of Formulae (I) - (IV) : (I) 5’ - (3’ IF) - (L) n-Z1- (L) n-TI- (5’ IF) -3’ ; (II) 5’ - (3’ IF) - (L) n-TI- (L) n-Z1- (5’ IF) -3’ ; (III) 5’ - (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) -3’ ; (IV) 5’ - (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) -3’ ; and wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI the a translation  initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is the therapeutic protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2.

[0021] In some embodiments, the nucleic acid disclosed herein comprise a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. In some embodiments, the 5’ homology arm, the 3’ homology arm, or both are 15 to 60 nucleotides in length.

[0022] In some embodiments of the nucleic acid disclosed herein, (1) the 3’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NO: 42-52 and 228; or (2) the 5’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229; or both (1) and (2) .

[0023] In some embodiments, nucleic acids provided herein are DNA. In some embodiments, nucleic acids provided herein are RNA. In some embodiments, nucleic acids provided herein are double stranded. In some embodiments, nucleic acids provided herein are single stranded. In some embodiments, nucleic acids provided herein are circular. In some embodiments, nucleic acids provided herein are linear.

[0024] Provided herein are also circRNAs produced by the self-splicing of the RNA encoded by the nucleic acids described herein.

[0025] Provided herein are also vectors comprising the nucleic acids described herein.

[0026] Provided herein are also cells comprising the nucleic acids described herein, the circRNA described herein, or the vector described herein. In some embodiments, the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI. In some embodiments, provided herein are also methods of expressing a protein comprising culturing the cell described herein under conditions and for a sufficient time for the expression of the protein.

[0027] Provided herein are also pharmaceutical compositions comprising the nucleic acids described herein, the circRNA described herein, the vectors described herein, or the cells described herein, and a pharmaceutically acceptable carrier, wherein the circRNA further comprises Z1 operatively linked to the TI, wherein Z1 encodes a therapeutic protein.

[0028] Provided herein are also methods of expressing of a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition described herein.

[0029] Provided herein are also methods of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical compositions described herein. Provided herein are also uses of the pharmaceutical compositions described herein for treating or preventing a disease or disorder.4. Brief Description of Drawings

[0030] FIGs. 1A-1C depict the secondary structure of group II introns. FIG. 1A provides a schematic diagram of group II intron’s general structure in which the sequence elements that provide the long-range interactions essential for their tertiary structure and function are denoted with Greek letters. As shown, a typical group II intron can have six stem-loop structures, referred to as Domains 1-6, or D1-D6. The 6 domains are sequentially arranged and comprise multiple exon binding sequences (EBSs) , such as EBS1, EBS2, and EBS3. These EBS sequences interact, such as complementarily pair, with the intron binding sequences (IBSs) in exon regions (such as IBS1, IBS2, and IBS3) , to trigger self-splicing. Note that a single nucleotide, the δ nucleotide, which is located directly upstream of EBS1 in domain 1 can also pair with IBS3, and the interaction between δ and IBS3 is referred to δ-IBS3 pairing. FIG. 1B provides the secondary structure of exemplary group II intron Cte. FIG. 1C provides the secondary structure of an exemplary synthetic group II intron based on sequence elements of Cte: Cte-Syn1.

[0031] FIGs. 2-5 provide schematic diagrams identifying the exon-intron interactions essential for the near-scarless or scarless splicing of the cRNAzymes disclosed herein.

[0032] FIG. 2 depicts near-scarless splicing based on the interactions between IBS1 and EBS1; and IBS3 and EBS3; optionally also between IBS2 and EBS2. A group II intron with flanking exon sequences E1 and E2 is split into two fragments at the D4 domain, with the 5’ intron fragment and the 3’ intron fragment swapped, and a target sequence inserted between the two fragments. Arrows indicate the interactions between IBS1 and EBS1; IBS2 and EBS2; and IBS3 and EBS3. As shown, the self-splicing of construct produces a circRNA consisting of the target sequence, E1 and E2.

[0033] FIG. 3 depicts near-scarless splicing based on the interactions between IBS1 and EBS1; and the δ nucleotide and IBS3; optionally also between IBS2 and EBS2. A group II intron with flanking exon sequences E1 and E2 is split into two fragments at the D4 domain, with the 5’ intron fragment and the 3’ intron fragment swapped, and a target sequence inserted between the two fragments. Arrows indicate the interactions between IBS1 and EBS1; IBS2 and EBS2; and IBS3 and δ. As shown, the self-splicing of construct produces a circRNA consisting of the target sequence, E1 and E2.

[0034] FIG. 4 depicts scarless splicing based on the interactions between IBS1’ and EBS1’ ; and IBS3’ and EBS3’ . A group II intron is split into two fragments at the D4 domain, with the 5’ intron fragment and the 3’ intron fragment swapped, and a target sequence inserted between the two fragments. Arrows indicate the interactions between IBS1’ and EBS1’ ; and IBS3’ and EBS3’ . As shown, the self-splicing of construct produces a circRNA consisting of the target sequence.

[0035] FIG. 5 depicts scarless splicing based on the interactions between IBS1’ and EBS1’ ; and the δ” nucleotide and IBS3’ . A group II intron is split into two fragments at the D4 domain, with the 5’ intron fragment and the 3’ intron fragment swapped, and a target sequence inserted between the two fragments. Arrows indicate the interactions between IBS1’ and EBS1’ ; and IBS3’ and δ” . As shown, the self-splicing of construct produces a circRNA consisting of the target sequence.

[0036] FIG. 6 provides schematic diagrams showing the re-ligation of the target sequence fragment upon self-splicing of the cRNAzymes provided herein. As shown, the target sequence can be segmented into a 5’ fragment and 3’ fragment, which can be swapped and cloned into the cRNAzyme. Upon self-splicing cRNAzyme, circularization links the 3’ -end of the 5’ target sequence fragment to the 5’ -end of the 3’ target sequence fragment.

[0037] FIGs. 7A (a) -7B (b) provide schematic diagrams of cRNAzymes in which the 3’ target sequence fragment comprises a therapeutic protein-coding sequence (Z1) and the 5’ target sequence fragment comprises a translation initiation sequence (TI) . The cRNAzymes can either omit linkers (FIGs. 7A (a) - (b) ) or include linkers flanking Z1 (FIGs. 7B (a) - (b) ) . Scarless splicing is depicted in FIGs. 7A (a) and 7B (a) . Near-scarless splicing is depicted in FIGs. 7A (b) and 7B (b) .

[0038] FIGs. 8A (a) -8B (b) provide schematic diagrams of cRNAzymes in which the 3’ target sequence fragment comprises a translation initiation sequence (TI) and the 5’ target sequence fragment comprises a therapeutic protein-coding sequence (Z1) . The cRNAzymes can either omit linkers (FIGs. 8A (a) - (b) ) or include linkers flanking TI (FIGs. 8B (a) - (b) ) . Scarless splicing is depicted in FIGs. 8A (a) and 8B (a) . Near-scarless splicing is depicted in FIGs. 8A (b) and 8B (b) .

[0039] FIGs. 9A (a) -9B (b) provide schematic diagrams of cRNAzymes in which the 3’ target sequence fragment comprises a 3’ fragment of TI (TIB) and a therapeutic protein-coding sequence (Z1) and the 5’ target sequence fragment comprises a translation initiation sequence (TI) . The cRNAzymes can either omit linkers (FIGs. 9A (a) - (b) ) or include linkers flanking Z1 (FIGs. 9B (a) - (b) ) . Scarless splicing is depicted in FIGs. 9A (a) and 9B (a) . Near-scarless splicing is depicted in FIGs. 9A (b) and 9B (b) .

[0040] FIGs. 10A-10B provide schematic diagrams of cRNAzymes in which the 3’ target sequence fragment comprises a 3’ fragment of Z1 (Z1B) and the 5’ target sequence fragment comprises a translation initiation sequence (TI) and a 5’ fragment of Z1 (Z1B) . The cRNAzymes can either omit linkers (FIG. 10A) or include linkers flanking TI (FIG. 10B) .

[0041] FIG. 11 shows in vitro evaluation of a luciferase-coding circular RNA for luciferase expression in different cell lines. The data was standardized by first calculating the average expression level, then normalizing by dividing them by the average expression of the reference. Key features of the data interpretation include: Cutoff at 0.1: This value represents 1 / 10th the expression level of the mean expression level of the reference control. Expression values above this cutoff indicate that the IRES is functional and capable of driving luciferase expression. Cutoff at 1.0: This value corresponds to the mean expression level of the reference control. Expression values exceeding this cutoff suggest that the IRES is more effective than the reference control at promoting luciferase expression, indicating a stronger or more efficient IRES activity. FIG. 12 shows an in vitro evaluation of a luciferase-coding circular RNA (circRNA) for luciferase expression across various human and mouse tissues (Human tissues included muscle (with A204 cells) , lung (with A549 cells) , stomach (with AGS cells) , neural tissue (with SY5Y and U87 cells) , colon (with HCT116 cells) , liver (with HepG2 cells) , skin (with HFF1 cells) , and immune tissue (with Jurkat cells) . Mouse tissues included neural tissue (with Neuro-2a cells) , immune tissue (with RAW264.7 cells) , and muscle (with C2C12 cells) ) . The data were processed by calculating the mean expression level of each sample group. To ensure data consistency, the relative expression values were normalized by dividing the measured expression levels by the mean expression level of the reference control. Key features of the data interpretation include: Cutoff at 0.1: This value represents 1 / 10th the expression level of the mean expression level of the reference control. Expression values above this cutoff indicate that the IRES is functional and capable of driving luciferase expression. Cutoff at 1.0: This value corresponds to the mean expression level of the reference control. Expression values exceeding this cutoff suggest that the IRES is more effective than the reference control at promoting luciferase expression, indicating a stronger or more efficient IRES activity. FIG. 13 shows in vitro evaluation of a HGF coding circular RNA for HGF expression and cell  viability across three cell lines: A673 (human rhabdomyosarcoma) , C2C12 (mouse myoblast) , and L6 (rat myoblast) . FIG. 14 shows in vivo evaluation of a HGF coding circular RNA for HGF expression in mouse  tissue.

[0042] Note, depicted in FIGs. 7A (a) , 7B (a) , 8A (a) , 8B (a) , 9A (a) , 9B (a) , 10A, and 10B is scarless splicing wherein the sequence elements within the target sequence serve as E1 and E2. In some embodiments, the 5’ terminal region of the target sequence (e.g., part of TI, Z1, or linker) can serve as E2. In some embodiments, the 3’ terminal region of the target sequence (e.g., part of TI, Z1, or linker) can serve as E1. Optionally, as depicted on FIGs. 7A (b) , 7B (b) , 8A (b) , 8B (b) , 9A(b) , and 9B (b) , (1) extra exon sequence E2 can be included between the 3’ intron fragments and the target sequence; (2) extra exon sequence E1 can be included between the target sequence and the 5’ intron fragments; or both (1) and (2) . As such, E1 and / or E2 remain with the target sequence in the circRNA after self-splicing.5. Detailed Description

[0043] Increasing protein expression is desirable to improve gene expression and / or protein expression and manufacture in therapeutic applications including, but not limited to, protein replacement therapy and vaccinations. The present disclosure overcomes problems associated with current technologies by providing novel synthetic Internal Ribosome Entry Site ( “IRES” ) sequences that can function more efficiently than natural IRES sequences.

[0044] Provided herein are nucleic acids containing the synthetic IRES sequences described herein for use in enhancing protein expression in eukaryotic cells. In some embodiments, the present disclosure provides nucleic acids comprising the synthetic IRES sequences and an expression sequence encoding a protein of interest, for use in improving protein manufacture and production. In some embodiments, the present disclosure provides nucleic acids comprising the IRES sequences and an expression sequence encoding a protein of interest, such as a therapeutic protein which can be suitable for use as a medicament or a vaccine, such as for application in gene therapy and / or genetic vaccination. In some embodiments, the nucleic acid is an RNA polynucleotide. In some embodiments, the RNA is a circular RNA.

[0045] Further, the use of a synthetic IRES sequences described herein provides the opportunity for fine-tuned control and modulation of protein expression, which is also advantageous for the precise control of the ratio of peptide expression in the manufacture and production of multimeric proteins and / or multicistronic cassettes for therapeutic applications, such as antibody production. Together, the nucleic acids provided herein overcome the disadvantages of in the prior art by providing a cost-effective, systematic and straight forward approach for modulating protein expression.

[0046] Before the present disclosure is further described, it is to be understood that the disclosure is not limited to the particular embodiments set forth herein, and it is also to be understood that the terminology used herein is for the purpose of describing particular embodiments, and is not intended to be limiting. 5.1 Definitions

[0047] Unless otherwise defined herein, scientific and technical terms used in the present disclosures shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclatures used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art.

[0048] As used herein in the specification, “a” or “an” may mean one or more. As used herein in the claim (s) , when used in conjunction with the word “comprising, ” the words “a” or “an” may mean one or more than one.

[0049] As used herein, the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or. ” As used herein “another” or “additional” may mean at least a second or more. As used herein, the term “about” is used to indicate that a value includes the inherent variation of  error for the device, the method being employed to determine the value, or the variation that exists among the study subjects. The term “about” encompasses the exact number recited. In some embodiments, “about” means within plus or minus 10%of a given value or range. In certain embodiments, “about” means that the variation is ±5%, ±4%, ±3%, ±2%, ±1%, ±0.5%, ±0.2%, or ±0.1%of the value to which “about” refers. In some embodiments, “about” means that the variation is ±1%, ±0.5%, ±0.2%, or ±0.1%of the value to which “about” refers.

[0050] As used herein, “essentially free, ” in terms of a specified component, is used herein to mean that none of the specified component has been purposefully formulated into a composition and / or is present only as a contaminant or in trace amounts. The total amount of the specified component resulting from any unintended contamination of a composition is therefore well below 0.1%, preferably below 0.05%, and more preferably below 0.01%. Most preferred is a composition in which no amount of the specified component can be detected with standard analytical methods.

[0051] The terms “peptide, ” “polypeptide” and “protein” are used interchangeably herein, and refer to a polymeric form of amino acids comprising at least two or more contiguous amino acids chemically or biochemically modified or derivatized amino acids. The term “peptide” as used herein refers to a class of short polypeptides. The term peptide may refer to a polymer of amino acids (natural or non-naturally occurring) having a length of up to about 100 amino acids. For example, peptides may be about 1 to about 10, about 10 to about 25, about 25 to about 50, about 50 to about 75, about 75 to about 100 amino acid residues in length. In some embodiments, the peptides may be about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, are about 5000 amino acid residues in length.

[0052] The terms “nucleic acid, ” “polynucleotide, ” and “oligonucleotide” are used interchangeably herein and refer to a polymer or oligomer of nucleotides of any length. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases (such as methylated, hydroxymethylated, or glycosylated) , non-natural nucleotides, non-nucleotide building blocks that exhibit similar structure and / or function as natural nucleotides (i.e., “nucleotide analogs” ) , and / or any substrate that can be incorporated into a polymer by DNA or RNA polymerase. The nucleic acids or polynucleotides can be heterogenous or homogenous in composition, can be isolated from naturally occurring sources, or can be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and can exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states. Nucleic acid structures also include, for instance, a DNA / RNA helix, peptide nucleic acid (PNA) , morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 4 (14) : 4503-4510 (2002) and U.S. Patent 5,034,506) , locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97: 5633-5638 (2000) ) , cyclohexenyl nucleic acids (see Wang, Am. Chem. Soc., 122: 8595-8602 (2000) ) , and / or a ribozyme.

[0053] As is understood in the art, a nucleic acid strand is inherently directional, as the carbon atoms in the sugar ring are numbered from 1’ to 5’ and the “5’ -end” has a free hydroxyl (or phosphate) on a 5’ carbon and the “3’ prime end” has a free hydroxyl (or phosphate) on a 3’ carbon. As used herein and understood in the art, a nucleic acid having certain sequence elements “from 5’ to 3’ ” means that these sequence elements are arranged linearly from the 5’ end to the 3’ end of the nucleic acid.

[0054] When referring to a nucleotide sequence or protein sequence, the term “identity” is used to denote similarity between two sequences. Sequence similarity or identity may be determined using standard techniques known in the art, including, but not limited to, the local sequence identity algorithm of Smith &Waterman, Adv. Appl. Math. 2, 482 (1981) , by the sequence identity alignment algorithm of Needleman &Wunsch, J Mol. Biol. 48,443 (1970) , by the search for similarity method of Pearson &Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988) , by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, WI) , the Best Fit sequence program described by Devereux et al., Nucl. Acid Res. 12, 387-395 (1984) , or by inspection. Another algorithm is the BLAST algorithm, described in Altschul et al., J Mol. Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993) . A particularly useful BLAST program is the WU-BLAST-2 program which was obtained from Altschul et al., Methods in Enzymology, 266, 460-480 (1996) ; blast. wustl / edu / blast / README. html. WU-BLAST-2 uses several search parameters, which are  optionally set to the default values. The parameters are dynamic values and are established by the program itself depending upon the composition of the particular sequence and composition of the particular database against which the sequence of interest is being searched; however, the values may be adjusted to increase sensitivity. Further, an additional useful algorithm is gapped BLAST as reported by Altschul et al., (1997) Nucleic Acids Res. 25, 3389-3402. Unless otherwise indicated, percent identity is determined herein using the algorithm available at the internet address: blast. ncbi. nlm. nih. gov / Blast. cgi.

[0055] As used herein, terms “complementary” and “complementarity” refers to the relationship between two nucleic acid molecules having the capacity to form hydrogen bond (s) with one another by either traditional Watson-Crick base-paring or other non-traditional types of pairing. The two DNA / RNA strands with complementary sequences bind to form a duplex that follows the Watson–Crick base-pairing rules: A binds to T (U) with two hydrogen bonds; G binds to C with three hydrogen bonds. The degree of complementarity between two nucleotide sequences can be indicated by the percentage of nucleotides in a nucleotide sequence which can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleotide sequence (e.g., about 50%, about 60%, about 70%, about 80%, about 90%, and 100%complementary) . Two nucleotide sequences are “perfectly complementary” or “100%complementary” if all the contiguous nucleotides of a nucleotide sequence will hydrogen bond with the same number of contiguous nucleotides in a second nucleotide sequence. Two nucleotide sequences are “substantially complementary” if the degree of complementarity between the two nucleotide sequences is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) over a region of at least 8 nucleotides (e.g., at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides) , or if the two nucleotide sequences hybridize under at least moderate, or, in some embodiments high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37℃ in a solution comprising 20%formamide, 5%SSC (150 mM NaCl, 15 mM trisodium citrate) , 50 mM sodium phosphate (pH 7.6) , 5x Denhardt’s solution, 10%dextran sulfate, and 20 mg / ml denatured sheared salmon sperm DNA, followed by washing the filters in 1*SSC at about 37-50℃, or substantially similar conditions, e.g., the moderately stringent conditions described in Sambrook, J., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press; 4th edition (June 15, 2012) . High stringency conditions are conditions that use, for example (1) low ionic strength and high temperature for washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1%sodium dodecyl sulfate (SDS) at 50℃C, (2) employ a denaturing agent during hybridization, such as formamide, for example, 50% (v / v) formamide with 0.1%bovine serum albumin (BSA)  / 0.1%Ficoll / 0.1%polyvinylpyrrolidone (PVP)  / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42℃, or (3) employ 50%formamide, 5xSSC (0.75 M NaCl, 0.075 M sodium citrate) , 50 mM sodium phosphate (pH 6.8) , 0.1%sodium pyrophosphate, 5x Denhardt’s solution, sonicated salmon sperm DNA (50 pg / ml) , 0.1%SDS, and 10%dextran sulfate at 42℃, with washes at (i) 42℃ in 0.2*SSC, (ii) 55℃ in 50%formamide, and (iii) 55℃ in 0.1*SSC (optionally in combination with EDTA) . Additional details and an explanation of stringency of hybridization reactions are provided in, e.g., Sambrook, supra, and Ausubel et al., eds., SHORT PROTOCOLS IN MOLECULAR BIOLOGY, 5th ed., John Wiley &Sons, Inc., Hoboken, N.J. (2002) .

[0056] The term “exogenous, ” as used herein and understood in the art in relation to a protein, gene, nucleic acid, or polynucleotide in a cell or organism refers to a protein, gene, nucleic acid, or polynucleotide that has been introduced into the cell or organism by artificial or natural means; or in relation to a cell, the term refers to a cell that was isolated and subsequently introduced into a cell population or to an organism by artificial or natural means. An exogenous nucleic acid may be from a different organism or cell, or it may be one or more additional copies of a nucleic acid that occurs naturally within the organism or cell. An exogenous cell may be from a different organism, or it may be from the same organism. By way of a non-limiting example, an exogenous nucleic acid is one that is in a chromosomal location different from where it would be in natural cells, or is otherwise flanked by a different nucleic acid sequence than that found in nature.

[0057] The term “operably linked” as used herein and understood in the art with reference to sequence elements in nucleic acid molecules means that these sequence elements (e.g., an intron fragment, a target sequence, a promoter, and a coding sequence) are functionally related to each other. For example, a promoter is operatively linked to a coding sequence if it controls the transcription of the sequence; or a ribosome binding site is operatively linked to a coding sequence if it is positioned so as to permit translation.

[0058] The term “hybridization” or “hybridized” when referring to nucleotide sequences is the association formed between and / or among sequences having complementarity.

[0059] The term “homology” refers to the percent of identity between the nucleic acid residues of two polynucleotides or the amino acid residues of two polypeptides. The correspondence between one sequence and another can be determined by techniques known in the art. For example, homology can be determined by a direct comparison of the sequence information between two polypeptides by aligning the sequence information and using readily available computer programs. Two polynucleotide (e.g., DNA) or two polypeptide sequences are “substantially homologous” to each other when at least about 80%, preferably at least about 90%, and most preferably at least about 95%of the nucleotides, or amino acids, respectively match over a defined length of the molecules, as determined using the methods above.

[0060] The term “Cte” as used herein refers to a group IIB intron C. te. I1, found in the human pathogen Clostridium tetani. (McNeil et al., RNA, 20 (6) : 855-866 (2014) ) .

[0061] The term “Pli” as used herein refers to a group IIB intron Pli, found in the mitochondrial genome of a filamentous brown alga pathogen Pylaiella littoralis. (Zhao and Pyle, Trends in Biochem. Sci., 42.6 (2017) : 470-482) .

[0062] The term “Oi” as used herein refers to a group IIC intron O. i., found in the Oceanobacillus iheyensis. (Toor et al. (2010) , RNA 16, 57–69) .

[0063] The term “LtrB” as used herein refers to a group IIA intron Ll. LtrB, found in the Lactococcus lactis. (Qu, G. et al. (2016) Nat. Struct. Mol. Biol. 23, 549–557) .

[0064] The term “group I intron self-splicing activity” or “group I intron activity” refers to the self-splicing activity derived from a group I intron. The term “group II intron self-splicing activity” or “group II intron activity” refers to the self-splicing activity derived from a group II intron. The term “group IIA intron self-splicing activity” or “group IIA intron activity” refers to the self-splicing activity derived from a group IIA intron. The term “group IIB intron self-splicing activity” or “group IIB intron activity” refers to the self-splicing activity derived from a group IIA intron. The term “group IIC intron self-splicing activity” or “group IIC intron activity” refers to the self-splicing activity derived from a group IIA intron.

[0065] The term “cRNAzymes” as used herein refer to RNAs that are engineered ribozymes with self-splicing activity which, upon self-splicing, forms circRNAs.

[0066] The terms “scar, ” as used herein, refers to the non-target sequence region in the circRNA splicing product. The term “scarless splicing” as used herein refers to the self-splicing of the cRNAzymes which produce circRNAs that do not contain additional sequence elements beyond the target sequence. As such, a “scarless” circRNA contains no scar, meaning that it solely consists of the target sequence. The term “near-scarless splicing” as used herein refers to the self-splicing of the cRNAzymes which produce circRNAs that include no more than 20 nucleotides besides the target sequence. A “near-scarless circRNA” is a circRNA resulting from “near-scarless” self-splicing of a cRNAzyme, which include no more than 20 nucleotides besides the target sequence.

[0067] The pairing between the exon-binding sequences (or “EBSs” ) in the intron and the intron-binding sequences (or “IBSs” ) in the flanking exons are critical during the splicing. The term “EBS” as used herein in connection with a group II intron refers to an exon binding sequence in the intron, which interact (e.g., complementarily pair) with the intron binding sequences ( “IBSs” ) flanking the exon regions to trigger splicing. A group II intron can have multiple EBSs, such as EBS1, EBS2, and EB2s, which interact with IBS1, IBS2, and IBS3, respectively. In addition to EBS1, EBS2 and EBS3, the single nucleotide located directly upstream of EBS1 in domain 1, the “δ nucleotide, ” can also pair with IBS3, and the interaction between δ and IBS3 is referred to δ-IBS3 pairing. As used herein, the term “EBS’ ” refers to an EBS modified to allow scarless splicing of a group II intron. Correspondingly, the sequence elements within the target sequence that pair with the EBS’s are referred to herein as the IBS’s. According, EBS1’ , EBS2’ , and EBS3’ refer to the EBS1, EBS2, and EBS3 sequences that are modified to allow scarless splicing, respectively. IBS1’ , IBS2’ , and IBS3’ refer to the sequences in the target sequence that function as the IBS1, IBS2, and IBS3 in the native exon sequences flanking a group II intron to locate splicing site by interacting with EBS1’ , EBS2’ , and EBS3’ , respectively. Additionally, the “δ” nucleotide” refers to the nucleotide upstream of EBS1’ that pairs with IBS3’ , and the interaction between δ” and IBS3’ is referred to as the δ” -IBS3’ pairing.

[0068] As used herein, the term “E1” and “E2” refer to the exon fragments flanking the target sequence in the cRNAzymes, which remain with the target sequence after self-splicing of the cRNAzymes. E2 is linked to 5’ end of the target sequence and E1 is linked to the 3’ end of the target sequence. Both E1 and E2 comprise an IBS and facilitate the self-splicing of the cRNAzyme. In some embodiments, the E1 and / or E2 can be the exon sequences flanking naturally existing group II intron. In some embodiments, the E1 and / or E2 can be artificial sequences that are engineered into cRNAzymes to, e.g., enhance the accuracy and / or efficiency of self-splicing. In some embodiments, E1, E2, or both can be absent, and part of the target sequence (e.g., linker, TI, or Z1, or combination thereof) can comprise IBS and serve as E1, E2, or both. In scarless splicing, both E1 and E2 are absent, and part of the target sequence (e.g., linker, TI, or Z1, or combination thereof) can comprise IBS and serve as E1 and / or E2.

[0069] The term “in vitro transcription, ” or “IVT, ” refers to versatile method to produce RNA in vitro that uses an RNA polymerase, ribonucleotides, and appropriate buffer conditions to synthesize RNA from a DNA template.

[0070] The term “resulting target sequence, ” as used herein, refers to the target sequence as it is formed in the circRNA upon self-splicing of the RNAs (or cRNAzymes) provided herein.

[0071] The term “expression construct” or “expression cassette, ” as used herein, means a nucleotide sequence that directs translation.

[0072] The terms “coding sequence, ” “coding sequence region, ” “coding region, ” and “CDS, ” as used interchangeably here refer to the portion of a nucleic acid (e.g., a DNA or an RNA) that is or can be translated to protein.

[0073] The terms “reading frame, ” “open reading frame, ” and “ORF” as used interchangeably herein refer to a nucleotide sequence that begins with an initiation codon (e.g., ATG) and, in some embodiments, ends with a termination codon (e.g., TAA, TAG, or TGA) .

[0074] The term “control elements” as used herein refers collectively to promoter regions, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites (IRES) , enhancers, splice junctions, and the like, which collectively provide for the replication, transcription, post-transcriptional processing, and translation of a coding sequence in a recipient cell.

[0075] The term “promoter” as used herein refers to a nucleotide region comprising a DNA regulatory sequence, wherein the regulatory sequence is derived from a gene that is capable of binding to an RNA polymerase and allowing for the initiation of transcription of a downstream (3' direction) coding sequence. It may contain genetic elements at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors, to initiate the specific transcription of a nucleic acid sequence. A promoter that is “operatively positioned, ” “operatively linked” mean that a promoter is in a correct functional location and / or orientation in relation to a nucleic acid sequence to control transcriptional initiation and / or expression of that sequence, which is “under control” and “under transcriptional control” of the promoter.

[0076] The term “enhancer” as used herein means a nucleic acid sequence that, when positioned proximate to a promoter, confers increased transcription activity relative to the transcription activity resulting from the promoter in the absence of the enhancer domain.

[0077] The terms “internal ribosome entry site, ” “internal ribosome entry site sequence, ” “IRES, ” “IRES sequence, ” and “IRES sequence region” as used interchangeably herein refer to cis elements of viral or human cellular RNAs (e.g., messenger RNA (mRNA) and / or circRNAs) that bypass the steps of canonical eukaryotic cap-dependent translation initiation. The IRES sequence can be a natural sequence, i.e., a naturally occurring IRES. The IRES can be synthetic, i.e., non-naturally occurring. The IRES sequence can be an RNA sequence capable of engaging a ribosome (e.g., eukaryotic ribosome) . That is, the IRES attracts a ribosomal (e.g., eukaryotic ribosomal) translation initiation complex and promotes translation initiation. The IRES sequence permits the translation of one or more open reading frames from a circular RNA (e.g., open reading frames that form the expression sequence) .

[0078] The term “vector” or “construct” (sometimes referred to as a gene delivery system or gene transfer “vehicle” ) refers to a vehicle that is used to carry genetic material (e.g., a nucleotide sequence) , which can be introduced into a host cell, where it can be replicated and / or expressed.

[0079] The term “treat” as used herein refers to executing a protocol or plan, which can include administering one or more drugs or active agents to a patient, in an effort to alleviate signs or symptoms of the disease or the recurrence of the disease. Desirable effects of treatment include decreasing the rate of disease progression, ameliorating or palliating the disease state, and remission, increased survival, improved quality of life or improved prognosis. Alleviation or prevention can occur prior to signs or symptoms of the disease or condition appearing, as well as after their appearance. As used herein, a “treatment” does not require complete alleviation of signs or symptoms, and does not require a cure.

[0080] As used herein, the term “therapeutic beneficial” or “therapeutically effective” when used in connection with a therapeutic refers to the property of the therapeutic that promotes or enhances the well-being of the subject. This includes, but is not limited to, a reduction in the frequency, severity, or rate of progression of the signs or symptoms of a disease. For example, treatment of cancer may involve, for example, a reduction in the size of a tumor, a reduction in the invasiveness of a tumor, reduction in the growth rate of the cancer, or a reduction in the rate of metastasis or recurrence. Treatment of cancer can also refer to prolonging survival of a subject with cancer.

[0081] As used herein, the term “pharmaceutical or pharmacologically acceptable” refers to molecular entities and compositions that do not produce an adverse, allergic, or other untoward reaction when administered to an animal, such as a human, as appropriate. For animal (e.g., human) administration, it will be understood that preparations should meet sterility, pyrogenicity, general safety, and purity standards as required, e.g., by the FDA Office of Biological Standards.

[0082] As used herein, the term “pharmaceutically acceptable carrier” includes any and all aqueous biocompatible solvents (e.g., saline solutions, phosphate buffered saline, parenteral vehicles, such as sodium chloride, Ringer's dextrose, etc. ) , antioxidants, preservatives (e.g., antibacterial or antifungal agents, anti-oxidants, chelating agents, and inert gases) , isotonic agents, such like materials and combinations thereof, as would be known to one of ordinary skill in the art. The pH and exact concentration of the various components in a pharmaceutical composition are adjusted according to well-known parameters.

[0083] As used herein, the term “target cell” refers to the cell or type of cells to which the RNAs (or cRNAzymes) or circRNAs disclosed herein are intended to deliver.

[0084] The terms “transfection, ” “transformation, ” and “transduction” are used interchangeably herein and refer to the introduction of one or more exogenous polynucleotides into a host cell by using physical or chemical methods.

[0085] As used herein, the term “subject” as used herein refers to any animal (e.g., a mammal) , including, but not limited to, humans, non-human primates, canines, felines, rodents, and the like, which is to be the recipient of a particular treatment. A subject can be a human. A subject can have a particular disease or condition.

[0086] Nomenclature for nucleotides, nucleic acids, nucleosides, and amino acids used herein is consistent with International Union of Pure and Applied Chemistry (IUPAC) standards (see, e.g., bioinformatics. org / smsylupac. html) . Exemplary genes and polypeptides are described herein with reference to GenBank numbers, GI numbers and / or SEQ ID NOS. It is understood that one skilled in the art can readily identify homologous sequences by reference to sequence sources, including but not limited to Uniprot (https:  / / www. uniprot. org / ) , GenBank (ncbi. nlm. nih. gov / genbank / ) and EMBL (embl. org / ) .

[0087] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range. 5.2 Cap-Independent Translation Initiation

[0088] In some embodiments, provided herein is a nucleic acid comprising an engineered translation initiation element (TI) comprising a non-naturally occurring IRES described herein. The nucleic acids can be double-stranded or single-stranded. The nucleic acids can be circular or linear. The nucleic acids can be DNA or RNA. In some embodiments, the nucleic acids are linear RNAs comprising the IRES described herein that can mediate cap-independent translation initiation. In some embodiments, the nucleic acids are single stranded linear RNA. In some embodiments, the nucleic acids are mRNA. In some embodiments, the nucleic acids are circular RNAs comprising the IRES described herein that can mediate cap-independent translation initiation. In some embodiments, the nucleic acids are single stranded circular RNA. In some embodiments, the nucleic acids are double stranded circular RNA. In some embodiments, the nucleic acids are DNAs that encode these RNAs. In some embodiments, the nucleic acids are double stranded DNA. In some embodiments, the nucleic acids are single stranded DNA.

[0089] Translation initiation of mRNAs in eukaryotic cells is a complex process that involves the concerted interaction of numerous factors (Pain (1996) Eur. J. Biochem. 236, 747-771) . For most mRNAs, the first step is the recruitment of ribosomal 40S subunits onto the mRNA at or near the capped 5′end. Association of 40S to mRNA is greatly facilitated by the cap-binding protein complex eIF4F. Factor eIF4F is composed of three subunits: the RNA helicase eIF4A, the cap-binding protein eIF4E, and the multi-adaptor protein eIF4G which acts as a scaffold for the proteins in the complex and has binding sites for eIF4E, eIF4a, eIF3, and poly (A) binding protein. 5.3 Synthetic IRESs

[0090] In some embodiments, provided herein are IRES sequences identified as IRES-n, wherein n is an integer between 001 and 351, as shown in Table 26 (SEQ ID NOs: 264-614) . In some embodiments, provided herein are variants or subsequences of these IRES sequences, wherein these variants or subsequences can also mediate cap-independent translation initiation. In some embodiments, variants of the IRES sequences provided herein have the conserved cores of the reference IRES sequences. In some embodiments, variants of the IRES sequences provided herein have the same structural motif as the reference IRES sequences. The variant IRES sequences can have nucleotide changes that do not disrupt the core functional sequences or the overall structure. For example, the variant IRES sequences can have nucleotide changes, or point mutations. The nucleotide changes can be additions, deletions, or substitutions, or combinations thereof. Minor insertions or deletions outside of the critical regions can be tolerated, provided that they do not significantly alter the RNA’s secondary structure. In some embodiments, provided herein are also subsequences of the IRES sequences. Some IRES sequences, shorter subsequences that contain the essential elements for ribosome binding and initiation can still function independently. In some embodiments, deletion of certain domains within the IRES does not abolish activity. For example, certain subsequences with verified activity are provided in Table 27 (SEQ ID NOs: 615-622) , which represent exemplary subsequences of IRES-200, IRES-249, IRES-129, IRES-091, IRES-343, IRES-282, IRES-007, and IRES-314, respectively.

[0091] Accordingly, in some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are at least 90%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are at least 95%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are at least 98%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are at least 99%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614.

[0092] In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are at least 90%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are at least 95%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are at least 98%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are at least 99%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.

[0093] In some embodiments, provided herein is IRES-200, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 463, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 463. The subsequence of SEQ ID NO: 463 can be SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 615. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 463. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 615.

[0094] In some embodiments, provided herein is IRES-249, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 512, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 512. The subsequence of SEQ ID NO: 512 can be SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 616. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 512. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 616.

[0095] In some embodiments, provided herein is IRES-129, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 392, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 392. The subsequence of SEQ ID NO: 392 can be SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 617. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 392. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 617.

[0096] In some embodiments, provided herein is IRES-091, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 354, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 354. The subsequence of SEQ ID NO: 354 can be SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 618. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 354. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 618.

[0097] In some embodiments, provided herein is IRES-343, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 606, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 606. The subsequence of SEQ ID NO: 606 can be SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 619. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 606. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 619.

[0098] In some embodiments, provided herein is IRES-282, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 545, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 545. The subsequence of SEQ ID NO: 545 can be SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 620. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 545. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 620.

[0099] In some embodiments, provided herein is IRES-007, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 270. The subsequence of SEQ ID NO: 270 can be SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 621. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 270. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 621.

[0100] In some embodiments, provided herein is IRES-314, as well as its variants and subsequences that maintain its activity to mediate cap-independent translation initiation. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 577, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a subsequence of SEQ ID NO: 577. The subsequence of SEQ ID NO: 577 can be SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are at least 90%identical to SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are at least 95%identical to SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are at least 98%identical to SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are at least 99%identical to SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are 100%identical to SEQ ID NO: 622. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of a subsequence of SEQ ID NO: 577. In some embodiments, provided herein are IRES sequences that are at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a reverse complementary sequence of SEQ ID NO: 622. 5.4 Naturally occurring IRES

[0101] In some embodiments, one or more synthetic IRES sequences provided herein can be combined with a natural IRES sequence to form a translation initiation site. A multitude of natural IRES sequences are available and include sequences derived from or isolated from a wide variety of viruses, such as from leader sequences of piconaviruses such as the encephalomyocarditis virus (EMCV) UTR, the polio leader sequence, the hepatitis A virus leader sequence, the hepatitis C virus IRES, human rhinovirus type 2 IRES, an IRES element from the foot and mouth disease virus, a giardiavirus IRES, and the like.

[0102] In some embodiments, the natural IRES sequence is isolated or derived from an IRES sequence of Taura syndrome virus, Tiiatoma virus, Theiler's encephalomyelitis virus, Simian Virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, Reticuloendotheliosis virus, Human poliovirus 1, Plautia stall intestine virus, Kashmir bee virus, Human rhinovirus 2, Homalodisca coagulata virus-1, Human Immunodeficiency Virus type 1, Himetobi P virus, Hepatitis C virus, Hepatitis A virus, Hepatitis A Virus HA 16, Hepatitis GB virus, Foot and mouth disease virus, Human enterovirus 71, Equine rhinitis virus, Ectrapis obliqua picoma-like virus, Encephalomyocarditis virus, Drosophila C Virus, Human coxsackievirus B3, Crucifer tobamovirus, Cricket paralysis virus, Bovine viral diarrhea virus 1, Black Queen Cell Virus, Aphid lethal paralysis virus, Avian encephalomyelitis virus, Acute bee paralysis virus, Hibiscus chlorotic ringspot virus, Classical swine fever virus, tobacco etch virus, turnip crinkle virus, EMCV-A, EMCV-B, EMCV-Bf, EMCV-Cf, EMCV pEC9, Picobimavirus, HCV QC64, Human Cosavirus E / D, Human Cosavirus F, Human Cosavirus JMY, Rhinovirus NAT001, HRV14, HRV89, HRVC-02, HRV-A21, Salivirus A SHI, Salivirus FHB, Salivirus NG-J1, Human Parechovirus 1, Crohivirus B, Yc-3, Rosavirus M-7, Shanbavirus A, Pasivirus A, Pasivirus A 2, Echovirus E14, Human Parechovirus 5, Aichi Virus, Phopivirus, CVA10, Enterovirus C, Enterovirus D, Enterovirus J, Human Pegivirus 2, GBV-C GT110, GBV-C K1737, GBV-C Iowa, Pegivirus A 1220, Pasivirus A 3, Sapelovirus, Rosavirus B, Bakunsa Virus, Tremovirus A, Swine Pasivirus 1, PLV-CHN, Pasivirus A, Sicinivirus, Hepacivirus K, Hepacivirus A, BVDV1, Border Disease Virus, BVDV2, CSFV-PK15C, SF573 Dicistravirus, Hubei Picoma-like Virus, CRPV, Salivirus A BN5, Salivirus A BN2, Salivirus A 02394, Salivirus A GUT, Salivirus A CH, Salivirus A SZ1, Salivirus FHB, Coxsackievirus (e.g., CVA3, CVA12, CVB1, CVB3, CVB5) , Echovirus 7, Enterovirus A71, and / or EV24.

[0103] In some embodiments, a natural IRES sequence is isolated or derived from a eukaryotic IRES element selected from Human FGF2, Human SFTPA1, Human AML1 / RUNX1, Drosophila antennapedia, Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAPl, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIFl alpha, Human n. myc, Mouse Gtx, Human p27kipl, Human PDGF2 / c-sis, Human p53, Human Pim-1, Mouse Rbm3, Drosophila reaper, Canine Scamper, Drosophila Ubx, Human UNR, Mouse UtrA, Human VEGF-A, Human XIAP, Drosophila hairless, S. cerevisiae TFIID, and S. cerevisiae YAP1.

[0104] In some embodiments, a natural IRES sequence is an endogenous IRES sequence, which is derived from or isolated from Homo sapiens. In some embodiments, a natural IRES sequence is an endogenous IRES sequence, which is derived from or isolated from human tissue or human sample.

[0105] In some embodiments, an IRES sequence is isolated or derived from a cellular IRES element selected from AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, ATlR varl, ATlR_var2, ATlR_var3, ATlR_var4, BAGl_p36delta236nt, BAGl_p36, BiP_-222_-3, C-IAP1 285-1399, c IAP1 13 13-1462, c-jun, Cat-l_224, CCND1, eIF4GI-ext, eIF4GII, eIF4GII-long, FGF1A, FMR1, Gtx-l33-l4l, Gtx-l-l66, Gtx-l-l20, Gtx-l-l96, HAP4, HIFla, hSNMl, HsplOl, hsp70, hsp70, Hsp90, IGF2_leader2, L-myc, MNT 75-267, MNT 36-160, MTG8a, MYB, MYT2 997-1 152, NRF_-653_-l7, NtHSFl, ODC1, p27kipl, p53_l28-269, PDGF2 / c-sis, PITSLRE_p58, Rbm3 , reaper, Scamper, TFIID, TIF4631, Ubx_l-966, Ubx_373-96l, UNR, Ure2, XIAP 5-464, XIAP 305-466, YAP1, (GAAA) l6, (PPTl9) 4 and XI.

[0106] In some embodiments, an IRES sequence is isolated or derived from a viral IRES element selected from ABPV IGRpred, AEV, ALPV IGRpred, BQCV IGRpred, BVDV1 1-385, BVDV1 29-391, CrPV 5NCR, CrPV IGR, crTMV_IRESmp228, CSFV, DCV IGR, EoPV_5NTR, ERBV_l62-920, EV7l_l-748, FMDV type C, GBV-A, GBV-C, HAV HM175, HiPVJGRpred, HIV-l, HoCVlJGRpred, IAPVJGRpred, idefix, KBV IGRpred, PSIV IGR, PV typel Mahoney, PV_type3_Leon, REV-A, RhPV 5NCR, RhPV IGR , SINV l IGRpred, SV40 661-830, TMEV, TMV_UI_IRESmp228, TRV 5NTR, TrV IGR, TSV, and IGR.

[0107] In some embodiments, the synthetic IRES sequence of the present disclosure is combined with a sequence isolated or derived from a natural IRES sequence to further improve the expression of a protein of interest. 5.5 Translation Initiation (TI) Element

[0108] Provided herein are translation initiation (TI) elements that can be used to drive the cap-independent expression of a protein of interest. The TI elements comprise one or more synthetic IRES sequences disclosed herein, including those exemplified in Table 26 and their variants and subsequences (e.g., sequences in Table 27) . In some embodiments, the TI element provided here can comprise at least two IRES sequences described herein. In some embodiments, the TI elements provided herein also comprise a natural IRES sequence or a fragment thereof. In some embodiments, a complete natural IRES sequence is combined with a synthetic IRES sequence disclosed herein. In some embodiments, a fragment of a natural IRES sequence is combined with a synthetic IRES sequence disclosed herein. In some embodiments, the TI elements provided herein also comprise an additional synthetic IRES sequence known in the art, for example, one disclosed in PCT Application No. PCT / CN2022 / 095949.

[0109] The TI elements provided herein can initiate cap-independent translation when operably linked to a therapeutic protein-coding sequence. In some embodiments, the TI elements provided herein are optimized for initiating translation in a circular RNA. In some embodiments, the TI elements provided herein are optimized for initiating translation in a linear RNA. In some embodiments, TI elements provided herein possess high translation efficiency due to the optimized structure of the synthetic IRES sequences described herein, which enhances the ribosome's access and interaction, thereby facilitating robust protein production. The optimization extends to minimizing secondary structures and adjusting for codon preferences, which collectively boost the initiation rate of translation. The sustainability of protein expression is ensured through the stability of these synthetic IRES sequences, which minimize metabolic burden and cellular stress, thereby maintaining the expression over extended periods. Importantly, the expressed proteins achieve correct folding and functional structures, ensuring the bioactivity and integrity of the produced proteins. The carefully engineered TI sequences minimize off-target effects and are scalable, making them ideal for a wide range of biotechnological applications, from therapeutic protein production to vaccine development.

[0110] In some embodiments, the TI elements provided herein can sustain the protein expression in a cell for at least about 1 hr to about 30 days, or at least about 2 hrs, 6 hrs, 12 hrs, 18 hrs, 24 hrs, 2 days, 3, days, 4 days, 5 days, 6 days, 7 days, 8 days, 9 days, 10 days, 11 days, 12 days, 13 days, 14 days, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 21 days, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 60 days, or longer or any time therebetween. In some embodiments, the TI elements provided herein can sustain the protein expression for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days in vivo. In some embodiments, the TI elements provided herein can sustain the protein expression for at least 3 days. In some embodiments, the TI elements provided herein can sustain the protein expression for at least 5 days. In some embodiments, the TI elements provided herein can sustain the protein expression for at least 7 days. In some embodiments, the TI elements provided herein can sustain the protein expression for at least 14 days.

[0111] In some embodiments, the TI elements provided herein help achieve correct folding and functional structures of the expressed protein, ensuring their bioactivity and integrity. As misfolded protein or other incorrect translation products can elicit immune response in a host organism, enhancing the accuracy of translation and structure of the protein products can avoid unwanted activation of the host immune system. Accordingly, in some embodiments, the expressed protein (e.g., a therapeutic protein) is immunologically inert to the host. In some embodiments, the expressed protein (e.g., a therapeutic protein) is immunologically inert after the nucleic acid comprising the therapeutic protein-coding sequence operably linked to a TI element disclosed herein is administered to a human.

[0112] Assays described herein or otherwise known in the art can be used to confirm the efficiency and accuracy of the translation driven by the TI elements disclosed herein. 5.6 Therapeutic protein-coding sequence (Z1)

[0113] Provided herein are nucleic acids comprising the TI elements described herein. In some embodiments, the TI element is operably linked to a therapeutic protein-coding sequence (Z1) . The Z1 sequence can encode any protein for which expression is desired. In some embodiments, the Z1 sequence can encode any protein, e.g., a protein selected from a functional protein, an antigenic protein, a signal peptide, a tag protein, and the like. In some embodiments, the Z1 sequence encodes a therapeutic protein.

[0114] In some embodiments, the therapeutic protein is a polypeptide, a protein, an enzyme or an antibody. In some embodiments, the therapeutic protein comprises one or more polypeptide, protein, enzyme, antibody, or a combination thereof. In some embodiments, the protein or enzyme is associated with a genetic disease (e.g., a disease in which a genetic alteration (e.g., mutation) and / or protein dysregulation plays a role in the initiation, development, and / or manifestation of the disease) . In some embodiments, the therapeutic protein is an antigen or agent which can induce  vaccine-induced memory and / or enable the immune system to act quickly to protect the body from any of these agents in later encounters. In some embodiments, the therapeutic protein is an antigen or agent which can stimulate the body's immune system to recognize the agent as a foreign invader, generate antibodies against it, destroy it and develop a memory of it. In some embodiments, the polypeptide or protein resembles a weakened or non-viable form of a disease-causing agent (e.g., an infectious agent such as pathogen) , which can be selected from a microorganism, such as a bacterium, virus, fungus, parasite, or one or more components of such microorganism, such as toxins, proteins (e.g., surface proteins) , and / or cell walls. In some embodiments, the therapeutic protein is an antigen or agent which can stimulate the body's immune system to recognize the antigen or agent, generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop an immunological memory of the antigen or agent. In some embodiments, the therapeutic protein is an antigen or agent which can induce and / or strengthen vaccine-induced memory and / or enable the immune system to respond rapidly and effectively to the antigen or agent in later encounters.

[0115] In some embodiments, the therapeutic protein is derived from an infectious agent. In some embodiments, the infectious agent is selected from a member of the group consisting of strains of viruses and strains of bacteria.

[0116] In any of the embodiments provided herein, the infectious agent is a strain of virus selected from the group consisting of adenovirus; herpes simplex, type 1; herpes simplex, type 2; encephalitis virus, papillomavirus, varicella-zoster virus; Epstein-Barr virus; human cytomegalovirus; human herpes virus, type 8; human papillomavirus; BK virus; JC virus; smallpox; hepatitis B virus; human bocavirus; parvovirus B19; human astrovirus; Norwalk virus; coxsackievirus; hepatitis A virus; poliovirus; rhinovirus; severe acute respiratory syndrome virus; hepatitis C virus; yellow fever virus; dengue virus; West Nile virus; rubella virus; hepatitis E virus; human immunodeficiency virus (HIV) ; Guanarito virus; Junin virus; Lassa virus; Machupo virus; Sabiá virus; Crimean-Congo hemorrhagic fever virus; Ebola virus; Marburg virus; measles virus; mumps virus; parainfluenza virus; respiratory syncytial virus (RSV) ; human metapneumovirus; Hendra virus; Nipah virus; rabies virus; hepatitis D; rotavirus; orbivirus; coltivirus; Banna virus; human enterovirus; hantavirus; corona virus, severe acute respiratory syndrome (SARS) -associated coronavirus (SARS-CoV) , SARS-CoV-2 virus (COVID-19 associated) ; Middle East respiratory syndrome corona virus; Japanese encephalitis virus; vesicular exanthernavirus; Eastern equine encephalitis; and influenza virus. In some embodiments, the infectious agent is a strain of bacteria selected from Tuberculosis (Mycobacterium tuberculosis) , clindamycin-resistant Clostridium difficile, fluoroquinolon-resistant Clostridium difficile, methicillin-resistant Staphylococcus aureus (MRSA) , multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistance Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA) .

[0117] In some embodiments, the infectious agent is associated with humans, non-human primates, or other animals, such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.

[0118] In some embodiments, the therapeutic protein is an antibody. Antibodies encoded by the Z1 sequence include, but are not limited to, monoclonal antibodies, polyclonal antibodies, recombinantly produced antibodies, human antibodies, humanized antibodies, chimeric antibodies, synthetic antibodies, tetrameric antibodies comprising two heavy chain and two light chain molecules, antibody light chain monomers, antibody heavy chain monomers, antibody light chain dimers, antibody heavy chain, antibody heavy chain dimers, antibody light chain-heavy chain pairs, intrabodies, heteroconjugate antibodies, monovalent antibodies, antigen-binding fragments of full-length antibodies, and fusion proteins of the above. Such antigen-binding fragments include, but are not limited to, single-domain antibodies (variable domain of heavy chain antibodies (VHHs) or nanobodies) , Fabs, F (ab’ ) 2S, and scFvs (single-chain variable fragments) .

[0119] In some embodiments, the present compositions and methods are used to produce therapeutic proteins. In some embodiments, the is selected from the group consisting of Abarelix, Abatacept, Abciximab, Adalimumab, Aflibercept, Agalsidase beta, Albiglutide, Aldesleukin, Alefacept, Alemtuzumab, Alglucerase, Alglucosidase alfa, Alirocumab, Aliskiren, Alpha-1-proteinase inhibitor, Alteplase, Anakinra, Ancestim, Anistreplase, Anthrax immune globulin human, Antihemophilic Factor, Antithrombin Alfa, Antithrombin III human, Antithymocyte globulin, Anti-thymocyte Globulin (Equine) , Anti-thymocyte Globulin (Rabbit) , Aprotinin, Arcitumomab, Asfotase Alfa, Asparaginase, Asparaginase erwinia chrysanthemi, Atezolizumab, Autologous cultured chondrocytes, Basiliximab, Becaplermin, Belatacept, Belimumab, Beractant, Bevacizumab, Bivalirudin, Blinatumomab, Botulinum Toxin Type A, Botulinum Toxin Type B, Brentuximab vedotin, Brodalumab, Buserelin, Cl Esterase Inhibitor (Human) , Cl Esterase Inhibitor, Canakinumab, Canakinumab, Capromab, Certolizumab pegol, Cetuximab, Choriogonadotropin alfa, Chorionic Gonadotropin (Human) , Chorionic Gonadotropin, Coagulation factor IX, Coagulation factor Vila, Coagulation factor X human, Coagulation Factor XIII A-Subunit, Collagenase, Conestat alfa, Corticotropin, Cosyntropin, Daclizumab, Daptomycin, Daratumumab, Darbepoetin alfa, Defibrotide, Denileukin diftitox, Denosumab, Desirudin, Dinutuximab, Dornase alfa, Drotrecogin alfa, Dulaglutide, Eculizumab, Efalizumab, Efmoroctocog alfa, Elosulfase alfa, Elotuzumab, Enfuvirtide, Epoetin alfa, Epoetin zeta, Eptifibatide, Etanercept, Evolocumab, Exenatide, Factor IX Complex (Human) , Fibrinogen Concentrate (Human) , Fibrinolysin aka plasmin, Filgrastim, Filgrastim-sndz, Follitropin alpha, Follitropin beta, Galsulfase, Gastric intrinsic factor, Gemtuzumab ozogamicin, Glatiramer acetate, Glucagon recombinant, Glucarpidase, Golimumab, Gramicidin D, Hepatitis A Vaccine, Hepatitis B immune globulin, Human calcitonin, Human Clostridium tetani toxoid immune globulin, Human rabies virus immune globulin, Human Rho (D) immune globulin, Human Serum Albumin, Human Varicella-Zoster Immune Globulin, Hyaluronidase, Hyaluronidase, Ibritumomab, Ibritumomab tiuxetan, Idarucizumab, Idursulfase, Imiglucerase, Immune Globulin Human, Infliximab, Insulin aspart, Insulin Beef, Insulin Degludec, Insulin detemir, Insulin Glargine, Insulin glulisine, Insulin Lispro, Insulin Pork, Insulin Regular, Insulin Regular, Insulin porcine, Insulin isophane, Interferon Alfa-2a, Interferon alfa-2b, Interferon alfacon-1, Interferon alfa-nl, Interferon alfa-n9, Interferon beta-la, Interferon beta-lb, Interferon gamma-lb, Intravenous Immunoglobulin, Ipilimumab, Ixekizumab, Laronidase, Lenograstim, Lepirudin, Leuprolide, Liraglutide, Lucinactant, Lutropin alfa, Lutropin alfa, Mecasermin, Menotropins, Mepolizumab, Epoetin beta, Metreleptin, Muromonab, Natalizumab, alpha interferon, Necitumumab, Nesiritide, Nivolumab, Obiltoxaximab, Obinutuzumab, Ocriplasmin, Ofatumumab, Omalizumab, Oprelvekin, OspA lipoprotein, Oxytocin, Palifermin, Palivizumab, Pancrelipase, Panitumumab, Pembrolizumab, Pertuzumab, Poractant alfa, Pramlintide, Preotact, Protein S human, Ramucirumab, Ranibizumab, Rasburicase, Raxibacumab, Reteplase, Rilonacept, Rituximab, Romiplostim, Sacrosidase, Salmon Calcitonin, Sargramostim, Satumomab Pendetide, Sebelipase alfa, Secretin, Secukinumab, Sermorelin, Serum albumin, Serum albumin iodonated, Siltuximab, Simoctocog Alfa, Sipuleucel-T, Somatotropin Recombinant, Somatropin recombinant, Streptokinase, Sulodexide, Susoctocog alfa, Taliglucerase alfa, Teduglutide, Teicoplanin, Tenecteplase, Teriparatide, Tesamorelin, Thrombomodulin alfa, Thymalfasin, Thyroglobulin, Thyrotropin Alfa, Thyrotropin Alfa, Tocilizumab, Tositumomab, Trastuzumab, Tuberculin Purified Protein Derivative, Turoctocog alfa, Urofollitropin, Urokinase, Ustekinumab, Vasopressin, Vedolizumab, and Velaglucerase alfa.

[0120] A therapeutic protein can also be a protein related to enzyme replacement, such as Agalsidase beta, Agalsidase alfa, Imiglucerase, Taligulcerase alfa, Velaglucerase alfa, Alglucerase, Sebelipase alpha, Laronidase, Idursulfase, Elosulfase alpha, Galsulfase, Alglucosidase alpha, Factor VIII, C3 inhibitor, Hurler and Hunter corrective factors. In some embodiments, a POI is a nucleosidase, an NAD+ nucleosidase, a hydrolase, a glycosylase, a glycosylase that hydrolyzes N-glycosyl compounds, an NAD+ glycohydrolase, an NADase, a DPNase, a DPN hydrolase, an NAD hydrolase, a diphosphopyridine nucleosidase, a nicotinamide adenine dinucleotide nucleosidase, an NAD glycohydrolase, an NAD nucleosidase, or a nicotinamide adenine dinucleotide glycohydrolase.

[0121] In some embodiments, the therapeutic protein can be selected from the group consisting of RPE65, SM1, BCR-ABL, Factor IX, Factor VIII, CFTR, DMD, ABCA4, C9orf72, SOD1, TARDBP, P53, KRAS, EGFR, LCA, HBB, Huntingtin, Parkin (PRKN) , PINK1, LRRK2, ABCG5 / ABCG8, and MYO7A.

[0122] In some embodiments, the resulting target sequence disclosed herein can be codon-optimized, for example, via any codon-optimization technique known to one of skill in the art (see, e.g., review by Quax et al., 2015, Mol. Cell 59: 149-161) . A codon optimized sequence can be one in which codons in a polynucleotide encoding a therapeutic protein have been substituted in order to increase the expression, stability and / or activity of the therapeutic protein. Factors that influence codon optimization include, but are not limited to one or more of: (i) variation of codon biases between two or more organisms or genes or synthetically constructed bias tables, (ii) variation in the degree of codon bias within an organism, gene, or set of genes, (iii) systematic variation of codons including context, (iv) variation of codons according to their decoding tRNAs, (v) variation of codons according to GC %, either overall or in one position of the triplet, (vi) variation in degree of similarity to a reference sequence for example a naturally occurring sequence, (vii) variation in the codon frequency cutoff, (viii) structural properties of mRNAs transcribed from the DNA sequence, (ix) prior knowledge about the function of the DNA sequences upon which design of the codon substitution set is to be based, and / or (x) systematic variation of codon sets for each amino acid. In some embodiments, a codon optimized polynucleotide can minimize ribozyme collisions and / or limit structural interference between the expression sequence and the IRES.

[0123] For exemplary purposes, in some embodiments, TI has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 222-225. In some embodiments, TI has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence selected of SEQ ID NO: 222. In some embodiments, TI has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence selected of SEQ ID NO: 223. In some embodiments, TI has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence selected of SEQ ID NO: 224. In some embodiments, TI has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence selected of SEQ ID NO: 225.

[0124] In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 107-112, 214-221 and 258-259. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 107. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 108. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 109. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 110. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 111. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 112. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 214. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 215. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 216. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 217. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 218. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 219. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 220. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 221. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 258. In some embodiments, Z1 has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to the nucleotide sequence of SEQ ID NO: 259.

[0125] In some embodiments, Z1 comprises a nucleic acid sequence encoding the amino acid sequence selected from the group consisting of: (a) SEQ ID NO: 113; (b) SEQ ID NO: 114; (c) SEQ ID NO: 115; (d) SEQ ID NO: 116; (e) SEQ ID NO: 117; and (f) SEQ ID NO: 118.

[0126] In some embodiments, the IRES sequence and therapeutic protein-coding sequence are connected by a linker sequence (L) . In some embodiments, the linker sequence is 3-300 nucleic acid residues in length. In some embodiments the linker sequence is about 3-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-125, 125-150, 150-175, 175-200, 200-225, 225-250, 250-275 or 275-300 nucleic acid sequences in length. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-100. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-50. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-20. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-10. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is 1, 2, 3, 4 or 5. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is 1, 2 or 3. In some embodiments, the linker sequence is about 3 nucleic acid residues in length.

[0127] In some embodiments, the linker sequence comprises the nucleic acid sequence of RCC, wherein R is a guanine or an adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCC, wherein R is a guanine or an adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCCRCC, wherein R is a guanine or an adenine. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 226. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 227. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 248.

[0128] In some embodiments, the linker comprises a nucleic acid sequence that encodes a 5’ UTR, 3’ UTR, poly-A sequence, polyA-C sequence, poly-C sequence, poly-U sequence, poly-G sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulation elements (e.g., a Kozak sequence) , a protein binding site (e.g., PTBP1 or HUR) a non-natural nucleotide, or a non-nucleotide chemical-linker sequence.

[0129] In some embodiments, the nucleic acids provided herein comprise one expression cassette having a TI element and an Z1 sequence. In some embodiments, the nucleic acids provided herein comprise more than one expression cassettes, each having a TI element and an Z1 sequence, wherein at least one of the TI elements comprise an IRES sequence disclosed herein. In some embodiments, the nucleic acids provided herein comprise two, three, four, five, six, or more expression cassettes. In some embodiments, the nucleic acids provided herein comprise two expression cassettes. In some embodiments, the nucleic acids provided herein comprise three expression cassettes. In some embodiments, the nucleic acids provided herein comprise four expression cassettes. In some embodiments, the nucleic acids provided herein comprise five expression cassettes. In some embodiments, the nucleic acids provided herein comprise six expression cassettes.

[0130] Different expression cassettes on the same nucleic acid can be separated by, for example, a cleavable sequence, such as a 2A element. A 2A element, as understood in the art, encoding self-cleaving short peptides (about 20 amino acids) that provide a mechanism for subsequent separation of equimolarly produced polypeptides of interest. Illustrative 2A self-cleaving peptides include P2A (SEQ ID NO: 623) , E2A (SEQ ID NO: 624) , F2A (SEQ ID NO: 625) , and T2A (SEQ ID NO: 626) . 5.7 Vectors and methods of production

[0131] Provided herein is a nucleic acid comprising an engineered TI comprising a non-naturally occurring IRES described herein. In some embodiments, provided herein are also vectors that comprise the nucleic acids disclosed herein.

[0132] As used herein and understood in the art, the term “vector” or “construct” (sometimes referred to as a gene delivery system or gene transfer “vehicle” ) refers to a vehicle that is used to carry genetic material (e.g., a nucleotide sequence) , which can be introduced into a host cell, where it can be replicated and / or expressed.

[0133] Many vectors can be used, including, for example, expression vectors, plasmids, phage vectors, viral vectors, episomes and artificial chromosomes, which can include selection sequences or markers operable for stable integration into a host cell’s chromosome. A “plasmid” is a common type of a vector, which is an extra-chromosomal DNA molecule separate from the chromosomal DNA that is capable of replicating independently of the chromosomal DNA. In certain cases, it is circular and double-stranded. Exemplary artificial chromosomes such as yeast artificial chromosome (YAC) , bacterial artificial chromosome (BAC) , or P1-derived artificial chromosome (PAC) . Exemplary bacteriophages include such as lambda phage or M13 phage. Examples of categories of animal viruses useful as vectors include, without limitation, retrovirus (including lentivirus) , adenovirus, adeno-associated virus (AAV) , herpesvirus (e.g., herpes simplex virus) , poxvirus, baculovirus, papillomavirus, and papovavirus (e.g., SV40) . Examples of expression vectors are pClneo vectors (Promega) for expression in mammalian cells; pLenti4 / V5-DESTTM, pLenti6 / V5-DESTTM, and pLenti6.2 / V5-GW / lacZ (Invitrogen) for lentivirus-mediated gene transfer and expression in mammalian cells. Exemplary AAV serotypes include AAV1, AAV2, AAV4, AAV5, AAV6, AAV9 AAV8, and AAV9. In some embodiments, provided herein are viral vectors encoding the RNAs (cRNAzymes) provided herein. In some embodiments, provided herein are AAVs encoding the RNAs (cRNAzymes) provided herein.

[0134] In some embodiments, the vector is an episomal vector or a vector that is maintained extrachromosomally. As used herein, the term “episomal” refers to a vector that is able to replicate without integration into host’s chromosomal DNA and without gradual loss from a dividing host cell also meaning that said vector replicates extrachromosomally or episomally. The vector is engineered to harbor the sequence coding for the origin of DNA replication or “ori” from a lymphotrophic herpes virus or a gamma herpesvirus, an adenovirus, SV40, a bovine papilloma virus, or a yeast, specifically a replication origin of a lymphotrophic herpes virus or a gamma herpesvirus corresponding to oriP of EBV. In some embodiments, the lymphotrophic herpes virus may be Epstein Barr virus (EBV) , Kaposi's sarcoma herpes virus (KSHV) , Herpes virus saimiri (HS) , or Marek's disease virus (MDV) . Epstein Barr virus (EBV) and Kaposi's sarcoma herpes virus (KSHV) are also examples of a gamma herpesvirus. Typically, the host cell comprises the viral replication transactivator protein that activates the replication.

[0135] Additionally, vectors can include one or more selectable marker genes and appropriate expression control sequences. Selectable marker genes that can be included, for example, provide resistance to antibiotics or toxins, complement auxotrophic deficiencies, or supply critical nutrients not in the culture media. “Expression control sequences, ” “control elements, ” or “regulatory sequences” present in an expression vector are those non-translated regions of the vector-origin of replication, selection cassettes, promoters, enhancers, translation initiation signals (Shine Dalgarno sequence or Kozak sequence) introns, a polyadenylation sequence, 5'a nd 3'untranslated regions-which interact with host cellular proteins to carry out transcription and translation. Such elements can vary in their strength and specificity. Depending on the vector system and host utilized, any number of suitable transcription and translation elements, including ubiquitous promoters and inducible promoters can be used.

[0136] Illustrative ubiquitous expression control sequences that can be used in present disclosure include, but are not limited to, a cytomegalovirus (CMV) immediate early promoter, a viral simian virus 40 (SV40) promoter (e.g., early or late) , a Moloney murine leukemia virus (MoMLV) LTR promoter, a Rous sarcoma virus (RSV) LTR, a herpes simplex virus (HSV) (thymidine kinase) promoter, H5, P7.5, and P11 promoters from vaccinia virus, an elongation factor 1-alpha (EF1a) promoter, early growth response 1 (EGR1) , ferritin H (FerH) , ferritin L (FerL) , Glyceraldehyde 3-phosphate dehydrogenase (GAPDH) , eukaryotic translation initiation factor 4A1 (EIF4A1) , heat shock 70kDa protein 5 (HSPA5) , heat shock protein 90kDa beta, member 1 (HSP90B1) , heat shock protein 70kDa (HSP70) , β-kinesin (β-KIN) , the human ROSA 26 locus (Irions et al., Nature Biotechnology 25, 1477 -1482 (2007) ) , a Ubiquitin C promoter (UBC) , a phosphoglycerate kinase-1 (PGK) promoter, a cytomegalovirus enhancer / chicken β-actin (CAG) promoter, and a β-actin promoter.

[0137] Illustrative examples of inducible promoters / systems include, but are not limited to, steroid-inducible promoters such as promoters for genes encoding glucocorticoid or estrogen receptors (inducible by treatment with the corresponding hormone) , metallothionine promoter (inducible by treatment with various heavy metals) , MX-1 promoter (inducible by interferon) , the “GeneSwitch” mifepristone-regulatable system (Sirin et al., 2003, Gene, 323: 67) , the cumate inducible gene switch (WO 2002 / 088346) , tetracycline-dependent regulatory systems, etc.

[0138] The vectors provided herein can be made using standard techniques of molecular biology. For example, the various elements of the vectors provided herein can be obtained using recombinant methods, such as by screening cDNA and genomic libraries from cells, or by deriving the polynucleotides from a vector known to include the same. The various elements of the vectors provided herein can also be produced synthetically, rather than cloned, based on the known sequences. The complete sequence can be assembled from overlapping oligonucleotides prepared by standard methods and assembled into the complete sequence. See, e.g., Edge, Nature (1981) 292: 756; Nambair et al., Science (1984) 223 : 1299; and Jay et al., J. Biol. Chem. (1984) 259: 631 1.

[0139] Thus, particular nucleotide sequences can be obtained from vectors harboring the desired sequences or synthesized completely, or in part, using various oligonucleotide synthesis techniques known in the art, such as site-directed mutagenesis and polymerase chain reaction (PCR) techniques where appropriate. One method of obtaining nucleotide sequences encoding the desired vector elements is by annealing complementary sets of overlapping synthetic oligonucleotides produced in a conventional, automated polynucleotide synthesizer, followed by ligation with an appropriate DNA ligase and amplification of the ligated nucleotide sequence via PCR. See, e.g., Jayaraman et al., Proc. Natl. Acad. Sci. USA (1991) 88: 4084-4088. Additionally, oligonucleotide-directed synthesis (Jones et al., Nature (1986) 54: 75-82) , oligonucleotide directed mutagenesis of preexisting nucleotide regions (Riechmann et al., Nature (1988) 332: 323-327 and Verhoeyen et al., Science (1988) 239: 1534-1536) , and enzymatic filling-in of gapped oligonucleotides using T4 DNA polymerase (Queen et al., Proc. Natl. Acad. Sci. USA (1989) 86: 10029-10033) can be used.

[0140] The RNAs (or cRNAzymes) provided herein can be generated by incubating a vector provided herein under conditions permissive of transcription of the RNAs encoded by the vector. For example, in some embodiments, RNAs (or cRNAzymes) provided herein can be synthesized by incubating a vector provided herein that comprises an RNA polymerase promoter upstream of its 5’ duplex forming region and / or expression sequence with a compatible RNA polymerase enzyme under conditions permissive of in vitro transcription. In some embodiments, the vector is incubated inside of a cell by a bacteriophage RNA polymerase or in the nucleus of a cell by host RNA polymerase P.

[0141] The practice of the invention employs, unless otherwise indicated, conventional techniques in molecular biology, microbiology, genetic analysis, recombinant DNA, organic chemistry, biochemistry, PCR, oligonucleotide synthesis and modification, nucleic acid hybridization, and related fields within the skill of the art. These techniques are described in the references cited herein and are fully explained in the literature. See, e.g., Maniatis et al. (1982) MOLECULAR CLONING: A LABORATORY MANUAL, Cold Spring Harbor Laboratory Press; Sambrook et al. (1989) , MOLECULAR CLONING: A LABORATORY MANUAL, Second Edition, Cold Spring Harbor Laboratory Press; Sambrook et al. (2001) MOLECULAR CLONING: A LABORATORY MANUAL, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Ausubel et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley &Sons (1987 and annual updates) ; CURRENT PROTOCOLS IN IMMUNOLOGY, John Wiley &Sons (1987 and annual updates) Gait (ed. ) (1984) OLIGONUCLEOTIDE SYNTHESIS: A PRACTICAL APPROACH, IRL Press; Eckstein (ed. ) (1991) OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH, IRL Press; Birren et al. (eds. ) (1999) GENOME ANALYSIS: A LABORATORY MANUAL, Cold Spring Harbor Laboratory Press; Borrebaeck (ed. ) (1995) ; each of which is incorporated herein by reference in its entirety. 5.8 cRNAzymes and circRNAs

[0142] Circular RNA is a type of single-stranded, covalently closed-loop RNA, and does not contain the 5’ cap that is commonly known to be required for cap-dependent translation. Thus, circular RNA translation utilizes alternate mechanisms to initiate cap-independent translation, such as the use of an IRES sequence that is recognized by ribosomes. See, e.g., Wesselhoeft, R.A. et al., Nat. Commun. 9, 2629 (2018) and Chinese Patent Application No. 202110594352.4. In some embodiments, provided herein are circular RNAs comprising the IRES sequences disclosed herein.

[0143] The circRNAs can also be produced by ribozyme-catalyzed RNA splicing. As used herein and understood in the art, the term “ribozyme” refers to an RNA molecule with an enzymatic activity. Some ribozymes can catalyze self-splicing independent of the spliceosome, which are referred to as “ribozymes with self-splicing activity, ” “self-splicing ribozymes, ” or “self-splicing introns. ” Naturally occurring self-splicing ribozymes can be divided into group I and group II introns. Although the splicing products of the two categories of ribozymes are similar, the structures and splicing mechanisms of the ribozymes themselves are quite different. The group I intron has a 9-helix structure, which requires an external hydroxyl group in guanosine monophosphate (pG-OH) to trigger the reaction during catalytic splicing, and are highly dependent on the sequences of exons located at both ends of the group I intron. The group II intron relies on its own hydroxyl groups within the nucleotide sequence to trigger splicing. This splicing mechanism is closer to the splicing reaction mediated by a spliceosome and better simulate splicing in higher organisms. The term “group I intron self-splicing activity” or “group I intron activity” refers to the self-splicing activity derived from a group I intron; and the term “group II intron self-splicing activity” or “group II intron activity” refers to the self-splicing activity derived from a group II intron. Method for preparing circRNAs based on group II intron self-splicing activities have at least the following advantages: reduction of the use of biological and chemical reagents (such as ligase and associated reagents) , ease of operation, and simple design.

[0144] RNAs that are engineered ribozymes with self-splicing activity which, upon self-splicing, forms circRNAs are also referred to herein as “cRNAzymes. ” The cRNAzymes can have in vitro self-splicing activity. Disclosed herein are novel non-naturally occurring RNAs that have group II intron self-splicing activity, or “group II cRNAzymes, ” which, upon self-splicing, forms circRNAs that comprise an IRES sequence disclosed herein. Further provided herein are also vectors comprising nucleic acids encoding these group II cRNAzymes.

[0145] In some embodiments, provided herein are non-naturally occurring nucleic acids comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment; wherein the circRNA comprises a TI that comprises a synthetic IRES disclosed herein. In some embodiments, the synthetic IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof. In some embodiments, the synthetic IRES is 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614. In some embodiments, the synthetic IRES is 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. The group II self-splicing activities of the RNAs (cRNAzymes) provided herein are provided by the 3’ and 5’ intron fragments. The target sequences included in the RNAs (or cRNAzymes) provided herein are provided in detail in sections below. 5.9 Structures of naturally occurring group II introns

[0146] The structure and catalytic mechanism of group II introns have recently been elucidated through a combination of genetics, chemical biology, solution biochemistry, and crystallography (e.g., Koch et al. Mol. and Cel. Bio. 12.5 (1992) : 1950-1958; Qin and Pyle, Curr. Opin. Struc. Bio. 8.3 (1998) : 301-308; Toor et al., Science 320.5872 (2008) : 77-82; McNeil et al., Nucleic acids research 42.3 (2014) : 1959-1969; Zhao and Pyle, Trends in Biochem. Sci., 42.6 (2017) : 470-482; Chan et al. Nature Comm. 9.1 (2018) : 1-10. ) .

[0147] Group II introns catalyze self-splicing through an autocatalytic two-step reaction in which the introns excise themselves from surrounding RNA (exons) and stitch the resulting pieces back together. As shown, self-splicing requires only two essential components (intron RNA and a cation) , and it can proceed using purified components in vitro. The cation can be selected from the group consisting of Ba2+, Ca2+, Mg2+, Mn2+, Fe2+, Cu2+, Zn2+, Cd2+, Pb2+, Li+, Cs+, Na+, K+, Rb+, and NH4+, or a combination thereof. In some embodiments, the cation is a bivalent cation, such as Ba2+, Ca2+, Mg2+, Mn2+, Fe2+, Cu2+, Zn2+, Cd2+, and Pb2+. In some embodiments, the cation is a monovalent cation, such as Li+, Na+, K+, Rb+ and Cs+. In some embodiments, the self-splicing requires only the intron RNA and Mg2+.

[0148] The self-splicing reaction is a multistep process that can occur through one of two pathways, and for many introns (such as the yeast mitochondrial intron ai5γ) , both of these pathways are operative. In the branching pathway, the nucleophile for the first step of splicing is a specific bulged adenosine within intron Domain 6 (D6) , whereas in the hydrolysis pathway, the nucleophile during the first step is a water molecule. Both of these reactions lead to productive splicing, and their mechanisms have been extensively investigated.

[0149] All group II introns share a conserved secondary structure that is based on a common set of six (6) radiating domains connected by linker nucleotides, all having the stem-loop structure. Stem-loop structure is a type of a RNA secondary structure, which can be determined by any suitable polynucleotide folding algorithm. The 6 stem-loop structures of naturally occurring group II introns are called domains 1 to 6 (D1 to D6) , and arranged sequentially from 5’ to 3’ . Naturally occurring group II introns comprise multiple exon binding sequences (EBSs) , such as EBS1, EBS2, and EBS3, which interact, such as complementarily pair, with the intron binding sequences (IBSs) in exon regions, triggering splicing by virtue of their own hydroxyl groups within the EBS nucleic acid sequences. Additionally, group II introns also share a common tertiary structure, particularly within the catalytic core. Most of the domains can be transcribed as separate molecules the fold independently and which, when combined with other sections of the intron, retain the catalytic activity. Typically, group II intron structural elements and their role in reaction chemistry can be described by referring to regions within the intron secondary structure.

[0150] Intron Domain 1 (or “D1” ) is the largest domain. It provides the recognition sites for sequence-specific exon binding, and is essential for recognizing the exon in splicing reactions. In addition, D1 contains the active site constituents that form the molecular framework with which the other intronic domain associate. As provided in FIG. 1A, a set of intramolecular pairings are highly conservative and functionally important, including the B-B’ pair, the ε-ε’ pair, the λ-λ’ pair, the α-α’ pair, the ζ-ζ’ pair, the κ-κ’ pair, and the δ-δ’ pair.

[0151] The pairing between the exon-binding sequences (or “EBSs” ) in the intron and the intron-binding sequences (or “IBSs” ) in the flanking exons are critical during the splicing. D1 also contains the key EBSs for binding with exon interaction. EBSs, such as EBS1, EBS2, and EBS3 interact, such as complementarily pair, with the IBSs in exon regions (such as IBS1, IBS2, and IBS3) , whereby the hydroxyl groups within the EBS trigger splicing at the splicing site. In addition to EBS1, EBS2 and EBS3, the single nucleotide located directly upstream of EBS1 in domain 1, the δ nucleotide, can also pair with IBS3, and the interaction between δ and IBS3 is referred to δ-IBS3 pairing. The EBS1-IBS1 interaction, optionally combined with the EBS2-IBS2 interaction are important for specifying the 5’ -splice site. The EBS3 / δ -IBS3 interaction is important for specifying the 3’ -splice site.

[0152] Domain 2 (or “D2” ) can promote the assembly of the active intron structure, forming multiple interactions that control the position of D6 and the branch site. D2 is not phylogenetically conserved and the deletion of D2 was shown to have little effect on the efficiency of self-splicing.

[0153] Domain 3 (or “D3” ) can stimulate reaction chemistry by forming a network of important interactions with D5. Like D2, D3 is also not required for catalysis. D2 and D3 serve to orient their conserved, intervening junction (J2 / 3) within the active site, where it forms a part of the core. Like D1 and D5, D3 can be transcribed as a separate molecule and added to splicing reactions in trans. The lower stem-loop structure of D3 is phylogenetically conserved.

[0154] Domain 4 (or “D4” ) is the least conserved region of the intron, which does not appear to affect the splicing efficiency. In many group II introns, D4 is found to contain an open reading frame from which a maturase is translated, which can bind to stem-loop structures near the basal stem of D4.

[0155] Domain 5 (or “D5” ) is the heart of the active site, and it contains the most highly conserved nucleotides within the intron. D5 is characterized by a terminal loop and stem regions that form critical tertiary interactions with D1, and by a dynamic, asymmetric bulge that is essential for binding of catalytic metal ions. D5 is a small hairpin-loop structure, containing a two-nucleotide bulge and it is capped by a conserved GNRA tetraloop (where N is any nucleotide and R represents a purine) . Studies have found that eight 2’ -hydroxyl groups on D5 have a strong effect specifically on either binding or chemical catalysis, while four pro-Rp phosphate oxygens of D5 affect the overall rate of self-splicing. A variety of nucleotides are important for its function, including the most conserved AGC triad in the first helix (the catalytic triad) , in which the guanine of the triad is invariant and critical for self-splicing in vitro and in vivo.

[0156] Domain 6 (or “D6” ) usually takes the form of a hairpin loop that presents the highly conserved branch-site adenosine, and the rest of the domain facilitates presentation of the branch-site nucleophile during the first step of splicing. The 2'-hydroxyl of this adenosine acts as the nucleophile in the first step of splicing by transesterification, resulting in a 2'-5'linkage between tile branch point adenosine and the first nucleotide of the intron. This-lariat molecule is uniquely characteristic of group II and nuclear spliceosomal introns. Splicing in vivo and in vitro can occur without lariat formation –through a pathway in which water is the nucleophile during the first step of splicing.

[0157] Of all six domains, D1 and D5 are the only two required for catalysis. The presence of D2, D3, and D6 can improve slicing accuracy and / or efficiency in some instances.

[0158] Group II introns share an almost identical catalytic core, and they utilize the same basic mechanism for chemical catalysis. However, they can be divided into several families, including Group IIA, IIB, IIC, and others that display distinct structural and functional differences. The various classes have different 5’ -exon recognition strategies, and they display diversification of architectural scaffolding, protein interaction networks, and some aspects of chemical reactivity. The most prominent difference among IIA, IIB, and IIC ribozymes is the mechanism of exon recognition, because each class uses a distinct combination of pairing interactions to recognize the 5′ and 3′ exons (that is, different combinations of IBS1-EBS1, IBS2-EBS2, IBS3-EBS3, and δ-δ′ pairings) .

[0159] The Group IIB Class are generally believed to be a highly evolved, modern form of the group II intron. The network of hydrogen bonds and metal ion interactions within the catalytic core of IIB introns (involving D5, J2 / 3, and D1) are almost identical to those visualized in group IIC intron. As demonstrated in the crystal structure of the P. littoralis intron, including the lariat form of the intron, the second exon recognition element EBS2 is not required for catalysis. Consistent with previous cross-linking studies, the β-β’ interaction is found to be present in most group IIA and IIB intron, which is proximal to the exon-recognition motifs (EBS1 and EBS2) and appears to facilitate exon orientation within the core. In the three-dimensional structure, the β-β’ kissing loop forms a brace that joins upstream and downstream halves of D1, and it is interesting because it appears to have evolved sequentially over time as group II intron families developed.

[0160] The Group IIA Class share almost all major structural features with IIB introns, although the two classes use different exon recognition strategies. The pairing interactions that are important for exon recognition are for group IIA introns are IBS1-EBS1, IBS2-EBS2, and δ-δ, ′ but not IBS3-EBS3.

[0161] The Group IIC Class are the smallest and most streamlined class of group II introns. Group IIC intron contains only a single, short EBS1 and self-splice through hydrolysis of the 5’ -splice site, rather than by branching in vitro. The most conserved features that are shared by all group II introns can be found in group IIC introns. As demonstrated by the crystal structure of a group IIC intron from the bacterium Oceanobacillus iheyensis, D1 provides a supportive exoskeleton for the active-site domains, and D5 is docked at the center of the D1 shell, where it is secured through an elaborate network of highly conserved interactions, which include the κ-κ’ , λ-λ’ , andζ-ζ’ interactions that are also present in IIB introns. Nucleotides within the D5 bulge are twisted in a manner that brings their backbone phosphates into extremely close proximity, resulting in a highly specific binding site for the two divalent metal ions that are critical for chemical catalysis. The strained conformation of the D5 bulge is made possible by simultaneous interactions with an adjacent triple helix, which is formed by the major groove edge of a D5 stem and nucleotides at the junction between D2 and D3 (J2 / 3) .

[0162] The Group IIE and IIF Classes are known additional group II families. The two families were originally distinguished from the other classes by sequence divergence within the protein maturase domains. Albeit smaller than the IIB class, the IIE and IIF introns appear to have much higher target sequence specificity than IIC introns. 5.10 Structures of cRNAzymes

[0163] Provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment. The circRNA comprises a TI that comprises a synthetic IRES disclosed herein. The “3’ intron fragment” and “5’ intron fragment” are referred herein as such because the linkage of the 5’ -end of the 3’ intron fragment to the 3’ -end of the 5’ intron fragment would form a linear cRNAzyme, wherein the cRNAzyme has the sequence elements that form the essential structural elements of a naturally occurring group II intron and that are sequentially arranged as they would in the naturally occurring group II intron. See PCT / CN2022 / 135585 for additional structures of cRNAzymes, which is incorporated herein in its entirety. For illustrative purposes, the 3’ intron fragment of the RNAs provided herein can comprise a 3’ fragment of Cte (agroup IIB intron that is disclosed in greater detail below) comprising, e.g., D5 and D6 of Cte, or sequence elements that form the essential structural elements of D5 and D6, and a 5’ intron fragment can comprise a 5’ fragment of Cte comprising, e.g., D1 of Cte, optionally in combination with D2 and / or D3 of Cte.

[0164] In some embodiments, the RNAs (or cRNAzyme) provided herein further comprise two homology arms, including a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. The homology arms can help shorten the spatial distance between the 5’ intron fragment and the 3’ intron fragment, thereby facilitating the self-splicing (circularization) reaction. In some embodiments, the presence of the homology arms can enhance the self-splicing efficiency of the RNAs provided herein.

[0165] Accordingly, provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 5’ homology arm, (2) a 3’ intron fragment; (3) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; (4) a 5’ intron fragment; and (5) a 3’ homology arm; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment (FIG. 6) .

[0166] In some embodiments, the two homology arms can be 100%complementary to each other. In some embodiments, the two homology arms can have up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15%base mismatches. In some embodiments, the two homology arms are at least 85%complementary. In some embodiments, the two homology arms are 90%complementary. In some embodiments, the two homology arms are 95%complementary. In some embodiments, the two homology arms are 98%complementary. In some embodiments, the two homology arms are 99%complementary.

[0167] In some embodiments, the 5’ homology arm or 3’ homology arm is 15 to 60 nucleotides in length. In some embodiments, the 5’ homology arm or 3’ homology arm is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length. In some embodiments, both the 5’ homology arm and the 3’ homology arm are 15 to 60 nucleotides in length. In some embodiments, both the 5’ homology arm and the 3’ homology arm are 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length. 5.11 “Scarless” splicing & “near-scarless” splicing

[0168] The exons flanking a naturally occurring group II intron can play important roles for the self-splicing. The 5’ exon refers to the naturally occurring exon sequence on the 5’ end of the group II intron and the 3’ exon refers to the naturally occurring exon sequence on the 3’ end of the group II intron. The 5’ and 3’ flanking exons contain intron binding sequences (IBS) that interact, such as complementarily pair, with the EBS sequence within the intron, which allows the hydroxyl groups within the EBS to trigger splicing at the splicing site. As described above, naturally occurring group II introns commonly contain EBS1, EBS2, and EBS3 within D1 which interact, such as complementarily pair, with IBS1, IBS2, and IBS3 that are present in the flanking exons, respectively. In addition to the pairing between EBS1 and IBS1, between EBS2 and IBS2, and between EBS3 and IBS3, a single nucleotide, the δ nucleotide, which is located directly upstream of EBS1 can also pair with IBS3, and the interaction between δ and IBS3 is referred to δ-IBS3 pairing. While the IBS1-EBS1 and IBS3-EBS3 (δ-IBS3) interaction are generally required for efficient self-splicing of the group II introns, the EBS2 / IBS2 interaction can be removed without significantly affecting the self-splicing activity.

[0169] As such, to effect self-splicing, the EBS1 and EBS3 (or δ) of the RNAs (or cRNAzymes) provided herein need to interact, such as complementarily pair, with IBS1 and IBS3 in the exon elements. The exon elements can be either contained within the target sequence (typically at the terminal regions of the target sequence) or flanking the target sequence. In some embodiments, the target sequence of the RNAs (cRNAzymes) provided herein is flanked by the exon elements E1 and E2, wherein E1 comprises IBS1 and E2 comprises IBS3, and wherein the 3’ end of E2 is linked to 5’ end of the target sequence and the 5’ end of E1 is linked to the 3’ end of the target sequence. Accordingly, provided herein are non-naturally occurring RNAs (cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment, (2) E2; (3) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; (4) E1; and (5) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment, flanked by E1 and E2.

[0170] In some embodiments of the RNAs (or cRNAzymes) provided herein, the exon elements are contained within the target sequence. In other words, certain sequence elements within the target sequence which, for example, can be present at the terminal regions of the target sequence, contain IBS1 and IBS3 and can serve as E1 and E2 in self-splicing. As such, the self-splicing of these RNAs produce circRNAs that do not contain additional sequence elements beyond the target sequence and is therefore referred to herein as “scarless” splicing. A circRNAs that solely consists of the target sequence is referred to herein as a “scarless” circRNA. The terms “scar, ” as used herein, refers to the non-target sequence region in the circRNA splicing product. As such, a “scarless” circRNA contains no scar.

[0171] In some embodiments of the precursor RNA (or cRNAzymes) provided herein that include E1 and E2, the self-splicing produce circRNAs that contain E1 and E2 in addition to the target sequence. In some embodiments, E1 is 0 to 20 nucleotides in length. In some embodiments, E2 is 0 to 20 nucleotides in length. In some embodiments, E1 is 0 to 10 nucleotides in length. In some embodiments, E2 is 0 to 10 nucleotides in length. In some embodiments, E1 is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides in length. In some embodiments, E2 is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides in length. In some embodiments, the E1 and E2 of the RNAs (or cRNAzyme) provided herein combined have no more than 20 nucleotides in length, such as no more than 10 nucleotides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides, and the circRNAs resulting from self-splicing of the RNA (or cRNAzyme) include no more than 20 nucleotides besides the target sequence, which are herein referred to as “near-scarless” circRNAs. The self-splicing of the precursor RNA (or cRNAzymes) provided herein that produces a “near-scarless” circRNA is referred to as “near-scarless” splicing. In some embodiments, the near-scarless circRNA has a scar region equal to or less than 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides in length.

[0172] In some embodiments, to ensure scarless splicing, the naturally occurring EBSs of a group II intron are modified to be complementary to sequence elements of a corresponding length with the target sequence that serve as the IBSs. Accordingly, in some embodiments, the RNAs (or cRNAzymes) disclosed herein are modified to have a modified EBS region which is complementary to a region of a corresponding length in a target sequence. As used herein, an EBS modified to allow scarless splicing is referred to as an EBS’ . Correspondingly, the sequence elements within the target sequence that pair with the EBS’s are referred to herein as the IBS’s. Accordingly, EBS1’ , EBS2’ , and EBS3’ refer to the EBS1, EBS2, and EBS3 sequences that are modified to allow scarless splicing, respectively. IBS1’ , IBS2’ , and IBS3’ refer to the sequences in the target sequence that function as the IBS1, IBS2, and IBS3 in the native exon sequences flanking a group II intron to locate splicing site by interacting with EBS1’ , EBS2’ , and EBS3’ , respectively. Similarly, δ” refers to the nucleotide upstream of EBS1’ that pairs with IBS3’ , and the interaction between δ” and IBS3’ is referred to as the δ” -IBS3’ pairing. In some embodiments, the EBS or the EBS’ , can be 3 to 20 nucleotides in length, preferably 5 to 15 nucleotides, more preferably 6 to 10 nucleotides, such as 6, 7, 8, 9 or 10 nucleotides.

[0173] The region of the target sequence that is complementary paired with the EBS’ , can exist anywhere in the target sequence that allows it to pair with the EBS’ to form a secondary structure necessary for self-splicing. In general, sequences at both ends of the target sequence can be used as they correspond to the location of the IBS sequences of E1 and E2 that naturally interact with EBS. In some embodiments, the EBS’ regions include modified EBS1 (or EBS1’ ) and modified EBS3 (EBS3’ ) regions. In some embodiments, the modified EBS, e.g., EBS1’ , EBS3’ , or both, is (are) complementary to a stretch of sequence located at the 3’ and / or 5’ end of the target sequence.

[0174] In some embodiments, the EBS (or EBS’ ) can be modified so that it is complementarily paired with a stretch of sequence in the target sequence (the IBS or IBS’ ) , thereby allowing interaction. In some embodiments, the modification includes substitution of one or more nucleotides. Certain degrees of mismatch can be tolerated as long as sufficient interaction between the EBS and the IBS exists. In some embodiments, the modified EBS (or EBS’ ) is complementarily paired with a region of a corresponding length in the target sequence on at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100%of the nucleotide positions, or is at least 60%identical, at least 70%, at least 80%, at least 90%, at least 95%, or 100%identical to a complementary paired sequence of a region of a corresponding length in the target sequence.

[0175] As such, provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment. The circRNA comprises a TI that comprises a synthetic IRES disclosed herein. Based on the detailed splicing mechanism, four sets of RNAs (or cRNAzymes) are expressly contemplated in the present disclosure.

[0176] Set 1: Near-scarless splicing based on the EBS1-IBS1 pairing and the EBS3-IBS3 pairing (FIG. 2) . In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein comprises EBS1 and EBS3 that are each at least 60%complementarily paired with a region of a corresponding length flanking the target sequence. The target sequence can be flanked by E1 on its 3’ end and E2 on its 5’ end, wherein E1 and E2 comprise IBS1 and IBS3, respectively, and wherein the EBS1-IBS1 interaction and the EBS3-IBS3 interaction allow the self-splicing and the production of a near-scarless circRNA.

[0177] In some embodiments, EBS1 and EBS3 are complementarily paired with IBS1 and IBS3, respectively on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, the EBS1 and / or EBS3 can be modified from their naturally occurring counterpart (s) . The modification can include a substitution, a deletion and / or an addition of one or more nucleotides.

[0178] In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein further comprises EBS2 that is at least 60%complementarily paired with a region of a corresponding length within the target sequence, namely, the IBS2, and the EBS2-IBS2 interaction can further promotes the efficiency and accuracy of the self-splicing. In some embodiments, EBS2 is complementarily paired with IBS2 on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, EBS2 can be modified from their naturally occurring counterpart (s) . The modification can include a substitution, a deletion and / or an addition of one or more nucleotides.

[0179] Set 2: Near-scarless splicing based on the EBS1-IBS1 pairing and the δ-IBS3 pairing (FIG. 3) . In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein comprises EBS1 and the δ nucleotide, wherein EBS1 is 60%complementarily paired with a region of a corresponding length flanking the target sequence and the δ nucleotide is complementarily paired with a nucleotide within a sequence that flanks the target sequence. The target sequence can be flanked by E1 on its 3’ end and E2 on its 5’ end, wherein E1 and E2 comprise IBS1 and the δ nucleotide, respectively, and wherein the EBS1-IBS1 interaction and the δ-IBS3 interaction allow the self-splicing and the production of a near-scarless circRNA.

[0180] In some embodiments, the EBS1 is complementarily paired with IBS1 on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, the 4-10 nucleotides immediate upstream of the δ nucleotide (referred to herein as the “δ upstream” ) are complementarily paired with IBS3 on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, the EBS1 and / or the δ upstream can be modified from their naturally occurring counterpart (s) . The modification can include a substitution, a deletion and / or an addition of one or more nucleotides.

[0181] In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein further comprises EBS2 that is at least 60%complementarily paired with a region of a corresponding length within the target sequence, namely, the IBS2, and the EBS2-IBS2 interaction can further promote the efficiency and accuracy of the self-splicing. In some embodiments, EBS2 is complementarily paired with IBS2 on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, EBS2 can be modified from their naturally occurring counterpart (s) . The modification can include a substitution, a deletion and / or an addition of one or more nucleotides.

[0182] Set 3: Scarless splicing based on the EBS1’ -IBS1’ pairing and the EBS3’ -IBS3’ pairing (FIG. 4) . In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein comprises EBS1’ and EBS3’ that are each at least 60%complementarily paired with a region of a corresponding length in the target sequence. The complementarily paired regions can be located at one or both ends of the target sequence. As such, the target sequence can contain a sequence element at its 3’ terminal region that can serve as E1 and another sequence element at its 5’ terminal region that can serve as E2, wherein E1 and E2 comprise IBS1’ and IBS3’ , respectively, and wherein the EBS1’ -IBS1’ interaction and the EBS3’ -IBS3’ interaction allow the self-splicing and the production of a scarless circRNA.

[0183] To achieve EBS1’ -IBS1’ pairing and the EBS3’ -IBS3’ pairing, in some embodiments, the intron sequence elements EBS1’ and EBS3’ are modified from their naturally occurring counterparts to pair with the corresponding regions in the target sequence (e.g., the terminal regions) . In some embodiments, EBS1’ is modified to pair with IBS1’ at the 3’ terminal region of the target sequence. As such, the 3’ terminal region of the target sequence serves as E1. In some embodiments, EBS3’ is modified to pair with IBS3’ at the 5’ terminal region of the target sequence. As such, the 5’ terminal region of the target sequence serves as E2. In some embodiments, the sequence elements of the terminal region of the target sequence can be modified to pair with the EBS sequences in D1. For example, serving as E1, the 3’ terminal region of the target sequence can be modified to contain IBS1’ , the sequence element to pair with EBS1’ in D1. Similarly, serving as E2, the 5’ terminal region of the target sequence can be modified to contain IBS3’ , the sequence element to pair with EBS3’ in D1.

[0184] In some embodiments, the EBS1’ and EBS3’ are complementarily paired with IBS1’ and IBS3', respectively on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions.

[0185] Set 4: Scarless splicing based on the EBS1’ -IBS1’ pairing and the δ” -IBS3’ pairing (FIG. 5) . In some embodiments, the D1-like domain of the RNAs (or cRNAzymes) provided herein comprises EBS1’ and the δ” nucleotide, wherein EBS1’ is 60%complementarily paired with a region of a corresponding length in the target sequence and the δ nucleotide is complementarily paired with a nucleotide within the target sequence. The complementarily paired regions can be located at one or both ends of the target sequence. As such, the target sequence can contain a sequence element at its 3’ terminal region that can serve as E1 and another sequence element at its 5’ terminal region that can serve as E2, wherein E1 and E2 comprise IBS1’ and δ” , respectively, and wherein the EBS1’ -IBS1’ interaction and the δ” -IBS3’ interaction allow the self-splicing and the production of a scarless circRNA.

[0186] To achieve EBS1’ -IBS1’ pairing and the δ” -IBS3’ pairing, in some embodiments, the intron sequence elements EBS1’ and δ” are modified from their naturally occurring counterparts to pair with the corresponding regions in the target sequence (e.g., the terminal regions) . In some embodiments, EBS1’ is modified to pair with IBS1’ at the 3’ terminal region of the target sequence. As such, the 3’ terminal region of the target sequence can serve as E1. In some embodiments, the δ” nucleotide (optionally with its upstream) is modified to pair with IBS3’ at the 5’ terminal region of the target sequence. As such, the 5’ terminal region of the target sequence can serve as E2. In some embodiments, the sequence elements of the terminal region of the target sequence can be modified to pair with the EBS sequences in D1. For example, serving as E1, the 3’ terminal region of the target sequence can be modified to contain IBS1’ , the sequence element to pair with EBS1’ in D1. Similarly, serving as E2, the 5’ terminal region of the target sequence can be modified to contain the δ” nucleotide (optionally with its upstream) , the sequence element to pair with EBS3’ in D1.

[0187] In some embodiments, the EBS1’ is complementarily paired with IBS1’ on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. In some embodiments, the 4-10 nucleotides on the immediate upstream of the δ” nucleotide (referred to herein as the “δ” upstream” ) are complementarily paired with IBS3’ on at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%of the nucleotide positions. 5.12 Segmenting

[0188] The 3’ and 5’ intron fragments of RNAs (or cRNAzymes) provided herein can be generated by segmenting a naturally occurring group II intron at an unpaired region into two fragments: the 5’ fragment and the 3’ fragment, respectively serving as the 3’ and 5’ intron fragments of RNAs (or cRNAzymes) provided herein. In other words, in the RNAs (or cRNAzyme) provided herein, the 5’ fragment and the 3’ fragment of the naturally occurring group II intron are swapped and re-ligated with the target sequence inserted in between, resulting in the following construct from 5’ to 3’ : 3’ intron fragment-target sequence-5’ intron fragment. the 5’ fragment and the 3’ fragment.

[0189] In some embodiments, the naturally occurring group II intron can be modified. The modified group II intron can include a substitution, a deletion and / or an addition of one or more nucleotides. In some embodiments, the modification does not affect the self-splicing activity of the group II intron, especially the in vitro self-splicing activity.

[0190] In some embodiments, the 5’ fragment and the 3’ fragment of the naturally occurring group II intron are mutated, swapped and re-ligated with the target sequence inserted in between to form the RNAs (or cRNAzymes) disclosed herein. The mutation can comprise modification of one or more nucleotides, such as an addition, a deletion, and a substitution of one or more nucleotides, relative to their naturally occurring wild-type sequences. In some embodiments, the modification promotes the accuracy and / or efficacy of the self-splicing of the resulting RNAs (or cRNAzymes) disclosed herein.

[0191] In some embodiments, the modification includes deletion of the intron encoded protein (IEP) sequence in D4. The IEP sequence or similar structures in D4 are present in all group II introns, and known to be not required for in vitro transcription. In some embodiments, the modification includes deletion of the IEP of D4, whereas RNAs (or cRNAzymes) disclosed herein still comprise the 5’ and 3’ stem sequences of D4. The complementarity of the 5’ and 3’ stem sequences of D4 can help shorten the spatial distance between the 5’ intron fragment and the 3’ intron fragment, thereby facilitating the circularization reaction. In some embodiments, the modification includes deletion of the entire D4.

[0192] In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at an unpaired region into two fragments. In some embodiments, an unpaired region is a linear region between two adjacent domains of the group II intron. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D1. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D2. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D3. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D4. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D5. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a loop region of a stem-loop structure of D6.

[0193] In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a linear region between D1 and D2. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a linear region between D2 and D3. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a linear region between D3 and D4. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a linear region between D4 and D5. In some embodiments, the 5’ intron fragment and the 3’ intron fragment are obtained by segmenting a group II intron at a linear region between D5 and D6.

[0194] As a person of ordinary skill in the art would understand, each of the four sets of splicing mechanism described above can also apply in RNAs (or cRNAzymes) disclosed herein of which the 5’ and 3’ intron fragments are generated by segmenting, swapping and re-ligating the 5’ and 3’ fragments of naturally occurring group II introns. Accordingly, expressly contemplated herein are RNAs (or cRNAzymes) described in the instant section that are capable of (1) near-scarless splicing based on the EBS1-IBS1 pairing and the EBS3-IBS3 pairing; (2) near-scarless splicing based on the EBS1-IBS1 pairing and the δ-IBS3 pairing; (3) scarless splicing based on the EBS1’ -IBS1’ pairing and the EBS3’ -IBS3’ pairing; or (4) scarless splicing based on the EBS1’ -IBS1’ pairing and the δ” -IBS3’ pairing. 5.13 Exemplary group II introns and fragments thereof

[0195] Group II introns are found in eubacteria, archaebacteria, and the organelles of plants, fungi, and various lower eukaryotes. While the RNAs (or cRNAzymes) exemplified below focus on specific group II introns, namely, Cte, Oi, Pli, LtrB, and Syn1, guided with the teachings of instant disclosure, a person of ordinary skill in the art would be able to prepare additional RNAs (or cRNAzymes) based on using sequence elements from other group II introns.

[0196] RNAs (or cRNAzymes) disclosed herein can be derived from any naturally occurring group II intron. In some embodiments, RNAs (or cRNAzymes) disclosed herein can contain sequence elements from any naturally occurring group II intron disclosed herein or otherwise known in the art. A list of group II introns and their nucleotide sequences are provided in Tables 2 and 16.

[0197] In some embodiments, for example, provided herein are RNAs (or cRNAzymes) that contain sequence elements derived from exemplary group II introns Cte (SEQ ID NO: 135) . The secondary structure of Cte is provided in FIG. 1B. In some embodiments, RNAs (or cRNAzymes) provided contain sequence elements derived from a modified Cte (e.g., SEQ ID NO: 136 or 137) or a synthetic Cte (e.g., Cte-Syn1; SEQ ID NO: 143) .

[0198] In some embodiments, for example, provided herein are RNAs (or cRNAzymes) that contain sequence elements derived from exemplary group II introns Oi (SEQ ID NO: 138) . In some embodiments, RNAs (or cRNAzymes) provided contain sequence elements derived from a modified Oi (e.g., SEQ ID NO: 139) or a synthetic Oi. In some embodiments, for example, provided herein are RNAs (or cRNAzymes) that contain sequence elements derived from exemplary group II introns Pli (SEQ ID NO: 140) . In some embodiments, RNAs (or cRNAzymes) provided contain sequence elements derived from a modified Pli or a synthetic Pli (e.g., Pli-Syn1; SEQ ID NO: 144) . In some embodiments, for example, provided herein are RNAs (or cRNAzymes) that contain sequence elements derived from exemplary group II introns LtrB (SEQ ID NO: 140 or 141) . In some embodiments, RNAs (or cRNAzymes) provided contain sequence elements derived from a modified LtrB or a synthetic LtrB (e.g., LtrB-Syn1; SEQ ID NO: 145) .

[0199] In some embodiments, the group II intron from which the RNAs (or cRNAzymes) provided herein can be derived has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 33-41 and 135-145. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 33. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 34. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 35. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 36. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 37. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 38. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 39. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 40. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 41. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 135. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 136. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 137. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 138. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 139. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 140. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 141. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 142. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 143. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 144. In some embodiments, the group II intron has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 145.

[0200] In some embodiments, the nucleotide sequence of the group II intron from which the RNAs (or cRNAzymes) provided herein can be derived consists essentially of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 33-41 and 135-145. In some embodiments, the nucleotide sequence of the group II intron from which the RNAs (or cRNAzymes) provided herein can be derived consists of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 33-41 and 135-145. In some embodiments, the nucleotide sequence of the group II intron from which the RNAs (or cRNAzymes) provided herein can be derived consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 33-41 and 135-145. In some embodiments, the nucleotide sequence of the group II intron from which the RNAs (or cRNAzymes) provided herein can be derived consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 33-41 and 135-145. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 33. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 34. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 35. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 36. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 37. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 38. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 39. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 40. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 41. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 135. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 136. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 137. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 138. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 139. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 140. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 141. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 142. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 143. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 144. In some embodiments, the group II intron has the nucleotide sequence of SEQ ID NO: 145.

[0201] As disclosed above and understood in the art, certain sequence elements form the essential structural elements required for the self-splicing activities of group II introns. A person of ordinary skill in the art would be able to identify such sequence elements with the aid of the sequence analysis tools for RNAs disclosed herein or otherwise known in the art. For illustrative purposes, the sequence elements for exemplary naturally occurring group II introns Cte, Oi, Pli, and LtrB as well as synthetic cRNAzymes derived therefrom with group II intron activities (Cte-syn1, Oi-syn1, Pli-syn1, and LtrB-syn1 are summarized in Table 18.

[0202] Provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment, wherein the circRNA comprises a synthetic IRES disclosed herein.

[0203] Exemplary 3’ intron fragment:

[0204] In some embodiments, the 3’ intron fragment of the RNAs (cRNAzymes) provided herein comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 42-52 and 228. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 42. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 43. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 44. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 45. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 46. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 47. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 48. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 49. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 50. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 51. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 52. In some embodiments, the 3’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 228.

[0205] In some embodiments, the 3’ intron fragment consists essentially of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 42-52 and 228. In some embodiments, the 3’ intron fragment consists of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 42-52 and 228. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 42. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 43. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 44. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 45. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 46. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 47. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 48. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 49. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 50. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 51. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 52. In some embodiments, the 3’ intron fragment has the nucleotide sequence of SEQ ID NO: 228.

[0206] Exemplary 5’ intron fragment:

[0207] In some embodiments, the 5’ intron fragment of the RNAs (cRNAzymes) provided herein comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 75. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 76. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 77. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 78. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 79. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 80. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 81. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 82. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 83. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 84. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 85. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 86. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 87. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 88. In some embodiments, the 5’ intron fragment comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 229.

[0208] In some embodiments, the 5’ intron fragment consists essentially of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229. In some embodiments, the 5’ intron fragment consists of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 75. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 76. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 77. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 78. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 79. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 80. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 81. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 82. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 83. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 84. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 85. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 86. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 87. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 88. In some embodiments, the 5’ intron fragment has the nucleotide sequence of SEQ ID NO: 229.

[0209] Exemplary E2 and E1:

[0210] In some embodiments, the E2 of the RNAs (cRNAzymes) provided herein comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 53-63. In some embodiments, the E2 of the precursor RNA provided herein consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 53-63. In some embodiments, the E2 consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 53-63. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 53. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 54. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 55. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 56. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 57. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 58. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 59. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 60. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 61. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 62. In some embodiments, the E2 has the nucleotide sequence of SEQ ID NO: 63.

[0211] In some embodiments, the E1 of the of the RNAs (cRNAzymes) provided herein comprises a nucleotide sequence selected from the group consisting of: SEQ ID NOs: 64-74. In some embodiments, the E1 consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 64-74. In some embodiments, the E1 consists of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 64-74. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 64. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 65. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 66. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 67. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 68. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 69. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 70. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 71. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 72. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 73. In some embodiments, the E1 has the nucleotide sequence of SEQ ID NO: 74.

[0212] In some embodiments, IBS3 (or IBS3’ ) , the region of a corresponding length of EBS3 or EBS3’ , which either flanks a target sequence (IBS3) or is within the target sequence (IBS3’ ) , optionally with its down sequence is selected from the group consisting of: (a) SEQ ID NO: 131, (b) SEQ ID NO: 132, (c) SEQ ID NO: 133, and (d) SEQ ID NO: 134. In some embodiments, the δnucleotide and the δ upstream (or the δ” nucleotide and the δ” upstream) comprises a nucleotide sequence selected from the group consisting of: (a) SEQ ID NO: 127, (b) SEQ ID NO: 128, (c) SEQ ID NO: 129, and (d) SEQ ID NO: 130.

[0213] Homology arms: In some embodiments, the RNAs provided herein further comprise two homology arms, including a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. The homology arms can help shorten the spatial distance between the 5’ intron fragment and the 3’ intron fragment, thereby facilitating the self-splicing (circularization) reaction. In some embodiments, the presence of the homology arms can enhance the self-splicing efficiency of the RNAs provided herein.

[0214] Accordingly, provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 5’ homology arm, (2) a 3’ intron fragment; (3) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; (4) a 5’ intron fragment; and (5) a 3’ homology arm; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment (FIG. 6) .

[0215] In some embodiments, the two homology arms can be 100%complementary to each other. In some embodiments, the two homology arms can have up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15%base mismatches. In some embodiments, the two homology arms are at least 85%complementary. In some embodiments, the two homology arms are 90%complementary. In some embodiments, the two homology arms are 95%complementary. In some embodiments, the two homology arms are 98%complementary. In some embodiments, the two homology arms are 99%complementary.

[0216] In some embodiments, the 5’ homology arm or 3’ homology arm is 15 to 60 nucleotides in length. In some embodiments, the 5’ homology arm or 3’ homology arm is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length. In some embodiments, both the 5’ homology arm and the 3’ homology arm are 15 to 60 nucleotides in length. In some embodiments, both the 5’ homology arm and the 3’ homology arm are 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length.

[0217] In some embodiments, the 5’ homology arm comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 105. In some embodiments, the 5’ homology arm consists essentially a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 105. In some embodiments, the 5’ homology arm consists of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 105. In some embodiments, the 5’ homology arm has the nucleotide sequence of SEQ ID NO: 105.

[0218] In some embodiments, the 3’ homology arm comprises a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 106. In some embodiments, the 3’ homology arm consists essentially a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 106. In some embodiments, the 3’ homology arm consists of a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to SEQ ID NO: 106. In some embodiments, the 3’ homology arm has the nucleotide sequence of SEQ ID NO: 106.

[0219] Modified nucleotides / nucleosides

[0220] In some embodiments, the RNAs (or cRNAzymes) provided herein can comprise modified nucleotide / nucleoside. The modified RNA nucleotide and / or modified nucleoside can be introduced, for example, at in vitro transcription (IVT) . As used herein, in vitro transcription, or “IVT, ” refers to versatile method to produce RNA in vitro that uses an RNA polymerase, ribonucleotides, and appropriate buffer conditions to synthesize RNA from a DNA template.

[0221] In some embodiments, the RNAs (or cRNAzymes) provided herein comprise 10%to 100%modified RNA nucleotide and / or modified nucleoside. In some embodiments, the RNAs (or cRNAzymes) provided herein comprise 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%modified RNA nucleotide and / or modified nucleoside.

[0222] In some embodiments, the modified RNA nucleotide and / or modified nucleoside is m5C (5-methylcytidine) . In some embodiments, the polynucleotide construct of any one of Embodiments 47-48, wherein at least one of the modified RNA nucleotide and / or modified nucleoside is m5U (5-methyluridine) . In some embodiments, the modified RNA nucleotide and / or modified nucleoside is m6A (N6-methyladenosine) . In some embodiments, the modified RNA nucleotide and / or modified nucleoside is Y (pseudouridine) . In some embodiments, the modified RNA nucleotide and / or modified nucleoside is m1A (1-methyladenosine) . In some embodiments, the modified nucleoside is selected from the group consisting of: m5C (5-methylcytidine) , m5U (5-methyluridine) , m6A (N6-methyladenosine) , s2U (2-thiouridine) , Y (pseudouridine) , Um (2 '-O-methyluridine) , m1A (1-methyladenosine) , m2A (2-methyladenosine) , Am (2’ -0-methyladenosine) , ms2 m6A (2-methylthio-N6-methyladenosine) , i6A (N6-isopentenyladenosine) , ms2i6A (2-methylthio-N6 isopentenyladenosine) , io6A (N6- (cis-hydroxyisopentenyl) adenosine) , ms2io6A (2-methylthio-N6- (cis-hydroxyisopentenyl) adenosine) , g6A (N6-glycinylcarbamoyladenosine) , t6A (N6-threonylcarbamoyladeno sine) , ms2t6A (2-methylthio-N6-threonyl carbamoyladenosine) , m6t6A (N6-methyl-N6-threonylcarbamoyladenosine) , hn6A (N6-hydroxynorvalylcarbamoyladenosine) , ms2hn6A (2-methylthio-N6-hydroxynorvalyl carbamoyladenosine) , Ar (p) (2’ -0-ribosyladenosine (phosphate) ) , I (inosine) , m1I (1-methylinosine) , m1hn (1, 2’ -O-dimethylinosine) , m3C (3-methylcytidine) , Cm (2’ -0-methylcytidine) , s2C (2-thiocytidine) , ac4C (N4-acetylcytidine) , (5-formylcytidine) , m5Cm (5, 2 '-O-dimethylcytidine) , ac4Cm (N4-acetyl-2’ -O-methylcytidine) , k2C (lysidine) , m! G (1-methylguanosine) , m2G (N2-methylguanosine) , m7G (7-methylguanosine) , Gm (2'-0-methylguanosine) , m2 2G (N2, N2-dimethylguanosine) , m2Gm (N2, 2’ -O-dimethylguanosine) , m2 aGm (N2, N2, 2’ -O-trimethylguanosine) , Gr (p) (2’ -0-ribosylguanosine (phosphate) ) , yW (wybutosine) , oayW (peroxywybutosine) , OHyW (hydroxy wybutosine) , OHyW* (undermodified hydroxywybutosine) , imG (wyosine) , mimG (methylwyosine) , Q (queuosine) , oQ (epoxyqueuosine) , galQ (galactosyl-queuosine) , manQ (mannosyl-queuosine) , preQo (7-cyano-7-deazaguanosine) , preQi (7-aminomethyl-7-deazaguanosine) , G+ (archaeosine) , D (dihydrouridine) , m5Um (5, 2’ -0-dimethyluridine) , s4U (4-thiouridine) , m5s2U (5-methyl-2-thiouridine) , s2Um (2-thio-2’ -0-methyluridine) , acp3U (3- (3-amino-3-carboxypropyl) uridine) , ho5U (5-hydroxyuridine) , mo5U (5-methoxyuridine) , cmo5U (uridine 5-oxy acetic acid) , mcmo5U (uridine 5-oxy acetic acid methyl ester) , chm5U (5- (carboxyhydroxymethyl) uridine) ) , mchm5U (5- (carboxyhydroxymethyl) uridine methyl ester) , mcm5U (5-methoxycarbonylmethyluridine) , mcm5Um (5-methoxycarbonylmethyl-2’ -0-methyluridine) , mcm5s2U (5-methoxycarbonylmethyl-2-thiouridine) , nm5S2U (5-aminomethyl-2-thiouridine) , mnm5U (5-methylaminomethyluridine) , mnm5s2U (5-methylaminomethyl-2-thiouridine) , mnm5se2U (5-methylaminomethyl-2-selenouridine) , ncm5U (5-carbamoylmethyluridine) , ncm5Um (5-carbamoylmethyl-2 '-O-methyluridine) , cmnm5U (5-carboxymethylaminomethyluridine) , cmnm5Um (5-carboxymethylaminomethyl-2'-0-methyluridine) , cmnm5s2U (5-carboxymethylaminomethyl-2-thiouridine) , m6 2A (N6, N6-dimethyladenosine) , Im (2’ -0-methylinosine) , m4C (N4-methylcytidine) , m4Cm (N4, 2’ -0-dimethylcytidine) , hm5C (5-hydraxymethylcytidine) , m3U (3-methyluridine) , cm5U (5-carboxymethyluridine) , m6Am (N6, 2’ -O-dimethyladenosine) , m6 2Am (N6, N6, 0-2’ -trimethyladenosine) , m2, 7G (N2, 7-dimethylguanosine) , m2, 2, 7G (N2, N2, 7-trimethylguanosine) , m3Um (3, 2’ -0-dimethyluridine) , m5D (5-methyldihydrouridine) , f5Cm (5-formyl-2’ -0-methylcytidine) , m'Gm (l, 2’ -0-dimethylguanosine) , m'A m (l, 2’ -0-dimethyladenosine) , rm 5U (5-taurinomethyluridine) , τm5s2U (5-taurinomethyl-2-thiouridine) ) , imG-14 (4-demethylwyosine) , imG2 (isowyosine) , or ac6A (N6-acetyladenosine) , pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, l-taurinomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-m ethoxy-2-thio-pseudouridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2, 6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2, 6-diaminopurine, 7-deaza-8-aza-2, 6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6- (cis-hydroxyisopentenyl) adenosine, 2-methylthio-N6- (cis-hydroxyisopentenyl) adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonyl carbamoyladenosine, N6, N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2, N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2, N2-dimethyl-6-thio-guanosine, 5-methylcytosine, pseudouridine, and 1-methylpseudouridine. 5.14 Target sequence

[0223] Provided herein are non-naturally occurring RNAs (or cRNAzymes) comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circRNA that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment (FIG. 6) ; wherein the circRNA comprises a TI that comprises a synthetic IRES disclosed herein.

[0224] In some embodiments, the circRNA also includes therapeutic protein-coding sequence (Z1) operatively linked to the TI. Additional elements, such as a transcription termination signal, can also be included. The Z1 sequence can encode any protein whose expression is desired, such as those disclosed herein.

[0225] In some embodiments, the resulting targeting sequence comprises an expression cassette. In some embodiments, the resulting targeting sequence comprises a translation initiation sequence (TI) and a therapeutic protein-coding sequence (Z1) , wherein the 3’ -end of TI is operatively linked to the 5’ -end of Z1 (FIGs. 7A (a) - (b) , 7B (a) - (b) , 8A (a) - (b) and 8B (a) - (b) ) . In some embodiments, the resulting circRNA comprises the resulting target sequence. In some embodiments, the resulting circRNA consists of the resulting target sequence.

[0226] Set 1: In some embodiments of the RNAs provided herein, the target sequence consists of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; wherein the 3’ target sequence fragment comprises Z1 and the 5’ target sequence fragment comprises TI (FIGs. 7A (a) - (b) ) . In some embodiments, the 3’ target sequence fragment further comprises one or two linkers flanking Z1 (FIGs. 7B (a) - (b) ) . Accordingly, in some embodiments, circRNAs produced by the self-splicing of the RNAs provided herein comprise TI and Z1, wherein the 3’ -end of TI is operatively linked to the 5’ -end of Z1 (FIGs. 7A (a) - (b) ) . In some embodiments of the circRNAs, the Z1 is flanked by one or two linkers (FIGs. 7B (a) - (b) ) .

[0227] In some embodiments, RNAs provided herein can have a structure of Formula (I) : 5’ - (3’ IF) - (L) n-Z1- (L) n-TI- (5’ IF) -3’ ; wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence, Z1 is a therapeutic protein-coding sequence, and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 7A (a) - (b) and 7B (a) - (b) , upper) .

[0228] In some embodiments, the RNAs provided herein further comprise two homology arms, a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. Accordingly, in some embodiments, RNAs provided herein can have a structure of Formula (I) ’ : 5’ - (5’ HA) - (3’ IF) -(L) n-Z1- (L) n-TI- (5’ IF) - (3’ HA) -3’ ; wherein 5’ HA is the 5’ homology arm; 3’ HA is the 3’ homology arm; 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence, Z1 is a therapeutic protein-coding sequence, and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 7A (a) - (b) and 7B (a) - (b) , upper) .

[0229] FIGs. 7A (a) and 7B (a) show scarless splicing wherein the sequence elements within the target sequence serve as E1 and E2, the 5’ terminal region of the target sequence (e.g., part of Z1 or linker) can serve as E2. In some embodiments, the 3’ terminal region of the target sequence (e.g., part of TI) can serve as E1. Optionally, as depicted on FIGs. 7A (b) and 7B (b) , (1) extra exon sequence E2 can be included between the 3’ intron fragments and the target sequence; (2) extra exon sequence E1 can be included between the target sequence and the 5’ intron fragments; or both (1) and (2) . As such, E1 and / or E2 remain with the target sequence in the circRNA after self-splicing.

[0230] Set 2: In some embodiments of the RNAs provided herein, the target sequence consists of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; wherein the 3’ target sequence fragment comprises TI and the 5’ target sequence fragment comprises Z1 (FIGs. 8A (a) - (b) ) . In some embodiments, the 3’ target sequence fragment further comprises two linkers (L) flanking TI (FIGs. 8B (a) - (b) ) . Accordingly, in some embodiments, circRNAs produced by the self-splicing of the RNAs provided herein comprise TI and Z1, wherein the 3’ -end of TI is operatively linked to the 5’ -end of Z1 (FIGs. 8A (a) - (b) ) . In some embodiments of the circRNAs, the TI is flanked by one or two linkers (FIGs. 8B (a) - (b) ) .

[0231] Accordingly, in some embodiments, RNAs provided herein can have a structure of Formula (II) : 5’ - (3’ IF) - (L) n-TI- (L) n-Z1- (5’ IF) -3’ ; wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence; Z1 is a therapeutic protein-coding sequence, and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 8A (a) - (b)  and 8B (a) - (b) , upper) .

[0232] In some embodiments, the RNAs provided herein further comprise two homology arms, a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. Accordingly, in some embodiments, RNAs provided herein can have a structure of Formula (II) ’ : 5’ - (5’ HA) - (3’ IF) -(L) n-TI- (L) n-Z1- (5’ IF) - (3’ HA) -3’ ; wherein 5’ HA is the 5’ homology arm; 3’ HA is the 3’ homology arm; 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence; Z1 is a therapeutic protein-coding sequence, and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 8A (a) - (b) and 8B (a) - (b) , lower) .

[0233] FIGs. 8A (a) and 8B (a) show scarless splicing wherein the sequence elements within the target sequence serve as E1 and E2. In some embodiments, the 5’ terminal region of the target sequence (e.g., part of TI or linker) can serve as E2. In some embodiments, the 3’ terminal region of the target sequence (e.g., part of Z1) can serve as E1. Optionally, as depicted on FIGs. 8A (b) and 8B (b) , (1) extra exon sequence E2 can be included between the 3’ intron fragments and the target sequence; (2) extra exon sequence E1 can be included between the target sequence and the 5’ intron fragments; or both (1) and (2) . As such, E1 and / or E2 remain with the target sequence in the circRNA after self-splicing.

[0234] Set 3: The translation initiation sequence TI of the RNAs described herein can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) . In some embodiments of the RNAs provided herein, the target sequence consists of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; wherein the 3’ target sequence fragment comprises, from 5’ to 3’ , a 3’ fragment of TI (TIB) and Z1; and wherein the 5’ target sequence fragment comprises a 5’ fragment of TI (TIA) (FIGs. 9A (a) - (b) ) . In some embodiments, the 3’ target sequence further comprises two linkers (L) flanking Z1 (FIGs. 9B (a) - (b) ) . Accordingly, in some embodiments, circRNAs produced by the self-splicing of the RNAs provided herein comprise TIA, TIB and Z1, wherein the 3’ -end of TIA is operatively linked to the 5’ -end of TIB (FIGs. 9A (a) - (b) ) . In some embodiments of the circRNAs, Z1 is flanked by one or two linkers (FIGs. 9B (a) - (b) ) .

[0235] Accordingly, in some embodiments, the RNAs provided herein can have a structure of Formula (III) : 5’ - (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) -3’ ; wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is a therapeutic protein-coding sequence; and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 9A (a) - (b) and 9B (a) - (b) , upper) .

[0236] In some embodiments, the RNAs provided herein further comprise two homology arms, a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. Accordingly, in some embodiments, RNAs provided herein can have a structure of Formula (III) ’ : 5’ - (5’ HA) - (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) - (3’ HA) -3’ ; wherein 5’ HA is the 5’ homology arm; 3’ HA is the 3’ homology arm; 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is a therapeutic protein-coding sequence; and each L is independently a linker sequence, wherein n=0, 1 or 2 (FIGs. 9A (a) - (b) and 9B (a) - (b) , lower) .

[0237] FIGs. 9A (a) and 9B (a) show scarless splicing wherein the sequence elements within the target sequence serve as E1 and E2. In some embodiments, the 5’ terminal region of the target sequence (e.g., part of TIB) can serve as E2. In some embodiments, the 3’ terminal region of the target sequence (e.g., part of TIA) can serve as E1. Optionally, as depicted on FIGs. 9A (b) and 9B (b) , (1) extra exon sequence E2 can be included between the 3’ intron fragments and the target sequence; (2) extra exon sequence E1 can be included between the target sequence and the 5’ intron fragments; or both (1) and (2) . As such, E1 and / or E2 remain with the target sequence in the circRNA after self-splicing.

[0238] Set 4: The protein coding sequence Z1 of the RNAs described herein can be segmented into a 5’ fragment (Z1A) and a 3’ fragment Z1 (Z1B) . In some embodiments of the RNAs provided herein, the 3’ target sequence fragment comprises a 3’ fragment of Z1 (Z1B) ; and wherein the 5’ target sequence fragment comprises, from 5’ to 3’ , TI and a 5’ fragment of Z1 (Z1A) (FIG. 10A) . In some embodiments, the 3’ target sequence further comprises two linkers (L) flanking TI (FIG. 10B) . Accordingly, in some embodiments, circRNAs produced by the self-splicing of the RNAs provided herein comprise TI, Z1A, and Z1B and wherein the 3’ -end of Z1A is operatively linked to the 5’ -end of Z1B (FIG. 10A) . In some embodiments of the circRNAs, TI is flanked by one or two linkers (FIG. 10B) .

[0239] Accordingly, in some embodiments, the RNAs provided herein can have a structure of Formula (IV) : 5’ - (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) -3’ ; wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence; Z1 is a therapeutic protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2 (FIGs. 10A and 10B, upper) .

[0240] In some embodiments, the RNAs provided herein further comprise two homology arms, a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. Accordingly, in some embodiments, RNAs provided herein can have a structure of Formula (IV) ’ : 5’ - (5’ HA) - (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) - (3’ HA) -3’ ; wherein 5’ HA is the 5’ homology arm; 3’ HA is the 3’ homology arm; 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI is a translation initiation sequence; Z1 is a therapeutic protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2 (FIGs. 10A and 10B, lower) .

[0241] FIGs. 10A and 10B show scarless splicing wherein the sequence elements within the target sequence serve as E1 and E2. In some embodiments, the 5’ terminal region of the target sequence (e.g., part of Z1B) can serve as E2. In some embodiments, the 3’ terminal region of the target sequence (e.g., part of Z1A) can serve as E1.

[0242] E1 and E2: In some embodiments, the RNAs provided herein can further comprise one or two exon (s) flanking the target sequence that remains present in the resulting circRNA. Without being bound by theory, the presence of the exon (s) may improve the self-splicing efficiency. As such, in some embodiments, provided herein are RNAs comprising (1) an exon fragment 2 (E2) between the 3’ intron fragment and the target sequence; (2) an exon fragment 1 (E1) between the target sequence and the 5’ intron fragment; or (3) both (1) and (2) . Accordingly, in some embodiments, provided herein are non-naturally occurring RNAs comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) E2; (3) a target sequence; (4) E1; and (5) a 5’ intron fragment. In some embodiments, provided herein are non-naturally occurring RNAs comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) E2; (3) a target sequence; and (4) a 5’ intron fragment. In some embodiments, provided herein are non-naturally occurring RNAs comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence; (3) E1; and (4) a 5’ intron fragment.

[0243] Scarless splicing generates circRNAs consisting of the resulting target sequences. Near scarless splicing generates resulting circRNAs that also contain E1 and / E2, in addition to the resulting targeting sequences.

[0244] In some embodiments, RNAs (or cRNAzymes) provided herein comprises a linker sequence (L) . In some embodiments, the linker sequence is 3-300 nucleic acid residues in length. In some embodiments the linker sequence is about 3-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-125, 125-150, 150-175, 175-200, 200-225, 225-250, 250-275 or 275-300 nucleic acid sequences in length. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-100. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-50. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-20. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is an integer selected from 1-10. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is 1, 2, 3, 4 or 5. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, wherein N is 1, 2 or 3. In some embodiments, the linker sequence is about 3 nucleic acid residues in length.

[0245] In some embodiments, the linker sequence comprises the nucleic acid sequence of RCC, wherein R is a guanine or an adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCC, wherein R is a guanine or an adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCCRCC, wherein R is a guanine or an adenine. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 226. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 227. In some embodiments, the linker has the polynucleotide sequence of SEQ ID NO: 248.

[0246] In some embodiments, the linker comprises a nucleic acid sequence that encodes a 5’ UTR, 3’ UTR, poly-A sequence, polyA-C sequence, poly-C sequence, poly-U sequence, poly-G sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulation elements (e.g., a Kozak sequence) , a protein binding site (e.g., PTBP1 or HUR) a non-natural nucleotide, or a non-nucleotide chemical-linker sequence.

[0247] In some embodiments, the resulting target sequences comprise a 3’ UTR. In some embodiments, the 3’ UTR can be the 3’ UTR from human beta globin, human alpha globin xenopus beta globin, xenopus alpha globin, human prolactin, human GAP-43, human eEFlal, human Tau, human TNFa, dengue virus, hantavirus small mRNA, bunyavirus small mRNA, turnip yellow mosaic virus, hepatitis C virus, rubella virus, tobacco mosaic virus, human IL-8, human actin, human GAPDH, human tubulin, hibiscus chlorotic linsgspot virus, woodchuck hepatitis virus post translationally regulated element, sindbis virus, turnip crinkle virus, tobacco etch virus, or Venezuelan equine encephalitis virus.

[0248] In some embodiments, the resulting target sequences comprise a 5’ UTR. In some embodiments, the 5’ UTR can be the 5’ UTR from human beta globin, Xenopus laevis beta globin, human alpha globin, Xenopus laevis alpha globin, rubella virus, tobacco mosaic virus, mouse Gtx, dengue virus, heat shock protein 70 kDa protein 1A, tobacco alcohol dehydrogenase, tobacco etch virus, turnip crinkle virus, or the adenovirus tripartite leader.

[0249] In some embodiments, the resulting target sequences comprise a polyA region. In some embodiments the polyA region is at least 30 nucleotides or at least 60 nucleotides in length.

[0250] In some embodiments, the resulting target sequences of the RNAs (or cRNAzymes) provided herein comprise one expression sequence. In some embodiments, the resulting target sequences of the RNAs (or cRNAzymes) provided herein comprise more than one expression sequence, e.g., 2, 3, 4, or 5 expression sequences.

[0251] In some embodiments, the resulting target sequences of the RNAs (or cRNAzymes) provided herein encode a protein that is made up of subunits that are encoded by more than one gene. For example, the protein can be a heterodimer, wherein each chain or subunit of the protein is encoded by a separate gene. It is possible that more than one circRNAs are delivered in the transfer vehicle and each contains a resulting target sequence encoding a separate subunit of the protein. In some embodiments, separate circRNAs encoding the individual subunits can be administered in separate transfer vehicles. Alternatively, the resulting target sequences of the RNAs (or cRNAzymes) provided herein can contain more than one expression cassette and encode more than one subunit. 5.15 Circularization and circRNAs

[0252] Provided herein are also circRNAs prepared by the self-splicing of the RNAs (cRNAzymes) disclosed herein. In some embodiments, the circRNAs provided herein comprise one or more IRES sequences disclosed herein and are configured for persistent expression in a cell of a subject in vivo. In some embodiments, the circRNAs are configured such that expression of the one or more expression sequences in the cell at a later time point is equal to or higher than an earlier time point. In such embodiments, the expression of the one or more expression sequences can be either maintained at a relatively stable level or can increase over time. The expression of the expression sequences can be relatively stable for an extended period of time. For instance, in some cases, the expression of the one or more expression sequences in the cell over a time period of at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days does not decrease by 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%. In some cases, in some cases, the expression of the one or more expression sequences in the cell is maintained at a level that does not vary by more than 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5%for at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days.

[0253] In some embodiments, the circRNAs are scarless. In some embodiments, the circRNA are near-scarless.

[0254] In some embodiments, the circRNAs disclosed herein can be of any length or size. In some embodiments the circRNA is between 300 and 10000, 400 and 9000, 500 and 8000, 600 and 7000, 700 and 6000, 800 and 5000, 900 and 5000, 1000 and 5000, 1100 and 5000, 1200 and 5000, 1300 and 5000, 1400 and 5000, and / or 1500 and 5000 nucleotides in length. In some embodiments, the circRNAs disclosed herein can be at least 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1000 nt, 1100 nt, 1200 nt, 1300 nt, 1400 nt, 1500 nt, 2000 nt, 2500 nt, 3000 nt, 3500 nt, 4000 nt, 4500 nt, or 5000 nt in length. In some embodiments, the circRNA is no more than 3000 nt, 3500 nt, 4000 nt, 4500 nt, 5000 nt, 6000 nt, 7000 nt, 8000 nt, 9000 nt, or 10000 nt in length. In some embodiments, circRNAs disclosed herein can be about 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1000 nt, 1100 nt, 1200 nt, 1300 nt, 1400 nt, 1500 nt, 2000 nt, 2500 nt, 3000 nt, 3500 nt, 4000 nt, 4500 nt, 5000 nt, 6000 nt, 7000 nt, 8000 nt, 9000 nt, or 10000 nt in length. In some embodiments, circRNAs disclosed herein can be at least 500 nucleotides in length, at least 1000 nucleotides in length, or at least 1500 nucleotides in length.

[0255] The circRNAs provided herein are produced by self-splicing of the RNAs (or cRNAzymes) provided herein, which has group II intron activity. In some embodiments, RNAs (or cRNAzymes) provided herein are produced by transcription using a vector provided herein as a template. In some embodiments, RNAs (or cRNAzymes) provided herein are produced by run-off transcription. In some embodiments, RNAs (or cRNAzymes) provided herein are produced by in vitro transcription.

[0256] Self-splicing of group II introns needs to be accomplished under high-salinity conditions, and does not require the introduction of GTP. As such, in some embodiments, the buffer used in the self-splicing reaction comprises 10 mM to 100 mM, such as 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, and 100 mM divalent magnesium ions, such as MgCl2. The self-splicing buffer can also comprise 10 mM to 100 mM, such as 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, and 100 mM NaCl.

[0257] In some embodiments, the self-splicing reaction is performed in vitro for about 5 min to about 1 h, such as about 5 min, about 10 min, about 15 min, about 20 min, about 25 min, about 30 min, about 35 min, about 40 min, about 45 min, about 50 min, about 55 min, and about 1 h.

[0258] In some embodiments, the self-splicing reaction is performed at a temperature between 20 and 60 ℃, between 20 and 50 ℃, between 20 and 40 ℃, between 20 and 30 ℃, between 30 and 40 ℃, between 40 and 50 ℃, or between 50 and 60 ℃.

[0259] In some embodiments, the precursor RNAs disclosed herein are capable of achieving a circularization rate of at least 30%, such as a circularization rate of at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, and at least 95%.

[0260] Provided herein are circRNAs prepared by self-splicing the RNAs (or cRNAzymes) disclosed herein. In some embodiments, the circRNAs disclosed herein are purified. Purification includes, but is not limited to, the removal of non-circularized linear RNAs, dsRNAs, and other unwanted components. In some embodiments, the circRNAs disclosed herein are purified before being transfected into cells. The phosphate groups at both ends of a linear RNA and some dsRNAs might activate the RIG-1 signaling pathway, and the immune response resulted from RIG-1 signaling can lead to the degradation of exogenous RNAs, thus affecting the function of circular RNAs.

[0261] For illustrative purposes, the methods of purification can include any of the following: enzymatic treatment; chromatography, including but not limited to affinity column chromatography, reversed-phase silica gel column liquid chromatography, gel filtration chromatography, and gel exclusion liquid chromatography; and electrophoresis, including but not limited to gel electrophoresis such as agarose gel electrophoresis, and capillary electrophoresis; and any combination thereof. Methods for removing linear RNAs, for example, include enzymatic treatment, such as treatment with RNase R; and chromatography, such as high performance liquid chromatography (HPLC) . Methods for removing terminal phosphate groups, for example, include treatment with alkaline phosphatases, such as calf intestinal alkaline phosphatase (CIP) .

[0262] In some embodiments, purification comprises one or more of the following steps: phosphatase treatment, HPLC size exclusion purification, and RNase R digestion. In some embodiments, purification comprises the following steps in order: RNase R digestion, phosphatase treatment, and HPLC size exclusion purification. In some embodiments, purification comprises reverse phase HPLC. In some embodiments, a purified composition contains less double stranded RNA, DNA splints, triphosphorylated RNA, phosphatase proteins, protein ligases, capping enzymes and / or nicked RNA than unpurified RNA. 5.16 Methods of uses

[0263] The nucleic acids and vectors disclosed herein can be used for a variety of purposes. For example, the nucleic acids and vectors disclosed herein can be used for cap-independent protein expression in vitro or in vivo.

[0264] In some embodiments, the methods for protein expression comprises translation of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%of the total length of the circRNAs into polypeptides. In some embodiments, the methods for protein expression comprises translation of the circRNAs into polypeptides of at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 50 amino acids, at least 100 amino acids, at least 150 amino acids, at least 200 amino acids, at least 250 amino acids, at least 300 amino acids, at least 400 amino acids, at least 500 amino acids, at least 600 amino acids, at least 700 amino acids, at least 800 amino acids, at least 900 amino acids, or at least 1000 amino acids. In some embodiments, the methods for protein expression comprises translation of the circRNAs into polypeptides of about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 50 amino acids, about 100 amino acids, about 150 amino acids, about 200 amino acids, about 250 amino acids, about 300 amino acids, about 400 amino acids, about 500 amino acids, about 600 amino acids, about 700 amino acids, about 800 amino acids, about 900 amino acids, or about 1000 amino acids.

[0265] In some embodiments, the translation of the at least a region of the circRNAs disclosed herein takes place in vitro, such as rabbit reticulocyte lysate. In some embodiments, the translation of the at least a region of the circRNAs disclosed herein takes place in vivo, for instance, after transfection of a eukaryotic cell, or transformation of a prokaryotic cell such as a bacterial cell.

[0266] In some embodiments, the nucleic acids and vectors disclosed herein can be used for expressing a protein in a cell. Accordingly, in some embodiments, provided herein are methods of protein expression comprising (a) subjecting the RNAs (or cRNAzymes) disclosed herein to a self-splicing circularization reaction to form a circRNA, wherein the circRNA comprises a protein-encoding target sequence operably linked to an IRES sequence disclosed herein, and (b) transfecting the cell with the circRNA. Additionally, in some embodiments, provided herein are methods of protein expression in a cell comprising transfecting the cell with a vector that encodes the RNAs (or cRNAzymes) disclosed herein, which can self-splice in the cell to form a circRNA comprising a protein-encoding target sequence operably linked to an IRES sequence disclosed herein. Furthermore, in some embodiments, the cRNAzymes can be prepared in vitro by e.g., in vitro transcription and transfected into the cell. Upon self-splicing in the cell, the cRNAzymes form circRNAs that encode the protein of interest, and the protein is then expressed in vivo.

[0267] The cell can be a eukaryotic cell. In some embodiments, the cell is a hepatocyte, epithelial cell, hematopoietic cell, epithelial cell, endothelial cell, lung cell, bone cell, stem cell, mesenchymal cell, neural cell (e.g., meninge, astrocyte, motor neuron, cell of the dorsal root ganglia and anterior horn motor neuron) , photoreceptor cell (e.g., rod and cone) , retinal pigmented epithelial cell, secretory cell, cardiac cell, adipocyte, vascular smooth muscle cell, cardiomyocyte, skeletal muscle cell, beta cell, pituitary cell, synovial lining cell, ovarian cell, testicular cell, fibroblast, B cell, dendritic cell, reticulocyte, granulocyte, tumor cell, NK cell, liver starlet cell, HEK293, HEK293T, HeLa, MCF7, PC3, A549, NCI-H727, HCT-116, MCF10A, HPReC, FHC, immortalized cell lines, primary cell, yeast cell, Saccharomyces cerevisiae, Pichia pastoris, bacteria cell, Escherichia coli, insect cell, Spodoptera frugiperda sf9, Mimic Sf9, sf21, or Drosophila S2.

[0268] In some embodiments, the cell is an epithelial cell, a myoblast, a neuroblast, a monocyte or a fibroblast. In some embodiments, the cell is an epithelial cell. In some embodiments, the cell is a myoblast. In some embodiments, the cell is a neuroblast. In some embodiments, the cell is a monocyte. In some embodiments, the cell is a fibroblast.

[0269] The nucleic acids disclosed herein can be delivered into cells or animals using any of a variety of delivery systems. For example, the delivery system is selected from one or more of a group of: liposomes, polyethyleneimine (PEI) , metal-organic frameworks (MOFs) , lipid nanoparticles (LNPs) , polycations, blood glycoproteins, red blood cell transport vehicles, Au nanoparticle (AuNP) vehicles, magnetic nanoparticle vehicles, carbon nanotubes, graphene molecular vehicles, quantum dot material vehicles, upconversion nanoparticles, layered double hydroxide material vehicles, silica nanoparticles, and calcium phosphate. In some embodiments, the nucleic acids disclosed herein can be transfected into a cell using, for example, lipofection or electroporation.

[0270] In some embodiments, the present disclosure provides methods of in vivo expression of a protein of interest in a subject, comprising administering a nucleic acid to a cell of the subject wherein the nucleic acid comprises the one or more expression sequences encoding the protein of interest operably linked to an IRES sequence disclosed herein; wherein the protein of interest is expressed from the nucleic acid in the cell. In some embodiments, the nucleic acid is a circRNA. In some embodiments, the circRNA is configured such that expression of the one or more expression sequences in the cell at a later time point is equal to or higher than an earlier time point. Additionally, in some embodiments, provided herein are methods of in vivo expression of a protein of interest in a subject, comprising: administering to the subject a vector that encodes the cRNAzyme which can self-splice to form the circRNA that encodes the protein. Once the vector is administered, it can be transcribed in vivo, and the transcribed cRNAzyme can self-splice to form the circRNA that encodes the protein of interest, allowing the in vivo expression of the protein. Furthermore, in some embodiments, the cRNAzymes can be prepared in vitro and administered to a subject. Upon in vivo self-splicing, the cRNAzymes form circRNAs that comprises a sequence encode the protein of interest operably linked to an IRES sequence disclosed herein, and the protein is then expressed in vivo. Methods of the administering the nucleic acids or vectors disclosed herein to a subject are known to a person of ordinary skill in the art. Some exemplary methods are provided herein.

[0271] In some embodiments, the nucleic acid is configured such that expression of the one or more expression sequences in the cell over a time period of at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days does not decrease by greater than about 40%. In some embodiments, the nucleic acid is configured such that expression of the one or more expression sequences in the cell is maintained at a level that does not vary by more than about 40%for at least 7, 8, 9, 10, 12, 14, 16, 18, 20, 22, 23 or more days. In some embodiments, the administration of the nucleic acid is conducted using any delivery method described herein. In some embodiments, the circular polyribonucleotide is administered to the subject via intravenous injection. In some embodiments, the administration of the circRNA includes, but is not limited to, prenatal administration, neonatal administration, postnatal administration, oral, by injection (e.g., intravenous, intraarterial, intraperotoneal, intradermal, subcutaneous and intramuscular) , by ophthalmic administration and by intranasal administration.

[0272] In some embodiments, the methods for protein expression comprise modification, folding, or other post-translation modification of the translation product. In some embodiments, the methods for protein expression comprise post-translation modification in vivo, e.g., via cellular machinery. 5.17 Pharmaceutical Compositions

[0273] In some embodiments, the nucleic acids or vectors disclosed herein can be used as a therapeutic in the treatment of a disease or a condition.

[0274] As used herein, the term “treat” refers to executing a protocol or plan, which can include administering one or more drugs or active agents to a patient, in an effort to alleviate signs or symptoms of the disease or the recurrence of the disease. Desirable effects of treatment include decreasing the rate of disease progression, ameliorating or palliating the disease state, and remission, increased survival, improved quality of life or improved prognosis. Alleviation or prevention can occur prior to signs or symptoms of the disease or condition appearing, as well as after their appearance. As used herein and understood in the art, a “treatment” does not require complete alleviation of signs or symptoms, and does not require a cure.

[0275] As used herein, the term “therapeutic beneficial” or “therapeutically effective” when used in connection with a therapeutic refers to the property of the therapeutic that promotes or enhances the well-being of the subject. This includes, but is not limited to, a reduction in the frequency, severity, or rate of progression of the signs or symptoms of a disease. For example, treatment of cancer may involve, for example, a reduction in the size of a tumor, a reduction in the invasiveness of a tumor, reduction in the growth rate of the cancer, or a reduction in the rate of metastasis or recurrence. Treatment of cancer can also refer to prolonging survival of a subject with cancer.

[0276] As used herein, the term “pharmaceutical or pharmacologically acceptable” refers to molecular entities and compositions that do not produce an adverse, allergic, or other untoward reaction when administered to an animal, such as a human, as appropriate. For animal (e.g., human) administration, it will be understood that preparations should meet sterility, pyrogenicity, general safety, and purity standards as required, e.g., by the FDA Office of Biological Standards.

[0277] As used herein, the term “pharmaceutically acceptable carrier” includes any and all aqueous biocompatible solvents (e.g., saline solutions, phosphate buffered saline, parenteral vehicles, such as sodium chloride, Ringer's dextrose, etc. ) , antioxidants, preservatives (e.g., antibacterial or antifungal agents, anti-oxidants, chelating agents, and inert gases) , isotonic agents, such like materials and combinations thereof, as would be known to one of ordinary skill in the art. The pH and exact concentration of the various components in a pharmaceutical composition are adjusted according to well-known parameters.

[0278] In some embodiments, provided herein are compositions (e.g., pharmaceutical compositions) comprising a therapeutic agent provided herein. In some embodiments, the therapeutic agent is a nucleic acid (e.g., circRNA) provided herein. In some embodiments, the therapeutic agent is a vector provided herein. In some embodiments, the therapeutic agent is a cell comprising a nucleic acid or vector provided herein. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the compositions provided herein comprise a therapeutic agent provided herein in combination with other pharmaceutically active agents or drugs. In some embodiments, the pharmaceutical composition comprises a cell provided herein or populations thereof.

[0279] With respect to pharmaceutical compositions, the pharmaceutically acceptable carrier can be any of those conventionally used and is limited only by chemico-physical considerations, such as solubility and lack of reactivity with the active agent (s) , and by the route of administration. The pharmaceutically acceptable carriers described herein, for example, vehicles, adjuvants, excipients, and diluents, are well-known to those skilled in the art and are readily available to the public. It is preferred that the pharmaceutically acceptable carrier be one which is chemically inert to the therapeutic agent (s) and one which has no detrimental side effects or toxicity under the conditions of use.

[0280] The choice of carrier will be determined in part by the particular therapeutic agent, as well as by the particular method used to administer the therapeutic agent. Accordingly, there are a variety of suitable formulations of the pharmaceutical compositions provided herein.

[0281] In some embodiments, the pharmaceutical composition comprises a preservative. In some embodiments, suitable preservatives may include, for example, methylparaben, propylparaben, sodium benzoate, and benzalkonium chloride. Optionally, a mixture of two or more preservatives may be used. The preservative or mixtures thereof are typically present in an amount of about 0.0001%to about 2%by weight of the total composition.

[0282] In some embodiments, the pharmaceutical composition comprises a buffering agent. In some embodiments, suitable buffering agents may include, for example, citric acid, sodium citrate, phosphoric acid, potassium phosphate, and various other acids and salts. A mixture of two or more buffering agents optionally may be used. The buffering agent or mixtures thereof are typically present in an amount of about 0.001%to about 4%by weight of the total composition.

[0283] In some embodiments, the concentration of therapeutic agent in the pharmaceutical composition can vary, e.g., less than about 1%, or at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50%or more by weight, and can be selected primarily by fluid volumes, and viscosities, in accordance with the particular mode of administration selected.

[0284] The following formulations for oral, aerosol, parenteral (e.g., subcutaneous, intravenous, intraarterial, intramuscular, intradermal, intraperitoneal, and intrathecal) , and topical administration are merely exemplary and are in no way limiting. More than one route can be used to administer the therapeutic agents provided herein, and in some instances, a particular route can provide a more immediate and more effective response than another route.

[0285] Formulations suitable for oral administration can comprise or consist of (a) liquid solutions, such as an effective amount of the therapeutic agent dissolved in diluents, such as water, saline, or orange juice; (b) capsules, sachets, tablets, lozenges, and troches, each containing a predetermined amount of the active ingredient, as solids or granules; (c) powders; (d) suspensions in an appropriate liquid; and (e) suitable emulsions. Liquid formulations may include diluents, such as water and alcohols, for example, ethanol, benzyl alcohol and the polyethylene alcohols, either with or without the addition of a pharmaceutically acceptable surfactant. Capsule forms can be of the ordinary hard or soft shelled gelatin type containing, for example, surfactants, lubricants, and inert fillers, such as lactose, sucrose, calcium phosphate, and com starch. Tablet forms can include one or more of lactose, sucrose, mannitol, com starch, potato starch, alginic acid, microcrystalline cellulose, acacia, gelatin, guar gum, colloidal silicon dioxide, croscarmellose sodium, talc, magnesium stearate, calcium stearate, zinc stearate, stearic acid, and other excipients, colorants, diluents, buffering agents, disintegrating agents, moistening agents, preservatives, flavoring agents, and other pharmacologically compatible excipients. Lozenge forms can comprise the therapeutic agent with a flavorant, usually sucrose, acacia or tragacanth. Pastilles can comprise the therapeutic agent with an inert base, such as gelatin and glycerin, or sucrose and acacia, emulsions, gels, and the like containing, in addition to, such excipients as are known in the art.

[0286] Formulations suitable for parenteral administration include aqueous and nonaqueous isotonic sterile injection solutions, which can contain antioxidants, buffers, bacteriostats, and solutes that render the formulation isotonic with the blood of the intended recipient, and aqueous and nonaqueous sterile suspensions that can include suspending agents, solubilizers, thickening agents, stabilizers, and preservatives. In some embodiments, the therapeutic agents provided herein can be administered in a physiologically acceptable diluent in a pharmaceutical carrier, such as a sterile liquid or mixture of liquids including water, saline, aqueous dextrose and related sugar solutions, an alcohol such as ethanol or hexadecyl alcohol, a glycol such as propylene glycol or polyethylene glycol, dimethylsulfoxide, glycerol, ketals such as 2, 2-dimethyl-1, 3 -dioxolane-4-methanol, ethers, poly (ethyleneglycol) 400, oils, fatty acids, fatty acid esters or glycerides, or acetylated fatty acid glycerides with or without the addition of a pharmaceutically acceptable surfactant such as a soap or a detergent, suspending agent such as pectin, carbomers, methylcellulose, hydroxypropylmethylcellulose, or carboxymethylcellulose, or emulsifying agents and other pharmaceutical adjuvants.

[0287] Oils, which can be used in parenteral formulations in some embodiments, include petroleum, animal oils, vegetable oils, or synthetic oils. Specific examples of oils include peanut, soybean, sesame, cottonseed, com, olive, petrolatum, and mineral oil. Suitable fatty acids for use in parenteral formulations include oleic acid, stearic acid, and isostearic acid. Ethyl oleate and isopropyl myristate are examples of suitable fatty acid esters.

[0288] Suitable soaps for use in some embodiments of parenteral formulations include fatty alkali metal, ammonium, and triethanolamine salts, and suitable detergents include (a) cationic detergents such as, for example, dimethyl dialkyl ammonium halides and alkyl pyridinium halides, (b) anionic detergents such as, for example, alkyl, aryl, and olefin sulfonates, alky, olefin, ether, and monoglyceride sulfates, and sulfosuccinates, (c) nonionic detergents such as, for example, fatty amine oxides, fatty acid alkanolamides, and polyoxyethylenepolypropylene copolymers, (d) amphoteric detergents such as, for example, alkyl-b -aminopropionates , and 2-alkyl-imidazoline quaterary ammonium salts, and (e) mixtures thereof.

[0289] In some embodiments, the parenteral formulations will contain, for example, from about 0.5%to about 25%by weight of the therapeutic agent in solution. Preservatives and buffers may be used. In order to minimize or eliminate irritation at the site of injection, such compositions may contain one or more nonionic surfactants having, for example, a hydrophile-lipophile balance (HLB) of from about 12 to about 17. The quantity of surfactant in such formulations will typically range, for example, from about 5%to about 15%by weight. Suitable surfactants include polyethylene glycol, sorbitan fatty acid esters such as sorbitan monooleate, and high molecular weight adducts of ethylene oxide with a hydrophobic base formed by the condensation of propylene oxide with propylene glycol. The parenteral formulations can be presented in unit-dose or multi-dose sealed containers, such as ampoules or vials, and can be stored in a freeze-dried (lyophilized) condition requiring only the addition of a sterile liquid excipient, for example, water, for injections, immediately prior to use. Extemporaneous injection solutions and suspensions can be prepared from sterile powders, granules, and tablets of the kind previously described.

[0290] In some embodiments, injectable formulations are provided herein. The requirements for effective pharmaceutical carriers for injectable compositions are well-known to those of ordinary skill in the art (see, e.g., PHARMACEUTICS AND PHARMACY PRACTICE, J. B. Lippincott Company, Philadelphia, PA, Banker and Chalmers, eds., pages 238-250 (1982) , and ASHP Handbook on Injectable Drugs, Toissel, 4th ed., pages 622-630 (1986) ) .

[0291] In some embodiments, topical formulations are provided herein. Topical formulations, including those that are useful for transdermal drug release, are suitable in the context of certain embodiments provided herein for application to skin. In some embodiments, the therapeutic agent alone or in combination with other suitable components, can be made into aerosol formulations to be administered via inhalation. These aerosol formulations can be placed into pressurized acceptable propellants, such as dichlorodifluoromethane, propane, nitrogen, and the like. They also may be formulated as pharmaceuticals for non-pressured preparations, such as in a nebulizer or an atomizer. Such spray formulations also can be used to spray mucosa.

[0292] In some embodiments, the therapeutic agents provided herein can be formulated as inclusion complexes, such as cyclodextrin inclusion complexes, or liposomes. Liposomes can serve to target the therapeutic agents to a particular tissue. Liposomes also can be used to increase the half-life of the therapeutic agents. Many methods are available for preparing liposomes, as described in, for example, Szoka et al., Ann. Rev. Biophys. Bioeng., 9, 467 (1980) and U.S. Patents 4,235,871, 4,501,728, 4,837,028, and 5,019,369.

[0293] In some embodiments, the therapeutic agents provided herein are formulated in time-released, delayed release, or sustained release delivery systems such that the delivery of the composition occurs prior to, and with sufficient time to cause, sensitization of the site to be treated. Such systems can avoid repeated administrations of the therapeutic agent, thereby increasing convenience to the subject and the physician, and can be particularly suitable for certain composition embodiments provided herein. In one embodiment, the compositions of the disclosure are formulated such that they are suitable for extended-release of the circRNA contained therein. Such extended-release compositions may be conveniently administered to a subject at extended dosing intervals. For example, in one embodiment, the compositions of the present disclosure are administered to a subject twice a day, daily or every other day. In an embodiment, the compositions of the present disclosure are administered to a subject twice a week, once a week, every ten days, every two weeks, every three weeks, every four weeks, once a month, every six weeks, every eight weeks, every three months, every four months, every six months, every eight months, every nine months or annually.

[0294] In some embodiments, a protein encoded by the nucleic acid described herein is produced by a target cell for sustained amounts of time. For example, the protein can be produced for more than one hour, more than four, more than six, more than 12, more than 24, more than 48 hours, or more than 72 hours after administration. In some embodiments the therapeutic protein is expressed at a peak level about six hours after administration. In some embodiments the expression of the therapeutic protein is sustained at least at a therapeutic level. In some embodiments the therapeutic protein is expressed at least at a therapeutic level for more than one, more than four, more than six, more than 12, more than 24, more than 48, or more than 72 hours after administration. In some embodiments, the therapeutic protein is detectable at a therapeutic level in patient serum or tissue (e.g., liver or lung) . In some embodiments, the level of detectable therapeutic protein is from continuous expression from the nucleic acid composition over periods of time of more than one, more than four, more than six, more than 12, more than 24, more than 48, or more than 72 hours after administration.

[0295] In some embodiments, a protein encoded by a nucleic acid described herein is produced at levels above normal physiological levels. The level of protein can be increased as compared to a control. In some embodiments, the control is the baseline physiological level of the therapeutic protein in a normal individual or in a population of normal individuals. In other embodiments, the control is the baseline physiological level of the therapeutic protein in an individual having a deficiency in the relevant protein or polypeptide or in a population of individuals having a deficiency in the relevant protein or polypeptide. In some embodiments, the control can be the normal level of the relevant protein or polypeptide in the individual to whom the composition is administered. In other embodiments, the control is the expression level of the therapeutic protein upon other therapeutic intervention, e.g., upon direct injection of the corresponding therapeutic protein, at one or more comparable time points.

[0296] In some embodiments, the levels of a protein encoded by a nucleic acid described herein are detectable at 3 days, 4 days, 5 days, or 1 week or more after administration. Increased levels of secreted protein may be observed in the serum and / or in a tissue (e.g., liver or lung) .

[0297] In some embodiments, the method yields a sustained circulation half-life of a protein encoded by a nucleic acid described herein. For example, the protein can be detected for hours or days longer than the half-life observed via subcutaneous injection of the protein or mRNA encoding the protein. In some embodiments, the half-life of the protein is 1 day, 2 days, 3 days, 4 days, 5 days, or 1 week or more.

[0298] Many types of release delivery systems are available and known to those of ordinary skill in the art. They include polymer based systems such as poly (lactide-glycolide) , copolyoxalates, polycaprolactones, polyesteramides, polyorthoesters, polyhydroxybutyiic acid, and polyanhydrides. Microcapsules of the foregoing polymers containing drugs are described in, for example, U.S. Patent 5,075,109. Delivery systems also include non-polymer systems that are lipids including sterols such as cholesterol, cholesterol esters, and fatty acids or neutral fats such as mono-di-and tri-glycerides; hydrogel release systems; sylastic systems; peptide based systems: wax coatings; compressed tablets using conventional binders and excipients; partially fused implants; and the like. Specific examples include, but are not limited to: (a) erosional systems in which the active composition is contained in a form within a matrix such as those described in U.S. Patents 4,452,775, 4,667,014, 4,748,034, and 5,239,660 and (b) diffusional systems in which an active component permeates at a controlled rate from a polymer such as described in U.S. Patents 3,832,253 and 3,854,480. In addition, pump-based hardware delivery systems can be used, some of which are adapted for implantation.

[0299] In some embodiments, the therapeutic agent can be conjugated either directly or indirectly through a linking moiety to a targeting moiety. Methods for conjugating therapeutic agents to targeting moieties is known in the art. See, for instance, Wadwa et al., J. Drug Targeting 3: 111 (1995) and U.S. Patent 5,087,616.

[0300] In some embodiments, the therapeutic agents provided herein are formulated into a depot form, such that the manner in which the therapeutic agent is released into the body to which it is administered is controlled with respect to time and location within the body (see, for example, U.S. Patent 4,450,150) . Depot forms of therapeutic agents can be, for example, an implantable composition comprising the therapeutic agents and a porous or non-porous material, such as a polymer, wherein the therapeutic agents are encapsulated by or diffused throughout the material and / or degradation of the non-porous material. The depot is then implanted into the desired location within the body and the therapeutic agents are released from the implant at a predetermined rate. 5.18 Target Cells

[0301] Provided herein are cells comprising the nucleic acids or vectors provided herein. Provided herein are also cells comprising the nucleic acids or vectors disclosed herein. Also provided herein are methods of delivering the nucleic acids or vectors disclosed herein to a cell. As used herein, the term “target cell” refers to the cell or type of cells to which the nucleic acids or vectors disclosed herein are intended to deliver.

[0302] In some embodiments, the nucleic acids or vectors disclosed herein can be formulated using liposomes, lipoplexes, lipid nanoparticles, polymer based delivery systems, and viral vectors. In some embodiments, the nucleic acids or vectors can be formulated in a lipid nanoparticle such as those described in WO2012170930, herein incorporated by reference in its entirety. In some embodiments, the lipid can be a cleavable lipid such as those described in WO2012170889, herein incorporated by reference in its entirety. In some embodiments, the pharmaceutical compositions of the nucleic acids or vectors can include at least one of the PEGylated lipids described in WO2012099755, herein incorporated by reference. In some embodiments, a lipid nanoparticle formulation can be formulated by the methods described in WO2011127255 or WO2008103276, each of which is herein incorporated by reference in its entirety. A lipid nanoparticle can be coated or associated with a co-polymer such as, but not limited to, a block co-polymer, such as a branched polyether-polyamide block copolymer described in WO2013012476, herein incorporated by reference in its entirety. Liposomes, lipoplexes, or lipid nanoparticles can be used to improve the efficacy of the nucleic acids or vectors directed protein production as these formulations may be able to increase cell transfection efficiency, increase the in vivo or in vitro half-life of the nucleic acids or vectors, and / or allow for controlled release.

[0303] The present disclosure contemplates the discriminatory targeting of target cells and tissues by both passive and active targeting means. The phenomenon of passive targeting exploits the natural distribution patterns of a transfer vehicle in vivo without relying upon the use of additional excipients or means to enhance recognition of the transfer vehicle by target cells. For example, transfer vehicles which are subject to phagocytosis by the cells of the reticulo-endothelial system are likely to accumulate in the liver or spleen, and accordingly, may provide a means to passively direct the delivery of the compositions to such target cells.

[0304] Alternatively, the present disclosure contemplates active targeting, which involves the use of targeting moieties that can be bound (either covalently or non-covalently) to the transfer vehicle to encourage localization of such transfer vehicle at certain target cells or target tissues. For example, targeting can be mediated by the inclusion of one or more endogenous targeting moieties in or on the transfer vehicle to encourage distribution to the target cells or tissues. Recognition of the targeting moiety by the target tissues actively facilitates tissue distribution and cellular uptake of the transfer vehicle and / or its contents in the target cells and tissues (e.g., the inclusion of an apolipoprotein-E targeting ligand in or on the transfer vehicle encourages recognition and binding of the transfer vehicle to endogenous low density lipoprotein receptors expressed by hepatocytes) . As provided herein, the composition can comprise a moiety capable of enhancing affinity of the composition to the target cell. Targeting moieties can be linked to the outer bilayer of the lipid particle during formulation or post-formulation. These methods are well known in the art. In addition, some lipid particle formulations can employ fusogenic polymers such as PEAA, hemagglutinin, other lipopeptides (see U.S. Pat. No. 6,417,326, which is incorporated herein by reference) and other features useful for in vivo and / or intracellular delivery. In other some embodiments, the compositions of the present disclosure demonstrate improved transfection efficacies, and / or demonstrate enhanced selectivity towards target cells or tissues of interest. Contemplated therefore are compositions which comprise one or more moieties (e.g., peptides, aptamers, oligonucleotides, a vitamin or other molecules) that are capable of enhancing the affinity of the compositions and their nucleic acid contents for the target cells or tissues. Suitable moieties can optionally be bound or linked to the surface of the transfer vehicle. In some embodiments, the targeting moiety can span the surface of a transfer vehicle or be encapsulated within the transfer vehicle. Suitable moieties and are selected based upon their physical, chemical or biological properties (e.g., selective affinity and / or recognition of target cell surface markers or features) . Cell-specific target sites and their corresponding targeting ligand can vary widely. Suitable targeting moieties are selected such that the unique characteristics of a target cell are exploited, thus allowing the composition to discriminate between target and non-target cells. For example, compositions of the disclosure may include surface markers (e.g., apolipoprotein-B or apolipoprotein-E) that selectively enhance recognition of, or affinity to hepatocytes (e.g., by receptor-mediated recognition of and binding to such surface markers) . As an example, the use of galactose as a targeting moiety would be expected to direct the compositions of the present disclosure to parenchymal hepatocytes, or alternatively the use of mannose containing sugar residues as a targeting ligand would be expected to direct the compositions of the present disclosure to liver endothelial cells (e.g., mannose containing sugar residues that may bind preferentially to the asialoglycoprotein receptor present in hepatocytes) . (See Hillery A M, et al.“Drug Delivery and Targeting: For Pharmacists and Pharmaceutical Scientists” (2002) Taylor &Francis, Inc. ) The presentation of such targeting moieties that have been conjugated to moieties present in the transfer vehicle (e.g., a lipid nanoparticle) therefore facilitate recognition and uptake of the compositions of the present disclosure in target cells and tissues. Examples of suitable targeting moieties include one or more peptides, proteins, aptamers, vitamins and oligonucleotides.

[0305] In some embodiments, the nucleic acids disclosed herein are formulated using viral vectors. Viral vectors can be derived from a variety of viruses including adenovirus, adeno-associated virus, lentivirus (e.g., HIV, FIV, and EIAV) , and herpes virus. Examples of commercially available viral vectors include pSilencer adeno (Ambion, Austin, Tex. ) and pLenti6 / BLOCK-iTTM-DEST (Invitrogen, Carlsbad, Calif. ) . Selection of viral vectors, methods for expressing the RNAs (or cRNAzymes) from the vector and methods of delivering the viral vector are within the ordinary skill of one in the art. In some embodiments, the viral vector is a recombinant AAV (rAAV) vector, such as those known in the art (PMID: 30245471, PMID: 33614232) .

[0306] The present disclosure also provides a delivery system comprising the nucleic acids or vectors disclosed herein. In some embodiments, the delivery system is any one of a liposome, a nanoparticle, a polymer based delivery system or a ligand-conjugate delivery system. In some embodiments, the ligand-conjugate delivery system comprises one or more of an antibody, a peptide, a sugar moiety, or a combination thereof.

[0307] In some embodiments, the delivery system of the present disclosure comprises nanoparticles comprising the nucleic acids or vectors disclosed herein.

[0308] In some embodiments, the nanoparticle comprises a polymer-based nanoparticle, a lipid-polymer based nanoparticle, a metal based nanoparticle, a carbon nanotube based nanoparticle, a nanocrystal or a polymeric micelle. In some embodiments, the polymer-based nanoparticle comprises a multiblock copolymer, a deblock copolymer, a polymeric micelle or a hyperbranched macromolecule. In some embodiments, the polymer-based nanoparticle comprises a multiblock copolymer a diblock copolymer. In some embodiments, the polymer-based nanoparticle is pH responsive. In some embodiments, the polymer-based nanoparticle further comprises a buffering component.

[0309] In some embodiments, the delivery system comprises a liposome. Liposomes are spherical vesicles having at least one lipid bilayer, and in some embodiments, an aqueous core. In some embodiments, the lipid bilayer of the liposome may comprise phospholipids. An exemplary but non-limiting example of a phospholipid is phosphatidylcholine, but the lipid bilayer may comprise additional lipids, such as phosphatidylethanolamine. Liposomes may be multilamellar, i.e. consisting of several lamellar phase lipid bilayers, or unilamellar liposomes with a single lipid bilayer. Liposomes can be made in a particular size range that makes them viable targets for phagocytosis. Liposomes can range in size from 20 nm to 100 nm, 100 nm to 400 nm, 1 μM and larger, or 200 nm to 3 μM. Examples of lipidoids and lipid-based formulations are provided in U.S. Pub. No. 20090023673. In some embodiments, the one or more lipids are one or more cationic lipids.

[0310] In some embodiments, the delivery system of the present disclosure is a polymer based nanoparticle. Polymer based nanoparticles comprise one or more polymers. In some embodiments, the one or more polymers comprise a polyester, poly (ortho ester) , poly (ethylene imine) , poly (caprolactone) , polyanhydride, poly (acrylic acid) , polyglycolide or poly (urethane) . In some embodiments, the one or more polymers comprise poly (lactic acid) (PLA) or poly (lactic-co-glycolic acid) (PLGA) . In some embodiments, the one or more polymers comprise poly (lactic-co-glycolic acid) (PLGA) . In some embodiments, the one or more polymers comprise poly (lactic acid) (PLA) . In some embodiments, the one or more polymers comprise polyalkylene glycol or a polyalkylene oxide. In some embodiments, the polyalkylene glycol is polyethylene glycol (PEG) or the polyalkylene oxide is polyethylene oxide (PEO) .

[0311] In some embodiments, the target cells can be deficient in a protein or enzyme of interest. For example, where it is desired to deliver a nucleic acid to a hepatocyte, the hepatocyte represents the target cell. In some embodiments, some cells can be preferentially, or specifically, targeted. In some embodiments, the preferentially or specifically targeted cells can be, for example, hepatocytes, epithelial cells, hematopoietic cells, epithelial cells, endothelial cells, lung cells, bone cells, stem cells, mesenchymal cells, neural cells (e.g., meninges, astrocytes, motor neurons, cells of the dorsal root ganglia and anterior horn motor neurons) , photoreceptor cells (e.g., rods and cones) , retinal pigmented epithelial cells, secretory cells, cardiac cells, adipocytes, vascular smooth muscle cells, cardiomyocytes, skeletal muscle cells, beta cells, pituitary cells, synovial lining cells, ovarian cells, testicular cells, fibroblasts, B cells, dendritic cells, reticulocytes, granulocytes and tumor cells, NK cells, liver starlet cells, HEK293, HEK293T, HeLa, MCF7, PC3, A549, NCI-H727, HCT-116, MCF10A, HPReC, FHC or other immortalized cell lines or primary cell lines. In some embodiments, the target cell is an epithelial cell, a myoblast, a neuroblast, a monocyte or a fibroblast. In some embodiments, the target cell is an epithelial cell. In some embodiments, the target cell is a myoblast. In some embodiments, the target cell is a neuroblast. In some embodiments, the target cell is a monocyte. In some embodiments, the target cell is a fibroblast.

[0312] In some embodiments, the compositions provided herein that comprise the nucleic acids or vectors disclosed herein can be optimized for certain target cells. The compositions provided herein can be prepared to preferentially distribute to and / or optimized for target cells such as in the heart, lungs, kidneys, liver, and spleen. In some embodiments, the compositions provided herein can distribute into the cells of the liver to facilitate the delivery and the subsequent expression of the circRNA comprised therein by the cells of the liver (e.g., hepatocytes) . The targeted cells can function as a biological “reservoir” or “depot” capable of producing, and systemically excreting a functional protein or enzyme. Accordingly, in some embodiments, the transfer vehicle can target hepatocytes and / or preferentially distribute to the cells of the liver upon delivery. In some embodiments, following transfection of the target hepatocytes, the circRNA loaded in the vehicle are translated and a functional protein product is produced, excreted and systemically distributed. In other embodiments, cells other than hepatocytes (e.g., lung, spleen, heart, ocular, or cells of the central nervous system) can serve as a depot location for protein production.

[0313] In some embodiments, the compositions provided herein facilitate a subject’s endogenous production of one or more functional proteins and / or enzymes. In some embodiments, the transfer vehicles comprise the nucleic acids disclosed herein which encode a deficient protein or enzyme. Upon distribution of such compositions to the target tissues and the subsequent transfection of such target cells, the exogenous nucleic acids loaded into the transfer vehicle (e.g., a lipid nanoparticle) can be translated in vivo to produce a functional protein or enzyme encoded by the exogenously administered nucleic acids (e.g., a protein or enzyme in which the subject is deficient) . Accordingly, the compositions provided herein exploit a subject's ability to translate exogenously-or recombinantly-prepared nucleic acids to produce an endogenously-translated protein or enzyme, and thereby produce (and where applicable excrete) a functional protein or enzyme. The expressed or translated proteins or enzymes can also be characterized by the in vivo inclusion of native post-translational modifications which may often be absent in recombinantly-prepared proteins or enzymes, thereby further reducing the immunogenicity of the translated protein or enzyme.

[0314] In some embodiments, the target cells can be yeast cells, which include, but not limited to, Saccharomyces cerevisiae and Pichia pastoris. In some embodiments, the target cells can be bacteria cells, which include, but not limited to, Escherichia coli. In some embodiments, the target cells can be insect cells, which include, but not limited to, Spodoptera frugiperda sf9, Mimic Sf9, sf21, Drosophila S2. In some embodiments, the compositions provided herein are optimized for yeast cells, which include, but not limited to, Saccharomyces cerevisiae, Pichia pastoris.

[0315] In some embodiments, the compositions provided herein are optimized for a variety of bacteria cells, which include, but not limited to, Escherichia coli.

[0316] In some embodiments, the compositions provided herein are optimized for a variety of insect cells, which include, but not limited to, Spodoptera frugiperda sf9, Mimic Sf9, sf21, Drosophila S2.

[0317] The recombinant nucleic acids or vectors provided herein can be introduced into a cell by any method, including, for example, by transfection, transformation, or transduction. The terms “transfection, ” transformation, ” and “transduction” are used interchangeably herein and refer to the introduction of one or more exogenous polynucleotides into a host cell by using physical or chemical methods. Many transfection techniques are known in the art and include, for example, calcium phosphate DNA co-precipitation (see, e.g., Murray E.J. (ed. ) , METHODS IN MOLECULAR BIOLOGY, Vol. 7, Gene Transfer and Expression Protocols, Humana Press (1991) ) ; DEAE-dextran; electroporation; cationic liposome-mediated transfection; tungsten particle-facilitated microparticle bombardment (Johnston, Nature, 346: 776-777 (1990) ) ; strontium phosphate DNA co-precipitation (Brash et al., Mol. Cell. Biol., 7: 2031-2034 (1987) ; and magnetic nanoparticle-based gene delivery (Dobson, J. Gene Ther, 13 (4) : 283-7 (2006) ) . 5.19 Articles of manufacture or kits Articles of manufacture or kits comprising RNAs (or cRNAzymes) , vectors, or circRNAs  disclosed herein are also provided herein. An article of manufacture or kit can further comprise a package insert comprising instructions. Suitable containers include, for example, bottles, vials, bags and syringes. The container may be formed from a variety of materials such as glass, plastic (such as polyvinyl chloride or poly olefin) , or metal alloy (such as stainless steel or hastelloy) . In some embodiments, the container holds the formulation and the label on, or associated with, the container may indicate directions for use. The article of manufacture or kit may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts with instructions for use. In some embodiments, the article of manufacture further includes one or more of another agent (e.g., a chemotherapeutic agent, and anti-neoplastic agent) . Suitable containers for the one or more agent include, for example, bottles, vials, bags and syringes. 5.20 Assays

[0318] A variety of methods and assays are available in the art to predict the secondary and even tertiary structure and determine whether an RNA molecule has group II intron self-splicing activity. Some of these methods and assays are provided below. Guided by teachings in the present disclosure, a person of ordinary skill in the art would be able to determine whether an RNA containing the sequence elements disclosed herein have self-splicing activity to form circRNAs. 5.21 In vitro self-splicing assay

[0319] Various in vitro assays are available in the art to test and confirm the self-splicing activity of a given cRNAzyme. The procedure below is described for illustrative purposes. The DNA sequence encoding a subject cRNAzyme is into a proper vector, e.g., pUC57. Vectors are amplified and purified and linearized using BamHI for in vitro transcription. Radiolabeled transcripts are prepared using T7 RNA polymerase, 5 mM MgCl2, 40 mM Tris-HCl pH 7.5, 0.05%Triton X-100, 10 μCi [α-31P] UTP (3000 Ci mmol-1) , 0.5mM UTP, and 1mM other NTPs. In vitro transcription reactions are done for 1 h at 37 ℃ followed by purification of the RNA product on a denaturing 4% (19: 1) polyacrylamide, 8M urea, 1X TBE gel. Radiolabeled RNA product is refolded in 40 mM Tris-HCl pH 7.5 through heating at 90 ℃ for 1 min followed by incubation in 40 mM Tris-HCl pH 7.5, 10mM MgCl2 for 15 min. To initiate self-splicing, RNA was mixed with an equal volume of 2X splicing buffer (2M NH4Cl, 40 mM Tris-HCl pH 7.5) . Reactions are quenched by mixing with an equal volume of 80%formamide, 100 mM EDTA. Splicing products are resolved using denaturing 4% (19: 1) polyacrylamide, 8M urea, 1× TBE gels. All splicing assays are done in triplicate.

[0320] The self-splicing products can be quantified using any methods known in the art. For example, splicing gels are exposed to storage phosphor screens and splicing products are quantitated using Quantity Analysis Software. Unequal loading can be accounted for by normalizing to an internal control RNA. Band intensity is determined by dividing the background subtracted intensity of each band by the number of uridine residues in the RNA sequence corresponding to the band. All band intensities are then normalized to an unspliced control to give fractional values of input RNA. 5.22 Electron Microscopy (EM)

[0321] EM is an effective tool to visualize the structure of a cRNAzyme. An exemplary EM procedure is provided below to determine the structure of a group II intron.

[0322] For negative staining EM, 4 μl of the solution containing the cRNAzyme is applied to a glow-discharged holey carbon grid that is pre-coated with a thin layer of continuous carbon film over the holes. After 1 min, the grid is washed consecutively with three droplets of 2% (w / v) uranyl acetate solution for total of 30 sec. After another 1 min, the residual stain is blotted off and the grid was air-dried. The EM data for the negatively-stained specimen is collected on an FEI Tecnai-12 transmission electron microscope, equipped with a LaB6 filament and operated at 120 kV acceleration voltage. Images are collected at a nominal magnification of 68,000×, using a total dose of about  and defocus value ranging from -1.0 to -1.5 μm on a Gatan Ultrascan4000 CCD camera, with a pixel size of  on the object scale. Total of 66 tilt-pair images of the specimen at 0° and 50° are recorded manually for the random-conical tilt 3D reconstruction.

[0323] For cryo-EM, frozen-hydrated specimens are prepared using the FEI Vitrobot Mark IV plunger. 4 μl of the diluted sample is placed on a glow-discharged holey carbon grid (Quantifoil Cu-Rh R1.2 / 1.3) pre-coated with continuous carbon film (with ~2 nm thickness) . The excess of solution from the grid is blotted for 2.0 s at 100%humidity at 22 ℃ before the grid is flash frozen into liquid ethane slush cooled at liquid nitrogen temperature. Cryo-EM data are collected on an FEI Titan Krios electron microscope, equipped with a Gatan K2 Summit direct-electron counting camera. The microscope is operated at 300 kV and images of the specimen are recorded with a defocus range of -1.2 to -3 μm at a calibrated magnification in super resolution mode of the K2 camera, yielding a pixel size of about  on the object scale. 1000 to 5000 movie stacks, each containing 32 sub-frames, are recorded using the semi-automated low-dose acquisition program UCSF-Image4, with an electron dose rate of 6.25 electrons per  per second and total exposure time of 8 seconds. 5.8 Illustrative Embodiments

[0324] Embodiment 1: A non-naturally occurring nucleic acid comprising a translation initiation sequence (TI) comprising an internal ribosome entry site (IRES) that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof.

[0325] Embodiment 2: The nucleic acid of Embodiment 1, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.

[0326] Embodiment 3: The nucleic acid of Embodiment 1, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof.

[0327] Embodiment 4: The nucleic acid of Embodiment 3, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.

[0328] Embodiment 5: The nucleic acid of any one of Embodiments 1 to 4, wherein the TI comprises at least two IRESs.

[0329] Embodiment 6: The nucleic acid of any one of Embodiments 1 to 4, wherein the TI further comprises a natural IRES sequence or a fragment thereof.

[0330] Embodiment 7: The nucleic acid of any one of Embodiments 1 to 6 that is a DNA.

[0331] Embodiment 8: The nucleic acid of any one of Embodiments 1 to 6 that is an RNA.

[0332] Embodiment 9: The nucleic acid of any one of Embodiments 1 to 8 that is double stranded.

[0333] Embodiment 10: The nucleic acid of any one of Embodiments 1 to 8 that is single stranded.

[0334] Embodiment 11: The nucleic acid of any one of Embodiments 1 to 10 that is circular.

[0335] Embodiment 12: The nucleic acid of any one of Embodiments 1 to 10 that is linear.

[0336] Embodiment 13: The nucleic acid of Embodiment 8 that is mRNA.

[0337] Embodiment 14: The nucleic acid of any one of Embodiments 1 to 13, further comprising a therapeutic protein-coding sequence (Z1) operatively linked to the TI.

[0338] Embodiment 15: The nucleic acid of Embodiment 14, wherein Z1 is linked to TI via a linker.

[0339] Embodiment 16: The nucleic acid of Embodiment 14 or 15, wherein Z1 encodes a therapeutic protein.

[0340] Embodiment 17: The nucleic acid of Embodiment 16 that expresses the protein for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid is administered to a human.

[0341] Embodiment 18: The nucleic acid of Embodiment 16 or 17, wherein the expression of the therapeutic protein is immunologically inert after the nucleic acid is administered to a human.

[0342] Embodiment 19: A protein expressed by the nucleic acid of any one of Embodiment 14 to 18.

[0343] Embodiment 20: A vector comprising the nucleic acid of any one of Embodiments 1 to 18.

[0344] Embodiment 21: A cell comprising the nucleic acid of any one of Embodiments 1 to 18.

[0345] Embodiment 22: The cell of Embodiment 21, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.

[0346] Embodiment 23: A method of expressing a protein comprising culturing the cell of Embodiment 22 under conditions and for a sufficient time for the expression of the protein.

[0347] Embodiment 24: A pharmaceutical composition comprising the nucleic acid of any one of Embodiments 1 to 18, the vector of Embodiment 20, or the cell of Embodiment 21, and a pharmaceutically acceptable carrier, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI, wherein Z1 encodes a therapeutic protein.

[0348] Embodiment 25: A method of expressing a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 24.

[0349] Embodiment 26: A method of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 24.

[0350] Embodiment 27: Use of the pharmaceutical composition of Embodiment 24 for treating or preventing a disease or disorder.

[0351] Embodiment 28: A non-naturally occurring nucleic acid encoding an RNA comprising the following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circular RNA (circRNA) that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment; and wherein the circRNA comprises a TI that comprises an IRES that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof.

[0352] Embodiment 29: The nucleic acid of Embodiment 28, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.

[0353] Embodiment 30: The nucleic acid of Embodiment 28, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof.

[0354] Embodiment 31: The nucleic acid of Embodiment 30, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.

[0355] Embodiment 32: The nucleic acid of any one of Embodiments 28 to 31, wherein the TI comprises at least two IRESs.

[0356] Embodiment 33: The nucleic acid of any one of Embodiments 28 to 32, wherein the TI further comprises a natural IRES sequence or a fragment thereof.

[0357] Embodiment 34: The nucleic acid of any one of Embodiments 28 to 33, wherein the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.

[0358] Embodiment 35: The nucleic acid of Embodiment 34, wherein Z1 encodes a therapeutic protein.

[0359] Embodiment 36: The nucleic acid of Embodiment 35, wherein the therapeutic protein is expressed for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid or the circRNA is administered to a human.

[0360] Embodiment 37: The nucleic acid of Embodiment 35 or 36, wherein the expression of the therapeutic protein is immunologically inert after the nucleic acid or the circRNA is administered to a human.

[0361] Embodiment 38: The nucleic acid of any one of Embodiments 34 to 37, having a structure selected from the group consisting of Formulae (I) - (IV) : (I) 5’ - (3’ IF) - (L) n-Z1- (L) n-TI- (5’ IF) -3’ ; (II) 5’ - (3’ IF) - (L) n-TI- (L) n-Z1- (5’ IF) -3’ ; (III) 5’ - (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) -3’ ; (IV) 5’ - (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) -3’ ; and wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI the a translation  initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is the therapeutic protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2.

[0362] Embodiment 39: The nucleic acid of any one of Embodiments 28 to 38, further comprising a 5’ homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment.

[0363] Embodiment 40: The nucleic acid of Embodiment 39, wherein the 5’ homology arm, the 3’ homology arm, or both are 15 to 60 nucleotides in length.

[0364] Embodiment 41: The nucleic acid of any one of Embodiments 28 to 40, wherein (1) the 3’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NO: 42-52 and 228; or (2) the 5’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229; or both (1) and (2) .

[0365] Embodiment 42: The nucleic acid of any one of Embodiments 28 to 41 that is a DNA.

[0366] Embodiment 43: The nucleic acid of any one of Embodiments 28 to 41 that is an RNA.

[0367] Embodiment 44: The nucleic acid of any one of Embodiments 28 to 43 that is double stranded.

[0368] Embodiment 45: The nucleic acid of any one of Embodiments 28 to 43 that is single stranded.

[0369] Embodiment 46: The nucleic acid of any one of Embodiments 28 to 45 that is circular.

[0370] Embodiment 47: The nucleic acid of any one of Embodiments 28 to 45 that is linear.

[0371] Embodiment 48: A circRNA produced by the self-splicing of the RNA encoded by the nucleic acid of any one of Embodiments 28 to 47.

[0372] Embodiment 49: A vector comprising the nucleic acid of any one of Embodiments 28 to 47.

[0373] Embodiment 50: A cell comprising the nucleic acid of any one of Embodiments 28 to 47, the circRNA of Embodiment 48, or the vector of Embodiment 49.

[0374] Embodiment 51: The cell of Embodiment 50, wherein the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.

[0375] Embodiment 52: A method of expressing a protein comprising culturing the cell of Embodiment 51 under conditions and for a sufficient time for the expression of the protein.

[0376] Embodiment 53: A pharmaceutical composition comprising the nucleic acid of Embodiments 28 to 47, the circRNA of Embodiment 48, the vector of Embodiment 49, or the cell of Embodiment 50, and a pharmaceutically acceptable carrier, wherein the circRNA further comprises Z1 operatively linked to the TI, wherein Z1 encodes a therapeutic protein.

[0377] Embodiment 54: A method of expressing of a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 53.

[0378] Embodiment 55: A method of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 53.

[0379] Embodiment 56: Use of the pharmaceutical composition of Embodiment 53 for treating or preventing a disease or disorder. Another Illustrative Embodiments Embodiment 1’ : A non-naturally occurring nucleic acid comprising a translation initiation  sequence (TI) comprising an internal ribosome entry site (IRES) that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof. Embodiment 2’ : The nucleic acid of Embodiment 1’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof. Embodiment 3’ : The nucleic acid of Embodiment 1’ , wherein the IRES is at least 85%, at least  88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof. Embodiment 4’ : The nucleic acid of Embodiment 3’ , wherein the IRES is at least 85%, at least  88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. Embodiment 5’ : The nucleic acid of any one of Embodiments 1’ to 4’ , wherein the TI comprises  at least two IRESs. Embodiment 6’ : The nucleic acid of any one of Embodiments 1’ to 4’ , wherein the TI further  comprises a natural IRES sequence or a fragment thereof. Embodiment 7’ : The nucleic acid of any one of Embodiments 1’ to 6’ , that is a DNA. Embodiment 8’ : The nucleic acid of any one of Embodiments 1’ to 6’ that is an RNA. Embodiment 9’ : The nucleic acid of any one of Embodiments 1’ to 8’ that is double stranded. Embodiment 10’ : The nucleic acid of any one of Embodiments 1’ to 8’ that is single stranded. Embodiment 11’ : The nucleic acid of any one of Embodiments 1’ to 10’ , that is circular. Embodiment 12’ : The nucleic acid of any one of Embodiments 1’ to 10’ , that is linear. Embodiment 13’ : The nucleic acid of Embodiment 8’ , that is mRNA. Embodiment 14’ : The nucleic acid of any one of Embodiments 1’ to 13’ , further comprising a  therapeutic protein-coding sequence (Z1) operatively linked to the TI. Embodiment 15’ : The nucleic acid of Embodiment 14’ , wherein Z1 is linked to TI via a linker.  Embodiment 16’ : The nucleic acid of Embodiment 14’ or 15’ , wherein Z1 encodes a therapeutic protein. Embodiment 17’ : The nucleic acid of Embodiment 16’ that expresses the protein for at least 3  days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid is administered to a human. Embodiment 18’ : The nucleic acid of Embodiment 16’ or 17’ , wherein the expression of the  therapeutic protein is immunologically inert after the nucleic acid is administered to a human. Embodiment 19’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for muscle tissues,  and the muscle tissue is from a human or a mouse. Embodiment 20’ : The nucleic acid of Embodiment 19’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 425, 287, 286, 278, 542, 423, 277, 436, 269, 294, 608, 415, 535, 541, 289, 358, 498, 292, 266, 291, 360, 281, 284, 293, 369, 550, 271, 272, 590, 267, 430, 390, 411, 297, 357, 362, 499, 283, 491, 546, 290, 282, 478, 265, 410, 432, 413, 414, 315, 392, 426, 268, 302, 614, 288, 318, 357, 575, 534, 322, 339, 576, 544, 324, 354, 336, 355, 589, 588, 295, 296, 356, 560, 514, 553, 417, 573, 409, 329, 280, 341, 316, 431, 463, 511, 600, 282, 472, 558, 421, and 538, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-1 and Table 30-3 or a reverse complementary sequence thereof. Embodiment 21’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for lung tissues, and  the lung tissue is from a human or a mouse. Embodiment 22’ : The nucleic acid of Embodiment 21’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 287, 425, 541, 535, 269, 266, 542, 277, 291, 498, 358, 292, 294, 286, 297, 390, 550, 608, 289, 546, 436, 282, 499, 278, 423, 293, 369, 265, 411, 478, 590, 553, 415, 272, 357, 392, 480, 360, 271, 413, 534, 281, 280, 356, 491, 284, 350, 285, 547, 484, 270, 290, 514, 267, 288, 511, 430, 614, 283, 537, 575, 339, 417, 410, 544, 432, 533, and 357, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-2 or a reverse complementary sequence thereof. Embodiment 23’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for stomach tissues,  and the stomach tissue is from a human or a mouse. Embodiment 24’ : The nucleic acid of Embodiment 23’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 269, 287, 297, 286, 291, 475, 265, 266, 390, 436, 280, 542, 267, 541, 498, 608, 277, 283, 546, 272, 499, 278, 293, 271, 392, 535, 553, 478, 550, 268, 358, 425, 289, and 590, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-3 or a reverse complementary sequence thereof. Embodiment 25’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for nervous tissues,  and the nervous tissue is from a human or a mouse. Embodiment 26’ : The nucleic acid of Embodiment 25’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 289, 608, 614, 590, 600, 297, 287, 285, 281, 423, 357, 266, 293, 425, 358, 414, 277, 575, 542, 392, 589, 436, 294, 541, 286, 535, 360, 576, 550, 269, 369, 290, 546, 390, 280, 268, 413, 415, 288, 283, 354, 480, 499, 534, 291, 329, 350, 347, 362, 315, 410, 282, 292, 498, 475, 278, 588, 339, 411, 272, 431, 374, 265, 271, 430, 453, 348, 478, 466, 267, 590, 284, 296, 355, 322, 357, 417, 352, 373, 533, 432, 492, 408, 472, 356, 560, 485, 318, 544, 295, 361, 316, 302, 514, 384, 488, and 511, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-4 and Table 30-1 or a reverse complementary sequence thereof. Embodiment 27’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for colonic tissues,  and the colonic tissue is from a human or a mouse. Embodiment 28’ : The nucleic acid of Embodiment 27’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 608, 358, 266, 575, 289, 541, 369, 475, 614, 542, 287, 293, 269, 350, 285, 281, 357, 535, 478, 480, 600, 354, 294, 576, 282, 360, 286, 290, 297, 546, 278, 277, 291, 550, 413, 292, 425, and 295, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-5 or a reverse complementary sequence thereof. Embodiment 29’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for liver tissues, and  the liver tissue is from a human or a mouse. Embodiment 30’ : The nucleic acid of Embodiment 29’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 287, 436, 293, 281, 415, 294, 269, 425, 392, 423, 290, 411, 277, 278, 413, 288, 266, 289, 268, 286, 608, 297, 614, 390, 357, 410, 358, 265, 475, 498, 432, 542, 280, 291, 466, 283, 414, 590, 360, 430, 271, 357, 292, 417, 589, 354, 267, 431, 541, 546, 535, 362, 356, 499, 285, 514, 480, 355, 350, 272, 369, 550, 295, 284, 527, 318, 339, 296, 531, 511, 329, 600, 478, 302, 590, 270, 322, 503, 491, 553, 463, 315, 575, 453, and 456, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-6 or a reverse complementary sequence thereof. Embodiment 31’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for skin tissues, and  the skin tissue is from a human or a mouse. Embodiment 32’ : The nucleic acid of Embodiment 31’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 278, 608, 425, 475, 289, 277, 286, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-7 or a reverse complementary sequence thereof. Embodiment 33’ : The nucleic acid of Embodiment 1’ , wherein the IRES is for immune tissues,  and the immune tissue is from a human or a mouse. Embodiment 34’ : The nucleic acid of Embodiment 33’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 297, 286, 287, 535, 281, 475, 277, 296, 542, 269, 541, 358, 390, 550, 369, 360, 436, 265, 357, 354, 350, 291, 357, 266, 546, 608, 271, 289, 355, 280, 590, 278, 425, 356, 362, 484, 272, 293, 292, 283, 491, 282, 480, 285, 294, 392, 498, 511, 287, 269, 277, 278, 390, 542, 291, 286, 289, 541, 608, 265, 266, 546, 294, 550, 498, 292, 511, 475, 531, 267, 514, 318, 600, 614, 508, 387, 284, 415, 423, 503, 588, 270, 295, 553, 288, 575, 408, 534, 430, 410, 411, 432, and 385, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-8 and Table 30-2 or a reverse complementary sequence thereof. Embodiment 35’ : A protein expressed by the nucleic acid of any one of Embodiments 14’ to 34’ . Embodiment 36’ : A vector comprising the nucleic acid of any one of Embodiments 1’ to 34’ . Embodiment 37’ : A cell comprising the nucleic acid of any one of Embodiments 1’ to 34’ . Embodiment 38’ : The cell of Embodiment 37’ , wherein the nucleic acid further comprises a  therapeutic protein-coding sequence (Z1) operatively linked to the TI. Embodiment 39’ : A method of expressing a protein comprising culturing the cell of Embodiment  38’ under conditions and for a sufficient time for the expression of the protein. Embodiment 40’ : A pharmaceutical composition comprising the nucleic acid of any one of  Embodiments 1’ to 34’ , the vector of Embodiment 36’ , or the cell of Embodiment 37’ , and a pharmaceutically acceptable carrier, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI, wherein Z1 encodes a therapeutic protein. Embodiment 41’ : A method of expressing a therapeutic protein in a subject in need thereof,  comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 40’ . Embodiment 42’ : A method of treating or preventing a disease or disorder in a subject in need  thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of Embodiment 40’ . Embodiment 43’ : Use of the pharmaceutical composition of Embodiment 40’ for treating or  preventing a disease or disorder. Embodiment 44’ : A non-naturally occurring nucleic acid encoding an RNA comprising the  following operably linked elements from 5’ to 3’ : (1) a 3’ intron fragment; (2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence  fragment, from 5’ to 3’ ; and (3) a 5’ intron fragment; wherein the RNA has group II intron activity and, upon self-splicing, can form a circular RNA  (circRNA) that comprises both the 5’ and 3’ target sequence fragments with the 3’ -end of the 5’ target sequence fragment linked to the 5’ -end of the 3’ target sequence fragment; and wherein the circRNA comprises a TI that comprises an IRES that is at least 85%, at least 88%, at  least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof. Embodiment 45’ : The nucleic acid of Embodiment 44’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof. Embodiment 46’ : The nucleic acid of Embodiment 44’ , wherein the IRES is at least 85%, at least  88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof. Embodiment 47’ : The nucleic acid of Embodiment 46’ , wherein the IRES is at least 85%, at least  88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622. Embodiment 48’ : The nucleic acid of any one of Embodiments 44’ to 47’ , wherein the TI  comprises at least two IRESs. Embodiment 49’ : The nucleic acid of any one of Embodiments 44’ to 48’ , wherein the TI further  comprises a natural IRES sequence or a fragment thereof. Embodiment 50’ : The nucleic acid of any one of Embodiments 44’ to 49’ , wherein the circRNA  further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI. Embodiment 51’ : The nucleic acid of Embodiment 50’ , wherein Z1 encodes a therapeutic protein. Embodiment 52’ : The nucleic acid of Embodiment 51’ , wherein the therapeutic protein is  expressed for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid or the circRNA is administered to a human. Embodiment 53’ : The nucleic acid of Embodiment 51’ or 52’ , wherein the expression of the  therapeutic protein is immunologically inert after the nucleic acid or the circRNA is administered to a human. Embodiment 54’ : The nucleic acid of any one of Embodiments 50’ to 53’ , having a structure  selected from the group consisting of Formulae (I) - (IV) : (I) 5’ - (3’ IF) - (L) n-Z1- (L) n-TI- (5’ IF) -3’ ; (II) 5’ - (3’ IF) - (L) n-TI- (L) n-Z1- (5’ IF) -3’ ; (III) 5’ - (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) -3’ ; (IV) 5’ - (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) -3’ ; and wherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI the a translation  initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is the protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2. Embodiment 55’ : The nucleic acid of any one of Embodiments 44’ to 54’ , further comprising a 5’  homology arm operatively linked to the 5’ -end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’ -end of the 5’ intron fragment. Embodiment 56’ : The nucleic acid of Embodiment 55’ , wherein the 5’ homology arm, the 3’  homology arm, or both are 15 to 60 nucleotides in length. Embodiment 57’ : The nucleic acid of any one of Embodiments 44’ to 56’ , wherein (1) the 3’  intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NO: 42-52 and 228; or (2) the 5’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229; or both (1) and (2) . Embodiment 58’ : The nucleic acid of any one of Embodiments 44’ to 57’ , that is a DNA. Embodiment 59’ : The nucleic acid of any one of Embodiments 44’ to 57’ , that is an RNA. Embodiment 60’ : The nucleic acid of any one of Embodiments 44’ to 59’ , that is double stranded. Embodiment 61’ : The nucleic acid of any one of Embodiments 44’ to 59’ , that is single stranded. Embodiment 62’ : The nucleic acid of any one of Embodiments 44’ to 61’ , that is circular. Embodiment 63’ : The nucleic acid of any one of Embodiments 44’ to 61’ , that is linear. Embodiment 64’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for muscle tissues,  and the muscle tissue is from a human or a mouse. Embodiment 65’ : The nucleic acid of Embodiment 64’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 425, 287, 286, 278, 542, 423, 277, 436, 269, 294, 608, 415, 535, 541, 289, 358, 498, 292, 266, 291, 360, 281, 284, 293, 369, 550, 271, 272, 590, 267, 430, 390, 411, 297, 357, 362, 499, 283, 491, 546, 290, 282, 478, 265, 410, 432, 413, 414, 315, 392, 426, 268, 302, 614, 288, 318, 357, 575, 534, 322, 339, 576, 544, 324, 354, 336, 355, 589, 588, 295, 296, 356, 560, 514, 553, 417, 573, 409, 329, 280, 341, 316, 431, 463, 511, 600, 282, 472, 558, 421, and 538, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-1 and Table 30-3 or a reverse complementary sequence thereof. Embodiment 66’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for lung tissues, and  the lung tissue is from a human or a mouse. Embodiment 67’ : The nucleic acid of Embodiment 66’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 287, 425, 541, 535, 269, 266, 542, 277, 291, 498, 358, 292, 294, 286, 297, 390, 550, 608, 289, 546, 436, 282, 499, 278, 423, 293, 369, 265, 411, 478, 590, 553, 415, 272, 357, 392, 480, 360, 271, 413, 534, 281, 280, 356, 491, 284, 350, 285, 547, 484, 270, 290, 514, 267, 288, 511, 430, 614, 283, 537, 575, 339, 417, 410, 544, 432, 533, and 357, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-2 or a reverse complementary sequence thereof. Embodiment 68’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for stomach tissues,  and the stomach tissue is from a human or a mouse. Embodiment 69’ : The nucleic acid of Embodiment 68’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 269, 287, 297, 286, 291, 475, 265, 266, 390, 436, 280, 542, 267, 541, 498, 608, 277, 283, 546, 272, 499, 278, 293, 271, 392, 535, 553, 478, 550, 268, 358, 425, 289, and 590, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-3 or a reverse complementary sequence thereof. Embodiment 70’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for nervous tissues,  and the nervous tissue is from a human or a mouse. Embodiment 71’ : The nucleic acid of Embodiment 70’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 289, 608, 614, 590, 600, 297, 287, 285, 281, 423, 357, 266, 293, 425, 358, 414, 277, 575, 542, 392, 589, 436, 294, 541, 286, 535, 360, 576, 550, 269, 369, 290, 546, 390, 280, 268, 413, 415, 288, 283, 354, 480, 499, 534, 291, 329, 350, 347, 362, 315, 410, 282, 292, 498, 475, 278, 588, 339, 411, 272, 431, 374, 265, 271, 430, 453, 348, 478, 466, 267, 590, 284, 296, 355, 322, 357, 417, 352, 373, 533, 432, 492, 408, 472, 356, 560, 485, 318, 544, 295, 361, 316, 302, 514, 384, 488, and 511, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-4 and Table 30-1 or a reverse complementary sequence thereof. Embodiment 72’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for colonic tissues,  and the colonic tissue is from a human or a mouse. Embodiment 73’ : The nucleic acid of Embodiment 72’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 608, 358, 266, 575, 289, 541, 369, 475, 614, 542, 287, 293, 269, 350, 285, 281, 357, 535, 478, 480, 600, 354, 294, 576, 282, 360, 286, 290, 297, 546, 278, 277, 291, 550, 413, 292, 425, and 295, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-5 or a reverse complementary sequence thereof. Embodiment 74’ : The nucleic acid of Embodiment 44’ , wherein the IRES is for liver tissues, and  the liver tissue is from a human or a mouse. Embodiment 75’ : The nucleic acid of Embodiment 74’ , wherein the IRES is identical to a  nucleotide sequence selected from the group consisting of SEQ ID NOs: 287, 436, 293, 281, 415, 294, 269, 425, 392, 423, 290, 411, 277, 278, 413, 288, 266, 289, 268, 286, 608, 297, 614, 390, 357, 410, 358, 265, 475, 498, 432, 542, 280, 291, 466, 283, 414, 590, 360, 430, 271, 357, 292, 417, 589, 354, 267, 431, 541, 546, 535, 362, 356, 499, 285, 514, 480, 355, 350, 272, 369, 550, ...

Claims

1.A non-naturally occurring nucleic acid comprising a translation initiation sequence (TI) comprising an internal ribosome entry site (IRES) that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof.2.The nucleic acid of claim 1, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.3.The nucleic acid of claim 1, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof.4.The nucleic acid of claim 3, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.5.The nucleic acid of any one of claims 1 to 4, wherein the TI comprises at least two IRESs.6.The nucleic acid of any one of claims 1 to 4, wherein the TI further comprises a natural IRES sequence or a fragment thereof.7.The nucleic acid of any one of claims 1 to 6, that is a DNA.8.The nucleic acid of any one of claims 1 to 6, that is an RNA.9.The nucleic acid of any one of claims 1 to 8, that is double stranded.10.The nucleic acid of any one of claims 1 to 8, that is single stranded.11.The nucleic acid of any one of claims 1 to 10, that is circular.12.The nucleic acid of any one of claims 1 to 10, that is linear.13.The nucleic acid of claim 8, that is mRNA.14.The nucleic acid of any one of claims 1 to 13, further comprising a therapeutic protein-coding sequence (Z1) operatively linked to the TI.15.The nucleic acid of claim 14, wherein Z1 is linked to TI via a linker.16.The nucleic acid of claim 14 or 15, wherein Z1 encodes a therapeutic protein.17.The nucleic acid of claim 16 that expresses the protein for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid is administered to a human.18.The nucleic acid of claim 16 or 17, wherein the expression of the therapeutic protein is immunologically inert after the nucleic acid is administered to a human.19.The nucleic acid of claim 1, wherein the IRES is for muscle tissues, and the muscle tissue is from a human or a mouse.20.The nucleic acid of claim 19, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 425, 287, 286, 278, 542, 423, 277, 436, 269, 294, 608, 415, 535, 541, 289, 358, 498, 292, 266, 291, 360, 281, 284, 293, 369, 550, 271, 272, 590, 267, 430, 390, 411, 297, 357, 362, 499, 283, 491, 546, 290, 282, 478, 265, 410, 432, 413, 414, 315, 392, 426, 268, 302, 614, 288, 318, 357, 575, 534, 322, 339, 576, 544, 324, 354, 336, 355, 589, 588, 295, 296, 356, 560, 514, 553, 417, 573, 409, 329, 280, 341, 316, 431, 463, 511, 600, 282, 472, 558, 421, and 538, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-1 and Table 30-3 or a reverse complementary sequence thereof.21.The nucleic acid of claim 1, wherein the IRES is for lung tissues, and the lung tissue is from a human or a mouse.22.The nucleic acid of claim 21, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 287, 425, 541, 535, 269, 266, 542, 277, 291, 498, 358, 292, 294, 286, 297, 390, 550, 608, 289, 546, 436, 282, 499, 278, 423, 293, 369, 265, 411, 478, 590, 553, 415, 272, 357, 392, 480, 360, 271, 413, 534, 281, 280, 356, 491, 284, 350, 285, 547, 484, 270, 290, 514, 267, 288, 511, 430, 614, 283, 537, 575, 339, 417, 410, 544, 432, 533, and 357, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-2 or a reverse complementary sequence thereof.23.The nucleic acid of claim 1, wherein the IRES is for stomach tissues, and the stomach tissue is from a human or a mouse.24.The nucleic acid of claim 23, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 269, 287, 297, 286, 291, 475, 265, 266, 390, 436, 280, 542, 267, 541, 498, 608, 277, 283, 546, 272, 499, 278, 293, 271, 392, 535, 553, 478, 550, 268, 358, 425, 289, and 590, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-3 or a reverse complementary sequence thereof.25.The nucleic acid of claim 1, wherein the IRES is for nervous tissues, and the nervous tissue is from a human or a mouse.26.The nucleic acid of claim 25, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 289, 608, 614, 590, 600, 297, 287, 285, 281, 423, 357, 266, 293, 425, 358, 414, 277, 575, 542, 392, 589, 436, 294, 541, 286, 535, 360, 576, 550, 269, 369, 290, 546, 390, 280, 268, 413, 415, 288, 283, 354, 480, 499, 534, 291, 329, 350, 347, 362, 315, 410, 282, 292, 498, 475, 278, 588, 339, 411, 272, 431, 374, 265, 271, 430, 453, 348, 478, 466, 267, 590, 284, 296, 355, 322, 357, 417, 352, 373, 533, 432, 492, 408, 472, 356, 560, 485, 318, 544, 295, 361, 316, 302, 514, 384, 488, and 511, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-4 and Table 30-1 or a reverse complementary sequence thereof.27.The nucleic acid of claim 1, wherein the IRES is for colonic tissues, and the colonic tissue is from a human or a mouse.28.The nucleic acid of claim 27, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 608, 358, 266, 575, 289, 541, 369, 475, 614, 542, 287, 293, 269, 350, 285, 281, 357, 535, 478, 480, 600, 354, 294, 576, 282, 360, 286, 290, 297, 546, 278, 277, 291, 550, 413, 292, 425, and 295, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-5 or a reverse complementary sequence thereof.29.The nucleic acid of claim 1, wherein the IRES is for liver tissues, and the liver tissue is from a human or a mouse.30.The nucleic acid of claim 29, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 287, 436, 293, 281, 415, 294, 269, 425, 392, 423, 290, 411, 277, 278, 413, 288, 266, 289, 268, 286, 608, 297, 614, 390, 357, 410, 358, 265, 475, 498, 432, 542, 280, 291, 466, 283, 414, 590, 360, 430, 271, 357, 292, 417, 589, 354, 267, 431, 541, 546, 535, 362, 356, 499, 285, 514, 480, 355, 350, 272, 369, 550, 295, 284, 527, 318, 339, 296, 531, 511, 329, 600, 478, 302, 590, 270, 322, 503, 491, 553, 463, 315, 575, 453, and 456, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-6 or a reverse complementary sequence thereof.31.The nucleic acid of claim 1, wherein the IRES is for skin tissues, and the skin tissue is from a human or a mouse.32.The nucleic acid of claim 31, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 278, 608, 425, 475, 289, 277, 286, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-7 or a reverse complementary sequence thereof.33.The nucleic acid of claim 1, wherein the IRES is for immune tissues, and the immune tissue is from a human or a mouse.34.The nucleic acid of claim 33, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 297, 286, 287, 535, 281, 475, 277, 296, 542, 269, 541, 358, 390, 550, 369, 360, 436, 265, 357, 354, 350, 291, 357, 266, 546, 608, 271, 289, 355, 280, 590, 278, 425, 356, 362, 484, 272, 293, 292, 283, 491, 282, 480, 285, 294, 392, 498, 511, 287, 269, 277, 278, 390, 542, 291, 286, 289, 541, 608, 265, 266, 546, 294, 550, 498, 292, 511, 475, 531, 267, 514, 318, 600, 614, 508, 387, 284, 415, 423, 503, 588, 270, 295, 553, 288, 575, 408, 534, 430, 410, 411, 432, and 385, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-8 and Table 30-2 or a reverse complementary sequence thereof.35.A protein expressed by the nucleic acid of any one of claims 14 to 34.36.A vector comprising the nucleic acid of any one of claims 1 to 34.37.A cell comprising the nucleic acid of any one of claims 1 to 34.38.The cell of claim 37, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.39.A method of expressing a protein comprising culturing the cell of claim 38 under conditions and for a sufficient time for the expression of the protein.40.A pharmaceutical composition comprising the nucleic acid of any one of claims 1 to 34, the vector of claim 36, or the cell of claim 37, and a pharmaceutically acceptable carrier, wherein the nucleic acid further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI, wherein Z1 encodes a therapeutic protein.41.A method of expressing a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of claim 40.42.A method of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of claim 40.43.Use of the pharmaceutical composition of claim 40 for treating or preventing a disease or disorder.44.A non-naturally occurring nucleic acid encoding an RNA comprising the following operably linked elements from 5’ to 3’:(1) a 3’ intron fragment;(2) a target sequence consisting of (i) a 3’ target sequence fragment and (ii) a 5’ target sequence fragment, from 5’ to 3’; and(3) a 5’ intron fragment;wherein the RNA has group II intron activity and, upon self-splicing, can form a circular RNA (circRNA) that comprises both the 5’ and 3’ target sequence fragments with the 3’-end of the 5’ target sequence fragment linked to the 5’-end of the 3’ target sequence fragment; andwherein the circRNA comprises a TI that comprises an IRES that is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, a subsequence thereof, or a reverse complementary sequence thereof.45.The nucleic acid of claim 44, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 264-614, or a reverse complementary sequence thereof.46.The nucleic acid of claim 44, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to SEQ ID NO: 270, 354, 392, 463, 512, 545, 577, or 606, a subsequence thereof, or a reverse complementary sequence thereof.47.The nucleic acid of claim 46, wherein the IRES is at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, at least 97%, at least 98%, at least 99%or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 615-622.48.The nucleic acid of any one of claims 44 to 47, wherein the TI comprises at least two IRESs.49.The nucleic acid of any one of claims 44 to 48, wherein the TI further comprises a natural IRES sequence or a fragment thereof.50.The nucleic acid of any one of claims 44 to 49, wherein the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.51.The nucleic acid of claim 50, wherein Z1 encodes a therapeutic protein.52.The nucleic acid of claim 51, wherein the therapeutic protein is expressed for at least 3 days, at least 4 days, at least 5 days, at least 6 days, or at least 7 days after the nucleic acid or the circRNA is administered to a human.53.The nucleic acid of claim 51 or 52, wherein the expression of the therapeutic protein is immunologically inert after the nucleic acid or the circRNA is administered to a human.54.The nucleic acid of any one of claims 50 to 53, having a structure selected from the group consisting of Formulae (I) - (IV) :(I) 5’- (3’ IF) - (L) n-Z1- (L) n-TI- (5’ IF) -3’;(II) 5’- (3’ IF) - (L) n-TI- (L) n-Z1- (5’ IF) -3’;(III) 5’- (3’ IF) -TIB- (L) n-Z1- (L) n-TIA- (5’ IF) -3’;(IV) 5’- (3’ IF) -Z1B- (L) n-TI- (L) n-Z1A- (5’ IF) -3’; andwherein 3’ IF is the 3’ intron fragment; 5’ IF is the 5’ intron fragment; TI the a translation initiation sequence, which can be segmented into a 5’ fragment (TIA) and a 3’ fragment TI (TIB) ; Z1 is the protein-coding sequence, which can be segmented into a 5’ fragment (Z1A) and a 3’ fragment (Z1B) ; and each L is independently a linker sequence, and n=0, 1 or 2.55.The nucleic acid of any one of claims 44 to 54, further comprising a 5’ homology arm operatively linked to the 5’-end of the 3’ intron fragment, and a 3’ homology arm operatively linked to the 3’-end of the 5’ intron fragment.56.The nucleic acid of claim 55, wherein the 5’ homology arm, the 3’ homology arm, or both are 15 to 60 nucleotides in length.57.The nucleic acid of any one of claims 44 to 56, wherein (1) the 3’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NO: 42-52 and 228; or (2) the 5’ intron fragment has a nucleotide sequence that is at least 95%, at least 98%, at least 99%, or 100%identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 75-88 and 229; or both (1) and (2) .58.The nucleic acid of any one of claims 44 to 57, that is a DNA.59.The nucleic acid of any one of claims 44 to 57, that is an RNA.60.The nucleic acid of any one of claims 44 to 59, that is double stranded.61.The nucleic acid of any one of claims 44 to 59, that is single stranded.62.The nucleic acid of any one of claims 44 to 61, that is circular.63.The nucleic acid of any one of claims 44 to 61 that is linear.64.The nucleic acid of claim 44, wherein the IRES is for muscle tissues, and the muscle tissue is from a human or a mouse.65.The nucleic acid of claim 64, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 425, 287, 286, 278, 542, 423, 277, 436, 269, 294, 608, 415, 535, 541, 289, 358, 498, 292, 266, 291, 360, 281, 284, 293, 369, 550, 271, 272, 590, 267, 430, 390, 411, 297, 357, 362, 499, 283, 491, 546, 290, 282, 478, 265, 410, 432, 413, 414, 315, 392, 426, 268, 302, 614, 288, 318, 357, 575, 534, 322, 339, 576, 544, 324, 354, 336, 355, 589, 588, 295, 296, 356, 560, 514, 553, 417, 573, 409, 329, 280, 341, 316, 431, 463, 511, 600, 282, 472, 558, 421, and 538, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-1 and Table 30-3 or a reverse complementary sequence thereof.66.The nucleic acid of claim 44, wherein the IRES is for lung tissues, and the lung tissue is from a human or a mouse.67.The nucleic acid of claim 66, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 287, 425, 541, 535, 269, 266, 542, 277, 291, 498, 358, 292, 294, 286, 297, 390, 550, 608, 289, 546, 436, 282, 499, 278, 423, 293, 369, 265, 411, 478, 590, 553, 415, 272, 357, 392, 480, 360, 271, 413, 534, 281, 280, 356, 491, 284, 350, 285, 547, 484, 270, 290, 514, 267, 288, 511, 430, 614, 283, 537, 575, 339, 417, 410, 544, 432, 533, and 357, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-2 or a reverse complementary sequence thereof.68.The nucleic acid of claim 44, wherein the IRES is for stomach tissues, and the stomach tissue is from a human or a mouse.69.The nucleic acid of claim 68, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 269, 287, 297, 286, 291, 475, 265, 266, 390, 436, 280, 542, 267, 541, 498, 608, 277, 283, 546, 272, 499, 278, 293, 271, 392, 535, 553, 478, 550, 268, 358, 425, 289, and 590, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-3 or a reverse complementary sequence thereof.70.The nucleic acid of claim 44, wherein the IRES is for nervous tissues, and the nervous tissue is from a human or a mouse.71.The nucleic acid of claim 70, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 289, 608, 614, 590, 600, 297, 287, 285, 281, 423, 357, 266, 293, 425, 358, 414, 277, 575, 542, 392, 589, 436, 294, 541, 286, 535, 360, 576, 550, 269, 369, 290, 546, 390, 280, 268, 413, 415, 288, 283, 354, 480, 499, 534, 291, 329, 350, 347, 362, 315, 410, 282, 292, 498, 475, 278, 588, 339, 411, 272, 431, 374, 265, 271, 430, 453, 348, 478, 466, 267, 590, 284, 296, 355, 322, 357, 417, 352, 373, 533, 432, 492, 408, 472, 356, 560, 485, 318, 544, 295, 361, 316, 302, 514, 384, 488, and 511, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-4 and Table 30-1 or a reverse complementary sequence thereof.72.The nucleic acid of claim 44, wherein the IRES is for colonic tissues, and the colonic tissue is from a human or a mouse.73.The nucleic acid of claim 72, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 608, 358, 266, 575, 289, 541, 369, 475, 614, 542, 287, 293, 269, 350, 285, 281, 357, 535, 478, 480, 600, 354, 294, 576, 282, 360, 286, 290, 297, 546, 278, 277, 291, 550, 413, 292, 425, and 295, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-5 or a reverse complementary sequence thereof.74.The nucleic acid of claim 44, wherein the IRES is for liver tissues, and the liver tissue is from a human or a mouse.75.The nucleic acid of claim 74, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 287, 436, 293, 281, 415, 294, 269, 425, 392, 423, 290, 411, 277, 278, 413, 288, 266, 289, 268, 286, 608, 297, 614, 390, 357, 410, 358, 265, 475, 498, 432, 542, 280, 291, 466, 283, 414, 590, 360, 430, 271, 357, 292, 417, 589, 354, 267, 431, 541, 546, 535, 362, 356, 499, 285, 514, 480, 355, 350, 272, 369, 550, 295, 284, 527, 318, 339, 296, 531, 511, 329, 600, 478, 302, 590, 270, 322, 503, 491, 553, 463, 315, 575, 453, and 456, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-6 or a reverse complementary sequence thereof.76.The nucleic acid of claim 44, wherein the IRES is for skin tissues, and the skin tissue is from a human or a mouse.77.The nucleic acid of claim 76, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 278, 608, 425, 475, 289, 277, 286, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-7 or a reverse complementary sequence thereof.78.The nucleic acid of claim 44, wherein the IRES is for immune tissues, and the immune tissue is from a human or a mouse.79.The nucleic acid of claim 78, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 297, 286, 287, 535, 281, 475, 277, 296, 542, 269, 541, 358, 390, 550, 369, 360, 436, 265, 357, 354, 350, 291, 357, 266, 546, 608, 271, 289, 355, 280, 590, 278, 425, 356, 362, 484, 272, 293, 292, 283, 491, 282, 480, 285, 294, 392, 498, 511, 287, 269, 277, 278, 390, 542, 291, 286, 289, 541, 608, 265, 266, 546, 294, 550, 498, 292, 511, 475, 531, 267, 514, 318, 600, 614, 508, 387, 284, 415, 423, 503, 588, 270, 295, 553, 288, 575, 408, 534, 430, 410, 411, 432, and 385, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-8 and Table 30-2 or a reverse complementary sequence thereof.80.A circRNA produced by the self-splicing of the RNA encoded by the nucleic acid of any one of claims 44 to 79.81.A vector comprising the nucleic acid of any one of claims 44 to 79.82.A cell comprising the nucleic acid of any one of claims 44 to 79, the circRNA of claim 80, or the vector of claim 81.83.The cell of claim 82, wherein the circRNA further comprises a therapeutic protein-coding sequence (Z1) operatively linked to the TI.84.A method of expressing a protein comprising culturing the cell of claim 83 under conditions and for a sufficient time for the expression of the protein.85.A pharmaceutical composition comprising the nucleic acid of claims 44 to 79, the circRNA of claim 80, the vector of claim 81, or the cell of claim 82, and a pharmaceutically acceptable carrier, wherein the circRNA further comprises Z1 operatively linked to the TI, wherein Z1 encodes a therapeutic protein.86.The pharmaceutical composition of claim 85, wherein the expression of a therapeutic protein is in a tissue.87.The pharmaceutical composition of claim 86, wherein the expression is in vivo.88.The pharmaceutical composition of claim 86, wherein the expression is in vitro.89.The pharmaceutical composition of claim 86, wherein the expression is in muscle tissues, and the muscle tissue is from a human or a mouse.90.The pharmaceutical composition of claim 89, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 425, 287, 286, 278, 542, 423, 277, 436, 269, 294, 608, 415, 535, 541, 289, 358, 498, 292, 266, 291, 360, 281, 284, 293, 369, 550, 271, 272, 590, 267, 430, 390, 411, 297, 357, 362, 499, 283, 491, 546, 290, 282, 478, 265, 410, 432, 413, 414, 315, 392, 426, 268, 302, 614, 288, 318, 357, 575, 534, 322, 339, 576, 544, 324, 354, 336, 355, 589, 588, 295, 296, 356, 560, 514, 553, 417, 573, 409, 329, 280, 341, 316, 431, 463, 511, 600, 282, 472, 558, 421, and 538, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-1 and Table 30-3 or a reverse complementary sequence thereof.91.The pharmaceutical composition of claim 86, wherein the expression is in lung tissues, and the lung tissue is from a human or a mouse.92.The pharmaceutical composition of claim 91, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 475, 287, 425, 541, 535, 269, 266, 542, 277, 291, 498, 358, 292, 294, 286, 297, 390, 550, 608, 289, 546, 436, 282, 499, 278, 423, 293, 369, 265, 411, 478, 590, 553, 415, 272, 357, 392, 480, 360, 271, 413, 534, 281, 280, 356, 491, 284, 350, 285, 547, 484, 270, 290, 514, 267, 288, 511, 430, 614, 283, 537, 575, 339, 417, 410, 544, 432, 533, and 357, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-2 or a reverse complementary sequence thereof.93.The pharmaceutical composition of claim 86, wherein the expression is in stomach tissues, and the stomach tissue is from a human or a mouse.94.The pharmaceutical composition of claim 93, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 269, 287, 297, 286, 291, 475, 265, 266, 390, 436, 280, 542, 267, 541, 498, 608, 277, 283, 546, 272, 499, 278, 293, 271, 392, 535, 553, 478, 550, 268, 358, 425, 289, and 590, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-3 or a reverse complementary sequence thereof.95.The pharmaceutical composition of claim 86, wherein the expression is in nervous tissues, and the nervous tissue is from a human or a mouse.96.The pharmaceutical composition of claim 95, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 289, 608, 614, 590, 600, 297, 287, 285, 281, 423, 357, 266, 293, 425, 358, 414, 277, 575, 542, 392, 589, 436, 294, 541, 286, 535, 360, 576, 550, 269, 369, 290, 546, 390, 280, 268, 413, 415, 288, 283, 354, 480, 499, 534, 291, 329, 350, 347, 362, 315, 410, 282, 292, 498, 475, 278, 588, 339, 411, 272, 431, 374, 265, 271, 430, 453, 348, 478, 466, 267, 590, 284, 296, 355, 322, 357, 417, 352, 373, 533, 432, 492, 408, 472, 356, 560, 485, 318, 544, 295, 361, 316, 302, 514, 384, 488, and 511, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-4 and Table 30-1 or a reverse complementary sequence thereof.97.The pharmaceutical composition of claim 86, wherein the expression is in colonic tissues, and the colonic tissue is from a human or a mouse.98.The pharmaceutical composition of claim 97, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 608, 358, 266, 575, 289, 541, 369, 475, 614, 542, 287, 293, 269, 350, 285, 281, 357, 535, 478, 480, 600, 354, 294, 576, 282, 360, 286, 290, 297, 546, 278, 277, 291, 550, 413, 292, 425, and 295, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-5 or a reverse complementary sequence thereof.99.The pharmaceutical composition of claim 86, wherein the expression is in liver tissues, and the liver tissue is from a human or a mouse.100.The pharmaceutical composition of claim 99, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 287, 436, 293, 281, 415, 294, 269, 425, 392, 423, 290, 411, 277, 278, 413, 288, 266, 289, 268, 286, 608, 297, 614, 390, 357, 410, 358, 265, 475, 498, 432, 542, 280, 291, 466, 283, 414, 590, 360, 430, 271, 357, 292, 417, 589, 354, 267, 431, 541, 546, 535, 362, 356, 499, 285, 514, 480, 355, 350, 272, 369, 550, 295, 284, 527, 318, 339, 296, 531, 511, 329, 600, 478, 302, 590, 270, 322, 503, 491, 553, 463, 315, 575, 453, and 456, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-6 or a reverse complementary sequence thereof.101.The pharmaceutical composition of claim 86, wherein the expression is in skin tissues, and the skin tissue is from a human or a mouse.102.The pharmaceutical composition of claim 101, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 278, 608, 425, 475, 289, 277, 286, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-7 or a reverse complementary sequence thereof.103.The pharmaceutical composition of claim 86, wherein the expression is in immune tissues, and the immune tissue is from a human or a mouse.104.The pharmaceutical composition of claim 103, wherein the IRES is identical to a nucleotide sequence selected from the group consisting of SEQ ID NOs: 297, 286, 287, 535, 281, 475, 277, 296, 542, 269, 541, 358, 390, 550, 369, 360, 436, 265, 357, 354, 350, 291, 357, 266, 546, 608, 271, 289, 355, 280, 590, 278, 425, 356, 362, 484, 272, 293, 292, 283, 491, 282, 480, 285, 294, 392, 498, 511, 287, 269, 277, 278, 390, 542, 291, 286, 289, 541, 608, 265, 266, 546, 294, 550, 498, 292, 511, 475, 531, 267, 514, 318, 600, 614, 508, 387, 284, 415, 423, 503, 588, 270, 295, 553, 288, 575, 408, 534, 430, 410, 411, 432, and 385, or a reverse complementary sequence thereof; or the IRES is identical to a nucleotide sequence selected from the IRES shown in Table 29-8 and Table 30-2 or a reverse complementary sequence thereof.105.A method of expressing of a therapeutic protein in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of any one of claims 85 to 104.106.A method of treating or preventing a disease or disorder in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of any one of claims 85 to 104.107.Use of the pharmaceutical composition of any one of claims 85 to 104 for treating or preventing a disease or disorder.

Citation Information

Patent Citations

  • Circular RNA for translation in eukaryotic cells

    CN112399860A

  • IRES Elements for Expression of Polypeptides and Methods of Using the Same

    US20160040177A1

  • Defective interfering viral genomes

    WO2021191688A1

  • Circular RNA compositions and methods

    WO2022261490A2

  • Compositions and methods for improved protein translation from recombinant circular rnas

    WO2022271965A2