Genetic elements that drive circular RNA translation and methods of use

JP2026143428APending Publication Date: 2026-09-08THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026079592
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-10
Filing Date
2026-05-11
Publication Date
2026-09-08

Smart Images

  • Figure 2026143428000001_ABST
    Figure 2026143428000001_ABST
Patent Text Reader

Abstract

This invention provides recombinant circular RNA (circRNA) molecules containing an internal ribosome entry site (IRES) operably linked to a protein-coding nucleic acid sequence. [Solution] The IRES comprises at least one RNA secondary structure element and a sequence region complementary to 18S ribosomal RNA (rRNA). A method for producing proteins in cells using recombinant circRNA molecules is also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 186,507 filed on 10 May 2021 and U.S. Provisional Patent Application No. 63 / 043,964 filed on 25 June 2020, the contents of which are incorporated herein by reference in their entirety.

[0002] Sequence List The contents of the text file submitted electronically with respect to this specification are incorporated herein by reference in their entirety: a computer-readable formatted copy of the sequence listing (filename: CRCB_003_02WO_SeqList_ST25.txt, creation date: June 24, 2021, file size: approximately 12.8 megabytes).

[0003] The present invention relates to a recombinant circular RNA (circRNA) molecule containing an internal ribosome entry site (IRES) that includes RNA secondary structure elements and nucleic acid sequence regions complementary to 18S rRNA, and to a method of using the same.

[0004] Statements concerning federally funded research This invention was made with government support under contract CA209919 granted by the U.S. National Institutes of Health. The government has certain rights to this invention. [Background technology]

[0005] Over the past decade, deep sequencing and computer analysis have suggested that circular RNA (circRNA) is a large class of RNA in mammalian cells that plays a crucial role in various biological processes. Disruption of circRNA expression has been found to be associated with human diseases such as Alzheimer's disease, diabetes, and cancer. Furthermore, thanks to the exceptional stability and cell-specific expression patterns of circRNA, it has been used as a biomarker for diseases such as cancer and as an index of the effectiveness of certain treatments. While much of the research has demonstrated that circRNA functions as a non-coding RNA, such as a sponge for miRNA, a regulator of mRNA splicing mechanisms, a sequester of RNA-binding proteins (RBPs), a regulator of RBP interactions, and an activator of immune responses, the evidence suggests that some circRNAs code for peptides and / or proteins, thereby suggesting that they function through these coded polypeptides. Proteins known to be translated from circRNA regulate cell proliferation, differentiation, migration, and myogenesis. Dysregulation of circRNA-coding proteins is associated with tumorigenesis in certain cancers. Therefore, circRNA-coding proteins may be a crucial link between classes of biologically relevant circRNAs and cancer and possibly other diseases. Thus, understanding the mechanisms of circRNA translation could help develop circRNA biology and therapeutic and / or modalities that utilize its encoded proteins.

[0006] Since circRNAs are produced by spliceosome-mediated head-tail junction of premRNAs, they do not contain the 5' cap that is generally known to be necessary for cap-dependent translation. Therefore, circRNA translation utilizes alternative mechanisms to initiate cap-independent translation, such as the use of internal ribosome entry site (IRES) sequences recognized by ribosomes. The introduction of IRESs onto synthetically produced circRNAs is sufficient to initiate translation of the encoded circRNA protein, thereby suggesting that endogenous circRNAs harboring IRES sequences may be translatable once transported into the cytoplasm. [Overview of the project] [Problems that the invention aims to solve]

[0007] Given that this field is rapidly advancing but still in its infancy, the need remains not only for IRESs but also for the identification and characterization of genetic elements that can promote, initiate, direct, or regulate circRNA translation. In particular, there is a need to identify novel IRES sequences that can operatively promote the expression of circRNA-encoded proteins. [Means for solving the problem]

[0008] (Summary of the invention) This disclosure provides polynucleotides (e.g., DNA sequences) encoding one or more circular RNA (circRNA) molecules, wherein the circular RNA molecules comprise a payload sequence region (e.g., a protein-coding or non-coding sequence region) and an internal ribosome entry site (IRES) sequence region operably ligated to the payload sequence region. In some embodiments, the IRES comprises at least one sequence region having an RNA secondary structure element and a sequence region complementary to 18S ribosomal RNA (rRNA). In some embodiments, the IRES has a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C. Some embodiments of this disclosure include embodiments in which the RNA secondary structure sequence region or element is formed from nucleotides at positions approximately 40 to approximately 60 of the IRES, where the first nucleotide at the 5' end of the IRES is considered to be position 1.

[0009] The disclosure also provides a polynucleotide (e.g., a DNA sequence) encoding a circular RNA molecule, wherein the circular RNA molecule comprises a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES), and the IRES is encoded by one of the following: the nucleic acid sequences of SEQ ID NOs. 1-228 or SEQ ID NOs. 229-17201, or a nucleic acid sequence having at least 90% or at least 95% identity or homology thereto over at least 50% of the length of the nucleic acid sequence.

[0010] This disclosure also provides recombinant circular RNA molecules comprising a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence region, wherein the IRES comprises at least one sequence region having secondary structural elements and a sequence region complementary to 18S ribosomal RNA (rRNA), and the IRES has a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C. In some embodiments, the protein-coding nucleic acid sequence region is operably linked to the IRES in a non-natural configuration.

[0011] This disclosure also provides recombinant circular RNA molecules comprising a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence region, wherein the IRES is encoded by one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a nucleic acid sequence having at least 90% or at least 95% homology or identity thereto. In some embodiments, the protein-coding nucleic acid sequence region is operably linked to the IRES in a non-natural configuration.

[0012] A method for producing proteins in cells using the aforementioned recombinant circular RNA molecule, or a polynucleotide encoding it (e.g., a DNA molecule), is also provided.

[0013] A vector containing the aforementioned recombinant circular RNA molecule, or a DNA molecule encoding it, is also provided.

[0014] A host cell containing the aforementioned recombinant circular RNA molecule, or a DNA molecule encoding it, is also provided.

[0015] (i) a DNA sequence encoding a circular RNA, and (ii) a composition comprising a non-coding circular RNA or a DNA sequence encoding it are also provided.

[0016] A method for delivering a non-coding circular RNA to a cell is also provided, comprising the step of contacting the cell with a DNA sequence encoding a circular RNA, and (ii) a composition comprising a non-coding circular RNA or a DNA sequence encoding it, thereby delivering the non-coding circular RNA to the cell.

[0017] This disclosure further provides oligonucleotides comprising a nucleic acid sequence region that hybridizes to an internal ribosome entry site (IRES) sequence region present on a circular RNA molecule, thereby inhibiting translation of the coding sequence region of the circular RNA molecule. A method for inhibiting translation of the protein-coding nucleic acid sequence region (e.g., payload) of a circular RNA molecule using the aforementioned oligonucleotides is also provided.

[0018] These and other embodiments are described in more detail below and in the accompanying drawings.

[0019] A patent or application file includes at least one drawing made in color. A copy of this patent or patent application publication, including the color drawing, will be provided by the Patent Office upon request and payment of the required fees. [Brief explanation of the drawing]

[0020] [Figure 1A] Figure 1A shows a high-throughput identification of RNA sequences that can promote cap-independent translational activity on circRNA. Figure 1A provides a schematic diagram of a high-throughput split-eGFP circRNA reporter screening assay for identifying circRNA IRESs. A synthetic oligo library containing 55,000 oligos was cloned into a split-eGFP circRNA reporter. Since full-length eGFP is reconstituted only when back-spliced ​​onto circRNA, the eGFP fluorescence signal can only arise from cap-independent translational activity driven by the inserted oligo on the circRNA. eGFP(+) cells were sorted into seven expression bins by eGFP fluorescence intensity by FACS. The number of reads for each synthetic oligo in each bin was determined by next-generation DNA sequencing. Final eGFP expression for each synthetic oligo was quantified by the mean weighted bin number, according to the distribution of the number of reads across the seven expression bins from two independent biological replicas. [Figure 1B]Figure 1B is a diagram showing high-throughput identification of RNA sequences capable of promoting cap-independent translation activity on circRNA. Shown in Figure 1B is the eGFP expression distribution of 40,855 captured synthetic oligos. eGFP(+) oligos were defined as oligos with higher eGFP expression than the background eGFP threshold (eGFP expression of an eGFP circRNA reporter with no inserted oligo). The pie chart represents the composition of different oligo categories among the eGFP(+) oligos. [Figure 1C] Figure 1C is a diagram showing high-throughput identification of RNA sequences capable of promoting cap-independent translation activity on circRNA. Shown in Figure 1C is quantification of the percentage of captured eGFP(+) oligos among oligos derived from sequences of IRES, viral 5'UTRs, or human 5'UTRs reported in the screening assay. [Figure 1D] Figure 1D is a diagram showing high-throughput identification of RNA sequences capable of promoting cap-independent translation activity on circRNA. Shown in Figure 1D is the identification of circular and linear RNA-specific IRES. Normalized eGFP expression (log10) is shown for each captured oligo in screening assays performed using the same synthetic oligo library in circular RNA (as described herein) or linear RNA (Weingarten-Gabbay et al., 2016) screening systems. Circular IRES (green dots) or linear IRES (blue dots) were identified by comparing the IRES activity of oligos detected only in either the circular or linear RNA screening system, respectively. The red dashed line represents the normalized eGFP expression threshold. [Figure 2A] Figure 2A is a diagram showing that circRNA containing eGFP(+) oligos has higher cap-independent translation activity. Figure 2A shows a schematic diagram of the circRNA polysome profiling method for capturing translated circRNA. [Figure 2B]Figure 2B is a diagram showing that circRNA containing eGFP(+) oligonucleotides have higher cap-independent translation activity. Figure 2B shows (poly)ribosome fractions of cells transfected with a split-eGFP circRNA reporter containing a synthetic oligonucleotide library, followed by cycloheximide (CHX) treatment. Fractions 7-12 (shaded blue) were determined as (poly)ribosome fractions according to the Abs254 pattern. [Figure 2C] Figure 2C is a diagram showing that circRNA containing eGFP(+) oligonucleotides have higher cap-independent translation activity. Shown in Figure 2C is quantification of the percentage of (poly)ribosome-enriched oligonucleotides among captured eGFP(-) oligonucleotides with eGFP expression below the 20th percentile or eGFP(+) oligonucleotides with eGFP expression above the 80th percentile, respectively. [Figure 2D] Figure 2D is a diagram showing that circRNA containing eGFP(+) oligonucleotides have higher cap-independent translation activity. Figure 2D provides sequencing reads from Ribo-seq and QTI-seq plotted against genes showing eGFP(+) oligonucleotides with aTIS (top), nTIS (middle), and dTIS (bottom), along with overlapping annotated circRNAs (brown segments). [Figure 2E] Figure 2E is a diagram showing that circRNA containing eGFP(+) oligonucleotides have higher cap-independent translation activity. Figure 2E shows quantification of the percentage of eGFP(-) or eGFP(+) oligonucleotides without TIS (TIS(-)) (left) or with more than one TIS (TIS(+)) (right), and the percentage of aTIS, nTIS, or dTIS oligonucleotides among eGFP(+) / TIS(+) oligonucleotides. [Figure 3A] Figure 3A is a diagram showing that the 18S rRNA complementary sequence on the IRES promotes cap-independent translation activity of circRNA. Provided in Figure 3A is a schematic diagram of the sliding window design of synthetic oligonucleotides for mapping active regions on human 18S rRNA. [Figure 3B] Figure 3B shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3B also shows the quantification of mean eGFP expression of synthetic oligos overlapping with corresponding positions across human 18S rRNA. The dashed line indicates background eGFP expression. Identified active regions on 18S rRNA are shaded in green. [Figure 3C] Figure 3C shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3C provides a diagram of the secondary structure of human 18S rRNA, showing identified active regions and reported RNA contact regions on 18S rRNA. Identified active regions 1-6 are shaded in green. Boxes indicate regions on 18S rRNA that have been reported to contact mRNA (red) or IRES RNA (orange). [Figure 3D] Figure 3D shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3D shows the quantification of the number of 18S rRNA active heptamers or random heptamers in eGFP(+) or eGFP(-) oligos, plotted in a Tukey box plot (outliers are not shown). Ns: Not significant; ****: p-value < 0.001 by unpaired two-sample t-test. [Figure 3E] Figure 3E shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3E shows the quantification of IRES activity for oligos with higher or lower 18S rRNA complementarity, as determined by FACS (MFIeGFP / mRuby). *: p-value < 0.05 for wild-type (WT) oligo by unpaired two-sample t-test (n=4-6 independent replications). Error bars: SEM. [Figure 3F]Figure 3F shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3F also provides a schematic diagram of the design of synthetic oligos for systematic scanning mutagenesis. [Figure 3G] Figure 3G shows that the 18S rRNA complementary sequence on the IRES promotes circRNA cap-independent translational activity. Figure 3G shows eGFP expression for each oligo containing random substitution mutations at the corresponding positions on the HCV IRES. Black dots represent the start sites of each mutation on the IRES. Identified essential elements are shaded in blue. Red lines represent reported functional domains on the HCV IRES. eGFP expression for each oligo was normalized to the average eGFP expression of all oligos on the HCV IRES. [Figure 3H] Figure 3H shows that the 18S rRNA complementary sequence on the IRES promotes circRNA cap-independent translational activity. Figure 3H provides examples of circRNA IRESs with local and global sensitivity identified by scanning mutagenesis. Identified essential elements on the IRES are shaded in blue. [Figure 3I] Figure 3I shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. In Figure 3I, the average eGFP expression of all circRNA IRES oligos with overall sensitivity is shown at each mutation site across the IRES. Regions containing regulatory elements are shaded in 10 different colors (blue: 5–15 nt and 135–165 nt; red: 40–60 nt). [Figure 3J] Figure 3J shows that the 18S rRNA complementary sequence on IRES promotes circRNA cap-independent translational activity. Figure 3J shows the quantification of local MFE in a 15-nucleotide (nt) sliding window on IRES. Regions containing regulatory elements are shaded in different colors (blue: 5–15nt and 135–165nt; red: 40–60nt). [Figure 4A]Figure 4A shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4A shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4B] Figure 4B shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4B shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4C]Figure 4C shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4C shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4D] Figure 4D shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4D shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4E]Figure 4E shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4E shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4F] Figure 4F shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4F shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4G]Figure 4G shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4G shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4H] Figure 4H shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4H shows the secondary structures of mutant IRESs (SEQ ID NOs. 33925-33932) determined by M2-seq. The arrows indicate the high-confidence secondary structures identified by M2-net, and the corresponding positions are labeled with the same color in the RNA structure panel. The red arrows indicate the SuRE at the 40-60 nt position on the circular IRES. CircIRES-dis: Circular IRES with a disrupted SuRE by sequence substitution. CircIRES-rearrangement: Circular IRES with a rearranged SuRE in the 90-110 nt region. CircIRES-single and circIRES-comp: Circular IRESs with a single complementary mutation and a compensatory double complementary mutation, respectively. circIRES-BoxB: Circular IRES with a SuRE substituted in the BoxB stem-loop. Linear IRES-addition: A linear IRES having a 40-60nt region substituted with SuRE at the 40-60nt position on a cyclic IRES. [Figure 4I]Figure 4I shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can enhance circular IRES activity. Figure 4I shows the quantification of IRES activity for each mutant IRES determined by FACS (MFIeGFP / mRuby). The activity of each rearranged IRES was normalized relative to linear IRES. Ns: not significant; **: p-value < 0.01, ****: p-value < 0.001 relative to linear IRES by independent two-sample t-test (n=4-6 independent replications). Error bars: SEM. [Figure 4J] Figure 4J shows that a separate SuRE at the 40-60 nucleotide (nt) position on the IRES can enhance cyclic IRES activity. Figure 4J shows the quantification of the percentages of eGFP(+) oligos (left) and endogenous translated circRNA (right) that have 18S rRNA complementarity or a SuRE element. [Figure 4K] Figure 4K shows that a distinct SuRE at the 40-60 nucleotide (nt) position on the IRES can promote circular IRES activity. Figure 4K provides a diagram of two key regulatory elements that promote circRNA cap-independent translation: a complementary 18S rRNA sequence and a SuRE at the 40-60 nt position on the IRES. [Figure 5A] Figure 5A shows that the IRES element promotes translation initiation of endogenous circRNA. Figure 5A also shows a schematic diagram of disrupting key regulatory elements on the IRES of the oligo-split-eGFP-circRNA reporter by co-transfecting with antisense LNAs that target specific regions on the IRES. LNA-18S: LNA that targets the 18S rRNA complementary sequence on the IRES; LNA-SuRE: LNA that targets SuRE at the 40-60nt position on the IRES; LNA-Rnd: LNA that targets a random position downstream of LNA-18S or LNA-SuRE on the IRES. [Figure 5B]Figure 5B shows that the IRES element promotes translation initiation of endogenous circRNA. Figure 5B shows the quantification of normalized eGFP fluorescence signal intensity in cells co-transfected with an oligo-split-eGFP-circRNA reporter containing the corresponding LNA and IRES. Numbers represent the oligo index number. Ns: not significant; *: p-value < 0.05; **: p-value < 0.01, ***: p-value < 0.005 for mock transfection by unpaired two-sample t-test (n=3-5 independent replications). Error bars: SEM. [Figure 5C] Figure 5C shows that the IRES element promotes the initiation of translation of endogenous circRNA. Figure 5C provides a schematic diagram of the QTI-qRT-PCR quantification of the level of translation-initiating endogenous circRNA. [Figure 5D] Figure 5D shows that the IRES element promotes translation initiation of endogenous circRNA. Figure 5D also shows the quantification of translation initiation RNA levels of human endogenous circRNA containing the corresponding IRES after disruption of the IRES by the corresponding LNA transfection. CircRNA levels were normalized to GAPDH mRNA. Ns: not significant; *: p-value < 0.05; **: p-value < 0.01, ***: p-value < 0.005 for mock transfection by unpaired two-sample t-test (independent replications of n=4-6). Error bars: SEM. [Figure 5E] Figure 5E shows that the IRES element promotes the initiation of translation of endogenous circRNA. Figure 5E also shows a Western blot image indicating the level of protein produced from endogenous circRNA upon IRES disruption by transfection of the corresponding LNA. [Figure 6A]Figure 6A demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6A shows the quantification of the percentage of IRES-mapped human endogenous circRNAs that have one or more eGFP(+) oligo sequences (IRES(+)circRNAs) or that do not have eGFP(+) oligo sequences (IRES(-)circRNAs). [Figure 6B] Figure 6B illustrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6B also shows a quantification of the distribution of parental genes among IRES(+)circRNAs. Each portion of the pie chart represents a different gene. [Figure 6C] Figure 6C illustrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6C shows the quantification of potential cancer-associated IRES(+)circRNAs from CSCD. [Figure 6D] Figure 6D demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6D provides a histogram showing the distribution of the number of IRESs present in each individual IRES(+)circRNA (up to n=20). [Figure 6E] Figure 6E demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6E is a histogram showing the distribution of the number of circRNAs mapped for each individual eGFP(+) oligo (up to n=20). [Figure 6F] Figure 6F demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6F shows the top 12 representative biological processes from GO-term analysis that are rich in the parental genes of IRES(+)circRNAs. [Figure 6G] Figure 6G illustrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6G provides a schematic diagram for generating a putative endogenous circORF list. [Figure 6H]Figure 6H illustrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6H shows the top 15 representative conserved motifs from Pfam analysis, which are rich in polypeptides encoded by the predicted circRNAs. [Figure 6I] Figure 6I illustrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6I shows a schematic diagram of the peptidomic validation of putative circORFs. [Figure 6J] Figure 6J demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6J provides a heatmap showing the number of unique trypsin polypeptides detected in the peptidomic dataset of each MS-captured circORF. [Figure 6K] Figure 6K demonstrates the identification of proteins encoded by putative endogenous circRNAs. Figure 6K shows the MS1 and MS2 spectra of a representative trypsin BSJ polypeptide (SEQ ID NO: 33933) captured from circORF_575. [Figure 6L] Figure 6L demonstrates the identification of proteins encoded by putative endogenous circRNA. Figure 6L shows representative MS2 spectra and the top three ranks of PRM-MS transition ion spectra for spike-in heavy isotope-labeled polypeptide from circORF_19 (upper right (SEQ ID NO: 33934)) and sample trypsin polypeptide (lower right (SEQ ID NO: 33934)). [V]: Heavy isotope-labeled valine (13C5, 15N; +6Da). [Figure 7A] Figure 7A shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7A provides a schematic diagram of the CDS of FGFR1 and circFGFR1 transcripts. [Figure 7B]Figure 7B shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7B also shows a schematic diagram of the design of the junctional RT-PCR primer (black arrow) and the Sanger sequencing results for detecting the backsplicing junction of circFGFR1 (sequence number 33935) (yellow box). [Figure 7C] Figure 7C shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7C provides a schematic diagram of conserved motifs on FGFR1 and circFGFR1p. Ab (both): Antibody capable of detecting both FGFR1 and circFGFR1p. Ab-circFGFR1p: Custom circFGFR1p antibody. The blue lines indicate the position of the antigen peptide for each antibody. [Figure 7D] Figure 7D shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7D shows a schematic diagram of polypeptides captured by IP-LC-MS. MS (underlined) matching the specific region of circFGFR1p (red) and the region overlapping with FGFR1 (black) using a custom antibody against the specific region of circFGFR1p (bold) (circFGFR1p (SEQ ID NO: 33902); circFGFR1p fragment (SEQ ID NO: 33936)). Regions extracted on a Coomassie blue-stained SDS-PAGE gel (approximately 30-45 kDa) are indicated by red boxes. [Figure 7E] Figure 7E shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7E shows representative MS2 spectra and the top three ranks of PRM-MS transition ion spectra for the spike-in heavy isotope-labeled polypeptide (top (SEQ ID NO: 33937)) and BJ trypsin polypeptide (bottom (SEQ ID NO: 33937)) of circFGFR1p. [L]: Heavy isotope-labeled leucine (13C6, 15N; +7Da). [Figure 7F]Figure 7F shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7F provides images of FGFR1 (red), circFGFR1p (green), and DAPI (blue) in HEK-293T cells co-transfected with plasmids expressing HA-FGFR1 and FLAG-circFGFR1p without permeabilization. Scale bar: 10 micrometers. [Figure 7G] Figure 7G shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7G shows Western blots indicating circFGFR1p and FGFR1 protein levels (both Ab and Ab), as well as quantification of FGFR1 and circFGFR1 RNA levels by qRT-PCR of cells transfected with siRNA or LNA. siCtrl: untargeted siRNA; siCircFGFR1: circFGFR1-specific siRNA; circFGFR1-LNA: antisense LNA oligo targeting the 18S rRNA complementary sequence on circFGFR1 IRES. P-FGFR1: phosphorylated FGFR1. Ns: not significant; **p-value <0.01 for siCtrl by unpaired two-sample t-test (n=3 independent replications). Error bars: SEM. [Figure 7H] Figure 7H shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. (Figure 7H) Quantification of cell proliferation in cells with knockdown of circFGFR1 RNA (siCircFGFR1) or circFGFR1p (circFGFR1-LNA) from day 1 to day 4 by FGF1 addition is shown. *p value < 0.05; **p value < 0.01; ***p value < 0.005 relative to siCtrl by unpaired two-sample t-test (independent replications of n=3-5). Error bars: SEM. [Figure 7I]Figure 7I shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7I provides Western blot images (left panel) showing cells with overexpression of FGFR1, circFGFR1p, or FGFR1+circFGFR1p, and their corresponding cell proliferation from day 1 to day 4 after FGF1 addition (right panel). Ns: not significant; *p value < 0.05, **p value < 0.01, ****p value < 0.001 for mock transfections by unpaired two-sample t-test (independent replications of n=4-6). Error bars: SEM. [Figure 7J] Figure 7J shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7J provides Western blot images showing FGFR1 protein and circFGFR1p levels with and without heat shock. [Figure 7K] Figure 7K shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7K shows Western blot quantification of circFGFR1p protein levels for FGFR1 (all isoforms) under normal (WT) and heat shock (HS) conditions. Error bars: SEM from three independent blots. *p-value < 0.05 for WT by unpaired two-sample t-test (n=3 independent blots). [Figure 7L] Figure 7L shows that circFGFR1p, encoded by circRNA, suppresses cell proliferation under stress conditions. Figure 7L shows quantifications of Western blots illustrating changes in FGFR1 and circFGFR1p protein levels under heat shock conditions. Protein levels are normalized to the GAPDH protein-loading control for each condition. Error bars: SEM from three independent blots. Ns: Not significant; *p value < 0.05 relative to 1 by 1-sample t-test (n=3 independent blots). [Figure 8A]Figure 8A shows that the oligo-split-eGFP-circRNA reporter construct does not generate an eGFP signal from transsplicing. Figure 8A shows Northern blot images of cells transfected with the IRES-split-eGFP circRNA reporter using probes for mRuby, 3'eGFP, and the eGFP backsplicing junction region on the reporter transcript, with or without RNase R treatment, and then sorted by mRuby(+) / eGFP(+). [Figure 8B] Figure 8B shows that the oligo-split-eGFP-circRNA reporter construct does not generate an eGFP signal from trans-splicing. Figure 8B shows the quantification of eGFP circRNA or mRuby linear transcript RNA levels in total RNA of cells transfected with the IRES-split-eGFP circRNA reporter and sorted by mRuby(+) / eGFP(+), compared to RNaseR(-) samples. RNA levels were normalized to GAPDH mRNA levels in each sample. Error bars: SEM. Ns: Not significant; ****: p-value <0.001 (n=3 independent replications) compared to RNaseR(-) samples by independent two-sample t-test. Error bars: SEM. [Figure 8C] Figure 8C shows that the oligo-split-eGFP-circRNA reporter construct does not generate an eGFP signal from trans-splicing. Figure 8C also shows flow cytometry analysis of eGFP(+) cells transfected with the corresponding reporter construct. The eGFP(+) cells were gated according to the cells that underwent mock transfection. [Figure 9A]Figure 9A shows high-throughput identification of IRES sequences that can promote cap-independent translational activity on circRNA. Figure 9A provides a reproducibility measurement of eGFP expression for each captured oligo from two independent biological replicas from a screening assay. Only oligos recovered in both replicas were included in the analysis. R represents Pearson's correlation coefficient. [Figure 9B] Figure 9B shows a high-throughput identification of IRES sequences that can promote cap-independent translational activity on circRNA. Figure 9B provides a schematic diagram of primer designs for quantifying the expression levels of linear and cyclic transcripts of the reporter construct. Branched cyclic primers extending to the backsplicing junction of circRNA should detect only the cyclic transcript. [Figure 9C] Figure 9C shows high-throughput identification of IRES sequences that can promote cap-independent translational activity on circRNA. Figure 9C shows quantification of cyclization efficiency by qRT-PCR of seven randomly selected clones transfected with an oligo-split-eGFP reporter plasmid. Cyclization efficiency was calculated by normalizing the expression level of the cyclized transcript to the expression level of the linear transcript. Numbers indicate the oligo index. No IRES: Reporter plasmid without IRES insertion. Ns: Not significant against empty circRNA by unpaired two-sample t-test (n=3 independent replicas). Error bars: SEM. [Figure 9D] Figure 9D shows the high-throughput identification of IRES sequences that can promote cap-independent translational activity on circRNA. Figure 9D shows the distribution of read fractions across all seven bins of cells transfected with an IRES-split-eGFP circRNA reporter containing oligos that either do not contain IRES (background eGFP expression) or oligos that exhibit high (oligo #25674), medium (oligo #26338), or none (oligo #26961) cap-independent translational activity. The black line represents the polynomial trend line of the distribution. [Figure 9E]Figure 9E shows high-throughput identification of IRES sequences that can promote cap-independent translational activity on circRNA. Figure 9E provides images of Western blots showing the expression levels of eGFP, Cre, and CD4 from split-eGFP circRNA reporters that do not contain the IRES or contain the corresponding IRES. The numbers indicate the oligo index. [Figure 9F] Figure 9F shows the high-throughput identification of IRES sequences that can promote cap-independent translation activity on circRNAs. Figure 9F provides Western blot images showing eGFP expression levels for cap-dependent translation linear RNA (CMV promoter-driven) and cap-independent translation circRNA (IRES-driven; oligo #8788). [Figure 10A] Figure 10A shows that a high-throughput IRES screening assay can capture IRESs from both viral and human 5'UTR. This figure shows examples of IRESs captured in the screening assay with the top 10 eGFP-expressing IRESs (i.e., linear IRESs) reported (Figure 10A). [Figure 10B] Figure 10B shows that a high-throughput IRES screening assay can capture IRESs from viral and human 5'UTR. This figure shows examples of IRESs captured in a screening assay with the top 10 eGFP expressions from viral 5'UTR (Figure 10B). [Figure 10C] Figure 10C shows that a high-throughput IRES screening assay can capture IRESs from both viral and human 5'UTR. This figure shows examples of IRESs captured in a screening assay with the top 10 eGFP expressions from human 5'UTR (Figure 10C). [Figure 11A]Figure 11A shows the IRES composition of captured linear and circular IRESs. Figure 11A provides a Venn diagram showing the number of circular and linear-specific IRESs by comparing the results from the study (circular RNA system) with the results from the study (linear RNA system) described in Weingarten-Gabbay et al., Science 351, aad4939 (2016). [Figure 11B] Figure 11B shows the IRES composition of the captured linear and cyclic IRESs. Figure 11B shows the composition of captured viral and human IRESs in the cyclic IRES (Figure 11B). [Figure 11C] Figure 11C shows the IRES composition of the captured linear and cyclic IRESs. Figure 11C shows the composition of captured viral and human IRESs in linear IRESs (Figure 11C). [Figure 11D] Figure 11D shows the IRES composition of captured linear and cyclic IRESs. The composition shown is that of IRESs (Figure 11D) exhibiting cap-independent translational activity in both linear and cyclic RNA systems. [Figure 12A] Figure 12A shows that circRNAs containing eGFP(+) oligo sequences are more actively translated. Figure 12A shows the 40S and (poly)ribosome fractions of cells transfected with a split-eGFP reporter containing a synthetic oligo library, treated with puromycin (PMY; left panel) or cycloheximide (CHX; right panel), and subsequently subjected to sucrose gradient fractionation. [Figure 12B] Figure 12B shows that circRNAs containing the eGFP(+) oligo sequence are more actively translated. Figure 12B quantifies the ratio of eGFP circRNA levels to mRuby linear transcript levels in cells transfected with an IRES-split-eGFP circRNA reporter and sorted by mRuby(+) / eGFP(+) after RNase R treatment over time (20U of RNase R per 20μg of RNA). Error bars: SEM. [Figure 12C] Figure 12C shows that circRNAs containing eGFP(+) oligo sequences are more actively translated. Figure 12C shows the quantification of the total number of captured oligo reads in the 40S and (poly)ribosome fractions treated with PMY or CHX. [Figure 12D] Figure 12D shows that circRNAs containing eGFP(+) oligo sequences are translated more actively. Figure 12D provides the number of eGFP(-) and eGFP(+) oligos captured before and after RNase R treatment. [Figure 12E] Figure 12E shows that circRNAs containing eGFP(+) oligo sequences are translated more actively. Figure 12E also shows the normalized number of reads for all captured oligos and captured oligos in the poly(ribosome) fraction after RNase R treatment. [Figure 13A] Figure 13A shows that eGFP(+) oligos more frequently overlap with translation initiation sites (TIS) on the human genome. Figure 13A quantifies the number of TIS reads in eGFP(+) or eGFP(-) oligos on the human genome. **** indicates a p-value < 0.001 from an unpaired two-sample t-test. Error bars: SEM. [Figure 13B] Figure 13B shows that eGFP(+) oligos more frequently overlap with translation initiation sites (TIS) on the human genome. Figure 13B shows the mapped locations of TIS on each TIS(+) oligo plotted on the oligo. TIS locations were selected based on their mapped locations on the oligo. [Figure 13C] Figure 13C shows that eGFP(+) oligos more frequently overlap with translation initiation sites (TIS) on the human genome. Figure 13C shows the percentage of active heptamers at each position on the oligo out of all eGFP(+) oligos. [Figure 13D]Figure 13D shows that eGFP(+) oligos more frequently overlap with translation initiation sites (TIS) on the human genome. Figure 13D shows the cumulative frequency distribution of the number of RRACH motifs on eGFP(+) and eGFP(-) oligos. Ns: The Kolmogorov-Smirnov cumulative distribution test is not statistically significant. [Figure 14A] Figure 14A shows a comparison of the characteristics between linear and circular IRES sequences. Figure 14A shows the quantification of GC content (left) and MFE (right) of circular and linear IRES sequences, plotted as Tukey box plots (outliers are not shown). **** indicates a p-value < 0.001 from an independent two-sample t-test. [Figure 14B] Figure 14B shows a comparison of features between linear and circular IRES sequences. Figure 14B also shows the cumulative frequency distribution of the number of canonical translation start codons (ATGs) on circular and linear IRES sequences. The Kolmogorov-Smirnov cumulative distribution test did not show significant results. [Figure 14C] Figure 14C shows a comparison of features between linear and circular IRES sequences. Figure 14C shows the cumulative frequency distribution of the number of m6A motifs (RRACH, SEQ ID NO: 3394) on circular and linear IRES sequences. The cumulative distribution is not statistically significant according to the Kolmogorov-Smirnov cumulative distribution test. [Figure 14D] Figure 14D shows a comparison of features between linear and circular IRES sequences. Figure 14D quantifies the number of Kozak sequences (ACCATGG, SEQ ID NO: 33945) in circular and linear IRES sequences. Ns: Not significant by unpaired two-sample t-test. Error bars: SEM. [Figure 14E]Figure 14E shows a comparison of features between linear and circular IRES sequences. Figure 14E also shows the quantification of IRES activity of oligonucleotides in a circular RNA reporter (left) and a linear RNA reporter (right), respectively. IRES activity was determined by FACS by normalizing the fluorescence intensity of eGFP medium driven by the oligonucleotide (MFIeGFP) to the linear RNA expression level of the reporter construct, determined by the fluorescence intensity of mRuby medium (MFImRuby). The values ​​were further normalized to oligonucleotide-6472. Ns: not significant, *: p-value < 0.05, ***: p-value < 0.005, ****: p-value < 0.001 for oligonucleotide-6742 by independent two-sample t-test (n=4-6 independent replications). Error bars: SEM. [Figure 14F] Figure 14F shows a comparison of features between linear and circular IRES sequences. Figure 14F shows the secondary structures of exemplary circular IRESs (sequence numbers 33938-33940) and linear IRESs (sequence numbers 33941-33943) (three IRESs for each) determined by M2-seq. Arrows indicate reliable secondary structures identified by M2-net, and the corresponding positions are labeled in the same color on the RNA structure panel. Red arrows indicate SuREs at positions 40-60 nt on the circular IRES. [Figure 14G] Figure 14G shows a comparison of features between linear and circular IRES sequences. Figure 14G displays the secondary structures of exemplary circular IRESs (sequence numbers 33938-33940) and linear IRESs (sequence numbers 33941-33943) (three IRESs for each) determined by M2-seq. Arrows indicate reliable secondary structures identified by M2-net, and the corresponding positions are labeled in the same color on the RNA structure panel. Red arrows indicate SuREs at positions 40-60 nt on the circular IRES. [Figure 15A]Figure 15A shows that the IRES element promotes translation initiation of endogenous circRNA. Figure 15A shows the quantification of eGFP circRNA levels relative to mRuby linear transcript levels in cells co-transfected with oligo-split-eGFP-circRNA reporters containing the corresponding LNA and the corresponding IRES. Ns: not significant; *: p-value <0.05 for mock transfection by unpaired two-sample t-test (independent replications of n=4-6). Error bars: SEM. [Figure 15B] Figure 15B shows that the IRES element promotes translation initiation of endogenous circRNA. Figure 15B also shows the quantification of human endogenous circRNA levels with corresponding IRESs after disruption of the IRES by corresponding LNA transfection. CircRNA levels were normalized to GAPDH mRNA. Ns: Not significant compared to mock transfections by unpaired two-sample t-test (n=4-6 independent replications). Error bars: SEM. [Figure 16A] Figure 16A shows the identification of polypeptides encoded by putative endogenous circRNAs. Figure 16A shows the quantification of the percentage of all endogenous human circRNAs that do not have an oligo sequence, do not have an eGFP(+) oligo sequence, or have one or more eGFP(+) oligo sequences. [Figure 16B] Figure 16B shows the identification of polypeptides encoded by putative endogenous circRNAs. Provided in Figure 16B is a histogram showing the distribution of the number of IRESs present in each individual IRES(+)circRNA (up to n=20) among the circRNAs generated from 159 transcripts designed so that the oligos tile across the entire transcript. [Figure 16C]Figure 16C shows the identification of polypeptides encoded by putative endogenous circRNAs. Provided in Figure 16C is a histogram showing the distribution of the number of mapped circRNAs for each individual eGFP(+) oligo (up to n=20) among the circRNAs generated from 159 transcripts designed so that the oligos tile across the entire transcript. [Figure 16D] Figure 16D shows the identification of polypeptides encoded by putative endogenous circRNAs. Figure 16D shows the distribution of distances from the backsplicing junction to the mapped IRES on each circRNA (with an upper limit of nt=2000). This distance is calculated from the backsplicing junction to the first mapped nucleotide of the IRES. GC-matched oligos: RNA sequences on the circRNA that have the same length and GC content as the mapped IRES. The distance of GC-matched oligos was determined on each IRES-mapped circRNA by taking the average distance from the backsplicing junction to all GC-matched oligos on the circRNA. [Figure 16E] Figure 16E shows the identification of polypeptides encoded by putative endogenous circRNAs. Figure 16E also shows the quantification of the percentage of polypeptides encoded by circRNAs that have an ORF overlapping with the IRES region on the circRNA. [Figure 16F] Figure 16F shows the identification of polypeptides encoded by putative endogenous circRNAs. Figure 16F shows the quantification of the percentage of polypeptides encoded by circRNAs that have an infinitely recursive ORF on the circRNA among the IRES duplicated ORFs. [Figure 16G] Figure 16G shows the identification of polypeptides encoded by putative endogenous circRNA. Figure 16G shows a Western blot image of cells transfected with a split-eGFP circRNA reporter containing in-frame IRES (Oligo-2007) showing infinitely recursive eGFP translation. [Figure 16H] Figure 16H shows the identification of polypeptides encoded by the putative endogenous circRNA. Figure 16H provides a histogram showing the size distribution of the predicted circRNA-encoded polypeptides. [Figure 16I] Figure 16I shows the identification of polypeptides encoded by putative endogenous circRNA. Figure 16I shows the quantification of the percentage of circORFs with matching sORFs using mapped IRES ORF analysis or conventional ORF analysis. [Figure 16J] Figure 16J shows the identification of polypeptides encoded by putative endogenous circRNA. Figure 16J shows the quantification of the percentage of circRNAs identified by peptidomics containing at least one RFP fragment that uniquely overlaps with the backsplicing junction. [Figure 16K] Figure 16K shows the identification of polypeptides encoded by putative endogenous circRNA. Figure 16K shows the percentage coverage of polypeptides identified by MS-20 on each protein with different expression levels in human iPSCs. MS profiling data were obtained and mapped as described by Chen et al. (2020). [Figure 16L] Figure 16L shows the identification of polypeptides encoded by putative endogenous circRNA. Figure 16L shows polypeptide coverage identified by MS against low-expression EGFR protein in human iPSCs. The red boxes represent the mapped locations of polypeptides identified by MS on the protein. [Figure 17A] Figure 17A shows that circFGFR1p, encoded by circRNA, suppresses cell growth under stress conditions. Figure 17A also shows H3K4me3 levels obtained from encoding on genomic regions of FGFR1 and circFGFR1 that do not show enrichment of the promoter signature near the circFGFR1 IRES. [Figure 17B]Figure 17B shows that circFGFR1p, encoded by circRNA, suppresses cell growth under stress conditions. Figure 17B provides images of FGFR1 (red), circFGFR1p (green), and DAPI (blue) in HEK-293T cells co-transfected with permeabilized FGFR1 and plasmids expressing FLAG-circFGFR1p. Scale bar: 10 micrometers. [Figure 17C] Figure 17C shows that circFGFR1p, encoded by circRNA, inhibits cell growth under stress conditions. Figure 17C shows quantification of circFGFR1 expression levels in tumor and normal adjacent samples plotted as Tukey box plots (outliers not shown). Data were extracted from TCGA analysis without filtering (Nair et al., Oncotarget 7, 80967, (2016)). ERBC: estrogen receptor-positive breast cancer; TNBC: triple-negative breast cancer. [Figure 17D] Figure 17D shows that circFGFR1p, encoded by circRNA, inhibits cell growth under stress conditions. Figure 17D also shows quantification of circFGFR1 expression levels in non-transformed and cancer cell lines, plotted as Tukey box plots. Data were extracted from the CSCD database (Xia et al., 2018). [Figure 17E] Figure 17E shows that circFGFR1p, encoded by circRNA, inhibits cell growth under stress conditions. Figure 17E shows quantification of circFGFR1 IRES activity with and without heat shock. Ns: Not significant compared to normal conditions (WT) by unpaired two-sample t-test (n=3 independent replications). Error bars: SEM. [Figure 17F]Figure 17F shows that circFGFR1p, encoded by circRNA, inhibits cell growth under stress conditions. Figure 17F shows the quantification of the relative circulation efficiency of circFGFR1 RNA under normal conditions or heat shock (HS) conditions by qRT-PCR, by normalizing circFGFR1 RNA levels to linear FGFR1 RNA levels using linear and circular RNA-specific primers, respectively. Ns: Not significant compared to normal conditions by unpaired two-sample t-test (n=3 independent replications). [Figure 17G] Figure 17G shows that circFGFR1p, encoded by circRNA, suppresses cell growth under stress conditions. Figure 17G shows the quantification of circFGFR1 RNA levels by qRT-PCR with and without heat shock (normalized to GAPDH mRNA levels). Ns: Not significant compared to normal conditions by unpaired two-sample t-test (n=3 independent replications). [Figure 17H] Figure 17H ​​shows that circFGFR1p, encoded by circRNA, suppresses cell growth under stress conditions. Figure 17H ​​provides a schematic diagram showing that, under normal conditions, the addition of FGF causes FGFR1 to dimerize and autophosphorylate, activating downstream cell signaling pathways and promoting cell proliferation. [Figure 17I] Figure 17I shows that circFGFR1p, encoded by circRNA, suppresses cell growth under stress conditions. Figure 17I provides a schematic diagram showing that under stress conditions, FGFR1 RNA translation is downregulated, resulting in lower FGFR1 protein levels and stable cap-independent translational activity of circFGFR1 IRES. Upon addition of FGF, circFGFR1p dimerizes with FGFR1. However, because circFGFR1p lacks an autophosphorylation domain, the circFGFR1p-FGFR1 dimer cannot activate downstream cellular signaling pathways, resulting in suppressed cell proliferation. [Figure 18A]Figure 18A shows the mean free energy (MFE) of various IRESs identified and / or tested in the screening described herein. Figure 18A shows the MFE for all viral IRES-positive oligos in DNA format and all human IRES-positive oligos in DNA format. [Figure 18B] Figure 18B shows the mean free energy (MFE) of various IRESs identified and / or tested in the screening described herein. Provided in Figure 18B are histograms showing eGFP expression levels driven by representative viral IRESs from DNA-based IRES screening. [Figure 19A] Figure 19A provides a schematic diagram of the fixed position of a secondary structural element as part of the IRES, where the secondary structural element extends approximately 40–60 bp from the +1 start site of the IRES sequence. The 18S complementary sequence may be positioned 5' (Figure 19A) relative to the secondary structural element. In this figure, the secondary structural element is a hairpin, but the secondary structural element may have one or more alternative structures as described herein. [Figure 19B] Figure 19B provides a schematic diagram of the fixed position of a secondary structural element as part of the IRES, where the secondary structural element extends approximately 40–60 bp from the +1 start site of the IRES sequence. The 18S complementary sequence may be positioned 3' (Figure 19B) relative to the secondary structural element. In this figure, the secondary structural element is a hairpin, but the secondary structural element may have one or more alternative structures as described herein. [Modes for carrying out the invention]

[0021] This disclosure is based, at least in part, on the development of a high-throughput reporter assay capable of systematically screening and quantifying the IRES activity of RNA sequences that can promote circRNA translation. This assay can identify elements in the primary and secondary structures of circRNA IRESs that are crucial for promoting circRNA translation. This assay can also identify potential endogenous protein-coding circRNAs and further expand the currently understood proteome. For example, this disclosure demonstrates the identification of circFGFR1p, a circRNA-encoded protein that functions as a negative regulator of FGFR1 through a dominant-negative mechanism that suppresses cell proliferation under stress conditions. Embodiments described herein provide strategies for recognizing and manipulating circRNA translation and revealing a new range of the endogenous circRNA proteome, thereby providing insights into circRNA-related diseases and the development of novel therapies targeting circRNA-encoded proteins.

[0022] This disclosure further builds upon the discovery that circFGFR1p is a protein encoded by endogenous circRNA, which is a negative regulator of FGFR1 signaling and suppresses cell proliferation under stress conditions. While cells reduce overall translation under stress conditions, many IRESs, including the circZNF-609 IRES, can drive higher cap-independent translational activity under stress conditions. Embodiments described herein shed light on the key regulatory mechanisms by which cells utilize different translational mechanisms to respond to stress conditions and describe how circRNAs can be used to maintain protein translation under such conditions. While cells primarily utilize cap-dependent linear mRNA translation to produce proteins, under stress conditions, the RNA source of translation can be shifted to circRNAs by upregulating the cap-independent translational activity of circRNA IRESs. In human cancers, circFGFR1 depletion can occur, downregulating circFGFR1p and increasing proliferation signals through FGF signaling. CircRNA-encoded proteins are useful for expressing individual subunits or "modules" of multidomain proteins, which can give cells the ability to independently regulate their translation. This disclosure provides a novel model of how circRNA translation is regulated by a mechanism different from linear mRNA translation, and how cells utilize circRNA-encoded proteins to respond to the dynamic environment. This disclosure also provides recombinant circular RNAs, including protein-coding nucleic acid sequences and IRESs operably ligated to protein-coding nucleic acid sequences, which can be used to express one or more target proteins in cells.

[0023] definition To facilitate understanding of this technology, several terms and phrases are defined below. Additional definitions will be provided throughout the detailed explanation.

[0024] In the context describing this invention (particularly in the context of the following claims), the use of the terms “a,” “an,” “the,” “at least one,” and similar references should be interpreted as encompassing both singular and plural unless otherwise indicated herein or clearly contradicts the context. The use of the term “at least one” and a subsequent list of one or more items (e.g., “at least one of A and B”) should be interpreted as meaning one item (A or B) selected from the enumerated items or any combination of two or more enumerated items (A and B), unless otherwise indicated herein or clearly contradicts the context. The enumeration of value ranges herein is intended solely as a simplified notation to refer individually to each separate value that falls within that range, unless otherwise indicated herein, and each separate value is incorporated herein as if it were individually enumerated herein. All methods described herein may be carried out in any order unless otherwise indicated herein or clearly contradicts the context. The use of any examples or illustrative language provided herein (e.g., "etc.") is intended solely to facilitate a better understanding of the invention and, unless otherwise claimed, does not limit the scope of the invention. The language in the specification should not be construed as indicating that any non-claimed element is essential to the practice of the invention.

[0025] The terms “nucleic acid sequence,” “polynucleotide,” and “oligonucleotide” are used interchangeably herein and refer to polymers or oligomers of cytosine, thymine, and uracil, and pyrimidine and / or purine bases such as adenine and guanine, respectively (Albert L. Lehninger, Principles of Biochemistry, pp. 793-800 (Worth Pub. 1982)). This term encompasses any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The polymers or oligomers may be heterogeneous or homogeneous in composition, may be isolated from naturally occurring sources, or may be artificially or synthetically produced. Furthermore, nucleic acids may be DNA or RNA, or mixtures thereof, and may exist permanently or transiently in single-stranded or double-stranded forms, including homo-double-stranded, hetero-double-stranded, and hybrid states. Nucleic acids or nucleic acid sequences may include, for example, DNA / RNA helices, peptide nucleic acids (PNAs), morpholino nucleic acids (see, e.g., Braasch and Corey, Biochemistry, 41(14): pp. 4503-4510 (2002) and U.S. Patent No. 5,034,506), locked nucleic acids (LNAs; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97: pp. 5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: pp. 8595-8602 (2000)) and / or other types of nucleic acid structures such as ribozymes. The terms “nucleic acid” and “nucleic acid sequence” may also include non-natural nucleotides, modified nucleotides, and / or non-nucleotide components (e.g., “nucleotide analogs”) that can exhibit the same function as natural nucleotides. In this specification, the term "DNA sequence" is used to refer to nucleic acids that contain a series of DNA bases.

[0026] The terms “polypeptide” and “protein” are used interchangeably herein and refer to polymeric forms of amino acids comprising at least two consecutive chemically or biochemically modified or derivatized amino acids, and polypeptides having a modified peptide backbone. The term “peptide” as used herein refers to a class of short polypeptides. A peptide may also be a polymer of amino acids (natural or unnatural) having a length of up to about 100 amino acids. For example, a peptide may have a length of about 1 to about 10, about 10 to about 25, about 25 to about 50, about 50 to about 75, or about 75 to about 100 amino acids. In some embodiments, the peptide has a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids.

[0027] The nomenclature for nucleotides, nucleic acids, nucleosides, and amino acids used herein is consistent with the International Union of Pure and Applied Chemistry (IUPAC) standards (see, for example, bioinformatics.org / sms / iupac.html).

[0028] When referring to nucleic acid sequences or protein sequences, the term "identity" is used to describe the similarity between two sequences. Sequence similarity or identity may be determined using standard techniques known in the art, including but not limited to the local sequence identity algorithm of Smith & Waterman, Adv. Appl. Math. 2, 482 (1981), the sequence identity alignment algorithm of Needleman & Wunsch, J Mol. Biol. 48, 443 (1970), the search technique for similarity methods of Pearson & Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988), the computerized implementation of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, WI), the Best Fit sequence program described by Develeux et al., Nucl. Acid Res. 12, pp. 387-395 (1984), or the scrutiny technique. Another algorithm is the BLAST algorithm, described by Altschul et al., J Mol. Biol. 215, pp. 403-410 (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, pp. 5873-5787 (1993). A particularly useful BLAST program is the WU-BLAST-2 program, obtained from Altschul et al., Methods in Enzymology, 266, pp. 460-480 (1996); blast.wustl / edu / blast / README.html. WU-BLAST-2 uses several search parameters, which are set to default values ​​as needed. The parameters are dynamic values, established by the program itself depending on the composition of a particular sequence and the composition of the particular database in which the target sequence is being searched, although the values ​​may be adjusted to increase sensitivity. Furthermore, an additional useful algorithm is the gapped BLAST, reported by Altschul et al., (1997) Nucleic Acids Res. 25, pp. 3389-3402.Unless otherwise specified, percent identity is determined herein using the algorithm available at the internet address:blast.ncbi.nlm.nih.gov / Blast.cgi.

[0029] The terms “internal ribosome entry site,” “internal ribosome entry sequence,” “IRES,” and “IRES sequence region” are used interchangeably herein and refer to the cis-element of viral or human intracellular RNA (e.g., messenger RNA (mRNA) and / or circRNA) that bypasses the canonical eukaryotic cap-dependent translation initiation step. The canonical cap-dependent mechanism used by the vast majority of eukaryotic mRNAs requires the 5' end of the mRNA to be translatable at the start codon. 7 G-cap, initiator Met-tRNA met It requires more than 10 initiation factor proteins, directional scanning, and GTP hydrolysis. IRES typically consists of long, highly structured 5'-UTRs that mediate translation initiation complex binding and catalyze the formation of functional ribosomes.

[0030] When referring to nucleic acid sequences, the terms “coding sequence,” “coding region,” “coding region,” and “CDS” may also be used interchangeably herein to refer to a portion of a DNA or RNA sequence that is translated or can be translated into a protein, for example. The terms “reading frame,” “open reading frame,” and “ORF” may also be used interchangeably herein to refer to a nucleotide sequence that begins with a start codon (e.g., ATG) and, in some embodiments, ends with a stop codon (e.g., TAA, TAG, or TGA). An open reading frame may contain introns and exons; therefore, all CDSs are ORFs, but not all ORFs are CDSs.

[0031] The terms "complementary" and "complementarity" refer to the relationship between two nucleic acid sequences or nucleic acid monomers that have the ability to form hydrogen bonds (or more) with each other through traditional Watson-Crick base pairing or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences can be indicated by the percentage of nucleotides in the nucleic acid sequence that can form hydrogen bonds (e.g., Watson-Crick base pairing) with the second nucleic acid sequence (e.g., approximately 50%, 60%, 70%, 80%, 90%, and approximately 100% complementary). Two nucleic acid sequences are "perfectly complementary" if every consecutive nucleotide in the nucleic acid sequence forms hydrogen bonds with the same number of consecutive nucleotides in the second nucleic acid sequence. Two nucleic acid sequences are "substantially complementary" if the degree of complementarity between the two nucleic acid sequences is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) over a region of at least 8 nucleotides (e.g., at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides), or if the two nucleic acid sequences hybridize under conditions of at least moderate, or in some embodiments, high strictness.Exemplary moderate-strict conditions include overnight incubation at 37°C in a solution containing 20% ​​formamide, 5×SSC (150mM NaCl, 15mM trisodium citrate), 50mM sodium phosphate (pH 7.6), 5×Denhart solution, 10% dextran sulfate, and 20 mg / ml denatured shear salmon sperm DNA, followed by washing the filter in 1×SSC at approximately 37–50°C, or substantially similar conditions, including the moderately strict conditions described in Sambrook, J., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press; 4th edition (June 15, 2012). High-strictness conditions include, for example, (1) using low ionic strength and high temperature for washing, such as 0.015M sodium chloride / 0.0015M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50°C, (2) using a denaturing agent such as formamide during hybridization, for example, 0.1% bovine serum albumin (BSA) / 0.1% Ficol / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer 50% (v / v) formamide at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42°C, or (3) using 50% formamide, 5×SSC(0.75M NaCl, 0.075M The following conditions were used for washing: (i) 42°C with 0.2×SSC, (ii) 55°C with 50% formamide, and (iii) 55°C with 0.1×SSC (combined with EDTA if necessary). Further details and explanations of the rigor of the hybridization reaction are provided, for example, in Sambrook, above; and Ausubel et al., eds., Short Protocols in Molecular Biology, 5th edition, John Wiley & Sons, Inc., Hoboken, NJ (2002).When referring to nucleic acid sequences, the term "hybridization" or "hybridized" refers to the association formed between two complementary sequences (between) and / or between three or more sequences (among).

[0032] As used herein with respect to nucleic acid sequences (e.g., RNA, DNA, etc.), the terms “secondary structure,” “secondary structure element,” or “secondary structure region” refer to any nonlinear three-dimensional structure of a nucleotide or ribonucleotide unit. Such nonlinear three-dimensional structures may include base pairing interactions within a single nucleic acid polymer or between two polymers. Single-stranded RNA typically forms complex and intricate base pairing interactions due to its increased ability to form hydrogen bonds resulting from extra hydroxyl groups in the ribose sugar. Examples of secondary structures or secondary structure elements include, but are not limited to, stem-loops, hairpin structures, bulges, internal loops, multiloops, coils, random coils, helices, partial helices, and pseudoknots. In some embodiments, the term “secondary structure” may refer to a SuRE element. The term “SuRE” stands for stem-loop structured RNA element (SuRE).

[0033] As used herein, the term "free energy" refers to the energy released by folding an unfolded polynucleotide (e.g., RNA or DNA) molecule, or conversely, the amount of energy that must be added to unfold a folded polynucleotide (e.g., RNA or DNA). The "minimum free energy (MFE)" of a polynucleotide (e.g., DNA, RNA, etc.) describes the lowest free energy observed for a polynucleotide when evaluating its various secondary structures. The MFE of an RNA molecule may be used to predict the RNA or DNA secondary structure, and this MFE is influenced by the number, composition, and arrangement of RNA or RNA nucleotides. The more negative the free energy of a structure, the higher the likelihood of its formation, as more stored energy is released upon its formation.

[0034] The term "melting temperature (Tm)" refers to the temperature at which approximately 50% of a double-stranded nucleic acid structure (e.g., DNA / DNA, DNA / RNA, or RNA / RNA double helix) denatures and separates into single-stranded structures.

[0035] As used herein, the term “recombinant” means that a particular nucleic acid (DNA or RNA) is transformed through various combinations of cloning, restriction, polymerase chain reaction (PCR), and / or ligation steps to produce a construct having a structurally coding or non-coding sequence that is distinguishable from endogenous nucleic acids found in the natural system. A polypeptide-encoding DNA sequence can be assembled from cDNA fragments or a series of synthetic oligonucleotides to provide synthetic nucleic acids that can be expressed in cells or from recombinant transcription units contained in cell-free transcription and translation systems. Genomic DNA containing the relevant sequence can also be used to form recombinant genes or transcription units. The non-coding DNA sequence may be located at the 5' or 3' end of the open reading frame, where it does not interfere with the manipulation or expression of the coding region but may act to regulate the production of desired products by various mechanisms. Alternatively, a DNA sequence encoding untranslated RNA may also be considered recombinant. Therefore, the term “recombinant” nucleic acid also refers to nucleic acids that do not exist in nature and are created, for example, through human intervention by artificial combinations of two other-separated segments of a sequence. This artificial combination is often achieved by chemical synthesis or by artificial manipulation of isolated nucleic acid segments, such as genetic engineering techniques. Such techniques are typically performed to replace codons with codons encoding the same amino acid, a conserved amino acid, or a non-conserved amino acid. Alternatively, artificial combinations may be performed to combine nucleic acid segments of desired function to produce a desired combination of function. This artificial combination is often achieved by chemical synthesis or by artificial manipulation of isolated nucleic acid segments, such as genetic engineering techniques. When recombinant polynucleotides encode polypeptides, the sequence of the encoded polypeptide can be naturally occurring ("wild-type") or a variant of a naturally occurring sequence (e.g., a mutant). Therefore, the term "recombinant" polypeptide does not necessarily refer to a polypeptide whose sequence does not exist in nature.Instead, "recombinant" polypeptides are encoded by recombinant DNA sequences, but the polypeptide sequence can be naturally occurring ("wild-type") or not naturally occurring (e.g., variants, mutants, etc.). Thus, "recombinant" polypeptides are the result of human intervention, but may contain naturally occurring amino acid sequences.

[0036] As used herein, the terms “operatably linked” and “operatably linked” refer to an arrangement of elements configured to function or structure in a manner suitable for the intended purpose. For example, a given promoter operatably linked to a coding sequence can result in the expression of the coding sequence if the appropriate enzyme is present. Expression is intended to include the transcription of one or more recombinant nucleic acids that encode circular RNA or mRNA derived from DNA or an RNA template, and may further include the translation of proteins from recombinant circular RNA containing IRES sequences (e.g., non-natural IRESs). Thus, for example, an intervening, untranslated but transcribed sequence may exist between the promoter sequence and the coding sequence, and the promoter sequence can still be considered “operatably linked” to the coding sequence.

[0037] As used herein, the term “non-viral-like particles” may refer to any protein-based particles that are neither viruses nor virus-like particles. For example, in some embodiments, non-viral-like particles are protein nanogels or protein spheres that enable encapsulation.

[0038] As used herein with respect to lipid nanoparticles, the term “decorated” means lipid nanoparticles bound to one or more targeting agents (e.g., small molecules, peptides, polypeptides, carbohydrates, etc.). The targeting agents bind to one or more peptides, polypeptides, carbohydrates, cells, etc., enabling the lipid nanoparticles to be specifically directed to them.

[0039] Circular RNA Circular RNA (circRNA) is a single-stranded RNA molecule in which the head is attached to the tail. It was first discovered in pathogenic genomes such as those of hepatitis D virus (HDV) and plant viloids (Kos et al., Nature, 323: pp. 558-560 (1986); Sanger et al., PNAS USA, 73: pp. 3852-3856 (1976)). CircRNA has been recognized as a broad class of non-coding RNA in eukaryotic cells (Salzman et al., PLoS One, 7: e30733, (2012); Memczak et al., Nature, 495: pp. 333-338 (2013); Hansen et al., Nature, 495: pp. 384-388 (2013)). CircRNAs are typically generated through backsplicing and, due to their exceptional stability, have been hypothesized to function in intercellular communication or memory (Jeck, WR & Sharpless, NE, Nat Biotech, 32: pp. 453-461 (2014)).

[0040] Although the function of endogenous circRNAs is not fully understood, recent discoveries regarding the regulation of human circRNAs in viral resistance through NF90 / NF110 control and autoimmunity through PKR control demonstrate their large number and the necessity of a circRNA immune system in the presence of viral circRNA genomes. Circular RNAs can function as potent adjuvants to induce specific T and B cell responses. Furthermore, circRNAs may have the ability to induce both autoimmune and adaptive immune responses, inhibiting tumor establishment and growth.

[0041] The inventors have previously shown that intron identity directs circRNA immunity. See, for example, Chen, Y.G., et al., Mol. Cell (2019) 76(1): pp. 96-109.e9; Chen, Y.G., et al., Mol. Cell (2017) 67(2): pp. 228-238.e5. Since introns are not part of the final circRNA product, it has been hypothesized that introns may direct the deposition of one or more covalent chemical marks on circRNA. Among the more than 100 known RNA chemical modifiers, m 6 A is the most abundant modifier on linear mRNA and long non-coding RNA, and is present in 0.2%–0.6% of all adenosine in mammalian poly(A) tail transcripts (Roundtree et al., Cell, 169: pp. 1187–1200 (2017)). 6 A has recently been detected on mammalian circRNAs (Zhou et al., Cell Reports, 20: pp. 2262-2276 (2017)). Human endogenous circRNAs have one or more covalently bonded m based on the introns that program their backsplicing. 6 It is thought that the characteristics at birth are determined by the A modifier.

[0042] This disclosure provides recombinant circular RNA molecules comprising a protein-coding nucleic acid sequence and a non-native internal ribosome entry site (IRES) operably ligated to the protein-coding nucleic acid sequence, as well as the DNA sequence encoding them. Recombinant circRNA molecules may be produced or manipulated according to several methods. For example, recombinant circRNA molecules may be produced by back-splicing linear RNA. For example, in some embodiments, recombinant circular RNA is produced by back-splicing a downstream 5' splice site (splice donor) to an upstream 3' splice site (splice acceptor). The splice donor and / or splice acceptor is typically found, for example, in a human intron or a portion thereof used for circRNA production at an endogenous locus shown in Figure 1A. In some embodiments, recombinant circular RNA is produced by contacting a cell with a DNA plasmid, the DNA plasmid encoding linear RNA, and the linear RNA is back-spliced ​​to produce recombinant circular RNA. In some embodiments, the DNA plasmid contains an intron derived from the mammalian ZKSCAN1 gene.

[0043] Circular RNA can be produced by any non-mammalian splicing method. For example, linear RNA containing various types of introns, including self-splicing group I introns, self-splicing group II introns, spliceosome introns, and tRNA introns, can be circularized. In particular, group I and group II introns can undergo self-splicing due to their autocatalytic ribozyme activity, which has the advantage of being easily used to produce circular RNA in vitro and in vivo.

[0044] Alternatively, circular RNA can be prepared in vitro from linear RNA by chemical or enzymatic ligation of the 5' and 3' ends of the RNA. Chemical ligation can be carried out, for example, using cyanogen bromide (BrCN) or ethyl-3-(3'-dimethylaminopropyl)carbodiimide (EDC) to activate the nucleotide phosphate monoester group that enables phosphate diester bond formation (Sokolova, FEBS Lett, 232:153-155 (1988); Dolinnaya et al., Nucleic Acids Res., 19:3067-3072 (1991); Fedorova, Nucleosides Nucleotides Nucleic Acids, 15:1137-1147 (1996)). Alternatively, RNA can be circularized using enzymatic ligation. Exemplary ligases available include T4 DNA ligase (T4 Dnl), T4 RNA ligase 1 (T4 Rnl 1), and T4 RNA ligase 2 (T4 Rnl 2).

[0045] In other embodiments, circular RNA may be generated using sprint ligation. Sprint ligation involves joining the ends of a linear RNA for ligation using an oligonucleotide sprint that hybridizes with the two ends of the linear RNA. Hybridization of the sprint, which can be a deoxyribooligonucleotide or a ribooligonucleotide, directs the 5'-phosphate and 3'-OH of the RNA ends for ligation. Subsequent ligation can be carried out using chemical or enzymatic techniques as described above. Enzymatic ligation can be carried out, for example, using T4 DNA ligase (requiring DNA sprint), T4 RNA ligase 1 (requiring RNA sprint), or T4 RNA ligase 2 (requiring DNA or RNA sprint). Chemical ligation, such as using BrCN or EDC, is in some cases more efficient than enzymatic ligation if the structure of the hybridized sprint-RNA complex interferes with enzymatic activity (see, for example, Dolinnaya et al., Nucleic Acids Res, 21(23):5403-5407 (1993); Petkovic et al., Nucleic Acids Res, 43(4):2454-2465 (2015)).

[0046] Circular RNAs are generally more stable than their linear counterparts because they lack free ends, which are primarily required for exonuclease-mediated degradation. However, the stability of recombinant circRNAs described herein may be further improved by additional modifications. Further modifications of other types may improve circulation efficiency, circRNA purification, and / or protein expression from circRNAs. For example, recombinant circRNAs may be manipulated to include "homology arms" (i.e., 9-19 nucleotides in length placed at the 5' and 3' ends of the precursor RNA to bring the 5' and 3' splice sites closer together), spacer sequences, and / or phosphorothioate (PS) caps (Wesselhoeft et al., Nat. Commun., 9:2629 (2018)). Recombinant circRNAs may be modified to include a 2'-O-methyl-,-fluoro- or -O-methoxyethyl conjugate, phosphorothioate skeleton, or 2',4'-cyclic 2'-O-ethyl variant to enhance their stability (Holdt et al., Front Physiol., 9:1262 (2018); Krutzfeldt et al., Nature, 438(7068):685~9 (2005); and Crooke et al., Cell Metab., 27(4):714~739 (2018)). Recombinant circRNA molecules may contain one or more variants that reduce the innate immunogenicity of the circRNA molecule in the host, for example, at least one N6-methyladenosine (m 6 A) may be included.

[0047] In some embodiments, the recombinant circular RNA molecule is encoded by a nucleic acid containing at least two introns and at least one exon. In some embodiments, the DNA sequence encoding the circular RNA molecule includes a sequence encoding at least two introns and at least one exon. The term “exon,” as used herein, refers to a nucleic acid sequence present in a gene represented in the mature form of RNA after the removal of an intron during transcription. Exons can be translated into proteins (for example, in the case of messenger RNA (mRNA)). The term “intron,” as used herein, refers to a nucleic acid sequence present in a given gene that is removed by RNA splicing during the maturation of the final RNA product. Introns are generally found between exons. During transcription, introns are removed from precursor messenger RNA (pre-mRNA), and exons are joined by RNA splicing. In some embodiments, the recombinant circular RNA molecule includes a nucleic acid sequence containing one or more exons and one or more introns.

[0048] Therefore, circular RNA can be produced by splicing endogenous or exogenous introns, as described in WO2017 / 222911. As used herein, the term “endogenous intron” means an intron sequence specific to the host cell from which the circRNA is produced. For example, human introns are endogenous introns when circRNA is expressed in human cells. “Exogenous intron” means an intron that is heterogeneous to the host cell from which the circRNA is produced. For example, bacterial introns are considered exogenous introns when circRNA is expressed in human cells. A wide variety of intron sequences are known from organisms and viruses, including sequences derived from genes encoding proteins, ribosomal RNA (rRNA), or transfer RNA (tRNA). Representative intron sequences are available in various databases, including the Group I Intron Sequence and Structure Database (rna.whu.edu.cn / gissd / ), the Bacterial Group II Intron Database (webapps2.ucalgary.ca / ~groupii / index.html), the Mobile Group II Intron Database (fp.ucalgary.ca / group2introns), the Yeast Intron Database (emblS16 heidelberg.de / ExternalInfo / seraphin / yidb.html), the Ares Lab Yeast Intron Database (compbio.soe.ucsc.edu / yeast_introns.html), the U12 Intron Database (genome.crg.es / cgibin / u12db / u12db.cgi), and the Exon-Intron Database (bpg.utoledo.edu / ~afedorov / lab / eid.html).

[0049] In some embodiments, the DNA molecule encoding the recombinant circular RNA molecule contains self-splicing group I introns. Group I introns are a distinct class of RNA self-splicing introns that catalyze their own deletion from mRNA, tRNA, and rRNA precursors in a wide range of organisms. All known group I introns present in the eukaryotic cell nucleus disrupt functional ribosomal RNA genes located at ribosomal DNA loci. Cellular group I introns are widely present among eukaryotic microorganisms, and myxomycetes (slime molds) are rich in self-splicing introns. Self-splicing group I introns contained in circular RNA molecules may be present in or originate from any organism, such as bacteria, bacteriophages, and eukaryotic viruses. Self-splicing group I introns may be found in certain organelles, such as mitochondria and chloroplasts, and such self-splicing introns may be incorporated into circular RNA molecules.

[0050] In some embodiments, recombinant circular RNA molecules are generated from a DNA molecule containing the self-splicing group I intron of the phage T4 thymidylate synthase (td) gene. It has been well-documented that the group I intron of the phage T4 thymidylate synthase (td) gene is circularized and that the exons linearly splice each other (Chandry and Belfort, Genes Dev., 1:1028-1037 (1987); Ford and Ares, Proc. Natl. Acad. Sci. USA, 91:3117-3121 (1994); and Perriman and Ares, RNA, 4:1047-1054 (1998)). If the order of the td intron is altered while remaining adjacent to any exon sequence (i.e., the 5' half placed at the 3' position, and vice versa), the exon is cyclized by two autocatalytic esterification reactions (Ford and Ares, above; Puttaraju and Been, Nucleic Acids Symp. Ser., 33:49~51 (1995)).

[0051] In some embodiments, the recombinant circular RNA molecule is encoded by a DNA molecule containing a ZKSCAN1 intron. The ZKSCAN1 intron is described, for example, by Yao, Z. et al., Mol. Oncol. (2017) 11(4): pp. 422-437. In some embodiments, the recombinant circular RNA molecule is encoded by a DNA molecule containing a mini-ZKSCAN1 intron.

[0052] The recombinant circular RNA molecule may be of any length or size. For example, the recombinant circular RNA molecule may contain approximately 200 to 10,000 nucleotides (e.g., approximately 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, or 9,000 nucleotides, or any two of the aforementioned values). In some embodiments, the recombinant circular RNA molecule has approximately 500 to 6,000 nucleotides (approximately 550, 650, 750, 850, 950, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, The recombinant circular RNA molecule contains approximately 2,700, 2,800, 2,900, 3,100, 3,300, 3,500, 3,700, 3,800, 3,900, 4,100, 4,300, 4,500, 4,700, 4,900, 5,100, 5,300, 5,500, 5,700, or 5,900 nucleotides, or a range limited by any two of the aforementioned values. In one embodiment, the recombinant circular RNA molecule contains approximately 1,500 nucleotides.

[0053] In some embodiments, the recombinant circular RNA molecule comprises a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence region, the IRES comprising at least one sequence region having secondary structural elements and a sequence region complementary to 18S ribosomal RNA (rRNA), the IRES sequence region having a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C. In some embodiments, the IRES sequence is linked to the protein-coding nucleic acid sequence region in a non-natural configuration.

[0054] This disclosure also provides recombinant circular RNA molecules comprising a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence, wherein the IRES is encoded by one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a nucleic acid sequence having at least 90% or at least 95% identity or homology thereto. In some embodiments, the IRES sequence is linked to the protein-coding nucleic acid sequence region in a non-natural configuration.

[0055] CircRNA internal ribosome entry site The recombinant circular RNAs described herein contain an internal ribosome entry site (IRES) operably ligated to the protein-coding sequence of the circRNA in a non-native configuration. Inclusion of the IRES enables translation of one or more open reading frames from the circular RNA. The IRES element attracts the eukaryotic ribosome translation initiation complex, thereby promoting translation initiation. Two known mechanisms for translation initiation in eukaryotes are recognized. The first is the canonical cap-dependent mechanism used by the majority of eukaryotic mRNAs, which requires the 5' end of the mRNA to be ligated to place the translatable ribosome at the start codon. 7This requires a G cap, the initiator Met-tRNAmet, a dozen or so initiation factor proteins, directional scanning, and GTP hydrolysis. The second mechanism is cap-independent initiation, used by some mRNAs and many eukaryotic infective viruses. This mechanism avoids the need for a cap and, often, many protein factors, and uses a cis-acting IRES RNA element to recruit ribosomes and initiate protein synthesis. There is great diversity among viral IRES RNAs in terms of their sequence, proposed secondary structure, and functional requirements for protein factors, but all drive a translation initiation mechanism that depends on a specific RNA sequence and possibly a specific RNA structure within the IRES.

[0056] Accordingly, various IRES sequences that can drive protein translation when present in circRNA are provided herein. In some embodiments, the circRNA IRES can be operably ligated to a protein-coding nucleic acid sequence. In some embodiments, the circRNA IRES is operably ligated to a protein-coding nucleic acid sequence in a non-native configuration. In some embodiments, the IRES is a human IRES. In some embodiments, the IRES is a viral IRES.

[0057] As used herein, the term “non-native configuration” refers to a binding between an IRES and a protein-coding nucleic acid that is not present in naturally occurring circRNA molecules. For example, a viral IRES may be operably ligated to a protein-coding nucleic acid sequence in circular RNA, or an IRES not found in naturally occurring circRNA molecules may be operably ligated to a protein-coding nucleic acid sequence in circRNA. In some embodiments, an IRES found in a naturally occurring circRNA molecule operably ligated to a particular protein-coding nucleic acid may be operably ligated to a different protein-coding nucleic acid (i.e., a nucleic acid to which the IRES is not operably ligated in any naturally occurring circRNA). In some embodiments, an IRES found in naturally occurring linear mRNA may be operably ligated to a protein-coding sequence in circular RNA.

[0058] Numerous linear IRES sequences are known and may be included in the recombinant circular RNA molecules described herein. For example, linear IRES sequences may be derived from a wide variety of viruses, such as picornavirus leader sequences (e.g., encephalomyocarditis virus (EMCV) UTR) (Jang et al., J. Virol., 63:1651-1660 (1989)), polio leader sequences, hepatitis A virus leaders, hepatitis C virus IRESs, human rhinovirus type 2 IRESs (Dobrikova et al., Proc. Natl. Acad. Sci., 100(25):15125-15130 (2003)), foot-and-mouth disease virus IRES elements (Ramesh et al., Nucl. Acid Res., 24:2697-2700 (1996)), and giardiavirus IRESs (Garlapati et al., J. Biol. Chem., 279(5):3389-3397 (2004)). Yeast-derived IRES sequences, type 1 human angiotensin II receptor IRES (Martin et al., Mol. Cell Endocrinol., 212: pp. 51-61 (2003)), fibroblast growth factor IRES (e.g., FGF-1 IRES and FGF-2 IRES, Martineau et al., Mol. Cell. Biol., 24(17): pp. 7622-7635 (2004)), vascular endothelial growth factor IRES (Baranick et al., Proc. Natl. Acad. Sci. USA, 105(12): pp. 4733-4738 (2008); Stein et al., Mol. Cell. Biol., 18(6): pp. 3112-3119 (1998); Bert et al., RNA, 12(6): pp. 1074-1083 (2006)), and insulin-like growth factor 2 Various non-viral IRES sequences, including but not limited to IRESs (Pedersen et al., Biochem. J., 363(Pt 1):37-44 (2002)), can also be included in circular RNA molecules.

[0059] IRES sequences and vectors encoding IRES elements are commercially available from various sources, such as Clontech (Mountain View, CA), Invivogen (San Diego, CA), Addgene (Cambridge, MA), GeneCopoeia (Rockville, MD), and IRESite (a database of experimentally validated IRES structures (iresite.org)). Notably, these databases focus on the activity of IRES sequences in mRNA (i.e., linear RNA) and do not emphasize circRNA IRES activity profiles.

[0060] In some embodiments, an IRES contains at least one RNA secondary structure element. Intramolecular RNA base pairing is often the basis of RNA secondary structure and, in some situations, a determinant of critical significance for overall macromolecular folding. Together with cofactors and RNA-binding proteins (RBPs), secondary structure elements can form higher-order tertiary structures, thereby conferring catalytic, regulatory, and scaffold functions to RNA. Therefore, an IRES may contain any RNA secondary structure element that provides such structural or functional determinants. In some embodiments, the RNA secondary structure may be formed from nucleotides approximately 40 to 60 of the IRES relative to the 5' end of the IRES. The most common RNA secondary structures are helices, loops, bulges, and junctions, with stem-loops or hairpin loops being the most common elements of RNA secondary structure. Stem-loops are formed when an RNA strand folds back into itself to form a double helix bundle called a stem, and unpaired nucleotides form a single-stranded region called a loop. The bulge and inner loop are formed by the separation of the double helix on one strand (bulge) or both strands (inner loop) by unpaired nucleotides. A tetraloop is a 4-base pair hairpin RNA structure. Ribosomal RNA has three common families of tetraloops: UNCG, GNRA, and CUUG (where N is one of the four nucleotides and R is a purine). A pseudoknot is formed when nucleotides from a hairpin loop pair pair with a single-stranded region outside the hairpin to form a helical segment. RNA secondary structures are further described, for example, in Vandivier et al., Annu Rev Plant Biol., 67:463-488 (2016); and Tinoco and Bustamante, above). In some embodiments, the IRES of a recombinant circRNA molecule includes at least one stem-loop structure. At least one RNA secondary structure element may be at any position in the IRES, as long as translation is efficiently initiated from the IRES. In some embodiments, the stem portion of the stem loop may contain 3 to 7 base pairs, 4, 5, 6, 7, 8, 9, 10, 11, or 12 base pairs or more.The loop portion of the stem-loop may contain 3 to 12 nucleotides, including 4, 5, 6, 7, 8, 9, 10, 11, 12 or more nucleotides. The stem-loop structure may have one or more bulges (mismatches) on either side of the stem. In some embodiments, the RNA secondary structure element is formed from nucleotides at approximately positions 40 to 60 of the IRES, where the first nucleic acid at the 5' end of the IRES is considered to be position 1. In some embodiments, the sequence complementary to 18S rRNA is located on the 5' side of at least one RNA secondary structure element (i.e., in the range of approximately positions 1 to 40 of the IRES, see Figure 19A). In some embodiments, the sequence complementary to 18S rRNA is located on the 3' side of at least one RNA secondary structure element (i.e., in the range of approximately position 61 to the terminal of the IRES, see Figure 19B). Sequences encoding exemplary secondary structure-forming RNA sequences that may be included in the IRESs described herein are provided in SEQ ID NOs: 17202-28976.

[0061] In some embodiments, at least one RNA secondary structure element of the IRES is a stem-loop. In some embodiments, at least one RNA secondary structure element is encoded by one of the nucleic acid sequences of SEQ ID NOs: 17202-28976. In some embodiments, at least one RNA secondary structure element is encoded by a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NOs: 17202-28976. In some embodiments, at least one RNA secondary structure element is encoded by a nucleic acid sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotide substitutions with one of SEQ ID NOs: 17202-28976.

[0062] RNA secondary structure can typically be predicted from experimental thermodynamic data in addition to chemical mapping, nuclear magnetic resonance (NMR) spectroscopy, and / or sequence comparison. In some embodiments, RNA secondary structure is predicted by machine learning / deep learning algorithms (e.g., CNN) (Zhao, Q. et al., "Review of Machine-Learning Methods for RNA Secondary Structure Prediction," Sept 1, 2020 (available on the world wide web at:arxiv.org / abs / 2009.08868)). Various algorithms and software packages for RNA secondary structure prediction and analysis are known in the art and can be used in the context of this disclosure (e.g., Hofacker IL (2014) Energy-Directed RNA Structure Prediction. In: Gorodkin J., Ruzzo W. (eds.) RNA Sequence, Structure, and Function: Computational and Bioinformatic Methods. Methods in Molecular Biology (Methods and Protocols), Vol. 1097. Humana Press, Totowa, NJ; Mathews et al., above; Mathews et al., "RNA secondary structure prediction," Current Protocols in Nucleic Acid Chemistry, Chapter 11 (2007): Unit 11.2. doi:10.1002 / 0471142700.nc1102s28; Lorenz et al., Methods, 103: pp. 86-98 (2016); Mathews et al., Cold Spring Harb Perspect Biol., 2(12): a003665 (2010) (see references).

[0063] In some embodiments, the recombinant circRNA IRES may contain nucleic acid sequences complementary to 18S ribosomal RNA (rRNA). Eukaryotic ribosomes, also known as "80S," have two non-identical subunits named according to their sedimentation coefficient: a small subunit (40S) (also called "SSU") and a large subunit (60S) (also called "LSU"). Both subunits contain dozens of ribosomal proteins arranged on a scaffold composed of ribosomal RNA (rRNA). In eukaryotes, eukaryotic 80S ribosomes contain nucleotides larger than 5500 of rRNA: 18S rRNA in the small subunit and 5S, 5.8S, and 25S rRNA in the large subunit. The small subunit monitors complementarity between tRNA anticodons and mRNA, while the large subunit catalyzes peptide bond formation. Ribosomes typically contain about 60% rRNA and about 40% protein. The primary structure of rRNA sequences can vary among organisms, but the base pairings within these sequences generally form a stem-loop configuration.

[0064] In some embodiments, the recombinant circRNA IRES may contain any nucleic acid sequence complementary to any eukaryotic 18S rRNA sequence. In some embodiments, the nucleic acid sequence complementary to 18S rRNA is encoded by one of the nucleic acid sequences shown in Table 1. In some embodiments, the nucleic acid sequence complementary to 18S rRNA is encoded by a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity or homology to the sequences shown in Table 1. In some embodiments, the nucleic acid sequence complementary to 18S rRNA is encoded by a nucleic acid sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotide substitutions compared to the nucleic acid sequences shown in Table 1.

[0065] [Table 1]

[0066] The most commonly used criterion for predicting RNA secondary structure is minimum free energy (MFE). This is because, according to thermodynamics, the MFE structure is not only the most stable but also the most likely structure in thermodynamic equilibrium. The MFE of an RNA or DNA molecule is influenced by three properties of nucleotides in the RNA / DNA sequence: number, composition, and arrangement. For example, longer sequences are generally more stable because they can form more stacking and hydrogen bonding interactions; guanine-cytosine (GC)-rich RNAs are typically more stable than adenine-uracil (AU)-rich sequences; and the order of nucleotides affects folding stability because it determines the number and elongation of loops and double helix structures. Unlike other non-coding RNAs, mRNA and microRNA precursors have been found to have a larger negative MFE than would be expected considering their nucleotide number and composition. Therefore, free energy can also be used as a criterion for identifying functional RNA.

[0067] The IRES of a recombinant circRNA molecule may contain a minimum free energy (MFE) of less than approximately -15 kJ / mol (e.g., less than approximately -16 kJ / mol, less than approximately -17 kJ / mol, less than approximately -18.5 kJ / mol, less than approximately -19 kJ / mol, less than approximately -18.9 kJ / mol, less than approximately -20 kJ / mol, less than approximately -30 kJ / mol). In some embodiments, the MFE is greater than approximately -90 kJ / mol (e.g., greater than approximately -85 kJ / mol, greater than approximately -80 kJ / mol, greater than approximately -70 kJ / mol, greater than approximately -60 kJ / mol, greater than approximately -50 kJ / mol, greater than approximately -40 kJ / mol). In some embodiments, the IRES has a minimum free energy (MFE) of approximately -18.9 kJ / mol or less. In some embodiments, the IRES has a MFE in the range of about -15.9 kJ / mol to about -79.9 kJ / mol. In some embodiments, the IRES may contain an MFE in the range of about -12.55 kJ / mol to about -100.15 kJ / mol. In some embodiments, the IRES is a viral IRES and has an MFE in the range of about -15.9 kJ / mol to about -79.9 kJ / mol. In some embodiments, the IRES is a human IRES and has an MFE in the range of about -12.55 kJ / mol to about -100.15 kJ / mol.

[0068] In some embodiments, at least one secondary structural element of the IRES may have a minimum free energy (MFE) of less than about -0.4 kJ / mol, less than about -0.5 kJ / mol, less than about -0.6 kJ / mol, less than about -0.7 kJ / mol, less than about -0.8 kJ / mol, less than about -0.9 kJ / mol, or less than about -1.0 kJ / mol. In some embodiments, at least one secondary structural element of the IRES may have an MFE of less than about -0.7 kJ / mol.

[0069] In some embodiments, RNA sequences containing nucleotides at approximately positions 40 to 60 of the IRES of the circRNA described herein may have a minimum free energy (MFE) of less than approximately -0.4 kJ / mol, less than approximately -0.5 kJ / mol, less than approximately -0.6 kJ / mol, less than approximately -0.7 kJ / mol, less than approximately -0.8 kJ / mol, less than approximately -0.9 kJ / mol, or less than approximately -1.0 kJ / mol. In some embodiments, RNA sequences containing nucleotides at approximately positions 40 to 60 of the IRES may have an MFE of less than approximately -0.7 kJ / mol.

[0070] As discussed above, the minimum free energy of a particular RNA (e.g., RNA constructed from a DNA sequence) may be determined using various computer calculation methods and algorithms. The most commonly used software programs for predicting secondary RNA or DNA structure using the MFE algorithm utilize the so-called nearest neighbor energy model. This model uses free energy rules based on empirical thermodynamic parameters (Mathews et al., J Mol Biol, 288:911-940 (1999); and Mathews et al., Proc Natl Acad Sci USA, 101:7287-7292 (2004)) and calculates the overall stability of the RNA or DNA structure by summing the independent contributions of local free energy interactions due to adjacent base pairs and loop regions. For sequences with homogeneous nucleotide arrangement and composition, the additional and independent characteristics of local free energy contributions suggest a linear relationship between the calculated MFE and sequence length (Trotta, E., PLoS One, 9(11):e113380 (2014)). Algorithms for determining MFE are further described, for example, in Hajiaghayi et al., BMC Bioinformatics, 13:22 (2012); Mathews, DH, Bioinformatics, Vol. 21, No. 10: pp. 2246-2253 (2005); and Doshi et al., BMC Bioinformatics, 5:105 (2004) doi 10.1186 / 1471-2105-5-105).

[0071] A person skilled in the art would recognize that the melting temperature (T m ) of a specific circRNA molecule can also indicate stability. In fact, RNA sequences with high T m generally contain thermostable and functionally important RNA structures (see, for example, Nucleic Acids Res., 45(10):pp. 6109-6118 (2017)). Accordingly, in some embodiments, the IRES of the recombinant circRNA molecule has a melting temperature of at least 35.0°C. In some embodiments, the IRES of the recombinant circRNA molecule has a melting temperature of at least 35.0°C but not more than 85°C. In some embodiments, the RNA secondary structure has a melting temperature of at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 43°C, at least 44°C, at least 45°C, at least 46°C, at least 47°C, at least 48°C, at least 49°C, or higher. In some embodiments, the melting temperature is about 85°C or lower, about 75°C or lower, about 70°C or lower, about 65°C or lower, about 60°C or lower, about 55°C or lower, about 50°C or lower, or lower.

[0072] The melting temperature of a specific nucleic acid molecule can be determined using thermodynamic analyses and algorithms described herein and known in the art (see, for example, Kibbe W.A., Nucleic Acids Res., 35(Web Server issue):W43-W46(2007). doi: 10.1093 / nar / gkm234; and Dumousseau et al., BMC Bioinformatics, 13:p. 101(2012). doi.org / 10.1186 / 1471-2105-13-101).

[0073] In some embodiments, the IRES comprises at least one RNA secondary structure element and a nucleic acid sequence complementary to 18S ribosomal RNA (rRNA), and the IRES has a minimum free energy (MFE) of -18.9 kJ / mol or lower and a melting temperature of at least 35°C. In some embodiments, the RNA secondary structure element of the IRES has a minimum free energy (MFE) of less than -18.9 kJ / mol and is formed from nucleotides at positions approximately 40 to 60 of the IRES, with the first nucleic acid at the 5' end of the IRES being considered as position 1. In some embodiments, the RNA secondary structure element has a melting temperature of at least 35.0°C and is formed from nucleotides at positions approximately 40 to 60 of the IRES, with the first nucleic acid at the 5' end of the IRES being considered as position 1.

[0074] Since circRNA molecules are often generated from linear RNA by backsplicing of a downstream 5' splice site (splice donor) to an upstream 3' splice site (splice acceptor), recombinant circular RNA molecules may further contain backsplice junctions. In some embodiments, the IRES may be located within approximately 100 to 200 nucleotides of the backsplice junction. Furthermore, it has been observed that regions of RNA with higher GC content have a more stable secondary structure than RNA strands with lower GC content. Therefore, in some embodiments, recombinant circRNA molecules may further contain minimal levels of GC base pairs. For example, the non-natural IRES of a recombinant circRNA molecule may contain at least 25% (e.g., at least 30%, at least 35%, at least 40%, at least 45%, or more) but less than approximately 75% (e.g., approximately 70%, approximately 65%, approximately 60%, approximately 55%, approximately 50%, or less) GC content. In some embodiments, the IRES has at least 25% GC content.

[0075] The GC content of a given nucleic acid sequence may be measured using any method known in the art, such as chemical mapping (see, for example, Cheng et al., PNAS, 114(37):9876-9881 (2017); and Tian, ​​S. and Das, R., Quarterly Reviews of Biophysics, 49:e7 doi:10.1017 / S0033583516000020 (2016)).

[0076] Exemplary sequences encoding IRESs for use in the circRNA molecules of this disclosure are shown in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201. Accordingly, this disclosure further provides recombinant circular RNA molecules comprising a protein-coding nucleic acid sequence and an IRES operably ligated to the protein-coding nucleic acid sequence in a non-natural configuration, wherein the IRES is encoded by one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201.

[0077] In some embodiments, the IRES is encoded by one of the nucleic acid sequences shown in any one of sequence numbers 1 to 228. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with any one of sequence numbers 1 to 228. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotide substitutions compared to any one of sequence numbers 1 to 228.

[0078] In some embodiments, the IRES is encoded by one of the nucleic acid sequences shown in any one of sequence numbers 229 to 17201. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity or homology to any one of sequence numbers 229 to 17201. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotide substitutions compared to any one of sequence numbers 229 to 17201.

[0079] In some embodiments, the IRES is encoded by a nucleic acid sequence represented by index 876 (sequence number 531), 6063 (sequence number 2270), 7005 (sequence number 2602), 8228 (sequence number 3042), or 8778 (sequence number 3244). In some embodiments, the IRES is encoded by the nucleic acid sequence of sequence number 33948.

[0080] In some embodiments, the IRES is encoded by one of the nucleic acid sequences shown in Table 2. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity or homology to one of the nucleic acid sequences in Table 2. In some embodiments, the IRES is encoded by a nucleic acid sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotide substitutions compared to one of the sequences in Table 2.

[0081] [Table 2] TIFF2026143428000004.tif206170 TIFF2026143428000005.tif211170 TIFF2026143428000006.tif208170 TIFF2026143428000007.tif207170 TIFF2026143428000008.tif207170 TIFF2026143428000009.tif207170 TIFF2026143428000010.tif147170

[0082] IRES may be of any length or size. For example, IRES may be approximately 100 to 600 nucleotides in length (e.g., approximately 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, or 575 nucleotides, or limited by any two of the aforementioned values). In some embodiments, IRES may be approximately 200 to approximately 800 nucleotides in length (approximately 200, approximately 210, approximately 220, approximately 240, approximately 260, approximately 280, approximately 320, approximately 340, approximately 360, approximately 380, approximately 420, approximately 440, approximately 460, approximately 480, approximately 500, approximately 520, approximately 540, approximately 560, approximately 580, approximately 600, approximately 620, approximately 640, approximately 660, approximately 680, approximately 700, approximately 720, approximately 740, approximately 760, approximately 780, or approximately 800 nucleotides in length, or limited to any two of the aforementioned values). In some embodiments, IRES may be approximately 200 to approximately 400, approximately 400 to approximately 600, approximately 600 to approximately 700, or approximately 600 to approximately 800 nucleotides in length. In some embodiments, the IRES is approximately 210 nucleotides long. In some embodiments, the IRES may be approximately 100 to approximately 3000 nucleotides long.

[0083] In some embodiments, the circular RNA molecule includes an IRES consisting of a sequence encoded by one of the DNA sequences SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201. In some embodiments, the circular RNA molecule includes an IRES sequence encoded by one of the DNA sequences SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, and the IRES sequence further includes up to 1000 additional nucleotides. In some embodiments, the IRES sequence is encoded by one of the sequences SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, and further includes up to 1000 additional nucleotides located at the 5' end of that sequence. In some embodiments, the IRES sequence is encoded by one of the sequences SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, and further includes up to 1000 additional nucleotides located at the 3' end of that sequence. In some embodiments, the IRES sequence is encoded by one of the sequences SEQ ID NOs: 1-228 or 229-17201, and further comprises up to 1000 additional nucleotides located at the 5' end of that sequence and up to 1000 additional nucleotides located at the 5' end of that sequence.

[0084] In some embodiments, the circular RNA molecule includes an internal ribosome entry site (IRES) sequence region, the IRES sequence region includes a sequence encoded by one of the DNA sequences of SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, the sequence encoded by one of the DNA sequences of SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201 having a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C.

[0085] In some embodiments, the circular RNA molecule includes an internal ribosome entry site (IRES) sequence region, the IRES sequence region includes a sequence encoded by one of the DNA sequences of sequence numbers 1-228 or 229-17201, and the IRES sequence region has a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C over its entire length.

[0086] In some embodiments, the circular RNA molecule includes an internal ribosome entry site (IRES) sequence region, the IRES sequence region includes a sequence encoded by one of the DNA sequences of sequence numbers 1-228 or 229-17201, further including up to 1000 additional nucleotides located at the 5' end and up to 1000 additional nucleotides located at the 5' end, the IRES sequence region having a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C over its entire length.

[0087] In some embodiments, the recombinant circular RNA molecule comprises a protein-coding nucleic acid sequence operably linked to an IRES in a non-native configuration. Any protein or polypeptide of interest (e.g., peptides, polypeptides, protein fragments, protein complexes, fusion proteins, recombinant proteins, phosphorylated proteins, glycoproteins, or lipoproteins) can be encoded by the protein-coding nucleic acid sequence. In some embodiments, the protein-coding nucleic acid sequence encodes a therapeutic protein. Examples of suitable therapeutic proteins include cytokines, toxins, tumor suppressor proteins, growth factors, hormones, receptors, mitotic agents, immunoglobulins, neuropeptides, neurotransmitters, and enzymes. Alternatively, the protein-coding nucleic acid sequence can encode an antigen of a pathogen (e.g., bacteria, viruses, fungi, protists, or parasites), and the circRNA can be used as a vaccine or as a component of a vaccine. Therapeutic proteins and examples thereof are further described, for example, in Dimitrov, DS, Methods Mol Biol., 899: pp. 1-26 (2012); and Lagasse et al., F1000 Research, 6: pp. 113 (2017).

[0088] Ideally, IRESs are "in-frame" with respect to the protein-coding nucleic acid sequence; that is, the IRES is located in the precise reading frame for the encoded protein within the circRNA molecule. Examples of IRES elements found to be in-frame with one or more coding sequences are shown in SEQ ID NOs: 28984–32953. However, in some embodiments, the IRES may be "out-of-frame" with respect to the protein-coding nucleic acid sequence, such that the IRES's position disrupts the ORF of the protein-coding nucleic acid sequence. In other embodiments, the IRES may overlap with one or more ORFs of the protein-coding nucleic acid sequence. Furthermore, in some embodiments, the protein-coding nucleic acid sequence includes at least one stop codon, while in other embodiments, the protein-coding nucleic acid sequence may lack a stop codon. We have found that circRNA molecules containing a protein-coding nucleic acid sequence having an in-frame non-natural IRES and lacking a stop codon can initiate a recursive (i.e., infinite loop) translation mechanism. This recursive translation may produce a chain of protein multimers (e.g., >200 kDa). This particular circRNA design allows for the production of up to 10 repetitions of ORF units of single ORF size. Without being bound by any specific theory, the use of the circRNAs described herein for recursive gene coding may represent a novel “data compression” algorithm for genes that addresses the gene size limitations associated with many current gene therapy applications.

[0089] In some embodiments, the IRES comprises (i) at least one RNA secondary structure element and (ii) a sequence complementary to 18S rRNA. In some embodiments, the IRES comprises (i) at least one RNA secondary structure element and (ii) a sequence complementary to 18S rRNA, and the RNA secondary structure of the IRES is formed from nucleotides at approximately positions 40 to 60 of the IRES, where the first nucleic acid at the 5' end of the IRES is considered to be position 1. The relative positions of the at least one RNA secondary structure and the sequence complementary to 18S RNA can vary. For example, in some embodiments, the IRES comprises (i) at least one RNA secondary structure element and (ii) a sequence complementary to 18S rRNA, and the at least one RNA secondary structure is located 5' to the sequence complementary to 18S rRNA (see Figure 4K). In some embodiments, the IRES comprises (i) at least one RNA secondary structure element and (ii) a sequence complementary to 18S rRNA, wherein the at least one RNA secondary structure element is located 3' to the sequence complementary to 18S rRNA (see Figure 19B).

[0090] DNA molecules, vectors, and cells In some embodiments, this disclosure provides a DNA molecule comprising a nucleic acid sequence encoding any one of the recombinant circRNA molecules disclosed herein. Thus, DNA sequences that can be used to encode circular RNA are described herein. In some embodiments, the DNA sequence encodes a circular RNA comprising an IRES. In some embodiments, the DNA sequence encodes a circular RNA comprising a protein-coding nucleic acid. In some embodiments, the DNA sequence encodes a circular RNA molecule, the circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably ligated to the protein-coding nucleic acid sequence in a non-native configuration. In some embodiments, the DNA sequence encodes a protein-coding nucleic acid sequence, and the protein is a therapeutic protein.

[0091] The DNA sequences disclosed herein may, in some embodiments, include at least one non-coding functional sequence. For example, the non-coding functional sequence may be a microRNA (miRNA) sponge. The microRNA sponge may include a binding site complementary to the miRNA of interest. In some embodiments, the binding site of the sponge is specific to the miRNA seed region, thereby enabling the binding site of the sponge to block the entire family of miRNAs to which it relates. In some embodiments, the miRNA sponge is selected from one of the miRNA sponges shown in Table 3 below.

[0092] [Table 3]

[0093] In some embodiments, the non-coding sequence may also be an RNA-binding protein site. RNA-binding proteins and binding sites are therefore listed in numerous databases known to those skilled in the art, including RBPDB (rbpdb.ccbr.utoronto.ca). In some embodiments, the RNA-binding protein includes one or more RNA-binding domains selected from RNA-binding domains (RBD, RNP domain and RNA recognition motif, also known as RRM), K-homology (KH) domains (Type I and Type II), RGG (Arg-Gly-Gly) box, Sm domain; DEAD / DEAH box (SEQ ID NOs. 34036 and 34037), zinc finger (ZnF, mostly C-x8-X-x5-X-x3-H (SEQ ID NOs. 34038)), double-stranded RNA-binding domain (dsRBD), cold shock domain; Pumilio / FBF (PUF or Pum-HD) domain, and Piwi / Argonaute / Zwille (PAZ) domain.

[0094] In some embodiments, the DNA sequence includes an aptamer. An aptamer is a short, single-stranded DNA molecule that can selectively bind to a specific target. The target may be, for example, a protein, peptide, carbohydrate, small molecule, toxin, or living cell. Some aptamers can bind to DNA, RNA, self-aptamers, or other non-self-aptamers. Aptamers take on a variety of shapes due to their tendency to form helices and single-stranded loops. Exemplary DNA and RNA aptamers are listed in the aptamer database (scicrunch.org / resources / Any / record / nlx_144509-1 / SCR_001781 / resolver?q=*&l=).

[0095] In some embodiments, the DNA sequence encodes a circular RNA molecule containing approximately 200 to 10,000 nucleotides.

[0096] In some embodiments, the DNA sequence encodes a circular RNA molecule containing a spacer between the IRES and the start codon of the protein-coding nucleic acid sequence. The spacer may be of any length. For example, in some embodiments, the length of the spacer is selected to optimize the translation of the protein-coding nucleic acid sequence.

[0097] In some embodiments, the DNA sequence encodes a circular RNA molecule containing an IRES configured to facilitate rolling circle translation. In some embodiments, the DNA sequence encodes a circular RNA containing a protein-coding nucleic acid sequence lacking a stop codon. In some embodiments, the DNA sequence encodes a circular RNA molecule containing (i) an IRES configured to facilitate rolling circle translation, and (ii) a protein-coding nucleic acid sequence lacking a stop codon.

[0098] The DNA sequences described herein may be contained in one or more vectors. For example, in some embodiments, the viral vector includes a DNA sequence encoding circular RNA. The viral vector may be, for example, an adeno-associated virus (AAV) vector, an adenovirus vector, a retrovirus vector, a lentiviral vector, a vaccinia or herpesvirus vector.

[0099] In some embodiments, the viral vector is an AAV. As used herein, the term “adeno-associated virus” (AAV) includes, but is not limited to, AAV1, AAV2, AAV3 (including types 3A and 3B), AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, avian AAV, bovine AAV, canine AAV, equine AAV, sheep AAV, and any other AAV currently known or subsequently discovered. In some embodiments, the AAV vector may be one or more modified forms of AAV1, AAV2, AAV3 (including types 3A and 3B), AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, bird AAV, bovine AAV, canine AAV, equine AAV, or sheep AAV (i.e., forms containing one or more amino acid variants compared to them). Various AAV serotypes and their variants are described, for example, in BERNARD N. FIELDS et al., VIROLOGY, Vol. 2, Chapter 69 (4th edition, Lippincott-Raven Publishers). Several relatively new AAV serotypes and clades have been identified (see, for example, Gao et al., (2004) J Virology 78: pp. 6381-6388; Moris et al., (2004) Virology 33-: pp. 375-383). The genome sequences of various serotypes of AAV, as well as the sequences of native terminal repeats (TRs), Rep proteins, and capsid subunits, are publicly known in the art. These sequences can be found in the literature or in publicly available databases such as the GenBank® database.For example, GenBank trust numbers NC_044927, NC_002077, NC_001401, NC_001729, NC_001863, NC_001829, NC_001862, NC_000883, NC_001701, NC_001510, NC_006152, NC_006261, AF063497, U89790, AF043303, AF028705, AF028704, J02275, JO See 1901, J02275, X01457, AF288061, AH009962, AY028226, AY028223, NC_001358, NC_001540, AF513851, AF513852, AY530579; the disclosures of the aforementioned documents are incorporated herein by reference to teach parvovirus and AAV nucleic acid and amino acid sequences. For example, Srivistava et al., (1983) J Virology 45:555; Chiorini et al., (1998) J. Virology 71:6823; Chiorini et al., (1999) J Virology 73:1309; Bantel-Schaal et al., (1999) J. Virology 73:939; Xiao et al., (1999) J. Virology 73:3994; Muramatsu et al., (1996) Virology 221:208; Shade et al., (1986) J Virol.58:921; Gao et al., (2002) Proc.Nat.Acad.Sci.USA 99:1 1854; Moris et al., (2004) Virology See also pp. 33-375-383; International Publication Nos. 00 / 28061, 99 / 61601, 98 / 11244; and U.S. Patent No. 6,156,303; the disclosures of the aforementioned documents are incorporated herein by reference.

[0100] In some embodiments, the DNA sequences described herein are included in the AAV2 vector or a variant thereof. In some embodiments, the DNA sequences described herein are included in the AAV4 vector or a variant thereof. In some embodiments, the DNA sequences described herein are included in the AAV8 vector or a variant thereof. In some embodiments, the DNA sequences described herein are included in the AAV9 vector or a variant thereof.

[0101] In some embodiments, the DNA sequences described herein are contained in virus-like particles (VLPs). Virus-like particles closely resemble viruses but are non-infectious because they contain little to no viral genetic material. Virus-like particles can be synthesized through the expression of naturally occurring viral structural proteins, which then self-assemble into virus-like structures. VLPs can be created using combinations of structural capsid proteins derived from different viruses. For example, VLPs may be derived from AAVs, retroviruses, flaviviridae, paramyxoviridae, or bacteriophages. VLPs can be produced in a number of cell culture systems, including bacteria, mammalian cell lines, insect cell lines, yeast, and plant cells.

[0102] In some embodiments, the DNA sequences described herein are contained in a nonviral vector. The nonviral vector may be, for example, a plasmid containing the DNA sequence. In some embodiments, the nonviral vector is closed-end DNA. Closed-end DNA is a nonviral capsid-free DNA vector having a covalent closed end (e.g., WO2019 / 169233). In some embodiments, a miniintron plasmid vector contains the DNA sequences described herein. The miniintron plasmid is an expression system containing a bacterial origin of replication and a select marker that maintains the juxtaposition of the 5' and 3' ends of the transgene expression cassette as a small circular (e.g., see Lu, J. et al., Mol Ther (2013) 21(5) pp. 954-963).

[0103] In some embodiments, the DNA sequences described herein are contained in lipid nanoparticles. Lipid nanoparticles (or LNPs) are submicron-sized lipid emulsions and may offer one or more of the following advantages: (i) controlled and / or targeted drug release, (ii) high stability, (iii) biodegradability of the lipids used, (iv) avoidance of organic solvents, (v) ease of scale-up and sterilization, (vi) not as expensive as polymer / surfactant-based carriers, and (vii) easy to establish and obtain regulatory approval. In some embodiments, the lipid nanoparticles have a diameter in the range of about 10 to about 1000 nm.

[0104] In some embodiments, the DNA sequence encodes a circular RNA molecule, which comprises a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) that operably ligates to the protein-coding nucleic acid sequence in a non-native configuration, the IRES comprising at least one RNA secondary structure and a sequence complementary to 18S RNA (rRNA).

[0105] In some embodiments, the DNA sequence encodes a circular RNA molecule, the circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence in a non-native configuration, the IRES comprising at least one RNA secondary structure element and a sequence complementary to 18S RNA (rRNA), the IRES having a minimum free energy of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C, and the RNA secondary structure element being formed from nucleotides approximately 40 to approximately 60 of the IRES, where the first nucleic acid at the 5' end of the IRES is considered to be at position 1.

[0106] In some embodiments, the DNA sequence comprises a nucleic acid sequence encoding a circular RNA molecule, the circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence in a non-native configuration, the IRES being encoded by one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a nucleic acid sequence that is at least 90% or at least 95% identical thereto.

[0107] Cells containing recombinant circRNA molecules, DNA molecules, or vectors described herein are also provided herein. Any prokaryotic or eukaryotic cell that can be stably maintained in contact with a recombinant circRNA molecule, a DNA molecule encoding a recombinant circRNA molecule, or a vector containing a recombinant circRNA molecule may be used in the context of this disclosure. Examples of prokaryotic cells include, but are not limited to, cells derived from the genera Bacillus (such as Bacillus subtilis and Bacillus brevis), Escherichia (such as E. coli), Pseudomonas, Streptomyces, Salmonella, and Erwinia. In some embodiments, the host cell is a eukaryotic cell. Suitable eukaryotic cells are known in the art and include, for example, yeast cells, insect cells, and mammalian cells. Examples of yeast cells include those derived from the genera Hansenula, Kluyveromyces, Pichia, Rhinosporidium, Saccharomyces, and Schizosaccharomyces. Suitable insect cells include Sf-9 and HIS cells (Invitrogen, Carlsbad, Calif.), which are described, for example, Kitts et al., Biotechniques, 14:810-817 (1993); Lucklow, Curr. Opin. Biotechnol., 4:564-572 (1993); and Lucklow et al., J. Virol., 67:4566-4579 (1993).

[0108] In some embodiments, the cells are mammalian cells. Several mammalian cells are known in the art, many of which are available from the United States Cell Culture and Cell Lineage Preservation Center (ATCC, Manassas, Va.). Examples of mammalian cells include, but are not limited to, HeLa cells, HepG2 cells, Chinese hamster ovary cells (CHO) (e.g., ATCC No. CCL61), CHO DHFR- cells (Urlaub et al., Proc. Natl. Acad. Sci. USA, 97:4216-4220 (1980)), human embryonic kidney (HEK) 293 or 293T cells (e.g., ATCC No. CRL1573), and 3T3 cells (e.g., ATCC No. CCL92). Other mammalian cell lines include the monkey COS-1 (e.g., ATCC No. CRL1650) and COS-7 cell lines (e.g., ATCC No. CRL1651), and the CV-1 cell line (e.g., ATCC No. CCL70). Further exemplary mammalian host cells include primate and rodent cell lines, including transformed cell lines. Cell lines derived from normal diploid cells, primary tissues, and in vitro cultures of primary explants are also suitable. Other mammalian cell lines include, but are not limited to, mouse neuroblastoma N2A cells, HeLa, mouse L-929 cells, and BHK or HaK hamster cell lines, all of which are available from the United States Cell Culture and Cell Lineage Preservation Center (ATCC, Manassas, Va). Methods for selecting mammalian cells and for transforming, culturing, amplifying, screening, and purifying such cells are well known in the art (e.g., Ausubel et al., see above). In some embodiments, the mammalian cells are human cells.

[0109] Methods for producing proteins This disclosure further provides a method for producing a protein in a cell, comprising the step of contacting the cell with the recombinant circular RNA molecule, the DNA molecule containing a nucleic acid sequence encoding a recombinant circRNA molecule, or a vector containing a recombinant circRNA molecule, under conditions in which a protein-coding nucleic acid sequence is translated in the cell and a protein is produced.

[0110] In some embodiments, a method for producing a protein in cells includes the step of contacting cells with a DNA sequence described herein, or a vector containing a DNA sequence, under conditions in which a protein-coding nucleic acid sequence is translated in the cells and a protein is produced. Proteins produced by the disclosed method are also provided.

[0111] In some embodiments, protein production is tissue-specific. For example, the protein may be selectively produced in one or more of the following tissues: muscle, liver, kidney, brain, skin, pancreas, blood, or heart.

[0112] In some embodiments, the protein is recursively expressed in the cell.

[0113] In some embodiments, the half-life of circular RNA in cells is approximately 1 to 7 days. For example, the half-life of circular RNA may be approximately 1, 2, 3, 4, 5, 6, 7, or more days.

[0114] In some embodiments, the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer than when the protein-coding nucleic acid sequence is supplied to the cell using a viral vector encoding linear RNA or as linear RNA.

[0115] In some embodiments, the protein is produced in the cell at levels at least about 10%, at least about 20%, or at least about 30% higher than when the protein-coding nucleic acid sequence is supplied to the cell using a viral vector or as linear RNA.

[0116] Using the IRES sequences described herein to express proteins from circular RNA, in some embodiments, continuous expression of proteins from circular RNA may be possible in cells even under stress conditions. In response to one or more stress conditions, protein production from linear RNA is often suppressed. Therefore, in some embodiments, circRNA can be used as a substitute for protein production from linear RNA during stress conditions. In some embodiments, proteins expressed from circular RNA in cells are expressed under one or more stress conditions. In some embodiments, protein expression from circular RNA in cells is not substantially interfered with when cells are exposed to one or more stress conditions. For example, when cells are exposed to one or more stress conditions, protein expression from circular RNA may change by less than 15%, less than 10%, less than 5%, less than 3%, less than 1%, or less than 0.5%. In some embodiments, proteins expressed from circular RNA are expressed under one or more stress conditions at levels substantially the same as those expressed in the same cells in the absence of one or more stress conditions. In some embodiments, the expression level of protein from circular RNA in cells is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% compared to the expression level in the absence of one or more stress conditions. A non-limiting list of conditions that may cause cellular stress includes temperature changes (including exposure to extreme temperatures and / or heat shock), exposure to toxins (including viral or bacterial toxins, heavy metals, etc.), electromagnetic irradiation, mechanical damage, viral infection, etc.

[0117] In some embodiments, the circRNAs described herein (including their components such as IRES sequences) promote cap-independent translational activity from the circRNA. Canonical translation via cap-independent mechanisms may be reduced in some human diseases. Therefore, using circRNAs to express proteins may be particularly useful in treating such diseases. In some embodiments, the use of the circRNAs described herein promotes cap-independent translational activity from the circRNA under conditions where cap-dependent translation is reduced or blocked in cells.

[0118] As discussed above, the translation of a protein-coding nucleic acid sequence can occur in an infinite loop (i.e., recursively) if the IRES is in frame with the protein-coding nucleic acid sequence and the protein-coding nucleic acid sequence lacks a stop codon. Therefore, in some embodiments, the method of producing proteins in cells produces a chain of proteins.

[0119] Any prokaryotic or eukaryotic host cell described herein may be brought into contact with a recombinant circRNA molecule or a vector containing a circRNA molecule. The host cell may be a mammalian cell, such as a human cell. In some embodiments, the cell is in vivo. In some embodiments, the cell is in vitro. In some embodiments, the cell is ex vivo. In some embodiments, the cell is in a mammal, such as a human.

[0120] In some embodiments, regardless of the selected cell type, 5' cap-dependent translation is impaired in the cell (e.g., reduced, decreased, inhibited, or completely eliminated). In some embodiments, substantial 5' cap-dependent translation is absent in the cell.

[0121] A recombinant circular RNA molecule, a DNA molecule encoding a recombinant circular RNA molecule, or a vector containing a DNA molecule encoding a recombinant circular RNA molecule may be introduced into a cell by any method, including, for example, transfection, transformation, or transduction. The terms “transfection,” “transformation,” and “transduction” are used interchangeably herein and refer to the introduction of one or more exogenous polynucleotides into a host cell by physical or chemical means. Many transfection techniques are known in the art, including, for example, calcium phosphate DNA coprecipitation (see, e.g., Murray E.J (ed.), Methods in Molecular Biology, Vol. 7, Gene Transfer and Expression Protocols, Humana Press (1991)); DEAE dextran; electroporation, cation liposome-mediated transfection; tungsten particle-enhanced microparticle gun method (Johnston, Nature, 346:776-777 (1990)); strontium phosphate DNA coprecipitation (Brash et al., Mol. Cell. Biol., 7:2031-2034 (1987)); and magnetic nanoparticle-based gene delivery (Dobson, J., Gene Ther, 13(4):283-287 (2006)).

[0122] Naked RNA, a DNA molecule encoding a circular RNA molecule, or a vector containing circular RNA or DNA encoding circular RNA may be administered to cells in the form of a composition. In some embodiments, the composition includes a pharmaceutically acceptable carrier. The choice of carrier depends on the type of cell(s) into which the specific circular RNA molecule, DNA sequence, or vector and the circular RNA molecule, DNA sequence, or vector are introduced. Thus, various formulations of the composition are possible. For example, the composition may contain preservatives such as methylparaben, propylparaben, sodium benzoate, and benzalkonium chloride. A mixture of two or more preservatives may be used as needed. Furthermore, buffers may be used in the composition. Suitable buffers include, for example, citric acid, sodium citrate, phosphoric acid, potassium phosphate, and various other acids and salts. A mixture of two or more buffers may be used as needed. Methods for preparing compositions for pharmaceutical use are known to those skilled in the art and are described in more detail, for example, in Remington: The Science and Practice of Pharmacy, Lippincott Williams & Wilkins; 21st edition (May 1, 2005).

[0123] In some embodiments, compositions containing recombinant circular RNA molecules, DNA sequences, or vectors can be formulated as encapsulation complexes, such as cyclodextrin encapsulation complexes, or as liposomes. Using liposomes allows for targeting of host cells or an increase in the half-life of the circular RNA molecule. Methods for preparing liposome delivery systems are described, for example, in Szoka et al., Ann. Rev. Biophys. Bioeng., 9:467 (1980), and in U.S. Patents 4,235,871, 4,501,728, 4,837,028, and 5,019,369. Recombinant circRNA molecules may also be formulated as nanoparticles.

[0124] Host cells can be brought into contact in vivo or in vitro with recombinant circular RNA molecules, DNA sequences, or vectors, or compositions containing any of the aforementioned. The term “in vivo” refers to a method performed in a normal, intact state within a living organism, while an “in vitro” method uses components of an organism isolated from normal biological conditions. When the method is performed in vivo, in some embodiments, protein production is tissue-specific. “Tissue-specific” means that the protein is produced only in a subset of tissue types within the organism, or at a higher level in a subset of tissue types compared to baseline expression across all tissue types. The protein may be produced in any tissue type, such as muscle, liver, kidney, brain, lung, skin, pancreas, blood, or heart tissue.

[0125] Inhibition of circRNA translation This disclosure also provides oligonucleotide molecules containing nucleic acid sequences that hybridize to internal ribosome entry sites (IRESs) present on circular RNA molecules and inhibit translation of circRNA molecules. In some embodiments, the circular RNA molecule is a naturally occurring circular RNA molecule. In some embodiments, the circular RNA molecule is a recombinant circular RNA molecule, such as the recombinant circRNA molecules described herein. In some embodiments, as described herein, the recombinant circRNA molecule comprises a protein-coding nucleic acid sequence and an IRES operably ligated to the protein-coding nucleic acid sequence (in an unnatural configuration, if necessary), the IRES comprising at least one RNA secondary structure and a sequence complementary to 18S ribosomal RNA (rRNA), the IRES having a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C. In some embodiments, a recombinant circRNA molecule comprises a protein-coding nucleic acid sequence and an IRES (in an unnatural configuration, if necessary) operably ligated to the protein-coding nucleic acid sequence, the IRES encoding one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a nucleic acid sequence that is at least 90% or at least 95% identical thereto.

[0126] The oligonucleotide that hybridizes to the IRES on the recombinant circRNA molecule may be of any type and size. In some embodiments, the oligonucleotide may be about 8 to about 80 nucleotides long, such as about 15 to about 30 nucleotides long. In some embodiments, the oligonucleotide may be about 20, about 22, or about 24 nucleotides long. In some embodiments, the oligonucleotide may be an antisense oligonucleotide (also called "ASO"). The term “antisense oligonucleotide,” as used herein, refers to a short, synthetic, single-stranded oligodeoxyribonucleotide that is complementary to a target RNA sequence and can reduce, restore, or modify protein expression through several different mechanisms (Rinaldi, C., Wood, M., Nat Rev Neurol, 14: pp. 9-21 (2018); Crooke, ST, Nucleic Acid Ther., 27: pp. 70-77 (2017); and Chan et al., Clin. Exp. Pharmacol. Physiol. 33: pp. 533-540 (2006)). In some embodiments, the antisense oligonucleotide may be a locked nucleic acid oligonucleotide (LNA). The term "locked nucleic acid oligonucleotide (LNA)" refers to an oligonucleotide containing one or more nucleotide components in which an extra methylene crosslink fixes the ribose moiety in the C3'-end (beta-D-LNA) or C2'-end (alpha-L-LNA) three-dimensional structure (Grunweller A, Hartmann RK, BioDrugs, 21(4):235-243 (2007)). In some embodiments, the oligonucleotide is a small interfering RNA (siRNA), a small hairpin RNA (shRNA), CRISPR (sgRNA), or a microRNA (miRNA).

[0127] Oligonucleotides may contain one or more modifications that enhance the hybridization of the oligonucleotide to IRES of circRNA molecules and / or the inhibition of translation of circRNA molecules. The modifications may be at the 5' or 3' end of the oligonucleotide. Suitable modifications include, but are not limited to, modified nucleoside bonds, modified sugars, or modified nucleic acid bases. In some embodiments, oligonucleotides may be conjugated to fluorophores (e.g., Cy3, FAM, Alexa488, etc.) or to other molecules (e.g., biotin, alkaline phosphatase, antibody, nucleic acid aptamer, peptide, etc.). In some embodiments, oligonucleotides may be labeled with peptides or proteins, for example, using CLICK chemistry. The naturally occurring nucleoside bond between RNA and DNA is recognized as a 3'-5' phosphate diester bond. Oligonucleotides with one or more modifications, i.e., non-naturally occurring nucleoside bonds, are known to exhibit desirable properties such as enhanced intracellular uptake, enhanced affinity for target nucleic acids, and increased stability in the presence of nucleases. Modified nucleoside bonds include, for example, nucleoside bonds that retain a phosphorus atom and nucleoside bonds that do not contain a phosphorus atom. Typical phosphorus-containing nucleoside bonds include, but are not limited to, phosphate diesters, phosphate triesters, methylphosphonic acid, phosphoroamidic acid, and phosphorothioate. Methods for preparing phosphorus-containing and non-phosphorus-containing bonds are well known.

[0128] In some embodiments, oligonucleotide molecules may include a modified skeleton. Oligonucleotides having a modified skeleton include modified skeletons that retain a phosphorus atom in the skeleton and modified skeletons that do not have a phosphorus atom in the skeleton. Modified oligonucleotides that do not have a phosphorus atom in the internucleoside skeleton are sometimes called "oligonucleosides." Examples of modified oligonucleotide skeletons include, but are not limited to, phosphorothioates, chiral phosphorothioates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkyl phosphonates including 3'-alkylene phosphonates, 5'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidic acids including 3'-aminophosphoramic acids and aminoalkylphosphoramic acids, thionophosphoramic acids, thionoalkyl phosphonates, thionoalkyl phosphate triesters, selenophosphates, and normal 3'-5' bonds, their 2'-5' bond analogs, and boranophosphates having analogs with inverted polarity where one or more intranucleotide bonds are 3'-to-3', 5'-to-5', or 2'-to-2' bonds.

[0129] Oligonucleotide molecules may further contain one or more nucleotides having modified sugar moieties. Sugar modifications may impart nuclease stability, binding affinity, or any other beneficial biological properties to oligonucleotides. Modifications of the furanosyl sugar ring of a nucleoside include the addition of substituents, particularly at the 2' position; bridging of two non-geminal ring atoms to form a bicyclic nucleic acid (BNA); and the replacement of the ring oxygen at the 4' position with -S-, -N(R)-, or -C(R) 1 )(R 2 Modified sugars can be modified in several ways, including but not limited to the substitution of atoms or groups such as ). Modified sugars include, but are not limited to, substituted sugars, particularly 2'-substituted sugars having substituents such as 2'-F, 2'-OCH2(2'-OMe), or 2'-O(CH2)2-OCH3(2'-O-methoxyethyl or 2'-MOE); and bicyclic modified sugars (BNA), n=1 or n=2, having a 4'-(CH2)nO-2' bridge. Methods for preparing modified sugars are well known to those skilled in the art.

[0130] In some embodiments, oligonucleotides may be chemically modified at their 5' and / or 3' ends. In other words, one or more portions may be chemically (e.g., covalently) linked to the 5' and / or 3' ends of an oligonucleotide molecule. Such modifications include, for example, the chemical bonding of protein or sugar moieties to the 5' and / or 3' ends of an oligonucleotide molecule. Other modifications that enhance oligonucleotide affinity to target nucleic acids and / or increase oligonucleotide stability are known in the Art (see, for example, U.S. Patent Application Publication 2019 / 0323013) and may be used in the context of this disclosure.

[0131] The disclosure also provides a method for inhibiting the translation of a protein-coding nucleic acid sequence present on a circular RNA molecule, comprising the step of contacting the circular RNA molecule with the above-mentioned oligonucleotide molecule, thereby causing the oligonucleotide molecule to hybridize to a nucleic acid sequence complementary to the RNA secondary structure and / or 18S rRNA present on the IRES of the circular RNA molecule, thereby inhibiting the translation of the circular RNA molecule.

[0132] In some embodiments, oligonucleotides can hybridize to RNA secondary structure elements present on the IRES or to nucleic acid sequences complementary to 18S rRNA. In other embodiments, oligonucleotides can hybridize to both RNA secondary structure elements and nucleic acid sequences complementary to 18S rRNA. For example, a first oligonucleotide may hybridize to an RNA secondary structure element, and a second oligonucleotide may hybridize to a nucleic acid sequence complementary to 18S rRNA. Alternatively, a single oligonucleotide may hybridize to both RNA secondary structure elements present on the IRES and nucleic acid sequences complementary to 18S rRNA. Appropriate hybridization strictness conditions are described herein.

[0133] (i) DNA sequences disclosed herein and (ii) compositions comprising a non-coding circular RNA or a DNA sequence encoding it. The DNA sequence may encode a circRNA. In some embodiments, the non-coding circular RNA may comprise one or more of the following: a binding site for an RNA-binding protein, an aptamer, or a miRNA sponge. In some embodiments, the non-coding circular RNA may have one or more functions, such as absorbing miRNA, regulating mRNA splicing mechanisms, sequestering RNA-binding proteins (RBPs), regulating RBP interactions, or activating an immune response.

[0134] A method for delivering a non-coding circular RNA to a cell is also provided, comprising the step of contacting the cell with (i) a DNA sequence disclosed herein and (ii) a composition comprising a non-coding circular RNA and a DNA sequence encoding it, thereby delivering the non-coding circular RNA to the cell.

[0135] The following embodiments further illustrate the present invention, but should not be interpreted as limiting its scope. [Examples]

[0136] The following example illustrates the development of a high-throughput screening method for systematically identifying and quantifying RNA sequences that can direct circRNA translation. Over 17,000 circRNA internal ribosome entry sites (IRESs) were identified and validated, demonstrating that 18S rRNA complementarity and structured RNA elements on IRESs are crucial for promoting circRNA cap-independent translation. Genomic and peptidochemical analyses of IRESs identified nearly 1,000 putative endogenous protein-coding circRNAs, along with hundreds of translational units encoded by these circRNAs. The protein circFGFR1p, encoded by circFGFR1, was also characterized. This protein functions as a negative regulator of FGFR1 for suppressing cell growth under stress conditions.

[0137] [Example 1] This example describes the systematic identification of RNA sequences that promote cap-independent circRNA translation. Canonical translation via cap-independent mechanisms can be reduced in numerous human diseases. Therefore, the use of circRNAs for protein expression may be particularly useful in treating such diseases.

[0138] To systematically identify RNA sequences that can promote cap-independent translation on circRNA, we developed an oligo-split-eGFP-circRNA reporter construct that enables high-throughput screening and quantification of the cap-independent translational activity of synthetic oligonucleotide inserts (hereinafter referred to as "oligos") on circRNA (Figure 1A). Specifically, the construct contains a bicistronic mRuby reporter followed by a rearranged split-eGFP reporter flanked by a human ZKSCAN1 intron. During transcription, the construct's pre-mRNA undergoes spliceosome-mediated backsplicing to reconstitute full-length eGFP on circRNA. Since full-length eGFP is reconstituted only during backsplicing, the eGFP fluorescence signal can only originate from circRNA via cap-independent translation. Next, a synthetic oligonucleotide library was cloned into the construct to drive the expression of the eGFP reporter (Figure 1A). This library contained 55,000 oligonucleotide sequences from reported IRESs in the IRESite database (including human and non-human IRESs; see Mokrejs et al., Nucleic Acids Research 38, D131-D136 (2009)), native 5' untranslated regions (5'UTR) of viral and human genes, and native and synthetic sequences from viral and human transcripts (Figure 1A). The library design is detailed in Weingarten-Gabbay et al., Science 351, aad4939 (2016). For viral and human transcripts, we selected genes that have been reported to remain associated with polysomes when cap-dependent translation is suppressed, and genes that have alternative isoforms with different translation initiation sites (Figure 1A).

[0139] Two well-known concerns regarding bicistronic IRES screening are potential promoters or splice sites that activate transcription or readthrough of downstream open reading frames (ORFs), respectively. Since ectopic transcription of only the 5' fragment of split-eGFP cannot produce a fluorescent signal, the design used herein eliminates both concerns. Northern blotting, quantitative reverse transcription polymerase chain reaction (qRT-PCR), RNase R treatment, and reporter gene experiments confirmed that the detected eGFP signal did not originate from trans-splicing or nicking of eGFP circRNA (Figure 8A-8C). The reporter produced approximately 3000 nucleotides (nt) of primary linear transcript and approximately 900 nt of eGFP circRNA, and RNase R exonuclease treatment could efficiently remove the linear transcript but not the circRNA (Figure 8A). The mRuby gene allowed for normalization of transduction efficiency by translation of normal linear mRNA. After transfection of human fetal kidney (HEK) 293T cells, the transfected cells were sorted into seven bins based on the ratio of eGFP to mRuby fluorescence, and the frequency of oligo sequences in each pool was deconvolved by deep sequencing (Figure 1A). This system allows for high-throughput quantification of cap-independent translational activity on circRNA for each oligo in the library.

[0140] Of the 55,000 oligonucleotides, 40,855 were captured from the library (approximately 74.3%). To quantify the eGFP expression level of each oligonucleotide, the mean weighted rank distribution of reads across all bins was calculated. The weight for each bin is the fraction of the number of reads in this bin compared to the total number of reads in all seven bins. The rank is the number of bins from the bin with the lowest eGFP intensity (bin #1) to the bin with the highest eGFP intensity (bin #7) (Figure 1B). The quantification of translational activity was found to be highly reproducible between two independent biological replicates (Pearson correlation coefficient R = 0.70) (Figure 9A). Furthermore, it was confirmed that the results were not confused by changes in circRNA backsplicing efficiency due to different oligonucleotide inserts (Figures 9B and 9C). This screening assay revealed three groups of oligonucleotides based on their eGFP expression levels: a group of oligonucleotides that did not show eGFP expression (approximately 2,500) (eGFP expression (bin) = 0.0), and two groups of oligonucleotides showing a bimodal distribution of eGFP expression (eGFP expression (bin) = 0.8–2.2 and 2.4–7.0, respectively). To determine oligonucleotides with cap-independent translational activity (eGFP(+) oligonucleotides), the weighted rank distribution of eGFP intensity in cells transfected with a reporter plasmid without an inserted IRES (empty eGFP circRNA) was calculated as background eGFP expression. Oligonucleotides were defined as eGFP(+) oligonucleotides that exhibited higher eGFP expression than background eGFP expression (eGFP expression (bin) = 3.466387) (Figure 1B). Background eGFP expression was calculated based on the distribution of reads across the entire bin, rather than a simple cutoff value. This is a more conservative approach to avoid the possibility of false-positive events, as empty circRNA eGFP reporters may have weak translational activity (Figure 9D). Using this approach, 17,201 eGFP(+) oligos were identified from the screening assay (Figure 1B, SEQ ID NOs: 1-17201).Furthermore, the identified eGFP(+) oligos were able to initiate circRNA translation of reporters with different coding sequences (CDS), demonstrating that the screening results were not specific to eGFP reporters (Figure 9E). While circRNA translation was reproducibly detectable, circRNAs driven by eGFP(+) oligos exhibited considerably weaker cap-independent translation activity compared to linear RNA translation driven by cap-dependent translation (Figure 9F).

[0141] Previous studies have used the same synthetic oligo library in a linear bicistronic eGFP reporter screening assay to identify oligos with cap-independent translational activity for linear RNA on reporter plasmids without IRES insertions as a threshold (Weingarten-Gabbay et al., 2016). Therefore, it was possible to compare the cap-independent translational activity of each oligo sequence on linear and circular RNA. For each oligo, normalized eGFP expression was calculated from both circRNA and linear RNA templates. Of the oligos captured in both circRNA and linear RNA reporter screening (n=13,645), many oligos were found to exhibit cap-independent translational activity in both linear and circular RNA screening systems (n=7,424) (Figure 1D, Figure 11A). However, there was little correlation between the overall IRES activity of circular and linear RNA (Pearson's R=0.014; Spearman's R=0.010) (Figure 1D). Interestingly, we also captured several oligonucleotides that specifically exhibited IRES activity in either linear or circular screening systems (linear IRES and circular IRES, respectively) (Figure 1D). A more conservative approach was taken to define linear and circular IRESs: linear IRESs represent oligonucleotides that exhibit cap-independent translational activity only in linear RNA screening systems, and circular IRESs represent oligonucleotides that exhibit cap-independent translational activity only in circular RNA screening systems. Using this approach, we identified 4,582 circular IRESs and 1,639 linear IRESs (Figure 1D, Figure 11A).

[0142] Furthermore, when examining the distribution of human and viral IRESs among circular IRESs, linear IRESs, and IRESs exhibiting translational activity on both linear and circular RNAs, no significant differences were found among the IRESs (Figures 11B-11D). This result suggests that the recruitment or activity of circRNA-specific IRES transacting factors (ITAFs) in circulating IRESs can distinguish circRNAs from linear RNAs, depending on circRNA-specific biosynthesis such as circRNA backsplicing or circRNA nuclear export. In summary, these results demonstrate that high-throughput screening assays using circRNA reporter constructs can systematically identify RNA sequences with IRES activity that can promote cap-independent translation on circRNAs.

[0143] [Example 2] This example demonstrates that synthetic circRNAs containing eGFP+ oligo sequences are actively translated.

[0144] To validate the screening results of Example 1, polysome profiling was used to investigate whether the circRNAs containing the identified eGFP(+) oligo sequences were actively translated and involved with ribosomes. First, HEK-293T cells were transfected with an oligo-split-eGFP-circRNA reporter construct containing a synthetic oligo library. The transfected cells were treated with cycloheximide (CHX), and (poly)ribosome-associated RNA was isolated by sucrose gradient fractionation (Figures 2A, 2B, and 12A). Furthermore, fractions containing RNase R were processed to obtain highly enriched (poly)ribosome-associated circRNAs, and high-throughput sequencing was performed to identify the IRES sequences present in the circRNAs in each fraction. RNase R processing was performed under conditions that allowed RNase R to efficiently digest potential G quadruplexes containing linear RNA, and the RNase R digestion time was optimized to obtain more than 100 times circRNA enrichment compared to linear RNA (Figure 12B). Compared to CHX treatment, treatment of transfected cells with puromycin (PMY) shifted translated circRNA from the poly(ribosome)-associated fraction to the 40S fraction (Figure 12C), suggesting that CHX treatment can capture translated circRNA. To avoid confusion of results with weakly translated circRNA (data not shown), the ratio of (poly)ribosome-enriched oligos among eGFP(+) oligos with eGFP expression above the 80th percentile was calculated and compared with eGFP(-) oligos with eGFP expression below the 20th percentile. This result demonstrated that eGFP(+) oligos were more enriched in the (poly)ribosome fraction (57.2%) than eGFP(-) oligos (17.9%) (Figure 2C). It was confirmed that the higher enrichment of (poly)ribosome-associated eGFP(+) oligos was not caused by oligo capture efficiency or expression levels (Figures 12D and 12E). These results suggest that circRNAs containing eGFP(+) oligo sequences are translated more actively.However, because polysome profiling has low sensitivity for capturing weakly translated circRNAs (data not shown), quantitative translation initiation sequencing (QTI-seq) data were used for additional validation. Next, we examined data published from QTI-seq, a modified ribosome profiling (Ribo-seq) technique that maps translation initiation sites (TIS) across the genome (Gao et al., Nat Methods 12, 147-153, 2015). First, we investigated whether eGFP(+) oligo sequences overlapped with TIS identified on human transcripts. These results demonstrate that among oligos derived from human genomes with Ribo-seq coverage, the majority (approximately 76%) of eGFP(+) oligos overlap with identified TIS (TIS(+) oligos) on human transcripts identified by QTI-seq, while only 30% of eGFP(-) oligos are TIS(+) (Figures 2D and 2E, Figure 13A). This suggests that eGFP(+) oligos are more likely to initiate translation at these TIS than eGFP(-) oligos. Interestingly, by examining eGFP(+) / TIS(+) oligos in the human genome, three types of eGFP(+) / TIS(+) oligos were identified: (1) oligos containing an annotated translated start site on a linear transcript (annotated TIS; aTIS), (2) oligos containing an unannotated translated start site on a linear transcript that may be located in the 5'UTR, CDS, or 3'UTR region of the transcript (unannotated TIS; nTIS), and (3) oligos containing both aTIS and nTIS signals (dual TIS; dTIS) ​​(Figures 2D and 2E). These different types of TIS(+) oligos may suggest that the oligos utilize different mechanisms to initiate translation. aTIS oligos (approximately 30%) can utilize the same annotated translation initiation sites as linear transcripts for cap-independent translation, whereas, in contrast, nTIS oligos (approximately 41%) represent novel translation initiation sites different from those of linear transcripts for cap-independent translation, which have been observed when initiating the synthesis of alternative translation products.dTIS oligos can modulate dual activity between aTIS and nTIS, which requires further investigation, by utilizing regulatory mechanisms that are partially uncharacterized. Importantly, although the oligo library was enriched with oligos located upstream of annotated start codons, no bias was observed in the ratio of aTIS oligos, suggesting that this result was not confused by the design of the oligo library. Interestingly, eGFP(+) / TIS(+) oligos were found to be located within the genomic region encoding annotated circRNAs (Figure 2D), suggesting that these circRNAs can utilize TIS on the oligos to initiate endogenous circRNA translation. Nevertheless, when the location of the translation initiation site was mapped on each oligo, no translation initiation hotspots were observed on the oligos (Figure 13B), suggesting that translation initiation is not influenced by the location on the oligo. In summary, the above results provide strong evidence that the screening assay described herein can identify oligo sequences that can promote cap-independent translational activity on circRNAs.

[0145] [Example 3] This example describes the identification of 18S rRNA complementary sequences that promote circRNA translation.

[0146] Watson-Crick base pairing between IRESs and 18S ribosomal RNA (18S rRNA) has been demonstrated to promote cap-independent translation of linear mRNA. Therefore, we evaluated whether screening could identify regions on human 18S rRNA that can interact with circRNA IRESs and promote cap-independent translation. A synthetic oligo library was designed to contain 171 oligos with sequences complementary to human 18S rRNA, with a 10nt sliding window between each consecutive oligo that reconstructs the entire 1869nt full-length 18S rRNA sequence (Figure 3A, SEQ ID NOs. 28977-28983). For each position on 18S rRNA, the average eGFP expression of all oligos containing the corresponding complementary sequence was calculated. Using this sliding window method, circRNA IRES activity was determined for each 10nt window of the entire human 18S rRNA sequence (Figure 3A). Six "active regions" were identified on 18S rRNA, and complementary sequences within these active regions exhibited higher average eGFP expression than background eGFP expression (Figure 3B). Interestingly, active regions 2, 4, 5, and 6 possess helices that have been reported to contact mRNA in the eukaryotic ribosome initiation complex and interact with translated RNA (Pisarev et al., EMBO J 27, 1609-1621 (2008)) (Figure 3C). Furthermore, active region 4, as well as active regions 2 and 6, contained sequences that have been determined to promote cap-independent translation of IGF1R and HCV IRES, respectively, through Watson-Crick base pairing (Figure 3C). Active region 3 is located in one of the elongation segments of 18S rRNA (ES6S), which is involved in the recruitment of eukaryotic initiation factor 3 (eIF3). eIF3 directly binds to 5'UTR N6-methyladenosine (m6A) in linear mRNA, initiating cap-independent translation, suggesting that active region 3 on 18S r RNA may be important for eIF3-m6A-mediated cap-independent translation.Active region 1 interacts with active region 3 and is located in a separate elongation segment on the 18S rRNA (ES3S) that forms a tertiary structure, suggesting that active region 1 may also be involved in cap-independent translation mediated by eIF3-m6A. These results suggest that the identified active regions on the 18S rRNA do indeed play a role in promoting cap-independent translation.

[0147] Since heptomers derived from active region 4 have been shown to be enriched in IRESs reported for linear RNA, all heptomers were extracted from sequences complementary to the active region (active heptomer) of 18S rRNA, and the number of active heptomers in eGFP(+) and eGFP(-) oligos was compared. eGFP(+) oligos were found to be more enriched with active heptomers than eGFP(-) oligos (Figure 3D). In contrast, when comparing the number of matching random heptomers that do not overlap with active heptomers between eGFP(+) and eGFP(-) oligos, no significant difference was observed (Figure 3D), suggesting that the greater enrichment of active heptomers observed in eGFP(+) oligos is specific to the 18S rRNA complementary sequence. Nevertheless, no hotspot locations of active heptomers located on circRNA IRESs were observed (Figure 13C). To further verify the results, the 18S rRNA complementarity of IRESs was disrupted by replacing the 18S rRNA complementary sequence with a random heptamer or by adding an adjacent 18S rRNA complementary sequence to the IRES, and their circRNA translation activity was measured (Figure 3E). A decrease in IRES activity was observed when the 18S rRNA complementarity on the IRES was low, while conversely, IRES activity was more strongly programmed when the 18S rRNA complementarity added to the IRES was high (Figure 3E). These results suggest that circRNA IRESs containing RNA sequences complementary to the active region on 18S rRNA are one of the regulatory elements that can promote cap-independent translation on circRNAs.

[0148] [Example 4] This example describes the identification of essential elements on circ RNA IRES using systematic scanning mutagenesis.

[0149] We used scanning mutagenesis to define essential elements on circRNA IRESs. This analysis included 99 reported IRESs and 734 oligos designed for scanning mutagenesis of the natural 5'UTR in a synthetic oligo library. The oligos were designed as non-overlapping sliding windows of 14nt random substitution mutation tiling across the entire IRES or the 5'UTR (Figure 3F). Screening results allowed us to determine the impact of substitution mutations on each 14nt window on the IRES activity of the entire oligo sequence (Figure 3F). Essential elements (Figure 3G; highlighted in blue) were determined as the region from the start site of a mutation where IRES activity sharply decreased (Figure 3G; black dot) to the next start site of a mutation where IRES activity resumed or exceeded the mean eGFP expression level. By comparing the quantitative results with those of well-characterized IRESs and hepatitis C virus (HCV) IRESs, it was observed that known functional domains on HCV IRESs (Figure 3G; red line) colocalized with mutation sites where IRES activity was dramatically reduced. The specific reduction in IRES activity at these mutation sites suggests that the mutations disrupted essential elements of the IRES, rendering its cap-independent translational activity invalid. This result demonstrates that the assay can indeed capture all the essential elements reported on HCV IRESs, as well as elements that may be novel essential elements that have not yet been characterized (Figure 3G).

[0150] Furthermore, using the same scanning mutagenesis assay, essential elements on the circRNA IRES identified by scanning mutagenesis were further identified. The synthetic oligo library contains oligos that have a sliding window of 14nt random substitution mutation tiling across the entire circRNA IRES. Scanning mutagenesis captured two classes of circRNA IRES: locally sensitive circRNA IRES (Figure 3H; top), which show a decrease in circRNA IRES activity only when mutations occur at specific locations, and globally sensitive circRNA IRES (Figure 3H; bottom), where mutations at most locations can cause a decrease in IRES activity. Locally sensitive and globally sensitive IRESs were defined as whether IRES activity is affected by single or multiple mutations, respectively. Globally sensitive IRESs have more structured sequences (i.e., significantly lower minimum free energy (MFE) values) compared to locally sensitive IRESs. This suggests that the overall secondary structure of globally sensitive IRESs is important for their IRES activity, as more structured sequences are more likely to be affected by mutations regardless of the mutation site. On the other hand, locally sensitive IRESs have less structured sequences and are more resistant to mutations. This suggests that the IRES activity of locally sensitive IRESs can be regulated by short sequences as essential elements.

[0151] By overlaying the eGFP expression levels of all captured circRNA IRESs with the overall sensitivity, three regions (5–15nt, 40–60nt, and 135–165nt) were identified on the IRESs. When mutations hit these regions, IRES activity was significantly reduced (Figure 3I), suggesting that these regions may contain important elements for promoting cap-independent translation of circRNA IRESs. To further characterize whether the elements contained in these regions are structure-dependent, local MFE was calculated along circRNA IRESs with overall sensitivity in a non-overlapping 15nt window. The local MFE of the 40–60nt region on the circRNA IRES was found to be significantly low (Figures 3I and 3J; shaded in red), suggesting that this region may contain local structural elements that can drive circRNA translation. In contrast, the local MFEs in the 5–15nt and 135–165nt regions are no different from other regions on the IRES (Figures 3I and 3J; shaded in blue), suggesting that the regulatory elements located in these two regions are not involved in the local secondary structure.

[0152] In summary, this data demonstrates that scanning mutagenesis allows assays to determine circRNA IRESs with local or global sensitivity and systematically identify essential circRNA IRES elements required for IRES activity in a high-throughput manner.

[0153] [Example 5] This example demonstrates that stem-loop structures at different locations on circRNA IRES promote cap-independent translation.

[0154] Many natural and synthetic IRESs have been reported to promote cap-independent translation on circRNAs, but most IRESs have been characterized in linear RNA reporter systems, and the different IRES activities of linear and circular RNAs have not yet been reported. By comparing the screening results from the circRNA reporter system described herein with previous screening studies performed in linear RNA reporter systems using the same synthetic oligo library, we identified two distinct groups of oligos that possess IRES activity, particularly on either linear or circular RNA (linear IRESs and circular IRESs, respectively). To characterize the properties of these oligos that allow for the distinction between linear and circular IRESs, we analyzed the primary sequences of the oligos and found that circular IRESs have a higher GC content and a lower MFE than linear IRESs (Figure 14A). On the other hand, there was no difference in the number of canonical translation start codons (AUG), Kozak consensus sequences (ACCAUGG, SEQ ID NO: 34039), and m6A motifs (RRACH, SEQ ID NO: 33944) between linear and circular IRESs (Figures 14B-14D). Since low MFE often indicates that the RNA has a stable secondary structure, the low MFE of circular IRESs suggests that some structural elements may play a role in promoting cap-independent translation activity on circRNAs.

[0155] Next, the secondary structures of linear and circular IRESs were characterized using M2-seq, an assay that employs systematic mutation profiling and chemical structure probing to capture RNA secondary structures with a very low false-positive rate. Four circular and linear IRESs that specifically exhibit their respective IRES activities on circRNA or linear RNA were selected (Figure 14E), and their secondary structures were determined using M2-seq. These oligos exhibit potent IRES activity in either the circular RNA system (circular IRESs: 6742, 9128, 19420, and 18377) or the linear RNA system (linear IRESs: 6885, 6103, 5471, and 6527), and do not exhibit read-through translation activity, ribosome reinitiation, or hidden promoter activity on the linear bicistronic construct. By selecting these oligos, it was confirmed that the eGFP signal detected on the linear bicistronic construct originated from the cap-independent translation of the oligos. M2-seq results revealed that linear IRESs possess structuring elements, but circular IRESs are generally more structured than linear IRESs (Figures 4A and 4B, 14F and 14G). Of all four circular IRESs examined, all contained stem-loop structured RNA elements (SuREs) at different positions (40–60 nt from the +1 position (first nucleotide) of the IRES), while none of the linear IRESs examined contained such structures at this position (Figures 4A and 4B, 14F and 14G). Consistent with previous systematic scanning mutation profiling (Figures 3I and 3J), which also suggests that these different positions on the IRES contain structuring elements necessary to promote circRNA translation, it is proposed that SuREs at these different positions on the IRES can promote cap-independent translational activity on circular IRESs.

[0156] To test this hypothesis, we disrupted the SuRE at the 40-60 nt position on a cyclic IRES (oligo index: 6742) by substituting it with a sequence extracted from the same position on a linear IRES (oligo index: 6885), thereby forming a different secondary structure at this position (Figure 4C). Interestingly, disruption of the SuRE at this position on the cyclic IRES resulted in a decrease in its IRES activity (Figure 4I). Furthermore, to test whether the SuRE element is position-sensitive, we rearranged the SuRE element from the 40-60 nt position to the 90-110 nt position by swapping the sequences of these two regions on the IRES. A decrease in IRES translation activity was observed (Figures 4D and 4I). To further verify that SuRE is structure-dependent rather than sequence-dependent, we performed compensatory mutagenesis of the SuRE element. Specifically, we disrupted its double-strand structure by mutating each of the seven base pairs on the stem region of the SuRE element. A decrease in IRES translation activity was observed (Figures 4E and 4I). Furthermore, the translational activity of IRES can be rescued by compensatory double complementary mutations, allowing each of the seven base pairs on the stem region to be restored (Figures 4F and 4I). Interestingly, when SuRE was substituted in MS2 or BoxB, which have similar RNA structures, the same IRES activity as wild-type IRES was observed (Figures 4G and 4I), suggesting that the IRES activity regulated by SuRE is actually structure-dependent, not sequence-dependent. Finally, linear IRESs were converted to circular IRESs by transplanting SuRE at a position of 40-60 nt on a linear IRES (Figures 4H and 4I). The above results suggest that 40-60 nt of SuRE on an IRES can indeed promote circRNA translation.

[0157] In summary, the above results, along with 18S rRNA profiling, suggest that two key regulatory elements on circRNA IRESs—18S rRNA complementarity and 40-60 nt SuREs on IRESs—can promote cap-independent translation on circRNAs. Consistent with this model, of the 17,201 eGFP(+) oligos captured by screening, 12,091 (approximately 70%) possessed high 18S rRNA complementarity (18S rRNA complementarity(+)) or 40-60 nt SuREs (SuRE(+)) (Figure 4J), suggesting that these two regulatory elements can promote translation of exogenous reporter circRNAs. To further investigate whether these two regulatory elements can also promote endogenous circRNA translation, we examined polysome-associated circRNAs (translated circRNAs) captured in HEK-293 cells (Ragan et al., 2019). We found that 123 out of 165 endogenous translated circRNAs (approximately 75%) were either 18S rRNA complementarity (+) or 40-60 nt SuRE (+) circRNAs (Figure 4J), indicating that these two regulatory elements are common features among endogenous translated circRNAs. These results suggest that 18S rRNA complementarity and 40-60 nt SuRE can promote translation of both exogenous reporter circRNAs and endogenous circRNAs. Nevertheless, no preferential localization of the 18S complementary sequence to the 5' or 3' of the SuRE was observed (Figure 13C), suggesting that the SuRE on the IRES may cause a disruption in RNA unwinding, and that the 18S complementary sequence on the IRES may increase the likelihood of interaction with 18S 25 rRNA on the ribosome and promote cap-independent translation on the circRNA (Figure 4K).

[0158] [Example 6] This example demonstrates that the IRES element promotes the initiation of translation of endogenous circRNA.

[0159] To investigate whether key regulatory elements identified on IRESs, such as the 18S rRNA complementary sequence and SuRE at the 40–60 nt position, can promote translation of human endogenous circRNA, locked nucleic acids (LNAs) were used to disrupt these key elements on IRESs. This is because LNAs are used to specifically disrupt functional regions on HCV IRESs and inhibit HCV IRES activity. Antisense LNAs were designed to target (i) the 18S rRNA complementary sequence on IRESs (LNA-18S) to block 18S rRNA binding to IRESs, (ii) SuRE at the 40–60 nt position (LNA-SuRE) to disrupt SuREs on IRESs, and (iii) random positions downstream of LNA-18S or LNA-SuRE on IRESs for identified IRESs (LNA-Rnd) (Figure 5A). Next, LNAs were co-transfected with oligo-split-eGFP-circRNA reporter constructs containing corresponding IRESs, and the translational activity of the eGFP reporters was measured by their normalized fluorescence signal intensity. It was found that co-transfection with LNA-18S or LNA-SuRE could indeed disrupt the cap-independent translational activity of all IRESs (10 out of 10 LNAs), but most co-transfections with LNA-Rnd did not affect the translational activity of the IRESs (4 out of 5 LNAs) (Figure 5B). This result was also confirmed to be unconfounded by changes in circRNA expression levels, as co-transfection with LNA generally did not alter circRNA expression levels (Figure 15A). This result suggests that disrupting key elements on IRESs using LNA may affect the cap-independent translational activity of exogenous reporter circRNAs.

[0160] To further investigate whether the identified key regulatory elements on IRES can also promote the translation of human endogenous circRNA, cells were transfected with the corresponding antisense LNA, and translated circRNA was quantified by the QTI method. Specifically, to isolate translated RNA, LNA-transfected cells were treated with lactimidomycin (LTM), followed by puromycin (PMY), and ribosome-associated RNA was precipitated using a sucrose cushion to purify the translated RNA (Figure 5C). Next, the level of LNA-mediated translation of endogenous circRNA containing the targeted IRES was quantified by qRT-PCR using branching primers extending to the backsplicing junction of the circRNA. Disruption of key regulatory elements on the IRES of endogenous circRNA by LNA-18S or LNA-SuRE generally led to a decrease in circRNA translation activity (8 out of 10 LNAs), but all LNA-Rnds did not affect the translation activity of endogenous circRNA (5 out of 5 LNAs) (Figure 5D). It was also confirmed that LNA transfection did not alter the expression level of endogenous circRNA (Figure 15B). Since the QTI method specifically captured RNA at the translation initiation stage, the decrease in endogenous circRNA translation observed during LNA transfection was suggested to be due to a decrease in translation initiation. These results were further validated by quantifying the protein levels produced from endogenous circRNA by Western blotting. Western blotting results were consistent with those observed in QTI-qRT-PCR, namely, the general decrease in protein levels produced from circRNA when key regulatory elements on the IRES of endogenous circRNA were disrupted (3 out of 4 LNAs) (Figure 5E). These results suggest that identified key elements on the IRES, such as the 18S rRNA complementary sequence and the SuRE at the 40–60 nt position, are important for promoting translation initiation of endogenous circRNA.

[0161] [Example 7] This example describes the identification of circRNAs capable of encoding endogenous proteins. The introduction of synthetic IRESs onto circRNAs is sufficient to initiate cap-independent translation, suggesting that endogenous circRNAs with active circular IRESs may have the potential to generate proteins via cap-independent translation. Therefore, to determine the potential circRNA proteome, eGFP(+) oligo sequences were captured in a screening of the human circRNA database (circBase) (Glazar et al., RNA 20, 1666-1670 (2014)) to identify endogenous circRNAs with active IRESs. Data were gated for false positives by considering only circRNAs annotated by two different circRNA prediction algorithms and by including only circRNAs with high mapping scores in this analysis. These results suggested that a high proportion of endogenous circRNAs potentially encode proteins: of 2,052 endogenous circRNAs containing oligo sequences from the synthetic library used for the screening assay, 979 circRNAs (approximately 48%) contained one or more eGFP(+) oligo sequences (IRES(+) circRNAs) (Figure 6A, Figure 16A, Table 6). These circRNAs were generated from various parental genes that showed a fairly uniform distribution across the genome (Gini coefficient = 0.38) (Figure 6B). To further determine whether these IRES(+) circRNAs are associated with cancer progression, we examined the cancer-specific circRNA database (CSCD) (Xia et al., Nucleic Acids Res 46, D925-D929 (2018)), which contains a collection of potentially cancer-associated circRNAs, by analyzing RNA-seq data from 228 cancer and normal cell lines. Interestingly, 294 out of 979 IRES(+)circRNAs (approximately 30%) were found to be specifically expressed in either non-transformed cell lines (n=141 cell lines) or cancer cell lines (n=87 cell lines across 19 cancer types) (Figure 6C), suggesting their potential association with cancer progression.

[0162] Most IRES(+) circRNAs contained only one IRES (Figure 6D), and most eGFP(+) oligos mapped to only one circRNA (Figure 6E), suggesting a specific one-to-one relationship between these IRES(+) circRNAs and the proteins they encode. This result is partially predicted based on library design. In addition, a dominance of one IRES per circRNA was observed in 159 transcripts where oligotiling was designed across the entire transcript (Figures 16B and 16C). Thus, circRNA IRESs were difficult, if not impossible, to discover by comparative sequence analysis of the entire circRNA, but could be discovered by unbiased functional screening. This result also suggests that circRNA IRES activity may require longer RNA sequences that are likely to appear once per transcript, rather than very short or repetitive sequences that appear multiple times per transcript. Furthermore, while the most frequent location of mapped eGFP(+) oligos on circRNA was found near the backsplicing junction of the circRNA (within 100–200 nt of the junction), the average location of GC-matched oligos (174 nt of oligo on circRNA mapped by IRES with the same GC content as the mapped eGFP(+) oligos) showed a random distribution over a wide distance from the backsplicing junction on the circRNA (100–2000 nt of the junction) (Figure 16D). This result suggests that the cap-independent translational activity of IRES on circRNA is backsplicing-dependent, meaning that the IRES element or its downstream open reading frame (ORF) is assembled only during backsplicing. This requires the IRES to be positioned near the junction to promote its cap-independent translational activity. Finally, gene ontology (GO) analysis of the parental genes of these circRNAs suggested that they are enriched in stress response and translational regulation (Figure 6F).In particular, these results demonstrate that it is possible to identify endogenous circRNAs with potential cap-independent translational activity that can encode novel protein isoforms using the identified eGFP(+) oligo sequences.

[0163] [Example 8] This example illustrates the identification of polypeptides encoded by potential endogenous circRNAs.

[0164] To determine the polypeptide sequence of RNA-encoded proteins, we defined the protein-coding sequence on the RNA. This is typically achieved by ORF analysis. However, ORF analysis on circRNA often returns numerous results because it lacks information about where translation begins on the circRNA. The data presented here allows us to map the location of eGFP(+) oligo sequences on circRNA, which enables the determination of regions on the circRNA where the translation initiation site may be located. Therefore, to determine the potential polypeptide sequences of proteins encoded by endogenous circRNAs, we mapped the eGFP(+) oligo sequences to the respective sequences of individual high-reliability circRNAs within circBase (pregating the high-reliability circRNAs as described above) to determine the IRES location on each circRNA (Figure 6G). Next, we generated the predicted polypeptide sequences of proteins encoded by each circRNA by performing ORF analysis from the translation initiation codon (AUG) immediately downstream of the mapped IRES location (Figure 6G). Since many IRESs have been reported to be able to initiate translation from non-canonical start codons, ORF analysis was also performed on the top three frames (+1 to +3) using non-canonical start codons from mapped IRES locations (Figure 6G). This method generated a list of predicted polypeptide sequences encoded by human endogenous circRNAs (circORFs). For conservation purposes, micropeptides encoded by linear RNA were also examined, and any duplicate circORFs were excluded from the final list (n=5 duplicate circORFs). The final list contains 958 potential circORFs encoded by endogenous circRNAs (Figure 6G, SEQ ID NOs: 32954, -33911, Tables 7A and 7B).

[0165] By analyzing circORF sequences and mapped IRES locations on circRNAs, we discovered that several circRNAs contain IRES sequences that overlap with the translational region of the ORF (n=457; approximately 48%) (Figure 16E). IRES-overlapping ORFs have been observed in proteins encoded by several endogenous circRNAs, suggesting that several regulatory mechanisms may exist between the initiation and elongation of circRNA translation. Interestingly, among these circRNAs with IRES-overlapping ORFs, some contain in-frame ORFs without stop codons (n=82; approximately 18%), forming recurrent ORFs that may be a mechanism for amplifying the expression level of circRNA-encoded proteins (Figure 16F). It was further demonstrated that in-frame IRESs can indeed generate recurrent ORFs on the eGFP circRNA reporter (Figure 16G).

[0166] Next, the general functions of these potential circORFs were characterized by searching for conserved motifs in the predicted polypeptide sequences. Pfam analysis revealed that a significant number of circORFs contain conserved motifs. The top motifs were DNA-binding motifs, translation elongation factor-binding motifs, protein kinase domains, and protein dimerization domains (Figure 6H), suggesting that circORFs may play a role in regulating various biological functions, including signaling, transcription, and translation. Most of these potential circORFs are small in size (less than 100 amino acids) (Figure 16H), suggesting that the majority of them may be cleaved forms of proteins generated from parental linear transcripts.

[0167] To further investigate potential circORFs, we examined a short open reading frame (sORF) database (Olexiouk et al., Nucleic Acids Research 46, D497-D502 (2017)), which contains polypeptide sequences (less than 100 amino acids) from identified sORFs aggregated by multiple ribosome profiling studies to determine whether the polypeptide sequences of these sORFs could match circORFs. First, we mapped the sORFs to the current proteome database (UniProt) and excluded sORFs that were a perfect match to the ORFs of annotated linear transcripts. Next, we mapped the remaining sORFs to potential circORFs. We identified 317 predicted circORFs, which could match sORFs (approximately 33%) (Figure 16I), suggesting that the mapped IRES ORF analysis method can efficiently identify endogenous circORFs. On the other hand, conventional ORF analysis of the same circRNA, taking all possible translation initiation sites, yielded a vast number of predicted polypeptides (n=426,439), and only a small fraction of these polypeptides were captured by sORF studies (n=9,970; approximately 2%) (Figure 16I). Therefore, insights from circRNA IRES resulted in approximately a 15-fold improvement in predicting circRNA-derived sORFs. In summary, these results suggest that mapped IRES ORF analysis can identify endogenous circORFs more efficiently than conventional ORF analysis.

[0168] Next, peptide mix analysis was performed on tandem mass spectrometry (MS / MS) datasets to verify the endogenous expression of circORF. Specifically, the predicted list of circORF was added to the current proteome database (UniProt; linear proteome) to generate a combined proteome database (circORF + linear proteome). Then, raw MS / MS data were obtained from a wide range of cell lines, including K562, H358, U2OS, intracellular compartment SubCellBarCode (SCBC) databases, and 32 normal human tissues from the GTEx collection, and peptide-spectral matching (PSM) was performed against the combined proteome database (Figure 6I). To distinguish circORF from the linear proteome, circORF was excluded if it matched a trypsin polypeptide that could also match the linear proteome (Figure 6I). MS captured 118 circORFs possessing unique trypsin polypeptides (Figure 6J), of which 22 circORFs had MS-matched trypsin polypeptides extending to the circRNA backsplicing junction (BSJ) (Figure 6K). In addition to transformed cell lines, circORFs were captured in peptidomics of normal human tissue (Jiang et al., 2020), suggesting that these circORFs are expressed in normal human cells. Furthermore, concomitant reaction monitoring-MS (PRM-MS) was performed to obtain high-resolution verification of circORF expression in K562 and U2OS. Specifically, heavy isotope-labeled reference polypeptides for the unique regions of circORFs identified from K562 and U2OS MS / MS peptidomics were designed and synthesized. The labeled reference polypeptides were spiked into trypsin polypeptide samples, and precursor and transition ion detection was performed according to the labeled reference polypeptides (Figure 6L, SEQ ID NOs: 32954-33911, and Tables 7A and 7B). PRM-MS further verifies the presence of six of the eight targeted circORFs (Figure 6L, SEQ ID NOs: 32954-33911, and Tables 7A and 7B).MS / MS and PRM-MS peptide mixes provide strong evidence demonstrating that circORF is indeed endogenously expressed. As a supplementary approach, ribosome footprinting (RFP) data were examined in human iPSCs (Chen et al., Science (2020) 367(6482):1140~1146), and it was found that circORF detected by seven MS / MS assays contained at least one RFP fragment that uniquely overlapped with a circRNA backsplicing junction (Figure 16J). In summary, these results suggest that a putative circORF list can be constructed using circRNA IRES screening assays that can be validated by genomic and peptide analysis.

[0169] To further investigate whether circORFs are involved in antigen presentation, we analyzed human leukocyte antigen I (HLA1)-related peptidomics (Bassani-Sternberg et al., 2015). Two HLA1-related circORFs were identified (Figure 6J). In silico HLA1 binding predictor NetMHC4.1 analysis (Reynisson et al., 2020) suggests that these two circORFs are indeed potent HLA1 binding agents for HLA1 variants expressed in cell lines used in HLA1 peptidomics (HLA-A03:01 for circORF_674 in fibroblasts; HLA-C07:02 for circORF_917 in JY cells) (Tables 7A and 7B). This result indicates a novel functional role of circORFs, suggesting that some circORFs may enter the HLA-I presentation pathway and contribute to the antigen repertoire.

[0170] In particular, circORF detection by MS-based peptidomics is limited by (i) its insufficient ability to capture small amounts of circORF resulting from generally low circRNA expression levels, (ii) the potential instability of polypeptides encoded by circRNA, (iii) the inherent difficulty of detecting generally short circORFs, (iv) the number of available cell lines / types of peptidomics datasets, and (v) the narrow reference space of circORF-specific polypeptides, as all regions shared by circORF and linear proteomes are excluded. Given the limitations of circORF peptidomics, circORF identification should be interpreted as positive validation, and the absence of detection in MS proteomics data does not rule out the possibility of translation for candidate circRNAs. Consistent with the above limitations, it was found that when the same limitations were applied to proteins encoded by known mRNAs, with matching expression levels and cell lines examined, and the reference space downsampled, current peptidomics data could only recover about 5% of polypeptides of proteins encoded by mRNAs having the same RPKM as the mean circRNA RPKM (Figures 16K and 16L). Furthermore, since only the specific regions of circORFs are searched, the expected discovery rate of circORFs is estimated to be about 4%. The fact that approximately 12.3% (118 out of 958) of circORFs can be validated using peptidomics is far higher than the expected discovery rate for circORFs, further highlighting that the approach disclosed herein can efficiently identify candidate endogenous circORFs and supporting the claim that circRNAs broadly encode polypeptides similar to moderately expressed mRNAs.

[0171] [Example 9] This example demonstrates that circFGFR1p suppresses cell proliferation under stress conditions through dominant-negative regulation.

[0172] To evaluate the potential functions of the expanded circRNA proteome, we selected hsa_circ_0084007, an example of a potential protein-coding circRNA, and further investigated the function of its coding protein. Since circRNAs are generated from backsplicing of exons 2 and 7 of the human fibroblast growth factor receptor 1 (FGFR1) transcript, the names circFGFR1 and circFGFR1p were used to refer to this circRNA and its coding protein, respectively. Downregulation of circFGFR1 has been observed in samples from clinical cancer patients, which may suggest its role in regulating important biological processes. CircFGFR1 has an IRES located in the 5'UTR region of FGFR1, which shows strong eGFP expression in the screening assay (top 2%), followed immediately by an annotated AUG translation start codon (Figure 7A). ORF analysis using immediately downstream AUG revealed that backsplicing generates a de novo stop codon within the circFGFR1 IRES, resulting in an ORF (circORF_949) that partially overlaps with the IRES (Figure 7A). To better characterize the phenotype and function regulated by circFGFR1, we used BJ fibroblasts, a non-transformed human cell line, for subsequent analysis. This cell line has a diploid genome for good phenotypic analysis and high FGFR1 expression. First, we confirmed whether circFGFR1 expression could be detected in BJ cells by reverse transcriptase PCR (RT-PCR) and Sanger sequencing using branched primers adjacent to the backsplicing junctions of exons 2 and 7 on circFGFR1 (Figure 7B). This result demonstrated successful detection of circFGFR1 expression in BJ cells.

[0173] Analysis of the predicted protein sequence showed that circFGFR1p encodes a cleaved form of FGFR1, which possesses the intact extracellular fibroblast growth factor 1 (FGF1) ligand binding site and a portion of the dimerization domain (partial N' terminus of IgI, IgII, and IgIII), but lacks the intracellular FGFR1 tyrosine kinase domain (Figure 7C). CircFGFR1p also possesses a unique region resulting from backsplicing of circFGFR1, the polypeptide sequence of which is not found in the linear proteome (UniProt) database (Figure 7C). Western blotting using antibodies against the common region of circFGFR1p and FGFR1 (both Ab and Ab) showed signals of corresponding sizes for circFGFR1p (approximately 38 kDa) and FGFR1 (70–90 kDa) (Figure 7K). ENCODE data demonstrated the absence of a chromatin signature for the promoter (H3K4me3) near the circFGFR1p IRES (Figure 17A), suggesting that protein was not produced from the cleaved linear transcript due to a hidden promoter located in exon 2 of FGFR1. Consistent with the above observation, the identified circFGFR1 IRES (oligo index: 8228) did not show promoter activity from linear RNA IRES reporter screening (score = 0) (Weingarten-Gabbay et al., 2016).

[0174] To verify endogenous circFGFR1p expression, a custom antibody against a specific region of circFGFR1p was generated. circFGFR1p was isolated by immunoprecipitation (IP) using the custom antibody, and the protein (selected on a polyacrylamide gel with a size of approximately 30–45 kDa to separate circFGFR1p from FGFR1) was subjected to liquid chromatography with tandem mass spectrometry (LC-MS / MS) (Figure 7D). While circFGFR1p polypeptide was not detected in the IgG control sample, the trypsin polypeptide of the specific region of circFGFR1p, and the trypsin polypeptide overlapping with linear FGFR1, were detectable in the IP-LC-MS / MS sample (Figure 7D). This result suggests that circFGFR1p is indeed expressed and can be captured by the circFGFR1p antibody. To further confirm circFGFR1p expression at high resolution, PRM-MS was performed using a synthetic heavy isotope-labeled reference polypeptide of the specific region of circFGFR1p. We were able to identify the corresponding precursors and transition ions of the labeled reference polypeptide, as well as the sampled trypsin polypeptides from BJ cells (Figure 7E). In summary, IP-MS and PRM-MS provide strong evidence demonstrating endogenous circFGFR1p expression.

[0175] Upon binding to FGF, full-length FGFR1 dimerizes and autophosphorylates its kinase domain, further triggering downstream signaling pathways and promoting cell proliferation. By co-expressing HA-tagged FGFR1 and FLAG-tagged circFGFR1p in HEK-293T cells and co-staining the HA and FLAG tags to label FGFR1 and circFGFR1p, respectively, it was confirmed that circFGFR1p, like FGFR1, is localized to the cell membrane in patch domains and endosomes (Figures 7F and 17B). CircFGFR1 contains an FGFR1 dimerization / ligand-binding domain but lacks a kinase domain, suggesting that circFGFR1p may function as a dominant-negative regulator of FGFR1, suppressing cell proliferation. Furthermore, lower circFGFR1 expression levels were found in tumor samples of different breast cancer subtypes compared to normal adjacent samples from studies analyzing RNA sequencing data from The Cancer Genome Atlas (TCGA) (Figure 17C). In addition, circFGFR1 expression was specifically detected in non-transformed cell lines (n=5 specific non-transformed cell types out of 141 non-transformed cell line samples) but not in cancer cell lines from the CSCD database (n=87 cancer cell lines out of 87 cancer cell line samples) (Xia et al., 2018) (Figure 17D). These studies suggest that reduced circFGFR1 expression levels may be associated with cancer progression by upregulating cell proliferation. Therefore, circFGFR1p appears to function as a negative regulator of FGFR1 through a dominant-negative mechanism that suppresses cell proliferation.

[0176] To test this hypothesis, we first specifically knocked down circFGFR1 using siRNA targeting the backsplicing junction of circFGFR1 (Figure 7G). We found that knockdown of circFGFR1 could indeed promote cell proliferation upon FGF addition (Figure 7H), suggesting that circFGFR1 negatively modulates the FGFR1 function that promotes cell proliferation. To confirm that the observed cell proliferation phenotype was due to the downregulation of the circFGFR1p protein rather than circFGFR1 RNA, we further investigated the cell proliferation phenotype when we specifically downregulated the circFGFR1p protein by disrupting cap-independent translation of circFGFR1p IRES. Since translation initiation is typically the rate-limiting step in translation, we utilized an antisense LNA targeting the 18S rRNA complementary sequence on the circFGFR1 IRES (LNA-18S of IRES-8228). This was found to effectively block circFGFR1p translation initiation (Figures 5B and 5D), specifically knocking down circFGFR1p without altering the levels of circFGFR1 or FGFR1 RNA (Figure 7G). Inhibition of the circFGFR1 IRES mediated by the LNA resulted in decreased circFGFR1p expression levels and increased phosphorylated FGFR1 levels (Figure 7G), suggesting that knocking down circFGFR1p rather than circFGFR1 RNA may indeed lead to increased FGFR1 phosphorylation and higher levels of active FGFR1 (phosphorylated FGFR1). This is also consistent with the observation that knocking down circFGFR1p results in higher cell proliferation rates (Figure 7H). Interestingly, depletion of circFGFR1p is also accompanied by higher levels of total FGFR1 protein (Figure 7G). This result suggests that circFGFR1p not only functions as a dominant-negative in FGFR1 signaling, but also inhibits the accumulation of full-length FGFR1 in some way, possibly by increasing FGFR1 turnover or degradation.A similar FGFR1 degradation phenotype has also been observed when a dominant-negative variant of FGFR1 is expressed in vivo. Conversely, we investigated whether intracellular overexpression of circFGFR1 could suppress cell proliferation by encoding circFGFR1p, which has a FLAG epitope tag. Next, we cloned it into a linear mRNA expression plasmid driven by the CMV promoter to effectively overexpress circFGFR1p, and transfected BJ cells with the circFGFR1p expression plasmid (Figure 7I). This result demonstrated that overexpression of circFGFR1p (circFGFR1pOE) can indeed suppress cell proliferation (Figure 7I). In addition, simultaneous intracellular overexpression of FGFR1 and circFGFR1p (FGFR1OE + circFGFR1pOE) partially rescued the phenotype of the cell proliferation suspension (Figure 7I), further suggesting the antagonistic function of circFGFR1p against FGFR1. These results suggest that circFGFR1p, encoded by circRNA, can inhibit cell growth by interacting with FGFR1 via a dominant-negative mechanism.

[0177] Compared to FGFR1, circFGFR1p expression levels were relatively low (Figures 7J and 7K), suggesting that circFGFR1p may not be a potent regulator under normal conditions. Nevertheless, many IRESs, including those of several endogenous protein-coding circRNAs such as circZNF-609, have been reported to possess stable cap-independent translational activity under stress conditions. Therefore, we further investigated the cap-independent translational activity of circFGFR1 IRESs under stress conditions such as heat shock. First, we transfected cells with an oligo-split-eGFP-circRNA reporter construct driven by circFGFR1 IRESs and quantified circFGFR1 IRES activity with and without heat shock. The results demonstrated that the cap-independent translational activity of 15 FGFR1 IRESs remained stable during heat shock (Figure 17E). Next, we examined FGFR1 and circFGFR1p protein levels under heat shock conditions. FGFR1 protein levels were observed to be downregulated after heat shock (Figures 7J and 7L), which may be due to an overall decrease in cap-dependent translation caused by changes in the phosphorylation status of many eukaryotic initiation factors and the sequestration of eIF4G by Hsc70 during heat shock. On the other hand, circFGFR1p levels, regulated by cap-independent translation, remained stable after heat shock (Figures 7J, 7L, and 17F and 17G). These results suggest that the overall decrease in FGFR1 cap-dependent translation during heat shock is not directly caused by circFGFR1p levels, but that the decreased FGFR1 levels and stable circFGFR1p levels enhance the circFGFR1p dominant-negative effect, further reducing cell growth rate. Furthermore, FGFR1 has been shown to form homomultimers when induced by cell adhesion molecules. The oligomeric nature of FGFR1 may further enhance the dominant-negative effect of circFGFR1p.This is because one circFGFR1p can conjugate and "impair" signaling ability, or lead to the degradation of more FGFR1 molecules. These phenomena can explain how low-expression circFGFR1p effectively modulates high-expression FGFR1 and suppresses cell proliferation under heat shock or other forms of cellular stress (Figures 17H and 17I).

[0178] Interestingly, the circFGFR1 IRES (oligoindex: 8228) showed strong cap-independent translational activity towards circRNA (top 2%), while the same IRES showed very weak cap-independent translational activity towards linear RNA (bottom 10%) (Weingarten-Gabbay et al., 2016). This observation suggests that the cap-independent translational activity of the circFGFR1 IRES is preferentially activated on circFGFR1 rather than linear FGFR1 transcripts. This biased IRES activity of the circRNA-5 of the circFGFR1 IRES also explains how, under heat shock conditions, the circFGFR1 IRES can selectively generate a stable amount of circFGFR1p without increasing the level of FGFR1 protein isoforms generated from the cap-independent activity of the circFGFR1 IRES on linear FGFR1 transcripts, and how circFGFR1p can more effectively regulate FGFR1 function. In summary, the findings presented above demonstrate that the disclosed method failed to identify circFGFR1p, a novel circRNA-encoded protein that negatively regulates FGFR1 and suppresses cell proliferation via a dominant-negative mechanism under stress conditions. This study also reveals important regulatory mechanisms of circRNAs and their coded proteins.

[0179] Various embodiments of the present invention, including the best mode known to the inventors for carrying out the invention, are described herein. Variations of these embodiments may become apparent to those skilled in the art by reading the preceding description. The inventors expect that those skilled in the art will adopt such variations as appropriate, and the inventors intend that the invention may be carried out in ways other than those specifically described herein. Accordingly, the invention includes all modifications and equivalents of the subject matter described in the claims appended herein, as permitted by applicable law. Furthermore, unless otherwise indicated herein, or unless clearly inconsistent with the context, any combination of all possible variations of the above elements is incorporated into the invention.

[0180] Embedding by reference All references cited herein, including publications, patent applications, and patents, are thus incorporated by reference to the same extent as each reference is individually and clearly indicated as being incorporated by reference and is shown in whole herein.

[0181] Numbered Embodiments Notwithstanding the attached claims, the following numbered embodiments also form part of the present disclosure. 1. A polynucleotide sequence encoding a circular RNA molecule, wherein the circular RNA molecule comprises a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence, the IRES sequence region comprising at least one sequence region having an RNA secondary structure element; and a sequence region complementary to 18S ribosomal RNA (rRNA), the IRES sequence region having a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C; and the RNA secondary structure element is formed from nucleotides approximately 40 to 60 of the IRES, where the first nucleic acid at the 5' end of the IRES is considered to be at position 1. 2. The polynucleotide sequence according to Embodiment 1, wherein a protein-coding nucleic acid sequence region is operably linked to an IRES sequence region in a non-natural configuration. 3. A polynucleotide sequence according to Embodiment 1 or 2, which is a DNA sequence. 4. A polynucleotide sequence according to any one of Embodiments 1 to 3, wherein the sequence complementary to 4.18S rRNA is one of sequence numbers 28977 to 28983. 5. A polynucleotide sequence according to any one of embodiments 1 to 4, wherein at least one RNA secondary structure element is located at the 5' end of a sequence region complementary to 18S rRNA. 6. A polynucleotide sequence according to any one of embodiments 1 to 4, wherein at least one RNA secondary structure element is located at the 3' end of a sequence region complementary to 18S rRNA. 7. A polynucleotide sequence according to any one of embodiments 1 to 6, wherein at least one RNA secondary structure element is a stem-loop. 8. A polynucleotide sequence according to any one of Embodiments 1 to 7, wherein at least one RNA secondary structure element comprises one of the nucleic acid sequences listed in Table 2. 9. A polynucleotide sequence according to any one of Embodiments 1 to 8, wherein the IRES sequence region is approximately 100 to approximately 1000 nucleotides long. 10. A polynucleotide sequence according to any one of Embodiments 1 to 8, wherein the IRES sequence region is approximately 200 to approximately 800 nucleotides in length. 11. A polynucleotide sequence according to any one of Embodiments 1 to 8, wherein the IRES sequence is 150-200 nucleotides, 160-180 nucleotides, or 200-210 nucleotides in length. 12. A polynucleotide sequence according to any one of embodiments 1 to 11, comprising at least one non-coding functional sequence. 13. The polynucleotide sequence according to Embodiment 12, wherein the non-coding functional sequence includes one or more (a) microRNA binding sites or (b) RNA-binding protein binding sites. 14. A polynucleotide sequence according to any one of Embodiments 1 to 11, wherein the DNA sequence contains an aptamer. 15. A recombinant circular RNA molecule encoded by a polynucleotide sequence described in any one of Embodiments 1 to 14. 16. A DNA sequence encoding a circular RNA molecule, wherein the circular RNA molecule comprises a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably ligated to the protein-coding nucleic acid sequence, and the IRES sequence region comprises one of the nucleic acid sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a nucleic acid sequence having at least 90% or at least 95% identity or homology thereto. 17. The DNA sequence according to Embodiment 16, wherein a protein-coding nucleic acid sequence is operably linked to an IRES sequence region in a non-natural configuration. 18. A DNA sequence according to any one of embodiments 16 to 17, wherein the IRES sequence region has a GC content of at least 25%. 19. The DNA sequence according to any one of embodiments 16 to 18, wherein the IRES sequence region includes one of the nucleic acid sequences of sequence numbers 1 to 228. 20. The DNA sequence according to any one of embodiments 16 to 18, wherein the IRES sequence region includes one of the nucleic acid sequences of sequence numbers 229 to 17201. 21. The DNA sequence according to any one of Embodiments 16 to 20, wherein the IRES sequence region includes one nucleic acid sequence from among Sequence ID No. 531, 2270, 2602, 3042, 3244, and 33948. 22. A DNA sequence according to any one of embodiments 16 to 21, wherein the IRES sequence region includes a human IRES. 23. A DNA sequence according to any one of embodiments 16 to 22, wherein the protein-coding nucleic acid sequence region encodes a therapeutic peptide or protein. 24. A DNA sequence according to any one of Embodiments 16 to 23, wherein the circular RNA molecule contains approximately 200 to 10,000 nucleotides. 25. The DNA sequence according to any one of embodiments 13 to 24, wherein the circular RNA molecule includes a spacer between the IRES sequence region and the start codon of the protein-coding nucleic acid sequence region. 26. The DNA sequence according to Embodiment 25, wherein the length of the spacer is selected to increase the translation of the protein-coding nucleic acid sequence region compared to the translation of a circular RNA without a spacer or with a spacer different from the selected spacer. 27. A DNA sequence according to any one of embodiments 16 to 26, wherein the IRES sequence region is configured to facilitate rolling circle translation. 28. A DNA sequence according to any one of embodiments 16 to 26, wherein the protein-coding nucleic acid sequence region lacks a stop codon. 29. A DNA sequence according to any one of Embodiments 16 to 26, wherein (i) the IRES sequence region is configured to facilitate rolling circle translation, and (ii) the protein-coding nucleic acid sequence region lacks a stop codon. 30. A recombinant circular RNA molecule encoded by the DNA sequence described in any one of Embodiments 16 to 29. 31. A viral vector comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 32. A viral vector according to Embodiment 31, selected from the group consisting of adeno-associated virus (AAV) vectors, adenovirus vectors, retrovirus vectors, lentivirus vectors, vaccinia, and herpesvirus vectors. 33. A viral vector according to embodiment 31 or 32, which is AAV. 34. The viral vector according to Embodiment 33, wherein the AAV serotype is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV8, AAV9, AAVrh10, or any variant thereof having substantially the same tropism. 35. Virus-like particles comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 36. Non-viral-like particles comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 37. A closed DNA sequence comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 38. A plasmid comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 39. A miniintron plasmid vector comprising a polynucleotide as described in any one of Embodiments 1 to 14 or a DNA sequence as described in any one of Embodiments 16 to 29. 40. A composition comprising a polynucleotide according to any one of Embodiments 1 to 14, a DNA sequence according to any one of Embodiments 16 to 29, or a recombinant circular RNA molecule according to Embodiment 15 or 30. 41. The composition according to Embodiment 40, wherein lipid nanoparticles are used to decorate it. 42. A host cell comprising a polynucleotide described in any one of Embodiments 1 to 14, a DNA sequence described in any one of Embodiments 16 to 29, or a recombinant circular RNA molecule described in Embodiment 15 or 30. 43. A method for producing a protein in cells, comprising the step of contacting cells with (a) a polynucleotide according to any one of Embodiments 1 to 14, (b) a DNA sequence according to any one of Embodiments 16 to 29, (c) a circular RNA molecule according to any one of Embodiments 15 or 30, (d) a viral vector according to any one of Embodiments 31 to 34, (e) a virus-like particle according to Embodiment 35, (f) a non-virus-like particle according to Embodiment 36, (g) a closed DNA sequence according to Embodiment 37, (h) a plasmid according to Embodiment 38, (i) a miniintron plasmid vector according to Embodiment 39, or (j) a composition according to any one of Embodiments 40 to 41, under conditions in which a protein-coding nucleic acid sequence of circular RNA is translated in cells to produce a protein. 44. The method according to embodiment 43, wherein 5' cap-dependent translation in the cell is impaired or absent. 45. The method according to embodiment 43 or 44, wherein the cell is in vivo. 46. The method according to embodiment 45, wherein the cell is a mammalian cell. 47. The method according to embodiment 46, wherein the mammalian cell is derived from a human. 48. The method according to any one of embodiments 45 to 46, wherein protein production is tissue-specific. 49. The method according to embodiment 48, wherein the tissue specificity is localized to a tissue selected from the group consisting of muscle, liver, kidney, brain, lung, skin, pancreas, blood, and heart. 50. The method according to embodiment 43 or 44, wherein the cell is in vitro. 51. The method according to any one of embodiments 43 to 48, wherein the protein is recursively expressed in the cell. 52. The method according to any one of embodiments 43 to 51, wherein the half-life of the circular RNA in the cell is from about 1 day to about 7 days. 53. The method according to any one of embodiments 43 to 55, wherein the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer than when the protein-coding nucleic acid sequence is provided to the cell as linear format RNA or encoded for transcription as linear RNA. 54. A protein produced by the method according to any one of embodiments 43 to 53. 55. A recombinant circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence, wherein the IRES comprises at least one RNA secondary structure; and a sequence complementary to 18S ribosomal RNA (rRNA), and the IRES has a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C. 56. The recombinant circular RNA molecule according to embodiment 55, wherein the protein-coding nucleic acid sequence is operably linked to the IRES in a non-native configuration. 57. The recombinant circular RNA molecule according to embodiment 55 or 56, wherein the sequence complementary to 18S rRNA is encoded by any one of SEQ ID NOs: 28977 to 28983. 58. The recombinant circular RNA molecule according to any one of embodiments 55 to 57, wherein at least one RNA secondary structure is located 5' to the sequence complementary to 18S rRNA. 59. The recombinant circular RNA molecule according to any one of embodiments 55 to 57, wherein at least one RNA secondary structure is located 3' to the sequence complementary to 18S rRNA. 60. The recombinant circular RNA molecule according to any one of embodiments 55 to 59, wherein the at least one RNA secondary structure is a stem-loop. 61. The recombinant circular RNA molecule according to any one of embodiments 55 to 59, wherein the at least one RNA secondary structure comprises a sequence encoded by any one of the DNA sequences listed in Table 2. 62. The recombinant circular RNA according to any one of embodiments 55 to 61, wherein the IRES is about 100 to about 1000 nucleotides in length. 63. The recombinant circular RNA according to any one of embodiments 55 to 61, wherein the IRES is about 200 to about 200 nucleotides in length. 64. The recombinant circular RNA according to any one of embodiments 55 to 61, wherein the IRES sequence is 150 to 200 nucleotides, 160 to 180 nucleotides, or 200 to 210 nucleotides in length. 65. The recombinant circular RNA according to any one of embodiments 62 to 64, wherein the RNA secondary structure is formed from about position 40 to about position 60 nucleotides relative to the 5' end of the IRES. 66. The recombinant circular RNA according to any one of embodiments 55 to 65, comprising at least one non-coding functional sequence. 67. The recombinant circular RNA according to embodiment 66, wherein the non-coding functional sequence comprises one or more of (a) a microRNA binding site or (b) an RNA binding protein binding site. 68. The recombinant circular RNA according to any one of embodiments 55 to 65, wherein the circular RNA comprises an aptamer. 69. A recombinant circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence, wherein the IRES is encoded by one of the DNA sequences listed in SEQ ID NOs: 1-228 or SEQ ID NOs: 229-17201, or a DNA sequence having at least 90% or at least 95% identity or homology thereto. 70. A recombinant circular RNA molecule according to Embodiment 69, wherein a protein-coding nucleic acid sequence is operably linked to an IRES in a non-natural configuration. 71. Recombinant circular RNA according to any one of embodiments 55 to 70, wherein the recombinant circular RNA includes a back splice junction and the IRES is located within approximately 100 to 200 nucleotides from the back splice junction. 72. A recombinant circular RNA molecule according to any one of embodiments 55 to 71, wherein the IRES has a GC content of at least 25%. 73. A recombinant circular RNA molecule according to any one of embodiments 55 to 72, wherein IRES is encoded by one of the DNA sequences of sequence numbers 1 to 228. 74. A recombinant circular RNA molecule according to any one of embodiments 55 to 72, wherein IRES is encoded by one of the DNA sequences of sequence numbers 229 to 17201. 75. A recombinant circular RNA molecule according to Embodiment 74, wherein IRES is encoded by one of the DNA sequences shown in SEQ ID NOs: 531, 2270, 2602, 3042, 3244, and 33948. 76. A recombinant circular RNA molecule according to any one of embodiments 55 to 75, wherein IRES is a human IRES. 77. A recombinant circular RNA molecule according to any one of embodiments 55 to 76, wherein the protein-coding nucleic acid sequence encodes a therapeutic peptide or protein. 78. A recombinant circular RNA molecule according to any one of embodiments 55 to 77, wherein the circular RNA comprises approximately 200 to 10,000 nucleotides. 79. A recombinant circular RNA molecule according to any one of embodiments 55 to 78, wherein the circular RNA molecule includes a spacer between the IRES sequence region and the start codon of the protein-coding nucleic acid sequence region. 80. A recombinant circular RNA molecule according to Embodiment 79, wherein the length of the spacer is selected to increase the translation of a protein-coding nucleic acid sequence region compared to the translation of a circular RNA without a spacer or with a spacer different from the selected spacer. 81. A recombinant circular RNA molecule according to any one of embodiments 55 to 78, wherein the IRES sequence region is configured to facilitate rolling circle translation. 82. A recombinant circular RNA molecule according to any one of embodiments 55 to 78, wherein the protein-coding nucleic acid sequence region lacks a stop codon. 83. A recombinant circular RNA molecule according to any one of Embodiments 55 to 78, wherein (i) the IRES sequence region is configured to facilitate rolling circle translation, and (ii) the protein-coding nucleic acid sequence region lacks a stop codon. 84. A composition comprising a recombinant circular RNA molecule as described in any one of embodiments 55 to 83. 85. A host cell comprising a recombinant circular RNA molecule according to any one of Embodiments 55 to 83 or the composition according to Embodiment 84. 86. A method for producing a protein in cells, comprising the step of contacting the cells with a recombinant circular RNA molecule according to any one of Embodiments 55 to 83, or a composition according to Embodiment 84, under conditions in which a protein-coding nucleic acid sequence is translated in the cells and a protein is produced. 87. The method according to embodiment 86, wherein 5' cap-dependent translation in cells is impaired or absent. 88. The method according to embodiment 86 or 87, wherein the cells are in vivo. 89. The method according to embodiment 86 or 87, wherein the cells are mammalian cells. 90. The method according to Embodiment 89, wherein the mammalian cells are derived from humans. 91. The method according to any one of embodiments 55 to 90, wherein protein production is tissue-specific. 92. The method according to Embodiment 91, wherein the tissue specificity is localized to a tissue selected from the group consisting of muscle, liver, kidney, brain, lung, skin, pancreas, blood, and heart. 93. The method according to Embodiment 86 or 87, wherein the cells are in vitro. 94. The method according to any one of embodiments 86 to 93, wherein the protein is recursively expressed in a cell. 95. The method according to any one of embodiments 86 to 94, wherein the half-life of the circular RNA in the cell is approximately 1 to approximately 7 days. 96. The method according to any one of embodiments 86 to 94, wherein the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer than if the protein-coding nucleic acid sequence were provided to the cell in linear format RNA or encoded for transcription as linear RNA. 97. A protein produced by the method described in any one of embodiments 86 to 96. 98. Oligonucleotide molecules containing nucleic acid sequences that hybridize to the internal ribosome entry site (IRES) present on circular RNA molecules and inhibit the translation of circular RNA molecules. 99. The oligonucleotide molecule according to Embodiment 98, wherein the circular RNA is recombinant circular RNA. 100. The oligonucleotide molecule according to Embodiment 98, wherein the recombinant circular RNA is the recombinant circular RNA described in any one of Embodiments 55 to 83. 101. The oligonucleotide molecule according to Embodiment 98, wherein the circular RNA is a naturally occurring circular RNA. 102. An oligonucleotide molecule according to any one of embodiments 98 to 101, wherein the oligonucleotide is an antisense oligonucleotide. 103. The oligonucleotide according to Embodiment 102, wherein the antisense oligonucleotide is a locked nucleic acid oligonucleotide (LNA). 104. The oligonucleotide according to any one of embodiments 98 to 103, wherein the oligonucleotide is chemically modified at its 5' and / or 3' end. 105. A method for inhibiting the translation of a protein-coding nucleic acid sequence present on a circular RNA molecule, comprising the step of contacting the circular RNA molecule with an oligonucleotide molecule described in any one of embodiments 98 to 104, thereby hybridizing the oligonucleotide molecule to a nucleic acid sequence complementary to the RNA secondary structure and / or 18S rRNA present on the IRES of the circular RNA molecule, thereby inhibiting the translation of the circular RNA molecule. 106. The method according to Embodiment 105, wherein the oligonucleotide hybridizes to a nucleic acid sequence complementary to an RNA secondary structure or 18S rRNA. 107. The method according to Embodiment 105, wherein the oligonucleotide hybridizes to a nucleic acid sequence complementary to the RNA secondary structure and 18S rRNA. 108. The method according to Embodiment 105, wherein a first oligonucleotide hybridizes to an RNA secondary structure, and a second oligonucleotide hybridizes to a nucleic acid sequence complementary to 18S rRNA.

[0182] [Table 4] TIFF2026143428000013.tif181170 TIFF2026143428000014.tif180170 TIFF2026143428000015.tif174170 TIFF2026143428000016.tif174170 TIFF2026143428000017.tif183170 TIFF2026143428000018.tif180170 TIFF2026143428000019.tif176170 TIFF2026143428000020.tif182170 TIFF2026143428000021.tif184170 TIFF2026143428000022.tif174170 TIFF2026143428000023.tif181170 TIFF2026143428000024.tif180170 TIFF2026143428000025.tif176170 TIFF2026143428000026.tif182170 TIFF2026143428000027.tif181170 TIFF2026143428000028.tif181170 TIFF2026143428000029.tif181170 TIFF2026143428000030.tif180170 TIFF2026143428000031.tif174170 TIFF2026143428000032.tif175170 TIFF2026143428000033.tif175170 TIFF2026143428000034.tif181170 TIFF2026143428000035.tif173170 TIFF2026143428000036.tif181170 TIFF2026143428000037.tif176170 TIFF2026143428000038.tif180170 TIFF2026143428000039.tif176170 TIFF2026143428000040.tif198170

[0183] Table 5 TIFF2026143428000042.tif197170 TIFF2026143428000043.tif214170 TIFF2026143428000044.tif214170 TIFF2026143428000045.tif213170 TIFF2026143428000046.tif214170 TIFF2026143428000047.tif214170 TIFF2026143428000048.tif214170 TIFF2026143428000049.tif198170 TIFF2026143428000050.tif215170 TIFF2026143428000051.tif214170 TIFF2026143428000052.tif238170 TIFF2026143428000053.tif228170 TIFF2026143428000054.tif222170 TIFF2026143428000055.tif223170 TIFF2026143428000056.tif230170 TIFF2026143428000057.tif238170 TIFF2026143428000058.tif224170 TIFF2026143428000059.tif238170 TIFF2026143428000060.tif238170 TIFF2026143428000061.tif229170 TIFF2026143428000062.tif231170 TIFF2026143428000063.tif223170 TIFF2026143428000064.tif229170 TIFF2026143428000065.tif237170 TIFF2026143428000066.tif223170 TIFF2026143428000067.tif231170 TIFF2026143428000068.tif238170 TIFF2026143428000069.tif229170 TIFF2026143428000070.tif224170 TIFF2026143428000071.tif223170 TIFF2026143428000072.tif223170 TIFF2026143428000073.tif223170 TIFF2026143428000074.tif222170 TIFF2026143428000075.tif220170 TIFF2026143428000076.tif223170 TIFF2026143428000077.tif231170 TIFF2026143428000078.tif223170 TIFF2026143428000079.tif221170 TIFF2026143428000080.tif223170 TIFF2026143428000081.tif222170 TIFF2026143428000082.tif228170 TIFF2026143428000083.tif221170 TIFF2026143428000084.tif221170 TIFF2026143428000085.tif224170 TIFF2026143428000086.tif230170 TIFF2026143428000087.tif228170 TIFF2026143428000088.tif223170 TIFF2026143428000089.tif223170 TIFF2026143428000090.tif230170 TIFF2026143428000091.tif216170 TIFF2026143428000092.tif223170 TIFF2026143428000093.tif230170 TIFF2026143428000094.tif230170 TIFF2026143428000095.tif228170 TIFF2026143428000096.tif222170 TIFF2026143428000097.tif223170 TIFF2026143428000098.tif223170 TIFF2026143428000099.tif222170 TIFF2026143428000100.tif224170 TIFF2026143428000101.tif223170 TIFF2026143428000102.tif221170 TIFF2026143428000103.tif224170 TIFF2026143428000104.tif222170 TIFF2026143428000105.tif223170 TIFF2026143428000106.tif222170 TIFF2026143428000107.tif222170 TIFF2026143428000108.tif231170 TIFF2026143428000109.tif224170 TIFF2026143428000110.tif223170 TIFF2026143428000111.tif222170 TIFF2026143428000112.tif221170 TIFF2026143428000113.tif222170 TIFF2026143428000114.tif221170 TIFF2026143428000115.tif222170 TIFF2026143428000116.tif229170 TIFF2026143428000117.tif224170 TIFF2026143428000118.tif221170 TIFF2026143428000119.tif228170 TIFF2026143428000120.tif229170 TIFF2026143428000121.tif231170 TIFF2026143428000122.tif223170 TIFF2026143428000123.tif222170 TIFF2026143428000124.tif223170 TIFF2026143428000125.tif230170 TIFF2026143428000126.tif221170 TIFF2026143428000127.tif237170 TIFF2026143428000128.tif221170 TIFF2026143428000129.tif223170 TIFF2026143428000130.tif229170 TIFF2026143428000131.tif222170 TIFF2026143428000132.tif229170 TIFF2026143428000133.tif222170 TIFF2026143428000134.tif221170 TIFF2026143428000135.tif232170 TIFF2026143428000136.tif232170 TIFF2026143428000137.tif224170 TIFF2026143428000138.tif237170 TIFF2026143428000139.tif223170 TIFF2026143428000140.tif222170 TIFF2026143428000141.tif230170 TIFF2026143428000142.tif223170 TIFF2026143428000143.tif229170 TIFF2026143428000144.tif223170 TIFF2026143428000145.tif223170 TIFF2026143428000146.tif221170 TIFF2026143428000147.tif222170 TIFF2026143428000148.tif230170 TIFF2026143428000149.tif229170 TIFF2026143428000150.tif230170 TIFF2026143428000151.tif222170 TIFF2026143428000152.tif229170 TIFF2026143428000153.tif224170 TIFF2026143428000154.tif222170 TIFF2026143428000155.tif223170 TIFF2026143428000156.tif231170 TIFF2026143428000157.tif221170 TIFF2026143428000158.tif222170 TIFF2026143428000159.tif229170 TIFF2026143428000160.tif222170 TIFF2026143428000161.tif229170 TIFF2026143428000162.tif223170 TIFF2026143428000163.tif222170 TIFF2026143428000164.tif231170 TIFF2026143428000165.tif231170 TIFF2026143428000166.tif221170 TIFF2026143428000167.tif221170 TIFF2026143428000168.tif237170 TIFF2026143428000169.tif228170 TIFF2026143428000170.tif224170 TIFF2026143428000171.tif224170 TIFF2026143428000172.tif223170 TIFF2026143428000173.tif229170 TIFF2026143428000174.tif223170 TIFF2026143428000175.tif229170 TIFF2026143428000176.tif231170 TIFF2026143428000177.tif229170 TIFF2026143428000178.tif238170 TIFF2026143428000179.tif221170 TIFF2026143428000180.tif223170 TIFF2026143428000181.tif222170 TIFF2026143428000182.tif237170 TIFF2026143428000183.tif222170 TIFF2026143428000184.tif237170 TIFF2026143428000185.tif223170 TIFF2026143428000186.tif224170 TIFF2026143428000187.tif222170 TIFF2026143428000188.tif224170 TIFF2026143428000189.tif232170 TIFF2026143428000190.tif222170 TIFF2026143428000191.tif230170 TIFF2026143428000192.tif229170 TIFF2026143428000193.tif223170 TIFF2026143428000194.tif229170 TIFF2026143428000195.tif223170 TIFF2026143428000196.tif232170 TIFF2026143428000197.tif223170 TIFF2026143428000198.tif224170 TIFF2026143428000199.tif231170 TIFF2026143428000200.tif221170 TIFF2026143428000201.tif224170 TIFF2026143428000202.tif232170 TIFF2026143428000203.tif232170 TIFF2026143428000204.tif222170 TIFF2026143428000205.tif232170 TIFF2026143428000206.tif223170 TIFF2026143428000207.tif238170 TIFF2026143428000208.tif229170 TIFF2026143428000209.tif223170 TIFF2026143428000210.tif230170 TIFF2026143428000211.tif224170 TIFF2026143428000212.tif232170 TIFF2026143428000213.tif232170 TIFF2026143428000214.tif231170 TIFF2026143428000215.tif229170 TIFF2026143428000216.tif230170 TIFF2026143428000217.tif223170 TIFF2026143428000218.tif231170 TIFF2026143428000219.tif224170 TIFF2026143428000220.tif223170 TIFF2026143428000221.tif221170 TIFF2026143428000222.tif232170 TIFF2026143428000223.tif238170 TIFF2026143428000224.tif238170 TIFF2026143428000225.tif230170 TIFF2026143428000226.tif229170 TIFF2026143428000227.tif222170 TIFF2026143428000228.tif229170 TIFF2026143428000229.tif221170 TIFF2026143428000230.tif238170 TIFF2026143428000231.tif221170 TIFF2026143428000232.tif229170 TIFF2026143428000233.tif224170 TIFF2026143428000234.tif231170 TIFF2026143428000235.tif232170 TIFF2026143428000236.tif230170 TIFF2026143428000237.tif237170 TIFF2026143428000238.tif221170 TIFF2026143428000239.tif229170 TIFF2026143428000240.tif230170 TIFF2026143428000241.tif229170 TIFF2026143428000242.tif229170 TIFF2026143428000243.tif230170 TIFF2026143428000244.tif237170 TIFF2026143428000245.tif231170 TIFF2026143428000246.tif224170 TIFF2026143428000247.tif232170 TIFF2026143428000248.tif229170 TIFF2026143428000249.tif230170 TIFF2026143428000250.tif229170 TIFF2026143428000251.tif230170 TIFF2026143428000252.tif230170 TIFF2026143428000253.tif237170 TIFF2026143428000254.tif224170 TIFF2026143428000255.tif222170 TIFF2026143428000256.tif224170 TIFF2026143428000257.tif229170 TIFF2026143428000258.tif232170 TIFF2026143428000259.tif222170 TIFF2026143428000260.tif229170 TIFF2026143428000261.tif223170 TIFF2026143428000262.tif230170 TIFF2026143428000263.tif221170 TIFF2026143428000264.tif224170 TIFF2026143428000265.tif230170 TIFF2026143428000266.tif238170 TIFF2026143428000267.tif224170 TIFF2026143428000268.tif223170 TIFF2026143428000269.tif224170 TIFF2026143428000270.tif229170 TIFF2026143428000271.tif229170 TIFF2026143428000272.tif224170 TIFF2026143428000273.tif229170 TIFF2026143428000274.tif229170 TIFF2026143428000275.tif223170 TIFF2026143428000276.tif222170 TIFF2026143428000277.tif228170 TIFF2026143428000278.tif230170 TIFF2026143428000279.tif223170 TIFF2026143428000280.tif229170 TIFF2026143428000281.tif231170 TIFF2026143428000282.tif222170 TIFF2026143428000283.tif229170 TIFF2026143428000284.tif223170 TIFF2026143428000285.tif231170 TIFF2026143428000286.tif224170 TIFF2026143428000287.tif224170 TIFF2026143428000288.tif223170 TIFF2026143428000289.tif224170 TIFF2026143428000290.tif222170 TIFF2026143428000291.tif224170 TIFF2026143428000292.tif229170 TIFF2026143428000293.tif229170 TIFF2026143428000294.tif222170 TIFF2026143428000295.tif216170 TIFF2026143428000296.tif225170 TIFF2026143428000297.tif229170 TIFF2026143428000298.tif229170 TIFF2026143428000299.tif223170 TIFF2026143428000300.tif224170 TIFF2026143428000301.tif222170 TIFF2026143428000302.tif223170 TIFF2026143428000303.tif230170 TIFF2026143428000304.tif223170 TIFF2026143428000305.tif224170 TIFF2026143428000306.tif230170 TIFF2026143428000307.tif229170 TIFF2026143428000308.tif223170 TIFF2026143428000309.tif236170

[0184] [Table 6]

Claims

1. A polynucleotide sequence encoding a circular RNA molecule, A circular RNA molecule comprises a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence, wherein the IRES sequence region At least one sequence region having RNA secondary structure elements; and It contains a sequence region complementary to 18S ribosomal RNA (rRNA), The IRES sequence region has a minimum free energy (MFE) of less than -18.9 kJ / mol and a melting temperature of at least 35.0°C; A polynucleotide sequence in which RNA secondary structure elements are formed from nucleotides approximately 40 to 60 of IRES, where the first nucleic acid at the 5' end of IRES is considered to be position 1.

2. The polynucleotide sequence according to claim 1, wherein a protein-coding nucleic acid sequence region is operably linked to an IRES sequence region in a non-natural stereochemistry.

3. A polynucleotide sequence according to claim 1 or 2, which is a DNA sequence.

4. The polynucleotide sequence according to any one of claims 1 to 3, wherein the sequence complementary to 18S rRNA is one of sequence numbers 28977 to 28983.

5. The polynucleotide sequence according to any one of claims 1 to 4, wherein at least one RNA secondary structure element is located at the 5' end of a sequence region complementary to 18S rRNA.

6. The polynucleotide sequence according to any one of claims 1 to 4, wherein at least one RNA secondary structure element is located at the 3' end of a sequence region complementary to 18S rRNA.

7. The polynucleotide sequence according to any one of claims 1 to 6, wherein at least one RNA secondary structure element is a stem-loop.

8. A polynucleotide sequence according to any one of claims 1 to 7, wherein at least one RNA secondary structure element comprises one of the nucleic acid sequences listed in Table 2.

9. The polynucleotide sequence according to any one of claims 1 to 8, wherein the IRES sequence region is approximately 100 to approximately 1000 nucleotides long.

10. The polynucleotide sequence according to any one of claims 1 to 8, wherein the IRES sequence region is approximately 200 to approximately 800 nucleotides in length.

11. The polynucleotide sequence according to any one of claims 1 to 8, wherein the IRES sequence is 150 to 200 nucleotides long, 160 to 180 nucleotides long, or 200 to 210 nucleotides long.

12. A polynucleotide sequence according to any one of claims 1 to 11, comprising at least one non-coding functional sequence.

13. The polynucleotide sequence according to claim 12, wherein the non-coding functional sequence comprises one or more (a) microRNA binding sites or (b) RNA-binding protein binding sites.

14. A polynucleotide sequence according to any one of claims 1 to 11, wherein the DNA sequence includes an aptamer.

15. A recombinant circular RNA molecule encoded by a polynucleotide sequence according to any one of claims 1 to 14.

16. A DNA sequence that codes for a circular RNA molecule, A circular RNA molecule comprises a protein-coding nucleic acid sequence region and an internal ribosome entry site (IRES) sequence region operably linked to the protein-coding nucleic acid sequence. A DNA sequence in which the IRES sequence region contains one of the nucleic acid sequences listed in SEQ ID NOs: 1 to 228 or SEQ ID NOs: 229 to 17201, or one of the nucleic acid sequences having at least 90% or at least 95% identity or homology thereto.

17. The DNA sequence according to claim 16, wherein a protein-coding nucleic acid sequence is operably linked to an IRES sequence region in a non-natural stereochemistry.

18. The DNA sequence according to any one of claims 16 to 17, wherein the IRES sequence region has a G-C content of at least 25%.

19. The DNA sequence according to any one of claims 16 to 18, wherein the IRES sequence region includes one of the nucleic acid sequences of sequence numbers 1 to 228.

20. The DNA sequence according to any one of claims 16 to 18, wherein the IRES sequence region includes one of the nucleic acid sequences of sequence numbers 229 to 17201.

21. The DNA sequence according to any one of claims 16 to 20, wherein the IRES sequence region includes one nucleic acid sequence from among sequence numbers 531, 2270, 2602, 3042, 3244, and 33948.

22. The DNA sequence according to any one of claims 16 to 21, wherein the IRES sequence region includes human IRES.

23. The DNA sequence according to any one of claims 16 to 22, wherein the protein-coding nucleic acid sequence region encodes a therapeutic peptide or protein.

24. The DNA sequence according to any one of claims 16 to 23, wherein the circular RNA molecule comprises about 200 nucleotides to about 10,000 nucleotides.

25. The DNA sequence according to any one of claims 13 to 24, wherein the circular RNA molecule includes a spacer between the IRES sequence region and the start codon of the protein-coding nucleic acid sequence region.

26. The DNA sequence according to claim 25, wherein the length of the spacer is selected to increase the translation of a protein-coding nucleic acid sequence region compared to the translation of a circular RNA without a spacer or with a spacer different from the selected spacer.

27. The DNA sequence according to any one of claims 16 to 26, wherein the IRES sequence region is configured to promote rolling circle translation.

28. The DNA sequence according to any one of claims 16 to 26, wherein the protein-coding nucleic acid sequence region lacks a stop codon.

29. (i) The IRES sequence region is configured to facilitate rolling circle translation, and (ii) The protein-coding nucleic acid sequence region lacks a stop codon, according to any one of claims 16 to 26.

30. A recombinant circular RNA molecule encoded by the DNA sequence described in any one of claims 16 to 29.

31. A viral vector comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

32. The viral vector according to claim 31, selected from the group consisting of adeno-associated virus (AAV) vectors, adenovirus vectors, retrovirus vectors, lentivirus vectors, vaccinia and herpesvirus vectors.

33. The viral vector according to claim 31 or 32, which is AAV.

34. The viral vector according to claim 33, wherein the AAV serotype is selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV8, AAV9, AAVrh10, or any variant thereof having substantially the same tropism.

35. A virus-like particle comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

36. Non-viral-like particles comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

37. A closed DNA sequence comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

38. A plasmid comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

39. A miniintron plasmid vector comprising a polynucleotide according to any one of claims 1 to 14 or a DNA sequence according to any one of claims 16 to 29.

40. A composition comprising a polynucleotide according to any one of claims 1 to 14, a DNA sequence according to any one of claims 16 to 29, or a recombinant circular RNA molecule according to claim 15 or 30.

41. The composition according to claim 40, wherein lipid nanoparticles are used to decorate it.

42. A host cell comprising a polynucleotide according to any one of claims 1 to 14, a DNA sequence according to any one of claims 16 to 29, or a recombinant circular RNA molecule according to claim 15 or 30.

43. A method for producing a protein in a cell, comprising the step of contacting a cell with (a) a polynucleotide according to any one of claims 1 to 14, (b) a DNA sequence according to any one of claims 16 to 29, (c) a circular RNA molecule according to any one of claims 15 or 30, (d) a viral vector according to any one of claims 31 to 34, (e) a virus-like particle according to claim 35, (f) a non-virus-like particle according to claim 36, (g) a closed DNA sequence according to claim 37, (h) a plasmid according to claim 38, (i) a miniintron plasmid vector according to claim 39, or (j) a composition according to any one of claims 40 to 41, under conditions in which a protein-coding nucleic acid sequence of circular RNA is translated in the cell and a protein is produced.

44. The method according to claim 43, wherein 5' cap-dependent translation in cells is impaired or absent.

45. The method according to claim 43 or 44, wherein the cells are in vivo.

46. The method according to claim 45, wherein the cells are mammalian cells.

47. The method according to claim 46, wherein the mammalian cells are derived from humans.

48. The method according to any one of claims 45 to 46, wherein the production of the protein is tissue-specific.

49. The method according to claim 48, wherein the tissue specificity is localized to a tissue selected from the group consisting of muscle, liver, kidney, brain, lung, skin, pancreas, blood, and heart.

50. The method according to claim 43 or 44, wherein the cells are in vitro.

51. The method according to any one of claims 43 to 48, wherein the protein is recursively expressed in a cell.

52. The method according to any one of claims 43 to 51, wherein the half-life of the circular RNA in the cell is approximately 1 to approximately 7 days.

53. The method according to any one of claims 43 to 55, wherein the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer than when the protein-coding nucleic acid sequence is provided to the cell in linear format RNA or encoded for transcription as linear RNA.

54. A protein produced by the method described in any one of claims 43 to 53.

55. A recombinant circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence, wherein the IRES is At least one RNA secondary structure; and It contains a sequence complementary to 18S ribosomal RNA (rRNA), A recombinant circular RNA molecule having an IRES of less than -18.9 kJ / mol, a minimum free energy (MFE), and a melting temperature of at least 35.0°C.

56. The recombinant circular RNA molecule according to claim 55, wherein the protein-coding nucleic acid sequence is operably linked to IRES in a non-natural stereochemistry.

57. The recombinant circular RNA molecule according to claim 55 or 56, wherein a sequence complementary to 18S rRNA is encoded by any one of sequence numbers 28977 to 28983.

58. A recombinant circular RNA molecule according to any one of claims 55 to 57, wherein at least one RNA secondary structure is located at the 5' end of a sequence complementary to 18S rRNA.

59. A recombinant circular RNA molecule according to any one of claims 55 to 57, wherein at least one RNA secondary structure is located at the 3' end of a sequence complementary to 18S rRNA.

60. A recombinant circular RNA molecule according to any one of claims 55 to 59, wherein at least one RNA secondary structure is a stem-loop.

61. A recombinant circular RNA molecule according to any one of claims 55 to 59, wherein at least one RNA secondary structure comprises a sequence encoded by one of the DNA sequences listed in Table 2.

62. Recombinant circular RNA according to any one of claims 55 to 61, wherein IRES is approximately 100 to approximately 1000 nucleotides long.

63. Recombinant circular RNA according to any one of claims 55 to 61, wherein IRES is approximately 200 to approximately 200 nucleotides long.

64. Recombinant circular RNA according to any one of claims 55 to 61, wherein the IRES sequence is 150 to 200 nucleotides, 160 to 180 nucleotides, or 200 to 210 nucleotides long.

65. The recombinant circular RNA according to any one of claims 62 to 64, wherein the RNA secondary structure is formed from nucleotides approximately 40 to 60 relative to the 5' end of IRES.

66. A recombinant circular RNA according to any one of claims 55 to 65, comprising at least one non-coding functional sequence.

67. The recombinant circular RNA according to claim 66, wherein the non-coding functional sequence comprises one or more (a) microRNA binding sites or (b) RNA-binding protein binding sites.

68. Recombinant circular RNA according to any one of claims 55 to 65, wherein the circular RNA comprises an aptamer.

69. A recombinant circular RNA molecule comprising a protein-coding nucleic acid sequence and an internal ribosome entry site (IRES) operably linked to the protein-coding nucleic acid sequence, IRES is a recombinant circular RNA molecule encoded by one of the DNA sequences listed in SEQ ID NOs: 1-228 or 229-17201, or a DNA sequence having at least 90% or at least 95% identity or homology thereto.

70. The recombinant circular RNA molecule according to claim 69, wherein the protein-coding nucleic acid sequence is operably linked to IRES in a non-natural stereochemistry.

71. The recombinant circular RNA according to any one of claims 55 to 70, wherein the recombinant circular RNA includes a back splice junction, and the IRES is located within about 100 to about 200 nucleotides from the back splice junction.

72. A recombinant cyclic RNA molecule according to any one of claims 55 to 71, wherein IRES has a G-C content of at least 25%.

73. A recombinant circular RNA molecule according to any one of claims 55 to 72, wherein IRES is encoded by any of the DNA sequences of sequence numbers 1 to 228.

74. A recombinant circular RNA molecule according to any one of claims 55 to 72, wherein IRES is encoded by one of the DNA sequences of sequence numbers 229 to 17201.

75. The recombinant circular RNA molecule according to claim 74, wherein IRES is encoded by one of the DNA sequences shown in SEQ ID NOs: 531, 2270, 2602, 3042, 3244, and 33948.

76. A recombinant circular RNA molecule according to any one of claims 55 to 75, wherein IRES is human IRES.

77. A recombinant circular RNA molecule according to any one of claims 55 to 76, wherein the protein-coding nucleic acid sequence encodes a therapeutic peptide or protein.

78. A recombinant circular RNA molecule according to any one of claims 55 to 77, wherein the circular RNA comprises approximately 200 nucleotides to approximately 10,000 nucleotides.

79. A recombinant circular RNA molecule according to any one of claims 55 to 78, wherein the circular RNA molecule includes a spacer between the IRES sequence region and the start codon of the protein-coding nucleic acid sequence region.

80. The recombinant circular RNA molecule according to claim 79, wherein the length of the spacer is selected to increase the translation of the protein-coding nucleic acid sequence region compared to the translation of circular RNA without a spacer or with a spacer different from the selected spacer.

81. A recombinant circular RNA molecule according to any one of claims 55 to 78, wherein the IRES sequence region is configured to promote rolling circle translation.

82. A recombinant circular RNA molecule according to any one of claims 55 to 78, wherein the protein-coding nucleic acid sequence region lacks a stop codon.

83. (i) The IRES sequence region is configured to facilitate rolling circle translation, and (ii) The protein-coding nucleic acid sequence region lacks a stop codon, according to any one of claims 55 to 78.

84. A composition comprising a recombinant cyclic RNA molecule according to any one of claims 55 to 83.

85. A host cell comprising a recombinant circular RNA molecule according to any one of claims 55 to 83 or the composition according to claim 84.

86. A method for producing a protein in a cell, comprising the step of contacting the cell with a recombinant circular RNA molecule according to any one of claims 55 to 83 or the composition according to claim 84, under conditions in which a protein-coding nucleic acid sequence is translated in the cell and a protein is produced.

87. The method according to claim 86, wherein 5' cap-dependent translation in cells is impaired or absent.

88. The method according to claim 86 or 87, wherein the cells are in vivo.

89. The method according to claim 86 or 87, wherein the cells are mammalian cells.

90. The method according to claim 89, wherein the mammalian cells are derived from humans.

91. The method according to any one of claims 55 to 90, wherein the production of the protein is tissue-specific.

92. The method according to claim 91, wherein the tissue specificity is localized to a tissue selected from the group consisting of muscle, liver, kidney, brain, lung, skin, pancreas, blood, and heart.

93. The method according to claim 86 or 87, wherein the cells are in vitro.

94. The method according to any one of claims 86 to 93, wherein a protein is recursively expressed in a cell.

95. The method according to any one of claims 86 to 94, wherein the half-life of the circular RNA in the cell is approximately 1 to approximately 7 days.

96. The method according to any one of claims 86 to 94, wherein the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer than when the protein-coding nucleic acid sequence is provided to the cell in linear format RNA or encoded for transcription as linear RNA.

97. A protein produced by the method described in any one of claims 86 to 96.

98. An oligonucleotide molecule containing a nucleic acid sequence that hybridizes to the internal ribosome entry site (IRES) on a circular RNA molecule and inhibits the translation of the circular RNA molecule.

99. The oligonucleotide molecule according to claim 98, wherein the circular RNA is recombinant circular RNA.

100. The oligonucleotide molecule according to claim 98, wherein the recombinant circular RNA is the recombinant circular RNA according to any one of claims 55 to 83.

101. The oligonucleotide molecule according to claim 98, wherein the circular RNA is a naturally occurring circular RNA.

102. An oligonucleotide molecule according to any one of claims 98 to 101, wherein the oligonucleotide is an antisense oligonucleotide.

103. The oligonucleotide according to claim 102, wherein the antisense oligonucleotide is a locked nucleic acid oligonucleotide (LNA).

104. The oligonucleotide according to any one of claims 98 to 103, wherein the oligonucleotide is chemically modified at its 5' and / or 3' end.

105. A method for inhibiting the translation of a protein-coding nucleic acid sequence present on a circular RNA molecule, comprising the step of contacting the circular RNA molecule with an oligonucleotide molecule according to any one of claims 98 to 104, thereby hybridizing the oligonucleotide molecule to a nucleic acid sequence complementary to the RNA secondary structure and / or 18S rRNA present on the IRES of the circular RNA molecule, thereby inhibiting the translation of the circular RNA molecule.

106. The method according to claim 105, wherein the oligonucleotide hybridizes to a nucleic acid sequence complementary to the RNA secondary structure or 18S rRNA.

107. The method according to claim 105, wherein the oligonucleotide hybridizes to a nucleic acid sequence complementary to the RNA secondary structure and 18S rRNA.

108. The method according to claim 105, wherein a first oligonucleotide hybridizes to an RNA secondary structure, and a second oligonucleotide hybridizes to a nucleic acid sequence complementary to 18S rRNA.