circRNA expression constructs
The positioning of IRES at a specific distance from back-splicing sites enhances the production of circRNA, thereby improving back-splicing efficiency and stability of circRNA yield.
Patent Information
- Application Number
- JP2025533430
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-07
- Publication Date
- 2026-01-06
AI Technical Summary
Existing methods for producing circular RNAs (circRNAs) face challenges in achieving high-level stable expression and efficient backsplicing, particularly due to interference from internal ribosome entry sites (IRES) located close to back-splicing sites, which can interfere with the back-splicing efficiency, resulting in reduced circRNA production.
A nucleic acid molecule encoding a circRNA with an IRES positioned at a specific distance from back-splicing sites, separated by at least 50 nucleotides, and optionally, a first inverted repeat element, optionally, a promoter operably linked to the expression cassette to direct translation of the ORF.
Enhances backsplicing efficiency and stability of circRNA production, allowing for efficient protein translation and improved circRNA yield.
Smart Images

Figure 2026500226000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to nucleic acid molecules encoding circular RNAs (circRNAs). In another aspect, the present invention relates to nucleic acids, circRNAs encoded by vectors, and host cells containing the nucleic acids. Further, the present invention includes methods for producing circRNAs. [Background technology]
[0002] Circular RNAs (circRNAs) constitute a novel class of long non-coding RNAs characterized as covalently closed molecules. CircRNAs are typically produced by a non-linear "back-splicing" event, using a downstream splice donor (SD) and an upstream splice acceptor (SA), as opposed to conventional linear splicing. Endogenous circular RNAs have been intensively studied in the past decade, and although some controversy exists regarding their functional relevance, it is now widely recognized that, thanks to their circular nature, circRNAs are resistant to exonuclease degradation and therefore comprise a highly stable class of RNAs with half-lives that significantly exceed those of conventional linear mRNAs.
[0003] Over the past decade, significant advances have been made in our understanding of circRNA production both in vitro and in vivo, resulting in the therapeutic potential of the concept of durable circular RNAs with engineered functions. Broadly speaking, circRNA production can be achieved through two distinct mechanisms: 1) using group I intron-derived ribozyme-mediated circularization, which has been shown to be effective in in vitro production, or 2) using spliceosome-based backsplicing, similar to endogenous circRNA biogenesis, which is useful in in vivo production. In the latter mechanism, the insertion of adjacent inverted elements is known to dramatically stimulate backsplicing, which is proposed to occur by closely positioning the involved splice sites. CircRNAs lack a 5' cap and a 3' poly(A) tail, and therefore cannot serve as translation substrates by themselves. However, the insertion of an internal ribosome entry site (IRES) efficiently converts non-coding circRNAs into highly efficient protein-coding molecules. Therefore, the therapeutic use of circular RNAs as templates for durable protein production has attracted considerable attention.
[0004] The present invention provides an improved circRNA expression cassette that promotes high-level stable circRNA expression in vivo. Surprisingly, we discovered that the IRES located within the circRNA expression cassette is significantly involved in circRNA yield. Summary of the Invention
[0005] Disclosed herein is a nucleic acid encoding a circular RNA (circRNA), in which an IRES driving translation of the circRNA is located at a specific position relative to a splice site required for backsplicing of the circRNA, resulting in excellent backsplicing efficiency.
[0006] In a first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises: A) An expression cassette comprising: (i) a circRNA expression cassette comprising: (a) a nucleic acid comprising or consisting of a continuous or split open reading frame (ORF) encoding at least one protein; and (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF; (ii) a first back-splicing site located 5′ to the IRES; and (iii) a second back-splicing site located 3′ to the IRES; and, optionally, B) a first inverted repeat (IR) element located 5' to the first back splicing site and a second IR element located 3' to the second back splicing site; and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette; The IRES is located within the circRNA expression cassette as follows: (I) the 5' end of the IRES and the 3' end of the first back-splicing site are separated by at least 50 nucleotides, preferably at least 200, more preferably at least 300 nucleotides; (II) the 3' end of the IRES and the 5' end of the second back-splicing site are separated by at least 300 nucleotides, preferably at least 350 nucleotides; A nucleic acid molecule is provided. [Brief explanation of the drawings]
[0007] In the following, the contents of the drawings contained in this specification are described, and in this context reference is also made to the detailed description of the invention given above and / or below.
[0008] [Figure 1A](A) Schematic diagram of the continuous ORF (top) and split ORF (bottom) designs, showing the relative positions of the promoter, inverted repeat (IR), splice acceptor (SA), IRES, open reading frame (ORF), and splice donor (SD). [Figure 1B] (B) Western blot of ICOS-L protein expression from A375 cells transfected with either an empty vector control (EV) or circICOSL expression plasmids containing either the CVB3 or EMCV IRES, both continuous (Cont) and split ORFs. β-Actin was used as a loading control. [Figure 1C] (C) RT-qPCR quantification of circICOSL expression relative to GAPDH mRNA from A375 cells transfected with both continuous and split-ORF circICOSL expression plasmids containing either the CVB3 or EMCV IRES. Data from three biological replicates are shown. [Figure 1D] (D) Agarose gel showing RT-PCR amplification of circICOSL from A375 cells transfected with either an empty vector control (EV) or circICOSL expression plasmids containing either the CVB3 or EMCV IRES, both continuous and split ORF designs. [Figure 2A] (A) Schematic diagram of the eGFP split-ORF design. The circRNA diagram shows the relative position of the IRES within the split-eGFP ORF. [Figure 2B] (B) Agarose gel showing RT-PCR amplification of circEGFP derived from A375 cells transfected with the eGFP_FLAG_Split_ORF circRNA plasmid. [Figure 2C](C) RT-qPCR quantification of circEGFP expression relative to GAPDH mRNA from A375 cells transfected with the eGFP_FLAG_Split_ORF circRNA plasmid. Data from two biological replicates are shown. [Figure 2D] (D) Western blotting analysis of A375 cells transfected with the eGFP_FLAG_Split_ORF circRNA plasmid probed with the indicated antibodies for FLAG, GFP, and β-actin (loading control). [Figure 3A] (A) Schematic diagram of the eGFP_FLAG_Split_IRES circRNA design. The circRNA diagram shows the relative positions of the split_IRES and eGFP ORF. [Figure 3B] (B) Agarose gel showing RT-PCR amplification of circEGFP from A375 cells transfected with eGFP_FLAG_Split_IRES circRNA plasmid. [Figure 3C] (C) RT-qPCR quantification of circEGFP expression relative to GAPDH mRNA from A375 cells transfected with the eGFP_FLAG_Split_IRES circRNA plasmid. Data from two biological replicates are shown. [Figure 4A] (A) Schematic diagram of the eGFP_HIPK3_spacer circRNA design. The circRNA diagram shows the relative positions of the HIPK3_spacer (black), IRES (light gray), and eGFP ORF (dark gray). [Figure 4B] (B) Agarose gel showing RT-PCR amplification of circEGFP derived from A375 cells transfected with the eGFP_FLAG_HIPK3_spacer circRNA plasmid. [Figure 4C](C) RT-qPCR quantification of circEGFP expression relative to GAPDH mRNA from A375 cells transfected with the eGFP_FLAG_HIPK3_spacer circRNA plasmid. Data from two biological replicates are shown. [Figure 5A] (A) Schematic diagram of mRNA and circRNA design. The circRNA diagram shows the relative positions of the IRES and transgene ORF. [Figure 5B] (B-D) Western blotting analysis of A375 cells transfected with mRNA encoding ADA_FLAG (B), ICOSL (C), and Renilla (D), and circRNA plasmids probed with the indicated antibodies for FLAG, ICOSL, Renilla, and β-actin (loading control). [Figure 5C] Same as above. [Figure 5D] Same as above. [Figure 5E] (E-G) RT-qPCR quantification of circADA (E), circICOSL (F), and circRenilla (G) expression relative to mRNA encoding ADA_FLAG, ICOSL, and Renilla, and mRNA for GAPDH from A375 cells transfected with circRNA plasmids. Data from two biological replicates are shown. [Figure 5F] Same as above. [Figure 5G] Same as above. [Figure 6A] (A) Western blotting analysis of A375 cells transfected with the ICOSL_Split_CVB3 circRNA plasmid, probed with the indicated antibodies for ICOSL and β-actin (loading control). Schematic diagram of the ICOSL_Split_CVB3 circRNA expression cassette design showing the relative positions of the CVB3 IRES (dark gray) and ICOSL ORF (light gray). [Figure 6B](B) RT-qPCR quantification of ICOSL expression relative to GAPDH mRNA from A375 cells transfected with the ICOSL_Split_CVB3 circRNA plasmid. Data from two biological replicates are shown. [Figure 7A] (A) Western blotting analysis of A375 cells transfected with the eGFP_Split_CVB3 circRNA plasmid, probed with the indicated antibodies for GFP and β-actin (loading control). Schematic diagram of the eGFP_Split_CVB3 circRNA expression cassette design showing the relative positions of the CVB3 IRES (dark gray) and eGFP ORF (light gray). [Figure 7B] (B) RT-qPCR quantification of eGFP expression relative to GAPDH mRNA from A375 cells transfected with the eGFP_Split_CVB3 circRNA plasmid. Data from two biological replicates are shown. [Figure 7C] (C) Western blotting analysis of A375 cells transfected with the eGFP_Split_EMCV_circRNA plasmid probed with the indicated antibodies for GFP and β-actin (loading control). A schematic diagram of the eGFP_Split_EMCV circRNA expression cassette is shown. The circRNA diagram shows the relative positions of the EMCV IRES (dark gray) and eGFP ORF (light gray). [Figure 7D] (D) RT-qPCR quantification of eGFP expression relative to GAPDH mRNA from A375 cells transfected with the eGFP_Split_EMCV circRNA plasmid. Data from two biological replicates are shown. DETAILED DESCRIPTION OF THE INVENTION
[0009] Before describing the present invention in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols, and reagents described herein, as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which will be limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0010] Preferably, the terms used herein are defined as set forth in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, HGW, Nage, B. and Klbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010, Basel, Switzerland.
[0011] Throughout this specification and the claims that follow, unless the context requires otherwise, the word "comprise," and variations such as "comprises" and "comprising," will be understood to refer to the inclusion of the specified element, step, or group of elements or steps, and not to the exclusion of any other element, step, or group of elements or steps. In the following description, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or group of aspects, unless clearly indicated to the contrary. In particular, any feature indicated as optional, preferred, or advantageous may be combined with any other feature or any feature indicated as optional, preferred, or advantageous.
[0012] Throughout the text of this specification, several documents are cited. Each document cited herein (including all patents, patent applications, scientific literature, manufacturer's specifications, instructions, etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such disclosure by virtue of prior invention. Some of the documents cited herein are characterized as "incorporated by reference." In the event of a conflict between a definition or teaching of such an incorporated document and a definition or teaching set forth herein, the description herein controls.
[0013] Below, components of the present invention are described. These components are listed with specific embodiments; however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The various described examples and preferred embodiments should not be construed as limiting the invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments combining the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, unless the context indicates otherwise, any permutation and combination of all described components in this application should be considered disclosed by the present description.
[0014] definition
[0015] Below we provide definitions of some of the terms commonly used in this specification, which have their respective defined and preferred meanings in the remainder of the specification, at each instance of their use.
[0016] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise.
[0017] The term "about," when used in connection with a numerical value, is meant to encompass numerical values within a range having a lower limit of 5% less than the stated numerical value and an upper limit of 5% greater than the stated numerical value.
[0018] The term "circular RNA" or "circRNA" as used herein refers to a type of RNA in which the ends of an RNA strand are covalently linked to form a closed, continuous circle. CircRNAs are generally formed by covalent bonding between the 5' site of an upstream exon and the 3' site of the same or downstream exon.
[0019] The terms "5'" and "3'," as used herein, refer to either end of a single-stranded nucleic acid molecule. The 5' end is the end of a molecule that terminates in a 5' phosphate group, and the 5' direction is the direction toward the 5' end. Similarly, the 3' end is the end of a molecule that terminates in a 3' phosphate group, and the 3' direction is the direction toward the 3' end. In the context of the present invention, nucleic acid sequences are written with the 5' end at the left and the 3' end at the right (unless otherwise specified), with reference to the direction of DNA synthesis during replication (5' to 3'), RNA synthesis during transcription (5' to 3'), and reading of the mRNA sequence during translation (5' to 3').
[0020] The term "open reading frame," as used herein, refers to the portion of a nucleic acid sequence between the initiation and termination codons, but does not include the termination codon (which functions as a termination signal). A codon is a DNA or RNA sequence of three nucleotides that forms a unit of genetic information that encodes a specific amino acid or signals the end of protein synthesis (a termination codon). In the standard genetic code, three different termination codons are known: TAG, TAA, or TGA in DNA, or UAG, UAA, and UGA in RNA. Other termination codons may exist in variations of the standard genetic code. The initiation codon is the first codon of an RNA transcript (e.g., linear mRNA, circRNA) that is translated by the ribosome.
[0021] The term "back splicing," as used herein, refers to splicing in the reverse order (ie, back splicing) in which an upstream 3' splice site joins a downstream 5' splice site.
[0022] The term "back-splice site," as used herein, refers to the nucleotide sequence of an intron-exon boundary (i.e., a splice site). The back-splice site determines the location to which the pre-mRNA is spliced and allows for the recruitment of the spliceosome required for the back-splicing process. The sequence of a back-splice site typically contains two half-sites, one located in the intron and the other in the exon. The spliced nucleotide sequence is divided between the two half-sites of the back-splice site.
[0023] The term "internal ribosome entry site (IRES)," as used herein, refers to a region in RNA that allows for internal initiation of translation in a cap-independent manner. IRES elements typically comprise a highly structured stretch of RNA containing several stem-loop structures. IRES were first identified in picornaviruses, but are present in a variety of different viruses. More recently, IRES sequences have also been identified in many cellular mRNAs. Both types of IRES sequences (i.e., viral-derived and cellular IRES sequences) can generally be used to practice the present invention.
[0024] Viral IRES can be divided into four distinct classes based on two main criteria: first, the type of secondary and tertiary structure of their RNA elements, and second, their mode of action for translation initiation (see Mailliot and Martin, RNA 2017, e1458).
[0025] Class 1 IRESs typically require all initiation factors except eIF4E and contain a fairly basic secondary structure consisting of short and long hairpins. Class 1 IRESs typically recruit the ribosome upstream of the coding region and rely on classical 5' to 3' scanning to find the start codon. Class 1 IRESs are found, for example, in enterovirus A71 (EV-A71), coxsackievirus B3 (CVB3), poliovirus (PV), and human rhinovirus 2 (HRV).
[0026] In contrast, class 2 IRESs facilitate direct tethering of the translation initiation machinery to the start codon and do not have any scanning step. Class 2 IRESs are found in Picornaviridae, for example, encephalomyocarditis virus (EMCV), and foot-and-mouth disease virus (FMDV).
[0027] Class 3 IRES contain more sophisticated secondary and tertiary structures, such as pseudoknots. They require only a small subset of translation initiation factors, namely eIF2, eIF3, and eIF5, to recruit the ribosome and load it directly onto the AUG start codon without scanning. Class 3 IRES are found, for example, in the Flaviviridae family, such as hepatitis C virus (HCV) and classical swine fever virus (CSFV); Picornaviridae, such as porcine teschoviruses and porcine enterovirus 8 (PEV8), or simian virus 2 (SV2).
[0028] Class 4 IRESs are more compact and sophisticated in terms of structural complexity; they usually contain several pseudoknots and do not require any translation initiation factors. They are the smallest known IRESs (usually less than 200 nucleotides). Class 4 IRESs are found in Dicistroviridae, for example, in Cricket Paralysis Virus (CrPV), Israeli Acute Paralysis Virus (IAPV), Platia Stali Instestine Virus (PSIV), or Taura Syndrome Virus (TSV).
[0029] In the context of the present invention, the most preferred classes of IRES are class 1 and class 2 IRES.
[0030] The term "inverted repeat (IR) element," as used herein, refers to a nucleotide sequence followed downstream by its reverse complement, which may be identical to or may deviate from the reverse complement.
[0031] The term "promoter," as used herein, refers to a sequence of DNA to which a protein binds to initiate transcription of an RNA transcript from DNA downstream of the promoter. Promoters are typically located in the 5' region of a gene and do not encode a gene product but control its expression. Transcription is initiated at the promoter by the interaction of transcription factors and RNA polymerase. In general, it is advantageous to select a promoter that is active in the desired type of host cell.
[0032] Based on the pattern of promoter activity, promoters can be classified into different types. Some promoters are constitutive promoters, which are active in essentially all tissues and do not require any specific stimulus to be activated. Non-limiting examples of constitutive promoters are: CMV, EF1a, EFS, CAG, CBh, CBA, MSCV, PGK, SFFV, SV40, and UBC. In a preferred embodiment, the promoter is CMV. In another preferred embodiment, the CMV promoter has the sequence set forth in SEQ ID NO: 005.
[0033] Other promoters are active only in specific tissues and are therefore useful for restricting expression to specific tissues. Non-limiting examples of tissue-specific promoters that can be used in the present invention, and the tissues in which they are preferably active, are provided below: Embryonic tissue: Nanog, Brain / Nervous System / Spinal Cord: Nes, Tuba1a, Camk2a, SYN1, Hb9, Th, Thy1, NSE, GFAP (long), GFAP (short), Iba1, Prnp, Cnp, retina: ProA1, hRHO, hBEST1, Grm6, Epidermis: K14, BK5, mTyr, Heart / embryonic heart: cTnT, αMHC (long), αMHC (short), Hcn4, muscle: Myog, ACTA1, MHCK7, SM22a, EnSM22a, Desmin, Mb, Bone / cartilage: Runx2, OC, Col1a1, Col2a1, fat: aP2, Adipoq, Blood / bone marrow:Liver: Afp, Alb, TBG, mammary gland; MMTV, Wap, pancreas: HIP, Pdx1, Ins2, Elastase-1, kidney; NPHS2, lung; SPB, Endothelial cells: CD144, Flt-1, ICAM-2, Endoglin, Hematopoietic cells: WASP, IFN-B, B19, CD14, CD43, CD45, CD68, Tie1, CD11b, OG-2, tumor:TERT, E2F-1, OC, SLPI, Cox-2, CEA, AFP, LP-P. In a preferred embodiment, the promoter is selected from SYN1, GFAP.
[0034] Other promoters require a specific stimulus to become active and initiate transcription. Non-limiting examples of inducible promoters are: TRE (tet sensitive promoter), CRE (cumate sensitive promoter).
[0035] In preferred embodiments, the promoter is selected from a constitutive, tissue-specific, or inducible promoter. In preferred embodiments, the promoter is a constitutive promoter. In preferred embodiments, the promoter is a tissue-specific promoter. In preferred embodiments, the promoter is an inducible promoter.
[0036] The term "vector," as used herein, refers to a vehicle capable of introducing DNA or RNA sequences (e.g., foreign genes) into a host cell in order to transform the host and promote the expression (e.g., transcription and translation) of the introduced sequences.
[0037] The term "host cell" as used herein refers to a cell that contains the nucleic acid, vector, or circRNA described herein.Host cells can be eukaryotic cells, such as plants, animals, fungi, or algae, or prokaryotic cells, such as bacteria or protozoa.Host cells can be cultured cells or primary cells, i.e., cells directly isolated from living organisms, such as humans.Host cells can be adherent cells or suspension cells, i.e., cells grown in suspension.Preferably, host cells are mammalian cells.
[0038] The term "branch site," as used herein, refers to a nucleotide typically contained in the heptameric sequence of an intron. The heptameric sequence, except for the branch site, forms base pairs with the spliceosome. An unpaired branch site subsequently requires the formation of a lariat intermediate during the splicing event. Typically, the branch site is an adenine.
[0039] The term "polypyrimidine tract," as used herein, refers to a nucleotide sequence that is typically 15 to 20 base pairs in length and that can promote the assembly of the spliceosome, thereby assisting in splicing.
[0040] Embodiment
[0041] In the following, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or group of aspects, unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous.
[0042] In a first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises: A) An expression cassette comprising: (i) a circRNA expression cassette comprising: (a) a nucleic acid comprising or consisting of a continuous or split open reading frame (ORF) encoding at least one protein; and (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF; (ii) a first back-splicing site located 5′ to the IRES; and (iii) a second back-splicing site located 3′ to the IRES; and, optionally, B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site, and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette; The IRES is located within the circRNA expression cassette as follows: (I) the 5' end of the IRES and the 3' end of the first back-splicing site are separated by at least 50 nucleotides, preferably at least 200, more preferably at least 300 nucleotides; (II) the 3' end of the IRES and the 5' end of the second back-splicing site are separated by at least 300 nucleotides, preferably at least 350 nucleotides; A nucleic acid molecule is provided.
[0043] In an alternative first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises: A) An expression cassette comprising: (i) a circRNA expression cassette comprising: (a) a nucleic acid comprising or consisting of a continuous or split open reading frame (ORF) encoding at least one protein; and (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF; (ii) a first back-splicing site located 5′ to the IRES; and (iii) a second back-splicing site located 3′ to the IRES; and, optionally, B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site, and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette; A nucleic acid molecule is provided.
[0044] Below, the different elements of the nucleic acids of the different aspects of the invention are disclosed in more detail and particularly preferred embodiments are disclosed.
[0045] circRNA expression cassettes, and expression cassettes
[0046] Without wishing to be bound by theory, we believe that an IRES site located too close to a back-splicing site may interfere with back-splicing efficiency, resulting in reduced circRNA production. This interference may be due to any secondary structure formed by the IRES that does not allow optimal access of the spliceosome to the linear pre-mRNA. Therefore, locating the IRES away from the back-splicing site within the expression cassette may be advantageous for improving back-splicing efficiency.
[0047] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 350, at least 400, at least 450, at least 475, or at least 500 nucleotides. In a preferred embodiment, the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 160 nucleotides. In a preferred embodiment, the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 475 nucleotides.
[0048] In preferred embodiments, the IRES is positioned within the circRNA expression cassette such that the 3' end of the IRES and the 5' end of the second back-splicing site are separated by at least 300, at least 310, at least 320, at least 330, at least 340, or at least 350 nucleotides.
[0049] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 50 nucleotides, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, or at least 290 nucleotides. In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 50 nucleotides.
[0050] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 50 nucleotides, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, or at least 350 nucleotides (preferably at least 250 nucleotides), and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 300, at least 310, at least 320, at least 330, at least 340, or at least 350 nucleotides (preferably at least 300).
[0051] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 50 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 350 nucleotides.
[0052] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 50 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 680 nucleotides.
[0053] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 250 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 480 nucleotides.
[0054] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 360 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 380 nucleotides.
[0055] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 650 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 90 nucleotides.
[0056] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are separated by at least 300 nucleotides, and the 3' end of the IRES and the 5' end of the second back-splice site are separated by at least 300 nucleotides.
[0057] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that both the 5' end of the IRES and the 3' end of the first back-splice site, and the 3' end of the IRES and the 5' end of the second back-splice site, are at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, or at least 400 nucleotides apart from each other.
[0058] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splicing site are spaced 50 to 650 nucleotides (preferably 200 to 400 nt, more preferably 250 to 365 nt) apart from each other.
[0059] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splicing site are spaced 59 to 650 nucleotides (preferably 255 to 363 nt) apart from each other.
[0060] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 3' end of the IRES and the 5' end of the second back-splicing site are spaced 300 to 700 nucleotides (preferably 350 to 600 nt, more preferably 350 to 500 nt) apart from each other.
[0061] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 3' end of the IRES and the 5' end of the second back-splicing site are spaced 381 to 685 nucleotides (preferably 381 to 489 nt) apart from each other.
[0062] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are spaced 50 to 650 nucleotides (preferably 200 to 400 nt, more preferably 250 to 365 nt) apart, and the 5' end of the IRES and the 3' end of the first back-splice site are spaced 300 to 700 nucleotides (preferably 350 to 600 nt, more preferably 350 to 500 nt) apart.
[0063] In a preferred embodiment, the IRES is positioned within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splice site are spaced 59 to 650 nucleotides (preferably 255 to 363 nt) apart, and the 3' end of the IRES and the 5' end of the second back-splice site are spaced 381 to 685 nucleotides (preferably 381 to 489 nt) apart.
[0064] In a preferred embodiment, the IRES is positioned in the circRNA expression cassette so that the IRES is located at the center of the circRNA expression cassette. In one embodiment, the nucleotides that are equally spaced from the first and second back-splicing sites are part of the IRES sequence. In one embodiment, the sequences adjacent to the IRES sequence in the circRNA expression cassette differ by 50 nucleotides or less.
[0065] Another possibility for separating the IRES sequence from the back-splicing site is the use of a nucleotide sequence that is not part of the ORF, i.e., a spacer sequence. In a preferred embodiment, the circRNA expression cassette contains a spacer sequence, thus separating the IRES and the first back-splicing site. In a preferred embodiment, the spacer sequence is located 5' of the IRES sequence and 3' of the first back-splicing site. In another preferred embodiment, the spacer sequence is located 3' of the IRES sequence and 5' of the second back-splicing site. In yet another preferred embodiment, the circRNA expression cassette contains two spacer sequences. Preferably, the first spacer sequence is located 5' of the IRES sequence and 3' of the first back-splicing site, and the second spacer sequence is located 3' of the IRES sequence and 5' of the second back-splicing site. In a preferred embodiment, the spacer sequence has a length of 50 to 500 nucleotides, preferably 100 to 400 nucleotides, more preferably 150 to 300 nucleotides, and most preferably 200 to 250 nucleotides. In preferred embodiments, the spacer sequence has a length of 50, 100, 200, or 400 nucleotides.
[0066] In yet another embodiment, the spacer sequence may encode a protein.
[0067] In a preferred embodiment, the IRES is CVB3, and the IRES is positioned within the circRNA expression cassette as follows: (I) the 5' end of the IRES and the 3' end of the first back-splicing site are separated by at least 50 nucleotides, preferably at least 200, more preferably at least 300 nucleotides; (II) The 3' end of the IRES and the 5' end of the second back-splicing site are separated by at least 50 nucleotides, preferably at least 100, at least 150, at least 200, or at least 250 nucleotides.
[0068] Open reading frame (ORF) structure
[0069] The ORFs contained in the circRNA expression cassette may be arranged as a continuous ORF located 3' or 5' of the IRES, or as a split ORF in which the ORF is split (preferably at a sequence that creates half-sites for a back-splicing site at the split position) and the two parts of the split ORF are adjacent to the IRES in the circRNA expression cassette.
[0070] Split ORFs have several advantageous properties in the context of the present invention. First, the design of the split ORF separates the IRES from the back-splicing site, which is thought to improve back-splicing efficiency and thereby increase circRNA yield without the need for additional sequence elements, i.e., spacer sequences. Another advantageous property is that the protein encoded by the ORF can be translated only from the resulting circRNA, and not from the linear RNA that may have been generated from the nucleic acid molecule if circularization had failed or been incomplete. Thus, only the protein encoded by the circRNA is protected from translation.
[0071] In a preferred embodiment of the first aspect of the present invention, the split ORF comprises two parts, the first part of the split ORF comprises a stop codon and is located 5' to the IRES, and the second part of the split ORF comprises a start codon and is located 3' to the IRES. In circRNAs made from the nucleic acids of the first aspect of the present invention, the split ORFs will be reassembled by circularization of the circRNA, resulting in a continuous ORF within the closed circle of the circRNA.
[0072] In a preferred embodiment, where the circRNA expression cassette comprises a split ORF, the second half-site of the first back-splicing site and the first half-site of the second back-splicing site are located within the two parts of the split ORF. In other words, they are part of the protein-coding sequence in the ORF. This is advantageous because otherwise, any sequence left behind by the splicing event would be present in the translated sequence of the ORF, and the additional nucleotides in the sequence may interfere with protein translation and / or folding.
[0073] In another preferred embodiment, the ORF is a continuous ORF. Preferably, the circRNA expression cassette containing the continuous ORF further comprises a non-coding spacer sequence that separates the IRES from the back-splicing site. Examples of usable spacer sequences are described above.
[0074] In a preferred embodiment, the circRNA expression cassette comprises a continuous ORF and further comprises one or more non-coding nucleotide sequences located at the 5' and / or 3' ends of the continuous ORF.
[0075] In a preferred embodiment, the circRNA expression cassette comprises a split ORF and further comprises one or more non-coding nucleotide sequences located at the 5' end and / or 3' end of the IRES.
[0076] In a preferred embodiment, the circRNA expression cassette further comprises one or more (e.g., 1, 2, 3, 4) additional ORFs.
[0077] In a preferred embodiment, the circRNA expression cassette further comprises one or more (e.g., 1, 2, 3, 4) additional IRES(s).
[0078] In a preferred embodiment, the circRNA expression cassette further comprises an additional coding nucleotide sequence linked to one or more ORFs, preferably encoding a sequence that can be used to detect or capture a protein encoded by one or more ORFs.
[0079] In a preferred embodiment, the ORF encodes at least one therapeutic protein, therapeutic peptide, antigenic protein, or antigenic peptide. Generally, any type of peptide or protein may be encoded by the ORF.
[0080] Backsplicing Site
[0081] The nucleic acids of the present invention contain at least two back-splice sites located upstream (i.e., 5') and downstream (i.e., 3') of the IRES. The back-splice sites are important for enabling the formation (i.e., circularization) of circRNAs encoded by the nucleic acids of the first aspect of the present invention. The back-splice sites determine the position to which the pre-mRNA is spliced and enable the recruitment of spliceosomes required for the back-splicing process. Typically, a back-splice site contains two half-sites between which the nucleic acid is split during splicing. In this context, the term "splice acceptor (SA)" refers to the 3' end of the intron. The term "splice donor (SD)" refers to the 5' end of the intron. In a preferred embodiment, the first back-splice site contains a splice acceptor, and the second back-splice site contains a splice donor.
[0082] In a preferred embodiment, the back-splice site comprises a nucleotide sequence that enables spliceosome-assisted back-splicing of the encoded circRNA. In a preferred embodiment, the back-splice site requires the spliceosome to deliver the circRNA encoded by the circRNA expression cassette.
[0083] In a preferred embodiment, the first back-splicing site comprises two half-sites. In one embodiment, the first and second half-sites of the first back-splicing site are located on the 5' side of the circRNA expression cassette. In another embodiment, the first half-site of the first back-splicing site is located on the 5' side of the circRNA expression cassette, and the second half-site of the first back-splicing site is part of the circRNA expression cassette. This embodiment is preferred when the ORF is a split ORF. Preferably, the second half-site of the first back-splicing site is part of the ORF sequence.
[0084] In a preferred embodiment, the circRNA generated by backsplicing of the nucleic acid molecule of the first aspect of the present invention does not contain any sequence derived from the backsplicing site that is not part of the ORF or IRES sequence.
[0085] In a preferred embodiment, the second back-splicing site comprises two half-sites. In one embodiment, the first and second half-sites of the second back-splicing site are located 3' of the circRNA expression cassette. In another embodiment, the first half-site of the second back-splicing site is part of the circRNA expression cassette. This embodiment is preferred when the ORF is a split ORF. Preferably, in the case of a split ORF, the first half-site of the second back-splicing site is part of the ORF, and the second half-site of the second back-splicing site is located 3' of the circRNA expression cassette.
[0086] In a preferred embodiment, the first and second back-splicing sites each comprise two half sites. In one embodiment, the first and second half sites of the first back-splicing site are located on the 5' side of the circRNA expression cassette, and the first and second half sites of the second back-splicing site are located on the 3' side of the circRNA expression cassette. In another embodiment, the first half site of the first back-splicing site is located on the 5' side of the circRNA expression cassette, the second half site of the first back-splicing site is part of the circRNA expression cassette, the first half site of the second back-splicing site is part of the circRNA expression cassette, and the second half site of the second back-splicing site is located on the 3' side of the circRNA expression cassette.
[0087] In a preferred embodiment, the first half site of the first back splicing site comprises or consists of the nucleotide AG. In another preferred embodiment, the first half site of the first back splicing site comprises or consists of the nucleotide sequence CAG. In another preferred embodiment, the first half site of the first back splicing site is a splice acceptor. In another preferred embodiment, the second half site of the first back splicing site begins with the nucleotide G.
[0088] In a preferred embodiment, the second half site of the second back splicing site comprises or consists of the nucleotides GT. In another preferred embodiment, the second half site of the second back splicing site comprises or consists of the nucleotide sequence GTAAGT. In another preferred embodiment, the second half site of the second back splicing site is a splice donor. In another preferred embodiment, the first half site of the second back splicing site comprises or consists of the nucleotide sequence CAG or AAG.
[0089] In a preferred embodiment, the first half site of the first back splicing site comprises or consists of the nucleotides AG; and the second half site of the second back splicing site comprises or consists of the nucleotides GT. In a preferred embodiment, the first half site of the first back splicing site comprises or consists of the nucleotides CAG; and the second half site of the second back splicing site comprises or consists of the nucleotides GTAAGT.
[0090] In a preferred embodiment, the back-splice site is a consensus sequence for the U2 (major class) intron in the pre-mRNA, generally conforming to the following consensus sequences: 3' splice site: CAG|G, and 5' splice site: MAG|GTRAGT, where M is A or C, R is A or G, and | indicates the intron-exon boundary. In another preferred embodiment, the splice site has a sequence as shown by the similarity matrix for the human U2 intron, e.g., as disclosed in Zhang Hum Mol Genet. 1998 May;7(5):919-32, incorporated herein by reference.
[0091] In a preferred embodiment, the sequence of the ORF is modified to create a back-splicing site. Preferably, the modification introduces only silent mutations (i.e., does not change the amino acid sequence of the encoded protein).
[0092] In a preferred embodiment, the first back-splice site has the nucleotide sequence CAGGT. In another preferred embodiment, the first back-splice site has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the nucleotide sequence CAGGT.
[0093] In a preferred embodiment, the second back splice site has the nucleotide sequence AGGTA. In another preferred embodiment, the second back splice site has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the nucleotide sequence AGGTA.
[0094] Internal ribosome entry site (IRES)
[0095] circRNAs form covalently closed circles and therefore lack a 5' end, requiring a cap-independent mechanism to initiate translation of the encoded protein. Internal ribosome entry sites (IRES) are RNA sequences that recruit the 40S ribosomal subunit via a cap-independent mechanism, thus enabling translation initiation. These elements typically adopt complex secondary RNA structures and act as ribosome anchoring sites.
[0096] Therefore, the nucleic acid molecule of the present invention contains an IRES in the circRNA expression cassette. The IRES is operably linked to an ORF, and at least one protein encoded by the ORF can be translated from the circRNA. Without wishing to be bound by theory, the inventors believe that if the IRES is located very close to a back-splicing site, the complex secondary structure of the IRES or IRES-associated protein may interfere with the spliceosome. IRESs vary in their nucleic acid sequences, but they share the commonality that they typically have complex secondary structures. Therefore, the interference of the spliceosome by an IRES does not depend on the specific IRES, but rather on the common secondary structure of IRESs in general.
[0097] In a preferred embodiment, the IRES is a viral IRES or a cellular IRES, preferably a viral IRES. In a preferred embodiment, the IRES is a class 1 or class 2 IRES as defined above. Non-limiting examples of class 1 IRESs are found in enterovirus A71 (EV-A71), coxsackievirus B3 (CVB3), poliovirus (PV), and human rhinovirus 2 (HRV), preferably CVB3. Non-limiting examples of class 2 IRESs are found in picornaviruses such as encephalomyocarditis virus (EMCV) and foot-and-mouth disease virus (FMDV), preferably EMCV.
[0098] In a preferred embodiment, the IRES is a class 1 IRES. In another preferred embodiment, the IRES is a class 2 IRES. In another preferred embodiment, the IRES is a class 3 IRES. In another preferred embodiment, the IRES is a class 4 IRES.
[0099] In a preferred embodiment, the IRES is an IRES sequence from a Coxsackievirus (CVB), more preferably CVB3. In a more preferred embodiment, the IRES has the sequence set forth in SEQ ID NO: 001. In another preferred embodiment, the IRES has a sequence that is at least 50%, at least 75%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 001. In a preferred embodiment, the IRES has a sequence that is at least 75% identical to the sequence of SEQ ID NO: 001. In a preferred embodiment, the IRES has the sequence set forth in SEQ ID NO: 001.
[0100] In a preferred embodiment, the IRES is an IRES sequence from encephalomyocarditis virus (EMCV). In a more preferred embodiment, the IRES has the sequence set forth in SEQ ID NO: 002. In another preferred embodiment, the IRES has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 002.
[0101] In one embodiment, the IRES is a viral-derived IRES, preferably an IRES from the Adenoviridae, Arenaviridae, Birnaviridae, Chrysoviridae, Coronaviridae, Dicistroviridae, Filoviridae, Flaviviridae, Hepadnaviridae, Herpesviridae, Hypoviridae, Iflaviridae, Luteoviridae, Orthomyxoviridae e), Papillomaviridae, Paramyxoviridae, Parvoviridae, Picornaviridae, Pneumoviridae, Polyomaviridae, Potyviridae, Reoviridae, Retroviridae, Rhabdoviridae, Secoviridae, Tombusviridae, Totiviridae, and Virgaviridae.
[0102] In another preferred embodiment, the IRES is a variant of a viral IRES (preferably CVB3, or EMCV, more preferably CVB3) and has a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the nucleotide sequence of the viral IRES.
[0103] Inverted repeat (IR) elements
[0104] Although not necessarily required for circRNA backsplicing, the presence of inverted repeat elements adjacent to splice sites is known to improve circRNA yield. Inverted repeat elements are thought to bring splice sites into close proximity, facilitating circRNA backsplicing.
[0105] In the context of the present invention, an IR element located 5' of the first back splicing site (i.e., the first IR element) and an IR element located 3' of the second back splicing site (i.e., the second IR element) are used to further improve back splicing efficiency.
[0106] In general, any nucleotide sequence followed downstream by its reverse complement may be used as an IR element in the present invention. In some embodiments, the downstream sequence (i.e., the second IR element) is not identical to the reverse complement of the upstream sequence (i.e., the first IR element).
[0107] In some embodiments, the downstream nucleotide sequence (i.e., the second IR element) is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the reverse complement of the upstream sequence (i.e., the first IR element). In preferred embodiments, the first and second IR elements have sequences that are at least 80%, preferably at least 90%, and more preferably at least 95% identical to each other's reverse complements.
[0108] In preferred embodiments, the downstream nucleotide sequence differs from the reverse complement of the upstream sequence by 5, 4, 3, 2, or 1 nucleotide.
[0109] In a preferred embodiment, the IR element has a length of 50 to 1000, preferably 200 to 600, more preferably 300 to 500 nucleotides, In another preferred embodiment, the IR element has a length of 200 to 400 nucleotides.
[0110] The IR elements used in the present invention may be any IR element known in the art, or any of the following: AC004076.9, ACYP2, AMD1, ARHGAP10, ARHGAP12, ARHGEF12, ARHGEF28, ASAP1, ASH1L, ASXL1, ATXN2, BMPR2, BRWD1, BTBD10, CBFA2T2, CCDC126, CCDC134, CCDC66, CCDC7, CCDC9, CCNB1, CDK13, CFLAR, CHD9, CLIP2, CLNS1A, CNN2, COA1, CORO1C, CREBBP, CRKL, CTB-43P18.3, SNHG4, DCUN1D4, DEK, DHRS3, DLG1, DOPEY2, DYNC1H1, ELF2, EMC2, EPHB4, EPS15, ERC1, ETFA, EXOSC1, FAM13B, FARSA, FBXO7, FGD4, FGD6, FKBP3, FKBP8, FNTA, FOXK2, GAPVD1, GBAS, GDI2, GLIS2, GLS, GON4L, GRHPR, HERC1, HIPK3, HNRNPM, HOOK3, HP1BP3, HPS5, HTT, HUWE1, IARS, ILKAP, IQGAP 1, KDM1A, KIAA0368, KIAA1429, KIAA1841, KLHL8, KMT2C, Laccase, LMBR1, LRCH3, LZIC, MAP3K1, MARK4, MBOAT2, MCU, MED13L, METTL3, MGA, MGEA5, MITD 1, MORC3, MRPS35, MYO9B, NCOA2, NFAT5, NFATC3, NFX1, NUDC, NUP54, PAFAH1B2, PAIP2, PCMT1, PDCD11, PDE8A, PDS5A, PHC3, PHLDB2, PLEKHM1, PLEKHM3, P LOD2, PMS1, PNN, POLR2A, POMT1, PPP6R2, PROSC, PRRC2B, PSEN1, PSMA7, PTP4A2, PTPN12, QKI, R3HDM1, RAB6A, RALBP1, RARS, RBM23, RBM33, RBM39, RELL1 , REPS1, RERE, RHOBTB3, RLF, RNF19B, SDHAF2, ZRANB1, FAM228B, RPRD1B, RPS6KC1, RSF1, RSRC1, RTN4, RYK, SAFB2, SAMD4A, SCAF8, SCARF1, SCYL2, SDF4,SEC31A, SENP6, SIPA1L1, SKA3, SLC38A1, SLTM, SMARCA5, SMC3, SMO, SNX25, SOBP, SOS2, SPIDR, SPPL3, SRSF4, STAM, STK3, STX6 , TERF2, TIMMDC1, TMED2, TMEM138, TMEM165, TMEM181, TNPO1, TNPO3, TOP1, TTC39C, TTLL1, UBA2, UBAP2, UBE2K, UBQLN1, UBR5, The IR element may be derived from a gene such as UBXN2A, UBXN7, UHRF2, UIMC1, URI1, UTP18, UTRN, VAMP3, VAPB, VMP1, WDR78, XPO1, YTHDF2, YWHAE, YY1AP1, ZBTB46, ZCCHC11, ZCCHC6, ZFAND6, ZFX, ZKSCAN1, ZMYM4, ZNF124, ZNF236, ZNF394, ZNF430, ZNF652, ZNF720, or ZNF91. Preferably, the IR element is selected from HIPK3 and ZKSCAN1.
[0111] In a preferred embodiment, the IR element is derived from a gene harboring an IR element-derived circRNA.
[0112] In a preferred embodiment, the IR elements are short interspersed nuclear elements (SINEs).
[0113] In a preferred embodiment, the IR element is an Alu repeat element.
[0114] In a preferred embodiment, the first IR element has the sequence set forth in SEQ ID NO: 003. In another preferred embodiment, the first IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 003. In a preferred embodiment, the second IR element has the sequence set forth in SEQ ID NO: 004. In another preferred embodiment, the second IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 004.
[0115] In preferred embodiments, the first IR element has a sequence as contained in the disclosed constructs of the present invention. In other preferred embodiments, the first IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to a sequence as contained in the disclosed constructs of the present invention.
[0116] In preferred embodiments, the second IR element has a sequence as contained in the disclosed constructs of the present invention, hi other preferred embodiments, the second IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to a sequence as contained in the disclosed constructs of the present invention.
[0117] circRNA
[0118] The nucleic acids of the present invention may generally encode any type of circRNA. In a preferred embodiment, the circRNA encoded by the nucleic acids of the present invention or the circRNA of the second aspect of the present invention has a length of 50 nt to 5000 nt, preferably 200 nt to 2000 nt.
[0119] In a preferred embodiment, the circRNA encoded by the nucleic acid of the present invention, or the circRNA of the second aspect of the present invention, does not contain any sequence of a back-splicing site that does not form part of an ORF or IRES sequence.
[0120] In a preferred embodiment, the circRNA encoded by the nucleic acid of the present invention, or the circRNA of the second aspect of the present invention, does not encode a fluorescent protein, preferably does not encode a green or red fluorescent protein.
[0121] In a preferred embodiment, the circRNA encoded by the nucleic acid of the present invention, or the circRNA of the second aspect of the present invention, does not encode a therapeutic protein, a therapeutic peptide, an antigenic protein, or an antigenic peptide.
[0122] In a preferred embodiment, the circRNA has a GC content between 30%-70%, preferably 40% to 60%, more preferably 50% to 60%.
[0123] In a preferred embodiment, the circRNA does not contain an IR element.
[0124] In a preferred embodiment, the circRNA lacks a perfect miRNA target site, which is particularly effective in preventing miRNA-mediated endodegradation.
[0125] In a preferred embodiment, the circRNA lacks cryptic splice sites.
[0126] Further elements of the nucleic acids of the invention
[0127] The nucleic acid of the present invention may further comprise elements having functions such as expression regulation and stability.
[0128] In a preferred embodiment, the nucleic acid comprises a promoter operably linked to the expression cassette to direct expression of the expression cassette. In another embodiment, the nucleic acid further comprises an enhancer. Examples of promoters and enhancers used in the nucleic acids of the present invention include the SV40 early promoter and enhancer (Mizukami T. et al. 1987), the Moloney murine leukemia virus LTR promoter and enhancer (Kuwana Y et al. 1987), the cytomegalovirus (CMV) immediate early promoter and enhancer, and the like.
[0129] In another preferred embodiment, the nucleic acid molecule further comprises a branch site.
[0130] In another preferred embodiment, the nucleic acid molecule further comprises a polypyrimidine sequence.
[0131] Vectors of the present invention
[0132] A third aspect of the present invention refers to a vector comprising the nucleic acid molecule according to the first aspect of the present invention, or the circRNA according to the second aspect of the present invention.
[0133] In the present invention, a vector refers to any type of element that can contain the nucleic acid or circRNA of the present invention. Any suitable vector, such as a plasmid, cosmid, episome, liposome, exosome, artificial chromosome, phage, or virus-derived vector, can be used in the present invention.
[0134] The terms "vector," "cloning vector," and "expression vector" refer to a vehicle that can introduce DNA or RNA sequences (e.g., nucleic acids of the invention, or circRNA) into a host cell, transforming the host and promoting expression (e.g., transcription, circularization, and translation) of the introduced nucleic acid or circRNA.
[0135] Such vectors may contain control elements, such as promoters, enhancers, terminators, and the like, to induce or direct expression of the nucleic acid administered to a subject. Examples of promoters and enhancers used in expression vectors for animal cells include the SV40 early promoter and enhancer (Mizukami T. et al. 1987), the Moloney murine leukemia virus LTR promoter and enhancer (Kuwana Y et al. 1987), the cytomegalovirus (CMV) promoter (Mason JO et al. 1985) and enhancer (Gillies SD et al. 1983), and the like.
[0136] Any expression vector for animal cells can be used as long as it can insert and express the nucleic acid of the present invention. Examples of suitable vectors include pAGE107 (Miyaji H et al. 1990), pAGE103 (Mizukami T et al. 1987), pHSG274 (Brady G et al. 1984), pKCR (O'Hare K et al. 1981), pSG1 beta d2-4- (Miyaji H et al. 1990), and the like. Other examples of plasmids include replicating plasmids containing an origin of replication and the like.
[0137] Host Cells of the Invention
[0138] A fourth aspect of the present invention relates to a host cell (preferably by transformation, transduction or transfection of the host cell) comprising the nucleic acid, circRNA, and / or vector according to the present invention.
[0139] The term "transformation" originally refers to the naturally occurring process of gene transfer into a host cell, which involves the cell absorbing genetic material, such as nucleic acid, e.g., DNA or RNA, through the cell membrane, and the host cell expressing the introduced gene or sequence to produce the desired substance. There are two types, called natural transformation and artificial, or induced, transformation. Artificial, or induced, methods of transformation are performed under laboratory conditions.
[0140] A host cell that receives and expresses foreign nucleic acid, such as DNA or RNA, through the process of transformation has been "transformed."
[0141] The term "transfection," as used herein, refers to a form of gene transfer that involves creating holes in the host cell's plasma membrane, allowing the host cell to receive foreign genetic material. Typically, transfection refers to the transformation of eukaryotic cells, such as insect or mammalian cells. Chemical-mediated transfection involves the use of, for example, calcium phosphate, cationic polymers, or liposomes. Non-chemically mediated transfection methods are typically electroporation, sonoporation, impale infection, optical transfection, or hydrodynamic delivery. Particle-based transfection uses gene gun technology, which uses nanoparticles to introduce nucleic acids into host cells, or another method known as magnetofection. Nucleofection and the use of heat shock are other advanced methods for successful transfection. Host cells that receive foreign nucleic acids via transfection methods are "transfected."
[0142] The term "transduction," as used herein, is generally understood to refer to the introduction of foreign nucleic acid, such as DNA or RNA, into a cell by a virus or virus-derived vector. A host cell that receives and expresses foreign nucleic acid, such as DNA or RNA, by a virus or virus-derived vector has been "transduced."
[0143] Pharmaceutical Composition
[0144] A fifth aspect of the present invention refers to a pharmaceutical composition comprising the nucleic acid, circRNA, vector, and / or host cell of the present invention and a pharmaceutically acceptable carrier.
[0145] The term "pharmaceutical composition" or "therapeutic composition," as used herein, refers to a compound or composition capable of inducing a desired therapeutic effect when properly administered to a subject.
[0146] In some embodiments, a subject may refer to a patient.
[0147] Such therapeutic or pharmaceutical compositions may comprise a therapeutically effective amount of the nucleic acids, circRNAs, vectors, and / or host cells of the present invention, or may further comprise a therapeutic agent in admixture with a pharmaceutically or physiologically acceptable formulation selected to be appropriate for the mode of administration.
[0148] The nucleic acids, circRNAs, vectors, and / or host cells of the invention will typically be supplied as part of a sterile pharmaceutical composition, which will typically include a pharmaceutically acceptable carrier.
[0149] "Pharmaceutically" or "pharmaceutically acceptable" refers to molecular entities and compositions that do not produce adverse, allergic, or other untoward reactions when properly administered to mammals, especially humans. A pharmaceutically acceptable carrier, or excipient, refers to a non-toxic solid, semi-solid, or liquid excipient, diluent, encapsulating material, or auxiliary of any type.
[0150] A "pharmaceutically acceptable carrier" may also be referred to as a "pharmaceutically acceptable diluent" or a "pharmaceutically acceptable vehicle," and may include physiologically compatible solvents, bulking substances, stabilizing substances, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying substances, and the like. Thus, in one embodiment, the carrier is an aqueous carrier.
[0151] The form, route of administration, dose, and regimen of the pharmaceutical composition will naturally depend on the condition to be treated, the severity of the disease, the age, weight, and sex of the patient, the desired duration of treatment, etc. The pharmaceutical composition may be in any suitable form (depending on the desired method of administration to the patient). It may be provided in unit dosage form, which will generally be provided in a hermetically sealed container, or may be provided as part of a kit. Such a kit will usually (but not necessarily) include instructions for use. It may contain a plurality of said unit dosage forms.
[0152] Empirical considerations, such as biological half-life or bioavailability, will generally contribute to determining the dosage. The frequency of administration may be determined and adjusted over the course of treatment. Alternatively, a continuous sustained-release formulation may be appropriate.
[0153] Other aspects of the invention
[0154] The present invention further comprises a method for producing circular mRNA (circRNA), the method comprising introducing the nucleic acid molecule of the first aspect of the invention, or the vector of the third aspect of the invention, into a eukaryotic cell, and optionally purifying the circRNA from the cell.
[0155] In another aspect of the present invention, a method for producing circular mRNA (circRNA) comprises contacting the nucleic acid molecule of the first aspect of the present invention, or the vector described in the third aspect of the present invention, with isolated RNA polymerase II.
[0156] In yet another aspect, the present invention relates to a method for producing a recombinant protein, the method comprising introducing a nucleic acid molecule of the first aspect of the present invention, a circRNA according to the second aspect, or a vector according to the third aspect into a eukaryotic cell, and optionally purifying the recombinant protein encoded by said ORF.
[0157] In yet another aspect of the present invention, we claim circRNAs obtained by the methods of the present invention. [Example]
[0158] The present invention provides nucleic acids encoding circRNAs that utilize spliceosome-based backsplicing for cellular circRNA expression. To promote backsplicing, flanking regions of highly expressed circRNAs derived from the HIPK3 locus are used. Furthermore, the introduction of an IRES element is required to promote circRNA cap-independent protein translation. While both CVB and EMCV IRES elements resulted in improved circRNA expression using the nucleic acid design of the present invention, the CVB3 IRES was found to result in superior protein expression compared to the EMCV IRES. Surprisingly, the inventors found that the positioning of the IRES within the circRNA expression cassette significantly affected circRNA expression at both the RNA and protein levels, suggesting improved backsplicing. Positioning the IRES element close to the flanking regions of the circRNA cassette was found to suppress circRNA expression. Without wishing to be bound by theory, we therefore hypothesize that highly structured IRES elements may interfere with splice site recognition and spliceosome assembly, thereby inhibiting the back-splicing reaction and circRNA production. This is supported by the observation that adding a non-coding spacer sequence between the SA and IRES, increasing the distance of the IRES from the splice site, restores circRNA production in a distance-dependent manner. An increase in the distance between both features positively correlates with circRNA expression.
[0159] The introduction of additional non-coding sequences for transgene protein expression appears unnecessary because they provide no additional function other than promoting efficient backsplicing and may reduce the payload / insert capacity for expression of the gene of interest. To overcome this, split-ORF circRNA designs were tested. The transgene payload was split at either the CAGG or AAGG motif within the ORF, which best reflects the endogenous splicing reaction. We found that splitting the ORF significantly improved circRNA production compared to inserting a continuous ORF immediately upstream or downstream of the IRES. Furthermore, we found that splitting the ORF more centrally, thereby increasing the distance between the IRES and the splice site, improved circRNA production compared to a continuous ORF design and compared to splitting the ORF at the 5' and 3' ends. This supports the hypothesis that the proximity of the IRES to the splice site negatively impacts circRNA backsplicing. Further supporting this, we found that splitting the IRES element itself and inserting an ORF within the IRES also inhibited efficient circRNA production.
[0160] Finally, the split ORF design was found to be superior across multiple transgenes when compared to the continuous ORF design, suggesting that the circRNA expression design of the present invention is a powerful tool for transgene expression.
[0161] General Methods Used in the Examples
[0162] Constructs, cell culture, and treatments:
[0163] Plasmids were synthesized by Genscript using the pcDNA3.1 backbone vector for mammalian expression. The A375 human melanoma cell line was obtained from ATCC and maintained in Dulbecco's modified Eagle's medium (Thermo Fisher Scientific) supplemented with 10% fetal bovine serum (Thermo Fisher Scientific) and antibiotics (100 U / ml penicillin and 100 mg / ml streptomycin; Thermo Fisher Scientific, Inc.). All cell types were incubated at 37°C in a 5% CO2 humidified incubator.
[0164] Cells were seeded at ~150,000 cells per well in a 6-well dish 24 hours prior to transfection. Cells were transfected using Lipofectamine 3000 according to the manufacturer's guidelines. Briefly, 2 μg of circRNA or pcDNA3.1 linear expression vector was transfected per well. Cells were incubated overnight with the transfection mix before changing the medium. Unless otherwise specified, RNA and protein were harvested 48 hours post-transfection. In all experiments, cells were harvested by washing with 1x PBS followed by centrifugation at 1200 rpm for 4 minutes at 4°C. 66.6% of the harvested cells were used for RNA isolation, which was performed using TRIzol Reagent (Thermo Fisher Scientific) according to the manufacturer's protocol.
[0165] RT-PCR and RT-qPCR:
[0166] Using the M-MLV Reverse Transcriptase Kit (Thermo Fisher Scientific), 1 μg of DNase-treated total RNA was reverse transcribed using random hexamers to prime the reaction according to the manufacturer's protocol. For RT-PCR, reactions were performed with or without RT enzyme for 25 cycles. Products were visualized by 1% agarose gel electrophoresis and confirmed by Sanger sequencing. For quantitative PCR, cDNA was mixed with the Platinum SYBR Green I Master Kit (Invitrogen) and run on a 7500 Fast Real-Time PCR System (Applied Biosystems). Reactions were performed in technical triplicates. The Ct values obtained for each triplicate were transformed (2-Ct) and averaged (σ). All samples were normalized to GAPDH. Results were visualized as bar graphs using GraphPad (Prism7), showing individual biological replicates and standard deviations.
[0167] Western blotting:
[0168] Cells were collected in 1x PBS and centrifuged at 1200 rpm for 5 minutes at 4°C. For cell lysis, the cell pellet was collected and resuspended in 50 μL of RIPA buffer supplemented with protease inhibitors. The lysate was incubated for 20 minutes, followed by centrifugation at 12,000 x g for 15 minutes at 4°C to remove cell debris. The supernatant was collected, and protein levels were quantified using a BCA assay (Thermo Fisher Scientific). For Western blot analysis, 3 μg of protein was used. Prior to Western blotting, an equal volume of 2x SDS loading buffer [125mM Tris-HCl pH 6.8, 20% glycerol, 5% SDS, and 0.2M DTT] was added to 3µg of protein lysate and briefly boiled at 95°C for 5 minutes before loading onto a 12% Tris-glycine SDS-PAGE gel (Thermo Fisher Scientific) and running at 125V for approximately 1.5 hours. Proteins were transferred to a PVDF membrane (BioRad) by wet blotting for 2 hours at 4°C and 30V. The membrane was then preblocked with 10% skim milk for 1 hour at RT, followed by incubation with primary antibody for 1 hour and secondary antibody for 1 hour. After each antibody incubation, the membrane was rinsed 3 times for 5 minutes with 1x PBS + 0.05% Tween 20 and washed 1 time for 5 minutes with 1x PBS. Protein bands were developed using the SuperSignal West Femto Maximum Sensitivity Substrate kit (Thermo Fisher Scientific), and ChemiDoc XRS+ (Biorad).
[0169] Antibodies used:
[0170] ICOS-L(1:500;ab233151, Abcam) Renilla (1:2,000;ab185925, Abcam) eGFP (1:5,000; ab290, Abcam) FLAG(1:10,000;F3165, Sigma-Aldrich) Beta-Actin (1:20,000; A5441, Sigma-Aldrich) Rabbit Secondary(1:10,000;ab6721, Abcam) Mouse Secondary(1:10,000;ab90723, Abcam)
[0171] Example 1: Design of circRNA expression vectors
[0172] We designed several vectors containing circRNA expression cassettes using different IRES-ORF locations (see Figure 1A) to encode different transgenes from circular RNA transcripts. All expression cassettes shared a common promoter (CMV) and were flanked by inverted repeat elements (IR-HIPK3) to facilitate backsplicing and circRNA production. Because the HIPK3 gene encodes one of the most endogenously expressed circRNAs across multiple human cell types and tissue types, we utilized the flanking regions of HIPK3 in the design of expression vectors. Furthermore, to facilitate backsplicing, the circRNA exons were also flanked by splice sites. Other features, such as branch site sequences and polypyrimidine tracts, are involved in ensuring efficient splicing and are derived from the flanking regions of circHIPK3. The open reading frame (ORF) (the coding region of the circRNA) was designed using two different strategies: one was to place the ORF units contiguously downstream of the IRES, and the other was to split the ORF and flank it with an IRES element (Figure 1A). To design and generate split ORFs, ideally, a centrally located CAGG or AAGG motif within the ORF should be identified, splitting the ORF between two GG nucleotides (where one ORF ends with CAG and the other begins with G). The rationale for the design is based on the human consensus sequences of splice sites, splice donor (SD) splice sites, and splice acceptor (SA) splice sites recognized by the endogenous splicing machinery. In endogenous splicing, the splice donor binds to the splice acceptor, and the CAGG motif shows the highest level of conservation at these sites. An IRES element is then inserted between the split ORFs, with the C-terminal portion of the gene placed upstream and the N-terminus downstream of the IRES (Figure 1A). Therefore, the effect of specific properties, particularly IRES selection and IRES positioning, on circRNA production was examined.
[0173] Natural circRNAs lack the 5' cap and 3' poly(A) tail required for cap-dependent translation and are therefore considered non-coding. However, the introduction of an IRES element into a circRNA molecule has been shown to direct circRNA translation via a cap-independent mechanism. Therefore, the protein-coding ability of circRNAs containing IRES elements from CVB3 or EMCV was tested. As described above, circRNA expression vectors encoding ICOS-L were tested using both continuous and split ORF designs containing either the CVB3 or EMCV IRES (see Table 3 below for the ICOS-L split ORF). ICOS-L protein levels were examined 48 hours after transfection in A375 cells transfected with plasmids encoding circICOSL in continuous or split designs containing either the EMCV or CVB3 IRES. Western blot analysis showed that the CVB3 IRES stimulated superior protein production compared to the EMCV IRES (Figure 1B). Interestingly, the split ORF design combined with the CVB3 IRES induced significantly higher levels of ICOS-L protein expression compared to the continuous ORF (Figure 1B). Furthermore, the continuous ORF design combined with the EMCV IRES did not produce detectable protein. Furthermore, both RT-PCR (Figure 1C) and qRT-PCR (Figure 1D) analyses showed that circRNA levels were significantly higher from the split ORF design compared to the continuous ORF design. This suggests that the positioning of the IRES within the circRNA expression cassette has a significant impact on circRNA production.
[0174] Example 2: Distancing an IRES from a back-splicing site affects circRNA expression
[0175] To further investigate the impact of split ORF design on circRNA production, circEGFP expression vectors were synthesized. The expression vector containing all elements was described in Example 1 above. The eGFP ORF was split at several positions that varied the distance of the IRES (here, CVB3) from the adjacent splice site. The circEGFP designs were designated based on the distance of the IRES from the upstream SA site (i.e., the first backsplicing site). For example, in Split_59, the IRES was inserted 59 nt downstream of the SA (Figure 2A; Table 1).
[0176] A375 cells were transfected with plasmids encoding different split ORF designs of circEGFP. RNA and protein levels were examined 48 hours after transfection. CircRNA expression from each construct was confirmed by RT-PCR using divergent primers that only amplify the circular RNA transcript (Figure 2B). Quantitative RNA and protein expression analysis by qRT-PCR (Figure 2C) and Western blot (Figure 2D) respectively demonstrated a negative correlation between the proximity of the IRES to the splice site and circRNA expression. Splitting the ORF near the 5' end (Split_18 and Split_59) or 3' end (Split_650 and Split_732) resulted in reduced circRNA expression at both the RNA and protein levels compared with more central splits (Split_255 and Split_363) (Figure 2C-D). This suggests that a distance of ~250 nt between the IRES and splice site promotes high levels of circRNA expression. However, even with a distance of ~59-94 nt between the IRES and splice site, circRNA production is still detectable at proportionally higher levels than with a distance of ~13-18 nt. [Table 1] Table 1: Nucleotide distances between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP split-circRNA expression cassette design.
[0177] Example 3: Splitting of IRES
[0178] Furthermore, we tested whether splitting the IRES element itself affected circRNA expression. The expression vector containing all elements was described in Example 1 above. The IRES element was split at different positions of the CAGG / AAGG motif to enable circRNA splicing and circularization without introducing additional nucleotides into the IRES sequence. In this experiment, a continuous eGFP ORF was inserted into the split IRES element, and the IRES fragment would be immediately adjacent to the backsplicing site involved in circRNA production (Figure 3A). The split IRES designs were named based on the distance from the start of the IRES (the first nucleotide encoding the IRES) to the SD. The following expression constructs were tested: Split_i33 (SEQ ID NO: 012), Split_i219 (SEQ ID NO: 013), and Split_i317 (SEQ ID NO: 014). A375 cells were transfected with plasmids encoding different split IRES designs of circEGFP. CircRNA expression was assessed 48 hours after transfection by RT-PCR (Figure 3B) and qRT-PCR (Figure 3C). Western blotting was performed to assess eGFP protein expression (Figure 3D). Splitting the IRES significantly reduced circRNA expression compared to the split-ORF design (see Figures 3B-D). In fact, the split-IRES design consistently demonstrated the lowest levels of circRNA expression at both the RNA and protein levels. Therefore, splitting the IRES element and inserting an internal continuous ORF does not appear to be a suitable system for high-yield circRNA expression.
[0179] Example 4: Proximity of an IRES to adjacent back-splicing sites affects circRNA expression
[0180] To further examine the effect of IRES positioning relative to the backsplicing site on circRNA expression, we synthesized eGFP circRNA expression vectors based on the circONCOS backbone, i.e., containing identical flanking sequences and splice sites. Here, we examined the relationship between IRES proximity to splice sites. A circRNA expression cassette was synthesized in which a contiguous eGFP ORF was inserted downstream of the CVB3 IRES element. Additionally, a non-coding "spacer" sequence derived from HIPK3 exon 2 was inserted into the flanking region of the IRES-ORF cassette (Figure 4A). To ensure that each encoded circRNA was the same size, a 403-nt spacer sequence was added to each design, although the amount of spacer sequence inserted upstream and downstream of the IRES-ORF varied between circRNA designs (Figure 4A). Designs were named based on the size of the HIPK3 stuffer sequence inserted upstream of the IRES. Namely, circH1 had a 1-nt insertion upstream of the IRES and a 402-nt insertion downstream of the ORF (Table 2). A375 cells were transfected with a plasmid encoding circEGFP containing different HIPK3-derived spacer sequences upstream of the adjacent IRES / ORF. Both RT-PCR (Figure 4B) and qRT-PCR (Figure 4C) circRNA expression analysis clearly demonstrated a relationship between the distance of the IRES from the upstream SA and the level of circRNA expression, with increasing distance of the IRES from the SA positively correlating with circRNA expression. In other words, circH1, circH50, and circH100 showed significantly decreased levels of circRNA expression compared with circH200 and circH400 (Figure 4C). RT-PCR indicated that the levels of circRNA expression were comparable between circH200 and circH400. Overall, the data suggest that placing the IRES away from the SA increases circRNA yield. [Table 2] Table 2: Nucleotide distance between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP_HIPK3 spacer circRNA expression cassette design.
[0181] Example 5: Split ORF design provides improved circRNA yield and higher protein translation
[0182] To generalize our observation that the split-ORF design is a superior circRNA expression design independent of a specific ORF, we tested circRNA expression from constructs encoding different transgenes. Three additional transgenes were tested: those encoding the costimulatory molecule ICOS-L, the adenosine deaminase ADA, and the reporter gene Renilla luciferase. Three circRNA designs were tested: those in which the continuous transgene ORFs were placed upstream (circInv) or downstream (circCont) of the IRES, and a central split-ORF design (circSplit) in which at least 30% of the ORF was placed between the 3' end of the IRES and the SD (Figure 5A; Table 3). For ICOS-L, an additional split design was tested. For ICOS-L, both cassettes were split by the same motif, but version 2 contained a short spacer sequence upstream of the start codon. Furthermore, RNA and protein levels were compared with the corresponding linear mRNA.
[0183] A375 cells were transfected with plasmids encoding circular and linear forms of ICOS-L (SEQ ID NO: 21), Renilla (SEQ ID NO: 22), and FLAG-tagged ADA (SEQ ID NO: 20). Western blot analysis of protein expression from the circRNA cassettes showed that expression levels from the circSplit design were significantly higher than those derived from either the circInv or circCont designs (Figures 5B-D). Furthermore, across all three transgenes, qRT-PCR analysis showed that the circSplit design conferred superior circRNA expression compared to the circInv or circCont designs (Figures 5E-G). These observations support the hypothesis that splitting the ORF centrally and placing the CVB3 IRES within the ORF results in superior circRNA expression at both the RNA and protein levels. [Table 3] Table 3: Nucleotide distance between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP_HIPK3 spacer circRNA expression cassette design.
[0184] Furthermore, we observed that for each transgene, protein expression levels from the circSplit design were comparable (Figure 5B ADA; Figure 5D Renilla) or superior (Figure 5C ICOS-L) to their linear mRNA counterparts. Strikingly, these comparable / superior protein levels were observed even though the circRNA RNA levels were significantly lower compared to the corresponding mRNAs (Figure 5E-G). This suggests that the CVB3 IRES in combination with the split-ORF design is a powerful tool for transgene protein expression, potentially outperforming conventional linear mRNA expression cassettes even at early time points when mRNA expression is highest.
[0185] Example 6: Positioning of an IRES within a circRNA cassette affects circRNA expression regardless of the gene being expressed or the IRES used
[0186] As shown in the examples above, positioning an IRES within a circRNA dramatically enhances circRNA expression. These findings equally extend to a different gene, ICOSL. Similar to the above experiments, we split the ORF of ICOSL at different positions, placing the IRES between the stop codon (upstream of the IRES) and the start codon (downstream of the IRES), and adjusting the distance between the IRES and the splice acceptor (SA) and splice donor (SD) sites. The exact split positions are shown in Table 4.
[0187] Briefly, A375 cells were transfected with plasmids encoding different split ORF designs of circICOSL. RNA and protein levels were examined 48 hours after transfection. Protein and circRNA expression analyses were performed by Western blot (Figure 6A) and qRT-PCR (Figure 6B), respectively. As previously shown, there was a negative correlation between the proximity of the IRES to the splice site and circRNA expression. Splitting the ORF near the 5' end (Split_18) or 3' end (Split_870) reduced circRNA expression at both the RNA and protein levels compared with more central splits (Split_236 and Split_436). Furthermore, as previously shown, splitting the IRES (i33, i219, i317) reduced circRNA expression at both the protein and RNA levels (Figure 6A-B). Taken together, these results further support the observation that placing the IRES in the center of a split ORF, resulting in equivalent distances between the IRES and either the SA or SD, results in the highest protein expression levels. [Table 4] Table 4: Nucleotide distance between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circICOSL split-circRNA expression cassette design.
[0188] Further supporting this observation, when the CVB3 IRES (741 nt; group IRES) used in the previous example was replaced with the EMCV IRES (566 nt; group II IRES) in the eGFP-encoding circRNA vector (Table 5), a similar circEGFP expression pattern was observed. A375 cells were transfected with plasmids encoding different split-ORF designs of circEGFP carrying either the CVB3 (Figure 7A-B) or EMCV (Figure 7C-D) IRES. Protein and RNA levels were examined 48 hours after transfection. Protein and circRNA expression analyses were performed by Western blot (Figure 7A and C) and qRT-PCR (Figure 7B and D), respectively. This supported the importance of ORF-IRES positioning for high-yield circRNA expression, regardless of the IRES element used. Similar circRNA expression patterns were observed for both the CVB3 and EMCV circRNA designs, with more central splits (Split_363 and Split_436) resulting in the highest levels of circRNA expression at both the protein and RNA levels. Furthermore, similar to the results in previous examples, splitting the ORF closer to the 5' and 3' ends negatively impacts circRNA biogenesis. Furthermore, similar to what was observed in the CVB3 IRES split design (Figure 7A-B), splitting the EMCV IRES and placing the ORF within this region also inhibited high-yield circRNA expression (Figure 7C-D). [Table 5] Table 5: Nucleotide distance between the IRES, or eGFP_ORF, and the upstream splice acceptor (SA), or downstream splice donor (SD) in the design of each circEGFP split-circRNA expression cassette.
[0189] Overall, the data support the hypothesis that central IRES placement in a split ORF design confers the highest levels of circRNA expression. We hypothesize that because a central split allows for maximizing the distance of the IRES from the splice sites on either side, the increased expression is a result of increasing the distance of the IRES from adjacent splice sites required for backsplicing.
Claims
1. 1. A nucleic acid molecule encoding a circular RNA (circRNA), the nucleic acid molecule comprising: A) An expression cassette comprising: (i) a circRNA expression cassette comprising: (a) a nucleic acid comprising or consisting of a continuous or split open reading frame (ORF) encoding at least one protein; and (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF; (ii) a first back-splicing site located 5′ to the IRES; and (iii) a second back-splicing site located 3′ to the IRES; and, optionally, B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site; and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette; The IRES is located within the circRNA expression cassette as follows: (I) the 5' end of the IRES and the 3' end of the first back-splicing site are separated by at least 50 nucleotides, preferably at least 200, more preferably at least 300 nucleotides; (II) the 3' end of the IRES and the 5' end of the second back-splicing site are separated by at least 300 nucleotides, preferably at least 350 nucleotides; Nucleic acid molecule.
2. The nucleic acid molecule of claim 1, wherein the split ORF comprises two parts, the first part of the split ORF comprising a stop codon and located 5' to the IRES, and the second part of the split ORF comprising a start codon and located 3' to the IRES.
3. the first back-splicing site comprises two half sites; the first and second half-sites are located 5' to the circRNA expression cassette; or - the first half-site is located 5' to the circRNA expression cassette and the second half-site is part of the circRNA expression cassette, preferably part of the ORF in the case of a split ORF; the second back-splicing site comprises two half sites; the first and second half-sites are located 3' to the circRNA expression cassette; or the first half-site is part of the circRNA expression cassette, preferably part of the ORF in the case of a split ORF, and the second half-site is located 3' to the circRNA expression cassette; A nucleic acid molecule according to any one of the preceding claims.
4. the first half site of the first back-splicing site comprises or consists of the nucleotides AG; and / or the second half site of the second backsplicing site comprises or consists of the nucleotides GT, A nucleic acid molecule according to any one of the preceding claims.
5. the IRES is a virus-derived IRES, Preferably, the virus is selected from the group of Coxsackievirus (CVB), preferably CVB3, and Encephalomyocarditis virus (EMCV); or Preferably, the virus is selected from the group consisting of Adenoviridae, Arenaviridae, Birnaviridae, Chrysoviridae, Coronaviridae, Dicistroviridae, Filoviridae, and the like. Flaviviridae, Hepadnaviridae, Herpesviridae, Hypoviridae, Iflaviridae, Luteoviridae, Orthomyxoviridae, Papillomaviruses, Papillomaviridae, Paramyxoviridae, Parvoviridae, Picornaviridae, Pneumoviridae, Polyomaviridae, Potyviridae ), Reoviridae, Retroviridae, Rhabdoviridae, Secoviridae, Tombusviridae, Totiviridae, Virgaviridae; or Preferably, it is a variant of a viral IRES, having a nucleotide sequence that is at least 90% identical to the nucleotide sequence of the viral IRES. A nucleic acid molecule according to any one of the preceding claims.
6. 10. The nucleic acid molecule of any one of the preceding claims, wherein the first and second IR elements comprise sequences that are at least 80%, preferably at least 90%, more preferably at least 95% identical to each other's reverse complements.
7. 10. The nucleic acid molecule of any one of the preceding claims, wherein the IR element is selected from the group consisting of IR elements derived from the following genes: AC004076.9, ACYP2, AMD1, ARHGAP10, ARHGAP12, ARHGEF12, ARHGEF28, ASAP1, ASH1L, ASXL1, ATXN2, BMPR2, BRWD1, BTBD10, CBFA2T2, CCDC126, CCDC134, CCDC66, CCDC7, CCDC9, CCNB1, CDK13, CFLAR, CHD9, CLIP2, CLNS1A, CNN2, CO A1, CORO1C, CREBBP, CRKL, CTB-43P18.3, SNHG4, DCUN1D4, DEK, DHRS3, DLG1 , DOPEY2, DYNC1H1, ELF2, EMC2, EPHB4, EPS15, ERC1, ETFA, EXOSC1, FAM13B, F ARSA, FBXO7, FGD4, FGD6, FKBP3, FKBP8, FNTA, FOXK2, GAPVD1, GBAS, GDI2, G LIS2, GLS, GON4L, GRHPR, HERC1, HIPK3, HNRNPM, HOOK3, HP1BP3, HPS5, HTT, H UWE1, IARS, ILKAP, IQGAP1, KDM1A, KIAA0368, KIAA1429, KIAA1841, KLHL8, KMT2C, LMBR1, LRCH3, LZIC, MAP3K1, MARK4, MBOAT2, MCU, MED13L, METTL3, MG A, MGEA5, MITD1, MORC3, MRPS35, MYO9B, NCOA2, NFAT5, NFATC3, NFX1, NUDC, NUP54, PAFAH1B2, PAIP2, PCMT1, PDCD11, PDE8A, PDS5A, PHC3, PHLDB2, PLEKH M1, PLEKHM3, PLOD2, PMS1, PNN, POLR2A, POMT1, PPP6R2, PROSC, PRRC2B, PSE N1, PSMA7, PTP4A2, PTPN12, QKI, R3HDM1, RAB6A, RALBP1, RARS, RBM23, RBM33 , RBM39, RELL1, REPS1, RERE, RHOBTB3, RLF, RNF19B, SDHAF2, ZRANB1, FAM228 B, RPRD1B, RPS6KC1, RSF1, RSRC1, RTN4, RYK, SAFB2, SAMD4A, SCAF8, SCARF1,SCYL2, SDF4, SEC31A, SENP6, SIPA1L1, SKA3, SLC38A1, SLTM, SMARCA5, SMC3, SMO, SNX25, SOBP, SOS2, SPIDR, SPPL3, SRSF4, STAM, STK3, STX6, TERF2, TIMMDC1, TMED2, TMEM138, TMEM165, TMEM181, TNPO1, TNPO3, TOP1, TTC39C, TTLL1, UBA2, UBAP2, UBE2K, UBQLN 1, UBR5, UBXN2A, UBXN7, UHRF2, UIMC1, URI1, UTP18, UTRN, VAMP3, VAPB, VMP1, WDR78, XPO1, YTHDF2, YWHAE, YY1AP1, ZBTB46, ZCCHC11, ZCCHC6, ZFAND6, ZFX, ZKSCAN1, ZMYM4, ZNF124, ZNF236, ZNF394, ZNF430, ZNF652, ZNF720, ZNF91], preferably HIPK3, and ZKSCAN1.
8. 10. The nucleic acid molecule of any one of the preceding claims, wherein the promoter is selected from a constitutive, tissue-specific or inducible promoter, preferably a viral-derived promoter, preferably selected from the group consisting of: CMV immediate early promoter, SV40 promoter.
9. 10. The nucleic acid molecule of any one of the preceding claims, wherein the circRNA expression cassette further comprises: (i) one or more non-coding nucleotide sequences, where if the circRNA expression cassette comprises a continuous ORF, the non-coding nucleotide sequences are preferably located at the 5' and / or 3' end of the continuous ORF, or where if the circRNA expression cassette comprises a split ORF, the non-coding nucleotide sequences are preferably located at the 5' and / or 3' end of an IRES; (ii) one or more additional ORFs; (iii) one or more additional IRES; (iv) Additional coding nucleotide sequences associated with one or more ORFs, preferably encoding sequences that can be used for the detection or capture of proteins encoded by one or more ORFs.
10. 10. The nucleic acid molecule of any one of the preceding claims, wherein the nucleic acid molecule further comprises: (i) a branching site; and / or (ii) a polypyrimidine sequence;
11. 10. The nucleic acid molecule of any of the preceding claims, wherein the ORF encodes a therapeutic protein, a therapeutic peptide, an antigenic protein, or an antigenic peptide.
12. A circRNA encoded by a circRNA expression cassette of any one of claims 1 to 11.
13. A vector comprising a nucleic acid molecule according to any one of claims 1 to 11 or a circRNA according to claim 12.
14. A host cell comprising a nucleic acid molecule according to any one of claims 1 to 11, a circRNA according to claim 12, or a vector according to claim 13.
15. A pharmaceutical composition comprising a nucleic acid molecule according to any one of claims 1 to 11, a circRNA according to claim 12, or a vector according to claim 13.
16. A method for producing circular mRNA (circRNA), the method comprising: (i) introducing the nucleic acid molecule of any one of claims 1 to 11 or the vector of claim 13 into a eukaryotic cell, and optionally purifying circRNA from the cell; or (ii) contacting the nucleic acid molecule of any one of claims 1 to 11 or the vector of claim 13 with isolated RNA polymerase II.
17. 14. A method for producing a recombinant protein, the method comprising introducing a nucleic acid molecule of any one of claims 1 to 11, a circRNA of claim 12, or a vector of claim 13 into a eukaryotic cell, and optionally purifying the recombinant protein encoded by the ORF.