CircRNA expression constructs
By maintaining an appropriate distance between IRES and the splicing site in the circRNA expression cassette, the back-splicing efficiency was optimized, the problem of unstable circRNA expression in vivo was solved, and high-yield and efficient translation of the encoded protein was achieved.
Patent Information
- Application Number
- CN202380092849.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-07
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies make it difficult to achieve high-level stable circRNA expression in vivo, and the IRES position is crucial for circRNA production.
Expression cassettes containing IRES are designed by positioning the IRES at a specific position relative to the splice site required for circRNA backsplicing and maintaining a certain distance to optimize backsplicing efficiency.
The efficiency and yield of circRNA backsplicing are improved, ensuring that only the encoded protein is translated from the circRNA and avoiding the translation of linear RNA.
Smart Images

Figure CN120603946A_ABST
Abstract
Description
[0001] The present invention relates to a nucleic acid molecule encoding a circular RNA (circRNA). In another aspect, the present invention relates to a circRNA encoded by the nucleic acid, and a vector and a host cell comprising the nucleic acid. The present invention also includes a method for producing the circRNA. Background of the Invention
[0003] Circular RNAs (circRNAs) constitute a novel class of long noncoding RNAs characterized by covalently closed molecules. In contrast to traditional linear splicing, circRNAs are typically generated through nonlinear "backsplicing" events using a downstream splice donor (SD) and an upstream splice acceptor (SA). Endogenous circular RNAs have been extensively studied over the past decade. Although their functional relevance remains controversial, it is generally accepted that circRNAs, due to their circular nature, are resistant to exonuclease degradation. Consequently, circRNAs comprise a class of highly stable RNAs with a half-life far exceeding that of traditional linear mRNAs.
[0004] Over the past decade, significant progress has been made in understanding the mechanisms of circRNA production in vitro and in vivo, highlighting the therapeutic potential of engineered, functionally robust circular RNAs. Broadly speaking, circRNA production can be achieved through two distinct mechanisms: 1) by group I intron-derived ribozyme circularization, which has been shown to be effective for in vitro production, or 2) by spliceosome-based backsplicing, which, similar to the biogenesis of endogenous circRNAs, favors in vivo production. In the latter mechanism, insertion of flanking inverted elements is known to significantly stimulate backsplicing, presumably by positioning the relevant splice sites in close proximity. Although circRNAs lack a 5' cap structure and a 3' polyadenylation site and are therefore not themselves translational substrates, insertion of an IRES (internal ribosome entry site) effectively converts noncoding circRNAs into efficient protein-coding molecules. Consequently, the therapeutic use of circRNAs as templates for robust protein production has garnered considerable attention.
[0005] The present invention provides a circRNA expression cassette that is improved to promote high-level stable circRNA expression in vivo. Surprisingly, studies have found that the position of the IRES in the circRNA expression cassette has a significant correlation with the yield of the circRNA. Summary of the Invention
[0006] Disclosed herein is a nucleic acid encoding a circular RNA (circRNA), wherein an IRES that drives translation of the circRNA is positioned in a specific position relative to a splice site required for back-splicing of the circRNA, resulting in excellent back-splicing efficiency.
[0007] In a first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises:
[0008] A) An expression cassette comprising:
[0009] (i) a circRNA expression cassette comprising:
[0010] (a) a nucleic acid comprising or consisting of consecutive or separate open reading frames (ORFs) encoding at least one protein, and
[0011] (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF;
[0012] (ii) a first back-splicing site located at the 5' end of the IRES; and
[0013] (iii) a second back-splicing site located at the 3' end of the IRES;
[0014] and, optionally
[0015] B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site; and / or
[0016] a promoter operably linked to the expression cassette to direct expression of the expression cassette;
[0017] Wherein the IRES is located within the circRNA expression cassette, wherein:
[0018] (I) the 5' end of the IRES is separated from the 3' end of the first back-splicing site by at least 50 nucleotides, preferably at least 200 nucleotides, more preferably at least 300 nucleotides, and
[0019] (II) The 3' end of the IRES is separated from the 5' end of the second back-splicing site by at least 300 nucleotides, preferably at least 350 nucleotides. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The following describes the contents of the drawings included in this specification. In this case, please also refer to the detailed description of the present invention above and / or below.
[0021] Figure 1 : Schematic diagrams involving (A) continuous ORF (top) and separate ORF (bottom) designs. The relative positions of the promoter, inverted repeat (IR), splice acceptor (SA), IRES, open reading frame (ORF), and splice donor (SD) are shown. (B) Western blot of ICOS-L protein expression in A375 cells transfected with an empty vector control (EV) or circICOSL expression plasmids containing CVB3 or EMCV IRES in continuous and separate ORF designs. β-Actin was used as a loading control. (C) RT-qPCR quantification of circICOSL expression relative to GAPDH mRNA in A375 cells transfected with circICOSL expression plasmids containing CVB3 or EMCV IRES in continuous and separate ORF designs. Data from three biological replicates are shown. (D) Agarose gel image showing RT-PCR amplification of circICOSL in A375 cells transfected with an empty vector control (EV) or circICOSL expression plasmids containing CVB3 or EMCV IRES in continuous and separate ORF designs.
[0022] Figure 2 : Schematic diagram of the design of the (A) eGFP split ORF. CircRNA schematic diagram showing the relative position of the IRES within the split eGFP ORF. (B) Agarose gel image showing RT-PCR amplification of circEGFP in A375 cells transfected with the eGFP_FLAG_Split_ORFcircRNA plasmid. (C) RT-qPCR quantification of circEGFP expression relative to GAPDH mRNA in A375 cells transfected with the eGFP_FLAG_Split_ORFcircRNA plasmid. Data from two biological replicates are shown. (D) Western blot analysis of A375 cells transfected with the eGFP_FLAG_Split_ORF circRNA plasmid, detected with labeled antibodies to FLAG, GFP, and β-actin (loading control).
[0023] Figure 3Figure 1: Schematic diagram of (A) eGFP_FLAG_Split_IRES circRNA design. Schematic diagram of the circRNA showing the relative positions of the split_IRES and eGFP ORF. (B) Agarose gel showing RT-PCR amplification of circEGFP in A375 cells transfected with the eGFP_FLAG_Split_IRES circRNA plasmid. (C) RT-qPCR quantification of circEGFP expression relative to GAPDH mRNA in A375 cells transfected with the eGFP_FLAG_Split_IRES circRNA plasmid. Data from two biological replicates are shown.
[0024] Figure 4 Schematic diagram of (A) eGFP_HIPK3_spacer circRNA design. CircRNA schematic diagram showing the relative positions of the HIPK3_spacer (black), IRES (light gray), and eGFP ORF (dark gray). (B) Agarose gel image showing RT-PCR amplification of circEGFP in A375 cells transfected with the eGFP_FLAG_HIPK3_spacer circRNA plasmid. (C) RT-qPCR quantitative analysis of circEGFP expression relative to GAPDH mRNA in A375 cells transfected with the eGFP_FLAG_HIPK3_spacer circRNA plasmid. Data from two biological replicates are shown.
[0025] Figure 5 Schematic diagram of (A) mRNA and circRNA design. Schematic diagram of the circRNA shows the relative positions of the IRES and transgene ORF. (B-D) Western blot analysis of A375 cells transfected with mRNA and circRNA plasmids encoding ADA_FLAG (B), ICOSL (C), and Renilla (D), using antibodies against FLAG, ICOSL, Renilla, and β-actin (loading control). (E-G) Quantitative RT-qPCR analysis of circADA (E), circICOSL (F), and circRenilla (G) expression relative to GAPDH mRNA in A375 cells transfected with mRNA and circRNA plasmids encoding ADA_FLAG, ICOSL, and Renilla. Data from two biological replicates are shown.
[0026] Figure 6(A) Western blot analysis of A375 cells transfected with the ICOSL_Split_CVB3 circRNA plasmid, probed with labeled antibodies against ICOSL and β-actin (loading control). Schematic diagram of the ICOSL_Split_CVB3 circRNA expression cassette design showing the relative positions of the CVB3 IRES (dark gray) and ICOSL ORF (light gray). (B) RT-qPCR quantification of ICOSL expression relative to GAPDH mRNA in A375 cells transfected with the ICOSL_Split_CVB3 circRNA plasmid. Data from two biological replicates are shown.
[0027] Figure 7 : (A) Western blot analysis of A375 cells transfected with the eGFP_Split_CVB3 circRNA plasmid, probed with labeled antibodies against GFP and β-actin (loading control). Schematic diagram of the eGFP_Split_CVB3 circRNA expression cassette design showing the relative positions of the CVB3 IRES (dark gray) and the eGFP ORF (light gray). (B) RT-qPCR quantification of eGFP expression relative to GAPDH mRNA in A375 cells transfected with the eGFP_Split_CVB3 circRNA plasmid. Data from two biological replicates are shown. (C) Western blot analysis of A375 cells transfected with the eGFP_Split_EMCV circRNA plasmid, probed with labeled antibodies against GFP and β-actin (loading control). Schematic diagram of the eGFP_Split_EMCV circRNA expression cassette design. Schematic diagram of the circRNA showing the relative positions of the EMCV IRES (dark gray) and the eGFP ORF (light gray). (D) Quantitative analysis of eGFP expression relative to GAPDH mRNA by RT-qPCR in A375 cells transfected with the eGFP_Split_EMCV circRNA plasmid. Data from two biological replicates are shown. DETAILED DESCRIPTION
[0028] Before describing the present invention in detail below, it is to be understood that the present invention is not limited to the specific methods, protocols and reagents described herein, as these may vary. It is also to be understood that the terminology used herein is intended only to describe specific embodiments and is not intended to limit the scope of the present invention, which is defined solely by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art.
[0029] Preferably, the definitions of the terms used herein are as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, HGW, Nagel, B. and Klb1, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland).
[0030] Throughout this specification and the appended claims, unless the context requires otherwise, the word "comprise" and its variations (e.g., "comprising") are to be understood as including the stated integer, step, or group of integers or groups of steps, but not excluding any other integer, step, or group of integers or groups of steps. In the following paragraphs, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect unless expressly indicated to the contrary. In particular, any feature designated as optional, preferred, or advantageous may be combined with any other feature designated as optional, preferred, or advantageous.
[0031] This specification sheet quotes several documents in full. Each document cited herein (including all patents, patent applications, scientific publications, manufacturer specifications, specifications, etc.), whether above or below, is incorporated herein by reference in its entirety. Nothing herein should be construed as admitting that the present invention is not entitled to disclose earlier than this type of invention due to prior invention. Some documents cited herein are characterized as "incorporated by reference". If the definition or teaching in such incorporated references conflict with the definition or teaching described in this specification sheet, the text of this specification sheet shall prevail.
[0032] The elements of the present invention are described below. These elements list specific embodiments; however, it should be understood that they can be combined in any manner and in any number to produce additional embodiments. The various described embodiments and preferred embodiments should not be interpreted as limiting the present invention to only the embodiments explicitly described. This specification should be understood to support and cover embodiments that combine the explicitly described embodiments with any number of disclosed and / or preferred elements. In addition, any arrangement and combination of all described elements in this application should be deemed to be disclosed by the specification of this application unless the context indicates otherwise.
[0033] definition
[0034] The following provides the definitions of some terms frequently used in this specification. These terms have their respective defined meanings and preferred meanings each time they are used in the rest of the specification.
[0035] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise.
[0036] The term "about" when used in combination with a numerical value is intended to encompass a range of numerical values with a lower limit that is 5% less than the indicated numerical value and an upper limit that is 5% greater than the indicated numerical value.
[0037] As used herein, the term "circular RNA" or "circRNA" refers to a type of RNA in which the ends of an RNA chain are covalently linked to form a closed continuous loop. CircRNAs are typically formed by covalently binding the 5' site of an upstream exon to the 3' site of the same exon or a downstream exon.
[0038] As used herein, the terms "5" and "3" are used to identify one end of a single-stranded nucleic acid molecule. The 5' end is the end of the molecule that terminates at a 5' phosphate group, and the 5' direction is the direction pointing toward the 5' end. Similarly, the 3' end is the end of the molecule that terminates at a 3' phosphate group, and the 3' direction is the direction pointing toward the 3' end. In the context of the present invention, nucleic acid sequences are written with the 5' end on the left and the 3' end on the right (unless otherwise indicated), which refers to the direction of DNA synthesis during replication (from 5' to 3'), the direction of RNA synthesis during transcription (from 5' to 3'), and the direction of reading of mRNA sequences during translation (from 5' to 3').
[0039] As used herein, the term "open reading frame" refers to the portion of the nucleic acid sequence between the start codon and the stop codon, excluding the stop codon (the stop codon acts as a termination signal). A codon is a DNA or RNA sequence of three nucleotides that constitutes a genomic information unit encoding a specific amino acid or indicating the termination of protein synthesis (stop codon). In the standard genetic code, three different stop codons are known: TAG, TAA, or TGA at the DNA level and UAG, UAA, and UGA at the RNA level. In variants of the standard genetic code, other stop codons may be present. The start codon is the first codon of an RNA transcript (such as linear mRNA, circRNA) translated by the ribosome.
[0040] As used herein, the term "backsplicing" refers to splicing that occurs in the reverse order (ie, backsplicing), where an upstream 3' splice site is joined to a downstream 5' splice site.
[0041] As used herein, the term "back splice site" refers to a nucleotide sequence (i.e., splice site) at an intron-exon boundary. The back splice site determines the position of the spliced pre-mRNA and allows the recruitment of the spliceosome required for the back splicing process. The sequence of the back splice site typically comprises two half sites, wherein one half site is located on an intron and the other half site is located on an exon. The nucleotide sequence spliced is separated between the two half sites of the back splice site.
[0042] As used herein, the term "internal ribosome entry site (IRES)" refers to a region in RNA that allows internal initiation of translation in a cap-independent manner. An IRES element typically comprises a highly structured RNA containing multiple stem-loop structures. IRES were initially identified in picornaviruses, but they are present in a variety of different viruses. Recently, IRES sequences have also been found in many cellular mRNAs. Both types of IRES sequences (i.e., viral IRES sequences and cellular IRES sequences) are generally useful in practicing the present invention.
[0043] Viral IRESs can be divided into four different categories based on two main criteria: first, the type of secondary and tertiary structure of their RNA elements; and second, their mode of action in translation initiation (see Mailliot and Martin, RNA 2017, e1458).
[0044] Class 1 IRESs typically require all initiation factors except eIF4E and contain a relatively basic secondary structure consisting of a short and long hairpin. Class 1 IRESs typically recruit the ribosome upstream of the coding region and rely on classic 5' to 3' scanning to find the start codon. Class 1 IRESs are found in, for example, enterovirus A71 (EV-A71), coxsackievirus B3 (CVB3), poliovirus (PV), and human rhinovirus 2 (HRV).
[0045] In contrast, class 2 IRESs facilitate the direct anchoring of the translation initiation machinery to the start codon without any scanning step. Class 2 IRESs are found, for example, in Picornaviridae, such as encephalomyocarditis virus (EMCV) and foot-and-mouth disease virus (FMDV).
[0046] Class 3 IRESs contain more complex secondary and tertiary structures, such as pseudoknots. They require only a small subset of translation initiation factors, namely eIF2, eIF3, and eIF5, to recruit and load the ribosome directly to the AUG start codon without scanning. Class 3 IRESs are found, for example, in Flaviviridae, such as hepatitis C virus (HCV) and classical swine fever virus (CSFV); in Picornaviridae, such as porcine Teschovirus and porcine enterovirus 8 (PEV8); or in Simian virus 2 (SV2).
[0047] Class 4 IRESs are more compact and complex in terms of structural complexity; they typically contain multiple pseudoknots and do not require any translation initiation factors. They are the smallest known IRESs (usually less than 200 nucleotides). Class 4 IRESs are found, for example, in Dicistroviridae, such as Cricket Paralysis Virus (CrPV), Israel Acute Paralysis Virus (IAPV), Platia Stali Enterovirus (PSIV), or Taura Syndrome Virus (TSV).
[0048] In the context of the present invention, the most preferred classes of IRES are class 1 and class 2 IRES.
[0049] As used herein, the term "inverted repeat (IR) element" refers to a nucleotide sequence immediately downstream of its reverse complementary sequence. The downstream nucleotide sequence may be completely identical to the reverse complementary sequence, or may differ therefrom.
[0050] As used herein, the term "promoter" refers to a sequence of DNA to which proteins bind to initiate transcription of a single RNA transcript from the DNA downstream of the promoter. Promoters are typically located in the 5' region of a gene and do not encode a gene product, but rather control gene expression. In a promoter, transcription is initiated by the interaction of a transcription factor with an RNA polymerase. Generally, it is advantageous to select a promoter that is active in the desired host cell type.
[0051] Promoters can be divided into different categories based on their activity patterns. Some promoters are constitutive promoters, which are active in essentially all tissues and do not require any specific stimulus for activity. Non-limiting examples of constitutive promoters are: CMV, EF1a, EFS, CAG, CBh, CBA, MSCV, PGK, SFFV, SV40, and UBC. In a preferred embodiment, the promoter is CMV. In another preferred embodiment, the CMV promoter has a sequence according to SEQ ID NO: 005.
[0052] Other promoters are active only in certain tissues and thus can be used to restrict expression to specific tissues. Non-limiting examples of tissue-specific promoters that can be used in the present invention and the tissues in which they are preferably active are provided below: embryo fetal tissue :Nanog, Brain / nervous system / spinal cord : Nes, Tuba1a, Camk2a, SYN1, Hb9, Th, Thy1, NSE, GFAP (long), GFAP (short), Iba1, Prnp, Cnp, retina : ProA1, hRHO, hBEST1, Grm6, epidermis:K14、BK5、mTyr, Heart / embryonic heart :cTnT, αMHC (long), αMHC (short), Hcn4, muscle :Myog, ACTA1, MHCK7, SM22a, EnSM22a, Desmin, Mb, Bone / cartilage : Runx2, OC, Col1a1, Col2a1, Fat :aP2、Adipoq, Blood / bone marrow :Liver: Afp, Alb, TBG, breast :MMTV、Wap, pancreas : HIP, Pdx1, Ins2, Elastase-1, kidney : NPHS2, lung: SPB, endothelial cells : CD144, Flt-1, ICAM-2, Endoglin, hematopoietic cells : WASP, IFN-B, B19, CD14, CD43, CD45, CD68, Tie1, CD11b, OG-2, tumor : TERT, E2F-1, OC, SLPI, Cox-2, CEA, AFP, LP-P. In a preferred embodiment, the promoter is selected from SYN1, GFAP.
[0053] Other promoters require some stimulus to be activated so that they initiate transcription. Non-limiting examples of inducible promoters are: TRE (tetracycline responsive promoter), CRE (coumarin responsive promoter).
[0054] In a preferred embodiment, the promoter is selected from a constitutive promoter, a tissue-specific promoter, or an inducible promoter. In a preferred embodiment, the promoter is a constitutive promoter. In a preferred embodiment, the promoter is a tissue-specific promoter. In a preferred embodiment, the promoter is an inducible promoter.
[0055] The term "vector" used herein refers to a vehicle that can introduce DNA or RNA sequences (such as foreign genes) into host cells to transform the host and promote the expression (such as transcription and translation) of the introduced sequences.
[0056] As used herein, the term "host cell" refers to a cell comprising a nucleic acid, vector or circRNA as described herein. The host cell can be a eukaryotic cell, such as a plant, animal, fungal or algal cell, or can be a prokaryotic cell, such as a bacterial or protozoan cell. The host cell can be a cultured cell or a primary cell, i.e., a cell isolated directly from an organism (such as a human). The host cell can be an adherent cell or a suspension cell, i.e., a cell grown in suspension. Preferably, the host cell is a mammalian cell.
[0057] As used herein, the term "branch point" refers to a nucleotide typically contained within an intronic heptamer sequence. Beyond the branch point, the heptamer sequence undergoes base pairing with the spliceosome. Unpaired branch points are essential for the formation of the lariat intermediate during the splicing event. Typically, the branch point is an adenine.
[0058] As used herein, the term "polypyrimidine tract" refers to a nucleotide sequence that is typically 15 to 20 base pairs in length and that promotes the assembly of the spliceosome, thereby aiding splicing.
[0059] Implementation Plan
[0060] The following are more detailed definitions of different aspects of the present invention. Each aspect so defined can be combined with any other aspect unless explicitly indicated to the contrary. In particular, any feature indicated as preferred or advantageous can be combined with any other feature indicated as preferred or advantageous.
[0061] In a first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises:
[0062] A) An expression cassette comprising:
[0063] (i) a circRNA expression cassette comprising:
[0064] (a) a nucleic acid comprising or consisting of continuous or separate open reading frames (ORFs) encoding at least one protein, and
[0065] (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct
[0066] Directing the translation of the ORF;
[0067] (ii) a first back-splicing site located at the 5' end of the IRES; and
[0068] (iii) a second back-splicing site located at the 3' end of the IRES;
[0069] and optionally
[0070] B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site, and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette;
[0071] Wherein the IRES is located within the circRNA expression cassette, wherein:
[0072] (I) the 5' end of the IRES is separated from the 3' end of the first back-splicing site by at least 50 nucleotides, preferably at least 200 nucleotides, more preferably at least 300 nucleotides, and
[0073] (II) The 3' end of the IRES is separated from the 5' end of the second back-splicing site by at least 300 nucleotides, preferably at least 350 nucleotides.
[0074] In an alternative first aspect, the present invention provides a nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises:
[0075] A) An expression cassette comprising:
[0076] (i) a circRNA expression cassette comprising:
[0077] (a) a nucleic acid comprising or consisting of continuous or separate open reading frames (ORFs) encoding at least one protein, and
[0078] (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF;
[0079] (ii) a first back-splicing site located at the 5' end of the IRES; and
[0080] (iii) a second back-splicing site located at the 3' end of the IRES;
[0081] and optionally
[0082] B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site pair, and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette.
[0083] The different elements of the nucleic acid of the different aspects of the invention are disclosed in more detail below, and particularly preferred embodiments are disclosed.
[0084] circRNA expression cassettes and expression cassettes
[0085] Without wishing to be bound by theory, the inventors believe that an IRES site located very close to the back-splicing site may interfere with back-splicing efficiency, resulting in a reduced number of circRNAs produced. This interference may be due to any secondary structure formed by the IRES, which prevents the spliceosome from optimally accessing the linear pre-mRNA. Therefore, arranging the IRES within the expression cassette at a distance from the back-splicing site is believed to be beneficial for improving back-splicing efficiency.
[0086] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first back splice site by at least 50 nucleotides, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 350, at least 400, at least 450, at least 475, at least 500 nucleotides. In a preferred embodiment, the 5' end of the IRES is separated from the 3' end of the first back splice site by at least 160 nucleotides. In a preferred embodiment, the 5' end of the IRES is separated from the 3' end of the first back-splice site by at least 475 nucleotides.
[0087] In preferred embodiments, the IRES is located within the circRNA expression cassette, wherein the 3' end of the IRES is separated from the 5' end of the second backsplicing site by at least 300, at least 310, at least 320, at least 330, at least 340, or at least 350 nucleotides.
[0088] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 3' end of the IRES is separated from the 5' end of the second back splice site by at least 50 nucleotides, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290 nucleotides. In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 3' end of the IRES is separated from the 5' end of the second back splice site by at least 50 nucleotides.
[0089] In preferred embodiments, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplicing site by at least 50 nucleotides, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, at least 390, at least 400, at least 410, at least 420, at least 430, at least 440, at least 450, at least 460, at least 470, at least 480, at least 490, at least 500, at least 510, at least 520, at least 530, In some embodiments, the IRES comprises at least 200, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 350 nucleotides (preferably at least 250 nucleotides), and the 3' end of the IRES is separated from the 5' end of the second back splice site by at least 300, at least 310, at least 320, at least 330, at least 340, at least 350 nucleotides (preferably at least 300 nucleotides).
[0090] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 50 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 350 nucleotides.
[0091] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 50 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 680 nucleotides.
[0092] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 250 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 480 nucleotides.
[0093] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 360 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 380 nucleotides.
[0094] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 650 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 90 nucleotides.
[0095] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the 5' end of the IRES is separated from the 3' end of the first backsplice site by at least 300 nucleotides, and the 3' end of the IRES is separated from the 5' end of the second backsplice site by at least 300 nucleotides.
[0096] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 5' end of the IRES and the 3' end of the first back-splicing site and the distance between the 3' end of the IRES and the 5' end of the second back-splicing site are both at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, or at least 400 nucleotides.
[0097] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 5' end of the IRES and the 3' end of the first backsplicing site is 50 to 650 nucleotides (preferably 200 to 400 nt, more preferably 250 to 365 nt) from each other.
[0098] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 5' end of the IRES and the 3' end of the first backsplicing site is 59 to 650 nucleotides (preferably 255 to 363 nt) from each other.
[0099] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 3' end of the IRES and the 5' end of the second back-splicing site is 300 to 700 nucleotides (preferably 350 to 600 nt, more preferably 350 to 500 nt) from each other.
[0100] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 3' end of the IRES and the 5' end of the second back-splicing site is 381 to 685 nucleotides (preferably 381 to 489 nt) from each other.
[0101] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the distance between the 5' end of the IRES and the 3' end of the first back-splicing site is 50 to 650 nucleotides (preferably 200 to 400 nt, more preferably 250 to 365 nt), and the distance between the 5' end / 3' end of the IRES and the 3' end / 5' end of the first / second back-splicing site is 300 to 700 nucleotides (preferably 350 to 600 nt, more preferably 350 to 500 nt).
[0102] In a preferred embodiment, the IRES is located within the circRNA expression cassette such that the 5' end of the IRES and the 3' end of the first back-splicing site are at a distance of 59 to 650 nucleotides (preferably 255 to 363 nt) from each other, and the 3' end of the IRES and the 5' end of the second back-splicing site are at a distance of 381 to 685 nucleotides (preferably 381 to 489 nt) from each other.
[0103] In a preferred embodiment, the IRES is located within the circRNA expression cassette, wherein the IRES is located in the middle of the circRNA expression cassette. In one embodiment, the nucleotides equidistant from the first and second backsplicing sites are part of the IRES sequence. In one embodiment, the sequence differences on both sides of the IRES sequence in the circRNA expression cassette do not exceed 50 nucleotides.
[0104] Another possible way to make the IRES sequence and the back splice site a certain distance apart is to use a nucleotide sequence that does not belong to the ORF, that is, to use a spacer sequence. In a preferred embodiment, the circRNA expression cassette includes a spacer sequence, thereby separating the IRES from the first back splice site. In a preferred embodiment, the spacer sequence is located at the 5' end of the IRES sequence and the 3' end of the first back splice site. In another preferred embodiment, the spacer sequence is located at the 3' end of the IRES sequence and the 5' end of the second back splice site. In another preferred embodiment, the circRNA expression cassette comprises two spacer sequences. Preferably, the first spacer sequence is located at the 5' end of the IRES sequence and the 3' end of the first back splice site, and the second spacer sequence is located at the 3' end of the IRES sequence and the 5' end of the second back splice site. In a preferred embodiment, the length of the spacer sequence is 50 to 500 nucleotides, preferably 100 to 400 nucleotides, more preferably 150 to 300 nucleotides, and most preferably 200 to 250 nucleotides. In a preferred embodiment, the spacer sequence is 50, 100, 200 or 400 nucleotides in length.
[0105] In yet another embodiment, the spacer sequence may encode a protein.
[0106] In a preferred embodiment, the IRES is CVB3, and the IRES is located within the circRNA expression cassette, wherein:
[0107] (I) the 5' end of the IRES is separated from the 3' end of the first backsplicing site by at least 50 nucleotides, preferably at least 200 nucleotides, more preferably at least 300 nucleotides, and
[0108] (II) The 3' end of the IRES is separated from the 5' end of the second back-splicing site by at least 50 nucleotides, preferably at least 100, at least 150, at least 200, or at least 250 nucleotides.
[0109] Open reading frame (ORF) structure
[0110] The ORF contained in the circRNA expression cassette can be arranged as a continuous ORF located at the 3' end or 5' end of the IRES, or arranged as a split ORF, wherein the ORF is split (preferably at a sequence that forms a half-site of a back splice site at the split position) so that the two parts of the split ORF are located on both sides of the IRES in the circRNA expression cassette.
[0111] In the context of the present invention, separate ORFs have multiple advantageous properties. First, the separate ORF design places the IRES at a distance from the backsplicing site, which is thought to improve backsplicing efficiency, thereby increasing circRNA production, without the need for additional sequence elements (i.e., spacer sequences). Another advantageous property is that the protein encoded in the ORF can only be translated from the resulting circRNA, and not from the linear RNA that might be generated from the nucleic acid molecule if cyclization fails or is incomplete. Thus, it is ensured that only the protein encoded in the circRNA is translated.
[0112] In a preferred embodiment of the first aspect of the invention, the separated ORF comprises two parts, the first part of the separated ORF comprises a stop codon and is located at the 5' end of the IRES, and the second part of the separated ORF comprises a start codon and is located at the 3' end of the IRES. In the circRNA generated from the nucleic acid of the first aspect of the invention, the separated ORFs will be reassembled by circularization of the circRNA to form a continuous ORF in the closed loop of the circRNA.
[0113] In a preferred embodiment where the circRNA expression cassette comprises a separate ORF, the second half site of the first backsplicing site and the first half site of the second backsplicing site are located within the two parts of the separate ORF. In other words, they are part of the protein coding sequence in the ORF. This is advantageous because otherwise any residual sequence left by the splicing event will be located within the translated sequence of the ORF and may interfere with protein translation and / or folding due to the extra nucleotides in the sequence.
[0114] In another preferred embodiment, the ORF is a continuous ORF. Preferably, the circRNA expression cassette comprising the continuous ORF further comprises a non-coding spacer sequence that separates the IRES from the back splice site. Exemplary usable spacer sequences have been described above.
[0115] In a preferred embodiment, the circRNA expression cassette comprises a continuous ORF and further comprises one or more non-coding nucleotide sequences located at the 5' end and / or 3' end of the continuous ORF.
[0116] In a preferred embodiment, the circRNA expression cassette comprises a separate ORF and further comprises one or more non-coding nucleotide sequences located at the 5' end and / or 3' end of the IRES.
[0117] In a preferred embodiment, the circRNA expression cassette further comprises one or more (eg, 1, 2, 3, 4) additional ORFs.
[0118] In a preferred embodiment, the circRNA expression cassette further comprises one or more (eg, 1, 2, 3, 4) additional IRES.
[0119] In a preferred embodiment, the circRNA expression cassette further comprises an additional coding nucleotide sequence attached to the one or more ORFs, preferably encoding a sequence that can be used to detect or capture the protein encoded by the one or more ORFs.
[0120] In a preferred embodiment, the ORF encodes at least one therapeutic protein, therapeutic peptide, antigenic protein or antigenic peptide.In general, the ORF may encode any type of peptide or protein.
[0121] Back splice site
[0122] The nucleic acid of the present invention comprises at least two back splice sites, which are located upstream (i.e., 5' end) and downstream (i.e., 3' end) of the IRES. Back splice sites are crucial for the formation (i.e., cyclization) of circRNA encoded by the nucleic acid of the first aspect of the present invention. Back splice sites determine the splicing position of the pre-mRNA and allow the recruitment of the spliceosome required for the back splicing process. Typically, the back splice site comprises two half sites, wherein the nucleic acid is separated between these two half sites during splicing. In this context, the term "splice acceptor (SA)" refers to the 3' end of an intron, and the term "splice donor (SD)" refers to the 5' end of an intron. In a preferred embodiment, the first back splice site comprises a splice acceptor, and the second back splice site comprises a splice donor.
[0123] In a preferred embodiment, the back-splicing site comprises a nucleotide sequence that allows the spliceosome to assist the encoded circRNA in back-splicing. In a preferred embodiment, the back-splicing site requires the spliceosome to produce the circRNA encoded by the circRNA expression cassette.
[0124] In a preferred embodiment, the first back splice site comprises two half sites. In one embodiment, the first half site and the second half site of the first back splice site are located at the 5' end of the circRNA expression cassette. In another embodiment, the first half site of the first back splice site is located at the 5' end of the circRNA expression cassette, and the second half site of the first back splice site is a part of the circRNA expression cassette. In the case where the open reading frame (ORF) is a separate ORF, this embodiment is preferred. Preferably, the second half site of the first back splice site is a part of the ORF sequence.
[0125] In a preferred embodiment, the circRNA produced by back-splicing the nucleic acid molecule of the first aspect of the invention does not comprise any sequence derived from a back-splicing site that is not part of an ORF or IRES sequence.
[0126] In a preferred embodiment, the second back splice site comprises two half sites. In one embodiment, the first half site and the second half site of the second back splice site are located at the 3' end of the circRNA expression cassette. In another embodiment, the first half site of the second back splice site is a part of the circRNA expression cassette. In the case where the ORF is a separate ORF, this embodiment is preferred. Preferably, in the case where the ORF is a separate ORF, the first half site of the second back splice site is a part of the ORF, and the second half site of the second back splice site is located at the 3' end of the circRNA expression cassette.
[0127] In a preferred embodiment, the first back-splicing site and the second back-splicing site each comprise two half-sites. In one embodiment, the first half-site and the second half-site of the first back-splicing site are located at the 5' end of the circRNA expression cassette, and the first half-site and the second half-site of the second back-splicing site are located at the 3' end of the circRNA expression cassette. In another embodiment, the first half-site of the first back-splicing site is located at the 5' end of the circRNA expression cassette, and the second half-site of the first back-splicing site is part of the circRNA expression cassette, and the first half-site of the second back-splicing site is part of the circRNA expression cassette, and the second half-site of the second back-splicing site is located at the 3' end of the circRNA expression cassette.
[0128] In a preferred embodiment, the first half site of the first back-splice site comprises or consists of the nucleotides AG. In another preferred embodiment, the first half site of the first back-splice site comprises or consists of the nucleotide sequence CAG. In another preferred embodiment, the first half site of the first back-splice site is a splice acceptor. In another preferred embodiment, the second half site of the first back-splice site begins with the nucleotide G.
[0129] In a preferred embodiment, the second half site of the second back-splice site comprises or consists of the nucleotides GT. In another preferred embodiment, the second half site of the second back-splice site comprises or consists of the nucleotide sequence GTAAGT. In another preferred embodiment, the second half site of the second back-splice site is a splice donor. In another preferred embodiment, the first half site of the second back-splice site comprises or consists of the nucleotide sequence CAG or AAG.
[0130] In a preferred embodiment, the first half site of the first back-splice site comprises or consists of the nucleotides AG; and the second half site of the second back-splice site comprises or consists of the nucleotides GT. In a preferred embodiment, the first half site of the first back-splice site comprises or consists of the nucleotide sequence CAG; and the second half site of the second back-splice site comprises or consists of the nucleotide sequence GTAAGT.
[0131] In a preferred embodiment, the back splice site is the consensus sequence of the U2 (major class) intron in the pre-mRNA, generally conforming to the following consensus sequence: 3' splice site: CAG|G and 5' splice site: MAG|GTRAGT, where M is A or C, R is A or G, and | represents an intron-exon boundary. In another preferred embodiment, the sequence of the splice site is as shown in the similarity matrix of the human U2 intron, as disclosed in, for example, Zhang Hum Mol Genet. 1998 May; 7(5): 919-32, which is incorporated herein by reference.
[0132] In a preferred embodiment, the sequence of the ORF is modified to generate a back-splicing site. Preferably, the modification introduces only silent mutations (ie, does not change the amino acid sequence of the encoded protein).
[0133] In a preferred embodiment, the first back-splicing site has the nucleotide sequence CAGGT. In another preferred embodiment, the first back-splicing site has a sequence that is at least 85%, at least 90%, at least 95% or at least 99% identical to the nucleotide sequence CAGGT.
[0134] In a preferred embodiment, the second back-splicing site has the nucleotide sequence AGGTA. In another preferred embodiment, the second back-splicing site has a sequence that is at least 85%, at least 90%, at least 95% or at least 99% identical to the nucleotide sequence AGGTA.
[0135] Internal ribosome entry site (IRES)
[0136] CircRNAs are covalently linked to form closed loops, lacking a 5' end and requiring cap-independent mechanisms to initiate translation of their encoded proteins. Internal ribosome entry sites (IRESs) are RNA sequences that recruit the 40S ribosomal subunit via a cap-independent mechanism, allowing translation to begin. These elements often employ complex secondary RNA structures that serve as anchoring sites for the ribosome.
[0137] Therefore, nucleic acid molecules of the present invention include IRES in circRNA expression cassettes. IRES is operably linked to ORF so that at least one protein encoded in the ORF can be translated from circRNA. Without wishing to be bound by theory, the inventors believe that if IRES is close to the reverse splicing site, the complex secondary structure of IRES or IRES-related proteins may interfere with the spliceosome. Although IRES has diversity in terms of nucleic acid sequence, their commonality is that they usually have complex secondary structures. Therefore, the interference of IRES with the spliceosome does not depend on specific IRES, but depends on the generally shared secondary structure of IRES.
[0138] In a preferred embodiment, the IRES is a viral IRES or a cellular IRES, preferably a viral IRES. In a preferred embodiment, the IRES is a class 1 or class 2 IRES as defined above. Non-limiting examples of class 1 IRES are found in enterovirus A71 (EV-A71), coxsackievirus B3 (CVB3), poliovirus (PV), and human rhinovirus 2 (HRV), preferably CVB3. Non-limiting examples of class 2 IRES are found in the Picornaviridae family, such as encephalomyocarditis virus (EMCV) and foot-and-mouth disease virus (FMDV), preferably EMCV.
[0139] In a preferred embodiment, the IRES is a class 1 IRES. In another preferred embodiment, the IRES is a class 2 IRES. In another preferred embodiment, the IRES is a class 3 IRES. In another preferred embodiment, the IRES is a class 4 IRES.
[0140] In a preferred embodiment, the IRES is an IRES sequence from a coxsackievirus (CVB), more preferably CVB3. In a more preferred embodiment, the IRES has a sequence according to SEQ ID NO: 001. In another preferred embodiment, the sequence of the IRES has a sequence that is at least 50%, at least 75%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 001. In a preferred embodiment, the sequence of the IRES has a sequence that is at least 75% identical to the sequence of SEQ ID NO: 001. In a preferred embodiment, the IRES has a sequence according to the sequence of SEQ ID NO: 001.
[0141] In a preferred embodiment, the IRES is an IRES sequence from encephalomyocarditis virus (EMCV). In a more preferred embodiment, the IRES has a sequence according to SEQ ID NO: 002. In another preferred embodiment, the IRES has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 002.
[0142] In one embodiment, the IRES is a viral IRES, preferably an IRES from a virus of the following viral families: Adenoviridae, Arenaviridae, Birnaviridae, Chrysoviridae, Coronaviridae, Dicistroviridae, Filoviridae, Flaviviridae, Hepadnaviridae, Herpesviridae, Hypoviridae, Iflaviridae, Luteoviridae, Orthomyxoviridae (Orthomyxoviridae), Papillomaviridae, Paramyxoviridae, Parvoviridae, Picornaviridae, Pneumoviridae, Polyomaviridae, Potyviridae, Reoviridae, Retroviridae, Rhabdoviridae, Secoviridae, Tombusviridae, Totiviridae, Virgaviridae.
[0143] In another preferred embodiment, the IRES is a variant of a viral IRES (preferably CVB3 or EMCV, more preferably CVB3) having a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the nucleotide sequence of the viral IRES.
[0144] Inverted repeat (IR) elements
[0145] Although inverted repeat elements are not essential for back-splicing of circRNAs, the presence of inverted repeat elements on either side of the splice site is known to increase circRNA production. It is believed that inverted repeat elements bring the splice sites into close proximity to promote back-splicing of circRNAs.
[0146] In the context of the present invention, an IR element located 5' to the first back-splicing site (ie, first IR element) and an IR element located 3' to the second back-splicing site (ie, second IR element) are used to further improve back-splicing efficiency.
[0147] Generally, any nucleotide sequence immediately downstream of its reverse complement can be used as an IR element in the present invention. In some embodiments, the downstream sequence (i.e., the second IR element) is different from the reverse complement of the upstream sequence (i.e., the first IR element).
[0148] In some embodiments, the downstream nucleotide sequence (i.e., the second IR element) is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identical to the reverse complement of the upstream sequence (i.e., the first IR element). In a preferred embodiment, the first IR element and the second IR element comprise sequences that are at least 80%, preferably at least 90%, and more preferably at least 95% identical to the reverse complement of each other.
[0149] In a preferred embodiment, the downstream nucleotide sequence differs from the reverse complement of the upstream sequence by 5, 4, 3, 2 or 1 nucleotides.
[0150] In a preferred embodiment, the IR element is 50 to 1000 nucleotides in length, preferably 200 to 600 nucleotides, more preferably 300 to 500 nucleotides in length. In another preferred embodiment, the IR element is 200 to 400 nucleotides in length.
[0151] The IR elements used in the present invention may be IR elements known in the art, or IR elements derived from genes such as AC004076.9, ACYP2, AMD1, ARHGAP10, ARHGAP12, ARHGEF12, ARHGEF28, ASAP1, ASH1L, ASXL1, ATXN2, BMPR2, BRWD1, BTBD10, CBFA2T2, CCDC126, CCDC134, CCDC66, CCDC7, CCDC9, CCNB1, CDK13, CFLAR, CHD9, CLIP2, CLNS1A, CNN2, COA1, CORO1C, CREBBP, CRKL, CTB-43P18.3,SNHG4,DCUN1D4,DEK,DHRS3,DLG1,DOPEY2,DYNC1H1,ELF2,EMC2,EPHB4,EPS15,ERC1,ETFA,EXOSC1,FAM13B,FARSA,FBXO7,FGD4,FGD6, FKBP3, FKBP8, FNTA, FOXK2, GAPVD1, GBAS, GDI2, GLIS2, GLS, GON4L, GRHPR, HERC1, HIPK3, HNRNPM, HOOK3, HP1BP3, HPS5, HTT, HUWE1, IARS, ILKAP, IQGAP 1. KDM1A, KIAA0368, KIAA1429, KIAA1841, KLHL8, KMT2C, Laccase, LMBR1, LRCH3, LZIC, MAP3K1, MARK4, MBOAT2, MCU, MED13L, METTL3, MGA, MGEA5, MITD 1. MORC3, MRPS35, MYO9B, NCOA2, NFAT5, NFATC3, NFX1, NUDC, NUP54, PAFAH1B2, PAIP2, PCMT1, PDCD11, PDE8A, PDS5A, PHC3, PHLDB2, PLEKHM1, PLEKHM3, P LOD2, PMS1, PNN, POLR2A, POMT1, PPP6R2, PROSC, PRRC2B, PSEN1, PSMA7, PTP4A2, PTPN12, QKI, R3HDM1, RAB6A, RALBP1, RARS, RBM23, RBM33, RBM39, RELL1 , REPS1, RERE, RHOBTB3, RLF, RNF19B, SDHAF2, ZRANB1, FAM228B, RPRD1B, RPS6KC1, RSF1, RSRC1, RTN4, RYK, SAFB2, SAMD4A, SCAF8, SCARF1, SCYL2, SDF4,SEC31A, SENP6, SIPA1L1, SKA3, SLC38A1, SLTM, SMARCA5, SMC3, SMO, SNX25, SOBP, SOS2, SPIDR, SPPL3, SRSF4, STAM, ST K3, STX6, TERF2, TIMMDC1, TMED2, TMEM138, TMEM165, TMEM181, TNPO1, TNPO3, TOP1, TTC39C, TTLL1, UBA2, UBAP2, UBE2K 、UBQLN1、UBR5、UBXN2A、UBXN7、UHRF2、UIMC1、URI1、UTP18、UTRN、VAMP3、VAPB、VMP1、WDR78、XPO1、YTHDF2、YWHAE、YY1AP1、ZBTB46、ZCCHC11、ZCCHC6、ZFAND6、ZFX、ZKSCAN1、ZMYM4、ZNF124、ZNF236、ZNF394、ZNF430、ZNF652、ZNF720、ZNF91. Preferably, the IR element is selected from HIPK3 and ZKSCAN1.
[0152] In a preferred embodiment, the IR element is derived from a gene having an IR element-derived circRNA.
[0153] In a preferred embodiment, the IR elements are short interspersed nuclear elements (SINEs).
[0154] In a preferred embodiment, the IR element is an Alu repeat element.
[0155] In a preferred embodiment, the first IR element has a sequence according to SEQ ID NO: 003. In another preferred embodiment, the first IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 003. In a preferred embodiment, the second IR element has a sequence according to SEQ ID NO: 004. In another preferred embodiment, the second IR element has a sequence that is at least 85%, at least 90%, at least 95%, or at least 99% identical to the sequence of SEQ ID NO: 004.
[0156] In a preferred embodiment, the first IR element has a sequence as contained in the disclosed constructs of the present invention. In another preferred embodiment, the first IR element has a sequence that is at least 85%, at least 90%, at least 95% or at least 99% identical to a sequence as contained in the disclosed constructs of the present invention.
[0157] In a preferred embodiment, the second IR element has a sequence as contained in the disclosed constructs of the present invention. In another preferred embodiment, the second IR element has a sequence that is at least 85%, at least 90%, at least 95% or at least 99% identical to a sequence as contained in the disclosed constructs of the present invention.
[0158] circRNA
[0159] The nucleic acid of the present invention can generally encode any type of circRNA. In a preferred embodiment, the circRNA encoded by the nucleic acid of the present invention or the circRNA of the second aspect of the present invention has a length of 50 nt to 5000 nt, preferably a length of 200 nt to 2000 nt.
[0160] In a preferred embodiment, the circRNA encoded by the nucleic acid of the invention or the circRNA of the second aspect of the invention does not contain any sequence of a backsplicing site that does not form part of an ORF or IRES sequence.
[0161] In a preferred embodiment, the circRNA encoded by the nucleic acid of the present invention or the circRNA of the second aspect of the present invention does not encode a fluorescent protein, preferably does not encode a green or red fluorescent protein.
[0162] In a preferred embodiment, the circRNA encoded by the nucleic acid of the invention or the circRNA of the second aspect of the invention encodes a therapeutic protein, a therapeutic peptide, an antigenic protein or an antigenic peptide.
[0163] In a preferred embodiment, the GC content of the circRNA is between 30% and 70%, preferably 40% to 60%, more preferably 50% to 60%.
[0164] In a preferred embodiment, the circRNA does not contain an IR element.
[0165] In a preferred embodiment, the circRNA lacks a complete miRNA target site. This embodiment is particularly useful for preventing miRNA-mediated endonuclease cleavage.
[0166] In a preferred embodiment, the circRNA lacks cryptic splice sites.
[0167] Other elements of the nucleic acids of the present invention
[0168] The nucleic acid of the present invention may further comprise elements having functions in expression regulation, stability, etc.
[0169] In a preferred embodiment, the nucleic acid comprises a promoter operably linked to an expression cassette to direct expression of the expression cassette. In another embodiment, the nucleic acid further comprises an enhancer. Examples of promoters and enhancers used in the nucleic acids of the present invention include the early promoter and enhancer of SV40 (Mizukami T. et al., 1987), the LTR promoter and enhancer of Moloney murine leukemia virus (Kuwana Y et al., 1987), the immediate early promoter and enhancer of cytomegalovirus (CMV), and the like.
[0170] In another preferred embodiment, the nucleic acid molecule further comprises a branch point.
[0171] In another preferred embodiment, the nucleic acid molecule further comprises a polypyrimidine tract.
[0172] The vector of the present invention
[0173] The third aspect of the present invention relates to a vector comprising the nucleic acid molecule according to the first aspect of the present invention or the circRNA according to the second aspect of the present invention.
[0174] In the present invention, vector refers to any type of element that allows the inclusion of the nucleic acid or circRNA of the present invention. Any suitable vector (such as a plasmid, cosmid, episome, liposome, exosome, artificial chromosome, phage or viral vector) can be used in the present invention.
[0175] The terms "vector," "cloning vector," and "expression vector" refer to vehicles that can introduce DNA or RNA sequences (e.g., nucleic acids or circRNAs of the present invention) into host cells so as to transform the host and promote expression (e.g., transcription, cyclization, and translation) of the introduced nucleic acids or circRNAs.
[0176] Such vectors may contain regulatory elements (such as promoters, enhancers, terminators, etc.) to promote or direct the expression of the nucleic acid after administration to a subject. Examples of promoters and enhancers used in expression vectors for animal cells include the early promoter and enhancer of SV40 (Mizukami T. et al., 1987), the LTR promoter and enhancer of Moloney murine leukemia virus (Kuwana Y et al., 1987), the promoter of cytomegalovirus (CMV) (Mason JO et al., 1985) and enhancer (Gillies SD et al., 1983), etc.
[0177] Any expression vector for animal cells can be used as long as the nucleic acid of the present invention can be inserted and expressed. Examples of suitable vectors include pAGE107 (Miyaji H et al., 1990), pAGE103 (Mizukami T et al., 1987), pHSG274 (Brady G et al., 1984), pKCR (O'Hare K et al., 1981), pSG1βd2-4- (Miyaji H et al., 1990), and the like. Examples of other plasmids include replicative plasmids containing an origin of replication.
[0178] Host cells of the present invention
[0179] A fourth aspect of the present invention relates to a host cell comprising the nucleic acid, circRNA and / or vector according to the present invention (preferably obtained by transformation, transduction or transfection of the host cell).
[0180] The term "transformation" originally refers to the naturally occurring process of gene transfer into a host cell, which involves the cell taking up genetic material (such as nucleic acids, e.g., DNA or RNA) through the cell membrane so that the host cell will express the introduced gene or sequence to produce the desired substance. Transformation is divided into two types: natural transformation and artificial or induced transformation. Artificial or induced transformation methods are performed under laboratory conditions.
[0181] A host cell that receives and expresses exogenous nucleic acid (eg, DNA or RNA) through a process of transformation has been "transformed."
[0182] As used herein, the term "transfection" refers to a gene transfer method that involves forming pores in the cell membrane of a host cell so that the host cell can receive exogenous genetic material. Typically, transfection refers to the transformation of eukaryotic cells (such as insect cells or mammalian cells). Chemically mediated transfection includes the use of, for example, calcium phosphate or cationic polymers or liposomes. Non-chemically mediated transfection methods are typically electroporation, sonoporation, puncture transfection, optical transfection, or hydrodynamic delivery. Particle-based transfection uses gene gun technology (wherein nanoparticles are used to transfer nucleic acid to a host cell) or another method called magnetic transfection. The use of nuclear transfection and heat shock is other successful transfection methods that have been developed. Host cells that receive exogenous nucleic acids via a transfection method have been "transfected."
[0183] As used herein, the term "transduction" is generally understood to mean the transfer of exogenous nucleic acid (such as DNA or RNA) into a cell by a virus or viral vector. A host cell that receives and expresses exogenous nucleic acid (such as DNA or RNA) by a virus or viral vector has been "transduced."
[0184] Pharmaceutical composition
[0185] The fifth aspect of the present invention relates to a pharmaceutical composition comprising the nucleic acid, circRNA, vector and / or host cell of the present invention and a pharmaceutically acceptable carrier.
[0186] As used herein, the term "pharmaceutical composition" or "therapeutic composition" refers to a compound or composition capable of inducing a desired therapeutic effect when properly administered to a subject.
[0187] In some embodiments, a subject may also be referred to as a patient.
[0188] Such therapeutic or pharmaceutical compositions may comprise a therapeutically effective amount of the nucleic acid, circRNA, vector and / or host cell of the present invention, or further comprise a therapeutic agent, mixed with a pharmaceutically or physiologically acceptable formulation appropriately selected for the mode of administration.
[0189] The nucleic acids, circRNAs, vectors and / or host cells of the present invention are typically provided as part of a sterile pharmaceutical composition, which typically includes a pharmaceutically acceptable carrier.
[0190] "Pharmaceutically" or "pharmaceutically acceptable" refers to molecular entities and compositions that do not produce adverse, allergic or other untoward reactions when administered to mammals, particularly humans, as appropriate. Pharmaceutically acceptable carriers or excipients refer to any type of non-toxic solid, semisolid or liquid filler, diluent, encapsulating material or formulation excipient.
[0191] A "pharmaceutically acceptable carrier" may also be referred to as a "pharmaceutically acceptable diluent" or a "pharmaceutically acceptable vehicle" and may include solvents, bulking agents, stabilizers, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, etc., which are physiologically compatible. Thus, in one embodiment, the carrier is an aqueous carrier.
[0192] The form of the pharmaceutical composition, the route of administration, the dosage and regimen will naturally depend on the condition to be treated, the severity of the disease, the age, weight and sex of the patient, the desired duration of treatment, etc. The pharmaceutical composition may be in any suitable form (depending on the desired method of administration to the patient). It may be provided in unit dosage form, typically in a sealed container, and may be provided as part of a kit. Such kits typically (although not necessarily) include instructions for use. It may include a plurality of such unit dosage forms.
[0193] Empirical factors such as biological half-life or bioavailability will generally aid in determining dosage. The frequency of administration can be determined and adjusted during the course of treatment. Alternatively, a sustained continuous release formulation may be appropriate.
[0194] Other aspects of the invention
[0195] The present invention further includes a method for producing circular mRNA (circRNA), comprising the steps of introducing the nucleic acid molecule of the first aspect of the present invention or the vector of the third aspect of the present invention into a eukaryotic cell, and optionally purifying the circRNA from the cell.
[0196] In another aspect of the present invention, a method for producing circular mRNA (circRNA) comprises the following steps: contacting the nucleic acid molecule of the first aspect of the present invention or the vector according to the third aspect of the present invention with isolated RNA polymerase II.
[0197] In another aspect, the present invention relates to a method for producing a recombinant protein. The method comprises introducing the nucleic acid molecule of the first aspect of the present invention, the circRNA according to the second aspect, or the vector according to the third aspect into a eukaryotic cell, and optionally purifying the recombinant protein encoded by the ORF.
[0198] In yet another aspect of the present invention, the present invention relates to a circRNA obtained by the method of the present invention.
[0199] Example
[0200] The present invention provides a nucleic acid encoding a circRNA that utilizes spliceosome-based backsplicing for intracellular circRNA expression. Flanking regions of highly expressed circRNAs derived from the HIPK3 locus are used to promote backsplicing. Furthermore, an IRES element is introduced to promote cap-independent protein translation of the circRNA. Studies have shown that the CVB3 IRES outperforms the EMCV IRES in protein expression, although both IRES elements result in increased circRNA expression when using the nucleic acid design of the present invention. Surprisingly, the inventors observed that the position of the IRES within the circRNA expression cassette significantly affects circRNA expression at both the RNA and protein levels, suggesting improved backsplicing. Studies have shown that positioning the IRES element near the flanking regions of the circRNA expression cassette inhibits circRNA expression. Without wishing to be bound by theory, the inventors hypothesize that the highly structured IRES element interferes with splice site recognition and spliceosome assembly, thereby inhibiting backsplicing and circRNA production. This hypothesis is supported by the observation that increasing the distance between the IRES and the splice site by adding a non-coding spacer sequence between the SA and IRES restores circRNA production in a distance-dependent manner. An increase in the distance between these two features was positively correlated with the expression of circRNA.
[0201] For transgenic protein expression, incorporating additional noncoding sequences appears unnecessary, as they have no additional function beyond promoting efficient back-splicing and can reduce the payload / insertion capacity available for expressing the gene of interest. To overcome this issue, split ORF circRNA designs have been investigated. The transgenic payload is separated at CAGG or AAG motifs within the ORF to maximize mimicking of the endogenous splicing reaction. The inventors observed that splitting the ORF significantly increased circRNA production compared to inserting the continuous ORF directly upstream or downstream of the IRES. Furthermore, the inventors found that splitting the ORF at a more central position, thereby increasing the distance between the IRES and the splice site, further increased circRNA production, both when compared to a continuous ORF design and when compared to splitting the ORF at the 5' and 3' ends. This supports the hypothesis that the proximity of the IRES to the splice site negatively affects circRNA back-splicing. This is further supported by the finding that splitting the IRES element itself and inserting the ORF within the IRES inhibited efficient circRNA production.
[0202] Finally, the study found that the separate ORF design was superior in a variety of transgenes when compared to the continuous ORF design, indicating that the circRNA expression design of the present invention is a powerful tool for transgene expression.
[0203] General Methods Applied in the Examples
[0204] Constructs, cell culture and treatment:
[0205] Plasmids were synthesized by Genscript and expressed in mammalian cells using the pcDNA3.1 backbone vector. The A375 human melanoma cell line was obtained from ATCC and maintained in Dulbecco's modified Eagle's medium (Thermo Fisher Scientific) supplemented with 10% fetal bovine serum (Thermo Fisher Scientific) and antibiotics (100 U / ml penicillin and 100 mg / ml streptomycin; Thermo Fisher Scientific). All cell types were incubated at 37°C in a 5% CO2 humidified incubator.
[0206] 24 hours before transfection, cells were seeded at approximately 150,000 cells per well in a 6-well plate. Cell transfection was performed using Lipofectamine 3000 according to the manufacturer's guidelines. Briefly, 2 μg of circRNA or pcDNA3.1 linear expression vector was transfected per well. The transfection mixture was incubated with cells overnight and the culture medium was replaced. Unless otherwise stated, RNA and protein were harvested 48 hours after transfection. For all experiments, cells were harvested by washing with 1× PBS and then centrifuging at 1200 rpm for 4 minutes at 4°C. 66.6% of the harvested cells were used for RNA isolation, which was performed using TRIzol reagent (Thermo Fisher Scientific) according to the manufacturer's protocol.
[0207] RT-PCR and RT-qPCR :
[0208] According to the manufacturer's protocol, 1 μg of DNase-treated total RNA was reverse transcribed using an M-MLV reverse transcriptase kit (Thermo Fisher Scientific) by random hexamer initiation reaction. In the case of RT-PCR, 25 PCR cycles were performed with or without RT enzyme. The products were visualized by 1% agarose gel electrophoresis and verified using Sanger sequencing. For quantitative PCR, cDNA was mixed with a platinum SYBR Green I Master kit (Invitrogen) and run on a 7500Fast real-time PCR system (Applied Biosystem). The reaction was performed in technical replicates three times. The Ct values obtained for each replicate were converted (2-Ct) and averaged (σ). All samples were normalized to GAPDH. The results were displayed in a bar graph using GraphPad (Prism 7), showing individual biological replicates and standard deviations plotted as error bars.
[0209] Western Blot:
[0210] Cells were harvested in 1× PBS and centrifuged at 1200 rpm for 5 minutes at 4°C. For cell lysis, the cell pellet was collected and resuspended in 50 μL of RIPA buffer supplemented with protease inhibitors. The lysate was incubated for 20 minutes, and then cell debris was collected by centrifugation at 12,000 × g for 15 minutes at 4°C. The supernatant was collected, and protein levels were quantified using the BCA assay (Thermo Fisher Scientific). For Western blot analysis, 3 μg of protein was used for Western blot analysis. Prior to Western blotting, an equal volume of 2× SDS loading buffer [125 mM Tris–HCl pH 6.8, 20% glycerol, 5% SDS, and 0.2 M DTT] was added to 3 μg of protein lysate and briefly boiled at 95°C for 5 minutes before loading onto a 12% Tris-glycine SDS-PAGE gel (Thermo Fisher Scientific) and running at 125 V for approximately 1.5 hours. Proteins were transferred to PVDF membranes (BioRad) by wet blotting at 30 V at 4 ° C for 2 hours. Subsequently, the membranes were pre-blocked with 10% skim milk at room temperature (RT) for 1 hour and then incubated with primary antibodies for 1 hour and secondary antibodies for 1 hour. After each antibody incubation, the membranes were rinsed 3 × 5 min in 1 × PBS + 0.05% Tween20 and washed 1 × 5 min in 1 × PBS. Protein bands were developed using the SuperSignal WestFemto Maximum Sensitivity Substrate Kit (Thermo Fisher Scientific) and ChemiDoc XRS + (Biorad).
[0211] Antibodies used:
[0212] ICOS-L (1:500; 233151, Abcam)
[0213] Renilla (1:2,000; ab185925, Abcam)
[0214] eGFP (1:5,000; 290, Abcam)
[0215] FLAG (1:10,000; F3165, Sigma-Aldrich)
[0216] β-actin (1:20,000; A5441, Sigma-Aldrich)
[0217] Rabbit secondary antibody (1:10,000; ab6721, Abcam)
[0218] Mouse secondary antibody (1:10,000; ab90723, Abcam)
[0219] Example 1: Design of circRNA expression vector
[0220] Several vectors containing circRNA expression cassettes have been designed to encode different transgenes from circular RNA transcripts using different internal ribosome entry site (IRES)-open reading frame (ORF) positions (see Figure 1 A). All expression cassettes share a common promoter (CMV) and are flanked by inverted repeat elements (IR-HIPK3) to promote back-splicing and circRNA production. The flanking regions of HIPK3 were used in the design of the expression vectors because the circRNA encoded by the HIPK3 gene is one of the most endogenously expressed circRNAs in a variety of human cell and tissue types. In addition, the circRNA exons are also flanked by splice sites to promote back-splicing. Other features (such as branch point sequences and polypyrimidine tract sequences) are also involved to ensure efficient splicing, and these features are all taken from the adjacent flanking regions of circHIPK3. Open reading frames (ORFs) (coding regions of circRNAs) were designed using two different strategies: one in which the ORF unit is located downstream of the IRES and is continuous, and the other in which the ORF is split and flanked by the IRES element ( Figure 1 A).
[0221] To construct a split ORF design, ideally a CAGG or AAG motif should be found in the central region of the ORF and the ORF should be split between two G nucleotides (in this case one part of the ORF ends with CAG and the other starts with G). The principle of this design is based on the consensus sequence of human splice sites, namely the splice donor (SD) site and the splice acceptor (SA) site that can be recognized by the endogenous splicing machinery. In endogenous splicing, the splice donor is connected to the splice acceptor, and the CAGG motif is most conserved at these sites. Subsequently, an IRES element is inserted between the separated ORFs so that the C-terminal part of the gene is upstream of the IRES and the N-terminal part is downstream of the IRES ( Figure 1 A). In this way, the impact of specific features (particularly the choice and position of the IRES) on circRNA production can be studied.
[0222] Natural circRNAs lack the 5' cap and 3' polyadenylation tail required for cap-dependent translation and are therefore considered non-coding. However, it has been shown that the introduction of an IRES element into a circRNA molecule can induce the translation of the circRNA through a cap-independent mechanism. Therefore, this study examined the protein-coding potential of circRNAs containing IRES elements derived from coxsackievirus B3 (CVB3) or encephalomyocarditis virus (EMCV). As previously described, circRNA expression vectors encoding ICOS-L were tested using both continuous ORF and split ORF designs (see Table 3 below for details of the ICOS-L split ORF), which contained the IRES of CVB3 or EMCV. 48 hours after transfection, the levels of ICOS-L protein were measured in A375 cells transfected with plasmids encoding circICOSL (using a continuous ORF design, a split ORF design containing EMCVIRES, and a split ORF design containing CVB3IRES, respectively). Western blot analysis showed that CVB3IRES stimulated higher levels of protein production compared to EMCVIRES ( Figure 1 B). Interestingly, the split ORF design combined with CVB3IRES significantly increased the expression level of ICOS-L protein compared with the continuous ORF ( Figure 1 B). In addition, no protein output was detected by continuous ORF design combined with EMCVIRES. Figure 1 C) and real-time quantitative reverse transcription polymerase chain reaction (qRT-PCR) ( Figure 1 D) The analysis showed that the split ORF design produced significantly higher levels of circRNA than the continuous ORF design, indicating that the position of the IRES within the circRNA expression cassette has a significant impact on circRNA production.
[0223] Example 2: The distance between IRES and the backsplicing site affects circRNA expression
[0224] To further investigate the effects of split ORF design on circRNA production, a circEGFP expression vector was constructed. This expression vector contained all the elements described in Example 1 above. The eGFP ORF was split at multiple positions to vary the distance between the IRES (here, CVB3) and the flanking splice sites. The circEGFP design was labeled based on the distance between the IRES and the upstream SA site (i.e., the first reverse splice site). For example, for Split_59, the IRES was inserted 59 nt downstream of the SA site ( Figure 2 A; Table 1).
[0225] Plasmids encoding circEGFP with different split ORF designs were transfected into A375 cells, and RNA and protein levels were measured 48 hours after transfection. CircRNA expression of each construct was verified by RT-PCR using divergent primers to amplify only the circular RNA transcripts ( Figure 2 B). qRT-PCR ( Figure 2 C) and Western blot ( Figure 2 Quantitative RNA and protein expression analysis performed in D) showed a negative correlation between the proximity of the IRES to the splice site and circRNA expression. Splits located near the 5' end (split_18 and split_59) or near the 3' end (split_650 and split_732) of the ORF showed reduced circRNA expression at both the RNA and protein levels compared to splits located more centrally (split_255 and split_363). Figure 2 CD). This suggests that a distance of approximately 250 nt between the IRES and the splice site promotes the highest level of circRNA expression. However, when the distance between the IRES and the splice site is approximately 59-94 nt, circRNA production can still be detected, and its level is higher than that when the distance is approximately 13-18 nt.
[0226]
[0227] Table 1: Nucleotide distances between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP split circRNA expression cassette design.
[0228] Example 3: Isolation of IRES
[0229] In addition, the inventors also studied whether the separation of the IRES element itself affects circRNA expression. The expression vector contains all the elements described in Example 1 above. The IRES element is separated at different positions of the CAGG / AAGG motif to allow splicing and cyclization of circRNA without introducing additional nucleotides in the IRES sequence. In this experiment, the continuous eGFP ORF was inserted into the separated IRES element, and the IRES fragment was directly located on both sides of the back splicing site involved in circRNA production ( Figure 3A). Split IRES designs are named according to the distance from the start of the IRES (the first nucleotide encoding the IRES) to the SD. The following expression constructs were tested: split_i33 (SEQ ID NO: 012), split_i219 (SEQ ID NO: 013), and split_i317 (SEQ ID NO: 014). A375 cells were transfected with plasmids encoding circEGFP of different split IRES designs. RT-PCR ( Figure 3 B) and qRT-PCR ( Figure 3 C) Evaluation of circRNA expression. Western blotting was performed to evaluate the expression of eGFP protein ( Figure 3 D) IRES separation showed that compared with ORF separation design, circRNA expression was significantly reduced (see Figure 3 BD). In fact, at both the RNA and protein levels, the IRES split design consistently showed the lowest levels of circRNA expression. Therefore, splitting the IRES element and inserting an internal continuous ORF does not seem to be a suitable system for high-yield circRNA expression.
[0230] Example 4: Proximity of IRES to flanking backsplicing sites affects circRNA expression
[0231] To further investigate the effect of the position of the IRES relative to the backsplicing site on circRNA expression, an eGFP circRNA expression vector was generated based on the circONCOS backbone, i.e., containing identical flanking sequences and splicing sites. Here, the relationship between the proximity of the IRES and the splicing sites was investigated. A circRNA expression cassette was generated in which a continuous eGFP ORF was inserted downstream of the CVB3 IRES element. Furthermore, a noncoding "spacer" sequence derived from HIPK3 exon 2 was inserted into the flanking regions of the IRES-ORF cassette ( Figure 4 A). To ensure that each encoded circRNA was the same size, a 403nt spacer sequence was added to each design; however, the amount of spacer sequence inserted upstream and downstream of the IRES-ORF varied between each circRNA design ( Figure 4 A). These designs were named based on the size of the HIPK3 stuffer sequence inserted upstream of the IRES. For example, circH1 inserted 1 nt upstream of the IRES and 402 nt downstream of the ORF (Table 2). A375 cells were transfected with plasmids encoding circEGFP containing different HIPK3-derived spacer sequences flanking the upstream of the IRES / ORF. RT-PCR ( Figure 4 B) and qRT-PCR ( Figure 4C) Analysis of circRNA expression clearly demonstrated an association between the distance between IRES and upstream SA and circRNA expression levels, with an increase in the distance between IRES and SA positively correlated with circRNA expression. In other words, circH1, circH50, and circH100 exhibited significantly decreased circRNA expression levels compared to circH200 and circH400 ( Figure 4 C). RT-PCR showed that circRNA expression levels were similar between circH200 and circH400. Overall, these data suggest that positioning the IRES away from the SA increases circRNA production.
[0232]
[0233] Table 2: Nucleotide distances between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP_HIPK3 spacer sequence circRNA expression cassette design.
[0234] Example 5: Separate ORF design provides improved circRNA production and higher protein translation
[0235] To validate the observation that a split ORF design is an excellent circRNA expression design independent of a specific ORF, constructs encoding different transgenes were examined for circRNA expression. Three additional transgenes encoding the co-stimulatory molecule ICOS-L, the adenosine deaminase ADA, and the reporter gene Renilla luciferase were examined. Three circRNA designs were tested in which the contiguous transgene ORF was placed upstream (circInv) or downstream (circCont) of the IRES, in addition to a central split ORF design (circSplit) in which at least 30% of the ORF was located between the 3' end of the IRES and the SD ( Figure 5 A; Table 3). Additional split designs were also tested for ICOS-L. For ICOS-L, both cassettes split at the same motif, but version 2 contained a short spacer sequence upstream of the start codon. RNA and protein levels were also compared to the corresponding linear mRNA.
[0236] A375 cells were transfected with plasmids encoding circular and linear versions of ICOS-L (SEQ ID NO: 21), Renilla (SEQ ID NO: 22), and FLAG-tagged ADA (SEQ ID NO: 20). Western blot analysis of protein expression of the circRNA expression cassettes showed that the expression level of the circSplit design was significantly higher than that of the circInv or circCont designs ( Figure 5 BD). In addition, qRT-PCR analysis showed that circRNA expression was superior in the circSplit design compared with the circInv or circCont design in all three transgenics ( Figure 5 These observations support the hypothesis that splitting the ORF in the center and placing the CVB3IRES inside the ORF results in superior circRNA expression at both the RNA and protein levels.
[0237]
[0238] Table 3: Nucleotide distances between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP_HIPK3 spacer circRNA expression cassette design.
[0239] Furthermore, we observed that for each transgene, the circSplit designed proteins expressed at similar levels compared to the corresponding linear mRNA responses ( Figure 5 B ADA; Figure 5 D Renilla luciferase) or better ( Figure 5 C ICOS-L). It is worth noting that although the RNA level of circRNA is significantly lower than that of the corresponding mRNA ( Figure 5 This suggests that the CVB3 IRES, combined with a split ORF design, is a powerful tool for transgenic protein expression, potentially outperforming traditional linear mRNA expression cassettes even at early time points when mRNA expression levels are highest.
[0240] Example 6: The position of the IRES within the circRNA expression cassette affects circRNA expression independently of the expressed gene or the IRES used
[0241] As demonstrated in the above examples, the position of IRES within circRNA significantly enhances circRNA expression. The same results were obtained using a different gene (ICOSL). Similar to the above experiments, the ICOSL ORF was separated at different positions and the IRES was placed between the stop codon (upstream of the IRES) and the start codon (downstream of the IRES) to adjust the distance between the IRES and the splice acceptor (SA) and splice donor (SD) sites. The exact separation positions are reported in Table 4.
[0242] Briefly, A375 cells were transfected with plasmids encoding circICOSL with different split ORF designs. RNA and protein levels were examined 48 hours after transfection. Figure 6 A) and qRT-PCR ( Figure 6 B) Protein and circRNA expression analysis was performed. As previously shown, there is a negative correlation between the proximity of the IRES to the splice site and circRNA expression. When the split was performed near the 5' end (Split_18) or near the 3' end (Split_870) of the ORF, the expression of circRNAs at both the RNA and protein levels was reduced compared to the splits closer to the center (Split_236 and Split_436). In addition, as previously reported, splitting the IRES (i33, i219, i317) reduced circRNA expression at both the protein and RNA levels ( Figure 6 AB). Overall, these results further support the observation that placing the IRES in a central position within the dividing ORF, such that the distance between the IRES and the SA and SD is close, results in the highest levels of protein expression.
[0243]
[0244] Table 4: Nucleotide distances between the IRES element and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circICOSL split circRNA expression cassette design.
[0245] Further supporting this observation, a similar circEGFP expression pattern was observed when the CVB3 IRES (741 nt; class I IRES) used in the previous example was replaced with the EMCV IRES (566 nt; class II IRES) in the circRNA vector encoding eGFP (Table 5). Figure 7 AB) or EMCV IRES ( Figure 7A375 cells were transfected with circEGFP plasmids with different split ORF designs (CD), and protein and RNA levels were examined 48 hours after transfection. Figure 7 A and C) and qRT-PCR ( Figure 7 B and D) were used for protein and circRNA expression analysis. This supports that, regardless of the IRES element used, the positional relationship of ORF-IRES is crucial for high-yield circRNA expression. It was observed that the circRNA expression patterns of the CVB3 and EMCV circRNA designs were similar, with the separations closer to the center (separate_363 and separate_436) showing the highest levels of circRNA expression at both the protein and RNA levels. In addition, similar to the results in the above examples, separations closer to the 5' and 3' ends of the ORF had a negative impact on the biosynthesis of circRNA. Moreover, separating the EMCV IRES and placing the ORF in this region also inhibited the expression of high-yield circRNA ( Figure 7 CD), which is similar to the observations of CVB3IRES split design ( Figure 7 AB).
[0246]
[0247] Table 5: Nucleotide distances between IRES or eGFP_ORF and the upstream splice acceptor (SA) or downstream splice donor (SD) in each circEGFP split circRNA expression cassette design.
[0248] Overall, these data support the hypothesis that placing the IRES in the center of the split ORF design achieves the highest levels of circRNA expression. The inventors hypothesize that this increased expression is a result of increasing the distance between the IRES and the flanking splice sites required for backsplicing, as the central split maximizes the distance between the IRES and either splice site.
Claims
1. A nucleic acid molecule encoding a circular RNA (circRNA), wherein the nucleic acid molecule comprises: A) An expression cassette comprising: (i) a circRNA expression cassette comprising: (a) a nucleic acid comprising or consisting of continuous or separate open reading frames (ORFs) encoding at least one protein, and (b) an internal ribosome entry site (IRES) operably linked to the ORF to direct translation of the ORF; (ii) a first back-splicing site located at the 5' end of the IRES; and (iii) a second back-splicing site located at the 3' end of the IRES; and optionally B) a first inverted repeat (IR) element located 5' to the first back-splice site and a second IR element located 3' to the second back-splice site; and / or a promoter operably linked to the expression cassette to direct expression of the expression cassette; Wherein the IRES is located within the circRNA expression cassette, wherein: (I) the 5' end of the IRES is separated from the 3' end of the first back-splicing site by at least 50 nucleotides, preferably at least 200 nucleotides, more preferably at least 300 nucleotides, and (II) The distance between the 3' end of the IRES and the 5' end of the second back-splicing site is at least 300 nucleotides, preferably at least 350 nucleotides.
2. The nucleic acid molecule according to claim 1, wherein the split ORF comprises two parts, and the first part of the split ORF comprises a stop codon and is located 5' to the IRES, and the second part of the split ORF comprises a start codon and is located 3' to the IRES.
3. The nucleic acid molecule according to any one of the preceding claims, wherein the first back-splicing site comprises two half-sites, wherein ● the first and second half sites are located at the 5' end of the circRNA expression cassette; or The first half-site is located at the 5' end of the circRNA expression cassette, and the second half-site is part of the circRNA expression cassette, preferably part of the ORF in case of a separate ORF; And, wherein the second back-splicing site comprises two half-sites, wherein ● the first and second half sites are located at the 3' end of the circRNA expression cassette; or • The first half-site is part of the circRNA expression cassette, preferably part of the ORF in case of a separate ORF, and the second half-site is located at the 3' end of the circRNA expression cassette.
4. A nucleic acid molecule according to any one of the preceding claims, wherein the first half-site of the first back-splice site comprises or consists of the nucleotides AG; and / or • The second half-site of the second back-splice site comprises or consists of the nucleotides GT.
5. The nucleic acid molecule according to any one of the preceding claims, wherein the IRES is a viral IRES, Preferably, it is an IRES of a virus selected from the group consisting of Coxsackievirus (CVB), preferably CVB3 and encephalomyocarditis virus (EMCV); or Preferably, the IRES is selected from the group consisting of the following virus families: Adenoviridae, Arenaviridae, Birnaviridae, Chrysoviridae, Coronaviridae, Dicistroviridae, Filoviridae, Flaviviridae, Hepadnaviridae, Herpesviridae, Hypoviridae, Iflaviridae, Luteoviridae, Orthomyxoviridae. xoviridae), Papillomaviridae, Paramyxoviridae, Parvoviridae, Picornaviridae, Pneumoviridae, Polyomaviridae, Potyviridae, Reoviridae, Retroviridae, Rhabdoviridae, Secoviridae, Tombusviridae, Totiviridae, Virgaviridae; or Preferably, it is a variant of the viral IRES, the nucleotide sequence of which is at least 90% identical to the nucleotide sequence of the viral IRES.
6. The nucleic acid molecule according to any one of the preceding claims, wherein the first IR element and the second IR element comprise sequences that are at least 80%, preferably at least 90%, more preferably at least 95% identical to the reverse complement of each other.
7. The nucleic acid molecule according to any of the preceding claims, wherein the IR element is selected from the group consisting of IR elements derived from the following genes: AC004076.9, ACYP2, AMD1, ARHGAP10, ARHGAP12, ARHGEF12, ARHGEF28, ASAP1, ASH1L, ASXL1, ATXN2, BMPR2, BRWD1, BTBD10, CBFA2T2, CCDC126, CCDC134, CCDC66, CCDC7, CCDC9, CCNB1, CDK13, CFLAR, CHD9, CLIP2, CLNS1A, CNN2, COA1, CORO1C, CREBBP, C RKL, CTB-43P18.3, SNHG4, DCUN1D4, DEK, DHRS3, DLG1, DOPEY2, DYNC1H1, ELF2, EMC2, EPHB4, EPS15, ERC1, ETFA, EXOSC1, FAM13B, FARSA, FBXO7, FGD4, FG D6, FKBP3, FKBP8, FNTA, FOXK2, GAPVD1, GBAS, GDI2, GLIS2, GLS, GON4L, GRHPR, HERC1, HIPK3, HNRNPM, HOOK3, HP1BP3, HPS5, HTT, HUWE1, IARS, ILKAP, IQ GAP1, KDM1A, KIAA0368, KIAA1429, KIAA1841, KLHL8, KMT2C, LMBR1, LRCH3, LZIC, MAP3K1, MARK4, MBOAT2, MCU, MED13L, METTL3, MGA, MGEA5, MITD1, MORC 3. MRPS35, MYO9B, NCOA2, NFAT5, NFATC3, NFX1, NUDC, NUP54, PAFAH1B2, PAIP2, PCMT1, PDCD11, PDE8A, PDS5A, PHC3, PHLDB2, PLEKHM1, PLEKHM3, PLOD2, P MS1, PNN, POLR2A, POMT1, PPP6R2, PROSC, PRRC2B, PSEN1, PSMA7, PTP4A2, PTPN12, QKI, R3HDM1, RAB6A, RALBP1, RARS, RBM23, RBM33, RBM39, RELL1, REPS1 , RERE, RHOBTB3, RLF, RNF19B, SDHAF2, ZRANB1, FAM228B, RPRD1B, RPS6KC1, RSF1, RSRC1, RTN4, RYK, SAFB2, SAMD4A, SCAF8, SCARF1, SCYL2, SDF4, SEC31A,SENP6, SIPA1L1, SKA3, SLC38A1, SLTM, SMARCA5, SMC3, SMO, SNX25, SOBP, SOS2, SPIDR, SPPL3, SRSF4, STAM, STK3, STX6, TERF 2. TIMMDC1, TMED2, TMEM138, TMEM165, TMEM181, TNPO1, TNPO3, TOP1, TTC39C, TTLL1, UBA2, UBAP2, UBE2K, UBQLN1, UBR5, UBX N2A, UBXN7, UHRF2, UIMC1, URI1, UTP18, UTRN, VAMP3, VAPB, VMP1, WDR78, XPO1, YTHDF2, YWHAE, YY1AP1, ZBTB46, ZCCHC11, ZCCHC6, ZFAND6, ZFX, ZKSCAN1, ZMYM4, ZNF124, ZNF236, ZNF394, ZNF430, ZNF652, ZNF720, ZNF91], preferably IR elements derived from HIPK3 and ZKSCAN1.
8. The nucleic acid molecule according to any one of the preceding claims, wherein the promoter is selected from a constitutive, tissue-specific or inducible promoter, preferably a viral promoter, more preferably selected from the group consisting of the CMV immediate early promoter, the SV40 promoter.
9. The nucleic acid molecule of any one of the preceding claims, wherein the circRNA expression cassette further comprises: (i) one or more non-coding nucleotide sequences, if the circRNA expression cassette comprises a continuous ORF, the one or more non-coding nucleotide sequences are preferably located at the 5' end and / or 3' end of the continuous ORF; or if the circRNA expression cassette comprises a separate ORF, the one or more non-coding nucleotide sequences are preferably located at the 5' end and / or 3' end of the IRES; (ii) one or more additional ORFs; (iii) one or more additional IRES; (iv) an additional coding nucleotide sequence appended to one or more ORFs, which preferably encodes a sequence that can be used to detect or capture the protein encoded in the one or more ORFs.
10. The nucleic acid molecule according to any one of the preceding claims, wherein the nucleic acid molecule further comprises: (i) branch points; and / or (ii) Polypyrimidine strings.
11. The nucleic acid molecule of any preceding claim, wherein the ORF encodes a therapeutic protein, a therapeutic peptide, an antigenic protein or an antigenic peptide. 12 . A circRNA encoded by the circRNA expression cassette of the nucleic acid molecule according to claim 1 . 13 . A vector comprising the nucleic acid molecule according to claim 1 or the circRNA according to claim 12 . 14 . A host cell comprising the nucleic acid molecule according to claim 1 , the circRNA according to claim 12 , or the vector according to claim 13 . 15 . A pharmaceutical composition comprising the nucleic acid molecule according to any one of claims 1 to 11 , the circRNA according to claim 12 , or the vector according to claim 13 .
16. A method for producing circular mRNA (circRNA), the method comprising: (i) introducing the nucleic acid molecule according to any one of claims 1 to 11 or the vector according to claim 13 into a eukaryotic cell, and optionally purifying the circRNA from the cell; or (ii) contacting the nucleic acid molecule according to any one of claims 1 to 11 or the vector according to claim 13 with isolated RNA polymerase II.
17. A method for producing a recombinant protein, the method comprising introducing the nucleic acid molecule according to any one of claims 1 to 11, the circRNA according to claim 12, or the vector according to claim 13 into a eukaryotic cell, and optionally purifying the recombinant protein encoded by the ORF.