Method for producing self-circularized RNA
Patent Information
- Application Number
- JP2025535975
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-15
- Publication Date
- 2026-01-08
AI Technical Summary
The short half-life of messenger RNA (mRNA) in biological systems poses a limitation for its therapeutic and engineering applications due to exonuclease degradation, necessitating a method to extend protein expression duration.
A method for producing circular RNA by transcribing a vector that forms a precursor RNA with specific elements, including Group I self-splicing introns and internal ribosome entry sites (IRES), which self-cleave and self-ligate to form thermodynamically stable multidirectional junctions, enhancing translation and stability.
The method produces stable circular RNA that can efficiently translate protein coding regions and maintain biological activity, overcoming the limitations of linear mRNA half-life and facilitating efficient protein expression.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is an international (PCT) application claiming priority to U.S. Provisional Patent Application No. 63 / 476,864, filed December 22, 2022, the contents of which are specifically incorporated by reference in their entirety for all purposes. Incorporation of a sequence listing submitted as a conforming ASCII text file (.xml)
[0002] In accordance with the EFS-Web legal framework and 37 CFR § 1.821-825 (see MPEP § 2442.03(a)), a sequence listing submitted as an ASCII-compatible text file format (titled "3000076-009977_sequence_listing_ST26.xml", created on December 14, 2023, and 110,393 bytes in size) has been filed contemporaneously with the present application, and the entire contents of the sequence listing are incorporated herein by reference. [Background technology]
[0003] Messenger RNA (mRNA) has broad potential for a range of therapeutic and engineering applications. One important limitation to its use may be its relatively short half-life in biological systems. Therefore, there is a need to extend the duration of protein expression from full-length RNA messages. Summary of the Invention
[0004] In one aspect, the disclosure relates to a method of producing circular RNA, the method comprising transcribing a vector to form a precursor RNA, wherein the vector comprises the following elements operably linked to each other and arranged in the following order: a) a 5' element that does not contain or contains at least one stem-loop structure, b) a 3' Group I self-splicing intron fragment that contains a 3' splice site dinucleotide, c) an element that does not contain or contains an internal ribosome entry site (IRES) and a protein coding region or an element that contains a non-coding region, d) a 5' Group I self-splicing intron fragment that contains a 5' splice site dinucleotide, e) a 5' Group I self-splicing intron fragment that contains a 5' splice site dinucleotide, It comprises a 3' element that does not contain or contains at least one stem-loop structure, provided that if the 5' element does not contain a stem-loop structure, the 3' element contains at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element contains at least one stem-loop structure, wherein the 5' element and the 3' element form a thermodynamically stable multidirectional junction RNA structure, the precursor RNA of which can form a circular RNA that is translatable and / or biologically active in a cell.
[0005] In another embodiment, the thermodynamically stable multidirectional junction RNA structure may be a three-way junction (3WJ), a four-way junction (4WJ), a five-way junction (5WJ), a hand-in-hand interaction, a kissing loop, or a pseudoknot. In a hand-in-hand interaction structure, one loop sequence may have high affinity with another loop sequence to form a closed, interacting dimeric structure (Shu et al., "Stable RNA nanoparticles as potential new generation drugs for cancer therapy." Advanced drug delivery reviews. 66 (2014) 74-89; the contents of which are incorporated herein by reference in their entirety). A pseudoknot structure may comprise at least two stem-loop structures, in which half of one stem may interact with half of the other stem (Staple et al., "Pseudoknots: RNA structures with diverse functions." PLOS Bio 3(6):e213; the contents of which are incorporated herein by reference in their entirety).
[0006] In another embodiment, 3WJ may comprise a first branch of the 3WJ domain, which may be formed from the 5' portion of the 3WJa sequence and the 3' portion of the 3WJc sequence and may comprise a first helical region; a second branch of the 3WJ domain, which may be formed from the 3' portion of the 3WJa sequence and the 5' portion of the 3WJb sequence and may comprise a second helical region; and a third branch of the 3WJ domain, which may be formed from the 3' portion of the 3WJb sequence and the 5' portion of the 3WJc sequence and may comprise a third helical region, wherein each of the helical regions may comprise a plurality of RNA nucleotide pairs forming a standard Watson-Crick bond.
[0007] In another embodiment, the 3WJa, 3WJb and 3WJc sequences may be represented as follows: 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 4, 3WJb comprises or consists of SEQ ID NO: 5, and 3WJc comprises or consists of SEQ ID NO: 6, or 3WJa comprises or consists of SEQ ID NO: 10, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 12, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 13, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 17, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 20, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 22, or 3WJa comprises or consists of SEQ ID NO: 23, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 26, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 28, or 3WJa comprises or consists of SEQ ID NO: 29, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 32, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 33, and 3WJc comprises or consists of SEQ ID NO: 34, or 3WJa comprises or consists of SEQ ID NO: 41, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 44, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 50, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 53, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 56, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 58, or 3WJa comprises or consists of SEQ ID NO: 59, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 60, or 3WJa comprises or consists of SEQ ID NO: 61, 3WJb comprises or consists of SEQ ID NO: 63, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of SEQ ID NO: 66, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of UGUCACGGG, and 3WJc comprises or consists of SEQ ID NO: 68, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 69, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 70, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 71, and 3WJc comprises or consists of SEQ ID NO: 72, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 8, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 78, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 81, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 84, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 7, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 90, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 93, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 95, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 112, 3WJb comprises or consists of SEQ ID NO: 113, and 3WJc comprises or consists of SEQ ID NO: 114, or 3WJa comprises or consists of SEQ ID NO: 119, 3WJb comprises or consists of SEQ ID NO: 120, and 3WJc comprises or consists of SEQ ID NO: 121, or 3WJa includes AUGUGUA, 3WJb includes UACUUUG, and 3WJc includes AUCAUG, or 3WJa contains GCGUU, 3WJb contains UUCGC, and 3WJc contains GCCAUAGCG, or 3WJa includes GUAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUAC, or 3WJa includes AUGGUA, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGUGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUGG, 3WJb includes UCACGG, and 3WJc includes CCUCUUA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUG, 3WJb includes CAGGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UAUGU, 3WJb includes GCAGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UGUGU, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or 3WJa contains UGUGU, 3WJb contains ACUUU, and 3WJc contains AAAUCA.
[0008] In another aspect, 4WJ may comprise: a first branch of the 4WJ domain, which may be formed from the 5' portion of the 4WJa sequence and the 3' portion of the 4WJd sequence and may comprise a first helical region; a second branch of the 4WJ domain, which may be formed from the 3' portion of the 4WJa sequence and the 5' portion of the 4WJb sequence and may comprise a second helical region; a third branch of the 4WJ domain, which may be formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence and may comprise a third helical region; and a fourth branch of the 4WJ domain, which may be formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence and may comprise a fourth helical region, wherein each of the helical regions may comprise a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds.
[0009] In another embodiment, the 4WJa, 4WJb, 4WJc and 4WJd sequences may be represented as follows: 4WJa comprises or consists of SEQ ID NO: 7, 4WJb comprises or consists of SEQ ID NO: 8, 4WJc comprises or consists of SEQ ID NO: 9, and 4WJd comprises or consists of SEQ ID NO: 102, or 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO: 106, or 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO: 118, or 4WJa comprises UGCAGGUG, 4WJb comprises ACGGGC, 4WJc comprises CCAGCA, and 4WJd comprises SEQ ID NO: 67, or 4WJa comprises SEQ ID NO: 74, 4WJb comprises AACUG, 4WJc comprises SEQ ID NO: 75, and 4WJd comprises AUCAUG, or 4WJa contains SEQ ID NO: 122, 4WJb contains GAACU, 4WJc contains SEQ ID NO: 123, and 4WJd contains AAUCA.
[0010] In another embodiment, 5WJ comprises a first branch of the 5WJ domain that may be formed from a 5' portion of the 5WJa sequence and a 3' portion of the 5WJe sequence and may include a first helical region, a second branch of the 5WJ domain that may be formed from a 3' portion of the 5WJa sequence and a 5' portion of the 5WJb sequence and may include a second helical region, and a third branch of the 5WJ domain that may be formed from a 3' portion of the 5WJb sequence and a 5' portion of the 5WJc sequence and may include a third helical region. a third branch of the 5WJ domain that may be formed from a 3' portion of the 5WJc sequence and a 5' portion of the 5WJd sequence and may include a fourth helical region; and a fifth branch of the 5WJ domain that may be formed from a 3' portion of the 5WJd sequence and a 5' portion of the 5WJe sequence and may include a fifth helical region, wherein each of the helical regions may include multiple RNA nucleotide pairs that form standard Watson-Crick bonds.
[0011] In another aspect, 5WJa comprises or consists of SEQ ID NO: 107, 5WJb comprises or consists of SEQ ID NO: 108, 5WJc comprises or consists of SEQ ID NO: 109, 5WJd comprises or consists of SEQ ID NO: 110, and 5WJe comprises or consists of SEQ ID NO: 111; or 5WJa comprises GUGA, 5WJb comprises UUGC, 5WJc comprises GUGU, 5WJd comprises AUGC, and 5WJe comprises GUGC.
[0012] In another embodiment, the vector may further comprise an internal ribosome entry site (IRES) located at the 5' end of c), wherein the IRES is an IRES that is capable of encoding Taura syndrome virus, Triatoma virus, Theiler's encephalomyelitis virus, Simian virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, reticuloendotheliosis virus, human poliovirus 1, Plautia stali intestine virus, Kashmir bee virus, human rhinovirus 2, Homalodisca coagulata virus-1, human immunodeficiency virus type 1, leafhopper virus-1, Himetobi P virus, hepatitis C virus, hepatitis A virus, hepatitis G virus, virus), foot-and-mouth disease virus, human enterovirus 71, equine rhinitis virus, white-spotted picorna-like virus, encephalomyocarditis virus (EMCV), Drosophila C virus, cruciferous tobamovirus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen brood virus, aphid fatal paralysis virus, avian encephalomyelitis virus, bee acute paralysis enterovirus, hibiscus chlorotic ringspot virus, classical swine fever virus, human fibroblast growth factor 2 (FGF2), human surfactant Drug protein A1 (SFTPA1), human acute myeloid leukemia protein 1 / runt-related transcription factor 1 (AML1 / RUNX1), Drosophila antennapedia, human aquaporin-4 (AQP4), human angiotensin II receptor type 1 (AT1R), human BCL2-associated immortality gene 1 (BAG-1), human B-cell lymphoma 2 (BCL2), human immunoglobulin-binding protein (BiP), human inhibitor of apoptosis family protein 1 (c-IAP1), human c-myc, human eukaryotic translation initiation factor 4 G (eIF4G), mouse N-deacetylase and N-sulfotransferase 4 (NDST4L), human lymphoid enhancer-binding factor-1 (LEF1), mouse hypoxia-inducible factor 1 subunit alpha (HIF1α), human N-myc, mouse glial cell and testis-specific homeobox protein (Gtx), human cyclin-dependent kinase inhibitor 1B (p27kip1), human platelet-derived growth factor B / simian sarcoma virus homolog (PDGF2 / c-sis), human p53, human proviral integration site-1 of Moloney murine leukemia virus (Pim-1), mouse RNA-binding protein white 3 (Rbm3), Drosophila reaper, canine Scamper, Drosophila superbreast gene (Ultrabithorax (Ubx)), Salivirus, coronavirus, parechovirus, human N-ras upstream (UNR), mouse dystrophin-related protein (utrophin) A (UtrA), human vascular endothelial growth factor A (VEGF-A), human X-linked apoptosis protein (XIAP), Drosophila hairless, budding yeast transcription factor II The IRES sequence may be selected from IRES sequences from viruses or genes selected from the group consisting of TFIID, Saccharomyces cerevisiae Yes1-associated transcription factor (YAP1), human proto-oncogene tyrosine protein kinase Src (c-src), human fibroblast growth factor 1 (FGF-1), monkey virus, turnip crinkle virus, coxsackievirus B3 (CVB3), and coxsackievirus A (CVB1 / 2).
[0013] In another embodiment, the vector may further comprise an RNA polymerase promoter.
[0014] In another embodiment, the RNA polymerase promoter may be a T7 viral RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter.
[0015] In another embodiment, the 3' group I self-splicing intron fragment and the 5' group I self-splicing intron fragment may be derived from a Cyanobacterium anabaena sp. Pre-tRNA-Leu gene.
[0016] In another embodiment, the 3' group I self-splicing intron fragment and the 5' group I self-splicing intron fragment may be derived from the T4 bacteriophage Td gene.
[0017] In another embodiment, the methods of the present disclosure may further comprise forming the circular RNA by splint-mediated precursor RNA ligation.
[0018] In another embodiment, the vector may be transfected into cells using lipofection or electroporation prior to transcription.
[0019] In another embodiment, the vector may be transfected into cells using a nanocarrier prior to transcription.
[0020] In another embodiment, the nanocarrier may be a lipid, a polymer, or a lipid-polymer hybrid.
[0021] In another embodiment, the methods of the present disclosure may further comprise forming a circular RNA and purifying the circular RNA using a size exclusion chromatography column in tris-EDTA or ion-pair reverse phase HPLC.
[0022] In another embodiment, the methods of the present disclosure may further comprise forming a circular RNA and purifying the circular RNA in a high performance liquid chromatography (HPLC) system in a triethylammonium acetate (TEAA)-acetonitrile buffer solution having a pH range of about 4 to 10 at a flow rate of about 0.01 to 5 mL / min.
[0023] In another embodiment, the methods of the present disclosure may further comprise forming circular RNA and purifying the circular RNA using phosphatase treatment.
[0024] In another embodiment, the methods of the present disclosure may further comprise incubating the precursor RNA in the presence of (i) magnesium ions and / or (ii) guanosine nucleotides or guanosine nucleosides.
[0025] In another embodiment, the precursor RNA may be incubated at a temperature between about 20°C and about 60°C.
[0026] In another embodiment, transcription of the vector can occur in the presence of a nucleoside or a monophosphate or diphosphate nucleotide for incorporation of said nucleoside or nucleotide as the first nucleotide of a precursor RNA transcribed from said vector.
[0027] In another embodiment, the precursor RNA may contain a monophosphate 5' end that can be ligated using a ligase.
[0028] In another embodiment, transcription of the vector can occur in the presence of: a) a guanosine nucleoside or monophosphate or diphosphate nucleotide; b) a cytidine nucleoside or monophosphate or diphosphate nucleotide; c) a uracil nucleoside or monophosphate or diphosphate nucleotide; d) an adenosine nucleoside or monophosphate or diphosphate nucleotide; or e) a combination thereof, for incorporation of the nucleoside or monophosphate or diphosphate nucleotide as the first nucleotide in an RNA strand transcribed from the vector or in a transcript produced from the vector.
[0029] In another embodiment, the protein coding region may encode a non-naturally occurring protein that includes one or more synthetic protein elements.
[0030] In another embodiment, the vector may include a 5' spacer element located at the 3' end of b).
[0031] In another embodiment, the vector may include a 3' spacer element located at the 5' end of d).
[0032] In another embodiment, the 5' spacer element or the 3' spacer element may comprise a polyA sequence or a polyA-C sequence.
[0033] In another aspect, the non-coding region is an antisense RNA, transfer RNA (tRNA), transfer-messenger RNA (tmRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), Cajal body-specific small RNA (scaRNA), guide RNA (gRNA), Y RNA, splicing leader sequence RNA (SL RNA), microRNA (miRNA), small interfering RNA (siRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long non-coding RNA (lncRNA), Piwi-interacting RNA (piRNA), short hairpin RNA (shRNA), trans-acting siRNA (tasiRNA), repeat-associated siRNA (rasiRNA), 7SK The gene may include an element encoding one or more RNAs selected from the group consisting of RNA (7SK), telomerase RNA component (TERC), vault RNA (vRNA, vtRNA) and enhancer RNA (eRNA).
[0034] In one aspect, the disclosure relates to a precursor RNA comprising the following elements operably linked to each other and arranged in the following order: a) a 5' element that does not contain or contains at least one stem-loop structure, b) a 3' Group I self-splicing intron fragment comprising a 3' splice site dinucleotide, c) a protein coding or non-coding region, d) a 5' Group I self-splicing intron fragment comprising a 5' splice site dinucleotide, and e) a 3' element that does not contain or contains at least one stem-loop structure, provided that if the 5' element does not contain a stem-loop structure, the 3' element comprises at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element comprises at least one stem-loop structure, wherein the 5' element and the 3' element form a thermodynamically stable multidirectional junction, and wherein the precursor RNA is capable of forming a circular RNA that is translatable and / or biologically active in a cell.
[0035] In one embodiment, the vector may encode a precursor RNA of the present disclosure.
[0036] In alternative embodiments, the vector may be a plasmid, a viral vector, a polymerase chain reaction (PCR) product, a cosmid, a bacterial artificial chromosome (BAC), or a yeast artificial chromosome (YAC).
[0037] In one aspect, the disclosure relates to a method for producing a circular ribonucleic acid (RNA), the method comprising transcribing a vector to form a precursor RNA, the precursor RNA comprising the following elements operably linked to each other and arranged in tandem in a 5' to 3' direction: a) a 5' element, b) a 3' Group I self-splicing intron fragment comprising a 3' splice site dinucleotide, c) an element that does not include or includes an internal ribosome entry site (IRES) and a protein coding region, or an element that includes a non-coding region, d) a 5' Group I self-splicing intron fragment comprising a 5' splice site dinucleotide, and e) and a 3' element, wherein the 5' element and the 3' element form a stable structure with a Gibbs free energy (ΔG) of -190 kcal / mol to -9.0 kcal / mol, provided that the stable structure is not a duplex with at least 95% base pairs between the 5' element and the 3' element, and wherein the 3' Group I self-splicing intron fragment and the 5' Group I self-splicing intron fragment self-cleave and self-ligate to form a circular RNA molecule.
[0038] In one aspect, the disclosure relates to a precursor RNA comprising the following elements operably linked to each other and arranged in tandem in a 5' to 3' direction: a) a 5' element, b) a 3' Group I self-splicing intron fragment comprising a 3' splice site dinucleotide, c) an element that does not include or includes an IRES and a protein coding region, or an element that includes a non-coding region, d) a 5' Group I self-splicing intron fragment comprising a 5' splice site dinucleotide, and e) a 3' element, wherein the 5' element and the 3' element form a stable structure with a Gibbs free energy (ΔG) of -190 kcal / mol to -9.0 kcal / mol, provided that the stable structure is not a duplex with at least 95% base pairs between the 5' element and the 3' element, and wherein the 3' Group I self-splicing intron fragment and the 5' Group I self-splicing intron fragment self-cleave and self-ligate to form a circular RNA molecule, thereby generating a circular RNA.
[0039] In one aspect, the present disclosure relates to a method of producing a protein in a cell, the method comprising introducing a precursor RNA containing a protein coding region of the present disclosure or a vector containing the protein coding region into a cell and producing the protein.
[0040] In one aspect, the present disclosure relates to a method of editing a gene in a cell, the method comprising introducing into a cell a precursor RNA containing a gene-editable non-coding region of the disclosure or a vector containing a gene-editable non-coding region, and editing the gene. [Brief explanation of the drawings]
[0041] [Figure 1] 1 illustrates a "three-way junction" ("3WJ") domain according to one embodiment of the present disclosure. [Figure 2] 1 shows a 3WJ domain according to another embodiment of the present disclosure. [Figure 3] 1 shows a 3WJ domain according to another embodiment of the present disclosure. [Figure 4] 1 shows a method for producing circularized RNA (circRNA) according to one embodiment of the present disclosure. [Figure 5] 1 shows a method for producing circularized RNA (circRNA) according to another embodiment of the present disclosure. [Figure 6] 1 illustrates a 4WJ domain according to one embodiment of the present disclosure. [Figure 7] 1 shows a 4WJ domain according to another embodiment of the present disclosure. [Figure 8] 1 shows a method for producing circularized RNA (circRNA) according to one embodiment of the present disclosure. [Figure 9] 1 illustrates a 5WJ domain according to one embodiment of the present disclosure. [Figure 10] 1 shows a method for producing circularized RNA (circRNA) according to one embodiment of the present disclosure. [Figure 11] 1 shows a method for producing circularized RNA (circRNA) according to another embodiment of the present disclosure. [Figure 12] 1 shows the purity of circRNA produced by methods according to some embodiments of the present disclosure. [Figure 13] 1 shows the circRNA / precursor RNA ratio of circRNA produced by methods according to some embodiments of the present disclosure. [Figure 14A] 1 shows HPLC-purified circRNA produced by a method according to one embodiment of the present disclosure. [Figure 14B] This shows the purity of circRNA produced by a method according to one embodiment of the present disclosure. [Figure 15A] 1 shows HPLC-purified circRNA produced by a method according to another embodiment of the present disclosure. [Figure 15B] 1 shows the purity of circRNA produced by a method according to another embodiment of the present disclosure. [Figure 16] 1 shows gene expression of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 17] 1 shows gene expression of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 18] 1 shows the time course of gene expression of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 19] 1 shows gene expression over time of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 20] 1 shows the purity of circRNA produced by methods according to some embodiments of the present disclosure. [Figure 21A] 1 shows RNA circularization efficiency as determined by agarose gel electrophoresis and densitometry analysis, according to some embodiments of the present disclosure. [Figure 21B] 1 shows RNA circularization efficiency as determined by high performance liquid chromatography (HPLC), according to some embodiments of the present disclosure. [Figure 21C] 1 shows RNA circularization efficiency as determined by HPLC according to some embodiments of the present disclosure. [Figure 22A] 1 shows the time course of gene expression of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 22B] The fold change in gene expression in Figure 22A is shown. [Figure 23A] 1 shows gene expression over time of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 23B] The fold change in gene expression in Figure 23A is shown. [Figure 24A] 1 shows the copy number over time of circRNAs produced by methods according to some embodiments of the present disclosure. [Figure 24B] The copy number fold change in Figure 24A is shown. DETAILED DESCRIPTION OF THE INVENTION
[0042] mRNA has proven to be an excellent platform for developing RNA-based therapeutics and vaccines. The development of human mRNA vaccines may require a 5' cap and 3' polyA component for efficient mRNA expression in vivo, which may require a cap-dependent mechanism for ribosome recruitment for translation. These requirements may pose challenges for the production of high-quality mRNA, as the polyA tail may be lost during the plasmid production step. Furthermore, the stability of linear mRNA sequences may pose another limitation in in vitro storage and in vivo half-life due to exonuclease degradation.
[0043] Exogenous circular RNAs have been developed to extend the duration of protein expression from full-length RNA sequences. Exogenous RNA circularization can be achieved using three general strategies: chemical methods using cyanogen bromide or similar condensing agents, enzymatic methods using RNA or DNA ligases, and ribozyme methods using self-splicing introns. Ribozyme methods utilizing a substituted group I catalytic intron are more applicable to long RNA circularization and may require only the addition of GTP and Mg2+ as cofactors. Functional proteins can be produced from these circRNAs in eukaryotic cells, maximizing integration into various internal ribosome entry sites (IRES) and translation of internal polyadenosine fragments. This substituted intron-exon (PIE) splicing strategy can involve fused partial exons flanked by half-intron sequences. In vitro, these constructs can undergo transesterification reactions characteristic of group I catalytic introns, but because the exons are already fused, they are excised as covalently linked 5' to 3' loops. Using this strategy as a starting point for creating protein-coding circular RNAs, Wesselhoeft ("Engineering circular RNA for potent and stable translation in eukaryotic cells." Nat Commun. 2018 July 6;9(1):2629) described a method for producing circular RNAs in vitro using permuted intron and exon designs and utilizing homology arms in the sequences. However, this circularization reaction can be inefficient. Therefore, extensive HPLC purification and RNase R digestion may be required to remove unreacted precursor linear RNA to obtain relatively pure circular RNA for drug development.
[0044] Using PIE (permuted intron-exon) techniques, circular RNAs with sequences on both the 5' and 3' ends can be designed to form highly thermodynamically stable scaffolds with low Gibbs free energy (<-40 kcal / mol). The lower free energy drives ribozyme formation using intron sequences on the 5' and 3' ends of the circular RNA. Thermodynamically stable motifs on both the 5' and 3' ends of the precursor circular RNA may have high affinity to associate with each other to form complex structures, thereby promoting the formation of multidirectional junction motifs and ribozyme structures. For example, a three-way junction (3WJ) may have a fast "on" state for forming the 3WJ structure, with a kinetic rate of 1.37 x 10 5They have association rate constants of M-1 s-1 (Binzel et al., "Mechanism of three-component collision to produce ultrastable pRNA three-way junction of Phi29 DNA-packaging motor by kinetic assessment. RNA. 2016 Nov; 22(11):1710-1718, the contents of which are incorporated herein by reference in their entirety) and very low dissociation constants; they do not dissociate in the picomole range (Shu et al., "Thermodynamically stable RNA three-way junction for constructing multifunctional nanoparticles for delivery of therapeutics." Nature Nanotechnology 2011; 6(10):658-67, the contents of which are incorporated herein by reference in their entirety). This property potentially allows ribozymes to be formed from both ends of long precursor RNAs. Furthermore, the thermodynamic motif formed by this design further forms a high-affinity 3D structure (Zhang et al., "Crystal Structure of 3WJ Core Revealing Divalent Ion-promoted Thermostability and Assembly of the Phi29 Hexameric Motor pRNA." RNA 2013 Sep; 19(9):1226-37; the contents of which are incorporated herein by reference in their entirety), thereby further ensuring correct folding of the adjacent ribozyme structure by the intron.
[0045] As used herein, a "homology arm" or "homology region" or "duplex-forming region" is defined as: 1) a region that is predicted to base-pair with at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, about 100%, or 100% of another sequence (e.g., another homology arm) in the RNA; and 2) 3) is located before, adjacent to, or contained within the 3' intron fragment and / or after, adjacent to, or contained within the 5' intron fragment, and optionally, 4) is at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length and is 250 nt or less; The homology arm or homology region or duplex-forming region may be any contiguous sequence predicted to have less than 50% (e.g., less than 45%, less than 40%, less than 35%, less than 30%, less than 25%) base pairs with an unintended sequence in the RNA (e.g., a non-homologous arm sequence). In some examples, the length of the homology arm or homology region or duplex-forming region may be about 9 to about 50 nucleotides. In one example, the length of the homology arm or homology region or duplex-forming region may be about 9 to about 19 nucleotides. In some examples, the length of the homology arm or homology region or duplex-forming region may be about 20 to about 40 nucleotides. In some examples, the length of the homology arm or homology region or duplex-forming region may be about 30 nucleotides.
[0046] The 5' and 3' homology arms or homology regions or duplex-forming regions may be synthetic sequences and may be different but functionally similar to the internal homology regions. The length of the homology arms or homology regions or duplex-forming regions may be, for example, about 5-50 nucleotides, about 9-19 nucleotides, e.g., about 5, about 10, about 20, about 30, about 40, or about 50 nucleotides. In another example, the length of the homology arms or homology regions or duplex-forming regions may be 9 nucleotides. In a further example, the length of the homology arms or homology regions or duplex-forming regions may be 19 nucleotides. In some examples, the length of the homology arms or homology regions or duplex-forming regions is at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides. In some embodiments, the length of the homology arms or homology regions or duplex-forming regions is 50, 45, 40, 35, 30, 25, or 20 nucleotides or less. In some embodiments, the length of the homology arms or homology regions or duplex-forming regions is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides.
[0047] Unlike conventional 5' and 3' homology arms or homology regions or duplex-forming regions described, for example, in WO 2021236855 and US11447796 (the contents of which are incorporated herein by reference in their entireties), the 5' and 3' regions of the present disclosure can hybridize to form multidirectional junctions (WJs), such as 3WJs, 4WJs, and 5WJs, as described below.
[0048] As used herein, a 3' Group I intron fragment is a contiguous sequence that is at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, 100%) homologous to the 3'-proximal fragment of a naturally occurring Group I intron, includes the 3' splice site dinucleotide, and optionally, the length of the flanking exon sequence is at least 1 nucleotide (e.g., at least 5 nucleotides, at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 50 nucleotides). In one example, the length of the included flanking exon sequence is comparable to the length of the naturally occurring exon. In some embodiments, the 5' Group I intron fragment is a contiguous sequence that is at least 75% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, 100%) homologous to the 5' proximal fragment of a naturally occurring Group I intron, includes the 5' splice site dinucleotide, and optionally, the length of the flanking exon sequence is at least 1 nucleotide (e.g., at least 5 nucleotides, at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, at least 50 nucleotides). In one embodiment, the length of the included flanking exon sequence is comparable to the length of the naturally occurring exon.
[0049] Examples of Group I intron self-splicing sequences include, but are not limited to, the self-splicing replacement intron-exon sequence derived from the T4 bacteriophage gene td or the Cyanobacterium Anabaena sp. pre-tRNA-Leu gene.
[0050] As used herein, a "spacer" is any contiguous nucleotide sequence that: 1) is predicted to avoid interference with proximal structure, e.g., from an IRES, coding or non-coding region, or intron; 2) is at least 7 nucleotides in length (and optionally no more than 100 nucleotides); 3) is located downstream of and adjacent to a 3' intron fragment and / or is located upstream of and adjacent to a 5' intron fragment; and / or 4) contains one or more of the following: a) an unstructured region at least 5 nt in length, b) a region at least 5 nt in length that is predicted to base-pair with distal (i.e., non-adjacent) sequence (including another spacer), and / or c) a structured region at least 1 nt in length bounded by a spacer sequence.
[0051] As used herein, "interfering" with respect to a sequence is a sequence that is predicted or empirically determined to be capable of altering the folding of other structures in RNA, e.g., an IRES or group I intron-derived sequence.
[0052] As used herein, "unstructured" with respect to RNA is an RNA sequence that is not predicted by RNAFold software or similar prediction tools to form structures (e.g., hairpin loops) with itself or with other sequences in the same RNA molecule.
[0053] As used herein, "structured" with respect to RNA is an RNA sequence that is predicted by RNAFold software or similar prediction tools to form structures (e.g., hairpin loops) with itself or with other sequences in the same RNA molecule.
[0054] In some embodiments, the length of the spacer sequence may be, for example, at least 10 nucleotides, at least 15 nucleotides, or at least 30 nucleotides. In some embodiments, the length of the spacer sequence is at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, or 30 nucleotides. In some embodiments, the length of the spacer sequence is 100, 90, 80, 70, 60, 50, 45, 40, 35, or 30 nucleotides or less. In some embodiments, the length of the spacer sequence is between 20 and 50 nucleotides. In some embodiments, the spacer sequence is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length.
[0055] The spacer sequence may be a polyA sequence, a polyA-C sequence, a polyC sequence, or a polyU sequence, or the spacer sequence may be specifically engineered according to the IRES. The spacer sequences described herein may have two functions: (1) to promote circularization and (2) to promote functionality by properly folding the intron and IRES. More specifically, the spacer sequences engineered herein have three priorities: 1) they are inert with respect to the folding of the proximal intron and IRES structures, 2) they provide sufficient separation between the intron and IRES secondary structures, and 3) they contain a region of spacer-spacer complementarity to promote the formation of a "splice bubble." In one embodiment, the vector is compatible with many possible IRES and coding or noncoding regions and two spacer sequences.
[0056] In one example, the vector may include a 5' spacer sequence but not a 3' spacer sequence. In another example, the vector may include a 3' spacer sequence but not a 5' spacer sequence. In another example, the vector may not include a 5' spacer sequence or a 3' spacer sequence. In another example, the vector does not include an IRES sequence. In a further example, the vector does not include an IRES sequence, a 5' spacer sequence, or a 3' spacer sequence.
[0057] As used herein, "vector" refers to a segment of DNA synthesized (e.g., using PCR) or harvested from a virus, a plasmid, or the cell of a higher organism, into which a foreign DNA segment can be, or has been, inserted for cloning and / or expression purposes; in some embodiments, the vector can be stably maintained in the organism. A vector may include an origin of replication, a selectable marker or reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). The term may also include linear DNA segments (e.g., PCR products, linearized plasmid fragments), plasmid vectors, viral vectors, cosmids, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc. In one example, the vectors provided herein include a multiple cloning site (MCS). In another example, the vectors provided herein do not include an MCS.
[0058] Unless otherwise specified, cells according to the present disclosure may include any cell into which an exogenous nucleic acid described herein can be introduced or expressed. It should be understood that the basic concepts of the present disclosure described herein are not limited by cell type. Cells according to the present disclosure may include somatic cells, stem cells, eukaryotic cells, prokaryotic cells, animal cells, plant cells, fungal cells, archaeal cells, eubacterial cells, etc. Cells may include eukaryotic cells, such as yeast cells, plant cells, and animal cells. Specific cells may include mammalian cells, such as human cells. Furthermore, cells may include any cell in which expression of a circRNA is beneficial or desirable.
[0059] The protein coding region of the gene of interest (GOI) may encode a protein of eukaryotic or prokaryotic origin. In some embodiments, the protein may be any protein of therapeutic or diagnostic use. For example, the protein coding region may encode a human protein or antibody. In some embodiments, the protein may be selected from, but is not limited to, a chimeric antigen receptor (CAR), a T cell receptor (TCR), an antibody, human factor IX (hFIX), lung-associated surfactant protein B (SP-B), vascular endothelial growth factor A (VEGF-A), human methylmalonyl-CoA mutase (hMUT), CF-transmembrane signaling regulator (CFTR), a cancer autoantigen, and another gene editing enzyme, such as a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein (e.g., Cas9 and Cpf1), a zinc finger nuclease (ZFN), and a transcription activator-like effector nuclease (TALEN). In some embodiments, the vector or circRNA lacks a protein-coding sequence. In some embodiments, the precursor RNA is a necessary intermediate between the plasmid and the circRNA.
[0060] In some embodiments, the vector may include an IRES sequence located at the 5' end of a protein-encoding gene of interest (GOI), such as Taura syndrome virus, Assassin bug virus, Theiler's encephalomyelitis virus, Simian virus 40, red fire ant virus 1, wheat aphid virus, reticuloendotheliosis virus, human poliovirus 1, German stink bug enteric virus, Kashmir wasp virus, human rhinovirus 2, leafhopper virus-1, human immunodeficiency virus type 1, leafhopper virus-1, Himetovirus P virus, Hepatitis C virus, Hepatitis A virus, Hepatitis G virus, Foot-and-mouth disease virus, Human enterovirus 71, Equine rhinitis virus, Pale spotted picorna-like virus, Encephalomyocarditis virus (EMCV), Drosophila C virus, Cruciferous tobamovirus, Cricket paralysis virus, Bovine viral diarrhea virus 1, Black queen brood virus, Aphid fatal paralysis virus, Avian encephalomyelitis virus, Bee acute paralysis enterovirus, Hibiscus chlorotic ringspot virus, Classical swine fever virus, Human FGF2, Human SFTPA1, Human AML1 / RUNX1, Drosophila antennapedia, Human AQP4, Human AT1R, Human BAG-1, Human BCL2, Human BiP, Human c-IAP1, Human c-myc, Human eIF4G, Mouse NDST4L, Human LEF1, Mouse HIF1 The IRES sequence may be selected from, but is not limited to, the following: α, human N-myc, mouse Gtx, human p27kip1, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, monkey virus, turnip crinkle virus, an aptamer for eIF4G, coxsackievirus B3 (CVB3), or coxsackievirus A (CVB1 / 2). Wild-type IRES sequences may be modified and still be effective in the present invention. In some embodiments, the length of the IRES sequence is about 50 nucleotides.
[0061] In some embodiments, the gene of interest (GOI) may be replaced by a nucleic acid encoding one or more RNA molecules, including antisense RNA, transfer RNA (tRNA), transfer-messenger RNA (tmRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), Cajal body-specific small RNA (scaRNA), guide RNA (gRNA), Y RNA, splicing leader sequence RNA (SL RNA), microRNA (miRNA), small interfering RNA (siRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long non-coding RNA (lncRNA), Piwi-interacting RNA (piRNA), short hairpin RNA (shRNA), trans-acting siRNA (tasiRNA), repeat-associated siRNA (rasiRNA), 7SK RNA (7SK), telomerase RNA component (TERC), vault RNA (vRNA, vtRNA), and enhancer RNA (eRNA). In some embodiments, a vector comprising a nucleic acid encoding one or more RNA molecules may or may not contain an IRES sequence.
[0062] Circular RNA may be purified by running the RNA through a reverse-phase chromatography column in a triethylammonium acetate (TEAA) buffer in a high-performance liquid chromatography (HPLC) system. In one example, the RNA may be run through a RP-HPLC chromatography column in a TEAA-acetonitrile buffer with a pH range of about 4-10 at a flow rate of about 0.01-5 mL / min.
[0063] In some examples, the present specification provides methods for generating precursor RNA by in vitro transcription using a vector provided herein (e.g., a vector provided herein having an RNA polymerase promoter located upstream of a 5' homology arm) as a template.
[0064] In some examples, the nucleotides, nucleosides, or chemically modified nucleotides or nucleosides used in the in vitro transcription reactions described herein may be at excess concentrations relative to the analogous nucleotide triphosphates. "Excess concentrations" are defined as concentrations higher than the concentration of the analogous nucleotide triphosphates for the purpose of altering the 5'-terminal nucleotide, specifically to prevent inclusion of a 5' triphosphate motif to reduce immunogenicity of circRNA preparations, or to include the necessary 5' monophosphate motif to enable enzymatic circularization of precursor molecules.
[0065] In some embodiments, the nucleotide used in excess may be guanosine monophosphate (GMP). In other embodiments, the nucleotide used in excess may be GDP, ADP, CDP, UDP, AMP, CMP, UMP, guanosine, adenosine, cytidine, uridine, or any chemically modified nucleotide or nucleoside. In some embodiments, the excess may be about a 10-fold excess. In some embodiments, the excess may be about a 12.5-fold excess.
[0066] In one example, the nucleotide, nucleoside, or chemically modified nucleotide or nucleoside may be used in an in vitro transcription reaction at a concentration that is at least about 10 times greater than the analogous nucleotide triphosphate.
[0067] In some examples, circRNAs produced from precursor RNAs synthesized in an in vitro transcription reaction in the presence of at least about 10-fold excess of nucleotides, nucleosides, or chemically modified nucleotides or nucleosides over the analogous nucleotide triphosphates are then purified by HPLC to achieve minimal immunogenicity.
[0068] The purity of the circular RNA product is believed to be crucial for expression because residual precursor linear RNA or nicked RNA may compete with the circular RNA for ribosome recruitment, leading to reduced expression efficiency. A method for producing highly pure circular RNA in vitro can help translate circular RNA technology into the development of RNA therapeutics and vaccines. Examples of the present disclosure provide a solution for the in vitro design and production of circular RNAs with high self-circularization efficiency.
[0069] Embodiments of the present disclosure may include methods for producing precursor RNAs with thermodynamically stable multidirectional junctions next to splicing bubbles to more efficiently produce self-circularizing RNAs in vitro.
[0070] The efficiency of precursor RNA self-circularization designed using a substituted intron exon can be highly dependent on the efficiency of self-splicing bubble formation. Conventional methods rely on homology arms at the 5' and 3' ends of the precursor RNA to form a double-stranded RNA structure, which then folds the splicing bubble into its secondary structure to achieve splicing function. Due to the respiratory dynamics of the RNA duplex, the self-circularization efficiency can generally be less than 60%.
[0071] To solve the problem of efficiently forming the secondary structure of the self-splicing bubble, embodiments of the present disclosure may include a stable RNA motif consisting of at least two stems and one loop to lock the splicing bubble into its secondary structure, thereby promoting its correct folding and splicing efficiency. For example, the stable RNA motif may include an asymmetric three-way junction (3WJ) structure to promote correct folding and splicing efficiency by locking the splicing bubble, e.g., the thermodynamically stable multi-way junction structure of RNA described in the packaging RNA of the phi29 DNA packaging motor (Shu et al., "Thermodynamically stable RNA three-way junction for constructing multifunctional nanoparticles for delivery of therapeutics." Nat Nanotechnol. 2011 Sep. 11; 6(10):658-67, the contents of which are incorporated herein by reference in their entirety), and the three fragments 3WJa, 3WJb, and 3WJc rapidly and automatically assemble into a three-way junction (3WJ) structure to lock folding (Binzel et al., "Mechanism of three-component collision to produce ultrastable pRNA three-way junction of Phi29 DNA-packaging motor by kinetic assessment." RNA. November 2016; 22(11):1710-1718, the contents of which are incorporated herein by reference in their entirety).
[0072] A multidirectional junction structure may be formed by cleaving from one stem-loop position and assigning sequences to the 5' and 3' ends of the precursor RNA sequence, respectively. Three-way junction (3WJ) domain
[0073] Packaging (or front) ribonucleic acid (pRNA) three-way junction (3WJ) motifs may be used in biotechnology, for example, to target human immunodeficiency virus (HIV) and cancer. For example, the phi29 bacteriophage pRNA 3WJ nanomotif has been successfully used as a building block in the rational design of nanostructures with, for example, cancer-targeting functions. (U.S. Patent No. 9,297,013, the contents of which are incorporated herein by reference in their entirety.)
[0074] As used herein, the term "three-way junction" ("3WJ") or "tri-branched" scaffold (or domain) refers to a structure assembled from three RNA sequences. Figure 1 shows that a 3WJ domain (10) can be constructed from three (5'→3') RNA strands (termed 3WJa, 3WJb, and 3WJc), which are base-paired to each other. A first (5'→3') RNA oligonucleotide sequence designated as 3WJa, a second (5'→3') RNA oligonucleotide sequence designated as 3WJb, and a third (5'→3') RNA oligonucleotide sequence designated as 3WJc may be combined and base-paired to form a tri-branched 3WJ domain (10), wherein a first branch (11) of the 3WJ domain is formed from a 5' portion of the 3WJa sequence and a 3' portion of the 3WJc sequence, a second branch (12) of the 3WJ domain is formed from a 3' portion of the 3WJa sequence and a 5' portion of the 3WJb sequence, and a third branch (13) of the 3WJ domain is formed from a 3' portion of the 3WJb sequence and a 5' portion of the 3WJc sequence, wherein the first (11), second (12), and third (13) branches of the 3WJ domain are formed from a 3' portion of the 3WJb sequence and a 5' portion of the 3WJc sequence. Each of the branches may include a helical region with multiple RNA nucleotide pairs that form standard Watson-Crick bonds. One, two, and / or three branches of the 3WJ of the present disclosure may further include non-Watson-Crick nucleotide pairs or bulges, such as, but not limited to, GU wobble base pairs, or bulges with a small number of additional unpaired bases. In some embodiments, the 3' end of the 3WJb may be connected to the 5' end of the 3WJc via a linker sequence (14). In some embodiments, the linker sequence (14) connecting the 3WJb and 3WJc may not pair at all, for example, forming a stem-loop structure. In some embodiments, the linker sequence (14) may partially pair with the sequence, for example, forming a loop and a stem. In some embodiments, the linker sequence (14) may fully pair with the sequence, for example, forming a stem. In some embodiments, the linker sequence (14) may be absent, for example, the 3' end of 3WJb may be directly linked to the 5' end of 3WJc.
[0075] As used herein, a stem-loop structure can occur in single-stranded RNA. This structure may be referred to as a hairpin or hairpin loop, and the continuous sequence may include a stem and a (terminal) loop, where the stem may be formed by two adjacent completely or partially reversed complementary sequences, with the complementary sequences separated by a short sequence serving as a spacer to form a loop stem-loop structure. Two adjacent completely or partially reversed complementary sequences may be defined, for example, as elements of sequence 1 and sequence 2 of the stem-loop structure. When these two adjacent completely or partially reversed complementary sequences (e.g., sequence 1 and sequence 2 of the stem-loop structural element) base pair with each other, a stem-loop structure can be formed, leading to the formation of a double-stranded nucleic acid sequence, which contains an unpaired loop at its end formed from a short sequence located between elements of sequence 1 and sequence 2 of the stem-loop structure in the continuous sequence. Therefore, an unpaired loop is generally a nucleic acid region that cannot pair with any of these elements of the stem-loop structure. The resulting multidirectional junction domain folds into a compact structure with low Gibbs free energy (ΔG), thereby increasing the thermodynamic stability of the structure and promoting the formation of a true splicing bubble. The stability of the pairing element of the stem-loop structure is determined by the length-GC contrast ratio, the number of mismatches or loops contained therein (generally, a small number of mismatches is tolerated, especially in large double-stranded regions), and the base composition of the pairing region. In some embodiments of the present disclosure, loop lengths of 3 to 15 bases are tolerated, but more preferred loop lengths are 3 to 20 bases, preferably 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, or more preferably 4 to 5 bases, most preferably 4 bases. The length of the stem sequence forming the double-stranded structure may be 5 to 20 bases, preferably 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 bases.
[0076] The Gibbs free energy (ΔG) for the formation of each RNA scaffold was calculated using the M-fold (Mfold) calculation for the RNA folding format available from the UNAFold web server (Mfold web server for nucleic acid folding and hybridization prediction. Nucleic Acids Res. 31 (13), 3406-15, (2003)). ΔG was calculated based on RNA folding at 37°C and 1 M NaCl ion conditions.
[0077] The ΔG of the multi-way junction of the present disclosure is about −200 kcal / mol to about −5 kcal / mol, about −195 kcal / mol to about −6 kcal / mol, about −190 kcal / mol to about −7 kcal / mol, about −190 kcal / mol to about −8 kcal / mol, about −190 kcal / mol to about −9 kcal / mol, about −189 kcal / mol to about −9.5 kcal / mol, about −188.7 kcal / mol to about −9.8 ... / mol~about -27.1kcal / mol, about -188.7kcal / mol~about -35.8kcal / mol, about -188.7kcal / mol~about -143.7kcal / mol, about -188.7kcal / mol~about -62.7kcal / mol, about -188.7kca l / mol ~ approx. -35.4kcal / mol, approx. -188.7kcal / mol ~ approx. -31.9kcal / mol, approx. -188.7kcal / mol ~ approx. -31.7kcal / mol, approx. -188.7kcal / mol ~ approx. -31.4kcal / mol, approx. -188.7kca l / mol ~ approx. -32.5kcal / mol, approx. -188.7kcal / mol ~ approx. -32.3kcal / mol, approx. -188.7kcal / mol ~ approx. -32.0kcal / mol, approx. -188.7kcal / mol ~ approx. -29.2kcal / mol, approx. -188.7kca l / mol ~ approx. -25.9kcal / mol, approx. -188.7kcal / mol ~ approx. -22.6kcal / mol, approx. -188.7kcal / mol ~ approx. -21.7kcal / mol, approx. -188.7kcal / mol ~ approx. -17.8kcal / mol, approx. -188.7kca l / mol ~ approx. -15.0kcal / mol, approx. -188.7kcal / mol ~ approx. -13.5kcal / mol, approx. -188.7kcal / mol ~ approx. -13.3kcal / mol, approx. -188.7kcal / mol ~ approx. -10.10kcal / mol, approx. -188.7kc al / mol~about -32.0kcal / mol, about -188.7kcal / mol~about -28.10kcal / mol, about -188.7kcal / mol~about -29.30kcal / mol, about -188.7kcal / mol~about -26.70kcal / mol, about -188.7kcal / mol ~ approx. -23.40kcal / mol, approx. -188.7kcal / mol ~ approx. -20.10kcal / mol, approx. -188.7kcal / mol ~ approx. -19.20kca l / mol, about -188.7kcal / mol~about -18.20kcal / mol, about -188.7kcal / mol~about -13.10kcal / mol, about -188.7kcal / mol l~about -10.15kcal / mol, about -188.7kcal / mol~about -13.0kcal / mol, about -188.7kcal / mol~about -16.0kcal / mol, about -18 8.7kcal / mol~Approx.-12.7kcal / mol, Approx.-188.7kcal / mol~Approx.-23.65kcal / mol, Approx.-188.7kcal / mol~Approx.-19.40kc al / mol, about −188.7 kcal / mol to about −15.40 kcal / mol, about −188.7 kcal / mol to about −34.40 kcal / mol, about −188.7 kcal / mol to about −25.30 kcal / mol, about −188.7 kcal / mol to about −16.40 kcal / mol, about −188.7 kcal / mol to about −12.60 kcal / mol, about −188.7 kcal / mol to about −30.50 kcal / mol, about −188.7 kcal / mol to about −21.40 kcal / mol, about −188.7 kcal / mol to about −34.30 kcal / mol, about −188.7 kcal / mol to about −25.20 kcal / mol, or about −188.7 kcal / mol to about −16.30 kcal / mol. .
[0078] In some non-limiting examples, each of the 3WJa, 3WJb, and 3WJc oligonucleotide sequences of a 3WJ scaffold or domain of the present disclosure can independently comprise 8 to 36 nucleotides (e.g., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, or 36 nucleotides), excluding an RNA linker or an RNA portion attached to a biologically active portion of the 3WJ scaffold.
[0079] A four-way junction (4WJ) domain may be constructed from four (5'→3') RNA strands (designated 4WJa, 4WJb, 4WJc, and 4WJd), which are base-paired to each other. A first (5'→3') RNA oligonucleotide sequence designated as 4Wja, a second (5'→3') RNA oligonucleotide sequence designated as 4WJb, a third (5'→3') RNA oligonucleotide sequence designated as 4WJc, and a fourth (5'→3') RNA oligonucleotide sequence designated as 4WJd may be combined and base-paired to form a 4WJ domain. For example, Figure 6 shows that the first branch (first BR) of the 4WJ domain is formed from the 5' portion of the 4Wja sequence and the 3' portion of the 4WJd sequence, the second branch (second BR) of the 4WJ domain is formed from the 3' portion of the 4Wja sequence and the 5' portion of the 4WJb sequence, the third branch (third BR) of the 4WJ domain is formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence, and the fourth branch (fourth BR) of the 4WJ domain is formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence, where each of the first, second, third, and fourth branches may contain a helical region having multiple RNA nucleotide pairs that form standard Watson-Crick bonds. One, two, three, and / or four branches of each 4WJ of the present disclosure may further contain non-Watson-Crick nucleotide pairs, such as, but not limited to, GU. In some embodiments, the 3' end of 4Wja may be linked to the 5' end of 4WJb via a linker sequence, and the 3' end of 4WJc may be linked to the 5' end of 4WJd via a linker sequence. In some embodiments, one or more linker sequences connecting 4Wja and 4WJb and the linker sequence connecting 4WJc and 4WJd may not pair at all, for example, they may form a stem-loop structure. In some embodiments, one or more linker sequences may be partially paired with the sequence, for example, they may form a loop and a stem. In some embodiments, one or more linker sequences may be completely paired with the sequence, for example, they may form a stem.In some embodiments, one or more linker sequences may be absent, for example, the 3' end of 4Wja may be directly linked to the 5' end of 4WJb, and the 3' end of 4WJc may be directly linked to the 5' end of 4WJd.
[0080] A five-way junction (5WJ) domain may be constructed from five (5'→3') RNA strands (designated 5WJa, 5WJb, 5WJc, 5WJd, and 5WJe), which are base-paired to one another. A first (5'→3') RNA oligonucleotide sequence designated as 5WJa, a second (5'→3') RNA oligonucleotide sequence designated as 5WJb, a third (5'→3') RNA oligonucleotide sequence designated as 5WJc, a fourth (5'→3') RNA oligonucleotide sequence designated as 5WJd, and a fifth (5'→3') RNA oligonucleotide sequence designated as 5WJe may be combined and base-paired to form a 5WJ domain. For example, Figure 9 shows that the first branch (first BR) of the 5WJ domain is formed from the 5' portion of the 5WJa sequence and the 3' portion of the 5WJe sequence, the second branch (second BR) of the 5WJ domain is formed from the 3' portion of the 5WJa sequence and the 5' portion of the 5WJb sequence, the third branch (third BR) of the 5WJ domain is formed from the 3' portion of the 5WJb sequence and the 5' portion of the 5WJc sequence, the fourth branch (fourth BR) of the 5WJ domain is formed from the 3' portion of the 5WJc sequence and the 5' portion of the 5WJd sequence, and the fifth branch (fifth BR) of the 5WJ domain is formed from the 3' portion of the 5WJd sequence and the 5' portion of the 5WJe sequence, wherein each of the first, second, third, fourth, and fifth branches may include a helical region having multiple RNA nucleotide pairs that form canonical Watson-Crick bonds. One, two, three, four, and / or five branches of each 5WJ disclosed herein may further comprise a non-Watson-Crick nucleotide pair, such as, but not limited to, GU. In some examples, the 3' end of 5WJa may be linked to the 5' end of 5WJb via a linker sequence, the 3' end of 5WJc may be linked to the 5' end of 5WJd via a linker sequence, and the 3' end of 5WJd may be linked to the 5' end of 5WJe via a linker sequence. In some examples, one or more linker sequences linking 5WJa and 5WJb, linker sequences linking 5WJc and 5WJd, and linker sequences linking 5WJd and 5WJe may not be paired at all, and may form, for example, a stem-loop structure.In some embodiments, one or more linker sequences may be partially paired with the sequence, e.g., forming a loop and a stem. In some embodiments, one or more linker sequences may be fully paired with the sequence, e.g., forming a stem. In some embodiments, one or more linker sequences may be absent, e.g., the 3' end of 5WJa may be directly linked to the 5' end of 5WJb, the 3' end of 5WJc may be directly linked to the 5' end of 5WJd, and / or the 3' end of 5WJd may be directly linked to the 5' end of 5WJe.
[0081] In some embodiments, each multidirectional junction (e.g., 3WJ, 4WJ, and 5WJ) disclosed herein may contain a core structure. The duplex formed by each multidirectional junction may not affect the formation of the core structure (Khisamutdinov et al., "Enhancing immunomodulation on innate immunity by shape transition among RNA triangle, square, and pentagon nanovehicles." Nucleic Acids Res. 2014 Nov. 1; 42(15):9996-10004; the contents of which are incorporated herein by reference in their entirety). Changing the nucleotide (N) while forming a base-paired duplex with the corresponding arm may not affect the formation of the junction. For example, a 3WJ can form a 60-degree angle between the two arms, and a thermodynamically stable 3WJ structure can be achieved by designing the duplex sequences on both sides to different lengths to form a triangular, square, or pentagonal junction structure with good flexibility.
[0082] Table A shows the core structure sequences of 3WJ, 4WJ and 5WJ according to some embodiments of the present disclosure.
[0083] [Table 1-1] [Table 1-2] [Table 1-3]
[0084] Table 1 shows examples of 3WJ, 4WJ, and 5WJ sequences according to some embodiments of the present disclosure. The bolded sequences can be used to design 3WJ, 4WJ, and 5WJ motifs, forming a core structure with duplex arms that enhance their thermodynamic stability. The three fragments assemble into the 3WJ structure via a two-step reaction mechanism [ref1] with remarkable speed and affinity. For example, Figures 2 and 3 show the 3WJ motif formed from SEQ ID NOs: 1-3 (construct #1, ΔG = -27.3 Kcal / mol) and SEQ ID NOs: 4-6 (construct #2, ΔG = -35.8 Kcal / mol), respectively. Figures 6 and 7 show the 4WJ motif formed from SEQ ID NOs: 7-9 and 102 (construct #3, ΔG = -143.7 Kcal / mol) and the 4WJ motif formed from SEQ ID NOs: 103-106 (construct #57, ΔG = -62.7 Kcal / mol). Figure 9 shows the 5WJ motif formed from SEQ ID NOs: 107-111 (construct #58, ΔG = -188.7 Kcal / mol).
[0085] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] Example 1 3WJ-Circular RNA Production
[0086] As shown in Figure 4, by using 3WJa at the 5' end of the precursor circular RNA sequence and 3WJb and 3WJc at the 3' end, a thermodynamically stable 3WJ (ΔG = -27.7 kcal / mol) was formed, and self-splicing reactions were performed to efficiently generate circular RNAs with exon scars (exon scars) within the circRNA. An internal ribosome entry site (IRES) (e.g., from encephalomyocarditis virus (EMCV)) and a gene of interest (GOI) (e.g., from enhanced green fluorescent protein (eGFP)) can be inserted. Two short regions on either side correspond to the exon fragments (E1 and E2) of a replaced intron-exon (PIE) construct between the 3' and 5' introns (I1 and I2) of the replaced group I catalytic intron, e.g., the thymidylate synthase (Td) gene from T4 bacteriophage. Similarly, as shown in FIG. 5, thermodynamically stable 3WJ can be formed by using 3WJa and 3WJb at the 5′ end and 3WJc at the 3′ end.
[0087] The precursor RNA is synthesized by in vitro transcription using plasmid DNA linearized after the engineered 3' end of the precursor RNA under conditions known in the art, and then heated in the presence of magnesium ions and GTP to promote self-circularization before or after purification of the crude IVT reaction product. During self-circularization, the first branch of the 3WJ domain (first BR) may be formed from the 5' portion of the 3WJa sequence and the 3' portion of the 3WJc sequence, the second branch of the 3WJ domain (second BR) may be formed from the 3' portion of the 3WJa sequence and the 5' portion of the 3WJb sequence, and the third branch of the 3WJ domain (third BR) may be formed from the 3' portion of the 3WJb sequence and the 5' portion of the 3WJc sequence. Thus, the 3WJ domain can lock the splicing bubble to promote correct folding and splicing efficiency. In some embodiments, the third BR may be formed before self-circularization. For example, the 3' portion of the 3WJb sequence and the 5' portion of the 3WJc sequence may base pair to form a third BR before circularization occurs. In some embodiments, a third branch may form during self-circularization. Each branch may contain multiple RNA nucleotide pairs that form standard Watson-Crick bonds. During splicing, the 3' hydroxyl group of the guanosine nucleotide participates in a transesterification reaction at the 5' splice site. This half of the 5' intron (I1) is excised, and the free hydroxyl group at the end of the intermediate participates in a second transesterification reaction at the 3' splice site, circularizing the intermediate region and excising the 3' intron (I2) and the 3WJ domain together. Production of 4WJ-circular RNA
[0088] As shown in Figure 8, by using 4WJa and 4WJb at the 5' end of the precursor circular RNA sequence and 4WJc and 4WJd at the 3' end, the thermodynamically stable 4WJ (construct #57, ΔG = -62.7 kcal / mol) was formed, and a self-splicing reaction was carried out to efficiently generate circular RNAs with exon vestiges within the circRNA. An internal ribosome entry site (IRES) (e.g., from encephalomyocarditis virus (EMCV)) and a gene of interest (GOI) (e.g., from enhanced green fluorescent protein (eGFP)) can be inserted, and two short regions on either side correspond to the exon fragments (E1 and E2) of a replaced intron-exon (PIE) construct between the 3' and 5' introns (I1 and I2) of the replaced group I catalytic intron, e.g., the thymidylate synthase (Td) gene from T4 bacteriophage. Production of 5WJ-circular RNA
[0089] As shown in Figure 10, by using 5WJa, 5WJb, and 5WJc at the 5' end of the precursor circular RNA sequence and 5WJd and 5WJe at the 3' end, the thermodynamically stable 5WJ (construct #58, ΔG = -188.7 kcal / mol) was formed, and a self-splicing reaction was carried out to efficiently generate circular RNAs with exon vestiges within the circRNA. An internal ribosome entry site (IRES) (e.g., from encephalomyocarditis virus (EMCV)) and a gene of interest (GOI) (e.g., from enhanced green fluorescent protein (eGFP)) can be inserted, and two short regions on either side correspond to the exon fragments (E1 and E2) of a replaced intron-exon (PIE) construct between the 3' and 5' introns (I1 and I2) of the replaced group I catalytic intron, e.g., the thymidylate synthase (Td) gene from T4 bacteriophage. Similarly, as shown in FIG. 11, thermodynamically stable 5WJ can be formed by using 5WJa and 5WJb at the 5′ end and 5WJc, 5WJd, and 5WJe at the 3′ end. Example 2 Improving RNA circularization efficiency by 3WJ-circular RNA design
[0090] To compare the RNA circularization efficiency of the 3WJ-circular RNA design with that of conventional homologous RNA designs, precursor RNAs synthesized by in vitro transcription from a vector encoding eGFP, which had a 3WJa element at the 5' end and a 3WJb / 3WJc element at the 3' end (Figure 4), and precursor RNAs synthesized by in vitro transcription from a vector encoding eGFP, which had a 5' homology arm at the 5' end and a 3' homology arm at the 3' end, were analyzed by gel electrophoresis.
[0091] As shown in Figure 12, crude circular RNAs (lanes 1 and 3) formed from conventional homologous RNA designs without chromatography column purification (3WJ-) produced approximately 30% unreacted precursor RNA, as indicated by the clear middle band, and approximately 50% circular RNA, as indicated by the upper band. In contrast, circular RNAs (lane 5) formed from 3WJ- circular RNA designs without chromatography column purification (3WJ+) produced approximately 80% circularized RNA, as indicated by the upper band, and approximately 5% unreacted precursor RNA and approximately 15% nicked RNA, as indicated by the middle and lower bands, respectively. Furthermore, ion-pairing reverse-phase HPLC was more effective in purifying 3WJ- circular RNA (lane 6), which produced less unreacted precursor RNA, than homologous circular RNAs (lanes 2 and 4), which produced more unreacted precursor RNA. As shown in Figure 13, the circRNA / precursor RNA ratio produced from the 3WJ-circular RNA design (3WJ+) was higher than that of the conventional homologous circular RNA design (3WJ-), as measured by a bioanalyzer. The circRNAs produced from the 3WJ-circular RNA design were easily purified by reverse-phase HPLC. To compare the purity of circRNAs produced with or without 3WJ, we generated circular RNAs encoding eGFP using homologous arms (e.g., without 3WJ (3WJ(-))), and fractionated the crude product using reverse-phase HPLC (Figure 14A) and analyzed it by agarose gel electrophoresis (Figure 14B). In contrast, we generated circular RNAs encoding eGFP using 3WJ (3WJ(+)), and fractionated the crude product using reverse-phase HPLC (Figure 15A) and analyzed it by agarose gel electrophoresis (Figure 15B). As shown in these results, when produced using 3WJ, the abundance of precursor RNA (third peak, approximately 17.3 min) was reduced and the purity of the circular RNA portion was improved (see parts 3 and 4). These results indicated that the 3WJ-circular RNA design improved the efficiency of RNA circularization and purification compared to the conventional homology arm design.
[0092] To compare the purity of circRNAs produced with or without 3WJ and purified with or without HPLC, we generated circular RNAs encoding Gaussian luciferase (G-Luc) using homologous arms (e.g., without 3WJ) and fractionated the crude product using reverse-phase HPLC and analyzed it by agarose gel electrophoresis. We also generated circular RNAs encoding G-Luc using 3WJ, fractionated the crude product using reverse-phase HPLC, and analyzed it by agarose gel electrophoresis. As shown in Figure 20, the purity of the final crude circular RNAs encoding the G-Luc sequence using 3WJ was 69%-79% crude (3WJ + and purified -) (lane 3), and 81%-91% after HPLC purification (3WJ + and purified +) (lane 4). In contrast, the purity of the final crude circular RNA encoding the G-Luc sequence without 3WJ was 55%-67% crude (3WJ- and purified-) (lane 1), and 72%-83% after HPLC purification (3WJ- and purified+) (lane 2). These results demonstrate that the use of multiple ligations (e.g., 3WJ) can produce circular RNA with higher purity than the use of multiple ligations (e.g., conventional homology arm methods). Highly purified circular RNA could enable the therapeutic use of such circular RNA. Example 3 3WJ-circular RNA design improves circular RNA expression
[0093] To test whether the improved RNA circularization efficiency of the 3WJ-circular RNA design is associated with enhanced gene expression, we transfected the eGFP-encoding vector used in Figure 12 into A549 cells and performed eGFP expression analysis using a fluorescence microplate reader. As shown in Figure 16, the more highly purified circular RNA produced from the 3WJ-circular RNA design in A549 cells (3WJ+, purified+) exhibited superior circular RNA expression levels, as indicated by higher relative fluorescence units (RFU), compared with circular RNAs with conventional homology arm designs (3WJ-, purified+). Similarly, we transfected the eGFP-encoding vector used in Figure 12 into 293 cells and performed eGFP expression analysis using a fluorescence microplate reader. As shown in Figure 17, the more highly purified circular RNA produced from the 3WJ-circular RNA design in 293 cells (3WJ+, purified+) exhibited superior circular RNA expression levels, as indicated by higher relative fluorescence units (RFU), compared with circular RNAs with conventional homology arm designs (3WJ-, purified+). Mock transfection (Mock) was used as a control. To compare the time course of protein expression levels, A549 cells were transfected with eGFP-encoding circRNA or eGFP-encoding mRNA produced by - / +3WJ. After transfection, fluorescence was analyzed by flow cytometry for up to 6 days (10,000 cells were analyzed under each condition for each time point). As shown in Figures 18 and 19, the circRNA produced by 3WJ (3WJ+) had superior expression to those produced by the homologous arm (3WJ-) and mRNA, consistent with the HPLC-purified samples (Figures 14A, 14B, 15A, and 15B). Example 4 RNA circularization using 3WJ and 4WJ
[0094] [Table 3]
[0095] We compared the circularization of RNA containing the Coxsackievirus B3 (CVB3) IRES and eGFP sequence with that containing the CVB3 IRES and eSpCas9 nuclease (a mutant form of Cas9 nuclease) sequence using (1) 3WJ (Table 3, SEQ ID NOs: 119-121), (2) extended 3WJ (Ex3WJ, ΔG = -53.5 Kcal / mol) (Table 2), or (3) 4WJ (construct 57, ΔG = -62.7 Kcal / mol). Circularization was confirmed using agarose gel electrophoresis and densitometry analysis (Figure 21A) or HPLC (Figures 21B and 21C). As shown in these results, RNAs with (+)3WJ, Ex3WJ, or 4WJ produced higher RNA circularization than RNAs without (-)3WJ, Ex3WJ, or 4WJ. Time-dependent expression of Gaussian luciferase (Gluc) encoded by circular RNA
[0096] After transfection of A549 cells with Lipofectamine MessengerMAX reagent, we compared the expression of Gluc-encoding circRNA (unmodified) produced by the construct harboring 3WJ (Table 3, SEQ ID NO: 119-121, ΔG = -27.7 kcal / mol) with linear mRNA (generated by in vitro transcription using Cap1 (ribomethylation of the nucleotide adjacent to m7G) and 100% N1-methyl-pseudo-UTP modified). Luminescence was analyzed by a microplate reader (20 μL of medium was analyzed under each condition for each time point (Figure 22A) and normalized to well volume) (Figure 22B). As shown in these results, at 3 days post-transfection, Gluc expression produced by cells transfected with circRNA was approximately 4.5-fold higher than that of cells transfected with linear mRNA. Time-dependent expression of eGFP encoded by circular RNA
[0097] After transfection of A549 cells with Lipofectamine MessengerMAX reagent, we compared the expression of eGFP-encoding circRNA (unmodified) produced by the construct with 3WJ (Table 3, SEQ ID NOs: 119-121) with linear mRNA (Cap1, 100% N1-methyl-pseudo-UTP modified). After transfection, fluorescence was analyzed by flow cytometry (10,000 cells were analyzed under each condition for each time point (Figure 23A) and normalized to well volume (Figure 23B)). As shown in these results, at 3 days post-transfection, eGFP expression produced by cells transfected with circRNA was approximately 3.5-fold higher than that of cells transfected with linear mRNA.
[0098] After transfection with Lipofectamine MessengerMAX reagent, we compared the abundance of eGFP-encoding circRNA (unmodified) and linear mRNA (Cap1, 100% N1-methyl-pseudo-UTP modified) in A549 cells. RNA copy numbers were measured by quantitative PCR (using the standard curve method to analyze cDNA synthesized from extracted cellular RNA (Figure 24A) and normalized to well volume (Figure 24B)). As shown in these results, at 3 days post-transfection, the RNA copy number produced by cells transfected with circRNA is approximately four-fold higher than that of cells transfected with linear mRNA.
[0099] [Table 4]
[0100] The present invention can be defined by the following aspects. [Embodiment 1] A method for producing a circular RNA, the method comprising: transcribing a vector to form a precursor RNA, wherein the vector contains the following elements operably linked to each other and arranged in the following order: a) a 5' element that does not contain or that contains at least one stem-loop structure-forming sequence; b) a 3' Group I self-splicing intron fragment containing a 3' splice site dinucleotide; c) protein coding or non-coding regions; d) a 5' Group I self-splicing intron fragment containing a 5' splice site dinucleotide; e) comprises a 3' element that does not contain or contains at least one stem-loop structure-forming sequence, provided that if the 5' element does not contain a stem-loop structure, the 3' element contains at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element contains at least one stem-loop structure; wherein the 5' element and the 3' element form a thermodynamically stable multidirectional junction; wherein the precursor RNA is translatable within the cell and / or capable of forming a biologically active circular RNA. [Embodiment 2] The method according to embodiment 1, wherein the thermodynamically stable multi-way junction is a three-way junction (3WJ), a four-way junction (4WJ), or a five-way junction (5WJ). [Aspect 3] The 3WJ is a first branch of the 3WJ domain formed from the 5' portion of the 3WJa sequence and the 3' portion of the 3WJc sequence and including a first helical region; a second branch of the 3WJ domain formed from the 3' portion of the 3WJa sequence and the 5' portion of the 3WJb sequence and including a second helical region; a third branch of the 3WJ domain formed from the 3' portion of the 3WJb sequence and the 5' portion of the 3WJc sequence and including a third helical region; 3. The method of embodiment 2, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form canonical Watson-Crick bonds. [Aspect 4] 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 4, 3WJb comprises or consists of SEQ ID NO: 5, and 3WJc comprises or consists of SEQ ID NO: 6, or 3WJa comprises or consists of SEQ ID NO: 10, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 12, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 13, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 17, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 20, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 22, or 3WJa comprises or consists of SEQ ID NO: 23, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 26, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 28, or 3WJa comprises or consists of SEQ ID NO: 29, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 32, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 33, and 3WJc comprises or consists of SEQ ID NO: 34, or 3WJa comprises or consists of SEQ ID NO: 41, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 44, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 50, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 53, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 56, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 58, or 3WJa comprises or consists of SEQ ID NO: 59, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 60, or 3WJa comprises or consists of SEQ ID NO: 61, 3WJb comprises or consists of SEQ ID NO: 63, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of SEQ ID NO: 66, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of UGUCACGGG, and 3WJc comprises or consists of SEQ ID NO: 68, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 69, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 70, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 71, and 3WJc comprises or consists of SEQ ID NO: 72, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 8, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 78, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 81, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 84, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 7, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 90, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 93, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 95, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 112, 3WJb comprises or consists of SEQ ID NO: 113, and 3WJc comprises or consists of SEQ ID NO: 114, or 3WJa comprises or consists of SEQ ID NO: 119, 3WJb comprises or consists of SEQ ID NO: 120, and 3WJc comprises or consists of SEQ ID NO: 121, or 3WJa includes AUGUGUA, 3WJb includes UACUUUG, and 3WJc includes AUCAUG, or 3WJa contains GCGUU, 3WJb contains UUCGC, and 3WJc contains GCCAUAGCG, or 3WJa includes GUAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUAC, or 3WJa includes AUGGUA, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGUGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUGG, 3WJb includes UCACGG, and 3WJc includes CCUCUUA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUG, 3WJb includes CAGGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UAUGU, 3WJb includes GCAGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UGUGU, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or The method of embodiment 3, wherein 3WJa comprises UGUGU, 3WJb comprises ACUUU, and 3WJc comprises AAAUCA. [Aspect 5] The 4WJ is a first branch of the 4WJ domain formed from the 5' portion of the 4WJa sequence and the 3' portion of the 4WJd sequence and including a first helical region; a second branch of the 4WJ domain formed from the 3' portion of the 4WJa sequence and the 5' portion of the 4WJb sequence and including a second helical region; a third branch of the 4WJ domain formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence and including a third helical region; a fourth branch of the 4WJ domain formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence and including a fourth helical region; 3. The method of embodiment 2, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form canonical Watson-Crick bonds. [Aspect 6] 4WJa comprises or consists of SEQ ID NO: 7, 4WJb comprises or consists of SEQ ID NO: 8, 4WJc comprises or consists of SEQ ID NO: 9, and 4WJd comprises or consists of SEQ ID NO: 102, or 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO: 106, or 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO: 118, or 4WJa comprises UGCAGGUG, 4WJb comprises ACGGGC, 4WJc comprises CCAGCA, and 4WJd comprises SEQ ID NO: 67, or 4WJa comprises SEQ ID NO: 74, 4WJb comprises AACUG, 4WJc comprises SEQ ID NO: 75, and 4WJd comprises AUCAUG, or The method of embodiment 5, wherein 4WJa comprises SEQ ID NO: 122, 4WJb comprises GAACU, 4WJc comprises SEQ ID NO: 123, and 4WJd comprises AAUCA. [Aspect 7] The 5WJ is a first branch of the 5WJ domain formed from the 5' portion of the 5WJa sequence and the 3' portion of the 5WJe sequence and including a first helical region; a second branch of the 5WJ domain formed from the 3' portion of the 5WJa sequence and the 5' portion of the 5WJb sequence and including a second helical region; a third branch of the 5WJ domain formed from the 3' portion of the 5WJb sequence and the 5' portion of the 5WJc sequence and including a third helical region; a fourth branch of the 5WJ domain formed from the 3' portion of the 5WJc sequence and the 5' portion of the 5WJd sequence and including a fourth helical region; a fifth branch of the 5WJ domain formed from the 3' portion of the 5WJd sequence and the 5' portion of the 5WJe sequence and including a fifth helical region; 3. The method of embodiment 2, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form canonical Watson-Crick bonds. [Aspect 8] 5WJa comprises or consists of SEQ ID NO: 107, 5WJb comprises or consists of SEQ ID NO: 108, 5WJc comprises or consists of SEQ ID NO: 109, 5WJd comprises or consists of SEQ ID NO: 110, and 5WJe comprises or consists of SEQ ID NO: 111, or The method of embodiment 5, wherein 5WJa comprises GUGA, 5WJb comprises UUGC, 5WJc comprises GUGU, 5WJd comprises AUGC, and 5WJe comprises GUGC. [Embodiment 9] The vector further comprises an internal ribosome entry site (IRES) located at the 5' end of c), wherein the IRES is an IRES that is capable of binding to any of the following viruses: Taura syndrome virus, Assassin bug virus, Theiler's encephalomyelitis virus, Simian virus 40, red fire ant virus 1, Rhopalococcus aphid virus, reticuloendotheliosis virus, human poliovirus 1, German winged stink bug enteric virus, Kashmir wasp virus, human rhinovirus 2, leafhopper virus-1, human immunodeficiency virus type 1, leafhopper virus-1, Himetovirus P virus, Hepatitis C virus, Hepatitis A virus, Hepatitis G virus, Foot-and-mouth disease virus, Human enterovirus 71, Equine rhinitis virus, Pale spotted picorna-like virus, Encephalomyocarditis virus (EMCV), Drosophila C virus, Crucifer tobamovirus, Cricket paralysis virus, Bovine viral diarrhea virus 1, Black queen brood virus, Aphid fatal paralysis virus, Avian encephalomyelitis virus, Bee acute paralysis enterovirus, Hibiscus chlorotic ringspot virus, Classical swine fever virus, Human fibroblast growth factor 2 (HGF) FGF2), human surfactant drug protein A1 (SFTPA1), human acute myeloid leukemia protein 1 / runt-related transcription factor 1 (AML1 / RUNX1), Drosophila antennapedia, human aquaporin-4 (AQP4), human angiotensin II receptor type 1 (AT1R), human BCL2-associated immortality gene 1 (BAG-1), human B-cell lymphoma 2 (BCL2), human immunoglobulin-binding protein (BiP), human inhibitor of apoptosis family protein 1 (c-IAP1), human c-myc, human eukaryotic translation initiation factor 4 G (eIF4G), mouse N-deacetylase and N-sulfotransferase 4 (NDST4L), human lymphoid enhancer-binding factor-1 (LEF1), mouse hypoxia-inducible factor 1 subunit alpha (HIF1α), human N-myc, mouse glial cell and testis-specific homeobox protein (Gtx), human cyclin-dependent kinase inhibitor 1B (p27kip1), human platelet-derived growth factor B / simian sarcoma virus human homolog (PDGF2 / c-sis), human p53, human proviral integration site-1 of Moloney murine leukemia virus (Pim-1), mouse RNA-binding protein white 3 (Rbm3), Drosophila reaper, canine Scamper, Drosophila superbreast gene (Ubx), salivirus, coronavirus, parechovirus, human N-ras upstream (UNR), mouse dystrophin-related protein A (UtrA), human vascular endothelial growth factor A (VEGF-A), human X-linked apoptosis protein (XIAP), Drosophila hairless, budding yeast transcription factor IID (TFIID), budding yeast Yes1-related transcription factor (YAP1), human proto-oncogene tyrosine-linked protein 9. The method of any one of aspects 1 to 8, wherein the IRES sequence is selected from an IRES sequence from a virus or gene selected from the group consisting of kinase Src (c-src), human fibroblast growth factor 1 (FGF-1), monkey virus, turnip crinkle virus, coxsackievirus B3 (CVB3), and coxsackievirus A (CVB1 / 2). [Embodiment 10] The method according to any one of embodiments 1 to 9, wherein the vector further comprises an RNA polymerase promoter. [Aspect 11] The method according to Aspect 10, wherein the RNA polymerase promoter is a T7 viral RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter. [Embodiment 12] The method according to any one of embodiments 1 to 11, wherein the 3' group I self-splicing intron fragment and the 5' group I self-splicing intron fragment are derived from a Cyanobacterium Anabaena sp. Pre-tRNA-Leu gene. [Embodiment 13] The method of any one of embodiments 1 to 12, wherein the 3' group I self-splicing intron fragment and the 5' group I self-splicing intron fragment are derived from the Td gene of T4 bacteriophage. [Embodiment 14] The method of any one of embodiments 1 to 13, wherein the method further comprises forming a circular RNA by splint-mediated precursor RNA ligation. [Embodiment 15] The method according to any one of embodiments 1 to 14, wherein the vector is transfected into the cell using lipofection or electroporation before transcription. [Embodiment 16] The method of any one of embodiments 1 to 15, wherein the vector is transfected into the cell using a nanocarrier before transcription. [Embodiment 17] The method described in embodiment 16, wherein the nanocarrier is a lipid, a polymer, or a lipid-polymer hybrid. [Embodiment 18] The method of any one of embodiments 1 to 17, further comprising forming circular RNA and purifying the circular RNA using a size exclusion chromatography column in tris-EDTA or ion-pair reverse phase HPLC. [Embodiment 19] The method of any one of embodiments 1 to 17, further comprising forming a circular RNA and purifying the circular RNA in a high-performance liquid chromatography (HPLC) system in a triethylammonium acetate (TEAA)-acetonitrile buffer solution having a pH range of about 4 to 10 at a flow rate of about 0.01 to 5 mL / min. [Embodiment 20] The method of any one of embodiments 1 to 17, wherein the method further comprises forming circular RNA and purifying the circular RNA using phosphatase treatment. [Embodiment 21] The method of embodiment 20, further comprising incubating the precursor RNA in the presence of (i) magnesium ions and / or (ii) guanosine nucleotides or guanosine nucleosides. [Embodiment 22] The method of embodiment 21, wherein the incubation of the precursor RNA occurs at a temperature between about 20°C and about 60°C. [Aspect 23] A method according to any one of aspects 1 to 22, wherein transcription of the vector can occur in the presence of a nucleoside or a monophosphate or diphosphate nucleotide to incorporate the nucleoside or nucleotide as the first nucleotide of precursor RNA transcribed from the vector. [Embodiment 24] The method of embodiment 23, wherein the precursor RNA comprises a monophosphate 5' end that can be ligated using a ligase. [Aspect 25] The transcription of the vector comprises incorporating a nucleoside or a monophosphate or diphosphate nucleotide as the first nucleotide of an RNA strand transcribed from the vector or a transcript produced from the vector, a) guanosine nucleosides or monophosphate or diphosphate nucleotides; b) a cytidine nucleoside or a monophosphate or diphosphate nucleotide; c) uracil nucleosides or monophosphate or diphosphate nucleotides; d) an adenosine nucleoside or a monophosphate or diphosphate nucleotide, or e) combinations thereof; 25. The method of embodiment 24, wherein the method occurs in the presence of a substance comprising: [Embodiment 26] The method of any one of embodiments 1 to 25, wherein the protein coding region encodes a non-naturally occurring protein that includes one or more synthetic protein elements. [Embodiment 27] The method of any one of embodiments 1 to 26, wherein the precursor RNA comprises a nucleoside modification. [Embodiment 28] The nucleoside modification is N 6 -Methyladenosine (m6A), pseudouridine (Ψ), N 1 28. The method of embodiment 27, wherein the uridine is selected from the group consisting of 5-methylpseudouridine (m1Ψ) and 5-methoxyuridine (5moU). [Embodiment 29] The method according to any one of embodiments 1 to 28, wherein the vector comprises a 5' spacer element located at the 3' end of b). [Embodiment 30] The method according to any one of embodiments 1 to 29, wherein the vector comprises a 3' spacer element located at the 5' end of d). [Embodiment 31] The method according to embodiment 29 or 30, wherein the 5' spacer element or the 3' spacer element comprises a polyA sequence or a polyA-C sequence. [Aspect 32] The non-coding region may be an antisense RNA, transfer RNA (tRNA), transfer-messenger RNA (tmRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), Cajal body-specific small RNA (scaRNA), guide RNA (gRNA), Y RNA, splicing leader sequence RNA (SL RNA), microRNA (miRNA), small interfering RNA (siRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long non-coding RNA (lncRNA), Piwi-interacting RNA (piRNA), short hairpin RNA (shRNA), trans-acting siRNA (tasiRNA), repeat-associated siRNA (rasiRNA), 7SK 32. The method of any one of aspects 1 to 31, comprising elements encoding one or more RNAs selected from the group consisting of RNA (7SK), telomerase RNA component (TERC), vault RNAs (vRNA, vtRNA) and enhancer RNA (eRNA). [Embodiment 33] A precursor RNA, the precursor RNA comprising the following elements operably linked to each other and arranged in the following order: a) a 5' element that does not contain or contains at least one stem-loop structure; b) a 3' Group I self-splicing intron fragment containing a 3' splice site dinucleotide; c) protein coding or non-coding regions; d) a 5' Group I self-splicing intron fragment containing a 5' splice site dinucleotide; e) comprises a 3' element that does not contain or contains at least one stem-loop structure; provided that if the 5' element does not contain a stem-loop structure, the 3' element contains at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element contains at least one stem-loop structure; wherein the 5' element and the 3' element form a thermodynamically stable multidirectional junction; wherein the precursor RNA is capable of being translated in a cell and / or forming a biologically active circular RNA. [Embodiment 34] The precursor RNA according to embodiment 33, wherein the thermodynamically stable multidirectional junction is a three-way junction (3WJ), a four-way junction (4WJ), or a five-way junction (5WJ). [Embodiment 35] The 3WJ is a first branch of the 3WJ domain formed from the 5' portion of the 3WJa sequence and the 3' portion of the 3WJc sequence and including a first helical region; a second branch of the 3WJ domain formed from the 3' portion of the 3WJa sequence and the 5' portion of the 3WJb sequence and including a second helical region; a third branch of the 3WJ domain formed from the 3' portion of the 3WJb sequence and the 5' portion of the 3WJc sequence and including a third helical region; 35. The precursor RNA of embodiment 34, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds. [Aspect 36] 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 4, 3WJb comprises or consists of SEQ ID NO: 5, and 3WJc comprises or consists of SEQ ID NO: 6, or 3WJa comprises or consists of SEQ ID NO: 10, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 12, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 2, and 3WJc comprises or consists of SEQ ID NO: 13, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 13, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 14, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 1, 3WJb comprises or consists of SEQ ID NO: 16, and 3WJc comprises or consists of SEQ ID NO: 3, or 3WJa comprises or consists of SEQ ID NO: 17, 3WJb comprises or consists of SEQ ID NO: 15, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 20, and 3WJc comprises or consists of SEQ ID NO: 18, or 3WJa comprises or consists of SEQ ID NO: 19, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 22, or 3WJa comprises or consists of SEQ ID NO: 23, 3WJb comprises or consists of SEQ ID NO: 21, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 26, and 3WJc comprises or consists of SEQ ID NO: 24, or 3WJa comprises or consists of SEQ ID NO: 25, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 28, or 3WJa comprises or consists of SEQ ID NO: 29, 3WJb comprises or consists of SEQ ID NO: 27, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 32, and 3WJc comprises or consists of SEQ ID NO: 30, or 3WJa comprises or consists of SEQ ID NO: 31, 3WJb comprises or consists of SEQ ID NO: 33, and 3WJc comprises or consists of SEQ ID NO: 34, or 3WJa comprises or consists of SEQ ID NO: 41, 3WJb comprises or consists of SEQ ID NO: 11, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 44, and 3WJc comprises or consists of SEQ ID NO: 42, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 45, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 50, and 3WJc comprises or consists of SEQ ID NO: 48, or 3WJa comprises or consists of SEQ ID NO: 49, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 53, 3WJb comprises or consists of SEQ ID NO: 51, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 56, and 3WJc comprises or consists of SEQ ID NO: 54, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 58, or 3WJa comprises or consists of SEQ ID NO: 59, 3WJb comprises or consists of SEQ ID NO: 57, and 3WJc comprises or consists of SEQ ID NO: 60, or 3WJa comprises or consists of SEQ ID NO: 61, 3WJb comprises or consists of SEQ ID NO: 63, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of SEQ ID NO: 66, and 3WJc comprises or consists of SEQ ID NO: 64, or 3WJa comprises or consists of SEQ ID NO: 65, 3WJb comprises or consists of UGUCACGGG, and 3WJc comprises or consists of SEQ ID NO: 68, or 3WJa comprises or consists of SEQ ID NO: 43, 3WJb comprises or consists of SEQ ID NO: 69, and 3WJc comprises or consists of SEQ ID NO: 46, or 3WJa comprises or consists of SEQ ID NO: 47, 3WJb comprises or consists of SEQ ID NO: 70, and 3WJc comprises or consists of SEQ ID NO: 52, or 3WJa comprises or consists of SEQ ID NO: 55, 3WJb comprises or consists of SEQ ID NO: 71, and 3WJc comprises or consists of SEQ ID NO: 72, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 8, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 78, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 81, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 84, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 7, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 90, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 76, 3WJb comprises or consists of SEQ ID NO: 89, and 3WJc comprises or consists of SEQ ID NO: 9, or 3WJa comprises or consists of SEQ ID NO: 77, 3WJb comprises or consists of SEQ ID NO: 91, and 3WJc comprises or consists of SEQ ID NO: 79, or 3WJa comprises or consists of SEQ ID NO: 80, 3WJb comprises or consists of SEQ ID NO: 93, and 3WJc comprises or consists of SEQ ID NO: 82, or 3WJa comprises or consists of SEQ ID NO: 83, 3WJb comprises or consists of SEQ ID NO: 95, and 3WJc comprises or consists of SEQ ID NO: 85, or 3WJa comprises or consists of SEQ ID NO: 112, 3WJb comprises or consists of SEQ ID NO: 113, and 3WJc comprises or consists of SEQ ID NO: 114, or 3WJa comprises or consists of SEQ ID NO: 119, 3WJb comprises or consists of SEQ ID NO: 120, and 3WJc comprises or consists of SEQ ID NO: 121, or 3WJa includes AUGUGUA, 3WJb includes UACUUUG, and 3WJc includes AUCAUG, or 3WJa contains GCGUU, 3WJb contains UUCGC, and 3WJc contains GCCAUAGCG, or 3WJa includes GUAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUAC, or 3WJa includes AUGGUA, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGUGU, 3WJb includes ACUUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UGUGU, 3WJb includes ACGUU, and 3WJc includes AAUCA, or 3WJa includes UGGU, 3WJb includes ACUGU, and 3WJc includes AUCA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUGG, 3WJb includes UCACGG, and 3WJc includes CCUCUUA, or 3WJa includes UAUGGCAC, 3WJb includes GUCACGG, and 3WJc includes CUCUUA, or 3WJa includes UAUG, 3WJb includes CAGGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UAUGU, 3WJb includes GCAGGG, and 3WJc includes CUUG, or 3WJa includes UAUGU, 3WJb includes GCAGG, and 3WJc includes UCUUG, or 3WJa includes UGUGU, 3WJb includes ACUUUGU, and 3WJc includes AUCA, or 36. The precursor RNA of embodiment 35, wherein 3WJa comprises UGUGU, 3WJb comprises ACUUU, and 3WJc comprises AAAUCA. [Embodiment 37] The 4WJ is a first branch of the 4WJ domain formed from the 5' portion of the 4WJa sequence and the 3' portion of the 4WJd sequence and including a first helical region; a second branch of the 4WJ domain formed from the 3' portion of the 4WJa sequence and the 5' portion of the 4WJb sequence and including a second helical region; a third branch of the 4WJ domain formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence and including a third helical region; a fourth branch of the 4WJ domain formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence and including a fourth helical region; 35. The precursor RNA of embodiment 34, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds. [Aspect 38] 4WJa comprises or consists of SEQ ID NO: 7, 4WJb comprises or consists of SEQ ID NO: 8, 4WJc comprises or consists of SEQ ID NO: 9, and 4WJd comprises or consists of SEQ ID NO: 102, or 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO: 106, or 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO: 118, or 4WJa comprises UGCAGGUG, 4WJb comprises ACGGGC, 4WJc comprises CCAGCA, and 4WJd comprises SEQ ID NO: 67, or 4WJa comprises SEQ ID NO: 74, 4WJb comprises AACUG, 4WJc comprises SEQ ID NO: 75, and 4WJd comprises AUCAUG, or 38. The precursor RNA of embodiment 37, wherein 4WJa comprises SEQ ID NO: 122, 4WJb comprises GAACU, 4WJc comprises SEQ ID NO: 123, and 4WJd comprises AAUCA. [Embodiment 39] The 5WJ is a first branch of the 5WJ domain formed from the 5' portion of the 5WJa sequence and the 3' portion of the 5WJe sequence and including a first helical region; a second branch of the 5WJ domain formed from the 3' portion of the 5WJa sequence and the 5' portion of the 5WJb sequence and including a second helical region; a third branch of the 5WJ domain formed from the 3' portion of the 5WJb sequence and the 5' portion of the 5WJc sequence and including a third helical region; a fourth branch of the 5WJ domain formed from the 3' portion of the 5WJc sequence and the 5' portion of the 5WJd sequence and including a fourth helical region; a fifth branch of the 5WJ domain formed from the 3' portion of the 5WJd sequence and the 5' portion of the 5WJe sequence and including a fifth helical region; 35. The precursor RNA of embodiment 34, wherein each said helical region comprises a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds. [Aspect 40] 5WJa comprises or consists of SEQ ID NO: 107, 5WJb comprises or consists of SEQ ID NO: 108, 5WJc comprises or consists of SEQ ID NO: 109, 5WJd comprises or consists of SEQ ID NO: 110, and 5WJe comprises or consists of SEQ ID NO: 111, or The precursor RNA of embodiment 39, wherein 5WJa comprises GUGA, 5WJb comprises UUGC, 5WJc comprises GUGU, 5WJd comprises AUGC, and 5WJe comprises GUGC. [Aspect 41] The non-coding region is an antisense RNA, transfer RNA (tRNA), transfer-messenger RNA (tmRNA), ribosomal RNA (rRNA), signal recognition particle RNA (7SL RNA or SRP RNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), SmY RNA (SmY), Cajal body-specific small RNA (scaRNA), guide RNA (gRNA), Y RNA, splicing leader sequence RNA (SL RNA), microRNA (miRNA), small interfering RNA (siRNA), cis-natural antisense transcript (cis-NAT), CRISPR RNA (crRNA), long non-coding RNA (lncRNA), Piwi-interacting RNA (piRNA), short hairpin RNA (shRNA), trans-acting siRNA (tasiRNA), repeat-associated siRNA (rasiRNA), 7SK 41. The precursor RNA of any one of aspects 33 to 40, comprising one or more RNAs selected from the group consisting of RNA (7SK), telomerase RNA component (TERC), vault RNAs (vRNA, vtRNA) and enhancer RNA (eRNA). [Aspect 42] A vector encoding the precursor RNA according to any one of aspects 33 to 41. [Aspect 43] The vector according to Aspect 42, which is a plasmid, a viral vector, a polymerase chain reaction (PCR) product, a cosmid, a bacterial artificial chromosome (BAC), or a yeast artificial chromosome (YAC).
[0101] Advantages of the present disclosure may include incorporating RNA nanotechnology designs (e.g., thermodynamically stable 3WJ, 4WJ, and 5WJ structures) into the artificial circular RNA self-splicing process to promote correct folding and improve self-splicing efficiency, thereby paving the way for the mass production of high-quality circular RNAs for use in therapeutic development.
[0102] All references cited herein are incorporated by reference as if each reference were specifically and individually indicated to be incorporated by reference herein. Any reference is referenced as though disclosed prior to the filing date and should not be construed as an admission that such reference is not entitled to antedate the present disclosure by virtue of prior invention.
[0103] It should be understood that each one, two, or more of the above elements may find useful application in other types of methods different from those described above. Without further analysis, the above content sufficiently reveals the gist of the present disclosure so that others, by applying their current knowledge, can readily employ them in various applications without omitting features that, in the light of the prior art, fairly constitute essential features of the general or specific aspects of the present disclosure as set forth in the appended claims. The above embodiments are presented by way of example only, and the scope of the present disclosure is limited by the following claims.
Claims
1. 1. A method for producing circular ribonucleic acid (RNA), said method comprising: transcribing a vector to form a precursor RNA, said precursor RNA comprising the following elements operably linked to each other and arranged in tandem in a 5' to 3' direction: a) a 5' element, b) a 3' group I self-splicing intron fragment; c) elements that do not contain or contain internal ribosome entry sites (IRES) and protein coding regions or elements that contain non-coding regions; d) a 5' group I self-splicing intron fragment, and e) comprises a 3' element; the 5' element and the 3' element form a stable structure with a Gibbs free energy (ΔG) of -190 kcal / mol to -9.0 kcal / mol, and the stable structure is not a duplex with at least 95% base pairs between the 5' element and the 3' element; the 3′ Group I self-splicing intron fragment and the 5′ Group I self-splicing intron fragment form a self-cleaving and self-ligating RNA molecule to generate a circular RNA; the 5' element and the 3' element form a thermodynamically stable multidirectional junction RNA structure; the thermodynamically stable multidirectional junction RNA structure is a four-way junction (4WJ); The 4WJ is a first branch of the 4WJ domain formed from a 5' portion of the 4WJa sequence and a 3' portion of the 4WJd sequence and including a first helical region; a second branch of the 4WJ domain formed from the 3' portion of the 4WJa sequence and the 5' portion of the 4WJb sequence and including a second helical region; a third branch of the 4WJ domain formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence and including a third helical region; a fourth branch of the 4WJ domain formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence and including a fourth helical region; each said helical region comprises a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds; method.
2. 2. The method of claim 1, wherein the 5' element does not contain or contains a sequence that forms at least one stem-loop structure, and the 3' element does not contain or contains a sequence that forms at least one stem-loop structure, wherein if the 5' element does not contain a stem-loop structure, the 3' element contains at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element contains at least one stem-loop structure.
3. 4WJa comprises or consists of SEQ ID NO:7, 4WJb comprises or consists of SEQ ID NO:8, 4WJc comprises or consists of SEQ ID NO:9, and 4WJd comprises or consists of SEQ ID NO:102, or 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO: 106; or 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO: 118; or 4WJa comprises UGCAGGUG, 4WJb comprises ACGGGC, 4WJc comprises CCAGCA, and 4WJd comprises SEQ ID NO: 67; or 4WJa comprises SEQ ID NO: 74, 4WJb comprises AACUG, 4WJc comprises SEQ ID NO: 75, and 4WJd comprises AUCAUG; or 4WJa comprises SEQ ID NO: 122, 4WJb comprises GAACU, 4WJc comprises SEQ ID NO: 123, and 4WJd comprises AAUCA, or 4WJa comprises SEQ ID NO: 125, 4WJb comprises SEQ ID NO: 126, 4WJc comprises SEQ ID NO: 127, and 4WJd comprises SEQ ID NO: 128; The method of claim 1.
4. The IRES is a virus that is expressed in a variety of viruses, including Taura syndrome virus, Assassin bug virus, Theiler's encephalomyelitis virus, Simian virus 40, red fire ant virus 1, wheat aphid virus, reticuloendotheliosis virus, human poliovirus 1, German winged green bug enteric virus, Kashmir wasp virus, human rhinovirus 2, leafhopper virus-1, human immunodeficiency virus type 1, leafhopper virus-1, Himetovirus P virus, hepatitis C virus, hepatitis A virus, hepatitis G virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinitis virus, white-spotted picorna-like virus, encephalomyocarditis virus (EMCV), Drosophila C virus, cruciferous tobamovirus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen brood virus, aphid fatal paralysis virus, avian encephalomyelitis virus, bee acute paralysis enterovirus, hibiscus chlorotic ringspot virus, classical swine fever virus, human fibroblast growth factor 2 ( FGF2), human surfactant drug protein A1 (SFTPA1), human acute myeloid leukemia protein 1 / runt-related transcription factor 1 (AML1 / RUNX1), Drosophila antennapedia, human aquaporin-4 (AQP4), human angiotensin II receptor type 1 (AT1R), human BCL2-associated immortality gene 1 (BAG-1), human B-cell lymphoma 2 (BCL2), human immunoglobulin-binding protein (BiP), human inhibitor of apoptosis family protein 1 (c-IAP1), human c-myc, human eukaryotic translation initiation factor 4G (eIF4G), mouse N-deacetylase and N-sulfotransferase 4 (NDST4L), human lymphoid enhancer-binding factor-1 (LEF1), mouse hypoxia-inducible factor 1 subunit alpha (HIF1α), human N-myc, mouse glial cell and testis-specific homeobox protein (Gtx), human cyclin-dependent kinase inhibitor 1B (p27kip1), human platelet-derived growth factor B / simian sarcoma virus human homolog (PDGF2 / c-sis), human p53, human proviral integration site of Moloney murine leukemia virus-1 (Pim-1), mouse RNA-binding protein white 3 (Rbm3), Drosophila reaper, canine Scamper, Drosophila superbreast gene (Ubx), salivirus, coronavirus, parechovirus, human N-ras upstream (UNR), mouse dystrophin-related protein A (UtrA), human vascular endothelial growth factor A (VEGF-A), human X-linked apoptosis protein (XIAP), Drosophila hairless, budding yeast transcription factor II 2. The method of claim 1, wherein the IRES sequence is selected from IRES sequences from a virus or gene selected from the group consisting of TFIID, Saccharomyces cerevisiae Yes1-associated transcription factor (YAP1), human proto-oncogene tyrosine protein kinase Src (c-src), human fibroblast growth factor 1 (FGF-1), monkey virus, turnip crinkle virus, coxsackievirus B3 (CVB3), and coxsackievirus A (CVB1 / 2).
5. 2. The method of claim 1, wherein the 3' group I self-splicing intron fragment and the 5' group I self-splicing intron fragment are derived from a Cyanobacterium anabaena sp. Pre-tRNA-Leu gene and / or a T4 bacteriophage Td gene.
6. 2. The method of claim 1, wherein the method further comprises forming the circular RNA by splint-mediated precursor RNA ligation.
7. A precursor RNA, said precursor RNA comprising the following elements operably linked to each other and arranged in tandem in a 5' to 3' direction: a) a 5' element, b) a 3' group I self-splicing intron fragment; c) elements that do not contain or contain IRES and protein coding regions or elements that contain non-coding regions; d) a 5' group I self-splicing intron fragment, and e) containing a 3' element the 5' element and the 3' element form a stable structure with a Gibbs free energy (ΔG) of -190 kcal / mol to -9.0 kcal / mol, and the stable structure is not a duplex with at least 95% base pairs between the 5' element and the 3' element; the 3′ Group I self-splicing intron fragment and the 5′ Group I self-splicing intron fragment form a self-cleaving and self-ligating RNA molecule to generate a circular RNA; the 5' element and the 3' element form a thermodynamically stable multidirectional junction RNA structure; the thermodynamically stable multidirectional junction RNA structure is a four-way junction (4WJ), The 4WJ is a first branch of the 4WJ domain formed from a 5' portion of the 4WJa sequence and a 3' portion of the 4WJd sequence and including a first helical region; a second branch of the 4WJ domain formed from the 3' portion of the 4WJa sequence and the 5' portion of the 4WJb sequence and including a second helical region; a third branch of the 4WJ domain formed from the 3' portion of the 4WJb sequence and the 5' portion of the 4WJc sequence and including a third helical region; a fourth branch of the 4WJ domain formed from the 3' portion of the 4WJc sequence and the 5' portion of the 4WJd sequence and including a fourth helical region; wherein each said helical region comprises a plurality of RNA nucleotide pairs that form standard Watson-Crick bonds.
8. The precursor RNA of claim 7, wherein the 5' element does not contain or contains a sequence that forms at least one stem-loop structure, and the 3' element does not contain or contains a sequence that forms at least one stem-loop structure, and if the 5' element does not contain a stem-loop structure, the 3' element contains at least one stem-loop structure, and if the 3' element does not contain a stem-loop structure, the 5' element contains at least one stem-loop structure, and wherein the 5' element and the 3' element form a thermodynamically stable multidirectional junction RNA structure.
9. 4WJa comprises or consists of SEQ ID NO:7, 4WJb comprises or consists of SEQ ID NO:8, 4WJc comprises or consists of SEQ ID NO:9, and 4WJd comprises or consists of SEQ ID NO:102, or 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO: 106; or 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO: 118; or 4WJa comprises UGCAGGUG, 4WJb comprises ACGGGC, 4WJc comprises CCAGCA, and 4WJd comprises SEQ ID NO: 67; or 4WJa comprises SEQ ID NO: 74, 4WJb comprises AACUG, 4WJc comprises SEQ ID NO: 75, and 4WJd comprises AUCAUG; or 4WJa comprises SEQ ID NO: 122, 4WJb comprises GAACU, 4WJc comprises SEQ ID NO: 123, and 4WJd comprises AAUCA, or 4WJa comprises SEQ ID NO: 125, 4WJb comprises SEQ ID NO: 126, 4WJc comprises SEQ ID NO: 127, and 4WJd comprises SEQ ID NO: 128; The precursor RNA of claim 7.
10. A method for producing a protein in a cell, the method comprising introducing a precursor RNA described in claim 7 that contains a protein coding region into the cell and producing the protein.
11. 10. A method for editing a gene in a cell, the method comprising introducing into the cell the precursor RNA of claim 7, which comprises an editable non-coding region of the gene, and editing the gene.
12. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 7, 4WJb comprises or consists of SEQ ID NO: 8, 4WJc comprises or consists of SEQ ID NO: 9, and 4WJd comprises or consists of SEQ ID NO:
102.
13. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO:
106.
14. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO:
118.
15. The precursor RNA described in claim 9, wherein 4WJa includes or consists of UGCAGGUG, 4WJb includes or consists of ACGGGC, 4WJc includes or consists of CCAGCA, and 4WJd includes or consists of SEQ ID NO:
67.
16. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 74, 4WJb comprises or consists of AACUG, 4WJc comprises or consists of SEQ ID NO: 75, and 4WJd comprises or consists of AUCAUG.
17. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 122, 4WJb comprises or consists of GAACU, 4WJc comprises or consists of SEQ ID NO: 123, and 4WJd comprises or consists of AAUCA.
18. The precursor RNA described in claim 9, wherein 4WJa comprises or consists of SEQ ID NO: 125, 4WJb comprises or consists of SEQ ID NO: 126, 4WJc comprises or consists of SEQ ID NO: 127, and 4WJd comprises or consists of SEQ ID NO:
128.
19. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 7, 4WJb comprises or consists of SEQ ID NO: 8, 4WJc comprises or consists of SEQ ID NO: 9, and 4WJd comprises or consists of SEQ ID NO:
102.
20. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 103, 4WJb comprises or consists of SEQ ID NO: 104, 4WJc comprises or consists of SEQ ID NO: 105, and 4WJd comprises or consists of SEQ ID NO:
106.
21. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 115, 4WJb comprises or consists of SEQ ID NO: 116, 4WJc comprises or consists of SEQ ID NO: 117, and 4WJd comprises or consists of SEQ ID NO:
118.
22. The method of claim 3, wherein 4WJa comprises or consists of UGCAGGUG, 4WJb comprises or consists of ACGGGC, 4WJc comprises or consists of CCAGCA, and 4WJd comprises or consists of SEQ ID NO:
67.
23. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 74, 4WJb comprises or consists of AACUG, 4WJc comprises or consists of SEQ ID NO: 75, and 4WJd comprises or consists of AUCAUG.
24. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 122, 4WJb comprises or consists of GAACU, 4WJc comprises or consists of SEQ ID NO: 123, and 4WJd comprises or consists of AAUCA.
25. The method of claim 3, wherein 4WJa comprises or consists of SEQ ID NO: 125, 4WJb comprises or consists of SEQ ID NO: 126, 4WJc comprises or consists of SEQ ID NO: 127, and 4WJd comprises or consists of SEQ ID NO: 128.