Vector and method for preparing circular RNA free from non-target sequence
By modifying IGS with class I self-splicing introns, exon-free circular RNA was prepared, which solved the problem of increased immunogenicity and therapeutic risk caused by exons, and improved the stability and application range of circular RNA.
Patent Information
- Application Number
- PCT/CN2025/108778
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-01
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
The presence of exon sequences during the preparation of existing circular RNAs increases immunogenicity and therapeutic risks, limiting their application.
By modifying the internal guide sequence of class I self-splicing introns to eliminate the need for adjacent exons during circularization, exon-free circular RNA was prepared. Circularization of self-splicing RNA was then achieved using the recA gene of Bacillus anthracis and the IGS design from the 23S ribosome gene of Coxiella behnkeni.
It effectively removes unnecessary exon sequences from circular RNA, reduces the risk of immunogenicity, and improves the stability and application range of circular RNA.
Smart Images

Figure PCTCN2025108778-FTAPPB-I100001 
Figure PCTCN2025108778-FTAPPB-I100002 
Figure PCTCN2025108778-FTAPPB-I100003
Abstract
Description
Vectors and methods for making circular rnas free of non-desired sequences
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202410954230.5 filed on July 16, 2024, and Chinese Patent Application No. 202411556442.4 filed on November 1, 2024, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present application relates to an RNA molecule comprising a 3' self-cleaving intron fragment comprising a 3' splice site, a sequence of interest, and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein the elements are operably linked in order, wherein the 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment are capable of circularizing the RNA molecule and removing the RNA molecule during circularization, wherein the sequence of interest ends in a U at its 3' end sequence, the 5' self-cleaving intron fragment comprises an internal guide sequence which is reverse complementary to the 5' end sequence of the sequence of interest, and which 5' end sequence of the internal guide sequence starts with a G and is reverse complementary to the 3' end sequence of the sequence of interest. The present application also relates to a DNA molecule capable of transcribing the above-mentioned RNA molecule, and the use of the RNA molecule and the DNA molecule of the present application for the preparation of a circular RNA, in particular a circular RNA molecule free of an exon adjacent to the 3' self-cleaving intron fragment or the 5' self-cleaving intron fragment. BACKGROUND
[0004] With the breakthrough of technology, ribonucleic acid (RNA) therapy, including mRNA vaccine, has developed into a new field with strong potential in the biopharmaceutical industry. However, RNA, especially mRNA, has the problem of easy degradation. How to improve the stability of mRNA and thus prolong the duration of protein translation has attracted great attention and interest. An effective way is to convert mRNA into circular RNA (cirRNA), which is a class of closed circular single-stranded RNA molecules covalently connected at the head and tail, which lacks 5' and 3' ends, and thus has stronger resistance to exonuclease degradation, thereby having a longer half-life.
[0005] A gene segment usually contains introns and exons. After RNA splicing, the exon sequences are spliced together to form mature mRNA. The splicing removal of most introns requires multiple protein interactions to complete, but there are also some special introns that are removed by self-splicing. Such introns with self-splicing function can be divided into two categories according to structure and splicing mechanism: class I and class II. The processing of class I introns is initiated by an ester exchange reaction mediated by an exogenous guanosine cofactor (GTP) [3,4] . The splicing of class II introns is similar to that of eukaryotic mRNA precursors, in which the 5' splicing site is attacked by the -OH group of a nucleotide called a branch point bulge, first producing an intron lariat structure, and then the -OH released at the 5' exon attacks the 3' splicing site, finally allowing the two end exons to be spliced together [4,5] . Using the self-splicing function of class I and class II introns, researchers have designed sequences to split the self-splicing intron into two halves, which are loaded onto the two ends of the RNA sequence to be circularized, and the RNA circularization can be achieved by in vitro catalysis [1,6-8] . Daniel G. Anderson et al. [1,6-8] Using class I self-splicing introns such as the intron of Anabaena pre-tRNA-Leu gene or T4 bacteriophage Td gene (T4 td), in vitro circularization of RNA is achieved. Since the two ends of the circular RNA are covalently closed, there is no exposed 5' end, and therefore no cap structure to initiate RNA translation. To solve the problem of circular RNA translation, researchers add an internal ribosome entry site sequence (IRES) to the circular RNA, making it have stable in vitro and in vivo translation function.
[0006] The primary structure of class I self-splicing introns varies greatly, but the secondary and tertiary structures are relatively conserved. As shown in FIG. 1, the secondary structure of class I self-splicing introns is usually composed of paired (P) elements P1-P10 and single-stranded loop regions. The intron contains an internal guide sequence (IGS) near the 5' end, and the 5' end sequence of the IGS can complementarily pair with the adjacent exon at the 3' end of the intron, while the 3' end sequence of the IGS can complementarily pair with the adjacent exon at the 5' end of the intron. After the intron and the exon are complementarily paired, splicing occurs at the splice sites of P1 and P10, i.e., the junction between the intron and the exon (P1 and P10 regions pointed by triangles in FIG. 1), causing the intron to be removed from the overall structure, and the exons at the P1 and P10 regions to be connected [9] . When using self-splicing introns to prepare circular RNA, the adjacent exons on both sides of the intron need to be added to successfully complete splicing. After splicing is completed, the intron is removed from the final circular RNA, and the sequences of the adjacent exons on both sides remain in the final circular RNA product.
[0007] In addition, since the splice site of the substituted intron-exon element (PIE) is located near the IRES, and both sequences are highly structured, the IRES sequence can interfere with the folding of the splicing ribozyme. In order to enable these structures to fold independently, a series of spacer sequences were designed between the PIE splice site and the IRES, and by increasing the spacer sequence, the splicing efficiency was increased by nearly one-fold.
[0008] However, the adjacent exon sequences required for self-splicing, as well as the artificially added exogenous spacer sequence, which remain in the circular RNA end product, greatly increase the immunogenicity and related treatment risks, which may limit the application range of the circular RNA.
[0009] The citation of any document herein is not intended as an admission that such document is prior art with respect to the application. SUMMARY
[0010] The inventors of the present application found that in the self-splicing intron in the Bacillus anthracis recA gene shown in FIG. 2A, the 3' end sequence (P1) in the internal guide sequence (IGS) for reverse complementary pairing with the 5' exon sequence (P1 antisense strand), and / or the 5' end sequence (P10) for reverse complementary pairing with the 3' exon (P10 antisense strand), except for the G in P1 for forming a G:U pairing splice site with the last U of the 5' exon, can be modified (see FIG. 2B), so that a self-splicing intron without the 5' adjacent exon, or the 3' adjacent exon, or both exons can be used to prepare a circular RNA, and the prepared circular RNA naturally also does not contain the above 5' adjacent exon, 3' adjacent exon, or both exons.
[0011] Specifically, when the 3' end sequence of the IGS is designed to start with a nucleotide of base G and to be reverse complementarily paired with the 3' end sequence of the RNA to be circularized (as the P1 antisense strand), and the 5' end sequence of the IGS is designed to be reverse complementarily paired with the 5' end sequence of the RNA to be circularized (as the P10 antisense strand), the RNA circularization can be completed without using the 5' side adjacent exon of the 5' intron fragment and the 3' side adjacent exon of the 3' intron fragment. The circular RNA prepared does not contain the 5' side adjacent exon, the 3' side adjacent exon, or both of the above exons.
[0012] The inventors of the present application made similar modifications to the IGS of the class I self-splicing intron in the 23S ribosomal gene of Coxiella burnetii (as shown in FIG. 5), and found that the same effect can be achieved. Thus, the inventors of the present application believe that the class I self-splicing intron from different genes can be modified as described above to achieve the same effect.
[0013] In addition, the inventors of the present application also found that, in the case where the above-mentioned IGS modification is made and no artificial spacer sequence is added between the intron splicing site and the IRES, the RNA circularization efficiency can still meet the production requirements.
[0014] Thus, in a first aspect, the present application provides a single-stranded DNA molecule for preparing a circular RNA, which can comprise, in order from 5' to 3', a 3' self-splicing intron fragment comprising a 3' splicing site, a sequence of interest, and a 5' self-splicing intron fragment comprising a 5' splicing site, wherein each element is operably linked.
[0015] 3' self-cleaving intron fragment and 5' self-cleaving intron fragment can enable circularization of an RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule. During the circularization of the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule, the RNA fragment transcribed from the 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment is removed from the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule. The 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment can be derived from the same self-cleaving intron, particularly a Class I self-cleaving intron. In some embodiments, the 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment can be derived from a self-cleaving intron in the recA gene of Bacillus anthracis, or a self-cleaving intron in the 23S ribosomal gene of Coxiella burnetii. The 5' self-cleaving intron fragment and the 3' self-cleaving intron fragment, when arranged in this order, can join to form a self-cleaving intron having the self-cleaving function described above, such as a complete self-cleaving intron. In some embodiments, the 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment can be derived from a self-cleaving intron in the RecA gene of Bacillus anthracis, and the 3' self-cleaving intron fragment can comprise the nucleotide sequence set forth in SEQ ID NO: 10. In some embodiments, the 3' self-cleaving intron fragment and the 5' self-cleaving intron fragment can be derived from a self-cleaving intron in the 23S ribosomal gene of Coxiella burnetii, and the 3' self-cleaving intron fragment can comprise the nucleotide sequence set forth in SEQ ID NO: 21.
[0016] The 5' self-cleaving intron fragment can comprise an internal guide sequence.
[0017] The 5' end sequence of the internal guide sequence can be designed to be reverse complementary to the 5' end sequence of the sequence of interest. The 5' end sequence of the internal guide sequence can be reverse complementary to any number of nucleotides of the 5' end of the sequence of interest. For example, the 5' end sequence of the internal guide sequence can be reverse complementary to 2-15 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nt), e.g., 2-12 nt, of the 5' end of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to 3-6 nt of the 5' end of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to the 5' end sequence of the sequence of interest, and the single-stranded DNA molecule can further comprise a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 5' side of the 5' self-cleaving intron fragment. The naturally adjacent exon can comprise a partial or truncated sequence. The naturally adjacent exon or a sequence corresponding to a naturally adjacent exon can comprise nucleotides of length 3-20 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt), e.g., 3-10 nt. The single-stranded DNA molecule of the application can comprise at least no naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 3' side of the 3' self-cleaving intron fragment.
[0018] The 3' end sequence of the internal guide sequence can be designed to be reverse complementary to the 3' end sequence of the sequence of interest. The 3' end sequence of the internal guide sequence can start with a nucleotide that is a G and be reverse complementary to the 3' end sequence of the sequence of interest. The 3' end sequence of the sequence of interest can end with a nucleotide that is a T. The starting nucleotide that is a G at the 3' end of the internal guide sequence of the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule can form a G:U pair with the nucleotide that is a U at the end of the 3' end sequence of the sequence of interest that is transcribed. The 3' end sequence of the internal guide sequence can be reverse complementary to any number of nucleotides at the 3' end of the sequence of interest. For example, the 3' end sequence of the internal guide sequence can be reverse complementary to 3-20 nt (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt), e.g., 3-15 nt, of the 3' end of the sequence of interest. In some embodiments, the 3' end sequence of the internal guide sequence can be reverse complementary to 3-10 nt of the 3' end of the sequence of interest. In other embodiments, the 3' end sequence of the internal guide sequence can be reverse complementary to 3-9 nt of the 3' end of the sequence of interest. In some embodiments, the 3' end sequence of the internal guide sequence can start with a nucleotide that is a G and be reverse complementary to the end of the 3' end sequence of the sequence of interest that is a T, and the single-stranded DNA molecule can further comprise a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon 3' to the 3' spliceosomal intron fragment. The naturally adjacent exon can comprise a partial or truncated sequence. The naturally adjacent exon or a sequence corresponding to a naturally adjacent exon can comprise nucleotides that are 3-20 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt), e.g., 3-10 nt, in length. The single-stranded DNA molecule of the application can not comprise a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon at least 5' to the 5' spliceosomal intron fragment.
[0019] In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to 2-15 nt nucleotides of the 5' end of the sequence of interest, or the 3' end sequence of the internal guide sequence can be reverse complementary to 3-20 nt nucleotides of the 3' end of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to 2-15 nt nucleotides of the 5' end of the sequence of interest, and the 3' end sequence of the internal guide sequence can be reverse complementary to 3-20 nt nucleotides of the 3' end of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to 2-6 nt nucleotides of the 5' end of the sequence of interest, or the 3' end sequence of the internal guide sequence can be reverse complementary to 3-10 nt nucleotides of the 3' end of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to 2-6 nt nucleotides of the 5' end of the sequence of interest, and the 3' end sequence of the internal guide sequence can be reverse complementary to 3-10 nt nucleotides of the 3' end of the sequence of interest.
[0020] In some embodiments, the 5' end sequence of the internal guide sequence can be designed to be reverse complementary to the 5' end sequence of the sequence of interest, and the 3' end sequence of the internal guide sequence can be designed to be reverse complementary to the 3' end sequence of the sequence of interest. In some embodiments, the 5' end sequence of the internal guide sequence can be reverse complementary to the 5' end sequence of the sequence of interest, and the 3' end sequence of the internal guide sequence can start with a nucleotide of base G and be reverse complementary to the end of the 3' end sequence of the sequence of interest, and the sequence of interest can be a nucleotide of base T at the end of the 3' end sequence. In some embodiments, the 3' end sequence of the internal guide sequence can start with a nucleotide of base G and be reverse complementary to the 3' end sequence of the sequence of interest, wherein the sequence of interest can be a nucleotide of base T at the end of the 3' end sequence, wherein the initial nucleotide of base G of the 3' end sequence of the internal guide sequence of the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule can form a G:U pair with the nucleotide of base U at the end of the 3' end sequence of the sequence of interest transcribed. The single-stranded DNA molecule of the present application can not comprise a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 5' side of the 5' self-splicing intron fragment, and can not comprise a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 3' side of the 3' self-splicing intron fragment.
[0021] The sequence of interest can comprise an open reading frame encoding a peptide or protein of interest, a sequence of a translational functional element, a sequence of non-coding DNA, a monoclone site, or a polyclone site. The translational functional element can be selected from a translational initiation element, and a translational enhancer element. The translational initiation element can be a sequence that initiates RNA translation, such as an internal ribosome entry site (IRES). The translational enhancer element can be a sequence that enhances RNA translation, such as a poly(A) tail, a poly(C), a cap-independent translation enhancer (CITE), or an EIF4 ligand family recognition sequence, etc. In a circular RNA molecule prepared from a single-stranded DNA of the present application, the translational enhancer element can be located 5' or 3' to the open reading frame encoding the peptide or protein of interest. In a circular RNA molecule prepared from the complementary strand of a single-stranded DNA of the present application, the translational enhancer element can be located 5' or 3' to the translational initiation element. In some embodiments, the translational functional element can comprise a translational initiation element. In some embodiments, the translational functional element can be a translational initiation element. In some embodiments, the translational initiation element can be an IRES.
[0022] The target sequence may contain an open reading frame encoding the target peptide or protein, and a sequence of translational functional elements. The target sequence may include, from the 5' end to the 3' end, (a) the sequence of the translational functional element (e.g., IRES) and the open reading frame encoding the target peptide or protein, (b) the open reading frame encoding the target peptide or protein and the sequence of the translational functional element (e.g., IRES), (c) the 3' end sequence of the translational functional element (e.g., IRES), the open reading frame encoding the target peptide or protein, and the 5' end sequence of the translational functional element (e.g., IRES), wherein the 5' end sequence and the 3' end sequence of the translational functional element form the translational functional element when arranged in this order, or (d) the 3' end sequence of the open reading frame encoding the target peptide or protein, the sequence of the translational functional element (e.g., IRES), and the 5' end sequence of the open reading frame encoding the target peptide or protein, wherein the 5' end sequence and the 3' end sequence of the open reading frame encoding the target peptide or protein form the open reading frame encoding the target peptide or protein when arranged in this order. In some embodiments, the target sequence may comprise one or more open reading frames encoding a target peptide or protein, and a sequence of a translational functional element (e.g., IRES). The target sequence may comprise, from its 5' end to its 3' end, (a) the open reading frame encoding the one or more target peptide or protein, and the sequence of the translational functional element (e.g., IRES); (b) the sequence of the translational functional element (e.g., IRES), and the one or more open reading frames encoding the target peptide or protein; (c) the 3' end sequence of the translational functional element (e.g., IRES), the one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of the translational functional element (e.g., IRES); or (d) the 3' end sequence of one of the open reading frames encoding the target peptide or protein, other open reading frames encoding the target peptide or protein (which may exist), the sequence of the translational functional element (e.g., IRES), and the 5' end sequence of one of the open reading frames encoding the target peptide or protein.
[0023] In some embodiments, the sequence of interest can comprise a sequence of an open reading frame encoding a peptide or protein of interest, and a sequence of a translational functional element selected from the group consisting of a translational initiation element, and a translational enhancer element, wherein the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the sequence of the translational functional element, the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the sequence of the translational functional element, wherein the 5' end sequence of the translational functional element and the 3' end sequence of the translational functional element, when arranged in this order, form the translational functional element. In some embodiments, the sequence of interest can comprise a sequence of an open reading frame encoding a peptide or protein of interest, and a sequence of a translational functional element selected from the group consisting of a translational initiation element, and a translational enhancer element, wherein the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, the sequence of the translational functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest.
[0024] The sequence of interest can comprise a sequence of non-coding DNA. The sequence of interest can (a) comprise the sequence of the non-coding DNA, or (b) comprise, from 5' end to 3' end, a 3' end sequence of the sequence of the non-coding DNA, and a 5' end sequence of the sequence of the non-coding DNA, wherein the 5' end sequence of the non-coding DNA sequence and the 3' end sequence of the non-coding DNA sequence, when arranged in this order, form the non-coding RNA sequence.
[0025] The sequence of interest can comprise an open reading frame encoding a peptide or protein of interest. The sequence of interest can (a) comprise the open reading frame encoding the peptide or protein of interest, or (b) comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest. In some embodiments, the sequence of interest can comprise one or more open reading frames encoding a peptide or protein of interest. The sequence of interest can comprise, from 5' end to 3' end, (a) the one or more open reading frames encoding a peptide or protein of interest, or (b) a 3' end sequence of one of the open reading frames encoding a peptide or protein of interest, other open reading frames encoding a peptide or protein of interest (if present), and a 5' end sequence of one of the open reading frames encoding a peptide or protein of interest.
[0026] The sequence of interest can comprise a monoclonal site, or a polyclonal site. Any desired sequence, such as an open reading frame encoding a peptide or protein of interest, and a sequence of a translational functional element (e.g., IRES), etc., can be inserted into the single-stranded DNA molecule of the present application via the monoclonal site, or the polyclonal site.
[0027] The sequence of interest can comprise a monoclonal site or a polyclonal site, and a sequence of a translational functional element (e.g., IRES). The sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the sequence of the translational functional element (e.g., IRES), the monoclonal site or the polyclonal site, and a 5' end sequence of the sequence of the translational functional element (e.g., IRES). Any desired sequence, such as an open reading frame encoding a peptide or protein of interest, etc., can be inserted into the single-stranded DNA molecule of the present application via the monoclonal site, or the polyclonal site.
[0028] The sequence of interest can have a length of 50-8000 nucleotides.
[0029] The peptide or protein of interest can be a peptide or protein of eukaryotic or prokaryotic origin. The peptide or protein of interest can be a peptide or protein of human or non-human origin. In some embodiments, the peptide or protein of interest can be an antigenic protein, an antibody, or a protease, etc. In some embodiments, the peptide or protein of interest can be firefly luciferase, Gaussia luciferase, Renilla luciferase, green fluorescent protein, etc.
[0030] The translational functional element can be a translational initiation element, wherein the translational initiation element can be an internal ribosome entry site (IRES).
[0031] The internal ribosome entry site (IRES) can be an IRES derived from Coxsackie virus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), Human rhinovirus B6 (HRVB6), Coxsackie virus A (CVB1 / 2), Taura syndrome virus, Triatoma virus, Theiler's encephalomyelitis virus, Simian virus 40, Solenopsis invicta virus 1, Rhabditis pellio virus, Enamovirus-1, Human immunodeficiency virus type 1, Enamovirus-1, Boophilus microplus virus, Hepatitis C virus, Hepatitis A virus, Hepatitis GB virus, Foot-and-mouth disease virus, Human enterovirus 71, Equine rhinovirus, Euproctis pseudoconspersa virus, Encephalomyocarditis virus (EMCV), Drosophila C virus, Nicotiana virus, Laodelphax striatellus virus, Aedes pseudoscutellus virus, Aphid lethal paralysis virus, Avian encephalomyelitis virus, Acute bee paralysis virus, Rosa rugosa chlorotic ring spot virus, Swine fever virus, human FGF2, human SFTPA1, human AMM / RUNXl, Drosophila antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAPl, human c-myc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1 alpha, human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, Salivary virus, Coxsackievirus, Botha echovirus, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, Simian minute virus, Turnip crinkle virus, or eIF4G aptamer. In some embodiments, the IRES can be selected from an IRES of Coxsackie virus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), or Human rhinovirus B6 (HRVB6). In some embodiments, the IRES can be an IRES of Coxsackie virus B3 (CVB3). The sequence encoding the IRES can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 4. In some embodiments, the IRES can be an IRES of Enterovirus B107 (EVB107).The sequence encoding the IRES can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 61. In some embodiments, the IRES can be an IRES of human rhinovirus B3 (HRV B3). The sequence encoding the IRES can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 74. In some embodiments, the IRES can be an IRES of enterovirus A (EV-A). The sequence encoding the IRES can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 75 or SEQ ID NO: 76.
[0032] In some embodiments, the sequence of interest can comprise, from 5' end to 3' end, (a) a 3' end sequence of a sequence of an IRES, one or more open reading frames encoding a peptide or protein of interest, and a 5' end sequence of the sequence of the IRES, or (b) a 3' end sequence of the sequence of the IRES, the unique or multiple cloning site, and a 5' end sequence of the sequence of the IRES, wherein the IRES is an IRES of coxsackievirus B3 (CVB3) comprising the nucleotide sequence set forth in SEQ ID NO: 4, enterovirus B107 (EV-B107) comprising the nucleotide sequence set forth in SEQ ID NO: 61, human rhinovirus B3 (HRV B3) comprising the nucleotide sequence set forth in SEQ ID NO: 74, or enterovirus A (EV-A) comprising the nucleotide sequence set forth in SEQ ID NO: 75 or SEQ ID NO: 76. The 3' end sequence of the sequence encoding the IRES and the 5' end sequence of the sequence encoding the IRES comprise, respectively, (1) the nucleotide sequences set forth in SEQ ID NO: 26 and 27, (2) the nucleotide sequences set forth in SEQ ID NO: 40 and 41, (3) the nucleotide sequences set forth in SEQ ID NO: 46 and 47, (4) the nucleotide sequences set forth in SEQ ID NO: 55 and 56, (5) the nucleotide sequences set forth in SEQ ID NO: 62 and 63, or (6) the nucleotide sequences set forth in SEQ ID NO: 65 and 66.
[0033] In some embodiments, the open reading frame encoding the peptide or protein of interest is an open reading frame encoding a green fluorescent protein, wherein the open reading frame encoding the green fluorescent protein comprises the nucleotide sequence set forth in SEQ ID NO: 5. In some embodiments, the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, a sequence of a translational functional element (e.g., IRES), and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, wherein the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest can comprise AA and the nucleotide set forth in SEQ ID NO: 12, respectively. In some embodiments, the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, a sequence of a translational functional element (e.g., IRES), and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest can comprise AA and the nucleotide set forth in SEQ ID NO: 12, respectively.
[0034] In some embodiments, the open reading frame encoding the peptide or protein of interest is an open reading frame encoding a luciferase, wherein the open reading frame encoding the luciferase can comprise the nucleotide sequence set forth in SEQ ID NO: 28. In some embodiments, the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, a sequence of a translational functional element (e.g., IRES), and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, wherein the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest can comprise AA and the nucleotide sequence set forth in SEQ ID NO: 29, or SEQ ID NO: 32 and 33, respectively. In some embodiments, the sequence of interest can comprise, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, a sequence of a translational functional element (e.g., IRES), and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest can comprise AA and the nucleotide sequence set forth in SEQ ID NO: 29, or SEQ ID NO: 32 and 33, respectively.
[0035] In some embodiments, the 3' self-splicing intron fragment and the 5' self-splicing intron fragment can be derived from a self-splicing intron in the RecA gene of Bacillus anthracis, or a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii.
[0036] The single-stranded DNA molecule of the present application can also comprise a sequence of a 5' homology arm upstream or 5' to the 3' self-splicing intron fragment, and a sequence of a 3' homology arm downstream or 3' to the 5' self-splicing intron fragment. In some embodiments, the single-stranded DNA molecule can comprise, in order from 5' end to 3' end, a sequence of a 5' homology arm, a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, a 5' self-splicing intron fragment comprising a 5' splice site, and a sequence of a 3' homology arm. The 5' homology arm and the 3' homology arm can be complementary to pair, for example, to form at least about 15%, 25%, 35%, 45%, 55%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% base pairing. In some embodiments, the 5' homology arm and the 3' homology arm can be complementary to form a structure comprising one palindromic stem and one stem-loop, wherein the palindromic stem is connected to the stem of the stem-loop, wherein the 5' end of the 5' homology arm and the 3' end of the 3' homology arm are adjacent at the palindromic stem. In some embodiments, the 5' homology arm and the 3' homology arm can be complementary to form a structure comprising one palindromic stem and two stem-loops, wherein the palindromic stem is connected to the stems of the two stem-loops, wherein the stems of the two stem-loops are connected, wherein the 5' end of the 5' homology arm and the 3' end of the 3' homology arm are adjacent at the palindromic stem or the loop of one of the stem-loops. In some embodiments, the sequence of the 5' homology arm and the sequence of the 3' homology arm can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 2 and 7, respectively.
[0037] The single-stranded DNA molecule of the present application can also comprise a RNA polymerase promoter at the 5' end, for example, 5' to the sequence of the 3' self-splicing intron fragment or the 5' homology arm. The RNA polymerase promoter can be a RNA polymerase promoter derived from T7 virus, T6 virus, SP6 virus, T3 virus, or T4 virus. In some embodiments, the RNA polymerase promoter can be a T7 virus promoter, which can comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 1.
[0038] The single-stranded DNA molecule of the present application can also comprise a restriction enzyme site at the 3' end, for example, 3' to the sequence of the 5' self-splicing intron fragment or the 3' homology arm. The restriction enzyme site can be, for example, EcoRI, EcoRV.
[0039] The single-stranded DNA molecule of the present application can further comprise a 5' spacer sequence between the 3' self-splicing intron fragment and the sequence of interest, and / or a 3' spacer sequence between the sequence of interest and the 5' self-splicing intron fragment. In particular, the single-stranded DNA molecule of the present application can not comprise a 5' spacer sequence between the 3' self-splicing intron fragment and the sequence of interest, and / or a 3' spacer sequence between the sequence of interest and the 5' self-splicing intron fragment.
[0040] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a sequence of a 5' homology arm, a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, a 5' self-splicing intron fragment comprising a 5' splice site, and a sequence of a 3' homology arm, wherein each element is operably linked. In some embodiments, the single-stranded DNA molecule of the present application can consist of, in order from the 5' end to the 3' end: a sequence of a 5' homology arm, a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, a 5' self-splicing intron fragment comprising a 5' splice site, and a sequence of a 3' homology arm, wherein each element is operably linked. In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a sequence of a 5' homology arm, a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, a 5' self-splicing intron fragment comprising a 5' splice site, a sequence of a 3' homology arm, and a restriction enzyme site, wherein each element is operably linked. In some embodiments, the single-stranded DNA molecule of the present application can consist of, in order from the 5' end to the 3' end: a sequence of a 5' homology arm, a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, a 5' self-splicing intron fragment comprising a 5' splice site, a sequence of a 3' homology arm, and a restriction enzyme site, wherein each element is operably linked.
[0041] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site and a 3' exon, AA, a sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein (EGFP), and a 5' self-splicing intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-splicing intron fragment comprising a 3' splice site and a 3' exon, the sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein (EGFP), and the 5' self-splicing intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 3, 4, 12, and 9; or SEQ ID NOs: 18, 4, 12, and 20, respectively.
[0042] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, a sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein (EGFP), and a 5' exon and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein (EGFP), and the 5' exon and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 10, 4, 5, and 11, respectively; or SEQ ID NOs: 21, 4, 5, and 22, respectively.
[0043] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, AA, a sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein (EGFP), and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein (EGFP), and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 10, 4, 12, and 13, respectively; or SEQ ID NOs: 21, 4, 12, and 23, respectively.
[0044] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, AA, a sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein (EGFP), and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein (EGFP), and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 10, 4, 12, and 13, respectively; or SEQ ID NOs: 21, 4, 12, and 23, respectively.
[0045] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a 3' sequence encoding a CVB3 IRES, one or more open reading frames encoding a peptide or protein of interest, a 5' sequence encoding a CVB3 IRES, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site are derived from the 3' sequence and the 5' sequence of a self-splicing intron in the RecA gene of Bacillus anthracis or from the 3' sequence and the 5' sequence of a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii. In some embodiments, the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site can comprise the following nucleotide sequences: (1) SEQ ID NO: 10 and SEQ ID NO: 31; (2) SEQ ID NO: 21 and SEQ ID NO: 39; or (3) SEQ ID NO: 21 and SEQ ID NO: 48.
[0046] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a 3' sequence encoding an EVB107 IRES, one or more open reading frames encoding a peptide or protein of interest, a 5' sequence encoding an EVB107 IRES, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site are derived from the 3' sequence and the 5' sequence of a self-splicing intron in the RecA gene of Bacillus anthracis or from the 3' sequence and the 5' sequence of a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii. In some embodiments, the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site can comprise the following nucleotide sequences: (1) SEQ ID NO: 21 and SEQ ID NO: 57.
[0047] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a 3' sequence encoding a HRVB3 IRES, one or more open reading frames encoding a peptide or protein of interest, a 5' sequence encoding a HRVB3 IRES, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site are derived from the 3' sequence and the 5' sequence of a self-splicing intron in the RecA gene of Bacillus anthracis or the 3' sequence and the 5' sequence of a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii. In some embodiments, the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site can comprise the nucleotide sequences set forth in SEQ ID NO: 21 and SEQ ID NO: 64, respectively.
[0048] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a 3' sequence encoding an EV-A IRES, one or more open reading frames encoding a peptide or protein of interest, a 5' sequence encoding an EV-A IRES, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site are derived from the 3' sequence and the 5' sequence of a self-splicing intron in the RecA gene of Bacillus anthracis or the 3' sequence and the 5' sequence of a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii. In some embodiments, the 3' self-splicing intron fragment comprising a 3' splice site and the 5' self-splicing intron fragment comprising a 5' splice site can comprise the nucleotide sequences set forth in SEQ ID NO: 21 and SEQ ID NO: 67, respectively.
[0049] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a 3' sequence encoding a CVB3 IRES, a sequence encoding luciferase (Fluc), a 5' sequence encoding a CVB3 IRES, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-splicing intron fragment comprising a 3' splice site, the 3' sequence encoding a CVB3 IRES, the sequence encoding luciferase (Fluc), the 5' sequence encoding a CVB3 IRES, and the 5' self-splicing intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NO: 10, 26, 28, 27, and 31, respectively.
[0050] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, a 3' sequence encoding a luciferase (Fluc), a sequence encoding a CVB3 IRES, a 5' sequence encoding a luciferase (Fluc), and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the 3' sequence encoding a luciferase (Fluc), the sequence encoding a CVB3 IRES, the 5' sequence encoding a luciferase (Fluc), and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 10, 32, 4, 33, and 34, respectively.
[0051] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, a 3' sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein EGFP, a 5' sequence encoding a CVB3 IRES, and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the 3' sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein EGFP, the 5' sequence encoding a CVB3 IRES, and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 21, 40, 5, 41, and 39, respectively.
[0052] In some embodiments, the single-stranded DNA molecule of the present application can comprise, in order from the 5' end to the 3' end: a 3' self-cleaving intron fragment comprising a 3' splice site, a 3' sequence encoding a CVB3 IRES, a sequence encoding a green fluorescent protein EGFP, a 5' sequence encoding a CVB3 IRES, and a 5' self-cleaving intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-cleaving intron fragment comprising a 3' splice site, the 3' sequence encoding a CVB3 IRES, the sequence encoding a green fluorescent protein EGFP, the 5' sequence encoding a CVB3 IRES, and the 5' self-cleaving intron fragment comprising a 5' splice site comprise the nucleotide sequences set forth in SEQ ID NOs: 21, 40, 5, 41, and 39, respectively.
[0053] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding CVB3 IRES, a sequence encoding green fluorescent protein EGFP, a 5' sequence encoding CVB3 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding CVB3 IRES, the sequence encoding green fluorescent protein EGFP, the 5' sequence encoding CVB3 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 46, 5, 47, and 48.
[0054] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding CVB3 IRES, a sequence encoding luciferase (Fluc), a 5' sequence encoding CVB3 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding CVB3 IRES, the sequence encoding luciferase (Fluc), the 5' sequence encoding CVB3 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 46, 28, 47, and 48.
[0055] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding EVB107 IRES, a sequence encoding green fluorescent protein EGFP, a 5' sequence encoding EVB107 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding EVB107 IRES, the sequence encoding green fluorescent protein EGFP, the 5' sequence encoding EVB107 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO:21, 55, 5, 56, and 57.
[0056] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding EVB107 IRES, a sequence encoding Gaussian luciferase Gluc, a 5' sequence encoding EVB107 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding EVB107 IRES, the sequence encoding Gaussian luciferase Gluc, the 5' sequence encoding EVB107 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO:21, 55, 49, 56, and 57.
[0057] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding HRVB3 IRES, a sequence encoding Gaussian luciferase Gluc, a 5' sequence encoding HRVB3 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding HRVB3 IRES, the sequence encoding Gaussian luciferase Gluc, the 5' sequence encoding HRVB3 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 62, 49, 63, and 64.
[0058] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding HRVB3 IRES, a sequence encoding green fluorescent protein EGFP, a 5' sequence encoding HRVB3 IRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding HRVB3 IRES, the sequence encoding green fluorescent protein EGFP, the 5' sequence encoding HRVB3 IRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 62, 5, 63, and 64.
[0059] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding EV-AIRES, a sequence encoding Gaussian luciferase Gluc, a 5' sequence encoding EV-AIRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding EV-AIRES, the sequence encoding Gaussian luciferase Gluc, the 5' sequence encoding EV-AIRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 65, 49, 66, and 67.
[0060] In some embodiments, the single-stranded DNA molecule of this application may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a 3' sequence encoding EV-AIRES, a sequence encoding green fluorescent protein EGFP, a 5' sequence encoding EV-AIRES, and a 5' self-splicing intron fragment containing a 5' splice site, wherein each element is operatively linked, and the 3' self-splicing intron fragment containing a 3' splice site, the 3' sequence encoding EV-AIRES, the sequence encoding green fluorescent protein EGFP, the 5' sequence encoding EV-AIRES, and the 5' self-splicing intron fragment containing a 5' splice site respectively comprise nucleotide sequences as shown in SEQ ID NO: 21, 65, 5, 66, and 67.
[0061] This application also provides a double-stranded DNA molecule, which may comprise i) the single-stranded DNA molecule of this application, and ii) a second strand complementary to the single-stranded DNA molecule. In some embodiments, the double-stranded DNA molecule of this application may comprise i) the single-stranded DNA molecule of this application, and ii) a second strand completely complementary to the single-stranded DNA molecule. In some embodiments, the double-stranded DNA molecule is a double-stranded DNA molecule used for preparing circular RNA.
[0062] This application also provides a vector comprising the single-stranded DNA molecule or double-stranded DNA molecule of this application. In some embodiments, the vector may be a vector for preparing circular RNA, comprising the single-stranded DNA molecule of this application for preparing circular RNA, or the double-stranded DNA molecule for preparing circular RNA. The vector may be circular or linear. In some embodiments, the vector may be linear. In some embodiments, the vector may be circular and processed to become linear. The vector of this application can transcribe approximately 500 to approximately 10,000 nt of circular RNA.
[0063] In a second aspect, this application provides a circularizable RNA molecule that may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a target sequence, and a 5' self-splicing intron fragment containing a 5' splice site, wherein the elements are operatively linked.
[0064] The 3' and 5' self-splicing intron fragments enable the RNA molecule to circularize and be removed from the RNA molecule during the circularization process. The 3' and 5' self-splicing intron fragments can be derived from the same self-splicing intron, particularly a type I self-splicing intron. In some embodiments, the 3' and 5' self-splicing intron fragments can be derived from the self-splicing intron in the *Bacillus anthracis* recA gene or the self-splicing intron in the *Coxiella belladonna* 23S ribosome gene. When arranged in this order, the 5' and 3' self-splicing intron fragments can be joined to form a self-splicing intron with self-splicing function, such as a complete self-splicing intron. In some embodiments, the 3' and 5' self-splicing intron fragments can be derived from the self-splicing introns of the *Bacillus anthracis* RecA gene, and the 3' self-splicing intron fragment may contain the nucleotide sequence shown in SEQ ID NO:10. In some embodiments, the 3' and 5' self-splicing intron fragments can be derived from the self-splicing introns of the *Coxiella belladonna* 23S ribosome gene, and the 3' self-splicing intron fragment may contain the nucleotide sequence shown in SEQ ID NO:21.
[0065] 5' self-splicing intron fragments can contain internal guide sequences.
[0066] The 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence. The 5' end sequence of the internal guide sequence may be anticomplementary to any number of nucleotides at the 5' end of the target sequence. For example, the 5' end sequence of the internal guide sequence may be anticomplementary to 2-15 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nt) nucleotides at the 5' end of the target sequence, such as 2-12 nt nucleotides. For example, the 5' end sequence of the internal guide sequence may be anticomplementary to 3-6 nt (e.g., 3, 4, 5, or 6 nt) nucleotides at the 5' end of the target sequence. In some embodiments, the 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence, and the RNA molecule of this application may also include a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 5' side of the 5' self-splicing intron fragment. This naturally adjacent exon may contain a partial or truncated sequence. Naturally adjacent exons or sequences corresponding to naturally adjacent exons may contain nucleotides of length 3-20 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt), such as 3-10 nt. The single-stranded DNA molecules of this application may at least not contain naturally adjacent exons or sequences corresponding to naturally adjacent exons on the 3' side of the 3' self-splicing intron fragment.
[0067] The 3' end sequence of the internal guide sequence may be anticomplementary to the 3' end sequence of the target sequence. The 3' end sequence of the internal guide sequence may begin with a G-base nucleotide and be anticomplementary to the 3' end sequence of the target sequence. The 3' end sequence of the target sequence may end with a U-base nucleotide. The 3' end sequence of the internal guide sequence may be anticomplementary to any number of nucleotides at the 3' end of the target sequence. For example, the 3' end sequence of the internal guide sequence may be anticomplementary to 3-20 nt (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt) of the 3' end of the target sequence, for example, 3-15 nt nucleotides. In some embodiments, the 3' end sequence of the internal guide sequence may be anticomplementary to 3-10 nt (e.g., 3-9 nt) of the 3' end of the target sequence. In some embodiments, the 3' end sequence of the internal guide sequence may begin with a G-base nucleotide and be anticomplementary to the 3' end sequence of the target sequence, wherein the 3' end sequence of the target sequence may be a U-base nucleotide. The RNA molecule of this application may also contain a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon on the 3' side of the 3' self-splicing intron fragment. In some embodiments, the 3' end sequence of the internal guide sequence may begin with a G-base nucleotide and be anticomplementary to the 3' end sequence of the target sequence, wherein the 3' end sequence of the target sequence may be a U-base nucleotide. The G-base initiating nucleotide of the internal guide sequence can form a G:U pair with the U-base nucleotide at the 3' end sequence of the target sequence.
[0068] The naturally adjacent exon may contain a partial or truncated sequence. The naturally adjacent exon, or the sequence corresponding to the naturally adjacent exon, may contain nucleotides of length 3-20 nt (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nt), for example, 3-10 nt. The single-stranded DNA molecule of this application may at least not contain a naturally adjacent exon or the sequence corresponding to the naturally adjacent exon on the 5' side of the 5' self-splicing intron fragment.
[0069] In some embodiments, the 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence, and the 3' end sequence of the internal guide sequence may be anticomplementary to the 3' end sequence of the target sequence. In some embodiments, the 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence, the 3' end sequence of the internal guide sequence may begin with a G-base and be anticomplementary to the 3' end sequence of the target sequence, and the 3' end sequence of the target sequence may be a U-base. In some embodiments, the 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence, the 3' end sequence of the internal guide sequence may begin with a G-base and be anticomplementary to the 3' end sequence of the target sequence, and the 3' end sequence of the target sequence may be a U-base, where the G-base initiating nucleotide of the internal guide sequence can form a G:U pair with the U-base at the 3' end sequence of the target sequence. The RNA molecule of this application may not contain a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon at the 5' end of the 5' self-splicing intron fragment, and may not contain a naturally adjacent exon or a sequence corresponding to a naturally adjacent exon at the 3' end of the 3' self-splicing intron fragment.
[0070] The target sequence may contain an open reading frame encoding a target peptide or protein and a translational functional element. The translational functional element may be selected from translation initiation elements and translation enhancement elements. The translation initiation element may be a sequence that initiates RNA translation, such as an internal ribosome entry site (IRES). The translation enhancement element may be a sequence that enhances RNA translation, such as poly(A), poly(C), cap-independent translation enhancer (CITE), or EIF4 ligand family recognition sequences. In the circular RNA molecules prepared from circularizable RNA of this application, the translation enhancement element may be located on the 5' or 3' side of the open reading frame encoding the target peptide or protein. In the circular RNA molecules prepared from circularizable RNA of this application, the translation enhancement element may be located on the 5' or 3' side of the translation initiation element. In some embodiments, the translational functional element may include a translation initiation element. In some embodiments, the translational functional element may be a translation initiation element. In some embodiments, the translation initiation element may be an IRES.
[0071] The target sequence may contain an open reading frame encoding the target peptide or protein, and an internal ribosome entry site (IRES). The target sequence may include, from the 5' end to the 3' end, (a) the translational functional element (e.g., IRES) and the open reading frame encoding the target peptide or protein, (b) the open reading frame encoding the target peptide or protein and the translational functional element (e.g., IRES), (c) the 3' end sequence of the translational functional element (e.g., IRES), the open reading frame encoding the target peptide or protein, and the 5' end sequence of the translational functional element (e.g., IRES), wherein the 5' end sequence and the 3' end sequence of the translational functional element form the translational functional element when arranged in this order, or (d) the 3' end sequence of the open reading frame encoding the target peptide or protein, the translational functional element (e.g., IRES), and the 5' end sequence of the open reading frame encoding the target peptide or protein, wherein the 5' end sequence and the 3' end sequence of the open reading frame encoding the target peptide or protein form the open reading frame encoding the target peptide or protein when arranged in this order. In some embodiments, the target sequence may comprise one or more open reading frames encoding a target peptide or protein, and a translational functional element (e.g., IRES). The target sequence may comprise, from its 5' end to its 3' end, (a) the one or more open reading frames encoding the target peptide or protein, and the translational functional element (e.g., IRES); (b) the translational functional element (e.g., IRES), and the one or more open reading frames encoding the target peptide or protein; (c) the 3' end sequence of the translational functional element (e.g., IRES), the one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of the translational functional element (e.g., IRES); or (d) the 3' end sequence of one of the open reading frames encoding the target peptide or protein, other open reading frames encoding the target peptide or protein (if present), the translational functional element (e.g., IRES), and the 5' end sequence of one of the open reading frames encoding the target peptide or protein.
[0072] The target sequence may contain non-coding RNA. The target sequence may (a) contain the non-coding RNA, or (b) contain the 3' end sequence of the non-coding RNA and the 5' end sequence of the non-coding RNA from the 5' end to the 3' end, wherein the 5' end sequence of the non-coding RNA and the 3' end sequence of the non-coding RNA form the non-coding RNA sequence when arranged in this order.
[0073] The target sequence may contain an open reading frame (OPF) encoding a target peptide or protein. The target sequence may (a) contain the OPF encoding the target peptide or protein, or (b) contain, from the 5' end to the 3' end, the 3' end sequence of the OPF encoding the target peptide or protein, and the 5' end sequence of the OPF encoding the target peptide or protein, wherein the 5' end sequence of the OPF encoding the target peptide or protein and the 3' end sequence of the OPF encoding the target peptide or protein, when arranged in this order, form the OPF encoding the target peptide or protein. In some embodiments, the target sequence may contain one or more OPF encodings of a target peptide or protein. The target sequence may contain, from the 5' end to the 3' end, (a) the one or more OPF encodings of a target peptide or protein, or (b) the 3' end sequence of one of the OPF encodings of a target peptide or protein, other OPF encodings of a target peptide or protein (if present), and the 5' end sequence of one of the OPF encodings of a target peptide or protein.
[0074] The target sequence can be 50-10000 nucleotides in length. Alternatively, it can be 50-8000 nucleotides. For example, it can be 200, 500, 1000, 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 nucleotides.
[0075] The target peptide or protein can be of eukaryotic or prokaryotic origin. It can be human or non-human. In some embodiments, the target peptide or protein can be an antigen protein, antibody, or Cas9 endonuclease, etc. In some embodiments, the target peptide or protein can be firefly luciferase, long-horned luciferase, Gaussian luciferase, green fluorescent protein, etc.
[0076] Internal ribosome entry sites (IRES) can originate from Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), Human rhinovirus B6 (HRVB6), Coxsackievirus A (CVB1 / 2), Taura syndrome virus, blood-sucking assassin bug virus, Tyrell's encephalomyelitis virus, simian virus 40, red imported fire ant virus 1, rice constrictor aphid virus, reticuloendotheliosis virus, and Forman poliovirus. Toxin 1, Soybean Looper Virus, Kashmir Bee Virus, Human Rhinovirus 2, Glass Leafhopper Virus-1, Human Immunodeficiency Virus Type 1, Glass Leafhopper Virus-1, Lice P Virus, Hepatitis C Virus, Hepatitis A Virus, GB Hepatitis Virus, Foot-and-Mouth Disease Virus, Human Enterovirus 71, Equine Rhinovirus, Tea Looper-like Virus, Encephalomyelitis Virus (EMCV), Fruit Fly C Virus, Cruciferous Tobacco Virus, Cricket Paralysis Virus, Bovine Viral Diarrhea Virus 1, Black Queen Cell Virus, Aphid Lethal Paralysis Virus, Avian Encephalomyelitis Virus Acute bee paralysis virus, Hibiscus rosa-spot virus, classical swine fever virus, human FGF2, human SFTPA1, human AML1 / RUNX1, Drosophila antennae and legs, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAPl, human c-myc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, fruit fly reaper, canine scamper, fruit fly UBX, salivary virus, Coxsackievirus, bi-echovirus, human UNR, mouse UtrA, human VEGF-A, human XIAP, fruit fly hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, simian microRNA virus, turnip shrunkenness virus, or IRES of eIF4G aptamers. In some embodiments, the IRES may be selected from IRES of Coxsackievirus B3 (CVB3), enterovirus B107 (EVB107), human rhinovirus B3 (HRVB3), or human rhinovirus B6 (HRVB6). In some embodiments, the IRES may be an IRES of Coxsackievirus B3 (CVB3). The sequence encoding IRES may comprise a nucleotide sequence having at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO:4. In some embodiments, IRES may be an IRES of enterovirus B107 (EVB107).The sequence encoding IRES may comprise a nucleotide sequence having at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO:61. In some embodiments, IRES may be an IRES of human rhinovirus B3 (HRVB3). The sequence encoding IRES may comprise a nucleotide sequence having at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO:74. In some embodiments, IRES may be an IRES of enterovirus A (EV-A). The sequence encoding IRES may include a nucleotide sequence having at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO:75.
[0077] In some embodiments, the target sequence may comprise, from the 5' end to the 3' end, (a) the 3' end sequence of IRES, one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of IRES, wherein IRES is the IRES of Coxsackievirus B3 (CVB3) comprising the nucleotide sequence shown in SEQ ID NO:4. The 3' end sequence encoding IRES and the 5' end sequence encoding IRES comprise the nucleotide sequences shown in SEQ ID NO:26 and 27, respectively, the nucleotide sequences shown in SEQ ID NO:40 and 41, respectively, or the nucleotide sequences shown in SEQ ID NO:46 and 47, respectively.
[0078] In some embodiments, the target sequence may comprise, from the 5' end to the 3' end, (a) the 3' end sequence of IRES, one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of IRES, wherein the IRES is the IRES of EVB107 comprising the nucleotide sequence shown in SEQ ID NO:61. The 3' end sequence encoding the IRES and the 5' end sequence encoding the IRES comprise the nucleotide sequences shown in SEQ ID NO:55 and 56, respectively.
[0079] In some embodiments, the target sequence may comprise, from the 5' end to the 3' end, (a) the 3' end sequence of IRES, one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of IRES, wherein IRES is the IRES of HRVB3 comprising the nucleotide sequence shown in SEQ ID NO:74. The 3' end sequence encoding IRES and the 5' end sequence encoding IRES comprise the nucleotide sequences shown in SEQ ID NO:62 and 63, respectively.
[0080] In some embodiments, the target sequence may comprise, from the 5' end to the 3' end, (a) the 3' end sequence of the IRES, one or more open reading frames encoding the target peptide or protein, and the 5' end sequence of the IRES, wherein the IRES is the IRES of EV-A comprising the nucleotide sequence shown in SEQ ID NO:75. The 3' end sequence encoding the IRES and the 5' end sequence encoding the IRES comprise the nucleotide sequences shown in SEQ ID NO:65 and 66, respectively.
[0081] In some embodiments, the open reading frame encoding the target peptide or protein is an open reading frame encoding green fluorescent protein, wherein the open reading frame encoding green fluorescent protein comprises the nucleotide sequence shown in SEQ ID NO:5. In some embodiments, the target sequence may comprise, from the 5' end to the 3' end, the 3' end sequence encoding the target peptide or protein, IRES, and the 5' end sequence encoding the target peptide or protein, wherein the 3' end sequence encoding the target peptide or protein and the 5' end sequence encoding the target peptide or protein may respectively comprise the nucleotides shown in SEQ ID NO:12.
[0082] In some embodiments, the open reading frame encoding the target peptide or protein may be an open reading frame encoding luciferase, wherein the open reading frame encoding luciferase may contain the nucleotide sequence shown in SEQ ID NO:28. In some embodiments, the target sequence may include, from the 5' end to the 3' end, the 3' end sequence encoding the target peptide or protein, IRES, and the 5' end sequence encoding the target peptide or protein, wherein the 3' end sequence encoding the target peptide or protein and the 5' end sequence encoding the target peptide or protein may respectively contain AA and the nucleotide sequences shown in SEQ ID NO:29, or SEQ ID NO:32 and 33.
[0083] The RNA molecule of this application may also include a 5' homologous arm upstream of or 5' to the 3' self-splicing intron fragment, and a 3' homologous arm downstream of or 3' to the 5' self-splicing intron fragment. In some embodiments, the RNA molecule may sequentially include a 5' homologous arm, a 3' self-splicing intron fragment containing a 3' splice site, a target sequence, a 5' self-splicing intron fragment containing a 5' splice site, and a 3' homologous arm from the 5' end to the 3' end. The 5' homologous arm and the 3' homologous arm may be complementary, for example, forming at least about 15%, 25%, 35%, 45%, 55%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% base pairing. In some embodiments, the 5' and 3' homologous arms can complementarily form a structure containing a palindromic stem and a stem-loop, wherein the palindromic stem is connected to the stem portion of the stem-loop, and the 5' end of the 5' homologous arm and the 3' end of the 3' homologous arm are adjacent at the palindromic stem portion. In some embodiments, the 5' and 3' homologous arms can complementarily form a structure containing a palindromic stem and two stem-loops, wherein the palindromic stem is connected to the stem portions of two stem-loops, and the stem portions of two stem-loops are connected, wherein the 5' end of the 5' homologous arm and the 3' end of the 3' homologous arm are adjacent at the loop portion of the palindromic stem or one of the stem-loops. In some embodiments, the 5' and 3' homologous arms can each comprise a nucleotide sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO:2 and 7.
[0084] The RNA molecule of this application may include a 5' spacer sequence between the 3' self-splicing intron fragment and the target sequence, and / or include a 3' spacer sequence between the target sequence and the 5' self-splicing intron fragment. In particular, the RNA molecule of this application may not include a 5' spacer sequence between the 3' self-splicing intron fragment and the target sequence, and / or may not include a 3' spacer sequence between the target sequence and the 5' self-splicing intron fragment.
[0085] In some embodiments, the RNA molecule of this application may be transcribed from the complementary strand of a single-stranded DNA molecule, a double-stranded DNA molecule, or a vector of the first aspect of this application.
[0086] In a third aspect, this application provides a method for preparing circular RNA, comprising i) in vitro transcription of RNA molecules from the complementary strand of a single-stranded DNA molecule, a double-stranded DNA molecule, or a vector of the first aspect of this application under suitable conditions, and ii) incubating the RNA molecule under suitable conditions, wherein the suitable conditions for step ii) include the presence of magnesium ions and guanosine triphosphate (GTP). In some embodiments, suitable conditions include the presence of guanosine triphosphate and magnesium...2+ In the presence of [specific ingredient], incubate at approximately 37°C for approximately 2 hours. In some embodiments, suitable conditions include [specific conditions] in the presence of guanosine triphosphate and Mg [specific ingredient]. 2+ In the presence of [specific ingredient], incubate at approximately 37°C for about 2 hours, and then add DNase I and incubate at approximately 37°C for about 30 minutes. The concentration of GTP can be 2 mM, and the concentration of magnesium ions can be 10 mM-20 mM.
[0087] Alternatively, the method for preparing circular RNA according to this application may include incubating the circularizable RNA molecule of the second aspect of this application under suitable conditions, wherein suitable conditions include the presence of magnesium ions and guanosine triphosphate (GTP). In some embodiments, suitable conditions include the presence of guanosine triphosphate and magnesium... 2+ In the presence of [specific ingredient], incubate at approximately 37°C for approximately 2 hours. In some embodiments, suitable conditions include [specific conditions] in the presence of guanosine triphosphate and Mg [specific ingredient]. 2+ In the presence of [specific ingredient], incubate at approximately 37°C for about 2 hours, and then add DNase I and incubate at approximately 37°C for about 30 minutes. The concentration of GTP can be 2 mM, and the concentration of magnesium ions can be 10 mM-20 mM.
[0088] This application also protects circular RNA prepared from single-stranded DNA molecules, double-stranded DNA molecules or vectors of the first aspect of this application or circularizable RNA molecules of the second aspect of this application, as well as circular RNA prepared by the method for preparing circular RNA of this application.
[0089] Circular RNA molecules can be translated into proteins within cells or exist as biologically active non-coding RNAs. The length of circular RNA can be at least 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, or 9500 nucleotides.
[0090] In a fourth aspect, this application provides a circular RNA obtained according to a method for preparing circular RNA according to a third aspect of this application. In some embodiments, the circular RNA may not contain exons adjacent to either the 3' self-splicing intron fragment or the 5' self-splicing intron fragment. In other embodiments, the circular RNA may not contain exons adjacent to either the 3' self-splicing intron fragment or the 5' self-splicing intron fragment.
[0091] This application also provides a circular RNA, which may comprise a target sequence, a selectively translatable functional element, and exons selectively adjacent to either a 3' self-splicing intron fragment or a 5' self-splicing intron fragment. In some embodiments, the circular RNA may comprise a target sequence, a translatable functional element, and exons adjacent to a 3' self-splicing intron fragment. In some embodiments, the circular RNA may comprise a target sequence, a translatable functional element, and exons adjacent to a 5' self-splicing intron fragment. In some preferred embodiments, the circular RNA may consist of a target sequence and a translatable functional element. In other embodiments, the circular RNA may comprise a target sequence and exons adjacent to a 3' self-splicing intron fragment. In some embodiments, the circular RNA may comprise a target sequence and exons adjacent to a 5' self-splicing intron fragment. In some preferred embodiments, the circular RNA may consist only of the target sequence. In some preferred embodiments, the circular RNA may consist only of the target sequence and a translatable functional element; or consist only of the target sequence.
[0092] The translational functional element can be a translation initiation element, preferably an internal ribosome entry site (IRES). This IRES can be selected from IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human Rhinovirus B3 (HRVB3), Enterovirus A (EV-A), or Human Rhinovirus B6 (HRVB6). IRES from other sources described in this application are also applicable.
[0093] In a fifth aspect, this application provides a composition that may comprise the circular RNA molecule mentioned in the fourth aspect of this application and its open circular RNA, wherein the molar percentage of the circular RNA in the composition is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80%. In some embodiments, the molar percentage of the circular RNA in the composition may be at least 40%. In some embodiments, the molar percentage of the circular RNA in the composition may be detected by capillary electrophoresis.
[0094] The open circular RNA of this circular RNA can be any RNA whose closed circular structure is transformed into a linear structure by physical, chemical, or other factors.
[0095] The composition may also contain a circularizable RNA molecule used in the second aspect of this application for preparing the circular RNA. The composition may also contain RNA with residual intron sequences.
[0096] In a sixth aspect, this application provides a host cell that may contain the single-stranded DNA molecule, double-stranded DNA molecule, vector, circularizable RNA molecule, or circular RNA molecule of this application.
[0097] This application also provides a composition that may comprise the single-stranded DNA molecule, double-stranded DNA molecule, vector, or circularizable RNA molecule of this application, or a host cell containing the single-stranded DNA molecule, double-stranded DNA molecule, vector, or circularizable RNA molecule of this application. This composition can be used to prepare circularizable RNA and / or circular RNA, particularly to generate translatable proteins or biologically active circular RNA in vitro or in vivo. Biologically active circular RNA may be, for example, miRNA sponges or non-coding RNA. The composition of this application (comprising single-stranded DNA molecules and double-stranded DNA molecules) can be used to prepare a vector for preparing circular RNA.
[0098] This application also provides a composition that may comprise the circular RNA molecule of this application or a host cell containing the circular RNA molecule of this application. The composition may be a pharmaceutical composition and may also comprise a pharmaceutically acceptable carrier. The composition may be transfected into cells via, for example, liposome transfection, electroporation, or encapsulation with a nanocarrier.
[0099] This application also provides the use of the composition in the preparation of circular RNA or for in vivo therapy. For example, in one embodiment, this application provides a method for expressing a vaccine, therapeutic protein, or other type of protein in a subject in need, comprising administering the composition of this application containing circular RNA to the subject. The therapeutic protein may be, for example, an antibody, a fusion protein, etc.
[0100] In a seventh aspect, this application provides a method for treating or preventing a disease in a subject in need, comprising administering to the subject a pharmaceutical composition comprising a circular RNA molecule of this application. The circular RNA molecule comprises an open reading frame encoding a target peptide or protein. The target peptide or protein may be a disease-associated antigen or a therapeutic agent. Disease-associated antigens may be peptides or proteins located on the surface of microorganisms, such as viruses, bacteria, mycoplasma, etc., or tumor-associated antigens. The therapeutic agent may be, for example, an antibody.
[0101] When the target peptide or protein is a peptide or protein on the surface of a microorganism, such as a virus, bacteria, mycoplasma, etc., the method of this application can be used to treat or prevent diseases related to infection by that microorganism.
[0102] When the target peptide or protein is a tumor-associated antigen, or a protein such as an antibody that targets a tumor-associated antigen, the method of this application can be used to treat tumors associated with that tumor-associated antigen.
[0103] The target peptide or protein can also be a normal protein expressed in mammals, such as humans, which can be used to supplement subjects who lack this normal protein.
[0104] The subjects can be mammals, such as humans. In this application, the same nucleotide sequence, such as the nucleotide sequence represented by the same SEQ ID NO, can represent both a DNA sequence and an RNA sequence, the only difference being the substitution of T and U.
[0105] Other features and advantages disclosed herein will become readily apparent from the following detailed description and embodiments, which should not be construed as limiting. All references, Genbank registration numbers, patents, and published patent applications cited in this specification are incorporated herein by reference. Attached Figure Description
[0106] The following detailed description, given by way of example but not intended to limit the invention to the specific embodiments described, can be better understood in conjunction with the accompanying drawings.
[0107] Figure 1 shows a schematic diagram of the secondary structure of a type I self-splicing intron, where IGS is the internal guide sequence, and its 5' and 3' end sequences are complementary to the exons adjacent to the 3' and 5' sides of the intron, respectively. The triangle points to the 5' and 3' splicing sites of the intron.
[0108] Figures 2A and 2B show schematic structures of self-splicing introns in the Bacillus anthracis recA gene (2A) and nucleotide sites in the IGS that can be modified (2B).
[0109] Figures 3A-3D show schematic diagrams of vectors used to prepare circular RNA containing exons E3 and E5 adjacent to the self-splicing intron of the Bacillus anthracis recA gene (3A), containing only exon E3 adjacent to the self-splicing intron of the Bacillus anthracis recA gene (3B), containing only exon E5 adjacent to the self-splicing intron of the Bacillus anthracis recA gene (3C), and not containing exons E3 and E5 adjacent to the self-splicing intron of the Bacillus anthracis recA gene (3D).
[0110] Figures 4A-4D show the E-Gel of RNA molecules prepared using a vector containing only the E3 exon adjacent to the self-splicing intron in the Bacillus anthracis recA gene. TM EX gel electrophoresis images (4A) and capillary gel electrophoresis images (4C), and E-Gel images of RNA molecules prepared using a vector containing only the E5 exon adjacent to the self-splicing intron in the anthrax recA gene. TM EX gel electrophoresis image (4B) and capillary gel electrophoresis image (4D), in which each E-GelTM The three lanes in the EX gel electrophoresis image, from left to right, represent the RNA products obtained from in vitro transcription, RNA circularization, and RNA enzyme R treatment, respectively.
[0111] Figure 5 shows a schematic structure of the self-splicing intron in the Coxione 23S ribosomal gene.
[0112] Figures 6A-6D show schematic diagrams of vectors used to prepare circular RNA containing exons E3 and E5 adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene (6A), containing only exon E3 adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene (6B), containing only exon E5 adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene (6C), and not containing exons E3 and E5 adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene (6D).
[0113] Figures 7A-7D show the E-Gel of RNA molecules prepared using a vector containing only the E3 exon adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene. TM EX gel electrophoresis images (7A) and capillary gel electrophoresis images (7C), as well as E-Gel images of RNA molecules prepared using a vector containing only the E5 exon adjacent to the self-splicing intron of the Coxsell 23S ribosomal gene. TM E-gel electrophoresis image (7B) and capillary gel electrophoresis image (7D), in which each E-Gel TM The three lanes in the EX gel electrophoresis image, from left to right, represent the RNA products obtained from in vitro transcription, RNA circularization, and RNA enzyme R treatment, respectively.
[0114] Figures 8A-8C show schematic diagrams of vectors that do not contain E3 and E5 and contain IRES and open reading frames (ORF) sequentially between the 3' and 5' introns of the self-splicing introns in the Bacillus anthracis recA gene (8A), 3'IRES, ORF, and 5'IRES (8B), and 3'ORF, IRES, and 5'ORF (8C).
[0115] Figures 9A-9F show the E-Gel of RNA molecules prepared from vectors containing CVB3 IRES and open reading frames (ORFs) sequentially between the 3' and 5' intron fragments of the anthrax recA gene, which do not contain E3 and E5. TME-Gel electrophoresis image (9A) and capillary gel electrophoresis image (9D) of RNA molecules prepared from vectors containing 3'CVB3 IRES, ORF, and 5'CVB3 IRES sequentially between the 3' and 5' intron fragments, without E3 and E5. TM EX gel electrophoresis image (9B) and capillary gel electrophoresis image (9E), and E-Gel images of RNA molecules prepared from vectors that do not contain E3 and E5 and contain 3' ORF, CVB3 IRES and 5' ORF sequentially between the 3' and 5' intron fragments. TM EX gel electrophoresis pattern (9C) and capillary gel electrophoresis pattern (9F).
[0116] Figures 10A and 10B show the E-Gel of RNA molecules prepared from vectors containing 3' CVB3 IRES, ORF, and 5' CVB3 IRES sequentially between the 3' and 5' intron fragments of the self-splicing introns in the Coxsell 23S ribosomal gene, which do not contain E3 and E5. TM EX gel electrophoresis image (10A) and capillary gel electrophoresis image (10B).
[0117] Figures 11A-11I show the E-Gel of RNA molecules prepared from vectors containing no E3 and E5, consisting of a 3' CVB3 IRES, an ORF, and a 5' CVB3 IRES sequentially between the 3' and 5' introns of the self-splicing introns in the Coxsell 23S ribosomal gene, with the CVB3 IRES truncated after the 18th base. TM EX gel electrophoresis images (11A: Coxie-B18-Gluc; B: Coxie-B18-EGFP; 11C: Coxie-B18-Fluc) and capillary gel electrophoresis images (11D, 11E: Coxie-B18-Gluc; 11F, 11G: Coxie-B18-EGFP; 11H, 11I: Coxie-B18-Fluc), where each E-Gel TM The three lanes in the EX gel electrophoresis image, from left to right, represent the RNA products obtained from in vitro transcription, RNA circularization, and RNA enzyme R treatment, respectively.
[0118] Figures 12A-12D show the E-Gel of RNA molecules prepared from vectors containing 3' EVB107 IRES, ORF, and 5' EVB107 IRES sequentially between the 3' and 5' introns of the self-splicing introns in the Coxsell 23S ribosomal gene, with the EVB107 IRES truncated after the 18th base. TMEX gel electrophoresis images (12A: Coxie-EVB107-B18-Gluc; 12B: Coxie-EVB107-B18-EGFP) and capillary gel electrophoresis images (12C: Coxie-EVB107-B18-Gluc; 12D: Coxie-EVB107-B18-EGFP). Each E-Gel... TM The three lanes in the EX gel electrophoresis image, from left to right, represent the RNA products obtained from in vitro transcription, RNA circularization, and RNA enzyme R treatment, respectively.
[0119] Figures 13A-13C show the cell expression and immunogenicity test results of linear EGFP, circular Ana-EGFP, and circular coxie-B18-EGFP RNA prepared in the embodiments of this application. Figures 13A and 13B show the fluorescence intensity of HEK293T cells and A549 cells after 24 hours, 48 hours, 72 hours, 6 days (6D), and 8 days (8D), respectively. Figure 13C shows the immunogenicity test results of circular coxie-B18-EGFP transfected into A549 cells 48 hours later.
[0120] Figures 14A-14H show the E-Gel of RNA molecules prepared from vectors containing no E3 and E5, containing 3' HRVB3 IRES, ORF, and 5' HRVB3 IRES sequentially between the 3' and 5' intron fragments of the self-splicing intronic introns in the Coxsell 23S ribosomal gene, with the HRVB3 IRES truncated after the 18th base. TM EX gel electrophoresis images (14A: Coxie-HRVB3-B18-Gluc, 14B: Coxie-HRVB3-B18-EGFP) and capillary gel electrophoresis images (14E: Coxie-HRVB3-B18-Gluc, 14F: Coxie-HRVB3-B18-EGFP); E-Gel images of RNA molecules prepared from vectors containing no E3 and E5, containing 3' EV-A-mut IRES, ORF, and 5' EV-A-mut IRES sequentially between the 3' and 5' introns of the self-splicing introns in the Coxsell 23S ribosomal gene, with the EV-A-mut IRES truncated after the 374th base. TMEX gel electrophoresis images (14C: Coxie-EV-A-mut-B374-Gluc, 14D: Coxie-EV-A-mut-B374-EGFP) and capillary gel electrophoresis images (14G: Coxie-EV-A-mut-B374-Gluc, 14H: Coxie-EV-A-mut-B374-EGFP). Each E-Gel... TM The three lanes in the EX gel electrophoresis image, from left to right, represent the RNA products obtained from in vitro transcription, RNA circularization, and RNA enzyme R treatment, respectively. Detailed Implementation
[0121] Unless otherwise specified, the terms used herein have their common meanings as found in dictionaries, textbooks, and technical reference books, or as commonly understood by those skilled in the art. The following descriptions of some terms are for the purpose of understanding this application only and are not intended to impose any particular limitations on these terms, unless otherwise specified.
[0122] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include the plural form of the object referred to, unless the context clearly specifies otherwise.
[0123] The term "or" refers to a single element among the listed selectable elements, unless the context explicitly indicates otherwise.
[0124] The terms "comprising" or "including" mean that the stated elements, integers, or steps are included, but do not exclude the inclusion of any other elements, integers, or steps. In this document, when the terms "comprising" or "including" are used, unless otherwise specified, they also cover combinations of the stated elements, integers, or steps. The terms "consisting of" or "comprises of" generally mean that only the stated elements, integers, or steps are included, without the addition of other elements, integers, or steps.
[0125] The term "operably linked" means that the arrangement of the elements enables the completion of the required function. For example, the DNA strands in which the elements are "operably linked" in this application can perform RNA transcription in vitro or in vivo, and the RNA molecules in which the elements are "operably linked" in this application can perform circularization, etc.
[0126] The 5' end of a nucleic acid molecule can be a terminal with a free phosphate group, and the 3' end can be a terminal with a free hydroxyl group. "Upstream" generally refers to a position relatively closer to the 5' end in the nucleic acid sequence, while "downstream" generally refers to a position relatively closer to the 3' end in the nucleic acid sequence.
[0127] "In vitro transcription" or "IVT" refers to the process of forming RNA using DNA as a template in a cell-free system under conditions containing RNA transcriptase, NTPs, etc., mimicking the in vivo transcription process. When using a plasmid vector as a DNA template, the plasmid line is linearized by enzyme digestion sites before in vitro transcription.
[0128] "Precursor RNA", "IVT transcript", "messenger RNA", "mRNA" or "circularizable RNA" refers to the RNA product transcribed from the vector of this application, which can be circularized under suitable conditions to form circular RNA.
[0129] Exons are sequences in DNA that appear on mature RNA molecules. Exons are separated by introns and are joined together after transcription by intron removal. Conversely, introns are sequences that separate exons; they are transcribed into precursor RNA, removed by splicing, and ultimately not displayed in mature RNA molecules. Intron removal typically requires the participation of splice bodies, ribonucleoprotein complexes dynamically composed of nuclear small RNAs (snRNAs, U1, U2, U4, U5, U6, etc.) and protein factors (approximately 100 types), which recognize splice sites on the RNA precursor and catalyze the splicing reaction. However, a very small number of introns that form ribozymes undergo self-splicing, functioning as spliceases in RNA alone.
[0130] In this application, "self-splicing introns" refers to introns that can splice themselves without the involvement of splice bodies. Conditions for splicing of self-splicing introns include the presence of GTP and magnesium ions. In some embodiments, the conditions for splicing of self-splicing introns include incubating the purified RNA product after in vitro transcription at approximately 70°C for 5 minutes, placing it on ice for 3 minutes, and adding GTP and magnesium ions. 2+ Incubate at approximately 55°C for 5-10 minutes. Based on their structure and splicing mechanism, self-splicing introns are generally classified into two types: Type I and Type II.
[0131] In this document, "complementary," "pairing," "complementary pairing," or "base pairing" refers to the ability of two nucleotides or two bases to pair and bind according to the AT, AU, CG base complementarity principle. When one nucleotide sequence is "complementary" to another, it can mean that the two nucleotide sequences are 100% complementary, or that they are highly complementary, such as more than 90% complementary. The two "complementary" strands are each other's "complementary strands." "Anti-complementarity" refers to the ability of two fragments or two nucleotide chains on the same nucleotide chain to pair complementarily when read from the 5' to 3' direction and from the 3' to 5' direction, respectively. For example, the GGAAA fragment and the UUUCC fragment on the same nucleotide chain are anti-complementary, or the GGAAA fragment on one nucleotide chain and the UUUCC fragment on another nucleotide chain are anti-complementary. In particular, in this application, the GGAAAU fragment of the P1 sequence on the same nucleotide chain is anti-complementary to the GUUUCC fragment of the P1 antisense strand.
[0132] In this article, a "promoter" refers to a sequence that controls transcriptomic synthesis by providing recognition and binding sites for RNA polymerase. The promoter region may also include recognition or binding sites for other factors involved in transcriptional regulation. Promoters can be inducible, initiating transcription in response to an induction signal, or constitutive. Inducible promoters, in the absence of an induction signal, elicit very little or no transcription.
[0133] An "open reading frame" or "ORF" is a continuous sequence of nucleotides that begins with a start codon and ends with a stop codon, encoding a complete polypeptide chain. In an mRNA sequence, every three consecutive nucleotides (i.e., a triplet "codon") encode a corresponding amino acid. There is one start codon (AUG) and three stop codons (UAA, UAG, and UGA). The ribosomes begin translation from the start codon, synthesizing the polypeptide chain along the mRNA sequence and continuously elongating it. The elongation process terminates when a stop codon is encountered.
[0134] In this article, "translational functional elements" refers to sequences in RNA molecules that function in their own translation. These include sequences that promote translation, such as sequences that initiate translation (referred to as "translation initiation elements" in this article), like IRES, or sequences that enhance translation (referred to as "translation enhancement elements" in this article), such as poly(A) tails and cap-independent translation enhancers (CITEs). A poly(A) tail can contain consecutive adenosine nucleotides or consist of consecutive adenosine nucleotides. Alternatively, a poly(A) tail can contain 2-5 consecutive adenosine nucleotide segments separated by spacer sequences, where the spacer sequences contain 1-20 nucleotides and each consecutive adenosine nucleotide segment contains 10-100 consecutive adenosine nucleotides. A cap-independent translation enhancer (CITE) is a sequence that enables efficient RNA translation without relying on a cap, by recruiting translation initiation factors or ribosomal subunits. An "internal ribosome entry site" or "IRES" refers to an RNA sequence that forms a secondary structure to attract the precursor of the translation initiation complex to the translation initiation codon, such as AUG. The IRES from different sources mentioned herein include wild-type and corresponding variants, such as IRES derived from CVB3, EVB107, HRB3, or EV-A viruses, including but not limited to wild-type IRES and mutants of wild-type IRES derived from these viruses. Mutants of wild-type IRES can be IRES assembled from different functional regions of wild-type IRES, or IRES containing nucleotide substitutions, deletions, or insertions relative to wild-type IRES. In some embodiments, EV-AIRES comprise the nucleotide sequence shown in SEQ ID NO:76, and EV-AIRES mutants comprise a nucleotide sequence having at least 70% sequence identity with SEQ ID NO:76. In one specific embodiment, EV-AIRES mutants comprise the nucleotide sequence shown in SEQ ID NO:75.
[0135] In this application, terms such as "self-splicing intron," "open reading frame," and "translational element" are used interchangeably in both DNA and RNA molecules, without distinguishing between the different nucleotide sequence types after DNA molecules are transcribed into RNA molecules.
[0136] A "homologous arm" refers to an RNA sequence that can form at least about 15%, 25%, 35%, 45%, 55%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% base pairings with another RNA (e.g., another homologous arm).
[0137] "Spacer sequences" are sequences added to prevent interference between adjacent, nearby, or even distant elements during transcription, RNA folding, and other processes. Whether vector elements will interfere with each other during transcription, RNA folding, etc., can be predicted using computer software such as RNA folding software (e.g., RNAFold). Spacer sequences can also be designed using computer software such as RNA folding software (e.g., RNAFold).
[0138] The term "identity" or "sequence identity" as used herein refers to the percentage of nucleotides / amino acids in a sequence that are identical to those in a reference sequence after sequence alignment. If necessary, spaces are introduced in the sequence alignment to achieve the maximum percentage of sequence similarity between the two sequences. Those skilled in the art can use various methods, such as computer software, to perform pairwise or multiple sequence alignments to determine the percentage of sequence similarity between two or more nucleic acid or amino acid sequences. Such computer software includes, for example, ClustalOmega, T-coffee, Kalign, and MAFFT.
[0139] The term “subject” includes any human or non-human animal. The term “non-human animal” includes all vertebrates, such as mammals and non-mammalians, such as non-human primates, sheep, dogs, cats, cattle, horses, chickens, amphibians, and reptiles, although mammals, such as non-human primates, sheep, dogs, cats, cattle, and horses, are preferred.
[0140] The term "vector" refers to a naturally occurring or synthetic nucleotide fragment, such as a DNA or RNA fragment, including single-stranded and double-stranded DNA fragments, such as chemically synthesized DNA fragments, natural plasmids, or modified viral genomes. A foreign DNA fragment may be inserted into a vector for the cloning and / or expression of that foreign DNA fragment. Vectors may contain, for example, origins of replication, selectivity markers or reporter genes, multiple cloning sites (MCS), etc. The term includes linear DNA fragments (such as PCR products, linearized plasmid fragments, etc.), plasmid vectors, viral vectors, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc. When the vector is double-stranded DNA, the description of the element sequencing and the orientation of the element sequence refers to one of the DNA strands. In a double-stranded vector, the two DNA strands are substantially complementary or completely complementary, meaning that at least approximately 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% base pairing exists between the two DNA strands.
[0141] "Nucleoside" is a component of DNA and RNA, consisting of ribose (for RNA) or deoxyribose (for DNA) and a base. "Nucleotide" refers to a molecule composed of a nucleoside and a phosphate group, which may contain a hydroxyl group at the 5' position and a phosphate group at the 3' position, or vice versa. In this application, "nucleotide" and "base" may be used interchangeably in certain contexts.
[0142] The "internal guide sequence" or "IGS" refers to a sequence near the 5' end of a type I spliced intron that is required for self-splicing. This sequence can form complementary pairs with the 5' and 3' exons of the intron, enabling the intron to form the correct secondary and tertiary structures, thereby advancing the intron splicing process. The "splicing site" refers to the junction between the intron and the exon.
[0143] When preparing circular RNA using self-cleaving introns, it is inevitable that some non-target sequences will be introduced.
[0144] For example, self-splicing introns require adjacent exon sequences on both sides to form the necessary secondary and tertiary structures, thereby advancing the self-splicing process. Specifically, as shown in Figure 1, the internal guide sequence IGS near the 5' end of the intron needs to form complementary pairings with the 5' and 3' exons of the intron. Splicing occurs at the two triangular splice sites, resulting in intron removal and the joining of the exons on both sides. When constructing vectors for preparing circular RNA using self-splicing introns, the intron is typically divided into a 3' intron fragment (and 3' exon E3) and a 5' intron fragment (and 5' exon E5). The 3' intron fragment (and 3' exon E3) and the 5' intron fragment (and 5' exon E5) are then loaded onto the 5' and 3' sides of the RNA sequence to be circularized, respectively. Regarding how to split an intron in half, i.e., at which sites the intron is split in half, those skilled in the art can determine this based on the primary, secondary, tertiary, and / or quaternary structure of the intron. Furthermore, the methods listed in the embodiments of this application can be used to test whether the split intron fragments can successfully circularize the RNA to be circularized. During RNA circularization, the self-splicing intron and its adjacent exons undergo the aforementioned complementary pairing, and the intron detaches from the overall structure. E3 and E5 connect and remain in the circular RNA product. E3 and E5 are non-target sequences in the circular RNA; their presence may increase immunogenicity and related therapeutic risks, thus limiting the application scope of circular RNA.
[0145] Figures 2A and 5 show schematic structural diagrams of self-splicing introns in the *Bacillus anthracis* recA gene and the *Coxiella behnkeni* 23S ribosomal gene, respectively. The 5' and 3' sequences of the IGS are designated P10 and P1, respectively. The sequence complementary to P1 in the 5' exon is called the P1 antisense strand, while the sequence complementary to P10 in the 3' exon is called the P10 antisense strand. Near the 5' splice site, the G in the IGS forms a relatively conserved G:U pair with the last nucleotide U of the 5' exon.
[0146] The inventors of this application have achieved RNA circularization using a self-splicing intron that does not contain the 5' exon of the intron, by using the 3-20 nucleotide sequence at the 3' end of the RNA to be circularized as the P1 antisense strand and modifying the P1 in the intron IGS to be anticomplementary to the P1 antisense strand. The resulting circular RNA naturally does not contain the E5 sequence. In the modification, the conserved G:U pairing formed by P1 and the P1 antisense strand is preserved. Therefore, the 3' end of the RNA to be circularized should, or preferably should, be a U-terminated nucleotide.
[0147] In addition, the inventors also used the 2-15nt nucleotide sequence at the 5' end of the RNA to be circularized as the P10 antisense strand and modified the P10 in the intron IGS to be inversely complementary to the P10 antisense strand. This allowed them to use self-splicing introns that do not contain the 3' exon of the intron to achieve RNA circularization. Naturally, the resulting circular RNA does not contain the E3 sequence.
[0148] When performing the P1 and P10 modifications simultaneously, RNA circularization can be accomplished using self-splicing introns that do not contain E3 and E5. The resulting circular RNA will naturally not contain the non-target sequences E3 and E5.
[0149] The inventors modified two types of class I self-splicing introns—one from the recA gene of *Bacillus anthracis* and the other from the 23S ribosome gene of *Coxiella behnkeni*—using introns with IGS (Intracytoplasmic Spoken RNA). The modified introns were then used to prepare circularized RNA. Results showed that this intron IGS modification and the application of the modified introns are universally applicable; both can be used for RNA circularization with high efficiency, meeting production requirements. Therefore, the strategy of this invention can also be applied to introns from other sources, modifying their IGS and using the modified introns for the preparation of circularized RNA.
[0150] Therefore, this application provides a single-stranded DNA molecule for preparing circular RNA, which may sequentially comprise from the 5' end to the 3' end: a 3' self-splicing intron fragment containing a 3' splice site, a target sequence, and a 5' self-splicing intron fragment containing a 5' splice site, wherein the elements are operatively linked.
[0151] Both 3' and 5' self-splicing intron fragments can cause the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule to become circular. Both 3' and 5' self-splicing intron fragments can originate from the same self-splicing intron.
[0152] 5' self-splicing intron fragments can contain internal guide sequences.
[0153] The 5' end sequence of the internal guide sequence may be anticomplementary to the 5' end sequence of the target sequence. The 5' end sequence of the internal guide sequence may be anticomplementary to 2-15 nt, for example, 2-12 nt nucleotides at the 5' end of the target sequence. In some embodiments, the 5' end sequence of the internal guide sequence may be anticomplementary to 3-6 nt nucleotides at the 5' end of the target sequence.
[0154] The 3' end sequence of the internal guide sequence can be anticomplementary to the 3' end sequence of the target sequence. The 3' end sequence of the internal guide sequence can be anticomplementary to 3-20 nt (e.g., 3-15 nt) nucleotides at the 3' end of the target sequence. In some embodiments, the 3' end sequence of the internal guide sequence can be anticomplementary to 3-10 nt nucleotides at the 3' end of the target sequence. In other embodiments, the 3' end sequence of the internal guide sequence can be anticomplementary to 3-9 nt nucleotides at the 3' end of the target sequence. As described above, to maintain the conserved G:U pairing in the P1 and P1 antisense strands, the 3' end sequence of the internal guide sequence can be initiated by a G-based nucleotide, and the 3' end sequence of the target sequence can be an T-based nucleotide.
[0155] Once a self-splicing intron for RNA circularization is selected, the 3' intron fragment can be fixed in some cases, while the 5' intron fragment varies depending on the sequence of the RNA to be circularized, because the IGS within it needs to react complementaryly with the sequences at both ends of the RNA to be circularized. In some embodiments, the 3' and 5' self-splicing intron fragments can be derived from the self-splicing intron in the *Bacillus anthracis* RecA gene, and the 3' self-splicing intron fragment may contain the nucleotide sequence shown in SEQ ID NO:10. In some embodiments, the 3' and 5' self-splicing intron fragments can be derived from the self-splicing intron in the *Coxiella belladonna* 23S ribosome gene, and the 3' self-splicing intron fragment may contain the nucleotide sequence shown in SEQ ID NO:21.
[0156] The target sequence may contain an open reading frame encoding the target peptide or protein, a sequence of translational functional elements, a sequence encoding non-coding RNA, a single cloning site, or a multiple cloning site.
[0157] Translational functional elements can be selected from translation initiation elements and translation enhancement elements. In the circular RNA molecules prepared from single-stranded DNA of this application, the translation enhancement elements can be located on the 5' side or 3' side of the open reading frame encoding the target peptide or protein. For example, a poly(A) tail or 3'CITE can be located on the 3' side of the open reading frame encoding the target peptide or protein, while a translation enhancer derived from a Hox gene can be located on the 5' side of the open reading frame encoding the target peptide or protein. In the circular RNA molecules prepared from single-stranded DNA of this application, the translation enhancement elements can be located on the 5' side or 3' side of the translation initiation element. Those skilled in the art can arrange the positions of the translation initiation elements and translation enhancement elements in the single-stranded DNA molecules of this application according to the final structure of the circular RNA molecules. For example, in the final obtained circular RNA molecules, a poly(A) tail or 3'CITE is located on the 3' side of the open reading frame encoding the target peptide or protein, while a translation enhancer derived from a Hox gene is located on the 5' side of the open reading frame encoding the target peptide or protein.
[0158] The target sequence may contain an open reading frame encoding a target peptide or protein, and a sequence of a translational functional element (e.g., IRES). When the 3' end of the sequence of the translational functional element (e.g., IRES) is T, the target sequence can contain the open reading frame encoding the target peptide or protein and the sequence of the translational functional element (e.g., IRES) from the 5' end to the 3' end. When the 3' end of the sequence encoding IRES is not T, the sequence of the translational functional element (e.g., IRES) can be split into two segments, and the 3' end of the 5' end of each segment can be T, so that the target sequence contains the 3' end sequence of the sequence of the translational functional element (e.g., IRES), the open reading frame encoding the target peptide or protein, and the 5' end sequence of the sequence of the translational functional element (e.g., IRES) from the 5' end to the 3' end. When the 3' end of the open reading frame (OPF) encoding the target peptide or protein ends in a T, the target sequence can include the sequence of the translational functional element (e.g., IRES) and the ORF encoding the target peptide or protein from the 5' end to the 3' end. When the 3' end of the ORF encoding the target peptide or protein does not end in a T, the ORF encoding the target peptide or protein can be split into two segments, with the 3' end of each segment ending in a T. Thus, the target sequence can include the 3' end sequence of the ORF encoding the target peptide or protein, the sequence of the translational functional element (e.g., IRES), and the 5' end sequence of the ORF encoding the target peptide or protein from the 5' end to the 3' end. The target sequence can contain one or more ORFs encoding the target peptide or protein and sequences of translational functional elements (e.g., IRES). The target sequence may include, from the 5' end to the 3' end, (a) the open reading frame (OPF) encoding the target peptide or protein and the sequence of the translational element (e.g., IRES); (b) the sequence of the translational element (e.g., IRES) and the one or more ORF encoding the target peptide or protein; (c) the 3' end sequence of the translational element (e.g., IRES), the one or more ORF encoding the target peptide or protein and the 5' end sequence of the translational element (e.g., IRES); or (d) the 3' end sequence of one of the ORF encoding the target peptide or protein, other ORF encoding the target peptide or protein (if present), the sequence of the translational element (e.g., IRES), and the 5' end sequence of one of the ORF encoding the target peptide or protein. Similarly, the 3' end of the target sequence must be arranged as a nucleotide with a T base. The target sequence may contain a sequence encoding non-coding RNA. Depending on whether the 3' end of the sequence encoding non-coding RNA ends with a nucleotide containing a T, the target sequence may (a) contain the sequence encoding non-coding RNA, or (b) contain the 3' end of the sequence encoding non-coding RNA and the 5' end of the sequence encoding non-coding RNA from the 5' end to the 3' end.
[0159] The target sequence may contain an open reading frame (OPF) encoding a target peptide or protein. Depending on whether the 3' end of the OPF encoding the target peptide or protein is a T-terminated nucleotide, the target sequence may (a) contain the OPF encoding the target peptide or protein, or (b) contain the 3' end sequence of the OPF encoding the target peptide or protein and the 5' end sequence of the OPF encoding the target peptide or protein from the 5' end to the 3' end. In some embodiments, the target sequence may contain one or more OPF encoding the target peptide or protein. Depending on whether the 3' end sequence of one or more OPF encoding the target peptide or protein is a T-terminated nucleotide, the target sequence may contain (a) the one or more OPF encoding the target peptide or protein from the 5' end to the 3' end, or (b) the 3' end sequence of one OPF encoding the target peptide or protein, other OPF encoding the target peptide or protein (if present), and the 5' end sequence of one OPF encoding the target peptide or protein.
[0160] The target sequence can contain a single-cloning site or a multiple-cloning site. Any desired sequence, such as an open reading frame encoding a target peptide or protein, or a sequence of a translational functional element (e.g., IRES), can be inserted into the single-stranded DNA molecule of this application through a single-cloning site or multiple-cloning site. When inserting the desired sequence at a single-cloning or multiple-cloning site, it is necessary to consider whether the 3' end of the inserted sequence is a T-containing nucleotide.
[0161] The target sequence may contain a single cloning site or a multiple cloning site, and a sequence of a translational element (e.g., IRES). The target sequence may include, from the 5' end to the 3' end, the 3' end sequence of the translational element (e.g., IRES), the single cloning site or multiple cloning site, and the 5' end sequence of the translational element (e.g., IRES), wherein the 3' end of the 5' end sequence of the translational element (e.g., IRES) may end with a T-nucleotide. Any desired sequence, such as an open reading frame encoding a target peptide or protein, can be inserted into the single-stranded DNA molecule of this application through the single cloning site or multiple cloning site, without considering the type of nucleotide at the 3' end.
[0162] Accordingly, this application also protects a circularizable RNA molecule which may comprise, from the 5' end to the 3' end, the following in sequence: a 3' self-splicing intron fragment containing a 3' splice site, a target sequence, and a 5' self-splicing intron fragment containing a 5' splice site, wherein the elements are operatively linked.
[0163] The 3' self-splicing intron fragment and the 5' self-splicing intron fragment can cause the RNA molecule to circularize and be removed from the RNA molecule during the circularization process.
[0164] 5' self-splicing intron fragments can contain internal guide sequences.
[0165] The 5' end sequence of the internal guide sequence can be reverse complementary to the 5' end sequence of the target sequence. The 5' end sequence of the internal guide sequence can be reverse complementary to 2-15 nt (e.g., 2-12 nt) nucleotides at the 5' end of the target sequence.
[0166] The 3' end sequence of the internal guide sequence can be anticomplementary to the 3' end sequence of the target sequence. The 3' end sequence of the internal guide sequence can be anticomplementary to a 3-20 nt (e.g., 3-15 nt) nucleotide at the 3' end of the target sequence. As mentioned above, to maintain the conserved G:U pairing in the P1 and P1 antisense strands, the 3' end sequence of the internal guide sequence can begin with a G-based nucleotide, and the target sequence can end with a U-based nucleotide at its 3' end.
[0167] Similarly, depending on the nucleotides at the end of the 3' end of the target sequence, the sequence can be rearranged or split into two, with the 5' end placed at the 3' end of the target sequence, and the adjusted target sequence ending with a U nucleotide at the 3' end.
[0168] Thus, the DNA or RNA molecules described in this application can be selected as needed to prepare circular RNA molecules, ensuring that the circular RNA molecules do not contain the 5' exon sequence of self-splicing introns, or the 3' exon sequence, or neither the 5' exon sequence of self-splicing introns nor the 3' exon sequence, thereby reducing the non-target sequences in the circular RNA molecules. Circular RNA molecules without either the 5' exon sequence of self-splicing introns or the 3' exon sequence can be prepared from the DNA or RNA molecules described in this application at different splitting positions of the target sequence.
[0169] Furthermore, the DNA and RNA molecules of this application do not contain additional artificial spacer sequences, thereby further reducing non-target sequences. Moreover, the DNA and RNA molecules of this application maintain high RNA circularization efficiency during the preparation of circular RNA, enabling their application in production, especially large-scale production.
[0170] In the DNA and RNA molecules of this application, the target sequence may contain translational functional elements (IRES), such as CVB3 IRES, EVB107 IRES, HRVB3 IRES, and EV-AIRES (variants containing EV-AIRES, hereinafter referred to as EV-A-mut IRES). The IRES element can be adjusted or split into a 5' end sequence and a 3' end sequence. The 5' end sequence of the IRES is placed after the open reading frame encoding the target peptide or protein, and the 3' end sequence of the adjusted target sequence ends with a nucleotide containing the bases T or U. Splitting the IRES at different positions can prepare circular RNA molecules that do not contain either the 5' exon sequence or the 3' exon sequence of a self-splicing intron. In some preferred embodiments of this application, the cyclization rate of DNA or RNA molecules truncated into two after the 18th base of CVB3 IRES, EVB107 IRES or HRVB3 IRES, or after the 374th base of EV-A-mut IRES, is significantly improved. RNase R treatment of the cyclization product can yield cyclized RNA with higher purity.
[0171] This application also provides a composition that may comprise the single-stranded DNA molecule, double-stranded DNA molecule, vector, or circularizable RNA molecule of this application, or a cell comprising the aforementioned single-stranded DNA molecule, double-stranded DNA molecule, vector, or circularizable RNA molecule. This composition can be used to prepare circularizable RNA and / or circular RNA, particularly to generate translatable proteins or biologically active circular RNA in vitro or in vivo. Biologically active circular RNA may be, for example, miRNA sponges or non-coding RNA.
[0172] This application also provides a composition that may comprise the circular RNA molecule of this application or a cell comprising the circular RNA molecule. The composition may be a pharmaceutical composition and may also comprise a pharmaceutically acceptable carrier, excipient, or diluent.
[0173] Pharmaceutically acceptable carriers can be components in a composition other than the active ingredient (i.e., the carrier, cell, precursor RNA, or circular RNA of this application). Pharmaceutically acceptable carriers can include, but are not limited to, buffers, excipients, stabilizers, or preservatives. Examples of pharmaceutically acceptable carriers are physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delay agents, such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants, or emulsifiers, or combinations thereof. The amount of a pharmaceutically acceptable carrier in a drug composition can be determined experimentally based on the activity of the carrier and the desired properties of the formulation, such as stability and / or minimal oxidation.
[0174] In some embodiments, the compositions of this application may comprise buffer solutions such as acetic acid, citric acid, histidine, boric acid, formic acid, succinic acid, phosphoric acid, carbonic acid, malic acid, aspartic acid, Tris buffer, HEPPSO, HEPES, neutral buffered saline, phosphate buffered saline, etc.; carbohydrates such as glucose, sucrose, mannose or dextran, mannitol; proteins; peptides or amino acids such as glycine; antioxidants; chelating agents such as EDTA or glutathione; adjuvants (e.g., aluminum hydroxide); antibacterial and antifungal agents; and preservatives.
[0175] This application also provides the use of the composition in the preparation of circular RNA or for in vivo therapy. For example, in one embodiment, this application provides a method for expressing a vaccine, therapeutic protein, or other type of protein in a subject in need, including administering the composition of this application containing circular RNA to the subject. The therapeutic protein may be, for example, an antibody, a fusion protein, etc. In one embodiment, the subject is a subject with a loss of a functional gene, and the circular RNA in the composition of this application may perform the function of the RNA transcribed from the missing gene, or continuously express the protein translated from the missing gene in vivo.
[0176] When the circular RNA is translated to express a therapeutic protein or exerts a therapeutic effect, a therapeutically effective amount of the composition containing the circular RNA is administered to the subject. The "therapeutically effective amount" of the composition preferably causes a reduction in the severity of disease symptoms, an increase in the frequency and duration of symptom-free periods, or prevention of damage or disability caused by the disease. For example, in the treatment of a tumor-bearing subject, a "therapeuticly effective amount" means, relative to an untreated subject, preferably, tumor growth is inhibited by at least about 40%, more preferably by at least about 60%, more preferably by at least about 80%, and more preferably by at least about 99%. The therapeutically effective amount of the fusion protein of this application can reduce tumor volume or alleviate symptoms in a subject (typically a human, or possibly another mammal).
[0177] When necessary, the precursor RNA or circular RNA in this application can be purified. For example, when administering the composition of this application to a subject, the circular RNA is preferably purified. Purification of the circular RNA or precursor RNA can be performed by methods such as applying a size exclusion column in a high-performance liquid chromatography (HPLC) system.
[0178] The specific administration of the drug composition can be determined by medical professionals, such as doctors, based on the subject's specific circumstances, such as gender, age, and medical history.
[0179] The pharmaceutical compositions of this application can be formulated for oral, intravenous, intramuscular, subcutaneous, parenteral, spinal, or epidermal administration (e.g., by injection or bolus). "Parenteral administration" refers to methods other than intestinal and topical application, typically administered by injection, including but not limited to intravenous, intramuscular, intra-arterial, intramembranous, intracystic, intraorbital, intracardiac, intradermal, intraperitoneal, tracheal, subcutaneous, subepidermal, intra-articular, sub-bursular, subarachnoid, intraspinal, supradural, and intrasternal injections and boluses. In one embodiment, the composition can be formulated for infusion or intravenous administration. The compositions disclosed herein can be provided, for example, as sterile liquid formulations, such as isotonic aqueous solutions, emulsions, suspensions, dispersions, or viscous compositions that can be buffered to the desired pH. In some embodiments, the composition can be administered in any manner, such as via parenteral or non-parenteral administration, including via aerosol inhalation, injection, infusion, ingestion, transfusion, implantation, or transplantation. For example, the compositions described herein can be administered to patients via intravenous injection, intranasal administration, intrathecal administration, intrasheath administration, intraperitoneal administration via artery, intradermal administration, subcutaneous administration, intratumoral administration, intramedullary administration, intranodal administration, or intramuscular administration.
[0180] The compositions of this application can be transfected into cells via, for example, liposome transfection, electroporation, or encapsulation with nanocarriers. Nanocarriers can be, for example, lipids, polymers, or lipid-polymer hybrids.
[0181] The compositions of this application can be delivered to a subject via a release delivery system. Release delivery systems include polymer-based systems such as polylactide-glycolic acid, copolyoxalate, polyesteramide, polyorthoester, polycaprolactone, polyhydroxybutyrate, and polyanhydride. Delivery systems also include non-polymer systems, which are lipids, including sterols such as cholesterol and cholesterol esters, and fatty acids or neutral fats such as monoglycerides and triglycerides; peptide-based systems; hydrogel release systems; wax coatings; tablets using conventional adhesives and excipients; partially fused implants; and so on. In some embodiments, lipid nanoparticles or polymers are used as delivery carriers for the therapeutic circular RNA described herein, including the delivery of RNA to tissues.
[0182] The pharmaceutical composition can be a sustained-release agent, including implants and microcapsule delivery systems. Biodegradable and biocompatible polymers such as ethylene-vinyl acetate, polyanhydride, polyglycolic acid, collagen, polyorthoester, and polylactic acid can be used. The pharmaceutical composition can be administered via medical devices, such as (1) needle-free subcutaneous injection devices (e.g., U.S. Patents 5,399,163; 5,383,851; 5,312,335; 5,064,413; 4,941,880; 4,790,824; and 4,596,556); (2) microinfusion pumps (U.S. Patent 4,487,603); (3) transdermal drug delivery devices (U.S. Patent 4,486,194); (4) bolus injection devices (U.S. Patents 4,447,233 and 4,447,224); and (5) permeation devices (U.S. Patents 4,439,196 and 4,475,196).
[0183] The compositions of this application can be used in combination with other therapeutic agents, vaccines, etc. The combinations of therapeutic agents discussed herein can be administered simultaneously as a single composition in a pharmaceutically acceptable carrier, or as separate compositions, wherein each agent is contained in a pharmaceutically acceptable carrier. In another embodiment, the combinations of therapeutic agents can be administered sequentially.
[0184] Furthermore, if multiple combination therapies are administered and the drugs are administered sequentially, the order of administration at each time point can be reversed or kept the same, and sequential administration can be combined with simultaneous administration or any combination thereof.
[0185] This application will be further described with reference to the following non-limiting embodiments.
[0186] Example
[0187] Example 1. Modification of the internal guide sequence (IGS) of the self-splicing intron in the RecA gene of Bacillus anthracis.
[0188] Take Minsu KO's published article [2] The class I self-splicing intron of the RecA gene in *Bacillus anthracis* was modified, with the nucleotide sequence of its internal guide sequence (IGS) modified except for the G used for conserved G:U pairing. This was then used in a paper published by Deniell G. [1] A vector was constructed to prepare circular RNA, thereby obtaining circular RNA that does not contain self-splicing introns and their adjacent exons.
[0189] Figure 2A shows the structure of the self-splicing intron and its adjacent exons in the *Bacillus anthracis* RecA gene, including the P1 sequence (GUUUCC) and P10 sequence (UGG) in the IGS sequence. The region in the 5' adjacent exon that is anticomplementary to P1 in the IGS is called the P1 antisense strand (GGAAAU), and the region in the 3' adjacent exon that is anticomplementary to P10 in the IGS is called the P10 antisense strand (CCA). During the self-splicing reaction, P1 and P1 antisense strands, and P10 and P10 antisense strands, pair complementaryly. A break occurs after the U base of the conserved GU base pair in the complementary structure of P1 and P1 antisense strands, and a break occurs at the 5' end of the P10 antisense strand in the complementary structure of P1 and P10 antisense strands. The P1 antisense strand and the P10 antisense strand then connect. Figure 2B shows a schematic diagram of modifications to the IGS sequence and its complementary pairing regions, where nucleotides that can be modified are indicated by N.
[0190] Figure 3A shows the construction of an unmodified IGS vector (left) and the structure of the circular RNA obtained from the vector (right). The vector can sequentially contain a 3' intron of a self-splicing intron and its adjacent 3' exon (E3), a sequence encoding IRES, a sequence encoding EGFP, an adjacent 5' exon (E5), and the 5' intron of the self-splicing intron, wherein the IGS sequence is inversely complementary to the sequences of E3 and E5. Specifically, the vector can sequentially contain a 3' intron and a 3' exon (E3 or P10 antisense strand) (SEQ ID NO:3), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon (E5 or P1 antisense strand), and a 5' intron (SEQ ID NO:6). The sequences of the 3' intron and 3' exon... In (SEQ ID NO:3), the capitalized underscore sequence at the 3' end is the E3 or P10 antisense strand sequence. The 5' exon and 5' intron sequences... In (SEQ ID NO:6), the uppercase italic at the 5' end represents the E5 or P1 antisense sequence, and the underlined part represents the IGS sequence. The lowercase bold underlined part of the IGS sequence represents the sequence that is the anticomplement of E3, and the uppercase bold part of the IGS sequence represents the sequence that is the anticomplement of E5.
[0191] Figure 3B shows the construction (left) of a vector for preparing circular RNA with E5 removed and the IGS sequence modified to be inversely complementary to the 3' end sequence of E3 and the sequence encoding EGFP, and the structure of the circular RNA obtained from this vector (right). Specifically, the vector may sequentially contain a 3' intron and a 3' exon (E3 or P10 antisense strand) (SEQ ID NO:3), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks two AA segments at the 3' end compared to SEQ ID NO:5), and a 5' intron (SEQ ID NO:9). The sequences of the 3' intron and 3' exon... In (SEQ ID NO:3), the capitalized underscore sequence at the 3' end is the E3 or P10 antisense sequence. The 5' intron sequence... In (SEQ ID NO:9), the underlined part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is inversely complementary to E3 (i.e., the P10 sequence), and the uppercase bold part of IGS is the sequence that is inversely complementary to the 3' sequence of the sequence encoding EGFP (i.e., the P1 sequence).
[0192] Figure 3C shows the construction (left) of a vector for preparing circular RNA with E3 removed and the IGS sequence modified to be the 5' end sequence encoding IRES, and the reverse complement of E5, and the structure (right) of the circular RNA obtained from this vector. Specifically, the vector may sequentially contain a 3' intron (SEQ ID NO:10), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon, and a 5' intron (SEQ ID NO:11). The sequence of the 3' intron is shown as Aaagagatgaagagatagtccggactatagagatggaaaatctatagatagtg (SEQ ID NO:10). The sequences of the 5' exon and 5' intron are also shown. In (SEQ ID NO:11), the uppercase italic at the 5' end is the E5 or P1 antisense sequence, and the underlined one is the IGS sequence. The lowercase bold underlined one in the IGS is the sequence that is the reverse complement of the 5' end sequence encoding IRES, and the uppercase bold one in the IGS is the sequence that is the reverse complement of E5.
[0193] Figure 3D shows the construction (left) of a vector for preparing circular RNA, in which E3 and E5 are removed, the two AA nucleotides at the 3' end of the EGFP-encoding sequence are placed at the 5' end of the IRES-encoding sequence, and the IGS is modified to be inversely complementary to the 5' end sequence of the IRES-encoding sequence and its two AA nucleotides on its 5' side, as well as the 3' end sequence of the EGFP-encoding sequence, and the structure (right) of the circular RNA obtained from this vector. Specifically, the vector may sequentially contain a 3' intron (SEQ ID NO:10), two AA nucleotides, a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks the two AA nucleotides at the 3' end compared to SEQ ID NO:5), and a 5' intron (SEQ ID NO:13). The sequence of the 5' intron... In (SEQ ID NO:13), the underlined part is the IGS sequence, where the lowercase bold underlined part of IGS is the 5' end sequence of the sequence encoding IRES and the two AA sequences located on its 5' side are inversely complementary. The uppercase bold part of IGS is the sequence inversely complementary to the 3' end sequence of the sequence encoding EGFP.
[0194] 1.1 Carrier Construction
[0195] Construct the two carriers shown in Figures 3B and 3C.
[0196] Specifically, for the vector shown in Figure 3B, a synthesized coding DNA fragment is formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron and a 3' exon (SEQ ID NO:3), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks two AA segments at the 3' end compared to SEQ ID NO:5), a 5' intron (SEQ ID NO:9), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand are cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product is obtained by EcoRV digestion.
[0197] For the vector shown in Figure 3C, a synthesized coding DNA fragment was formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:10), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon and a 5' intron (SEQ ID NO:11), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary fragment were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion.
[0198] The gene synthesis, vector construction, and plasmid linearization described above were all performed using GenScript.
[0199] 1.2 In vitro transcription
[0200] The linearized plasmid vector obtained above was subjected to in vitro transcription (IVT) to obtain RNA product. Specifically, the reaction system shown in Table 1 was incubated at 37°C for 2 h, then 2 μl of DNase I (RNase-free, ON-109, Hongene) was added, mixed well, and incubated at 37°C for 30 min.
[0201] Table 1. Reaction system for in vitro transcription
[0202] 1.3 LiCl precipitation purification of RNA
[0203] The IVT stock solution obtained above was purified by LiCl precipitation to obtain the RNA product. Specifically, 22 μl of enzyme-free water (AM9937, Thermo Fisher) and 20 μl of 8M LiCl solution (DNase and RNase-free, ST498-100 ml, Beyotime) were added to the IVT stock solution to make the LiCl concentration 2.5M. The mixture was then incubated at -20℃ for at least 30 min. After centrifugation at 12000g for 15 min at 4℃, the supernatant was discarded. 500 μl of 70% ethanol was added, the mixture was inverted and mixed, and centrifuged at 12000g for 3 min at 4℃, and the supernatant was discarded; this step was repeated. After centrifugation at 12000g for 1 min at 4℃, the supernatant was aspirated, and 100 μl of enzyme-free water was added to dissolve the RNA.
[0204] The concentration of purified RNA was determined using a Nanodrop ONEC micro-volume UV-Vis spectrophotometer (ND-ONEC-W, Thermo Scientific).
[0205] 1.4 RNA Circulation
[0206] The purified RNA obtained above was then circularized (one-step circularization method). Specifically, 20 μg of purified RNA was taken and adjusted to approximately 0.5-2 μg / μl with enzyme-free water, incubated at 70°C for 5 min, and then immediately placed on ice for 3 min. GTP (100 mM) was added to a final concentration of 2 mM, and then Mg2+ was added. 2+ T4 RNA ligase buffer (10×) (B0216S, NEB) allows Mg 2+ The final concentration was 10 mM. Incubate at 55°C for 10 min.
[0207] The product was purified by LiCl precipitation. Specifically, 8M LiCl solution was added to the RNA circularization solution to make the LiCl concentration 2.5M. The concentration of the purified RNA could be detected by Nanodrop.
[0208] 1.5 RNase R digestion treatment
[0209] Take the purified RNA product obtained from LiCl in step 1.4, add 0.6 μl Tris-HCl, 1 μl KCl-3M, 0.003 μl MgCl2-1M, and 0.3 μl RNase R (48 U / μl), mix well by pipetting, and incubate at 37 °C for 15 min.
[0210] The product was purified by LiCl precipitation. Specifically, 8M LiCl solution was added to the RNA solution to make the LiCl concentration 2.5M. The concentration of the purified RNA could be detected by Nanodrop.
[0211] 1.6 Gel electrophoresis analysis
[0212] Take the RNA obtained in steps 1.3, 1.4, and 1.5 above, and process it using an E-Gel... TM E-Gel electrophoresis was performed for analysis. Specifically, 150-200 ng of purified RNA was taken, adjusted to 10 μl with enzyme-free water, and then 10 μl of gel loading buffer II (AM8546G, Thermo Fisher) was added and mixed well. The E-Gel was then removed. TM EX agarose gel (2%, G401002, Thermo Fisher), remove the comb, place the gel into the gel cartridge of the electrophoresis system, and add 20 μl of well-mixed sample to each well. Set the electrophoresis time to 10-20 min and begin electrophoresis. After electrophoresis, turn on the filter and cool for 5-10 min. Use E-Gel. TM The EX device (G8300, ThermoFisher) has a built-in image acquisition system for observing electrophoresis results and acquiring images.
[0213] E-Gel of RNA generated from the vector constructed according to Figure 3B TM The E-Gel electrophoresis gel diagram is shown in Figure 4A, representing the E-Gel of RNA generated from the vector constructed according to Figure 3C. TM The EX gel electrophoresis results are shown in Figure 4B. The three lanes in each figure, from left to right, represent the RNA obtained in lanes 1.3, 1.4, and 1.5, respectively. The uppermost band represents the linear RNA obtained by IVT. Due to the presence of self-splicing introns, it is slightly larger than the circular RNA band and moves more slowly. As can be seen from the figure, the RNA treated with RNase R contains almost no linear precursor RNA, while circular RNA is enriched.
[0214] 1.7 Capillary gel electrophoresis
[0215] To detect the efficiency of RNA circularization, the circularized RNA was subjected to capillary gel electrophoresis using an Agilent 5200 fragment analyzer (M53100AA, Agilent) and an Agilent DNF-471 RNA kit (15nt, DNF-471-0500, Agilent).
[0216] Specifically, take the purified RNA from step 1.4 and adjust its concentration to 100-150 ng / μl with enzyme-free water. Add 22 μL of dilution standard (RNA dilution standard), 2 μL of the sample to be tested, and 2 μL of RNA molecular weight standard to the sample plate. Mix and incubate at 80°C for 2 min, then immediately place on an ice box at 4°C for instantaneous cooling. Subsequently, perform capillary electrophoresis according to the Agilent 5200 fragment analyzer manual.
[0217] The capillary gel electrophoresis diagram of RNA generated from the vector constructed according to Figure 3B is shown in Figure 4C, and the capillary gel electrophoresis diagram of RNA generated from the vector constructed according to Figure 3C is shown in Figure 4D.
[0218] As can be seen from the band sizes in the figure, the band of the in vitro transcription product before circularization is slightly larger than the band of the circular RNA after circularization, because the linear precursor RNA contains self-splicing intron sequences. The purity of the circular RNA after circularization is above 50%, specifically 62.2% and 75.3%, respectively.
[0219] 1.8 Sequencing of circular RNA
[0220] The circular RNA obtained in step 1.4 was reverse transcribed, PCR-mediated, and gel-recovered, and the circular adapter sites were sequenced. II. First-strand cDNA Synthesis Kit (with gDNA removal agent) (R212-01, Novizan) was used to reverse transcribe the product obtained from the above in vitro self-splicing reaction. Specifically, the reaction system was prepared as shown in Table 2, mixed by pipetting, and heated at 42°C for 2 min.
[0221] Table 2. Reverse transcription reaction system
[0222] Next, add 4 μl of 5×HiScript II qRT SuperMix II to the above mixture, pipette to mix, and incubate at 50℃ for 15 min and 85℃ for 5 sec. After reverse transcription, add the obtained cDNA to the PCR reaction system shown in Table 3, pipette to mix, and perform the PCR reaction according to the procedure shown in Table 4.
[0223] Table 3. PCR reaction system
[0224] Table 4. PCR Program Settings
[0225] All reaction products were recovered by electrophoresis, and the target band was cut using a blue light gel cutter (OSE-470L, Tiangen Biotech Co., Ltd.) as a reference. Gel recovery was performed using the gel recovery kit (D2500-01 / D2500-02, omega) according to the instructions. The recovered product was sent for Sanger sequencing, and the sequencing primer sequences for the circular adapter were SEQ ID NO: 16 and 17.
[0226] The sequencing results of the circular RNA generated by the vectors constructed according to Figures 3B and 3C are shown in SEQ ID NO:44 and 45, respectively, where the italicized and bold text represents the residual adjacent exon sequences.
[0227] It can be seen that after removing E5 and making the P1 sequence in the intron IGS complementary to the 3' end sequence ACAAGT of the EGFP-encoding sequence, the resulting RNA can be circularized, and only the 3' exon sequence CCA remains in the circular RNA. After removing E3 and making the P10 in the IGS complementary to the 5' end TTA of the IRES-encoding sequence, the resulting RNA can be circularized, and only the 5' exon sequence GGAAAU remains in the circular RNA.
[0228] Example 2. Modification of the internal guide sequence of the self-splicing intron in the Coxione 23S ribosome gene.
[0229] Similarly, referring to Example 1, the self-splicing intron in the Coxionemal 23S ribosomal gene was modified by IGS. Figure 5 shows the structure of this intron and its adjacent exon, wherein the P1 sequence (GUAGUUACG) in the IGS is anticomplementary to the P1 antisense strand (CGUAACUAU) in the adjacent exon, and the P10 sequence (ACCGUU) in the IGS is anticomplementary to the P10 antisense strand (AACGGU) in the adjacent exon.
[0230] Figure 6A shows the structure of an unmodified IGS vector (left) and the structure of the circular RNA obtained from the vector (right). The vector can sequentially contain a 3' intron of a self-splicing intron and its adjacent 3' exon (E3), a sequence encoding IRES, a sequence encoding EGFP, an adjacent 5' exon (E5), and the 5' intron of the self-splicing intron, wherein the IGS sequence is inversely complementary to the sequences of E3 and E5. Specifically, the vector can sequentially contain a 3' intron and a 3' exon (E3 or P10 antisense strand) (SEQ ID NO:18), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon (E5 or P1 antisense strand), and a 5' intron (SEQ ID NO:19). The sequences of the 3' intron and 3' exon... In (SEQ ID NO:18), the capitalized underscore sequence at the 3' end is the E3 or P10 antisense strand sequence. The 5' exon and 5' intron sequences... In (SEQ ID NO:19), the uppercase italic at the 5' end represents the E5 or P1 antisense sequence, and the underlined part represents the IGS sequence. The lowercase bold underlined part of the IGS sequence represents the sequence that is the antisense complement of E3 (P10), and the uppercase bold part of the IGS sequence represents the sequence that is the antisense complement of E5 (P1).
[0231] Figure 6B shows the construction (left) of a vector for preparing circularizable RNA with E5 removed and the IGS sequence modified to be inversely complementary to the 3' end sequence of E3 and the sequence encoding EGFP, and the structure (right) of the circular RNA obtained from this vector. Specifically, the vector may sequentially contain a 3' intron and a 3' exon (E3 or P10 antisense strand) (SEQ ID NO:18), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks two AA segments at the 3' end compared to SEQ ID NO:5), and a 5' intron (SEQ ID NO:20). The 5' exon and 5' intron sequences... In (SEQ ID NO:20), the underlined part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is inversely complementary to E3, and the uppercase bold part of IGS is the sequence that is inversely complementary to the 3' end sequence of the EGFP encoding sequence.
[0232] Figure 6C shows the construction (left) of a vector for preparing circular RNA with E3 removed and the IGS sequence modified to be the 5' end sequence encoding the IRES sequence and E5 reverse complementary, and the structure (right) of the circular RNA obtained from the vector. Specifically, the vector may sequentially contain a 3' intron (SEQ ID NO:21), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon, and a 5' intron (SEQ ID NO:22). The sequence of the 3' intron is shown in SEQ ID NO:21. The sequences of the 5' exon and 5' intron are also shown. In (SEQ ID NO:22), the uppercase italic at the 5' end is the E5 or P1 antisense sequence, and the underlined one is the IGS sequence. The lowercase bold underlined one in the IGS is the sequence that is the reverse complement of the 5' end sequence encoding IRES, and the uppercase bold one in the IGS is the sequence that is the reverse complement of E5.
[0233] Figure 6D shows the construction (left) of a circular RNA preparation vector with E3 and E5 removed, the two AA nucleotides at the 3' end of the EGFP-encoding sequence placed on the 5' side of the IRES-encoding sequence, and the IGS modified to be reverse complementary to the 5' end sequence of the IRES-encoding sequence and its two AA nucleotides on the 5' side, as well as the 3' end sequence of the EGFP-encoding sequence, and the structure (right) of the circular RNA obtained from this vector. Specifically, the vector may sequentially contain a 3' intron (SEQ ID NO:21), two AA nucleotides, a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks the two AA nucleotides at the 3' end compared to SEQ ID NO:5), and a 5' intron (SEQ ID NO:23). The sequence of the 5' intron... In (SEQ ID NO:23), the underlined part is the IGS sequence, where the lowercase bold underlined part of IGS is the 5' end sequence of the sequence encoding IRES and the two AA sequences located on its 5' side, which are inversely complementary. The uppercase bold part of IGS is the sequence inversely complementary to the 3' end sequence of the sequence encoding EGFP.
[0234] Construct the vectors for preparing circularizable RNA as shown in Figures 6B and 6C.
[0235] Specifically, for the vector shown in Figure 6B, a synthesized coding DNA fragment is formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron and a 3' exon (SEQ ID NO:18), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:12, which lacks two AA segments at the 3' end compared to SEQ ID NO:5), a 5' intron (SEQ ID NO:20), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand are cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product is obtained by EcoRV digestion.
[0236] For the vector shown in Figure 6C, a synthesized coding DNA fragment was formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding green fluorescent protein (EGFP) (SEQ ID NO:5), a 5' exon and a 5' intron (SEQ ID NO:22), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion.
[0237] The gene synthesis, vector construction, and plasmid linearization described above were all performed using GenScript.
[0238] Referring to steps 1.2 to 1.7 in Example 1, the obtained linearized plasmid was subjected to in vitro transcription, circularization treatment, RNase R digestion, and E-Gel processing. TM EX gel electrophoresis and capillary electrophoresis detection.
[0239] E-Gel of RNA generated from the vector constructed according to Figure 6B TM The E-Gel electrophoresis gel diagram is shown in Figure 7A, showing the E-Gel of RNA generated from the vector constructed according to Figure 6C. TM The EX gel electrophoresis results are shown in Figure 7B. The three lanes in each figure, from left to right, represent the RNA obtained in lanes 1.3, 1.4, and 1.5, respectively. The uppermost band represents the linear RNA obtained by IVT. Due to the presence of self-splicing introns, it is slightly larger than the circular RNA band and moves more slowly. As can be seen from the figure, the RNA treated with RNase R contains almost no linear precursor RNA, while circular RNA is enriched.
[0240] Figure 7C shows the capillary gel electrophoresis results of RNA generated from the vector constructed according to Figure 6B, and Figure 7D shows the capillary gel electrophoresis results of RNA generated from the vector constructed according to Figure 6C. The band sizes indicate that the band of the in vitro transcription product before circularization is slightly larger than the band of the circular RNA after circularization, because the linear precursor RNA contains self-splicing intron sequences. The purity of the circular RNA after circularization is above 50%, specifically 52.2% and 60.6%, respectively.
[0241] Following the method steps in 1.8 of Example 1, the purified RNA obtained in 1.4 was subjected to reverse transcription, PCR, and gel recovery, and the circular adapter sites were sequenced. The primers used for PCR and sequencing were SEQ ID NO: 16 and 17.
[0242] The sequencing results of the circular RNA generated by the vectors constructed according to Figures 6B and 6C are shown in SEQ ID NO:24 and 25, respectively, where the italicized and bold text represents the residual adjacent exon sequences.
[0243] It can be seen that after removing E5 and making the P1 sequence in the intron IGS complementary to the 3' end sequence TGTACAAGT of the EGFP-encoding sequence, the resulting RNA can be circularized, and only the 3' exon sequence AACGGU remains in the circular RNA. Similarly, after removing E3 and making the P10 sequence in the IGS complementary to the 5' end TTAAAA of the IRES-encoding sequence, the resulting RNA can be circularized, and only the 5' exon sequence CGUAACUAU remains in the circular RNA.
[0244] Example 3. Construction of a vector for preparing circular RNA containing self-splicing introns of the RecA gene from Bacillus anthracis modified with IGS
[0245] A vector with the structure shown in Figure 8A was constructed to verify the effect of simultaneously removing RNA circularization from exons E3 and E5 adjacent to the self-splicing intron in the *Bacillus anthracis* RecA gene. The structural principle is the same as that shown in Figure 3D, except that the target gene encoded by *Fluc* was replaced with luciferase. Specifically, a coding DNA fragment was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:10), two AA nucleotides, a sequence encoding CVB3 IRES (SEQ ID NO:4), a sequence encoding luciferase (Fluc) (SEQ ID NO:29, lacking the two AA nucleotides at the 3' end compared to SEQ ID NO:28), a 5' intron (SEQ ID NO:30), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:30), the underlined part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is inversely complementary to the 5' end sequence of the sequence encoding IRES and the two AAs located at its 5'. The uppercase bold part of IGS is the sequence that is inversely complementary to the 3' end sequence of the sequence encoding Fluc.
[0246] Meanwhile, with the same aim of preparing circular RNA without adjacent exons E3 and E5, two other vectors were prepared and constructed, as shown in Figure 8B and Figure 8C.
[0247] The vector shown in Figure 8B places the IRES truncated at both ends of the open reading frame (ORF) encoding the target gene such as Fluc, ensuring that the last base of 5'IRES is T. At the same time, the P1 and P10 sequences in the IGS of the self-splicing intron are modified to be reverse complementary to the 3' end sequence of 5'IRES and the 5' end sequence of 3'IRES, respectively.
[0248] In the vector shown in Figure 8C, the ORF encoding the target gene such as Flu is truncated and placed at both ends of IRES, ensuring that the last base of the 5'ORF is T. At the same time, the P1 and P10 sequences in the IGS of the self-splicing intron are modified to be reverse complementary to the 3' end of the 5'ORF and the 5' end of the 3'ORF, respectively.
[0249] For the vector shown in Figure 8B, a synthesized coding DNA fragment was formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:10), a 3' CVB3 IRES coding sequence (SEQ ID NO:26), a sequence encoding luciferase (Fluc) (SEQ ID NO:28), a 5' CVB3 IRES assembly (SEQ ID NO:27), a 5' intron (SEQ ID NO:31), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:31), the underlined part is the IGS sequence, where the lowercase bold underlined part of the IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of the IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0250] For the vector shown in Figure 8C, a synthesized coding DNA fragment was formed, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:10), a 3' luciferase (Fluc) coding sequence (SEQ ID NO:32), a sequence encoding CVB3 IRES (SEQ ID NO:4), a 5' luciferase (Fluc) coding sequence (SEQ ID NO:33), a 5' intron (SEQ ID NO:34), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:34), the underlined part is the IGS sequence, where the lowercase bold underlined part of the IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'Fluc, and the uppercase bold part of the IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'Fluc.
[0251] Referring to steps 1.2 to 1.7 in Example 1, the obtained linearized plasmid was subjected to in vitro transcription, circularization (for intron splicing), RNase R digestion, and E-Gel digestion. TM EX gel electrophoresis and capillary electrophoresis detection.
[0252] E-Gel of RNA generated from vectors constructed according to Figures 8A-8C TM The gel electrophoresis results are shown in Figures 9A-9C. It can be seen that all three vectors can produce circular RNA.
[0253] The capillary electrophoresis results of the RNA produced by the vectors constructed according to Figures 8A-8C are shown in Figures 9D-9F, respectively. The purity of the circularized RNA was 45.9%, 57.3%, and 58.5%, respectively. It can be seen that the vectors constructed according to Figures 8B and 8C can produce circular RNA with high purity, meeting production requirements. The conventional production circularization rate of ORF-Fluc is around 50%.
[0254] Following the method steps in 1.8 of Example 1, the purified RNA obtained according to 1.4 was subjected to reverse transcription, PCR, and gel recovery, and the circular adapter positions were sequenced. Specifically, for the circular RNA produced by the vector constructed in Figure 8A, the primers used for PCR and sequencing are shown in SEQ ID NO:17 and 14; for the circular RNA produced by the vector constructed in Figure 8B, the primers used for PCR and sequencing are shown in SEQ ID NO:43 and 14; and for the circular RNA produced by the vector constructed in Figure 8C, the primers used for PCR and sequencing are shown in SEQ ID NO:15 and 35.
[0255] The sequencing results of the adapter sites of the circular RNAs generated by the vectors in Figures 8A-8C are shown in SEQ ID NO:36-38, respectively. As can be seen, none of the three circularized products contain exogenous residual sequences (i.e., no self-splicing type adjacent exon residues).
[0256] Gttttggagcacggaaagacgatgacggaaaaagagatcgtggattacgtcgccagtcaagtaacaaccgcgaaaaagttgcgcggaggagttgtgtttgtggacgaagtaccgaaaggtcttaccggaaaactcgacgcaagaaaaatcagagagatcctcataaaggccaagaagggcggaaagatcgccgtgt / aaTTAAAACAGCCTGTGGGTTGATCCCACCCACAGGCCCATTGGGCGCTAGCACTCTGGTATCACGGTACCTTTGTG(SEQ ID NO:36, lowercase is the Fluc3' end sequence, uppercase is the IRES5' end sequence, / is the linker)
[0257] (SEQ ID NO:37, lowercase for Fluc sequence, uppercase for IRES sequence, where the bold underlined part is the 5' end IRES sequence, and / indicates the junction)
[0258] Acatcacttacgctgagtacttcgaaatgtccgttcggttggcagaagctatgaaacgatatgggctgaatacaaatcacagaatcgtcgtatgcagtgaaaactctcttcaa ttctttatgccggtgttgggcgcgttatttatcggagttgcagttgcgcccgcgaacgacatttataatgaacgtgaattgctcaacagtatgggcatttcgcagcctaccgtg gtgttcgttt / CCAAAAAGGGGTTGCAAAAAATTTTGAACGTGCAAAAAAAGCTCCCAATCATCCAAAAAATTATTATCATGGATTCTAAAACGGATTACCAGGGATTTCAGTCGATGTACACGTTCGTCACATCTCATCTACCTCCCGGTTTTAATGAATACGATTTTGTGCCAGAGTCCTTCGATAGGGACAAGACAATTGCACTGATCATGAACTCCTCT(SEQ ID NO:38, lowercase represents the 5' end sequence of Fluc, uppercase represents the 3' end sequence of Fluc, / represents the linker)
[0259] Example 4. Construction of a vector for preparing circular RNA containing IGS-modified Coxiellar 23S ribosomal gene self-splicing introns.
[0260] Using the self-splicing intron from the Coxison 23S ribosomal gene, a vector as shown in Figure 8B was constructed to verify the RNA circularization effect of simultaneously removing exons E3 and E5 adjacent to the self-splicing intron.
[0261] Specifically, a synthetic coding DNA fragment was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3' CVB3 IRES (SEQ ID NO:40), a sequence encoding green fluorescent protein EGFP (SEQ ID NO:5), a sequence encoding 5' CVB3 IRES (SEQ ID NO:41), a 5' intron (SEQ ID NO:39), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:39), the underlined part is the IGS sequence, where the lowercase bold underlined part of the IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of the IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0262] Referring to steps 1.2 to 1.7 in Example 1, the obtained linearized plasmid was subjected to in vitro transcription, circularization (for intron splicing), RNase R digestion, and E-Gel digestion. TM EX gel electrophoresis and capillary electrophoresis detection.
[0263] The E-Gel of RNA produced by the constructed vector TM The gel electrophoresis results are shown in Figure 10A, which show the presence of circular RNA bands.
[0264] The results of capillary electrophoresis of the RNA produced by the constructed vector are shown in Figure 10B. The purity of the circularized RNA was 64.5%, which is high and meets the production requirements.
[0265] Following the method steps in 1.8 of Example 1, the purified RNA obtained in 1.5 was subjected to reverse transcription, PCR, and gel recovery, and the circular adapter site was sequenced. The primers used for PCR and sequencing were SEQ ID NO:16 and 17, and the sequencing results at the adapter site are shown in SEQ ID NO:42. It can be seen that the circular product does not contain exogenous residual sequences (i.e., no self-splicing type adjacent exon residues).
[0266] (SEQ ID NO:42, lowercase is the EGFP sequence, uppercase is the IRES sequence, where the bold underline is the 5' IRES sequence, and / is the linker)
[0267] Example 5. Construction of circular RNA preparation vectors based on different IRES truncation positions
[0268] To test whether circular RNA without exogenous residual sequences could still be prepared from different IRES breakpoints in the vector shown in Figure 8B, and to test the circularization effect of different IRES breakpoints, three vectors, Coxie-B18-Gluc, Coxie-B18-EGFP, and Coxie-B18-Fluc, were constructed using a self-splicing intron from the Coxsell 23S ribosomal gene and a CVB3 IRES truncated at the 18th base. These vectors were used to verify the RNA circularization effect of simultaneously removing exons E3 and E5 adjacent to the self-splicing intron.
[0269] Specifically, a DNA fragment encoding the Coxie-B18-Gluc vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3' CVB3 IRES (SEQ ID NO:46), a sequence encoding Gaussian luciferase Gluc (SEQ ID NO:49), a sequence encoding 5' CVB3 IRES (SEQ ID NO:47), a 5' intron (SEQ ID NO:48), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:48), the bold part is the IGS sequence, where the lowercase bold and underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0270] Specifically, the coding DNA fragment of the Coxie-B18-EGFP vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3' CVB3 IRES (SEQ ID NO:46), a sequence encoding green fluorescent protein EGFP (SEQ ID NO:5), a sequence encoding 5' CVB3 IRES (SEQ ID NO:47), a 5' intron (SEQ ID NO:48), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:48), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0271] Specifically, the coding DNA fragment of the Coxie-B18-Fluc vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3' CVB3 IRES (SEQ ID NO:46), a sequence encoding firefly luciferase Flu (SEQ ID NO:28), a sequence encoding 5' CVB3 IRES (SEQ ID NO:47), a 5' intron (SEQ ID NO:48), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:48), the bold part is the IGS sequence, where the lowercase bold and underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0272] Referring to steps 1.2 to 1.7 in Example 1, the obtained linearized plasmid was subjected to in vitro transcription, circularization (for intron splicing), RNase R digestion, and E-Gel digestion. TM EX gel electrophoresis and capillary electrophoresis detection.
[0273] The E-Gel of RNA produced by the constructed vector TM The gel electrophoresis results are shown in Figures 11A-11C. The location of the circular bands was determined by post-treatment with RNase R (RR), indicating that all constructed vectors can produce circular RNA.
[0274] The capillary electrophoresis results of the RNA produced by the constructed vectors are shown in Figures 11D-11I. The purity of the RNA after circularization with Coxie-B18-Gluc was 78.5%, and the purity after digestion with RNase R was 99.4%; the purity of the RNA after circularization with Coxie-B18-EGFP was 80.9%, and the purity after digestion with RNase R was 100%; the purity of the RNA after circularization with Coxie-B18-Fluc was 83.5%, and the purity after digestion with RNase R was 98.7%. The results show that the purity after circularization with different ORFs is superior to that of published strategies.
[0275] Following the method steps in 1.8 of Example 1, the purified RNA obtained according to 1.5 was subjected to reverse transcription, PCR, and gel recovery, and the circular adapter sites were sequenced. The primers used for PCR and sequencing of Circ-coxie-B18-Gluc (i.e., circularized coxie-B18-Gluc) were SEQ ID NO:53 and 54, and the sequencing results at the adapter sites are shown in SEQ ID NO:50. The primers used for PCR and sequencing of Circ-coxie-B18-EGFP (i.e., circularized coxie-B18-EGFP) were SEQ ID NO:16 and 54, and the sequencing results at the adapter sites are shown in SEQ ID NO:51. The primers used for PCR and sequencing of Circ-coxie-B18-Fluc (i.e., circularized coxie-B18-Fluc) were SEQ ID NO:14 and 54, and the sequencing results at the adapter sites are shown in SEQ ID NO:52. It can be seen that the circularized products do not contain exogenous residual sequences (i.e., no self-splicing type adjacent exon residues).
[0276] Circ-coxie-B18-Gluc sequencing results
[0277] (SEQ ID NO:50, lowercase represents the Gluc sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates a connection point)
[0278] Circ-coxie-B18-EGFP sequencing results
[0279] (SEQ ID NO:51, lowercase represents the EGFP sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates the linker)
[0280] Circ-Coxie-B18-Fluc sequencing results
[0281] (SEQ ID NO:52, lowercase represents the EGFP sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates the linker)
[0282] Based on the strategy in Figure 8B of this application for preparing circular RNA vectors without exogenous residual sequences, CVB3IRES can be truncated at different positions, and circular RNA without exogenous residual sequences can be produced. At the same time, the circularization rate is significantly improved after truncating the 18th base of CVB3IRES. After treatment with RNase R, circular RNA with higher purity can be obtained, which is significantly better than the published strategies.
[0283] Example 6. Construction of circular RNA preparation vectors for different types of IRES
[0284] To test whether the vector shown in Figure 8B is applicable to other IRES types, IRES variants of enterovirus B107 (EVB107), human rhinovirus B3 (HRVB3), and enterovirus A (EV-A) were screened.
[0010] The sequences are shown in SEQ ID NO:61, 74 and 75, respectively. We replaced the CVB3 IRES with IRES from enterovirus B107 (EVB107), human rhinovirus B3 (HRVB3) and the IRES variant of EV-A (EV-A-mut), respectively. At the same time, we used the self-splicing intron from the Coxiella 23S ribosomal gene to construct two vectors as shown in Figure 8B: Coxie-EVB107-B18-Gluc, Coxie-EVB107-B18-EGFP, Coxie-HRVB3-B18-Gluc, Coxie-HRVB3-B18-EGFP, Coxie-EV-A-mut-B374-Gluc and Coxie-EV-A-mut-B374-EGFP. To verify the effect of simultaneously removing exons E3 and E5 adjacent to the self-splicing intron, the IRES of EVB107 were truncated after the 18th base, the IRES of HRVB3 were truncated after the 18th base, and the IRES of EV-A-mut were truncated after the 374th base.
[0285] Specifically, the coding DNA fragment of the Coxie-EVB107-B18-Gluc vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'EVB107 IRES (SEQ ID NO:55), a sequence encoding Gaussian luciferase Gluc (SEQ ID NO:49), a sequence encoding 5'EVB107 IRES (SEQ ID NO:56), a 5' intron (SEQ ID NO:57), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:57), the bold part is the IGS sequence, where the lowercase bold and underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0286] Specifically, the coding DNA fragment of the Coxie-EVB107-B18-EGFP vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'EVB107 IRES (SEQ ID NO:55), a sequence encoding green fluorescent protein EGFP (SEQ ID NO:5), a sequence encoding 5'EVB107 IRES (SEQ ID NO:56), a 5' intron (SEQ ID NO:57), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:57), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0287] Specifically, the coding DNA fragment of the Coxie-HRVB3-B18-Gluc vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'HRVB3 IRES (SEQ ID NO:62), a sequence encoding Gaussian luciferase Gluc (SEQ ID NO:49), a sequence encoding 5'HRVB3 IRES (SEQ ID NO:63), a 5' intron (SEQ ID NO:64), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:64), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0288] Specifically, the coding DNA fragment of the Coxie-HRVB3-B18-EGFP vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'HRVB3 IRES (SEQ ID NO:62), a sequence encoding green fluorescent protein EGFP (SEQ ID NO:5), a sequence encoding 5'HRVB3 IRES (SEQ ID NO:63), a 5' intron (SEQ ID NO:64), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:64), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0289] Specifically, the coding DNA fragment of the Coxie-EV-A-mut-B374-Gluc vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'EV-A-mut IRES (SEQ ID NO:65), a sequence encoding Gaussian luciferase Gluc (SEQ ID NO:49), a sequence encoding 5'EV-A-mut IRES (SEQ ID NO:66), a 5' intron (SEQ ID NO:67), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:67), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0290] Specifically, the coding DNA fragment of the Coxie-EV-A-mut-B374-EGFP vector was synthesized, which, from 5' to 3', sequentially includes: a T7 promoter (SEQ ID NO:1), a 5' homologous arm (SEQ ID NO:2), a 3' intron (SEQ ID NO:21), a sequence encoding 3'EV-A-mut IRES (SEQ ID NO:65), a sequence encoding green fluorescent protein EGFP (SEQ ID NO:5), a sequence encoding 5'EV-A-mut IRES (SEQ ID NO:66), a 5' intron (SEQ ID NO:67), a 3' homologous arm (SEQ ID NO:7), and an EcoRV restriction enzyme site (GATATC) for plasmid linearization. The synthesized DNA fragment and its complementary strand were cloned into the pUC57 vector (SEQ ID NO:8), and the linearized vector product was obtained by EcoRV digestion. The 5' intron sequence... In (SEQ ID NO:67), the bold part is the IGS sequence, where the lowercase bold underlined part of IGS is the sequence that is the reverse complement of the 5' end sequence of the sequence encoding 3'IRES, and the uppercase bold part of IGS is the sequence that is the reverse complement of the 3' end sequence of the sequence encoding 5'IRES.
[0291] Referring to steps 1.2 to 1.7 in Example 1, the obtained linearized plasmid was subjected to in vitro transcription, circularization (for intron splicing), RNase R digestion, and E-Gel digestion. TM EX gel electrophoresis and capillary electrophoresis detection.
[0292] The E-Gel of RNA produced by the constructed Coxie-EVB107-B18-Gluc vector TM The gel electrophoresis results are shown in Figure 12A. The E-Gel of RNA produced by the constructed Coxie-EVB107-B18-EGFP vector is shown. TM The gel electrophoresis results are shown in Figure 12B; E-Gel of RNA produced by the Coxie-HRVB3-B18-Gluc vector. TM The gel electrophoresis results are shown in Figure 14A. The E-Gel of RNA produced by the constructed Coxie-HRVB3-B18-EGFP vector is shown. TM The gel electrophoresis results are shown in 14B; the E-Gel of RNA produced by the constructed Coxie-EV-A-mut-B374-Gluc vector. TMThe gel electrophoresis results are shown in Figure 14C. The E-Gel of RNA produced by the constructed Coxie-EV-A-mut-B374-EGFP vector... TM The gel electrophoresis results are shown in Figure 14D.
[0293] The three lanes in each figure, from left to right, represent the RNA obtained according to steps 1.3, 1.4, and 1.5 of Example 1. Circular RNA bands are present in both the circularized RNA (Circ) and the RNA treated with RNase R (RR). The capillary electrophoresis results of the RNA produced by the constructed vectors Coxie-EVB107-B18-Gluc and Coxie-EVB107-B18-EGFP are shown in Figures 12C and 12D, respectively. The purity of the circularized RNA from the Coxie-EVB107-B18-Gluc vector is 80.2%, and the purity of the circularized RNA from the Coxie-EVB107-B18-EGFP vector is 68.3%. The capillary electrophoresis results of RNA produced by Coxie-HRVB3-B18-Gluc, Coxie-HRVB3-B18-EGFP, Coxie-EV-A-mut-B374-Gluc, and Coxie-EV-A-mut-B374-EGFP are shown in Figures 14E, 14F, 14G, and 14H, respectively. The purity of RNA after circularization using the Coxie-HRVB3-B18-Gluc vector was 80.2%, the Coxie-HRVB3-B18-EGFP vector was 78%, the Coxie-EV-A-mut-B374-Gluc vector was 76.3%, and the Coxie-EV-A-mut-B374-EGFP vector was 78.9%. The purity of RNA after circularization using different ORFs was not significantly different from that of published strategies.
[0294] Following the method steps in 1.8 of Example 1, the purified RNA obtained according to 1.5 was subjected to reverse transcription, PCR, and gel recovery, and the circular adapter sites were sequenced. The primers used for PCR and sequencing of Circ-EVB107-B18-Gluc were SEQ ID NO:53 and 60, and the sequencing results at the adapter sites are shown in SEQ ID NO:58. The primers used for PCR and sequencing of Circ-EVB107-B18-EGFP were SEQ ID NO:16 and 60, and the sequencing results at the adapter sites are shown in SEQ ID NO:59. The primers used for PCR and sequencing of Circ-HRVB3-B18-Gluc were SEQ ID NO:53 and 72, and the sequencing results at the adapter sites are shown in SEQ ID NO:68. The primers used for PCR and sequencing of Circ-HRVB3-B18-EGFP were SEQ ID NO:16 and 72, and the sequencing results at the adapter sites are shown in SEQ ID NO:69. The primers used for PCR and sequencing of Circ-EV-A-mut-B374-Gluc are SEQ ID NO:53 and 73, and the sequencing results at the adapter are shown in SEQ ID NO:70. The primers used for PCR and sequencing of Circ-EV-A-mut-B374-EGFP are SEQ ID NO:16 and 73, and the sequencing results at the adapter are shown in SEQ ID NO:71. It can be seen that the cyclized product does not contain any exogenous residual sequences (i.e., no self-splicing type adjacent exon residues).
[0295] Circ-EVB107-B18-Gluc sequencing results
[0296] (SEQ ID NO:58, lowercase represents the Gluc sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates a connection point)
[0297] Sequencing results of Circ-EVB107-B18-EGFP
[0298] (SEQ ID NO:59, lowercase represents the Gluc sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates a connection point)
[0299] Circ-HRVB-B18-Gluc sequencing results
[0300] (SEQ ID NO:68, lowercase represents the Gluc sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates a connection point)
[0301] Circ-HRVB-B18-EGFP sequencing results
[0302] (SEQ ID NO:69, lowercase represents the EGFP sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates the linker)
[0303] Circ-EV-A-mut-B374-Gluc sequencing results
[0304] (SEQ ID NO:70, lowercase represents the Gluc sequence, uppercase represents the IRES sequence, where the bold underline is the 5' IRES sequence, and / indicates a connection point)
[0305] Circ-EV-A-mut-B374-EGFP sequencing results
[0306] (SEQ ID NO:71, lowercase is the EGFP sequence, uppercase is the IRES sequence, where the bold underline is the 5' IRES sequence, and / is the link) Based on the design strategy of preparing circular RNA vectors without exogenous residual sequences in this application, IRES can be replaced by IRES-CVB3 with IRES-EVB107, IRES-HRVB3 and IRES-EV-A-mut. The gene sequences of IRES-EVB107 and IRES-HRVB3 are truncated after the 18th base, and IRES-EV-A-mut is truncated at position 374. All three IRES can produce traceless circular RNA without exogenous residual sequences.
[0307] Example 7. Cellular expression assay and immunogenicity assay of circular RNA
[0308] To test the cellular expression level and immunogenicity of the circular RNA produced by the strategy in this application, we used linear mRNA-EGFP and circular RNA (circ-Ana-EGFP) prepared by a published strategy. [1] Cellular expression level and immunogenicity were tested using the circular RNA (circ-Coxie-B18-EGFP) prepared in Example 5 of this application.
[0309] Linear mRNA-EGFP, circ-Ana-EGFP prepared using a published strategy, and circ-Coxie-B18-EGFP prepared in Example 5 were transfected into HEK293T cells and A549 cells (purchased from the American Center for Type Culture Collection) for eukaryotic cell expression detection. Specifically, cells were seeded into 24-well plates using Dulbecco modified Eagle medium (DMEM, BI) supplemented with 10% fetal bovine serum (BI) and penicillin / streptomycin antibiotics (100 U / ml penicillin, 100 μg / ml streptomycin; Gibco), with 1 × 10⁶ cells per well. 5 Cells were cultured at 37°C, 5% CO2, and 90% relative humidity. On the second day, the above RNA samples (linear mRNA-EGFP, circ-Ana-EGFP, circ-Coxie-B18-EGFP) were added to 24-well plates at a concentration of 500 ng per well. The cells were then transfected with lipofectamine MessengerMax (Invitrogen, LRNA003). At this point, the cell culture must have >90% viability and reach 70% confluence.
[0310] Cells were placed on an inverted fluorescence microscope for fluorescence photography at 24, 48, 72, 6 and 8 days after transfection, and the fluorescence values were then read using ImageJ.
[0311] 48 h post-transfection, cell culture supernatant was collected. After centrifugation to remove A549 cell debris, 100 μL of the supernatant was taken and detected using the Human IL-6 ELISA kit according to the instructions (specific method: 100 μL of sample and standard were added to the ELISA plate and incubated at room temperature for 90 min; the liquid was discarded and 100 μL of biotinylated antibody was added and incubated for 60 min; the plate was washed; 100 μL of HRP enzyme conjugate was added and incubated for 30 min; the plate was washed; 90 μL of TMB was added and incubated for 15 min, and then stop solution was added to terminate the incubation). Finally, the OD value was detected at 450 nm using an ELISA reader, and the concentration of IL-6 in the sample was calculated by plotting a standard curve.
[0312] The results of cell transfection expression are shown in Figures 13A and 13B. The fluorescence intensity of circ-coxie-B18-EGFP prepared in Example 5 of this application was significantly higher than that of linear mRNA and circular RNA prepared using published strategies at 48 h and 8 days after transfection into HEK293T cells (Figure 13A). Similarly, the fluorescence intensity of circ-coxie-B18-EGFP prepared in Example 5 of this application was significantly higher than that of circular RNA prepared using published strategies at 48 h after transfection into A549 cells, and extremely significantly higher than that of linear mRNA and circular RNA prepared using published strategies at 8 days (Figure 13B).
[0313] The immunogenicity test results are shown in Figure 13C, where NC is the blank cell group and its relative IL-6 value is recorded as 1. The circ-coxie-B18-EGFP constructed using Coxella behnea-B18 in this application showed the lowest immunogenicity in A549 cells.
[0314] In summary, the circular RNA produced by constructing a vector using introns of the Coxsoe body used in this application and truncating IRES after the 18th base shows superior protein expression levels in mammalian cells compared to published strategies, and exhibits lower immunogenicity than circular RNAs from published strategies.
[0315] The sequences involved in this application are shown below.
[0316] Although this application has been described in conjunction with one or more embodiments, it should be understood that this application is not limited to these embodiments. The description in this application is intended to cover all variations and equivalents, all of which are included within the spirit and scope of the appended claims. All references cited herein are incorporated herein by reference in their entirety.
[0317] References
[0318] 1. Wesselhoeft, RA, Kowalski, PS & Anderson, DGEngineering circular RNA for potent and stable translation in eukaryotic cells. Nat. Commun. 9, 2629 (2018).
[0319] 2.Ko M, Choi H, Park C. Group I self-splicing intron in the recAgene of Bacillus anthracis. J Bacteriol. 2002Jul; 184(14):3917-22.
[0320] 3.Cech,TR1990.Self-splicing of group I introns.Annu.Rev.Biochem.59:543–568.
[0321] 4.Saldanha,R.,G.Mohr,M.Belfort,and A.M.Lambowitz.1993.Group I and group II introns.FASEB J.7:15–24
[0322] 5.Jacquier,A.1996.Group II introns:elaborate ribozymes.Biochimie 78:474–487.
[0323] 6.Chen R,Wang SK,Belk JA,Amaya L,Li Z,Cardenas A,Abe BT,Chen CK,Wender PA,Chang HY.Engineering circular RNA for enhanced protein production.Nat Biotechnol.2023Feb;41(2):262-272.
[0324] 7.Chen,C.,Wei,H.,Zhang,K.,Li,Z.,Wei,T.,Tang,C.,Yang,Y.,&Wang,Z.(2022).A flexible,efficient,and scalable platform to produce circular RNAs as new therapeutics.bioRxiv.
[0325] 8.Qiu,Z.,Zhao,Y.,Hou,Q.,Zhu,J.,Zhai,M.,Li,D.,Li,Y.,Liu,C.,Li,N.,Cao,Y.,Yang,J.,Sun,Z.,&Zuo,C.(2022).Clean-PIE:a novel strategy for efficiently constructing precise circRNA with thoroughly minimized immunogenicity to direct potent and durable protein expression.bioRxiv.
[0326] 9.Hausner G,Hafez M,Edgell DR.Bacterial group I introns:mobile RNA catalysts.Mob DNA.2014Mar 10;5(1):8.
[0327] 10.Yu H,Wen Y,Yu W,et al.Optimized circular RNA vaccines for superior cancer immunotherapy.Theranostics.2025;15(4):1420-1438.Published 2025Jan 2.doi:10.7150 / thno.104698
Claims
1. An RNA molecule comprising, in order from the 5' end to the 3' end: a 3' self-cleaving intron segment comprising a 3' splice site, a sequence of interest, and a 5' self-cleaving intron segment comprising a 5' splice site, wherein the elements are operably linked, wherein the 3' self-cleaving intron segment and the 5' self-cleaving intron segment are capable of circularizing the RNA molecule and removing the RNA molecule during circularization, wherein the 5' self-cleaving intron segment comprises an internal guide sequence, wherein i) the 5' end sequence of the internal guide sequence is reverse complementary to the 5' end sequence of the sequence of interest, or ii) the 3' end sequence of the internal guide sequence is initiated with a G base nucleotide and is reverse complementary to the 3' end sequence of the sequence of interest, wherein the 3' end sequence of the sequence of interest ends with a U base nucleotide, the initiating G base nucleotide of the 3' end sequence of the internal guide sequence is capable of forming a G:U pair with the U base nucleotide at the end of the 3' end sequence of the sequence of interest.
2. The RNA molecule of claim 1, wherein the 5' end sequence of the internal guide sequence is reverse complementary to 2-15 nt of the 5' end of the sequence of interest, or the 3' end sequence of the internal guide sequence is reverse complementary to 3-20 nt of the 3' end of the sequence of interest.
3. The RNA molecule of claim 1, wherein i) the 5' end sequence of the internal guide sequence is reverse complementary to the 5' end sequence of the sequence of interest, and ii) the 3' end sequence of the internal guide sequence is initiated with a G base nucleotide and is reverse complementary to the 3' end sequence of the sequence of interest, wherein the 3' end sequence of the sequence of interest ends with a U base nucleotide, the initiating G base nucleotide of the 3' end sequence of the internal guide sequence is capable of forming a G:U pair with the U base nucleotide at the end of the 3' end sequence of the sequence of interest.
4. The RNA molecule of claim 3, wherein, i) the sequence of interest comprises an open reading frame encoding a peptide or protein of interest, and a translation functional element, wherein the translation functional element is selected from the group consisting of a translation initiation element, and a translation enhancement element, wherein the sequence of interest comprises, in order from the 5' end to the 3' end (a) the translation functional element, and the open reading frame encoding the peptide or protein of interest, (b) the open reading frame encoding the peptide or protein of interest, and the translation functional element, (c) a 3' end sequence of the translation functional element, the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the translation functional element, wherein the 5' end sequence of the translation functional element and the 3' end sequence of the translation functional element, when arranged in this order, form the translation functional element, or (d) a 3' end sequence of the open reading frame encoding the peptide or protein of interest, the translation functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest, ii) the sequence of interest comprises a non-coding RNA sequence, which (a) comprises the non-coding RNA sequence, or (b) comprises, from 5' end to 3' end, a 3' end sequence of the non-coding RNA sequence, and a 5' end sequence of the non-coding RNA sequence, wherein the 5' end sequence of the non-coding RNA sequence and the 3' end sequence of the non-coding RNA sequence, when arranged in this order, form the non-coding RNA sequence, or iii) the sequence of interest comprises an open reading frame encoding a peptide or protein of interest, which (a) comprises the open reading frame encoding the peptide or protein of interest, or (b) comprises, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest.
5. The RNA molecule of claim 4, wherein the translation functional element is a translation initiation element, wherein the translation initiation element is an internal ribosome entry site (IRES).
6. The RNA molecule of claim 5, wherein the IRES is selected from the group consisting of a IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), or Human rhinovirus B6 (HRVB6).
7. The RNA molecule of claim 6, wherein the IRES comprises a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 4, SEQ ID NO: 61, SEQ ID NO: 74, SEQ ID NO: 75, or SEQ ID NO:
76.
8. The RNA molecule of claim 5, wherein the translation functional element is an IRES, wherein the IRES is selected from the group consisting of a IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), or Enterovirus A (EV-A), the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the IRES, the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the IRES, wherein the 3' end sequence of the IRES and the 5' end sequence of the IRES, respectively, comprise: (1) the nucleotide sequences set forth in SEQ ID NO: 26 and 27, (2) the nucleotide sequences set forth in SEQ ID NO: 40 and 41, (3) the nucleotide sequences set forth in SEQ ID NO: 46 and 47, (4) the nucleotide sequences set forth in SEQ ID NO: 55 and 56, (5) the nucleotide sequences set forth in SEQ ID NO: 62 and 63, or (6) the nucleotide sequences set forth in SEQ ID NO: 65 and 66.
9. The RNA molecule of claim 4, wherein the open reading frame encoding the peptide or protein of interest is an open reading frame encoding green fluorescent protein, wherein the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, the translation functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest comprise the nucleotide sequences set forth in AA and SEQ ID NO: 12, respectively, or wherein the open reading frame encoding the peptide or protein of interest is an open reading frame encoding luciferase, wherein the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, the translation functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, the 3' end sequence of the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest comprise the nucleotide sequences set forth in AA and SEQ ID NO: 29, or SEQ ID NO: 32 and 33, respectively.
10. The RNA molecule of claim 1, wherein the 3' self-splicing intron fragment and the 5' self-splicing intron fragment are derived from a self-splicing intron in the RecA gene of Bacillus anthracis, or a self-splicing intron in the 23S ribosomal gene of Coxiella burnetii.
11. The RNA molecule of claim 1, further comprising a 5' homology arm at the 5' end of the 3' self-splicing intron fragment comprising a 3' splice site, and a 3' homology arm at the 3' end of the 5' self-splicing intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 5' homology arm and the 3' homology arm are capable of complementary pairing.
12. A single-stranded DNA molecule comprising, in order from 5' end to 3' end: a 3' self-splicing intron fragment comprising a 3' splice site, a sequence of interest, and a 5' self-splicing intron fragment comprising a 5' splice site, wherein each element is operably linked, wherein the 3' self-splicing intron fragment and the 5' self-splicing intron fragment are capable of circularizing an RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule, wherein the 5' self-splicing intron fragment comprises an internal guide sequence, wherein i) a 5' end sequence of the internal guide sequence is reverse-complementary to a 5' end sequence of the sequence of interest, or ii) a 3' end sequence of the internal guide sequence is initiated with a nucleotide of base G and is reverse-complementary to a 3' end sequence of the sequence of interest, wherein the 3' end sequence of the sequence of interest ends with a nucleotide of base T, wherein the initiation nucleotide of base G at the 3' end sequence of the internal guide sequence of an RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule is capable of forming a G:U pair with the nucleotide of base U at the end of the 3' end sequence of the transcribed sequence of interest.
13. The single-stranded DNA molecule of claim 12, wherein i) the 5' end sequence of the internal guide sequence is reverse complementary to the 5' end sequence of the sequence of interest, and ii) the 3' end sequence of the internal guide sequence starts with a nucleotide that is a G and is reverse complementary to the 3' end sequence of the sequence of interest, wherein the 3' end sequence of the sequence of interest ends with a nucleotide that is a T, wherein the starting nucleotide that is a G of the 3' end sequence of the internal guide sequence of the RNA molecule transcribed from the complementary strand of the single-stranded DNA molecule is capable of forming a G:U pair with the nucleotide that is a U at the end of the 3' end sequence of the sequence of interest transcribed.
14. The single-stranded DNA molecule of claim 13, wherein, i) the sequence of interest comprises a sequence of a translational functional element and an open reading frame encoding a peptide or protein of interest, wherein the translational functional element is selected from the group consisting of a translational initiation element and a translational enhancer, wherein the sequence of interest comprises, from 5' end to 3' end, (a) the sequence of the translational functional element and the open reading frame encoding the peptide or protein of interest, (b) the open reading frame encoding the peptide or protein of interest and the sequence of the translational functional element, (c) the 3' end sequence of the sequence of the translational functional element, the open reading frame encoding the peptide or protein of interest, and the 5' end sequence of the sequence of the translational functional element, wherein the 5' end sequence of the sequence of the translational functional element and the 3' end sequence of the sequence of the translational functional element, when arranged in this order, form the sequence of the translational functional element, or (d) the 3' end sequence of the open reading frame encoding the peptide or protein of interest, the sequence of the translational functional element, and the 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest, ii) the sequence of interest comprises a sequence of non-coding DNA that (a) comprises the sequence of non-coding DNA, or (b) comprises, from 5' end to 3' end, the 3' end sequence of the sequence of non-coding DNA and the 5' end sequence of the sequence of non-coding DNA, wherein the 5' end sequence of the sequence of non-coding DNA and the 3' end sequence of the sequence of non-coding DNA, when arranged in this order, form the sequence of non-coding DNA, iii) the sequence of interest comprises an open reading frame encoding a peptide or protein of interest that (a) comprises the open reading frame encoding the peptide or protein of interest, or (b) comprises, from 5' end to 3' end, the 3' end sequence of the open reading frame encoding the peptide or protein of interest and the 5' end sequence of the open reading frame encoding the peptide or protein of interest, wherein the 5' end sequence of the open reading frame encoding the peptide or protein of interest and the 3' end sequence of the open reading frame encoding the peptide or protein of interest, when arranged in this order, form the open reading frame encoding the peptide or protein of interest, iv) the sequence of interest comprises a single cloning site or a multiple cloning site, or v) the sequence of interest comprises a monoclonal site or a polyclonal site, and a sequence of a translational functional element, which comprises, from 5' end to 3' end, a 3' end sequence of the sequence of the translational functional element, the monoclonal site or the polyclonal site, and a 5' end sequence of the sequence of the translational functional element, wherein the 5' end sequence of the sequence of the translational functional element and the 3' end sequence of the sequence of the translational functional element form the sequence of the translational functional element when arranged in this order.
15. The single-stranded DNA molecule of claim 14, wherein the translational functional element is a translational initiation element, wherein the translational initiation element is an internal ribosome entry site (IRES).
16. The single-stranded DNA molecule of claim 15, wherein the IRES is selected from the group consisting of an IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), or Human rhinovirus B6 (HRVB6).
17. The single-stranded DNA molecule of claim 15, wherein the translational functional element is an IRES, wherein the IRES is an IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), or Enterovirus A (EV-A), the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the sequence of the IRES, the open reading frame encoding the peptide or protein of interest, and a 5' end sequence of the sequence of the IRES, or comprises, from 5' end to 3' end, a 3' end sequence of the sequence of the IRES, the monoclonal site or the polyclonal site, and a 5' end sequence of the sequence of the IRES, wherein the 3' end sequence of the sequence of the IRES and the 5' end sequence of the sequence of the IRES comprise, respectively: (1) the nucleotide sequences set forth in SEQ ID NOs: 26 and 27, (2) the nucleotide sequences set forth in SEQ ID NOs: 40 and 41, (3) the nucleotide sequences set forth in SEQ ID NOs: 46 and 47, (4) the nucleotide sequences set forth in SEQ ID NOs: 55 and 56, (5) the nucleotide sequences set forth in SEQ ID NOs: 62 and 63, or (6) the nucleotide sequences set forth in SEQ ID NOs: 65 and 66.
18. The single-stranded DNA molecule of claim 15, wherein the open reading frame encoding the peptide or protein of interest is an open reading frame encoding green fluorescent protein, wherein the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, the sequence of the translational functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, the 3' end sequence of the open reading frame encoding the peptide or protein of interest and the 5' end sequence of the open reading frame encoding the peptide or protein of interest comprise, respectively, AA and the nucleotide set forth in SEQ ID NO: 12, or wherein the open reading frame encoding the peptide or protein of interest is an open reading frame encoding luciferase, and wherein the sequence of interest comprises, from 5' end to 3' end, a 3' end sequence of the open reading frame encoding the peptide or protein of interest, a sequence of the translation functional element, and a 5' end sequence of the open reading frame encoding the peptide or protein of interest; preferably, the 3' end sequence of the open reading frame encoding the peptide or protein of interest and the 5' end sequence of the open reading frame encoding the peptide or protein of interest comprise the nucleotide sequences set forth in AA and SEQ ID NO: 29, or SEQ ID NOs: 32 and 33, respectively.
19. The single-stranded DNA molecule of claim 12, wherein the 3' self-splicing intron fragment and the 5' self-splicing intron fragment are derived from a self-splicing intron in a RecA gene of Bacillus anthracis, or a self-splicing intron in a 23S ribosomal gene of Coxiella burnetii.
20. The single-stranded DNA molecule of claim 19, wherein the 3' self-splicing intron fragment comprises the nucleotide sequence set forth in SEQ ID NO: 10 or 21.
21. The single-stranded DNA molecule of claim 12, further comprising a sequence of a 5' homology arm at the 5' end of the 3' self-splicing intron fragment, and a sequence of a 3' homology arm at the 3' end of the 5' self-splicing intron fragment, wherein the 5' homology arm and the 3' homology arm are capable of complementary pairing.
22. The single-stranded DNA molecule of claim 20, further comprising a sequence of a RNA polymerase promoter at the 5' end of the sequence of the 5' homology arm; or further comprising a sequence of a restriction enzyme site at the 3' end of the sequence of the 3' homology arm.
23. A single-stranded DNA molecule, the complementary strand of which is capable of being transcribed into the RNA molecule of any one of claims 1-11.
24. A double-stranded DNA molecule comprising i) the single-stranded DNA molecule of any one of claims 12-23, and ii) a second strand complementary to the single-stranded DNA molecule.
25. A vector for preparing a circular RNA, comprising the RNA molecule of any one of claims 1-11, the single-stranded DNA molecule of any one of claims 12-23, or the double-stranded DNA molecule of claim 24.
26. A method for preparing a circular RNA, comprising: i) incubating the RNA molecule of any one of claims 1-11 under suitable conditions, wherein the suitable conditions comprise the presence of magnesium ions and guanosine triphosphate (GTP), or ii) (a) in vitro transcribing RNA from the complementary strand of the single-stranded DNA molecule of any one of claims 12-23, the double-stranded DNA molecule of claim 24, or the vector of claim 25 under suitable conditions, and (b) incubating the RNA under suitable conditions, wherein the suitable conditions in step (b) comprise the presence of magnesium ions and guanosine triphosphate (GTP).
27. A host cell comprising the RNA molecule of any one of claims 1-11, the single-stranded DNA molecule of any one of claims 12-23, the double-stranded DNA molecule of claim 24, or the vector of claim 25.
28. Use of the RNA molecule of any one of claims 1-11, the single-stranded DNA molecule of any one of claims 12-23, or the double-stranded DNA molecule of claim 24 in the manufacture of a circular RNA molecule that does not comprise an exon adjacent to a 3' self-cleaving intron segment or a 5' self-cleaving intron segment.
29. A circular RNA molecule produced according to the method of claim 26.
30. The circular RNA molecule of claim 29, which does not comprise an exon adjacent to a 3' self-cleaving intron segment or an exon adjacent to a 5' self-cleaving intron segment.
31. A circular RNA comprising a sequence of interest, optionally a translation functional element, and optionally an exon adjacent to a 3' self-cleaving intron segment or an exon adjacent to a 5' self-cleaving intron segment.
32. The circular RNA of claim 31, which comprises a sequence of interest, a translation functional element, and an exon adjacent to a 3' self-cleaving intron segment.
33. The circular RNA of claim 31, which comprises a sequence of interest, a translation functional element, and an exon adjacent to a 5' self-cleaving intron segment.
34. The circular RNA of claim 31, which has a sequence consisting of a sequence of interest and a translation functional element; or a sequence consisting of a sequence of interest.
35. The circular RNA of claim 34, wherein the translation functional element is a translation initiation element; preferably, the translation initiation element is an internal ribosome entry site (IRES).
36. The circular RNA of claim 35, wherein the IRES is selected from the group consisting of a IRES of Coxsackievirus B3 (CVB3), Enterovirus B107 (EVB107), Human rhinovirus B3 (HRVB3), Enterovirus A (EV-A), and Human rhinovirus B6 (HRVB6).
37. A composition comprising the circular RNA of any one of claims 29-36 and an open- circle RNA of the circular RNA, wherein, The molar percentage of the circular RNA in the composition is at least 30%, preferably at least 40%.
38. The composition of claim 37, wherein the molar percentage of the circular RNA in the composition is determined by capillary electrophoresis experiment.
Citation Information
Patent Citations
Circular RNA for translation in eukaryotic cells
CN112399860A
Circular RNA, composition and method for treating disease
CN118256490A
Circular rnas and preparation methods thereof
US20240093185A1
Compositions and methods for circular RNA affinity purification
WO2023242425A1