Universal'scar '-free RNA cyclization method

By modifying the internal guide sequence and UTR sequence screening of Anabaena type I ribozyme, a circularization system without "scar" was constructed, which solved the immunogenicity problem of PIE method, achieved efficient circularization and translation enhancement, and was suitable for the preparation of circular RNAs of different sequences.

CN120290557APending Publication Date: 2025-07-11PROXYBIO THERAPEUTICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510043781.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing circular RNA preparation methods such as PIE method have problems with high immunogenicity and nonspecific exogenous sequence introduction, which affects their biological activity and application.

Method used

By modifying the internal guide sequence of Anabaena-derived type I ribozyme, identifying cyclized fragments without exogenous sequences, and combining UTR sequence screening and IRES component optimization, a "scar"-free cyclization system is built to ensure high cyclization efficiency and translation enhancement effect.

Benefits of technology

It achieves efficient cyclization efficiency and translation enhancement, avoids immunogenicity, is suitable for cyclization needs of different sequences, and is suitable for industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290557A_ABST
    Figure CN120290557A_ABST
Patent Text Reader

Abstract

The invention discloses a universal and'scar '-free RNA cyclization method, which comprises a recombinant nucleic acid molecule for preparing circular RNA, the recombinant nucleic acid molecule comprises the following elements arranged in sequence along the direction from 5'to 3': an optional 5 'homologous arm, a 3' semi-intron fragment, a cyclization fragment, a 5 'semi-intron fragment and an optional 3' homologous arm; the 5'end of the cyclization fragment comprises a 3 'cyclization recognition fragment, and the 3' end of the cyclization fragment comprises a 5 'cyclization recognition fragment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention claims the priority of a prior application with the patent application number 2024100463312 and the invention title "A General and 'Scar'-Free RNA Cyclization Method", which was filed with the China National Intellectual Property Administration on January 11, 2024. The entire text of this prior application is incorporated herein by reference. Technical Field

[0002] This disclosure belongs to the field of biochemistry and specifically relates to a method for synthesizing circular RNA. Background Art

[0003] There are usually three methods for preparing circular RNA: chemical method, ligase method, and ribozyme method. The chemical method has been phased out due to its extremely low cyclization efficiency; the ligase method often has the problem of easy intermolecular ligation, which reduces the cyclization rate and also poses great challenges to product purification. In addition, the fragments that can be cyclized by the ligase method are often very short, and these disadvantages limit the use of the ligase method. The ribozyme method is a method that uses the self-catalytic action of ribozymes to produce circular RNA. Due to its own structure, ribozymes have an enzyme-like catalytic activity, and after back-splicing, circular RNA is produced. After years of development, the ribozyme method has undergone a re-design of its structure, and the positions of the introns and exons in the original ribozyme sequence have been swapped, so it is called the PIE method (permuted Intron-Exon). The PIE method is currently the most efficient method for in vitro preparation of circular RNA. The PIE method was initially proposed by Michael D. Been in 1992 [1] , and he successfully prepared circular RNA through the back-splicing process by modifying the type I intron from Anabeana. In 2018, Anderson et al. modified the PIE system [2] , adding elements such as homology arms and "spacers" that improve the cyclization efficiency, greatly enhancing the efficiency of the PIE method for preparing circular RNA and successfully cyclizing RNA with a maximum length of 5 kb.

[0004] However, in recent years, as this technology has received attention, some researchers have found that the circular RNA prepared by the PIE method has strong immunogenicity. For example, Howard Y. Chang et al. [3] in 2017 [4]It was found that the circular RNAs prepared by the PIE method generated strong immunogenicity in cells. By measuring cytokines such as RIG-I, it was found that the immunogenicity induced by the circular RNAs prepared by the PIE method was comparable to that of unmodified linear mRNAs. It was found that the immunogenicity was caused by the inevitable insertion of a foreign Exon sequence into the RNA during the preparation of circular RNAs by the PIE method; in addition, Ling-Ling Chen [5] et al. also found the same phenomenon and also found that the source of immunogenicity was the "scar" sequence introduced during the preparation by the PIE method. Further research showed that the immunogenicity was induced by the formation of an unconventional duplex structure by the "scar" sequence. Therefore, from the above research, although the PIE method has a high cyclization efficiency, it has the disadvantage of inducing strong immunogenicity, which will greatly hinder the application of circular RNAs. The immunogenicity is mainly caused by the "scar" sequence. In addition, more and more literature has found that circular RNAs have regulatory functions and can be used as potential therapeutic molecules [6] , for the synthesis of therapeutic circular RNA molecules, if the PIE method is used, it will introduce additional sequences of nearly 200 nt including the "scar" fragment and spacer sequence, which will obviously affect the biological activity of such therapeutic circular RNA molecules [7] .

[0005] Therefore, by re-engineering the PIE method and developing a "scar"-free cyclization method, it is possible to maintain the excellent characteristics of the PIE method, avoid high immunogenicity, and at the same time serve as a method for preparing therapeutic circular RNAs in vitro. At the same time, in subsequent designs, this system is designed as a general system, and no special design is required for different sequences to be cyclized, which is more convenient.

[0006] Circular nucleic acids are formed by the back-splicing of linear plasmids (cDNA templates) through biological reactions to form a closed-loop structure by covalent bonds. The main impurities mainly come from the in vitro transcription (IVT) stage: T7 enzyme, DNA residues, dNTPs, precursor linear, introns, open rings, polymers and circular isomers, etc. It can be known from animal cell experiments that precursor linear, introns, polymers and circular isomers affect cytotoxicity; open rings affect cell activity, and at the same time, too many open rings have strong immunogenicity. By controlling the reaction vessels (bioreactors and wave reactors, etc.), reaction temperature, stirring speed and reaction time used in the in vitro transcription (IVT) process, as well as optimizing the DNA template, T7 enzyme, dNTP and PEG concentrations added during the IVT reaction process, the production of Naked RNA can be controlled and the yield of circular nucleic acids can be determined.

[0007] circRNAs (Circular RNAs) are a class of non-coding RNA molecules that do not have a 5'-terminal cap and a 3'-terminal poly(A) tail and form a circular structure by covalent bonds. circRNA is a new type of RNA different from traditional linear RNA, with a closed circular structure and is abundantly present in the eukaryotic transcriptome. Most circular RNAs are composed of exon sequences, are conserved in different species, and have expression specificity in tissues and different developmental stages. Since circular RNAs are insensitive to nucleases, they are more stable than linear RNAs, which gives circular RNAs obvious advantages in the development and application of new clinical therapeutic drugs.

[0008] The laboratory level reported in the literature uses the A-Tailing / LiCl-buffer RNase R digestion method to obtain higher linear RNA removal and circRNAs enrichment efficiency, and obtain circRNAs with higher purity. However, the total yield of circRNAs < 8%, which is not suitable for large-scale production in the process. [8] 。 Summary of the Invention

[0009] Technical problems to be solved by the invention

[0010] By exploring the molecular mechanism of type I ribozymes from Anabaena, it was found that scarless cyclization can be achieved by modifying the IG sequence. At the same time, the IG sequence and the splicing site need to meet the following requirements:

[0011] ① The IG sequence needs to meet the sequence feature of "GNN";

[0012] ② When the IG sequence recognizes the 5'splice site, it needs to ensure base complementary pairing or "GU" wobble pairing;

[0013] ③ The 3'splice site is not conserved, but when it is base complementary paired with the sequence near the 5'splice site, it will greatly affect the cyclization efficiency;

[0014] ④ The two splicing sites need to have homologous arms with a minimum length of 3nt to stabilize the P1 duplex

[0015] Based on the above design principles, Firefly Luciferae (1653 nt) and eGFP (780 nt) were successfully circularized, and circular RNAs without "scar" containing only "IRES + ORF" were constructed. At the same time, it was proved by 2% EX-gel electrophoresis separation that they still retained a high circularization efficiency, and large fragment RNAs could be circularized effectively; both cell experiments and animal experiments proved that the constructed circular RNAs without "scar" were successfully expressed in vivo. This proved the universality and feasibility of this "scar"-free circularization method.

[0016] In addition, since the selection of splicing sites and the modification of IG sequences are required for circularization of different ORFs, which is not convenient. At the same time, we found in the experiment that some spacer sequences can not only improve the circularization efficiency but also enhance the translation of circular RNAs. We named these spacer sequences UTR sequences. We screened the UTR sequences and found that the sequence with the best enhancement effect on the translation of circular RNAs enhanced the translation by 3.7 times compared with the original circular RNA sequence of the PIE method. By modifying the IG sequence with the previous "scar"-free circularization method, a "scar"-free universal circularization system containing only "IRES + ORF + UTR" was constructed. The circular RNAs constructed by this method had a circularization efficiency comparable to that of the PIE method and a stronger translation ability, which has been verified at the cellular and animal levels.

[0017] The present invention also modified based on the CBV3 IRES sequence. First, the present invention explored the optimal position for inserting the IRES element into the CVB3 IRES sequence. The present invention designed 11 positions for inserting gene sequences, and it was found that 4 insertion positions had translation effects. The 4 specific insertion positions were respectively: the IRES element was inserted between domain Ⅰ and domain Ⅱ, named IRES-1; the IRES element was inserted into the stem-loop structure of domain Ⅱ, named IRES-2; the IRES element was inserted into the stem-loop structure (Distal loop) of domain Ⅳ, named IRES-3; the IRES element was inserted into the stem-loop structure (Proximal loop) of domain Ⅳ, named IRES-4. The IRES element can enhance the binding of the CVB3 IRES sequence to ribosomes, thereby promoting the translation of circRNA.

[0018] IRES elements are present in circRNA, fold into a structure similar to the initial tRNA, recruit more ribosomes, and bind translation regulatory factors such as ITAF. Then, the ribosomes are introduced into the interior of circRNA to bind and initiate protein translation. Ribosomes are highly diverse protein structures found in all cells. Under the action of translation regulatory factors such as ITAF, the ribosomes are introduced into the internal structure of circRNA to bind and initiate. Secondly, the present invention screens 4 IRES elements and then explores the influence of IRES elements with different sequences having the same enhancing effect on the translation effect of cicRNA. These 4 IRES elements all have the function of recruiting ribosomes, thus affecting the translation of circRNA. The 4 screened IRES element sequences are respectively cloned into a DNA template, and the composition of the DNA template is the same as above, and the optimal sequence is screened to obtain a gene sequence with high protein expression. Finally, based on the optimal position of the IRES element and the screening of the IRES element, a series of repetitive sequence designs are carried out on the IRES element to obtain the sequence for the IRES element to achieve its function. Compared with the original sequence, IRES-1-A1-X2 in the screened sequence is increased by about 2.3 times.

[0019] Finally, the conditions for the in vitro IVT reaction of Circular RNA were optimized to make the cyclization rate ≥75%, and the content of Naked RNA ≤3%; the purification method of circular nucleic acid was improved to ensure that the purity of circular nucleic acid ≥95%, the content of polymers, precursor linear, intron and isomers <1%, the total purification yield ≥30%, and the theoretical total yield of circular nucleic acid ≥70%, ensuring that the in vitro IVT reaction and purification of CircularRNA are suitable for industrial scale-up production.

[0020] Specific solution of the invention

[0021] To solve the deficiencies of the prior art, the present disclosure specifically provides:

[0022] The present disclosure first provides a recombinant nucleic acid molecule for preparing circular RNA. Along the 5' to 3' direction, the recombinant nucleic acid molecule includes elements that are operably linked in sequence:

[0023] An optional 5' homologous arm, a 3' semi-intron fragment, a cyclization fragment, a 5' semi-intron fragment, and an optional 3' homologous arm;

[0024] The 5' end of the cyclization fragment contains the 3' loop recognition fragment of the intron, and the 3' end contains the 5' loop recognition fragment of the intron;

[0025] The 5'-half intron fragment and the 3'-half intron fragment are derived from group I introns, preferably group I introns from Anabeana, preferably from Anabeana tRNA leu .

[0026] The 5'-half intron fragment and the 3'-half intron fragment are used to form an intron sequence in the 5' to 3' direction; the nucleotide sequence of the 5'-half intron fragment contains a partial sequence of the intron sequence close to the 5' direction, and the nucleotide sequence of the 3'-half intron fragment contains the remaining part of the intron sequence close to the 3' direction, and the 5'-half intron fragment contains an internal guide (IG) sequence;

[0027] Wherein:

[0028] (1) The IG sequence of the intron is 5'-GNN-3' of 3 nt, and the 5'-loop recognition fragment is 5'-NNY-3' of 3 nt, where Y is C or T / U; N is an optional nucleotide, and the nucleotides corresponding to the IG sequence and the 5'-loop recognition fragment have strict base complementary pairing or GU wobble pairing;

[0029] (2) At least 2 nt of the 3'-loop recognition fragment, and it is not complementary to the 5'-loop recognition fragment and its adjacent sequences;

[0030] (3) The sequences adjacent to the 5'-loop recognition fragment and the 3'-loop recognition fragment contain at least 3 nt of internal homology sequences.

[0031] In some specific embodiments of the present disclosure, the recombinant nucleic acid molecule satisfies at least one of the following conditions:

[0032] 1) The second nucleotide of the IG sequence is not A and / or the third nucleotide is not G;

[0033] 2) When the 3'-loop recognition fragment is located in the spacer sequence, it is not AAAA, AA, UUUU, UAAA, CAAA, or GAAA.

[0034] In the specific embodiments of the present disclosure, the 3'-loop recognition fragment does not have obvious homology with the sequence adjacent to the 5'-loop recognition fragment, that is, it is not complementary to the sequence adjacent to the 5'-loop recognition fragment.

[0035] In the specific embodiments of the present disclosure, the 5' end of the circularized fragment has a 3'-loop recognition fragment-internal homology sequence fragment, and the 3' end has an internal homology sequence fragment-5'-loop recognition fragment, and there are optionally 0, 1, 2, 3, 4, or 5 unpaired nucleotides between the loop recognition fragment and the internal homology sequence fragment.

[0036] In a further preferred embodiment, the homologous sequence is 3 nt, 4 nt or 5 nt.

[0037] In a further preferred embodiment, the length of the 3'-loop recognition fragment is 2 nt.

[0038] In certain specific embodiments of the present disclosure, the circularized fragment comprises a first part of the target polypeptide coding region or non-coding region, a translation initiation element, and a second part of the target polypeptide coding region or non-coding region. When designing the circularized fragment, the 5'-loop recognition fragment and the 3'-loop recognition fragment divide the target polypeptide coding region or non-coding region into two parts, that is, the target polypeptide coding region or non-coding region is cut as follows: the second part (including the 5'-loop recognition fragment) | (including the 3'-loop recognition fragment) the first part (see Attachments Figure 8 a and 8c), and then the following circularized fragment is formed: the 3'-loop recognition fragment and the first part of the target polypeptide coding region or non-coding region - translation initiation element - the second part of the target polypeptide coding region or non-coding region and the 5'-loop recognition fragment. The further formed recombinant nucleic acid molecule structure is:

[0039] Optional 5'-homologous arm, 3'-semi-intron fragment, 3'-loop recognition fragment and the first part of the target polypeptide coding region or non-coding region - translation initiation element - the second part of the target polypeptide coding region or non-coding region and the 5'-loop recognition fragment, 5'-semi-intron fragment and optional 3'-homologous arm (Attachments Figure 8 c);

[0040] After the 3'-loop recognition fragment of the first part and the second part including the 5'-loop recognition fragment are spliced by introns, they are reconnected into a complete target polypeptide coding region or non-coding region.

[0041] In certain specific embodiments of the present disclosure, the circularized fragment comprises a first part of the translation initiation element, the target polypeptide coding region or non-coding region, and a second part of the translation initiation element, that is, the 5'-loop recognition fragment and the 3'-loop recognition fragment divide the translation initiation element into two parts. After intron splicing, the 5' part and the 3' part of the translation initiation element are connected into a complete translation initiation element region.

[0042] In certain specific embodiments of the present disclosure, the circularized fragment comprises a 5'-spacer part, an optional translation initiation element, the target polypeptide coding region or non-coding region, and a 3'-spacer part.

[0043] In certain specific embodiments of the present disclosure, the 5'-spacer sequence is a 5'-UTR sequence, which has both the function of "spacing" and improving translation efficiency.

[0044] In certain specific embodiments of the present disclosure, the 3' spacer sequence is a 3' UTR sequence, which has the functions of both "spacer" and improving translation efficiency.

[0045] The 5' looping recognition fragment is located in the 5' spacer and is recognized by the modified IG sequence, and the 3' looping recognition fragment is located in the 3' spacer, thereby forming a universal cyclization system without the need to edit or recognize the IRES sequence or the target polypeptide coding region or non-coding region.

[0046] In certain specific embodiments of the present disclosure, the number of exogenously introduced nucleotides on the 5' spacer sequence and the 3' spacer sequence is less than 10.

[0047] In certain specific embodiments of the present disclosure, three nucleotides are inserted at the third or fourth nucleotide of the 5' spacer sequence as internal homologous sequences, and the two or three nucleotides of the 5' spacer sequence itself are used as the 3' loop recognition fragment; 6 nucleotides are inserted at the 3' end of the 3'UTR, wherein the 1st to 3rd nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, the 4th to 6th sequences are NNY, and the bases corresponding to the NNY and IG sequences have strict base complementary pairing or GU wobble pairing.

[0048] In certain specific embodiments of the present disclosure, 5 nucleotides are inserted at the 5' end of the 5' spacer sequence, wherein the 1st to 2nd nucleotides serve as 3' looping recognition fragments, the 3rd to 5th nucleotides serve as internal homologous sequences, and the two or three nucleotides of the 5' spacer sequence itself serve as 3' looping recognition fragments; 6 nucleotides are inserted at the 3' end of the 3' spacer sequence, wherein the 1st to 3rd nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, the 4th to 6th sequences are NNY, and the bases corresponding to the NNY and IG sequences have strict base complementary pairing or GU wobble pairing.

[0049] Preferred 5' spacer moieties are selected from:

[0050]

[0051] Preferred 3' spacer moieties are selected from:

[0052]

[0053] Preferably, the 5' spacer sequence and the 3' spacer sequence are used alone or together; preferably, the 5' spacer sequence and the 3' spacer sequence are both polyA+8CA, preferably the polyA fragment contains 10-100 A, more preferably 60-80 A, most preferably 61 A at the 5' end and 71 A at the 3' end.

[0054] Most preferably, it is a combination of 5'61A8CA and 3'71A8CA.

[0055] In another embodiment, if present, the IRES sequence is an IRES sequence of: Taura syndrome virus, Triatoma virus, Theiler's encephalomyelitis virus, simian virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, reticuloendotheliosis virus, Forman poliovirus 1, Autographa californica multiple nucleopolyhedrovirus, Kashmir bee virus, human rhinovirus 2, Homalodisca coagulata virus-1, human immunodeficiency virus type 1, Homalodisca coagulata virus-1, Pediculus humanus corporis virus, hepatitis C virus, hepatitis A virus, hepatitis G virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinovirus, Ectropis obliqua-like virus, encephalomyocarditis virus (EMCV), Drosophila C virus, tobacco mosaic virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, Hibiscus chlorotic ringspot virus, classical swine fever virus, human FGF2, human SFTPA1, human AML1 / RUNX1, Drosophila antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAP1, human cmyc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kip1, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, saliva virus, Coxsackievirus, Echovirus, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, picornavirus, turnip crinkle virus, aptamer of eIF4G, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVB1 / 2). In another embodiment, the IRES is the IRES sequence of Coxsackievirus B3 (CVB3). In another embodiment, the IRES is the IRES sequence of encephalomyocarditis virus.

[0056] In a specific embodiment of the present disclosure, the recombinant nucleic acid molecule comprises a CVB3 IRES sequence into which an IRES enhancer element is inserted.

[0057] The insertion positions of the IRES enhancer elements are as follows: between domain I and domain II of the IRES sequence (named IRES-1), at the stem-loop structure of domain II of the IRES sequence (named IRES-2), at the stem-loop structure (Distal loop) of domain IV of the IRES sequence (named IRES-3), and at the stem-loop structure (Proximal loop) of domain IV of the IRES sequence (named IRES-4).

[0058] The IRES enhancer elements are as follows:

[0059] Serial number IRES sequence A1 CCGGCGGGT A2 GTTTCATTTTATTCCTATAC A3 CAATTGAGAGATCGTTACCATATAGCTA A4 ATTTTATTCCTA

[0060] or repetitive sequences of the above sequences, for example:

[0061] Serial number IRES repeat sequence A1-X2 CCGGCGGGTCCGGCGGGT A1-X5 CCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGT A2-X2 GTTTCATTTTATTCCTATACGTTTCATTTTATTCCTATAC .

[0062] Optionally, the translation initiation element sequence comprises one or a combination of two or more of the following sequences: IRES sequence, 5’UTR sequence, Kozak sequence, sequence containing m6A modification, complementary sequence of ribosomal 18S rRNA.

[0063] In certain specific embodiments of the present disclosure, the 5’ spacer sequence is a 5’UTR sequence, which has the functions of both “spacer” and improving translation efficiency.

[0064] Optionally, the target polypeptide coding region or non-coding region is a protein coding region encoding a human protein or a non-human protein.

[0065] Optionally, the protein coding region encodes an antibody.

[0066] Optionally, the human protein or non-human protein is selected from hFIX, SP-B, VEGF-A, human methylmalonyl-CoA mutase (hMUT), CFTR, cancer autoantigens, and gene editing enzymes such as Cpf1, zinc finger nuclease (ZFN), and transcription activator-like effector nuclease (TALEN).

[0067] Optionally, the protein is a protein for therapeutic use.

[0068] Optionally, the antibody is a human anti-HIV antibody.

[0069] Optionally, the antibody is a bispecific antibody.

[0070] Optionally, the bispecific antibody binds CD3 and CLDN6 or binds CD19 and CD22.

[0071] Optionally, the protein is a protein for diagnostic use.

[0072] Optionally, the protein coding region encodes Gauss luciferase (Gluc), firefly luciferase (Fluc), enhanced green fluorescent protein (eGFP), human erythropoietin (hEPO), or Cas9 endonuclease.

[0073] In a specific embodiment of the present disclosure, the protein coding region comprises at least two coding regions, wherein a linker is connected between any two adjacent coding regions; preferably, the linker is a polynucleotide encoding a 2A peptide.

[0074] Optionally, a translation initiation element is connected between any two adjacent coding regions; optionally, the translation initiation element located between any two adjacent coding regions comprises one or more combinations of the following sequences: IRES sequence, 5'UTR sequence, Kozak sequence, sequence containing m6A modification, complementary sequence of ribosomal 18S rRNA.

[0075] Wherein, the recombinant nucleic acid molecule further comprises an insertion element, and the insertion element is located upstream of the translation initiation element; the insertion element is selected from at least one of the following groups (i)-(iii):

[0076] (i) Transcription level regulatory element, (ii) translation level regulatory element, (iii) purification element;

[0077] Optionally, the insertion element comprises a sequence of one or more combinations of the following:

[0078] Untranslated region sequence, polyA sequence, aptamer sequence, riboswitch sequence, sequence binding to a transcriptional regulatory factor.

[0079] In a specific embodiment of the present disclosure, the recombinant nucleic acid has the homologous arms, wherein the length of each homologous arm is about 5-50 nucleotides; preferably, the length of each homologous arm is about 9-19 nucleotides.

[0080] In a specific embodiment of the present invention, the recombinant nucleic acid molecule is a precursor RNA molecule for preparing circular RNA.

[0081] In another specific embodiment of the present invention, the recombinant nucleic acid molecule is a DNA molecule that can obtain the above precursor RNA molecule by transcription.

[0082] The second aspect of the present disclosure provides a recombinant expression vector, wherein the recombinant expression vector comprises the above recombinant nucleic acid molecule.

[0083] The third aspect of the present disclosure provides a circular precursor RNA molecule, which is obtained by transcribing the recombinant expression vector, and the circular precursor nucleic acid molecule fragment includes: an optional 5' homologous arm, a 3' semi-intron fragment, a circularization fragment, a 5' semi-intron fragment, and an optional 3' homologous arm;

[0084] The 5' end of the circularization fragment contains a 3' circularization recognition fragment, and the 3' end contains a 5' circularization recognition fragment.

[0085] The fourth aspect of the present disclosure provides a method for preparing circular RNA in vitro,

[0086] Transcription step: Transcribe a circular precursor RNA molecule according to the above recombinant expression vector;

[0087] Circularization step: The circular precursor nucleic acid molecule undergoes a circularization reaction to obtain circular RNA;

[0088] Optionally, the method further includes a step of purifying the circular RNA.

[0089] The present disclosure also provides a purification method for preparing circular RNA, which includes ultrafiltration, and a combination of the same or different means of affinity chromatography, molecular sieve chromatography, and CHT;

[0090] Preferred combinations are: two affinity chromatographies, affinity chromatography and molecular sieve chromatography, or affinity chromatography and CHT chromatography;

[0091] Preferably, arginine and PEG are added to the ultrafiltration reagent or chromatography reagent; preferably, the reagent is a buffer; preferably, the PEG is PEG2000-8000.

[0092] The fifth aspect of the present disclosure provides a circular RNA molecule obtained from the above recombinant nucleic acid molecule or vector, or the above method.

[0093] The sixth aspect of the present disclosure provides a circular RNA, which, along the 5' to 3' direction, contains elements arranged in the following order:

[0094] A translation initiation element, a target polypeptide coding region or a non-coding region;

[0095] Optionally, the circular RNA contains a 5' spacer sequence and a 3' spacer sequence located between the 5' end of the translation initiation element and the 3' end of the coding element;

[0096] Each of the elements is defined as above.

[0097] Preferably, the circular RNA contains the above spacer sequence and an IRES sequence;

[0098] Preferably, the circular RNA contains modifications; preferably, the modification is m6A.

[0099] The seventh aspect of the present disclosure provides a composition, wherein the composition comprises the above-mentioned recombinant nucleic acid molecule, the above-mentioned recombinant expression vector, or the above-mentioned circular RNA; preferably, it comprises the above-mentioned circular RNA;

[0100] Optionally, the composition further comprises one or more pharmaceutically acceptable carriers;

[0101] Optionally, the pharmaceutically acceptable carrier is selected from lipids, polymers, or lipid-polymer complexes.

[0102] The eighth aspect of the present disclosure provides a method for expressing a target polypeptide in a cell, wherein the method comprises the step of introducing the above-mentioned circular RNA, or the above-mentioned composition, into the cell.

[0103] The ninth aspect of the present disclosure provides a method for preventing or treating a disease, wherein the method comprises administering the above-mentioned circular RNA, or the above-mentioned composition, to a subject.

[0104] The present disclosure also provides a design method for a recombinant nucleic acid molecule for preparing circular RNA, and the method comprises:

[0105] 1) Providing a candidate target polypeptide coding region or non-coding region, a translation initiation element, and optionally a 5' spacer sequence and a 3' spacer sequence;

[0106] 2) Searching for or setting a 5' cyclization recognition fragment, a 3' cyclization recognition fragment, and an internal homology sequence in the target polypeptide coding region or non-coding region, the translation initiation element, or the 5' spacer sequence and the 3' spacer sequence to design a cyclization fragment:

[0107] The searching operation is:

[0108] S-1) Searching for the following sequence in the above elements: 5' internal homology sequence - 5' cyclization recognition fragment - 3' cyclization recognition fragment - 3' internal homology sequence;

[0109] S-2) Splitting the corresponding elements at the 3' junction of the 5' cyclization recognition fragment and the 3' cyclization recognition fragment according to the search result to form a cyclization fragment;

[0110] The setting is to respectively set a 5' internal homology sequence - 5' cyclization recognition fragment sequence and a 3' cyclization recognition fragment - 3' internal homology sequence at both ends of the sequence 5' spacer sequence - translation initiation element - target polypeptide coding region or non-coding region - 3' spacer sequence to form a cyclization fragment;

[0111] Wherein:

[0112] A) the 5' internal homologous sequence and the 3' internal homologous sequence contain at least 3 nt of reverse complementary sequence;

[0113] B) the 3' looping recognition fragment is a sequence of at least 2 nt and has no significant homology with the adjacent sequence of the 5' looping recognition fragment;

[0114] C) the 5' looping recognition fragment is a 3 nt sequence and is NNY, wherein N is an optional nucleotide and Y is C or T;

[0115] wherein "-" represents a phosphodiester bond or 1-5 (e.g., 1, 2, 3, 4 or 5) optional nucleotides;

[0116] 3) According to the circularized fragment obtained in step 2), the intron IG sequence to be used is modified, wherein the IG sequence is 3 nt "GNN" and has strict base complementary pairing or GU wobble pairing with the bases corresponding to the 5' circularization recognition fragment;

[0117] 4) dividing the modified intron sequence into a 3' half intron fragment and a 5' half intron fragment, and arranging them in the following order to form a recombinant nucleic acid molecule for preparing circular RNA:

[0118] an optional 5' homology arm, a 3' half intron fragment, a circularized fragment, a 5' half intron fragment, and an optional 3' homology arm;

[0119] The 5' end of the circularization fragment comprises a 3' circularization recognition fragment, and the 3' end comprises a 5' circularization recognition fragment.

[0120] Preferably, the setting is to insert 3 nucleotides as internal homologous sequences at the third or fourth nucleotide of the 5' spacer sequence, and use the two or three nucleotides of the 5' spacer sequence as the 3' loop recognition fragment; insert 6 nucleotides at the 3' end of the 3'UTR, wherein the 1st to 3rd nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, the 4th to 6th sequences are NNY, and have strict base complementary pairing or GU wobble pairing with the bases corresponding to the IG sequence;

[0121] Also preferably, the setting is to insert 5 nucleotides at the 5' end of the 5' spacer sequence, wherein the 1st to 2nd nucleotides serve as 3' looping recognition fragments, the 3rd to 5th nucleotides serve as internal homologous sequences, and the two or three nucleotides of the 5' spacer sequence itself serve as 3' looping recognition fragments; and insert 6 nucleotides at the 3' end of the 3' spacer sequence, wherein the 1st to 3rd nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, the 4th to 6th sequences are NNY, and have strict base complementary pairing or GU wobble pairing with the bases corresponding to the IG sequence.

[0122] The present disclosure also provides a computer module, which stores a program for implementing the above design method.

[0123] The present disclosure also provides a CVB3 IRES sequence into which an IRES enhancing element is inserted;

[0124] Preferably, the insertion positions of the IRES enhancing element are: between domain Ⅰ and domain Ⅱ of the IRES sequence (named IRES-1), at the stem-loop structure of domain Ⅱ of the IRES sequence (named IRES-2), at the stem-loop structure (Distal loop) of domain Ⅳ of the IRES sequence (named IRES-3), and at the stem-loop structure (Proximal loop) of domain Ⅳ of the IRES sequence (named IRES-4);

[0125] Preferably, the IRES enhancing element is:

[0126] Serial number IRES sequence A1 CCGGCGGGT A2 GTTTCATTTTATTCCTATAC A3 CAATTGAGAGATCGTTACCATATAGCTA A4 ATTTTATTCCTA

[0127] or a repeat sequence of the above sequence:

[0128] Serial number IRES repeat sequence A1-X2 CCGGCGGGTCCGGCGGGT A1-X5 CCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGT A2-X2 GTTTCATTTTATTCCTATACGTTTCATTTTATTCCTATAC 。

[0129] The present invention also provides a mutant of Anabaena tRNA leu class I intron or a combination of fragments with self-splicing activity derived from the mutant, and comprising at least one of the following groups of mutations:

[0130] 1) The second base in the original IG sequence 5'-GAG-3' is mutated to a base other than A;

[0131] or / and

[0132] 2) The third base (G at 3') in the original IG sequence 5'-GAG-3' is mutated to a base other than G;

[0133] Preferably, the mutant comprises at least a mutation of the third base (G at 3');

[0134] Preferably, the mutant does not contain the original E1 and E2 sequences;

[0135] Preferably, it is a combination of SEQ ID No.44 and 45 with the above mutations.

[0136] Beneficial technical effects

[0137] 1. By screening and modifying the sequence of Anabaena type I intron, it is verified that by modifying the IG sequence to recognize the circularization sequence, circularization without exogenous sequences can be achieved. Among them, the original IG sequence is "GAG", where the first "G" is strictly conserved and can only recognize the 5' splice site through "GC" pairing or "GU wobble" pairing; the second A and the third G are not conserved and can be modified according to the specific nucleotides of the circularization sequence, as long as strict base complementary pairing or wobble pairing relationships are ensured. In summary, the IG sequence of Anabaena needs to satisfy "GNN" and can be modified by the required circularization sequence. In addition, the 3' splice site of the circularization sequence is not conserved. The Anabaena ribozyme system is different from other type I ribozyme systems, and its IG sequence does not recognize the 3' splice site. The same conclusion is also found in this disclosure. However, this disclosure further finds that at least 3 nt of nucleotides at the 3' splice site need to be non-homologous to the 5' splice site, otherwise the circularization efficiency will be affected. In addition, since the Anabaena ribozyme is derived from tRNA, its original Exon sequence is the stem structure of "clover" and naturally has 5 nt nucleotide length homology. When circularizing without exogenous sequences, a certain length of sequence homology around the splice site also needs to be retained. This disclosure finds that a homology length of at least 3 nt will not affect the circularization efficiency.

[0138] This disclosure also screened the UTR sequence. Here, the UTR refers to the spacer sequence in the original PIE technology, and its main function is to play a "spacer" role between the ribozyme and the IRES, thereby improving the circularization efficiency. This disclosure finds that the UTR sequence also has the functions of "spacing" and improving translation efficiency. In this disclosure, 10 sequences were first screened at both the 5' UTR and the 3' UTR. It was found that 41A8CA and 51A8CA at the 5' UTR can both improve translation efficiency, and 15A8CA at the 3' UTR can also improve translation efficiency. At the same time, it was found that the improvement of translation efficiency is proportional to the number of A. Therefore, after increasing the number of A, the translation efficiency was further improved, and the optimal sequences of 5' 61A8CA and 3' 71A8CA were determined. This disclosure screened different 5' and 3' UTRs, and the combination of the optimal two ends of UTRs can improve the translation efficiency of circular RNA by about 5.7 times.

[0139] In addition, the present disclosure also modifies based on the CVB3 IRES sequence. First, the present disclosure explores the optimal position for inserting the IRES enhancer element into the CVB3 IRES sequence. The present disclosure designs 11 positions for inserting gene sequences, and it is experimentally found that there are 4 insertion positions with translation effects. The 4 specific insertion positions are respectively: the IRES enhancer element is inserted between domain I and domain II, named IRES-1; the IRES enhancer element is inserted into the stem-loop structure of domain II, named IRES-2; the IRES enhancer element is inserted into the stem-loop structure (Distal loop) of domain IV, named IRES-3; the IRES enhancer element is inserted into the stem-loop structure (Proximal loop) of domain IV, named IRES-4. The IRES enhancer element can enhance the binding of the CVB3 IRES sequence to ribosomes, thereby promoting the translation of circRNA. The IRES enhancer element exists in circRNA, folds into a structure similar to the initial tRNA, recruits more ribosomes, and binds to translation regulatory factors such as ITAF, and then introduces ribosomes into the interior of circRNA to bind and initiate protein translation. Ribosomes are highly diverse protein structures found in all cells. Under the action of translation regulatory factors such as ITAF, ribosomes are introduced into the internal structure of circRNA to bind and initiate. Secondly, the present disclosure screens 4 IRES enhancer elements, and then explores the influence of IRES enhancer elements with different sequences having the same enhancement effect on the translation effect of cicRNA. These 4 IRES enhancer elements all have the function of recruiting ribosomes, thus affecting the translation of circRNA. The 4 screened IRES enhancer element sequences are respectively cloned into the DNA template, and the composition of the DNA template is the same as above, and the optimal sequence is screened to obtain the gene sequence with high protein expression. Finally, based on the optimal position of the IRES enhancer element and the basis of screening the IRES enhancer element, the present disclosure conducts a series of repeat sequence designs on the IRES enhancer element to obtain the sequence for the IRES enhancer element to achieve its function. Compared with the original sequence, IRES-1-A1-X2 in the screened sequence has increased by about 2.3 times. BRIEF DESCRIPTION OF THE DRAWINGS

[0140] Figure 1 : Design diagram of Example 1 ( Figure 1 a: Main components of the PIE method and the structure of Design1; Figure 1 b: "GU" wobble pairing is added to Design2; Figure 1 c: Circular RNA is generated by 2% EX-gel electrophoresis analysis of Design2; Figure 1 d: The circular RNA generated by Design2 is introduced into Hela cells for expression measurement; Figure 1 e: Reverse transcription PCR sequencing of the product of Design2)

[0141] Figure 2 : Influence of the number of bases in the 3'-loop recognition fragment on the cyclization efficiency ( Figure 2 a: Trunc-Exon; Figure 2 b: Reverse transcription PCR sequencing of Trunc-Exon products; Figure 2 c: Fluorescent electrophoresis of products after single nucleotide deletion in the 5nt sequence; Figure 2 d: Cellular expression of products after single nucleotide deletion in the 5nt sequence)

[0142] Figure 3 : Conservative test of the 2nt sequence in the 3'-loop recognition fragment ( Figure 3 a-b: Construction of possible combinations based on the 2nt structure; Figure 3 c-d: Almost all 2nt sequences can be successfully cyclized without a decrease in cyclization efficiency; Figure 3 -f: Homology between the 2nt and the 5'-loop recognition fragment)

[0143] Figure 4 : Part of the IG sequence is conserved and recognizes the 5'-loop recognition fragment through base pairing or "GU" wobble pairing ( Figure 4 a: Original Exon1 sequence; Figure 4 b: Mutant Exon1 sequence; Figure 4 c: The IG sequence needs to satisfy the sequence feature of "GNN"; Figure 4 d: Influence of "GU" wobble pairing on splicing; Figure 4 e: Complete disruption of splicing when base pairing is not satisfied)

[0144] Figure 5 : Influence of internal homology ( Figure 5 a: Internal homologous sequence of the original Exon of Anabaena; Figure 5 b: 19nt-length homology arm in the studied sequence; Figure 5 c: Influence of 19nt removal; Figure 5 d-f: Influence of homologous sequences with different GC ratios and lengths on cyclization efficiency)

[0145] Figure 6 : Overall design principle of the present invention

[0146] Figure 7 : Design and results of cyclized IRES+Fluc ( Figure 7 .a: Circular RNA containing only "IRES+Fluc"; Figure 7 b: Splicing site setting; Figure 7 c: Cyclization efficiency; Figure 7 d: Translation efficiency; Figure 7e: In vitro reverse transcription PCR sequencing)

[0147] Figure 8 : Design and results of circularized IRES + eGFP( Figure 8 a: Splice site and internal homology sequence settings; Figure 8 b: Circularization efficiency; Figure 8 c: In vitro reverse transcription PCR sequencing; Figure 8 d: Translation efficiency)

[0148] Figure 9 : Effects of UTR on circularization efficiency and translation efficiency( Figure 9 a: 5’UTR screening; Figure 9 b: 3’UTR screening; Figure 9 c - d: PolyA sequence optimization; Figure 9 e - g: "IRES + ORF + UTR" scarless circularization system)

[0149] Figure 10 : Insertion position of CVB3 IRES

[0150] Figure 11 : Electrophoretogram of IRES enhancer element

[0151] Figure 12 : Transfection effect of IRES sequence inserted with IRES enhancer element in Example 4 of the present invention

[0152] Figure 13 : Production process flow of circular RNA

[0153] Figure 14 : Gel electrophoretogram of Gluc - cRNA production process

[0154] Figure 15 Gluc - cRNA first - step NHS affinity purification

[0155] Figure 16 Gluc - cRNA affinity purification sampling gel electrophoretogram (loading amount: 12.2 mg; total CRNA recovery: 10.23 mg; CRNA recovery rate: 83.85%; purity: 88.2%)

[0156] Figure 17 Gluc - cRNA - Core400 molecular sieve purification

[0157] Figure 18 Agilent 5200 capillary electrophoretogram in the two - step purification method of Gluc - cRNA (loading amount: 9.4 mg; total CRNA recovery: 8.01 mg; CRNA recovery rate: 85.21%; purity: 98.9%) Detailed implementation manners

[0158] In view of the current technical problems of circular RNAs, taking the ribozyme of Anabaena as the research object, the internal guide sequence in the ribozyme is modified to enable it to recognize the required circularization sequence, so as to achieve the synthesis of circular RNAs without the incorporation of exogenous sequences. At the same time, multiple UTR sequences are screened in the present disclosure, and the screened UTR sequences have an enhancing effect on the translation of circular RNAs. The optimal combination increases the translation efficiency by 5.7 times. Meanwhile, the IG sequence in the present disclosure recognizes the UTR sequence, and also makes the system a universal circularization system, and no other designs are required for different circularization sequences. In addition, the IRES sequence is screened and modified.

[0159] Through the study of the circularization mechanism of the PIE method in the present disclosure, it is found that the occurrence of the circularization process needs to meet the following principles: ① The 3 nucleotides of the IG sequence need to meet the characteristic of "GNN"; ② The IG sequence needs to ensure base complementary pairing or GU wobble pairing with the 3 nucleotides of the 5' splice site; ③ The 3' splice site should not have obvious homology with the 5' splice site and its adjacent nucleotides, as homology is not conducive to splicing and will greatly reduce the circularization efficiency; ④ At least 3 nucleotides adjacent to the 5' splice site need to have homology with at least 3 nucleotides adjacent to the 3' splice site. In addition, through the modification of the internal guide sequence to enable it to recognize the UTR sequence, a method for RNA circularization without a "scar" sequence is successfully developed, and the advantage of high circularization efficiency is retained. At the same time, due to the recognition of the spacer sequence, this method is universal for different sequences to be circularized and no additional design is required.

[0160] Secondly, 4 IRES enhancer elements are screened in the present disclosure, and then the influence of IRES enhancer elements with different sequences having the same enhancing effect on the translation effect of cicRNAs is explored. These 4 IRES enhancer elements all have the function of recruiting ribosomes and thus have an impact on the translation of circRNAs. The 4 screened IRES enhancer element sequences are respectively cloned into the DNA template, and the composition of the DNA template is the same as above, and the optimal sequence is screened to obtain the gene sequence with high protein expression. Finally, based on the optimal position of the IRES enhancer element and the screening of the IRES enhancer element, a series of repetitive sequence designs are carried out on the IRES enhancer element to obtain the sequence for the IRES enhancer element to achieve its function. The IRES-1-A1-X2 in the screened sequences is about 2.3 times higher than the original sequence.

[0161] Technical terms

[0162] When used in conjunction with the term "comprising" in the claims and / or the specification, the words "a" or "an" can mean "one", but can also mean "one or more", "at least one", and "one or more than one".

[0163] As used in the claims and specification, the words "comprising," "having," "including," or "containing" are meant to be inclusive or open-ended and do not exclude additional, unrecited elements or method steps.

[0164] Throughout the application, the term "about" means that a value includes the standard deviation of the error of the device or method used to determine that value.

[0165] Although the disclosed content supports a definition of the term "or" as being only alternatives and "and / or," unless expressly stated to be only alternatives or alternatives that are mutually exclusive, the term "or" in the claims means "and / or."

[0166] The terms "polypeptide," "peptide," and "protein" are used interchangeably in the present disclosure and are amino acid polymers of any length. The polymer can be linear or branched, it can contain modified amino acids, and it can be interrupted by non-amino acids.

[0167] In a specific embodiment of the present disclosure, the nucleic acid or peptide sequence of the present invention comprises a sequence with a homology or sequence identity of more than 90%, preferably more than 95%, more preferably 96%, 97%, 98%, 99%. Methods for determining sequence homology or identity known to those of ordinary skill in the art include, but are not limited to: Computational Molecular Biology, Lesk, A.M. ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W. ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part 1, Griffin, A.M. and Griffin, H.G. eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987 and Sequence Analysis Primer, Gribskov, M. and Devereux, J. eds. M Stockton Press, New York, 1991 and Carillo, H. and Lipman, D., SIAM J. Applied Math., 48:1073 (1988). Preferred methods for determining identity are to obtain the maximum match between the sequences being tested. Methods for determining identity are compiled in computer programs that are publicly available. Preferred computer program methods for determining identity between two sequences include, but are not limited to: the GCG program package (Devereux, J. et al., 1984), BLASTP, BLASTN, and FASTA (Altschul, S, F. et al., 1990). The BLASTX program is publicly available from NCBI and other sources (BLAST Manual, Altschul, S. et al., NCBI NLM NIH Bethesda, Md. 20894; Altschul, S. et al., 1990). The well-known Smith Waterman algorithm can also be used to determine identity.

[0168] Type I ribozyme or group I intron

[0169] As used in the present disclosure, a type I ribozyme or a group I intron refers to a "Group I Intron", which has a requirement for GTP and Mg 2+A self-splicing system that self-splices into a loop under existing conditions.

[0170] Group I introns are a class of very large ribozymes that can undergo self-splicing reactions and are usually widely present in many species, mainly participating in catalyzing the excision of precursors of mRNA, tRNA, and rRNA. According to the secondary structure, group I introns can be divided into 10 domains, named P1 - P10 respectively. Their catalytic centers are located at the junctions of the P4 - P6 (P4, P5, P6) and P3 - P9 (P3, P7, P8, P9) domains. P1 and P10 are the 5' and 3' splice sites respectively, while other sequences surrounding the conserved catalytic center are responsible for maintaining the stability and correct folding of the ribozyme. Its core secondary structure includes paired regions (P) and corresponding loop regions (L). The splicing of group I introns proceeds through two consecutive transesterification reactions. Exogenous guanosine or guanosine nucleotide (G) first docks at the active G-binding site located in P7, and its 3'-OH aligns to attack the phosphodiester bond at the 5' splice site in P1, resulting in a free 3'-OH group in the upstream exon, and the exogenous G is linked to the 5' end of the intron. Then the terminal G (omega G) of the intron exchanges with the exogenous G, occupying the G-binding site and organizing the second transesterification reaction: the 3'-OH group of the upstream exon in P1 aligns to attack the 3' splice site P10, resulting in the ligation (circularization) of the adjacent upstream and downstream exons and the release of the catalytic intron.

[0171] In the present disclosure, intron fragment I and intron fragment II, or 5'-half intron fragment and 3'-half intron fragment have the same meaning, referring to fragments derived from group I introns and respectively containing partial sequences close to the 5' direction and partial sequences close to the 3' direction that make up group I introns.

[0172] The internal cleavage sites of group I introns can be determined by those skilled in the art. For example, they can be determined by referring to the following documents: Puttaraju M., et al., (1992). Group I permuted intron-exon (PIE) sequences self-splice to produce circular exons; Puttaraju M., et al., (1996). Circular ribozymes generated in Escherichia coli using group I self-splicing permuted intron-exson sequences. For example, for group I introns, particularly for Anabaena group I introns, they can usually be cleaved at specific sites within their P6 region to form a PIE system. Therefore, in some embodiments, the 3'-terminal portion of a native group I intron (e.g., Anabaena group I intron) is the portion from a specific site within the P6 region of the native group I intron (e.g., Anabaena group I intron) to the 3'-end. In some embodiments, the 5'-terminal portion of the native group I intron (e.g., Anabaena group I intron) is the portion from a specific site within the P6 region of the native group I intron (e.g., Anabaena group I intron) to the 5'-end. For group I introns, particularly for Anabaena group I introns, the cleavage site can also be located within their P2, P5, P8, or P9 region to form a PIE system, as shown in WO2021236855A1. The group I introns also include various truncated forms or other mutants with self-splicing activity, such as the corresponding variants disclosed in WO2020 / 237227A1 (Tables 4 and 5), WO2023046153A1, or CN115997018A, etc.

[0173] In some embodiments, the self-splicing intron is a Group I intron of Anabaena, for example, the Group I intron of the Anabaena pre-tRNA-Leu gene. Correspondingly, in some embodiments, the 3' self-splicing intron fragment (3' Group I intron fragment) and the 5' self-splicing intron fragment (5' Group I intron fragment) are derived from the Group I intron of Anabaena, for example, the Group I intron of the Anabaena pre-tRNA-Leu gene. In some embodiments, the native Group I intron of the Anabaena pre-tRNA-Leu gene has the nucleotide sequence of SEQ ID NO:43. The P6 region corresponds to positions 98 to 157 of SEQ ID NO:43. The cleavage site can be any position between positions 122 and 138 of SEQ ID NO:43.

[0174] In some embodiments, the 3' self-splicing intron fragment (3' Group I intron fragment) is derived from the Group I intron of the Anabaena pre-tRNA-Leu gene and comprises the nucleotide sequence of SEQ ID NO:44 or a nucleotide sequence having at least 75%, for example, at least 80%, at least 85%, at least 90%, at least 95%, 100% identity to SEQ ID NO:44, or consists of the same. In some embodiments, the 5' self-splicing intron fragment (5' Group I intron fragment) is derived from the Group I intron of the Anabaena pre-tRNA-Leu gene and comprises the nucleotide sequence of SEQ ID NO:45 or a nucleotide sequence having at least 75%, for example, at least 80%, at least 85%, at least 90%, at least 95%, 100% identity to SEQ ID NO:45, or consists of the same.

[0175] Internal guide sequence or IGS

[0176] The internal guide sequence (IGS) generally refers to a nucleotide sequence in a Group I intron that pairs with the corresponding exon sequence through Watson-Crick pairing or wobble pairing, and is usually located in the P1 stem of the Group I intron. The original IG sequence of the Group I intron of the pre-tRNA-Leu gene is "GAG".

[0177] 5' loop recognition fragment and 3' loop recognition fragment

[0178] The 5'-loop recognition fragment is derived from the exon sequence (Exon 1, E1) connected to the 5'-end of the group I intron, and the 3'-loop recognition fragment is derived from the exon sequence (Exon 2, E2) connected to the 3'-end of the group I intron. Among them, the internal guide sequence recognizes the 5'-loop recognition fragment through Watson-Crick pairing or wobble pairing.

[0179] Among them, E1 is the adjacent exon sequence adjacent to the 5'-splice site, and its length is at least 1 nucleotide (for example, the length is at least 5 nucleotides, the length is at least 10 nucleotides, the length is at least 15 nucleotides, the length is at least 20 nucleotides, the length is at least 25 nucleotides, the length is at least 50 nucleotides); among them, E2 is the adjacent exon sequence adjacent to the 3'-splice site, and its length is at least 1 nucleotide (for example, the length is at least 5 nucleotides, the length is at least 10 nucleotides, the length is at least 15 nucleotides, the length is at least 20 nucleotides, the length is at least 25 nucleotides, the length is at least 50 nucleotides). WO2023046153A1 discloses alternative E1 fragments such as CUU or CUC of the group I intron of the pre-tRNA-Leu gene; and alternative E2 fragments such as AAAA, AA, UUUU, CAAA or GAAA.

[0180] In the prior art, usually E1 and E2 (including optimized E1 and E2) are retained in the loop-forming sequence to form a "scar" of 5'-E1-E2-3', also known as the residual loop-forming sequence; in the present disclosure, due to the modification of the IG sequence, the selection of the 5'-loop recognition fragment is thus expanded, and the corresponding 5'-loop recognition fragment can be set in the translation initiation element, coding region or spacer sequence; similarly, the present disclosure also optimizes the number of bases and base types of the 3'-loop recognition fragment, expanding the selectivity of the 3'-loop recognition fragment.

[0181] In the present disclosure, the 3'-loop recognition fragment does not have obvious homology with the 5'-loop recognition fragment and its adjacent nucleotides so as not to reduce the cyclization efficiency, where obvious homology means that two or more, such as three nucleoside bases, are complementary paired.

[0182] Homologous sequence or homologous arm

[0183] Homologous sequence or homologous arm has the same meaning in the present disclosure, including the 5'-homologous arm located at the 5'-end of the recombinant nucleic acid molecule and the 3'-homologous arm located at the 3'-end of the recombinant nucleic acid molecule. The nucleic acid sequence of the 5'-homologous arm is complementary to the nucleotide sequence of the 3'-homologous arm.

[0184] Internal homology or internal homologous sequence

[0185] In E1 and E2, or in the artificially inserted circularization fragment, several complementary fragments are required in the nucleotides adjacent to the 5'-circularization recognition fragment and the 3'-circularization recognition fragment to form a "splicing bubble". It is found in the present disclosure that at least 3 nucleotides adjacent to the 5'-circularization recognition fragment need to have homology with at least three nucleotides adjacent to the 3'-circularization recognition fragment.

[0186] In certain embodiments of the present disclosure, the internal homology is provided by a spacer sequence.

[0187] The term "adjacent region" or "adjacent sequence" refers to a region or sequence of 1-20 bases, preferably 1-10 bases, upstream or downstream of and connected to a described fragment such as a 5'-circularization recognition fragment or a 3'-circularization recognition fragment in the recombinant nucleic acid molecule; in the specific embodiments of the present disclosure, the adjacent sequence or adjacent region of the 5'-circularization recognition fragment or the 3'-circularization recognition fragment refers to the sequence or region adjacent to the above circularization recognition fragment and located in the circularization fragment, that is, it remains in the final circular RNA molecule together with the circularization fragment.

[0188] Spacer sequence:

[0189] In the present disclosure, a "spacer sequence" refers to any of the following continuous nucleotide sequences: 1) predicted to avoid interfering with proximal structures, such as those from IRES, coding or non-coding regions or introns, 2) at least 7 nucleotides in length (optionally not exceeding 100 nucleotides), 3) located downstream and adjacent to the 3' intron fragment and / or upstream and adjacent to the 5' intron fragment, and / or 4) containing one or more of the following: a) an unstructured region at least 3 nt long, b) a region predicted to base pair with a distal (i.e., non-adjacent) sequence at least 3 nt long, which includes another spacer sequence, and / or c) a structured region at least 7 nt long, the scope of which is limited to the sequence of the spacer sequence.

[0190] In certain embodiments, the spacer sequence can be, for example, at least 10 nucleotides in length, at least 15 nucleotides in length, or at least 30 nucleotides in length. In certain embodiments, the length of the spacer sequence is at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, or 30 nucleotides. In certain embodiments, the length of the spacer sequence does not exceed 100, 90, 80, 70, 60, 50, 45, 40, 35, or 30 nucleotides. In certain embodiments, the length of the spacer sequence is 20 to 50 nucleotides. In certain embodiments, the length of the spacer sequence is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides.

[0191] The spacer sequence can be a polyA sequence, a polyA-C sequence, a polyC sequence, or a poly-U sequence, or the spacer sequence can be specifically engineered according to the IRES. The spacer sequence described in the present disclosure can have two functions: (1) promoting cyclization and (2) improving translation efficiency. More specifically, the spacer sequence described in the present disclosure is designed to have the following functions: 1) being inert to the folding of the proximal intron and the IRES structure; 2) sufficiently separating the intron and the IRES secondary structure; 3) carrying splice sites; 4) containing a spacer sequence-spacer sequence complementary region to promote the formation of a "splicing bubble"; 5) improving translation efficiency.

[0192] The present disclosure preferably selects three types of 5' spacer sequences: the first type is the protien bingding sequence, which can recruit proteins by binding to improve translation efficiency, and the two sequences are 31A8CA and 41A8CA respectively; the second type is the IRES enhancing sequence, and the IRES binds to 18s rRNA, eIF4G, and eIF4A; the third type is the full-length UTR sequence.

[0193] The present disclosure preferably uses the 5' spacer sequence and the 3' spacer sequence together. Preferably, both are polyA+8CA. Preferably, the polyA fragment contains 10 - 100 As, more preferably 60 - 80 As, and most preferably 61 As at the 5' end and 71 As at the 3' end.

[0194] In a specific embodiment of the present disclosure, splice sites and internal homologous sequences are set on the 5' spacer sequence and the 3' spacer sequence to form a general cyclization system.

[0195] "Splicing bubble" refers to the region between the homologous arms and the internal homologous region, which contains a splicing ribozyme, see CN112399860A.

[0196] The term "Coding Region" refers to a gene sequence that can transcribe messenger RNA and ultimately be translated into a target polypeptide or protein.

[0197] The term "expression" includes any steps involved in polypeptide production, including but not limited to: transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0198] Unless otherwise defined or clearly indicated by the context, all technical and scientific terms in this disclosure have the same meaning as commonly understood by those of ordinary skill in the art to which this disclosure belongs.

[0199] This disclosure provides a general and "scar"-free RNA circularization method, which includes but is not limited to DNA constructs for preparing circular RNA, recombinant expression vectors including DNA constructs, circular precursor RNA molecules obtained by in vitro transcription using recombinant expression vectors, etc.

[0200] This disclosure first provides recombinant nucleic acid molecules for preparing circular RNA. Exemplarily, the recombinant nucleic acid molecule can be the above-mentioned DNA construct for preparing circular RNA or a circular precursor RNA molecule.

[0201] Circularization fragment

[0202] In this disclosure, the RNA molecule fragment that finally forms a ring or its corresponding DNA fragment is called a circularization fragment, which contains a translation initiation element, a target polypeptide coding region or a non-coding region, and other optional regulatory and insertion elements such as 5' or 3' UTR, and also contains a spacer fragment between intron fragments.

[0203] In certain specific embodiments of this disclosure, the circularization fragment contains the 5' part of the target polypeptide coding region or non-coding region, a translation initiation element, and the 3' part of the target polypeptide coding region or non-coding region, that is, the 5' circularization recognition fragment and the 3' circularization recognition fragment divide the target polypeptide coding region or non-coding region into two parts, and after intron splicing, they are reconnected into a complete target polypeptide coding region or non-coding region.

[0204] In certain specific embodiments of this disclosure, the circularization fragment contains the 5' part of the translation initiation element, the target polypeptide coding region or non-coding region, and the 3' part of the translation initiation element, that is, the 5' circularization recognition fragment and the 3' circularization recognition fragment divide the translation initiation element into two parts, and after intron splicing, the 5' part and the 3' part of the translation initiation element are connected into a complete translation initiation element region.

[0205] Translation initiation element

[0206] In the present disclosure, the translation initiation element can be any type of element capable of initiating the translation of a target polypeptide. In some embodiments, the translation initiation element is an element comprising any one or more than two of the following sequences: IRES sequence, 5'UTR sequence, Kozak sequence, a sequence containing m6A modification (N(6)-methyladenosine modification), a complementary sequence of ribosomal 18S rRNA. In some other embodiments, the translation initiation element can also be any other type of cap-independent translation initiation element.

[0207] In some embodiments, the translation initiation element is an IRES element, and the sources of the IRES element include, but are not limited to, viruses, mammals, Drosophila, etc. In some alternative embodiments, the IRES element is derived from a virus. Exemplarily, the IRES enhancer element contains an IRES sequence from picornavirus. In some alternative embodiments, the IRES element is derived from Taura syndrome virus, Triatoma virus, Theiler's encephalomyelitis virus, simian virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, reticuloendotheliosis virus, Forman poliovirus 1, Pseudoplusia includens virus, Kashmir bee virus, human rhinovirus 2, Homalodisca coagulata virus-1, human immunodeficiency virus type 1, Homalodisca coagulata virus-1, Pediculus humanus corporis virus, hepatitis C virus, hepatitis A virus, hepatitis G virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinovirus, Ectropis obliqua-like virus, encephalomyocarditis virus (EMCV), Drosophila C virus, tobacco mosaic virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, Hibiscus chlorotic ringspot virus, classical swine fever virus, human FGF2, human SFTPA1, human AMLl / RUNXl, Drosophila antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAPl, human cmyc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, saliva virus, Coxsackievirus, Echovirus, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, simian picornavirus, turnip crinkle virus, aptamer of eIF4G, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVB1 / 2); Further preferably, the IRES is the IRES sequence of Coxsackievirus B3 (CVB3); Still preferably, the IRES is the IRES sequence of encephalomyocarditis virus. Preferably, it can be selected from the specific sequences shown in Table 3 of WO2020 / 237227A1.

[0208] Target polypeptide coding region or non-coding region

[0209] In certain embodiments, the target polypeptide coding region or non-coding region is not a naturally occurring nucleotide sequence. In certain embodiments, the target polypeptide coding region encodes a natural or synthetic protein. In certain embodiments, the coding or non-coding region can be a natural or synthetic sequence. In certain embodiments, the coding region can encode a chimeric antigen receptor, an antibody, an immunomodulatory protein, and / or a transcription factor. In certain embodiments, the non-coding region can encode a sequence that can alter cell behavior (e.g., lymphocyte behavior). In certain embodiments, the non-coding sequence is antisense to the cellular RNA sequence.

[0210] In certain specific embodiments of the present disclosure, the circularized fragment comprises a 5' UTR, a target polypeptide coding region, and a 3' UTR, and the 5' circularization recognition fragment and the 3' circularization recognition fragment are located in the 5' UTR and the 3' UTR, respectively.

[0211] Insertion element

[0212] In some embodiments, the recombinant nucleic acid molecule comprises an insertion element, which can be used to regulate the transcription of the recombinant nucleic acid molecule, to regulate the translation of the circular RNA, to achieve the specific expression of the circular RNA between different tissues, or to purify the circular RNA, etc. Exemplarily, the insertion element is located in the coding element and the translation initiation element.

[0213] In some embodiments, the insertion element is selected from at least one of the following groups (i)-(iii): (i) a transcriptional level regulatory element, (ii) a translational level regulatory element, (iii) a purification element. Exemplarily, the insertion element comprises a sequence of one or any combination of two or more of the following: an untranslated region (UTR) sequence, a polyN sequence, an aptamer sequence, a riboswitch sequence, a sequence that binds to a transcriptional regulatory factor; in the polyN sequence, N is selected from at least one of A, T, G, and C.

[0214] In some alternative embodiments, the translation regulatory element comprises an untranslated region sequence, which can be used to regulate properties such as the stability, immunogenicity of the circular RNA, and the efficiency of the circular RNA to express the target polypeptide. The present disclosure does not specifically limit the untranslated region sequence, which can be selected from any type of sequence in the art that has properties such as regulating the transcription, translation, intracellular stability, immunogenicity, etc. of the circular RNA. Further, the untranslated region sequence is not limited to the 5' UTR sequence or the 3' UTR sequence.

[0215] In some alternative embodiments, the non-translation region sequence contains one or more miRNA recognition sequences, for example, 1, 2, 3, 4, 5, 6, 7, and so on. By adding one or more miRNA recognition sequences, specific expression of circular RNA in different tissues and cells can be achieved, and targeted delivery of circular RNA molecules can be realized.

[0216] In some alternative embodiments, the translation regulatory element contains a polyN sequence, where N can be at least one of A, T, G, and C. By increasing the translation regulatory element containing the polyN sequence, the efficiency of circular RNA expressing the target polypeptide, the immunogenicity, stability, etc. can be improved, or it can be used for the purification of circular RNA. The present disclosure does not specifically limit the length of the polyN sequence, the selection types of N in the polyN sequence, and the composition manner, as long as it is beneficial to improving the performance of circular RNA. Exemplarily, the polyN sequence is a polyA sequence, a polyAC sequence, and so on.

[0217] In some alternative embodiments, the translation regulatory element contains a riboswitch sequence. The riboswitch sequence is a type of non-translation sequence that has a regulatory function on the transcription and translation of RNA. In the present disclosure, the riboswitch sequence can affect the expression of circular RNA, including but not limited to transcription termination, translation initiation inhibition, mRNA self-cleavage, and alteration of the splicing pathway in eukaryotes. In addition, the riboswitch sequence can also control the expression of circular RNA by triggering the binding or removal of molecules. Exemplarily, the riboswitch sequence is a cobalamin riboswitch (also called B12-element), an FMN riboswitch (also called RFN element), a glmS riboswitch, a SAM riboswitch, a SAH riboswitch, a tetrahydrofolate riboswitch, a Moco riboswitch, and so on. The present disclosure does not restrictively limit the type and sequence of the riboswitch sequence, as long as it can achieve the regulation of the transcription and translation levels of circular RNA expressing the target polypeptide.

[0218] In some alternative embodiments, the translation regulatory element contains an aptamer sequence. In the present disclosure, the aptamer sequence can be used to regulate the transcription and translation of circular RNA, or for the in vitro purification and preparation of circular RNA.

[0219] A recombinant expression vector containing a recombinant nucleic acid molecule

[0220] As used in the disclosure, "vector" refers to a segment of DNA that is synthetic (e.g., using PCR), or isolated from a virus, plasmid, or cells of a higher organism, into which an exogenous DNA fragment can be inserted or has been inserted for cloning and / or expression purposes. In certain embodiments, the vector can be stably maintained in an organism. The vector can contain, for example, an origin of replication, a selectable marker or reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). The term includes linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, cosmids, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc.

[0221] In some embodiments, the vector allows the production of translatable and / or biologically active circular RNAs in eukaryotic cells.

[0222] In some embodiments, the recombinant nucleic acid molecule is present as part of a recombinant expression vector for the preparation of circular RNAs. During in vitro transcription and cyclization processes, circular RNAs that express the target polypeptide can be prepared.

[0223] In still other embodiments, the recombinant nucleic acid molecule can also be present as a circularization precursor RNA molecule or a part thereof obtained after linearization treatment and transcription reaction of the recombinant expression vector. That is, the recombinant nucleic acid molecule only needs to undergo a cyclization reaction to obtain circular RNAs that express the target polypeptide.

[0224] In some embodiments, the steps for preparing circular RNAs in vitro include:

[0225] Transcription step: The recombinant nucleic acid molecule as described in the present disclosure or the recombinant expression vector as described in the present disclosure is transcribed to form a circularization precursor nucleic acid molecule;

[0226] Cyclization step: The circularization precursor nucleic acid undergoes a cyclization reaction to obtain circular RNAs.

[0227] In some alternative embodiments, the method further includes a step of purifying the circular RNAs.

[0228] Circular RNA or circRNA

[0229] In the present disclosure, circular RNA or circRNA has the same meaning, both referring to RNA circular molecules without "scars" obtained by the methods of the present application.

[0230] Target polypeptide

[0231] The present disclosure does not limit the types of target polypeptides, which may be human proteins or non-human proteins. Exemplarily, the target polypeptide includes but is not limited to antigens, antibodies, antigen-binding fragments, fluorescent proteins, proteins with disease treatment activities, proteins with gene editing activities, etc.

[0232] In the present disclosure, the term "antibody" is used in the broadest sense to refer to a protein containing an antigen-binding site, covering natural antibodies and artificial antibodies of various structures, including but not limited to monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), single-chain antibodies, intact antibodies, and antibody fragments.

[0233] In the present disclosure, the term "antigen-binding fragment" is a part or segment of a complete or full antibody that has fewer amino acid residues than the complete or full antibody and can bind an antigen or compete with the complete antibody (i.e., the complete antibody from which the antigen-binding fragment is derived) for binding to the antigen. Antigen-binding fragments can be prepared by recombinant DNA technology or by enzymatic or chemical cleavage of a complete antibody. Antigen-binding fragments include but are not limited to Fv, Fab, Fab’, Fab’-SH, F(ab’)2; diabodies; linear antibodies; single-chain antibodies (e.g., scFv); single-domain antibodies; bivalent or bispecific antibodies or fragments thereof; camelid antibodies (heavy-chain antibodies); and bispecific or multispecific antibodies formed from antibody fragments.

[0234] In the present disclosure, proteins with disease treatment activities may include but are not limited to enzyme replacement proteins, proteins for supplementation, protein vaccines, antigens (e.g., tumor antigens, viruses, bacteria), hormones, cytokines, antibodies, immunotherapies (e.g., for cancer), cell reprogramming / transdifferentiation factors, transcription factors, chimeric antigen receptors, transposases or nucleases, immune effectors (e.g., affecting susceptibility to immune responses / signaling), regulated death effector proteins (e.g., inducers of apoptosis or necrosis), non-lytic inhibitors of tumors (e.g., oncoprotein inhibitors), epigenetic modifiers, epigenetic enzymes, transcription factors, DNA or protein modification enzymes, DNA intercalators, efflux pump inhibitors, nuclear receptor activators or inhibitors, proteasome inhibitors, enzyme competitive inhibitors, protein synthesis effectors or inhibitors, nucleases, protein fragments or domains, ligands or receptors, and CRISPR systems or their components, etc.

[0235] Coding element comprising tandem coding regions

[0236] In some embodiments, the coding element in the recombinant nucleic acid molecule comprises at least one coding region. Exemplarily, the coding element comprises 1, 2, 3, 4, 5, 10, 15, 20, 25, and so on. Optionally, in some alternative embodiments, the coding element comprises at least two coding regions, and each coding region independently encodes any type of target polypeptide.

[0237] In some alternative embodiments, the coding element comprises two or more coding regions. The recombinant nucleic acid molecule comprises elements arranged in the following order along the 5' to 3' direction: intron fragment II, translation initiation element truncated fragment II, at least two coding regions, translation initiation element truncated fragment I, intron fragment I. In some other alternative embodiments, the recombinant nucleic acid molecule consists of the elements arranged in the above order.

[0238] In some preferred embodiments, the coding element further comprises a linker located between any two adjacent coding regions. The linker separates adjacent coding regions, enabling the circular RNA prepared from the recombinant nucleic acid molecule to express two or more target polypeptides.

[0239] In the present disclosure, the linker can be a polynucleotide encoding a 2A peptide or other types of polynucleotides encoding a linker peptide for spacing target polypeptides. Among them, the 2A peptide is a short peptide (~18 - 25 amino acids) derived from a virus, and they are usually referred to as "self-cleaving" peptides, which can produce multiple proteins from one transcript. Exemplarily, the 2A peptide is P2A, T2A, E2A, F2A, and so on.

[0240] In some embodiments, the coding element comprises at least two coding regions, and a translation initiation element is connected between any two adjacent coding regions. Exemplarily, the coding element comprises 1, 2, 3, 4, 5, 10, 15, 20, 25, and so on. Moreover, within the coding element, a translation initiation element is arranged between any two adjacent coding regions. In the above manner, a translation initiation element can be connected upstream of each coding region in the circular RNA prepared in vitro from the recombinant nucleic acid molecule, enabling the expression of two or more target polypeptides.

[0241] In some alternative embodiments, the coding element comprises two or more coding regions. The recombinant nucleic acid molecule comprises elements arranged in the following order along the 5' to 3' direction: intron fragment II, translation initiation element truncated fragment II, at least two coding regions, translation initiation element truncated fragment I, intron fragment I. Among them, a translation initiation element is connected between any two adjacent coding regions. In some other alternative embodiments, the recombinant nucleic acid molecule consists of the elements arranged in the above order.

[0242] In the present disclosure, the translation initiation element can be any type of element capable of initiating the translation of a target polypeptide. In some embodiments, the translation initiation element is an element comprising any one or more than two of the following sequences: IRES sequence, 5'UTR sequence, Kozak sequence, a sequence containing m6A modification (N(6)-methyladenosine modification), and a complementary sequence of ribosomal 18S rRNA. In some other embodiments, the translation initiation element can also be any other type of cap-independent translation initiation element.

[0243] In the present disclosure, each coding region included in the coding element independently encodes any type of target polypeptide. Among them, the target polypeptides encoded by any two coding regions can be the same or different.

[0244] In the circular RNA prepared by using the above recombinant nucleic acid molecule, a translation initiation element is correspondingly connected to the 5' end of each coding region, and multiple coding regions are connected in series through multiple translation initiation elements to achieve the expression of at least two target polypeptides.

[0245] Drug composition / Administration

[0246] In an embodiment of the present disclosure, the circRNA product described and / or produced by using the vector and / or method described in the present disclosure can be provided in the form of a composition (such as a drug composition).

[0247] Therefore, in certain embodiments, the present disclosure also relates to a composition, such as a composition comprising circRNA (circRNA product) and a pharmaceutically acceptable carrier. On the one hand, the present disclosure provides a drug composition comprising an effective amount of the circRNA described in the present disclosure and a pharmaceutically acceptable excipient. The drug composition of the present disclosure can comprise the circRNA described in the present disclosure, as well as one or more pharmaceutically or physiologically acceptable carriers, excipients or diluents. In certain embodiments, the drug composition of the present disclosure can comprise circRNA-expressing cells, for example, various circRNA-expressing cells as described in the present disclosure, in combination with one or more pharmaceutically or physiologically acceptable carriers, excipients or diluents.

[0248] In certain embodiments, the pharmaceutically acceptable carrier can be a component other than the active ingredient that is non-toxic to the subject in the drug composition.

[0249] Pharmaceutically acceptable carriers can include, but are not limited to, buffers, excipients, stabilizers or preservatives. Examples of pharmaceutically acceptable carriers are physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents and absorption delaying agents, etc., such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants or emulsifying agents or combinations thereof. The amount of pharmaceutically acceptable carrier in a pharmaceutical composition can be determined experimentally based on the activity of the carrier and the desired properties of the formulation, such as stability and / or minimal oxidation.

[0250] In certain embodiments, such compositions can include buffers such as acetic acid, citric acid, histidine, boric acid, formic acid, succinic acid, phosphoric acid, carbonic acid, malic acid, aspartic acid, Tris buffer, HEPPSO, HEPES, neutral buffered saline, phosphate buffered saline, etc.; carbohydrates such as glucose, sucrose, mannose or dextran, mannitol; proteins; polypeptides or amino acids such as glycine; antioxidants; chelating agents such as EDTA or glutathione; adjuvants (e.g., aluminum hydroxide); antibacterial and antifungal agents; and preservatives.

[0251] In certain embodiments, the compositions of the present disclosure can be formulated for a variety of modes of parenteral or non-parenteral administration. In one embodiment, the composition can be formulated for infusion or intravenous administration. The compositions disclosed herein can be provided, for example, as sterile liquid formulations such as isotonic aqueous solutions, emulsions, suspensions, dispersions or viscous compositions, which can be buffered to the desired pH. Formulations suitable for oral administration can include liquid solutions, capsules, sachets, tablets, lozenges and troches, liquid suspensions powders and emulsions in a suitable liquid.

[0252] As used herein, the terms “complementary” or “hybridizing” are used to refer to “polynucleotides” and “oligonucleotides” (which are interchangeable terms referring to nucleotide sequences) that relate to base pairing rules. For example, the sequence “CAGT” is complementary to the sequence “GTCA”. Complementary or hybridization can be “partial” or “complete”. “Partial” complementarity or hybridization means that one or more nucleic acid bases are mismatched according to the base pairing rules, and “total” or “complete” complementarity or hybridization between nucleic acids means that each nucleic acid base matches another base under the base pairing rules. The degree of complementarity or hybridization between nucleic acid strands has an important impact on the hybridization efficiency and strength between nucleic acid strands. This is particularly important in amplification reactions as well as detection methods that rely on the binding between nucleic acids.

[0253] The term "recombinant nucleic acid molecule" refers to a polynucleotide having sequences that are not linked together in nature. The recombinant polynucleotide can be included in a suitable vector, and the vector can be used to transform into a suitable host cell. Then the polynucleotide is expressed in the recombinant host cell to produce, for example, "recombinant polypeptide", "recombinant protein", "fusion protein", etc.; RNA molecules can also be obtained by reverse transcription in vivo or in vitro.

[0254] The term "recombinant expression vector" refers to a DNA construct used to express, for example, a polynucleotide encoding a desired polypeptide. The recombinant expression vector can include, for example, a transcriptional subunit containing i) a collection of genetic elements that regulate gene expression, such as promoters and enhancers; ii) a structure or coding sequence that is transcribed into mRNA and translated into protein; and iii) appropriate transcriptional and translational start and stop sequences. The recombinant expression vector is constructed in any suitable manner. The nature of the vector is not important, and any vector can be used, including plasmids, viruses, bacteriophages, and transposons. Possible vectors for the present disclosure include, but are not limited to, chromosomal, non-chromosomal, and synthetic DNA sequences, such as viral plasmids, bacterial plasmids, bacteriophage DNA, yeast plasmids, and vectors derived from combinations of plasmid and bacteriophage DNA, DNA from viruses such as lentivirus, retrovirus, vaccinia, adenovirus, fowlpox, baculovirus, SV40, and pseudorabies.

[0255] The term "host cell" refers to a cell into which an exogenous polynucleotide has been introduced, including progeny of such cells. Host cells include "transformants" and "transformed cells", which include primary transformed cells and progeny derived therefrom. A host cell is any type of cell system that can be used to produce the antibody molecules of the present invention, including eukaryotic cells, such as mammalian cells, insect cells, yeast cells; and prokaryotic cells, such as Escherichia coli cells. Host cells include cultured cells, and also include cells inside transgenic animals, transgenic plants, or cultured plant or animal tissues. The term "recombinant host cell" encompasses a host cell that is different from the parental cell after introduction of a recombinant nucleic acid molecule, a recombinant expression vector, or circular RNA. The recombinant host cell is specifically achieved by transformation. The host cells of the present disclosure can be prokaryotic cells or eukaryotic cells, as long as they are cells capable of introducing the recombinant nucleic acid molecules, recombinant expression vectors, circular RNAs, etc. of the present disclosure.

[0256] As used in the present disclosure, the terms "individual", "patient", or "subject" include mammals. Mammals include, but are not limited to, domestic animals (such as cows, sheep, cats, dogs, and horses), primates (such as humans and non-human primates like monkeys), rabbits, and rodents (such as mice and rats).

[0257] As used in the present disclosure, the terms "transformation, transfection, transduction" have the meanings commonly understood by those skilled in the art, that is, the process of introducing exogenous DNA into a host. The methods of said transformation, transfection, transduction include any method of introducing nucleic acids into cells, and these methods include but are not limited to electroporation, calcium phosphate (CaPO4) precipitation, calcium chloride (CaCl2) precipitation, microinjection, polyethylene glycol (PEG) method, DEAE-dextran method, cationic liposome method, and lithium acetate-DMSO method.

[0258] As used in the present disclosure, "treatment" means: after suffering from a disease, exposing (such as administering) a subject to the circular RNA, circular precursor RNA, composition, etc. of the present invention, so that the symptoms of the disease are alleviated compared with when not exposed, and it does not mean that the symptoms of the disease must be completely suppressed. Suffering from a disease means: the body shows symptoms of the disease.

[0259] As used in the present disclosure, the term "effective amount" refers to such an amount or dose of the recombinant nucleic acid molecule, recombinant expression vector, circular precursor RNA, circular RNA, vaccine or composition of the present invention that, after being administered to a patient in a single or multiple doses, produces an expected effect in a patient in need of treatment or prevention. The effective amount can be easily determined by the attending physician, who is skilled in the art, by considering various factors such as: the species of the mammal; its size, age and general health; the specific disease involved; the degree or severity of the disease; the response of the individual patient; the specific antibody administered; the mode of administration; the bioavailability characteristics of the administered formulation; the selected dosing regimen; and the use of any concomitant therapies.

[0260] As used in the present disclosure, the terms "individual", "patient" or "subject" include mammals. Mammals include but are not limited to domestic animals (such as cows, sheep, cats, dogs and horses), primates (such as humans and non-human primates such as monkeys), rabbits, and rodents (such as mice and rats).

[0261] Unless otherwise defined or clearly indicated by the context, all technical and scientific terms in the present disclosure have the same meanings as commonly understood by those of ordinary skill in the art to which the present disclosure pertains.

[0262] Examples

[0263] The sequence modification method used in the present disclosure is basically based on the NEB kit as HiFi DNA Assembly Master Mix, and the operation is carried out according to the kit instructions. A small amount of plasmid construction is carried out by a small amount of sequence modification by the Gibson method, and the kit used is HiFi DNA Assembly Master Mix.

[0264] The linearized template for in vitro transcription was prepared by digestion with the restriction enzyme BspQI. After digestion, it was recovered using Sangon Biotech's PCR purification kit, and agarose gel electrophoresis or capillary electrophoresis was used to determine whether the digestion was complete.

[0265] The in vitro transcription kit used for in vitro transcription included T7 polymerase Mix, reaction buffer, and the four nucleotides AGCU, which were purchased from Shanghai Zhaowei Company. The original plasmid template was purchased from GenScript.

[0266] The preparation of LNP was carried out by encapsulation with cationic liposomes, and expression verification was performed in Hela cells or 293T cells. 100 ng of RNA was transfected into each well of a 96-well plate. The electrophoretic separation of circular RNA was carried out using E-gel EX 2% Agarose from Thermo Fisher.

[0267] All the RNAs used in this disclosure were prepared by in vitro transcription. The in vitro transcription templates used were purchased from GenScript, and the templates were linearized with BspQI. The in vitro transcription kit included T7 polymerase Mix, reaction buffer, and the four nucleotides AGCU, which were purchased from Shanghai Zhaowei Company. The sequence modification was mainly based on NEB's Site-Directed Mutagenesis Kit. A small amount of sequence modification was carried out by the Gibson method, and the kit used was HiFi DNA Assembly Master Mix. The gel electrophoresis separation of circular RNA was carried out using E-gel EX 2% Agarose from Thermo Fisher.

[0268] Example 1: Feasibility experiment on modifying the IG sequence to achieve scarless cyclization

[0269] The main components of the PIE method are as Figure 1 .a The upper part shows that by adding homology arms on both sides, the splicing sites are brought closer to each other, which is more conducive to splicing; by adding spacer sequences in the Intron and IRES sequences, the two sequences with complex secondary structures are separated to avoid mutual influence and thus improve the cyclization efficiency; the Intron sequences located on both sides play the function of ribozymes, and after the IG sequence in the Intron recognizes the heterologous Exon sequence, splicing cyclization occurs. The Anabaena tRNA used in the present invention leuThe IG sequence of the original type I ribozyme is 3 nt long, and the splicing site recognition and splicing are carried out through this 3 nt. Therefore, the first plan is to modify the IG sequence so that it can recognize a sequence located in the ORF, and then perform the cyclization reaction independently of the heterologous Exon sequence. Here, Firefly Luciferase (1653 nt) is used as a reporter gene for the experiment, and two designs are made. Design 1 first simulates the secondary structure and selects sites that are close to each other in space as splicing sites ( Figure 1 a), which is considered to be conducive to cyclization; design 2 is based on design 1, with the addition of the "GU" swing pairing feature ( Figure 1 b) This feature is conserved in many other type I ribozymes [9] , so it is considered to be a more important feature for the shear reaction.

[0270] After in vitro transcription and circularization, it was found that design 1 did not produce circular RNA, while design 2 produced circular RNA. However, by 2% EX-gel electrophoresis ( Figure 1 c) It was found that compared with the PIE method, the circular RNA produced by design2 was very low. After being introduced into Hela cells for expression measurement, it was found that ( Figure 1 d), which is only about 10% of the PIE method. In addition, reverse transcription PCR sequencing was performed on design2, and it was found that precise splicing was performed ( Figure 1 e). This proves that design2 does produce circular RNA and performs precise splicing, but the current yield is too low, which may be caused by the unclear molecular mechanism of ribozyme action, so we plan to explore the molecular mechanism. We found that in the PIE system, polyAC spacer can effectively prevent Intron from forming a complex secondary structure with its adjacent sequences. This secondary structure may affect the formation of the natural secondary structure of the ribozyme or destroy the P1 structure, which is not conducive to splicing. Therefore, in subsequent experiments, in order to study the molecular mechanism, the constructed sequence is similar to the original PIE method, but the heterologous Exon sequence is removed.

[0271] Example 2: Molecular mechanism of Anabaena type I ribozyme cyclization

[0272] 2.1 The 3' splice site is not recognized but cannot base-pair with the nucleotides near the 5' splice site

[0273] The structures of type I ribozymes are highly conserved. Almost all of them have an IG sequence that is used to recognize the splicing sites on both sides. The IG sequence forms P1 with the 5' splicing site recognition and P10 with the 3' splicing site recognition. However, compared with other type I ribozymes, the type I ribozyme from Anabaena has not been found to have P10. There is only a P1 structure formed by a 3-nt IG sequence and the 5' splicing site, which seems to indicate that the type I ribozyme from Anabaena does not need to recognize the 3' splicing site and only needs to be adjacent to it for splicing. However, there is controversy about this. Some researchers believe that if the type I ribozyme from Anabaena has P10, it may be the "AU" base pair

[10] . Since in the previous design, only the IG sequence was modified and the cyclization rate was too low. Therefore, it was speculated whether there is a recognition mechanism for the 3' splicing site. So first, the 5-nt sequence of the original Exon2, which serves as the 3' splicing site in the Ana system, was retained for in vitro cyclization. This sequence is called "Trunc-Exon"( Figure 2 a). This sequence was successfully cyclized, and the cyclization efficiency was not reduced compared with the PIE method. After RT-PCR and sequencing, it was found that this sequence also underwent correct splicing( Figure 2 b). Then, single nucleotide deletions were made on this 5-nt sequence. We found that when 4 nucleotides were deleted, the cyclization efficiency decreased significantly. When all 5 nucleotides were deleted, almost no circular RNA was produced( Figure 2 c), and the same conclusion was obtained from the cell expression experiment( Figure 2 d). Therefore, a 2-nt (AT) conserved region was first determined, which is consistent with the previous "AU" base pair conjecture.

[0274] But to further determine whether the nucleotides in this 2-nt conserved region are "AT" conserved, on the 2-nt structure( Figure 3 .a), 15 (2 4 -1) other possible 2-nt nucleotide combinations were constructed. The results showed that almost all 2-nt sequences could be successfully cyclized and the cyclization efficiency was not reduced( Figure 3.c d), the cyclization efficiency of only "CG", "GG", and "TG" decreased significantly, which is the same as the previous result of Exon2Δ5 because its 2-nt sequence as the 3' splice site is also "CG". In the sequence we constructed, the 2-nt sequence of Exon2Δ5 as the 3' splice site is actually located in the 19-nt "GC"-rich internal homology, which means that the nucleotide that should have been the 3' splice site has a relatively strong base pairing with the nucleotide adjacent to the 5' splice site, which is not conducive to the subsequent splicing process and may be the reason for the decrease in cyclization efficiency.

[0275] To verify the above conjecture, we planned two experiments, positive and negative. We modified the sequences of Exon2-AA with high cyclization rate and Exon2Δ5 that could not be cyclized. We added the "TT" sequence to the 5' splice site to pair with the "AA" of Exon2-AA. At the same time, we also added a segment of "AA" to the 5' splice site of Exon2Δ5 to disrupt the original pairing characteristics of "CG" and "GC" ( Figure 3 .e). The results showed that for Exon2Δ5, when this homology was disrupted, cyclization immediately recovered and showed the same cyclization efficiency as the original sequence. In contrast, the originally normally cyclizable Exon-AA lost its cyclization ability completely after mutation ( Figure 3 .f g).

[0276] This experiment proved that for the Ana system, the 3' splice site is non-conservative and there is no other sequence for recognition. However, when there is a pairing characteristic between the 3' splice site and the adjacent sequence of the 5' splice site, the cyclization efficiency will be greatly reduced. This may be because when the 3' splice site forms such a base complementary characteristic, a relatively tight secondary structure will be formed, making the second splicing process difficult to occur.

[0277] 2.2 The IG sequence is partially conserved and recognizes the 5' splice site through base pairing or "GU" wobble pairing

[0278] After a detailed study of the 3' splice site, we then studied the 5' splice site. The 5' splice site is recognized by the "GAG" of the IG sequence of Ana, and the original Exon1 sequence is "CTT". The Exon1 sequence was sorted from 1 to 3 in the order of 5'-3' ( Figure 4 a).

[0279] First, in the first experiment, while keeping the IG sequence unchanged, its recognition of the Exon1 sequence was mutated. For example, the "G" at the first position recognizes the "T" at the first position of Exon1. Thus, the "T" was mutated into the other three nucleotides "A, G, C". This experiment was to explore the recognition mechanism of the IG sequence for the Exon1 sequence. The electrophoresis results ( Figure 4 .b) show that the IG sequence recognizes the cleavage site through base complementary pairing and "GU" wobble pairing. The first and third nucleotides strictly follow this feature, while the second nucleotide does not fully comply with this rule. For example, the combination of "A-C" in the experiment can also be cleaved and still retains a high cyclization efficiency. This may mean that the second position of the IG sequence is a non-conservative base position, and the cell expression experiment also reflects the same rule.

[0280] The second experiment is the study of the conservation of the IG sequence. From the previous experiment, it can be known that ensuring base complementary pairing, that is, satisfying the recognition relationship, can cyclize normally. Then, when the IG sequence changes and base complementary pairing is ensured at the same time, will it affect the cyclization process, that is, whether the IG sequence itself has a certain degree of conservation. Thus, a series of subsequent mutations were constructed. For example, the first nucleotide of the IG sequence recognizes through "G-U" wobble pairing. Then, this pair of nucleotides was modified to "C-G", "A-T", "T-A", and "G-C" to explore whether it would affect cyclization. The electrophoresis results show ( Figure 4 .c) When strictly following base complementary pairing, for the second base, the mutation of the IG sequence has no effect on reducing the cyclization efficiency. Even when mutated to the base "GC", the cyclization efficiency even increases. From the experimental results of the mutation of the first base, it can be seen that the modification of the IG sequence has little effect on the cyclization efficiency. Even if there is some decrease, it has little impact. However, the mutations of the IG sequence at the third base all result in a significant decrease in the cyclization efficiency. In this experiment, except for the mutation of the third base, obvious differences can be seen in the electrophoresis pattern, and the effects of mutations at other positions are not obvious enough in the electrophoresis results. Therefore, the cell expression experiment is crucial for the characterization of this cyclization efficiency and can quantitatively analyze the cyclization efficiency more clearly.

[0281] In addition, "GU wobble pairing" is also a possible recognition rule during the recognition of the IG sequence. However, in the original sequence, "GU wobble pairing" only exists at the first base. Therefore, we studied the effect of "GU" wobble pairing on the cyclization efficiency at all other positions. Thus, the base pairs at different positions were also mutated. We found that the mutation of "T-G" at the first position to "G-T" has a huge impact on cyclization, and almost no circular RNA is produced. However, such mutations at other positions can all be cyclized, but the cyclization efficiency decreases to varying degrees ( Figure 4.d). In addition, we believe it is necessary to clarify that in the above research, we found that most mutations only reduced the cyclization efficiency, rather than completely preventing the generation of circular RNAs. We believe there is an essential difference here. To verify this theory, we constructed a set of mutant sequences, denoted as Mut1-4. This set of sequences does not follow the above recognition rules. Except for maintaining base pairing at the non-conserved second base, there is no base complementary pairing at the other two positions. After in vitro cyclization, we found that such sequences that do not follow the recognition rule did not generate any circular RNAs at all ( Figure 4 .e), which is different from the result of reducing the cyclization efficiency in the above experiment.

[0282] Therefore, we believe that when maintaining base complementary pairing or "G-U" wobble pairing, the decrease in cyclization efficiency is caused by the decrease in the cleavage rate, which is different from the complete absence of circular RNAs caused by not meeting the recognition rules. This decrease in cyclization efficiency should be able to be optimized by exploring the optimal cyclization temperature, metal ion concentration, and cyclization reaction time. However, since nicking RNA will be generated during the cyclization process, obviously a faster cleavage rate is more beneficial. Therefore, combining the above experiments, we believe that the IG sequence should satisfy the characteristic of "GNN" and is recognized through base complementary pairing and partial "GU" wobble pairing.

[0283] 2.3 Sequence homology near the splice site affects the cyclization effect

[0284] The type I ribozyme system derived from Anabaena was previously studied and found that there is 5-nt length homology in the adjacent sequences of the two ends of the splice site. This homology is derived from the original Exon of Anabaena ( Figure 5 a). The study pointed out

[10] that this 5-nt length homology has a great impact on the cyclization efficiency. This may be due to the too short IG sequence of the Ana system. In other type I ribozyme systems, the median length of the IG sequence is 5 nt. The too short IG sequence leads to an unstable helix formed with the 5' splice site and requires additional homologous sequences to stabilize it. Since this 5-nt homologous sequence is brought by the original heterologous Exon sequence, when performing scarless cyclization, the heterologous Exon is removed, and this homologous sequence will also be deleted accordingly. In the research sequences we constructed, there is a 19-nt length homology arm near the splice site ( Figure 5 b), which may also be the reason why the previous experiments were not affected. After removing it, the results showed that no circular RNAs were generated at all ( Figure 5 c).

[0285] Experiments have shown that the homology of sequences near the splicing sites is important for RNA cyclization. According to previous studies, the length of the homologous sequence of the original Exon is 5 nt. Therefore, homologous sequences with different lengths and different GC ratios were constructed on the basis of 5 nt ( Figure 5 d), to study their effects on the cyclization efficiency. The results ( Figure 5 .ef) showed that when there was a 3-nt-long homologous arm, the cyclization ability was restored and the cyclization efficiency was the same as that of the wild type. However, longer homologous arms such as 10 nt in length did not improve the cyclization ability of circular RNAs, and when the homologous arm was shortened to 2 nt, the cyclization efficiency decreased significantly. Moreover, there was no obvious difference in the cyclization efficiency among homologous arms of the same length with different GC ratios, indicating that as long as homology was ensured, there was no sequence specificity.

[0286] 2.4 Summary of the scarless cyclization rule of the modified IG sequence

[0287] Based on the above experiments, we explored in detail several principles that affect the cyclization efficiency of in vitro cyclization of the Anabaena system. The following scarless cyclization design principles were summarized:

[0288] ① The IG sequence needs to satisfy the sequence feature of "GNN";

[0289] ② When the IG sequence recognizes the 5'splice site, base complementary pairing or "GU" wobble pairing needs to be ensured;

[0290] ③ The 3'splice site is not conserved, but when it is base complementary paired with the sequence near the 5'splice site, it will greatly affect the cyclization efficiency;

[0291] ④ The splicing sites on both sides need to have homologous arms with a minimum length of 3 nt to stabilize the P1 duplex

[0292] For the specific structure, see Figure 6

[0293] Example 3 Cyclization of IRES+Fluc and IRES+eGFP according to the scarless cyclization strategy of the modified IG sequence

[0294] According to the design principles summarized from the previous exploration of the molecular mechanism, first, a suitable cyclization site was found in Firefly Luciferase, and then the IG sequence was modified to construct a circular RNA containing only "IRES+Fluc" ( Figure 7a). First, a sequence of "TGGTGCCTTTTCACCA" from positions 719 - 734 was discovered in Fluc. Among them, "CCT" can serve as the 5' splice site, and the IG sequence was modified to "GGG" for recognition; "TT" serves as the 3' splice site, and there is no obvious base complementary pairing with the nucleotides near the 5' splice site; moreover, the "TGGTG" and "CACCA" on both sides can form a homologous sequence with a length of 5 nt( Figure 7 b). This segment of the sequence was successfully circularized in vitro, and the circularization efficiency did not decrease( Figure 7 c), and through in vitro reverse transcription PCR sequencing, it was found that precise splicing occurred( Figure 7 e). However, perhaps due to the lack of the polyAC spacer sequence, the translation efficiency decreased slightly( Figure 7 d).

[0295] To verify the universality of this modification method, it was then planned to perform the same modification on eGFP to construct a circular RNA of "IRES + eGFP". A segment of sequence from positions 61 - 76 that also conforms to the modification rules was selected in the ORF of eGFP( Figure 8 a), with "CGT" as the 5' splice site, and the IG sequence was modified to "GCG". The results showed that circular "IRES + eGFP" was also obtained in vitro, and the circularization efficiency was the same as that of PIE( Figure 8 b), and the sequencing results also showed that precise cleavage occurred( Figure 8 .c). The cell expression experiment also proved that the obtained circular RNA could express eGFP in cells. However, similarly, compared with the PIE method circular RNA with polyAC, the translation was downregulated. Perhaps the spacer sequence is very important for enhancing the translation ability of circular RNA.

[0296] Based on the above experiments, it was proved that through the above method, "scar"-free circular RNA could be successfully prepared in vitro, and the circularization efficiency was comparable to that of the PIE method, and large - fragment circular RNA could be prepared with high efficiency. However, for different ORFs, suitable splice sites need to be selected, and the secondary structure of the circularized fragment may affect the function of ribozymes, which was mentioned by someone in previous literature [7] , which is obviously inconvenient. In addition, in our previous research, it was found that certain spacer sequences can improve the translation ability of circular RNA, and this was also confirmed in the research of Howard Chang

[11] , so it is planned to screen some spacer sequences that are helpful for circular RNA translation. Since it has the function of improving translation, which is similar to the function of the UTR of linear mRNA, it will be called the UTR sequence hereafter. Then, the IG sequence will be modified according to the UTR sequence to construct a general circularization method of "IRES + ORF + UTR" without "scar".

[0297] Example 4: Screening of UTR sequences and construction of a general circularization method of "IRES + ORF + UTR"

[0298] After studying the mechanism of "scar"-free circularization in the experiment of modifying the IG sequence, it is planned to screen UTR sequences with enhanced translation effects, modify the IG sequence to make it recognize the UTR region, and then construct a general "scar"-free system, which does not require re-design for different circularization sequences. The UTR sequence has two functions. First, it can separate the intron sequence from the IRES. The secondary structures of these two sequences are both relatively complex. If the adjacent IRES is too close, it will greatly affect the catalytic function of the ribozyme. The second function is to improve the translation effect of circular RNA.

[0299] Based on the above principles, the 5'UTR was screened first, and three types of sequences were selected. The first type is the IRES enhancer sequence: IRES binds to 18s rRNA, eIF4G, and eIF4A to initiate translation. If this binding can be enhanced, the translation efficiency can be enhanced. Some researchers

[12] found that there is a 9nt sequence in the 5'UTR of the homeodomain protein Gtx, which is completely paired with 1132 - 1124 of 18s rRNA. The article also pointed out that the tandem repeat of the 9nt sequence can greatly enhance the translation efficiency of IRES, so the 9nt sequence was repeated 6 times and denoted as "Gtx54" (54nt); in addition, the author also found in subsequent experiments

[13] that the 7nt sequence in this 9nt sequence has a more obvious enhancing effect on the translation effect. Similarly, after repeated combination, the enhancing effect on the translation effect is more significant, so this sequence was repeated 7 times and denoted as "Gtx56" (6nt); some researchers also found a short "IRES-like" sequence

[14] that can play the function of IRES to initiate translation and is denoted as "OR4F17" (56nt); some researchers also found an aptamer sequence of eIF4G, which can recruit eIF4G to improve expression, so this sequence was designed as a UTR sequence denoted as "Apt-eIF4G" (51nt)

[15] The second category is Protein-binding sequences. Some studies have found that, compared with traditional polyA, inserting cytosine at the tail of polyA can enhance protein expression. This sequence can recruit polyA-binding proteins more effectively, thus achieving the effect of enhancing translation. According to the research content of the literature, the two-terminal sequences 31A8CA (40nt) and 41A8CA (50nt) were selected for research; the third category is the full-length mRNA UTR sequences. Previously, three UTR sequences with obvious enhancing effects on mRNA were screened out of 12,000 sequences through high-throughput screening. Then, the enhancing effects of these three sequences on circular RNA were experimentally studied, denoted as NeoUTR1 (100nt), NeoUTR2 (100nt), and NeoUTR3 (100nt). In total, 9 sequences were constructed in this way. In addition, a sequence with a GC ratio of 50% and a length of 50nt was constructed, denoted as the random sequence, as a control; the original sequence is a polyAC sequence, denoted as "polyAC". After sequence construction, EX-gel 2% electrophoresis was used to analyze the effects of these sequences on circularization. The gel diagram shows that these sequences have no effect on the circularization efficiency. Through cell expression experiments, it can be seen that 31A8CA and 41A8CA have obvious enhancing effects on translation ( Figure 9 a).

[0300] After screening the 5’UTR, the 3’UTR was screened immediately. For the 3’UTR, three sequences with good effects in the 5’UTR were first selected: OR4F17, 15A8CA, and NeoUTR; then a protein-binding sequence, ARE, was selected. This sequence can bind to the HuR protein and thus play a role in enhancing translation; finally, 5 3’UTR sequences of highly abundant proteins expressed in mammals were selected, including GADPH, α-globin, β-globin, CYBA, and DECR1. Similarly, a random fragment with a 50% GC content was used as a control, and the original sequence was also "polyAC". EX-gel 2% electrophoresis was used to analyze the effects of these sequences on circularization. The gel diagram shows that these sequences have no effect on the circularization efficiency. Through cell expression experiments, it can be seen that the 15A8CA and α-globin sequences have obvious improvements in translation effects and obvious enhancing effects on translation ( Figure 9 b).

[0301] Among them, the two sequences 41A8CA and 15A8CA were reported previously to show that increasing the number of A can enhance its effect on translation. Therefore, the number of A was increased step by step. A sequence was designed every 10 A, and a total of 9 sequences were obtained. The results showed that in the 5’UTR, translation first increased and then decreased with the increase in the number of A. Among them, 61A8CA had the most significant enhancement effect on translation. In the 3’UTR, 3’71A8CA showed the strongest enhancement effect. Figure 9 .c). Finally, the selected 5’61A8CA was combined with 3’71A8CA, and the constructed sequence showed the strongest translation effect, which was nearly 3.7 times higher than the original sequence. Figure 9 .d)

[0302] According to the modification principle of previous experiments, the UTR sequence and the IG sequence were modified to recognize the UTR sequences at both ends. “CCC” was inserted at the third nucleotide at the 5’ end of the 5’UTR. The two “AA” in 5’61A8CA were used as the 3’ splicing site. “GGGCTT” six nucleotides were inserted at the 3’ end of the 3’UTR. Among them, “GGG” and “CCC” at the 5’ end formed the required shortest homology, and “CTT” was recognized by the “GAG” IG sequence of the ribozyme as the 5’ splicing site to complete splicing. According to this design, the original 108nt heterologous fragment including Exon and internal homology could be removed, and only 9nt of the additional designed sequence was needed to complete circularization, constructing an “IRES+ORF+UTR” scarless circularization system. Figure 9 e), and it was proved by 2% EX-gel separation that the circularization efficiency was the same as that of the PIE method. Figure 9 .f). This sequence was introduced into Balb / C mice. The results showed that compared with the original sequence, the translation effect was also significantly enhanced. Figure 9 g).

[0303] Example 5 Modifying the translation effect of IRES-enhanced circular RNA

[0304] Design the DNA sequence: Using PUC57-Kan as the plasmid vector, XbaLI and PciI were selected as the restriction enzyme sites on both sides of the target gene. The target gene sequence was T7 promoter sequence, homology arm (HA), intron, polyAC sequence, CVB3 IRES sequence, optimized luciferase sequence, polyAC sequence, intron, homology arm (HA) in turn.

[0305] 1. Optimal position of the IRES enhancer element: The present disclosure designed 11 positions for inserting gene sequences. The insertion positions of the IRES enhancer element are as shown in Figure 10 . The experimental results found that there were 4 positions with translation effects. The IRES enhancer element was inserted between domain I and domain II and named IRES-1. The IRES enhancer element was inserted into the stem-loop structure of domain II and named IRES-2. The IRES enhancer element was inserted into the stem-loop structure (Distal loop) of domain IV and named IRES-3. The IRES enhancer element was inserted into the stem-loop structure (Proximal loop) of domain IV and named IRES-4. The plasmid without the inserted IRES enhancer element was named CVB3 IRES.

[0306] 2. IRES enhancer elements screened from the natural gene pool: Four sequences were selected from the natural gene pool. Except that the DNA sequence of the IRES enhancer element was different from that of IRES-1, the other genes on the plasmid were the same as those of IRES-1. The four sequences were named IRES A1-4 in sequence [17-20] , and the specific primer sequences are shown in Table 3. The specific sequences of IRES A1-4 are shown in Table 4.

[0307] 3. Exploration of the repeated design of the IRES enhancer element: The screened IRES enhancer elements were repeatedly designed to explore the optimal translation efficiency.

[0308] 4. The repeated sequences were named IRES-Xn (n: number of repetitions) in sequence, and the specific sequences are shown in Table 5.

[0309] II. Insertion of the IRES enhancer element: Primers were designed, and then the plasmid with the inserted IRES enhancer element was obtained by PCR. The primer sequences for obtaining the vector by PCR are shown in Table 1, and the PCR reaction parameters are shown in Table 2.

[0310] Table 1:

[0311]

[0312] Table 2:

[0313]

[0314] III. Purification of the PCR reaction: A PCR reaction purification kit was used, and according to the operating steps of the instruction manual, the purified vector and IRES enhancer element were obtained. The DNA was qualitatively and quantitatively analyzed using an ultra-micro ultraviolet spectrophotometer, and agarose gel electrophoresis was performed to verify the integrity and purity of the DNA template.

[0315] IV. Transformation of the PCR product into competent cells

[0316] (1) The E. coli DH5α competent cells were thawed on ice before use.

[0317] (2) 2 μL of PCR product was added to 50 μL of the cells, and the tube wall was flicked gently to mix well.

[0318] (4) Incubate on ice for 30 min, quickly place in a 42 °C water bath for heat shock for 45 s, and then quickly transfer to ice and let stand for 2 min.

[0319] (5) Add SOC medium to make up the volume to 1 mL.

[0320] (6) Incubate with shaking at 37 °C for 1 hour.

[0321] (7) Concentration: Centrifuge at 5000 × g for 1 min, discard 900 μL of the supernatant, and pipette the remaining part to mix well.

[0322] (8) Spread 100 μL of the bacterial solution on an LB plate with 100 mg / L Kan resistance, and place it upright at 37 °C until the bacterial solution is absorbed.

[0323] (9) Incubate overnight at 37 °C in an inverted position.

[0324] (10) The next day, pick a single colony from the plate and inoculate it into an LB medium with Kan resistance, and incubate with shaking at 37 °C for 12 - 16 hours.

[0325] V. In vitro transcription of circRNA: Extract the plasmid, linearize the DNA template by digestion with enzymes, use an in vitro transcription kit, add transcription raw materials such as RNA polymerase, ribonucleoside triphosphates, and DNA template in vitro, incubate at 37 °C for 2 h, add enzyme-free water and GTP, incubate at 55 °C for 30 min, and transcribe circRNA using DNA as the template.

[0326] VI. Verification by precast gel electrophoresis: Detect the circularization efficiency of the synthesized circRNA, and the precast gel electrophoresis pattern is as Figure 11 shown.

[0327] VII. Preparation of liposomes: Prepare liposomes according to the following ratio. The molar ratio of cationic liposome: helper lipid: cholesterol: polyethylene glycol is 50:10:38.5:1.5. The cationic liposome is SM-102, the helper lipid is dioleoyl phosphatidylethanolamine, abbreviated as DOPE, and the polyethylene glycol is dimyristoyl glycerol-polyethylene glycol 2000, abbreviated as DMG-PEG2000.

[0328] VIII. Synthesis of lipid nanoparticles: Prepare the required volume of nanoparticles according to the volume ratio of liposome to mRNA of 3:1, mix gently by shaking, and prepare the lipid nanoparticles required for transfecting cells.

[0329] IX. Cell transfection: One day before transfection, seed the cells and observe their status. After the cell density reaches approximately 70%-90%, transfect the synthesized circRNA containing the insertion position of the IRES enhancer element into Hela and HEK-293T cells; transfect IRES-A1-X2, IRES-A1-X5, IRES-A2-X2, IRES-A3-X1, and IRES-A4-X1 circRNA into Hela and HEK-293T cells. After transfection, gently shake and mix well, and then place them in the cell culture incubator for culture.

[0330] X. Detect the fluorescence intensity after 24 h: After 24 h of transfection, take out the well plate, equilibrate it to room temperature, add the luciferin substrate, and allow the cells to lyse fully for 10 minutes. Then, detect the luminescence signal using a microplate reader. The results of exploring the optimal insertion position of the IRES enhancer element, the transfection results of IRES A1-4, and the transfection results of IRES-A1-X2, IRES-A1-X5, IRES-A2-X2, IRES-A3-X2, and IRES-A4-X2 are as Figure 12 .

[0331] In addition, the present disclosure also modifies based on the CVB3 IRES sequence. First, the present disclosure explores the optimal position for inserting the IRES enhancer element into the CVB3 IRES sequence. The present disclosure designs 11 positions for inserting the gene sequence, and it is experimentally found that there are 4 insertion positions with translation effects. The 4 specific insertion positions are respectively that the IRES enhancer element is inserted between domain Ⅰ and domain Ⅱ, named IRES-1; the IRES enhancer element is inserted into the stem-loop structure of domain Ⅱ, named IRES-2; the IRES enhancer element is inserted into the stem-loop structure (Distal loop) of domain Ⅳ, named IRES-3; the IRES enhancer element is inserted into the stem-loop structure (Proximal loop) of domain Ⅳ, named IRES-4. The IRES enhancer element can enhance the binding of the CVB3 IRES sequence to ribosomes, thereby promoting the translation of circRNA. The IRES enhancer element exists in circRNA, folds into a structure similar to the initial tRNA, recruits more ribosomes, and binds translation regulatory factors such as ITAF, and then introduces the ribosomes into the interior of circRNA to bind and initiate protein translation. Ribosomes are highly diverse protein structures found in all cells. Under the action of translation regulatory factors such as ITAF, the ribosomes are introduced into the internal structure of circRNA to bind and initiate. Secondly, the present disclosure screens 4 IRES enhancer elements, and then explores the influence of IRES enhancer elements with different sequences having the same enhancement effect on the translation effect of cicRNA. These 4 IRES enhancer elements all have the function of recruiting ribosomes, thus affecting the translation of circRNA. The 4 screened IRES enhancer element sequences are respectively cloned into the DNA template, and the composition of the DNA template is the same as above, and the optimal sequence is screened to obtain the gene sequence with high protein expression. Finally, based on the optimal position of the IRES enhancer element and the screening of the IRES enhancer element, a series of repeat sequence designs are carried out on the IRES enhancer element to obtain the sequence for the IRES enhancer element to achieve its function. Compared with the original sequence, IRES-1-A1-X2 in the screened sequence is increased by about 2.3 times.

[0332] The IRES enhancer element is:

[0333] Serial number IRES sequence A1 CCGGCGGGT A2 GTTTCATTTTATTCCTATAC A3 CAATTGAGAGATCGTTACCATATAGCTA A4 ATTTTATTCCTA

[0334] or a repeat sequence of the above sequence, for example:

[0335] Serial number IRES repeat sequence A1-X2 CCGGCGGGTCCGGCGGGT A1-X5 CCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGTCCGGCGGGT A2-X2 GTTTCATTTTATTCCTATACGTTTCATTTTATTCCTATAC

[0336] Example Six Development of Circular RNA Purification Process

[0337] There are relevant literature reports on the application of arginine in the purification process of biological products. The main mechanisms of action of arginine are as follows: (1) It can inhibit the formation of nucleic acid aggregates and promote the dissociation of complexes; (2) The addition of arginine increases the overall hydrophobicity of the solution, thereby promoting the hydrophobic interaction between the aggregate-nucleic acid complex and the affinity probe, and reducing the non-specific adsorption of the affinity chromatography packing

[21] 。

[0338] There are relevant literature reports on the application of PEG in the purification process of biological products. In addition to forming a depletion zone around the solute, it has been found that PEG is also "repelled" outside the hydrophilic surface, such as most chromatographic media: adding PEG to the mobile phase of molecular sieve will cause the retention time to be delayed, and at the same time, it will also enhance the binding of substances to ion exchange and affinity chromatography packing

[22] 。

[0339] During ultrafiltration and purification, adding a certain concentration of arginine and PEG to the buffer improved the resolution of the collected peak and reduced the non-specific adsorption of nucleic acids, enabling the purity and total yield of cRNA to meet the requirements of the process scale-up. Circular nucleic acid purification process flow Figure 13 。

[0340] The choice and precision of the IVT reaction equipment determine the yield of the IVT in vitro reaction. Commonly used IVT reaction systems include thermostatic shaking metal baths, reaction kettles, temperature-controlled tanks, and waver shakers, etc

[0341] Common brands of thermostatic shaking metal baths include Thermo Fisher (model: Thermo Scientific Compact), etc. Reaction kettles include Mettler reaction kettles (models EasyMax102 and EasyMax402) and Radeys (model Mya4), etc. Temperature-controlled tanks include LePure and Bailinke water-cooled temperature-controlled LeXiaobao (model compatible with 1000mL and 500mL disposable PC square bottles, magnetic stirring, customized), etc. Waver shakers include Satorius, LePure, and Bailinke, etc. (model Waver10), etc

[0342] The reaction condition control of thermostatic shaking metal baths, reaction kettles, and temperature-controlled tank instruments is as follows:

[0343]

[0344] The reaction condition control of the Waver 10 shaker is as follows:

[0345]

[0346] The circular nucleic acid IVT reaction system consists of template DNA, T7 enzyme, dNTP, pyrophosphatase (PPI), and PEG, etc. It is required that the homogeneity of the template DNA is ≥95%, and the metal residue is less than 10 PPM; for the T7 enzyme, the brands required are Promega, Thermo, and Kaika; the PEG is PGE2000 - 8000 from Thermo or Sigma, and its function is to increase the viscosity of the IVT reaction system and reduce the fluidity of the reaction solution, so as to prevent Mg 2+ from randomly attacking the base sequence of circular nucleic acid to control the generation of Naked RNA.

[0347] The IVT reaction conditions are controlled as follows in the table; the cyclization rate ≥75%, and Naked RNA ≤3%:

[0348]

[0349]

[0350] The cyclization steps are as follows:

[0351] The first step: ultrafiltration and buffer exchange using a membrane package or hollow fiber column

[0352] The ultrafiltration membrane package is made of PES material, and the pore size ranges from 100KD, 300KD, 500KD, 750KD to 1000KD

[0353] The hollow fiber column is made of PES material, and the pore size ranges from 100KD, 300KD, 500KD, 750KD to 1000KD

[0354] The manufacturers of ultrafiltration membrane packages and hollow fiber columns include Cytiva, Sartorius, Kobite, and Acephar, etc.

[0355] Through ultrafiltration or filtration using a hollow fiber column, impurities such as T7 enzyme, DNA residue, dNTP, and all or part of intron can be removed.

[0356] The solution for ultrafiltration replacement 1: NaCl + EDTA, pH = 7.4 ± 0.1

[0357] Control conditions Control range Case Sodium chloride concentration 20 mM - 1 M 100 mM EDTA concentration 5 mM - 20 mM 10 mM

[0358] The solution for ultrafiltration replacement 2: NaCl + arginine + EDTA, pH = 7.4 ± 0.1

[0359] Control conditions Control range Case Sodium chloride concentration 20 mM - 1 M 100 mM Arginine concentration 20 mM - 1 M 100 mM EDTA concentration 5 mM - 20 mM 10 mM

[0360] The volume of the solution for ultrafiltration replacement is 10 - 50 times (30 times), and the results are shown in Figure 14 .

[0361] Animal cell experiments show that: precursor linearity, intron, aggregates and circular isomers affect cytotoxicity; open rings affect cell activity, and excessive open rings have strong immunogenicity.

[0362] Step 2: NHS-activated beads based on 4FF or 6FF are coupled with Oligo probes. The Oligo probes are designed according to the structures of precursor linearity and intron generated by biological reactions (IVT), usually 1 - 3 in number. This filler is named NHS affinity, and the particle size of the affinity filler ranges from 5um to 90um. Generally, it is affinity anion selection, adopting an anion flow-through mode.

[0363] The NHS-activated beads based on 4FF or 6FF are from Cytiva, Chutian Microspheres, Boge Long, and Bailinke, etc.

[0364] Process condition 1: NHS affinity + NHS affinity

[0365] The elution solution is: NaCl + arginine + EDTA, pH = 7.4 ± 0.1

[0366] Control conditions Control range Case Sodium chloride concentration 20 mM - 1 M 100 mM Arginine concentration 20 mM - 1 M 100 mM EDTA concentration 5 mM - 20 mM 10 mM

[0367] In the first NHS affinity process, 5 - 20% of the A260 flow-through starting peak signal is discarded, and the solution collected in the 5% - 20% section mainly contains aggregates and precursor linearity.

[0368] All of the flow-through peaks of the second NHS affinity are collected.

[0369] The precursor linearity, intron, aggregates and circular isomers in the samples collected from the two affinity flow-throughs can be controlled to be <1%. The total yield of circular nucleic acids ≥ 30%, and the yield of circular nucleic acids with high interest rate ≥ 70%. The results are shown in Figure 15 ,16,18

[0370] Process condition 2: NHS affinity + molecular sieve

[0371] The molecular sieves include 4FF, 6FF, core400, core500, core700, core1000, SEC1000, and SEC2000, and the particle size of the filler ranges from 5um to 90um

[0372] In the first NHS affinity process, 5 - 20% of the A260 flow-through starting peak signal is discarded, and the solution collected in the 5% - 20% section mainly contains aggregates and precursor linearity.

[0373] Molecular sieve: Gradient elution is adopted

[0374] The condition of the equilibration solution is: PB + NaCl + EDTA, pH = 7.4 ± 0.1

[0375] Control conditions Control range Case PB concentration 50 - 100 mM 75 mM Sodium chloride concentration 0 mM - 200 mM 100 mM EDTA concentration 5 mM - 20 mM 10 mM

[0376] The elution solution conditions are: PB + NaCl + PEG + EDTA, pH = 7.4 ± 0.1

[0377]

[0378] For the sample precursor collected by one - step affinity + molecular sieve, the linear form, intron, polymer and cyclic isomers can be controlled within < 1%, and a small amount of open - loop can be removed. The total yield of circular nucleic acid ≥ 30%, and the theoretical yield of circular nucleic acid ≥ 70%

[0379] Process condition 3: NHS affinity + CHT

[0380] In the first NHS affinity process, 5 - 20% of the A260 flow - through starting peak signal is discarded, and the solution collected in the 5% - 20% section is mainly polymers and precursor linear forms.

[0381] CHT: Gradient elution is adopted

[0382] The equilibrium solution conditions are: PB + NaCl + EDTA, pH = 7.4 ± 0.1

[0383] Control conditions Control range Case PB concentration 50 - 100 mM 75 mM Sodium chloride concentration 0 mM - 200 mM 100 mM EDTA concentration 5 mM - 20 mM 10 mM

[0384] The elution solution conditions are: PB + NaCl + PEG + EDTA, pH = 7.4 ± 0.1

[0385]

[0386]

[0387] For the sample precursor collected by one - step affinity + CHT, the linear form, intron, polymer and cyclic isomers can be controlled within < 1%, and some open - loop can be removed. The total yield of circular nucleic acid ≥ 30%, and the theoretical yield of circular nucleic acid ≥ 70%.

[0388] References:

[0389] [1]Group I permuted intron - exon(PIE)sequences self - splice to produce circular exons[J].

[0390] [2]!!!INVALID CITATION!!![2],

[0391] [3]!!!INVALID CITATION!!![3],

[0392] [4]Chen Y G,Chen R,Ahmad S,et al.N6-Methyladenosine Modification Controls Circular RNA Immunity[J].Molecular Cell,2019,76(1):96-109.e9.

[0394] [5]!!!INVALID CITATION!!![5],

[0395] [6]!!!INVALID CITATION!!![6],

[0396] [7]Rausch J W,Heinz W F,Payea M J,et al.Characterizing and circumventing sequence restrictions for synthesis of circular RNA in vitro[J].Nucleic Acids Res,2021,49(6):e35.

[0397] [8]Beverly M,Hagen C,Slack O.Poly A tail length analysis of in vitro transcribed mRNA by LC-MS[J].Anal Bioanal Chem,2018,410(6):1667-77.

[0398] [9]<Representation of the secondary and tertiary structure of group I introns.pdf>[J].

[0399]

[10] Zaug A J,Mcevoy M M,Cech T R.Self-splicing of the group I intron from Anabaena pre-tRNA:requirement for base-pairing of the exons in the anticodon stem[J].Biochemistry,1993,32(31):7946-53.

[0400]

[11] Chen R, Wang S K, Belk J A, et al. Engineering circular RNA for enhanced protein production[J]. Nature Biotechnology, 2022,

[0401]

[12] <a 9-nt segment of a cellular mrna can function as an.pdf>[J].

[0402]

[13] <Biochemical and functional analysis of a 9-nt RNA.pdf>[J].

[0403]

[14] Fan X,Yang Y,Chen C,et al.Pervasive translation of circular RNAsdriven by short IRES-like elements[J].2020,

[0405]

[15] Tusup M,Kundig T,Pascolo S.An Eif4g-Recruiting Aptamer Increasesthe Functionality of in VitroTranscribed Mrna[J].EPH-International Journal ofMedical and Health Science,2018,4(2):29-34.

[0406]

[16] Lai W C,Zhu M,Belinite M,et al.Intrinsically UnstructuredSequences in the mRNA 3'UTR Reduce theAbility of Poly(A)Tail to EnhanceTranslation[J].J Mol Biol,2022,434(24):167877.

[0407]

[17] Chen C K,Cheng R,Demeter J,et al.Structured elements driveextensive circular RNA translation[J].MolCell,2021,81(20):4300-18e13.

[0408]

[18] Bhattacharyya S, Das S. Mapping of secondary structure of the spacer region within the 5′-untranslated region of the coxsackievirus B3 RNA: possible role of an apical GAGA loop in binding La protein and influencing internal initiation of translation[J]. Virus Research, 2005, 108(1-2): 89-100.

[0409]

[19] Gao W, Li Q, Zhu R, et al. La Autoantigen Induces Ribosome Binding Protein 1(RRBP1) Expression through Internal Ribosome Entry Site(IRES)-Mediated Translation during Cellular Stress Condition[J]. International Journal of Molecular Sciences, 2016, 17(7):

[0410]

[20] Prokaryotic-like cis elements in the cap-independent internal initiation of translation on picornavirus RNA[J].

[21] Shukla D, Zamolo L, Cavallotti C, et al. Understanding the Role of Arginine as an Eluent in Affinity Chromatography via Molecular Computations[J]. The Journal of Physical Chemistry B, 2011, 115(11): 2645-54.

[0412]

[22] Levanova A,Poranen M M.Application of steric exclusionchromatography on monoliths for separation andpurification of RNA molecules[J].Journal of Chromatography A,2018,1574(50-9.

[0413] Sequence information:

[0414]

[0415]

[0416]

[0417]

[0418]

[0419]

[0420]

[0421]

[0422]

[0423]

[0424]

[0425]

Claims

1. A recombinant nucleic acid molecule for preparing circular RNA, along the 5' to 3' direction, the recombinant nucleic acid molecule comprises elements operably linked in sequence: An optional 5' homologous arm, a 3' half-intron fragment, a circularization fragment, a 5' half-intron fragment, and an optional 3' homologous arm; The 5' end of the circularization fragment contains a 3' circularization recognition fragment, and the 3' end contains a 5' circularization recognition fragment; The 5' half intron fragment and the 3' half intron fragment are derived from group I introns, preferably group I introns from Anabeana, preferably derived from Anabaena tRNA leu ; The 5' half-intron fragment and the 3' half-intron fragment form an intron sequence along the 5' to 3' direction; the nucleotide sequence of the 5' half-intron fragment contains a partial sequence of the intron sequence closer to the 5' direction, and the nucleotide sequence of the 3' half-intron fragment contains the remaining part of the intron sequence closer to the 3' direction, and the 5' half-intron fragment contains an internal guide (IG) sequence; Wherein: (1) The IG sequence of the intron is 3nt of 5'-GNN-3', and the 5' circularization recognition fragment is 3nt of 5'-NNY-3', where Y is C or T / U; N is an optional nucleotide, and the bases corresponding to the IG sequence and the 5' circularization recognition fragment have strict base complementary pairing or GU wobble pairing; (2) At least 2nt of the 3' circularization recognition fragment, and it has no obvious homology with the 5' circularization recognition fragment and its adjacent sequences; (3) The adjacent sequences of the 5' circularization recognition fragment and the 3' circularization recognition fragment contain at least 3nt of internal homology sequences; The obvious homology refers to complementary pairing of two or more nucleotides.

2. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, the 5' end of the circularization fragment has a 3' circularization recognition fragment-internal homology sequence fragment, and the 3' end has an internal homology sequence fragment-5' circularization recognition fragment, and there are optionally 0, 1, 2, 3, 4, or 5 unpaired bases between the circularization recognition fragment and the internal homology sequence; Preferably, the internal homology sequence is 3nt, 4nt, or 5nt; and / or, the length of the 3' circularization recognition fragment is 2nt.

3. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, the circularization fragment comprises elements operably linked in sequence: The first part of the target polypeptide coding region or non-coding region containing the 3' circularization recognition fragment, a translation initiation element, and the second part of the target polypeptide coding region or non-coding region containing the 5' circularization recognition fragment; That is, the target polypeptide coding region or non-coding region is cleaved as follows: the second part (containing the 5' circularization recognition fragment)|(containing the 3' circularization recognition fragment) the first part.

4. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, the circularization fragment comprises elements operably linked in sequence: The first part of the translation initiation element containing the 3' circularization recognition fragment, the first part of the target polypeptide coding region or non-coding region, and the second part of the translation initiation element containing the 5' circularization recognition fragment; That is, the translation initiation element is cleaved as follows: the second part (containing the 5' circularization recognition fragment)|(containing the 3' circularization recognition fragment) the first part.

5. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the circularized fragment comprises sequentially operably linked elements: A 5' spacer portion, an optional translation initiation element, a target polypeptide coding region or a non-coding region, and a 3' spacer portion; Preferably, the exogenously introduced bases on the 5' spacer sequence and the 3' spacer sequence are less than 10; Preferably, three nucleotides are inserted at the third or fourth nucleotide of the 5' spacer sequence as internal homologous sequences, and two or three nucleotides of the 5' spacer sequence are used as 3' loop recognition fragments; six nucleotides are inserted at the 3' end of the 3' spacer sequence, wherein the first to third nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, the fourth to sixth sequences are NNY, and the bases corresponding to the NNY and IG sequences have strict base complementary pairing or GU wobble pairing; Also preferably, 5 nucleotides are inserted at the 5' end of the 5' spacer sequence, wherein the 1st to 2nd nucleotides serve as the 3' looping recognition fragment, the 3rd to 5th nucleotides serve as the internal homologous sequence, and the two or three nucleotides of the 5' spacer sequence itself serve as the 3' looping recognition fragment; and 6 nucleotides are inserted at the 3' end of the 3' spacer sequence, wherein the 1st to 3rd nucleotides are complementary to the nucleotides inserted in the 5' spacer sequence, and the 4th to 6th sequences are NNY.

6. The recombinant nucleic acid molecule for preparing circular RNA according to claim 5, The 5' spacer moiety is selected from: The 3' spacer moiety is selected from: Preferably, the 5' spacer sequence and the 3' spacer sequence are used alone or together; preferably, the 5' spacer sequence and the 3' spacer sequence are both polyA+8CA, preferably the polyA fragment contains 10-100 A, more preferably 60-80 A, most preferably 61 A at the 5' end and 71 A at the 3' end; Most preferably, it is a combination of 5'61A8CA and 3'71A8CA.

7. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the translation initiation element is an IRES sequence; Preferred are the following IRES sequences: Taura syndrome virus, Triatoma virus, Theiler's encephalomyelitis virus, simian virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, reticuloendotheliosis virus, Forman poliovirus 1, Autographa californica multiple nucleopolyhedrovirus, Kashmir bee virus, human rhinovirus 2, Homalodisca coagulata virus-1, human immunodeficiency virus type 1, Homalodisca coagulata virus-1, Pediculus humanus corporis virus, hepatitis C virus, hepatitis A virus, GB hepatitis virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinovirus, Ectropis obliqua-like virus, encephalomyocarditis virus (EMCV), Drosophila C virus, tobacco mosaic virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, Hibiscus chlorotic ringspot virus, classical swine fever virus, human FGF2, human SFTPA1, human AMLl / RUNXl, Drosophila antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAPl, human cmyc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, saliva virus, Coxsackievirus, Echovirus, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, picornavirus, turnip crinkle virus, aptamer of eIF4G, Coxsackievirus B3 (CVB3) or Coxsackievirus A (CVB1 / 2); More preferably, the IRES is the IRES sequence of Coxsackievirus B3 (CVB3); Even more preferably, the IRES is the IRES sequence of encephalomyocarditis virus.

8. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the translation initiation element is the CVB3 IRES sequence inserted with an IRES enhancer element; Preferably, the insertion positions of the IRES enhancer element are: between domain I and domain II of the IRES sequence (named IRES-1), at the stem-loop structure of domain II of the IRES sequence (named IRES-2), at the distal loop of domain IV of the IRES sequence (named IRES-3), at the proximal loop of domain IV of the IRES sequence (named IRES-4); Preferably, the IRES enhancer element is: or a repeat sequence of the above sequences: 。 9. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the translation initiation element sequence comprises one or more combinations of the following sequences: IRES sequence, 5'UTR sequence, Kozak sequence, sequence containing m6A modification, complementary sequence of ribosomal 18S rRNA.

10. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the target polypeptide coding region or non-coding region is a protein coding region encoding a human protein or a non-human protein; Optionally, the protein coding region encodes an antibody; Optionally, the human protein or non-human protein is selected from hFIX, SP-B, VEGF-A, human methylmalonyl-CoA mutase (hMUT), CFTR, cancer autoantigen and gene editing enzymes, such as Cpf1, zinc finger nuclease (ZFN) and transcription activator-like effector nuclease (TALEN); Optionally, the protein is a protein for therapeutic use; Optionally, wherein the antibody is a human anti-HIV antibody; Optionally, wherein the antibody is a bispecific antibody; Optionally, wherein the bispecific antibody binds CD3 and CLDN6 or binds CD19 and CD22. Optionally, wherein the protein is a protein for diagnostic use; Optionally, wherein the protein coding region encodes Gauss luciferase (Gluc), firefly luciferase (Fluc), enhanced green fluorescent protein (eGFP), human erythropoietin (hEPO) or Cas9 endonuclease; Optionally, the protein coding region comprises at least two coding regions, wherein a linker is connected between any two adjacent coding regions; preferably, the linker is a polynucleotide encoding 2A peptide; Optionally, a translation initiation element is connected between any two adjacent coding regions; optionally, the translation initiation element located between any two adjacent coding regions comprises one or more combinations of the following sequences: IRES sequence, 5'UTR sequence, Kozak sequence, sequence containing m6A modification, complementary sequence of ribosomal 18S rRNA.

11. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the recombinant nucleic acid molecule further comprises an insertion element, and the insertion element is located upstream of the translation initiation element; the insertion element is selected from at least one of the following groups (i)-(iii): (i) Transcription level regulatory element, (ii) Translation level regulatory element, (iii) Purification element; Optionally, the insertion element comprises a sequence of one or more combinations of the following: Untranslated region sequence, polyA sequence, aptamer sequence, riboswitch sequence, sequence binding a transcriptional regulatory factor.

12. The recombinant nucleic acid molecule for preparing circular RNA according to claim 1, wherein the target polypeptide coding region or non-coding region is a protein coding region encoding a human protein or a non-human protein; the recombinant nucleic acid molecule has the homologous arms, wherein, The length of each homologous arm is about 5-50 nucleotides; preferably, the length of each homologous arm is about 9-19 nucleotides.

13. A recombinant expression vector, which comprises the recombinant nucleic acid molecule according to claim 1.

14. A circularized precursor nucleic acid molecule, which is transcribed from the recombinant expression vector according to claim 13, wherein the circularized precursor nucleic acid molecule fragment comprises: Optional 5' homologous arm, 3' half-intron fragment, circularized fragment, 5' half-intron fragment and optional 3' homologous arm; the 5' end of the circularized fragment contains a 3' circularization recognition fragment, and the 3' end contains a 5' circularization recognition fragment.

15. A method for preparing circular RNA in vitro, comprising: 1) Transcription step: transcribing the recombinant expression vector according to claim 13 to form a circularized precursor nucleic acid molecule; 2) Circularization step: subjecting the circularized precursor nucleic acid molecule to a circularization reaction to obtain circular RNA; Optionally, the method further comprises a step of purifying the circular RNA.

16. A method for purifying circular RNA, the method comprising ultrafiltration, and a combination of the same or different means of affinity chromatography, molecular sieve chromatography, and CHT; The preferred combination is: two affinity chromatographies, affinity chromatography and molecular sieve chromatography, or affinity chromatography and CHT chromatography; Preferably, arginine and PEG are added to the ultrafiltration reagent or chromatography reagent; preferably, the reagent is a buffer; preferably, the PEG is PEG2000-8000.

17. The circular RNA molecule obtained by the recombinant nucleic acid molecule according to claim 1 or the vector according to claim 13, and the method according to claim 15.

18. A circular RNA, along the 5' to 3' direction, which comprises elements arranged in the following order: Translation initiation element, target polypeptide coding region or non-coding region; Optionally, the circular RNA comprises a 5' spacer sequence and a 3' spacer sequence located between the 5' end of the translation initiation element and the 3' end of the coding element; Each of the elements is defined as in claim 1; Preferably, the circular RNA comprises the spacer sequence according to claim 6; preferably, the circular RNA comprises less than 9 nt of exogenous inserted nucleotides; Preferably, the circular RNA comprises the IRES sequence according to claim 8; Preferably, the circular RNA comprises a modification; preferably, the modification is m6A.

19. A composition, the composition comprising the recombinant nucleic acid molecule according to claim 1, the recombinant expression vector according to claim 13, the circular RNA according to claim 17 or 18; preferably comprising the circular RNA according to claim 17 or 18; Optionally, the composition further comprises one or more pharmaceutically acceptable carriers; Optionally, the pharmaceutically acceptable carrier is selected from lipids, polymers or lipid-polymer complexes.

20. A method for expressing a target polypeptide in a cell, wherein, The method comprises the step of transferring the circular RNA according to claim 17 or 18, or the composition according to claim 19 into cells.

21. A method for preventing or treating a disease, wherein, The method comprises administering to a subject the circular RNA according to claim 17 or 18, or the composition according to claim 19.

22. A design method for a recombinant nucleic acid molecule for preparing circular RNA, the method comprising: 1) Providing a candidate target polypeptide coding region or non-coding region, translation initiation element, and optional 5' spacer sequence and 3' spacer sequence; 2) Search for or set 5'-loop recognition fragments, 3'-loop recognition fragments, and internal homology sequences in the target polypeptide coding region or non-coding region, translation initiation elements, or 5' and 3' spacer sequences to design a circularization fragment: The search operation is as follows: S-1) Search for the following sequences in the above elements: 5'-internal homology sequence - 5'-loop recognition fragment - 3'-loop recognition fragment - 3'-internal homology sequence; S-2) Split the corresponding element at the junction of the 5'-loop recognition fragment and the 3'-loop recognition fragment according to the search result to form a circularization sequence; The setting operation is as follows: Set 5'-internal homology sequence - 5'-loop recognition fragment sequence and 3'-loop recognition fragment - 3'-internal homology sequence at both ends of the sequence: 5'-spacer sequence - translation initiation element - target polypeptide coding region or non-coding region - 3'-spacer sequence to form a circularization fragment; Wherein: A) The 5'-internal homology sequence and the 3'-internal homology sequence are reverse complementary sequences containing at least 3 nt; B) The 3'-loop recognition fragment is a sequence of at least 2 nt and does not have significant homology with the adjacent sequence of the 5'-loop recognition fragment; C) The 5'-loop recognition fragment is a 3-nt sequence and is NNY, where N is an optional base and Y is C or T / U; 3) According to the circularization fragment obtained in step 2), modify the intron IG sequence to be used. The IG sequence is a 3-nt "GNN" and has strict base complementary pairing or GU wobble pairing with the nucleotide corresponding to the 5'-loop recognition fragment; 4) Split the modified intron sequence into a 3'-semi-intron fragment and a 5'-semi-intron fragment, and operably link them in the following order to form a recombinant nucleic acid molecule for preparing circular RNA: Optional 5'-homologous arm, 3'-semi-intron fragment, circularization fragment, 5'-semi-intron fragment, and optional 3'-homologous arm; the 5' end of the circularization fragment contains a 3'-loop recognition fragment, and the 3' end contains a 5'-loop recognition fragment; Preferably, the setting is to insert 3 nucleotides as the internal homology sequence at the third or fourth nucleotide of the 5'-spacer sequence, and use two or three nucleotides of the 5'-spacer sequence itself as the 3'-loop recognition fragment; insert 6 nucleotides at the 3' end of the 3'UTR, where the 1st - 3rd nucleotides are complementary to the nucleotides inserted in the 5'-spacer sequence, and the 4th - 6th sequence is NNY and has strict base complementary pairing or GU wobble pairing with the nucleotides corresponding to the IG sequence; More preferably, the setting is to insert 5 nucleotides at the 5' end of the 5'-spacer sequence, where the 1st - 2nd nucleotides are used as the 3'-loop recognition fragment, the 3rd - 5th nucleotides are used as the internal homology sequence, and use two or three nucleotides of the 5'-spacer sequence itself as the 3'-loop recognition fragment; insert 6 nucleotides at the 3' end of the 3'-spacer sequence, where the 1st - 3rd nucleotides are complementary to the nucleotides inserted in the 5'-spacer sequence, and the 4th - 6th sequence is NNY and has strict base complementary pairing or GU wobble pairing with the nucleotides corresponding to the IG sequence.

23. A computer module, the module stores a program for implementing the design method described in claim 22.

24. The CVB3 IRES sequence into which an IRES enhancing element is inserted, wherein: The insertion positions of the IRES enhancing element are: between domain I and domain II of the IRES sequence (named IRES-1), at the stem-loop structure of domain II of the IRES sequence (named IRES-2), at the stem-loop structure (Distalloop) of domain IV of the IRES sequence (named IRES-3), at the stem-loop structure (Proximal loop) of domain IV of the IRES sequence (named IRES-4); The IRES enhancing element is: Or a repeat sequence of the above sequence: 。 25. Derived from Anabaena tRNA leu Mutants of group I introns or fragments with self-splicing activity derived from the mutants, and comprising at least one set of mutations as follows: 1) The second base in the original IG sequence 5'-GAG-3' is mutated to a base other than A; Or / and 2) The third base (the G at the 3') in the original IG sequence 5'-GAG-3' is mutated to a base other than G Preferably, it contains at least a mutation of the third base (the G at the 3'); Preferably, it is a combination of SEQ ID No. 44 and 45 having the above mutations.

Citation Information

Patent Citations

  • Circular RNA for translation in eukaryotic cells

    CN112399860A

  • Self-cyclized RNA constructs

    CN115997018A

  • Circular RNA compositions and methods

    WO2020237227A1

  • Circular RNA compositions and methods

    WO2021236855A1

  • Circular RNA and preparation method thereof

    WO2023046153A1