Method for preparing circular RNA and nucleic acid sequence for use in said method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RINUAGENE BIOTECHNOLOGY CO LTD
- Filing Date
- 2024-07-05
- Publication Date
- 2026-05-05
AI Technical Summary
It is difficult to efficiently prepare high-purity circular RNA in the prior art, especially on the scale of industrial production, and conventional purification methods have problems such as low loading, high cost, complex process and low resolution.
By inserting tag sequences into both end arms of the linear RNA precursor, the specific binding of the tag sequence to the affinity chromatography medium is used to achieve the separation of circular RNA and by-products. The specific method includes generating linear RNA precursors with poly(A) tags, performing self-splicing and cyclization reactions, followed by contact with affinity chromatography media, separating and collecting the supernatant to obtain high-purity circular RNA.
This method can significantly improve the production efficiency, purity and yield of circular RNA, reduce production costs, simplify the process, improve the safety of circular RNA, basically eliminate linear RNA by-products, and meet the needs of industrial production.
Smart Images

Figure CN121986162A_ABST
Abstract
Description
Method for preparing circular RNA and nucleic acid sequence used in the method Technical Field
[0001] The present invention relates to methods and nucleic acid sequences for preparing and purifying circular RNA. Background Art
[0002] In recent years, the development and application of mRNA vaccines or mRNA drugs have played a vital role in controlling the COVID-19 pandemic, and have also significantly accelerated the development of mRNA drugs in other disease treatment areas. It has been noted that linear mRNA drugs face major challenges in vivo, such as low stability and short half-life. However, circular RNA, because it does not contain free ends, can effectively avoid degradation by exonuclease in cells, thus having significantly higher stability than linear mRNA molecules (see non-patent document 1), providing an effective candidate solution to address the above challenges.
[0003] Circular RNA (circRNA) is a class of covalently closed circular RNA molecules. Since the first discovery in 1976 that the genomes of plant viroids are circular RNA molecules, hundreds of natural circRNAs have been identified in viruses (such as hepatitis delta virus (HDV)) and eukaryotic cells (such as mammalian cells). These molecules have been found to perform a variety of different biological functions in vivo, including but not limited to protein encoding, gene transcription regulation, microRNA sponges, and protein scaffolds (see Non-Patent Document 2). Related molecular biology and functional research and analysis have promoted the in vitro synthesis of circular RNAs. Chemical synthesis, ligase methods (T4 DNA ligase, T4 RNA ligase 1, T4 RNA ligase 2), and ribozyme methods (including group I intron ribozymes and group II intron ribozymes) have been reported (see Non-Patent Document 1).
[0004] Among them, the ribozyme method is one of the main approaches to synthesize circular RNA in the current field of nucleic acid drugs. It uses the PIE system (Permuted Intron-Exon System) established by modifying type I or type II self-splicing introns to prepare circular RNA. It is reported that the splicing of type I introns does not depend on any protein and is not affected by the presence of Mg. 2+Under the conditions of guanosine and guanosine, splicing can be achieved through a two-step transesterification reaction: in the first step, guanosine attacks the 5' splicing site (5'SS), and the 3' hydroxyl group of guanosine undergoes an ester exchange reaction at this site, causing splicing between the 5' exon and the 5' intron; in the second step, the 3'-OH of the intermediate formed after the first reaction attacks the 3' splicing site (3'SS), and a second ester exchange reaction occurs at this site, resulting in the exons on both sides of the intron being directly connected, and the 5' and 3' ends of the intron itself are cyclized and then cut off. According to the above mechanism of action, the natural type I self-splicing intron is divided into two, and the arrangement order of "exon 1-intron fragment 1-intron fragment 2-exon 2" is artificially replaced with the arrangement of "intron fragment 2-exon 2-exon 1-intron fragment 1", which can be used in Mg 2+ and guanosine, prepare circular RNA obtained by cyclization of the "exon 2-exon 1" fragment. For example, non-patent literature 3 and 4 respectively reported the PIE system established by transforming the type I self-splicing intron of Anabaena (fish algae) pre-Trna or T4 phage Td gene by the above-mentioned replacement method. The above-mentioned system can successfully cyclize RNAs of varying lengths of ~100nt or ~550nt. The PIE system based on these two type I self-splicing intron ribozymes is still widely used in circular RNA synthesis. Non-patent literature 5 also reported the optimization of the PIE system based on the Anabaena pre-tRNA type I intron ribozyme. By adding external homology arms, internal homology arms, etc., the in vitro cyclization efficiency can be significantly improved, and the length of the cyclizable RNA fragment can be increased to ~5kb, laying a good foundation for the synthesis of circular RNA drugs.
[0005] The synthesis process of the above-mentioned circular RNA will produce a variety of RNA byproducts, including but not limited to: introns on both sides that are cut off after the circularization reaction, linear RNA precursors that have not yet formed a circle, and linear and circular high-molecular-weight polymerized RNA produced by the polymerization of multiple linear precursors. The prior art has reported that the presence of these byproducts has various adverse effects on the final circular RNA biological product, not only reducing the purity and potency of the product, but also significantly increasing the immunogenicity in vivo (see Non-Patent Document 2). Therefore, there is an urgent need for methods that can efficiently prepare high-purity circular RNA molecules, especially preparation methods that can meet the production level of industrialization.
[0006] The existing preparation methods for separating by-products from circular RNA mainly utilize the difference in molecular weight between the two, or the difference in physical properties between linear and circular structures, such as size exclusion chromatography (SEC) or high-performance liquid chromatography (HPLC) (see non-patent documents 2 to 6, patent document 1). However, the above purification methods are difficult to distinguish between the target circular RNA and linear RNA precursors with very close molecular weights, as well as nicked RNA with exactly the same molecular weight. Therefore, the conventional practice in the industry is to perform RNase R digestion before SEC or HPLC to remove the above impurities. At the same time, SEC and HPLC themselves have common problems such as low loading capacity, high cost, complex process, and low resolution. After adding additional enzyme treatment steps, it is even more difficult to achieve simultaneous improvements in yield, efficiency, and purity, let alone industrial-grade circular RNA production.
[0007] Patent Document 2 reports a method for purifying linear mRNA with a poly(A) structure using column chromatography using oligo dT as an affinity ligand. Although Patent Document 2 also mentions the possibility of artificially adding poly(A) sequences to circular RNA molecules, this patented method adds poly(A) to the circularized region of the RNA molecule. Therefore, theoretically, it is impossible to simultaneously separate all self-splicing byproducts from the target circular RNA molecule in a single column chromatography step. In particular, it is impossible to separate linear RNA precursors, nicked RNA, and the target circular RNA molecule that also have poly(A) sequences but have not yet formed a circular structure.
[0008] Patent Document 3 reports another method for purifying circular RNA, which includes adding a poly(A) tail to linear RNA mixed with the circular RNA, followed by removing the poly(A)-tailed RNA using oligo(dT)25-coupled magnetic beads. However, this method primarily aims to isolate natural circular RNA and its corresponding linear transcript fragments, rather than targeting the various byproducts of circular RNA synthesis.
[0009] Patent Document 4 reports a method for incorporating at least one purification tag into the 3' or 5' end of a linear precursor RNA, thereby enabling the use of antisense oligo affinity purification of circular RNA. However, Patent Document 4 has explicitly excluded the use of poly(A) and its variants as purification tags as a whole. Specifically, poly(A) is an essential element constituting the circular RNA of Patent Document 4. If it is also inserted as a tag into the 3' or 5' end of the linear RNA precursor, it is clear that according to the design principle of affinity purification, it will be impossible to distinguish between the desired isolated circular RNA and uncircularized impurities. Patent Document 4 also clearly states that "as a result of self-splicing, the corresponding circularized RNA no longer contains 3' and / or 5' end purification tags." Therefore, Patent Document 4 actually provides technical teachings that are completely contrary to the present invention.
[0010] In addition, Patent Document 4 also records the use of SEQ ID NO: 208 (tctttaccctcgtcttgacg) and 209 (tatgctgttatccgtcgatt) as oligos for affinity purification of circular RNA. However, Figure 6 of the document shows that the enrichment effect of circular RNA occurs after the step of separating the target circular RNA from incompletely circularized impurities. In other words, in Patent Document 4, adding a purification tag to the 3' or 5' end does not directly improve the efficiency of the circularization reaction and the purity of the circular RNA.
[0011] References:
[0012] Non-patent literature 1: Chen et al., Frontiers in Bioengineering and Biotechnology (2021). 9:787881;
[0013] Non-patent literature 2: Liu et al., Molecular Cell (2022), 82: 1-15;
[0014] Non-patent document 3: Been et al., Nucleic Acids Research (1992). 20: 5357-5364;
[0015] Non-patent document 4: Ford et al., Proceedings of National Academy of Science (1994). 91: 3117-3121;
[0016] Non-patent literature 5: Wesselhoeft et al., Nature Communications (2018). 9: 2629;
[0017] Non-patent literature 6: Wesselhoeft et al., Molecular Cell (2019). 74: 508-520;
[0018] Patent Document 1: WO2019 / 236673A1;
[0019] Patent document 2: CN114381454A;
[0020] Patent Document 3: CN110283895A;
[0021] Patent document: 4: WO2023 / 073228A1.
[0022] Summary of the Invention
[0023] Through in-depth research, the inventors discovered that by inserting tag sequences into the two end arms of a linear RNA precursor used to prepare circular RNA, they can exploit the specific binding of the tag sequences with corresponding affinity ligands to separate the prepared target circular RNA molecules from various RNA byproducts generated during the preparation process using affinity chromatography. Compared to traditional circular RNA preparation methods, the present method utilizes an easy-to-use and industrially feasible affinity chromatography process, enabling industrial production while reducing production costs. The inventors also surprisingly discovered that inserting polyadenylic acid (PA) or its functional variants at both ends of the linear RNA precursor significantly reduces, or even essentially eliminates, the production of byproducts such as nicked RNA. This significantly reduces the amount of RNase R enzyme required for purification of the circular RNA, and can even eliminate this enzyme treatment step. This not only simplifies the process and reduces costs, but also significantly improves the production efficiency, product purity, and yield of the circularized RNA. Furthermore, the present method eliminates the need to insert a tag unrelated to the target circular RNA, improving the safety of the circular RNA. Furthermore, the presence of the tag sequence in all impurities greatly simplifies purity testing and quality control. The circular RNA produced using the present method is essentially free of linear RNA byproducts, providing an excellent active pharmaceutical ingredient for the development of circular RNA drugs. The present invention also screened for suitable locations within linear RNA precursors and their encoding DNA for the insertion of purification tags. The results showed that inserting the purification tag into a specific region of the linear RNA precursor's encoding DNA did not reduce the yield of the linear RNA precursor produced by in vitro transcription (IVT) of the encoding DNA.
[0024] Therefore, in one aspect, the present invention provides a method for preparing circular RNA, comprising:
[0025] Step A: Producing a linear RNA precursor, wherein the linear precursor comprises a 5' end arm, a 3' self-splicing site, a circularization region, a 5' self-splicing site, and a 3' end arm in an operably connected manner in sequence from 5' to 3' direction, wherein a first tag is inserted into the 5' end arm and a second tag is inserted into the 3' end arm;
[0026] Step B: subjecting the linear RNA precursor to conditions suitable for self-splicing at the 3' self-splicing site and the 5' self-splicing site, thereby obtaining a mixture comprising linear RNA fragments carrying the first tag and / or the second tag and circular RNA obtained by circularization of the circularization region;
[0027] Step C: contacting the mixture with an affinity chromatography medium capable of simultaneously binding the first tag and the second tag for a period of time sufficient to allow the affinity chromatography medium to bind to the linear RNA fragment containing at least one of the tags; and
[0028] Step D: separating the affinity chromatography medium and the mixture after contact, collecting the supernatant, and obtaining circular RNA.
[0029] The first tag and / or the second tag are independently selected from a poly(A) tag consisting of 20 to 100, preferably 30 to 90, more preferably 40 to 80, further preferably 45 to 70, and most preferably 50 to 65 consecutive adenine nucleotides or a functional variant thereof, and the first tag and the second tag are not present in the cyclization region and the circular RNA.
[0030] In some embodiments, the functional variant of the poly(A) tag used in the method of the present invention is a poly(A) tag in which one or more non-A bases are inserted, preferably 1 to 20 non-A bases are inserted, and more preferably 1 to 10 non-A bases are inserted.
[0031] In some embodiments, functional variants of the poly(A) tag used in the methods of the present invention include the following:
[0032] (1) a single element a, at least one element b, and at least one element c,
[0033] (2) only one element a, at least one element b, and at least one element d; or
[0034] (3) a single element a, at least one element b, at least one element c and at least one element d,
[0035] wherein the element a is composed of 20 or more consecutive adenine nucleotides, the element b is composed of 3 or more and less than 20 consecutive adenine nucleotides, the element c is composed of a nucleotide selected from uracil nucleotides, cytosine nucleotides, and guanine nucleotides, and the element d is composed of 2 or more and 20 or less nucleotides, wherein the nucleotides are arbitrarily selected from adenine nucleotides, uracil nucleotides, cytosine nucleotides, and guanine nucleotides, and the element d does not contain 3 or more consecutive adenine nucleotides, and the 5' and 3' terminal nucleotides are not adenine nucleotides.
[0036] When the Poly(A) tag contains two or more elements b, c, or d, the sequences of each two elements b may be the same or different, the sequences of each two elements c may be the same or different, and the sequences of each two elements d may be the same or different.
[0037] Furthermore, the element a and the element b, the element c and the element d, the elements b, the elements c, and the elements d are not adjacent to each other.
[0038] In some embodiments, element a used in the method of the present invention consists of more than 20 and less than 80 consecutive adenine nucleotides, preferably consists of 30 to 70, 35 to 65, 40 to 60, or 45 to 55 consecutive adenine nucleotides, and more preferably consists of 60 consecutive adenine nucleotides.
[0039] In some embodiments, the element b used in the method of the present invention consists of 3 to 10, 10 to 19, 12 to 15, 14 nt to 17, or 16 to 19, preferably 19 consecutive adenine nucleotides. In some embodiments, the number of elements b used in the method of the present invention is 2 to 10, preferably 2 to 5, and more preferably 3.
[0040] In some embodiments, the element c used in the method of the present invention is a guanine nucleotide. In some embodiments, the number of elements c used in the method of the present invention is 2 to 10, 3 to 8, 4 to 6, or 2 to 5, preferably 2.
[0041] In some embodiments, the element d used in the methods of the present invention consists of 3 to 18, 5 to 16, 4 to 10, or 6 to 12 nucleotides, preferably 6 nucleotides. In preferred embodiments, the element d used in the methods of the present invention can be independently selected from any one of GAUAUC, GUAUAC, GAAUCU, GCAUAUGACU, or GAUAUCGUAUAC. In some embodiments, the number of elements d used in the methods of the present invention is 0 to 5, preferably 1 to 3, and more preferably 1.
[0042] In some embodiments, when element c and element d are present at the same time, the total number of element c and element d used in the method of the present invention is 2 to 15, preferably 3 to 5, and more preferably 3.
[0043] In a preferred embodiment, the functional variant of the poly(A) tag used in the method of the present invention has any one structure selected from the following structures: element a-element c-element b-element c-element b-element c-element b-element c-element b, element b-element c-element b-element c-element a-element d-element b-element c-element b-element c-element b, element b-element c-element b-element c-element b-element d-element a-element c, element a-element d-element b-element c-element b-element c-element b-element b, or, element b-element c-element b-element c-element b-element d-element a.
[0044] In an optional embodiment, the functional variant of the poly(A) tag used in the method of the present invention may further comprise an element e, wherein the element e is composed of one or two consecutive adenine nucleotides, wherein the element e is located at the 3' end of the Poly(A) tag sequence and is adjacent to the element d or the element c.
[0045] In some embodiments, the linear RNA precursor in the method of the present invention comprises a 5' external homology arm and a 3' intron fragment in the 5' end arm in the 5' to 3' direction, and the first tag is inserted into the 5' external homology arm, or inserted into the 5' end region of the 3' intron fragment near the 5' external homology arm, the 5' end region is preferably 20 nucleotides at the 5' end of the 3' intron fragment, more preferably 15 nucleotides, and further preferably 10 nucleotides, or inserted at the 5' end of the 5' external homology arm. Upstream, and the linear RNA precursor in the method of the present invention also comprises a 5' intron fragment and a 3' external homology arm in the 3' end arm in the 5' to 3' direction, and the second tag is inserted in the 3' external homology arm, or inserted in the 3' end region of the 5' intron fragment close to the 3' external homology arm, the 3' end region is preferably 20 nucleotides at the 3' end of the 5' intron fragment, more preferably 15 nucleotides, further preferably 10 nucleotides, or inserted downstream of the 3' end of the 3' external homology arm.
[0046] In a preferred embodiment, the first tag is inserted at any position in the 5' external homology arm, and the second tag is inserted downstream of the 3' end of the 3' external homology arm (for example, the position of the last nucleotide residue at the 3' end of the 3' external homology arm). In other preferred embodiments, the first tag is inserted at any position in the 5' external homology arm, and the second tag is inserted at any position in the 3' external homology arm. In other preferred embodiments, the first tag is inserted at any position in the 5' external homology arm, and the second tag is inserted at any position in the 3' terminal region of the 5' intron fragment. In a more preferred embodiment, the first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted downstream of the 3' end of the 3' external homology arm (for example, the position of the last nucleotide residue at the 3' end of the 3' external homology arm). In other more preferred embodiments, the first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted at any position in the 3' external homology arm. In other more preferred embodiments, the first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue from the 5' end, and the second tag is inserted at any position in the 3' end region of the 5' intron fragment.
[0047] In some embodiments, the first tag is inserted at residue position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, starting from the first nucleotide residue at the 5' end of the 5' outer homology arm in the 5' to 3' direction. In a preferred embodiment, the first tag is inserted at residue position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, starting from the first nucleotide residue at the 5' end of the 5' outer homology arm. In a more preferred embodiment, the first tag is inserted at residue position 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, starting from the first nucleotide residue at the 5' end of the 5' outer homology arm.
[0048] For example, the term "inserted at position n" or "inserted at residue position n" means that the 5'-most and 3'-most nucleotide residues of the described insertion sequence are covalently linked to the residues at positions n and n+1, respectively, prior to insertion. For example, "inserted at position 1" means that the inserted sequence is located between the original residues at positions 1 and 2.
[0049] In some embodiments, the second tag is inserted into the 3' outer homology arm at residue position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, starting from the last nucleotide residue at the 3' end in the 3' to 5' direction. In other more preferred embodiments, the second tag is inserted into the 3' outer homology arm at residue position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, starting from the last nucleotide residue at the 3' end in the 3' to 5' direction. In some embodiments, the second tag is inserted into the 5' intron fragment at residue position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, starting from the last nucleotide residue at the 3' end in the 3' to 5' direction. In other more preferred embodiments, the second tag is inserted into the 5' intron fragment at residue position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 from the last nucleotide residue at the 3' end in the 3' to 5' direction.
[0050] In some embodiments, the circularization region comprises a 3' coding region fragment, a translation initiation element, and a 5' coding region fragment in a manner that is operably linked to each other in the 5' to 3' direction. In other embodiments, the circularization region comprises a 3' exon fragment, a 5' internal homology arm, an insert fragment, a 3' internal homology arm, and a 5' exon fragment in a manner that is operably linked to each other in the 5' to 3' direction. In optional embodiments, the circularization region comprises a first spacer between the insert fragment and the 5' internal homology arm, and a second spacer between the insert fragment and the 3' internal homology arm. In other embodiments, the insert fragment comprises a translation initiation element, or comprises a translation initiation element and a coding region. In some embodiments, the translation initiation element is an IRES sequence. In a preferred embodiment, the IRES sequence can be selected from, but not limited to, the following IRES sequences: Taura syndrome virus, blood-sucking assassin bug virus, Theile's encephalomyelitis virus, simian virus 40, fire ant virus 1, cereal aphid virus, reticuloendotheliosis virus, Forman polio virus 1, soybean looper virus, Kashmir bee virus, human rhinovirus 2, glass leafhopper virus-1, human immunodeficiency virus type 1, glass leafhopper virus-1, lice P disease Virus, hepatitis C virus, hepatitis A virus, GB hepatitis virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinovirus, tea geometrid-like virus, encephalomyocarditis virus (EMCV), fruit fly C virus, crucifer tobacco virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, hibiscus yellow ringspot virus, swine fever virus, human FGF2, human SFTPA1, human AML1 / RUNX1, Drosophila antennae, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAP1, human c-myc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kip1, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, Saccharomyces cerevisiae YAP1, human c-src, human FGF-1, simian picornavirus, turnip shrivelled disease virus, an aptamer to eIF4G, coxsackievirus B3 (CVB3), or coxsackievirus A (CVB1 / 2). In some embodiments, the IRES sequence is a wild-type IRES sequence or a modified IRES sequence. In certain embodiments, the IRES sequence is about 50 nucleotides in length.
[0051] In some embodiments, the inserted fragment comprises the coding sequence of a structural gene or a functional fragment thereof, or the sequence of a non-coding RNA or its complementary sequence, wherein the structural gene is selected from a polypeptide, a protein subunit, a protein active center, a protein or a protein hybrid of a non-natural catalytic group, a recombinant protein active subunit or active center / , a recombinant artificial enzyme or other biological effect units mainly composed of amino acids, and the non-coding RNA is selected from microRNA (miRNA), small interfering RNA (siRNA), PIWI protein interacting RNA (piRNA), transfer RNA-derived small RNA (tsRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), long non-coding RNA (lncRNA), pseudogene, ceRNA (competing endogenous RNAs), microRNA sponge or other types of non-mRNA RNA.
[0052] In optional embodiments, the length of the 5' and 3' outer homology arms are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt. In optional embodiments, the lengths of the 5' and 3' intron fragments are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, less than 70 nt, less than 80 nt, less than 90 nt, less than 100 nt, less than 150 ... 0nt, less than 90nt, less than 100nt, less than 150nt, less than 200nt, 5-200nt, 10-150nt, 50-200nt, 50-150nt, 5-60nt, 10-55nt, 15-50nt, 20-45nt, 25-40nt, 30-35nt, 10nt, 15nt, 20nt, 25nt, 30nt, 35nt, 40nt, 45nt, or 50nt. In optional embodiments, the length of the 5' and 3' exonic fragments is each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt.In optional embodiments, the length of the 5' and 3' internal homology arms are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt. In optional embodiments, the length of the insert is greater than 50 nt, greater than 100 nt, greater than 150 nt, greater than 200 nt, greater than 250 nt, greater than 300 nt, greater than 400 nt, greater than 500 nt, greater than 600 nt, greater than 1k nt, greater than 1.5k nt, greater than 2k nt, greater than 3k nt, less than 50 nt, less than 100 nt, less than 150 nt, less than 200 nt, less than 250 nt, less than 300 nt, less than 400 nt, less than 500 nt, less than 600 nt, less than 600 nt, less than 1k nt, less than 1.5k nt, less than 2k nt, less than 3k nt, 50-5k nt, 50-5k nt, 50-4k nt, 50-3k nt, 50-2k nt, 50-1.5k nt, 50-1k nt, 50~600nt, 100~550nt, 150~500nt, 200~450nt, 250~400nt, 300~350nt.
[0053] In a preferred embodiment, the 5' outer homology arm has the sequence shown in SEQ ID NO: 1 or 56, and the 3' outer homology arm has the sequence shown in SEQ ID NO: 2 or 57. In a preferred embodiment, both the 3' intron segment and the 5' intron segment are derived from a type I intron, preferably from a pre-tRNA gene of the cyanobacterium Anabaena or a Td gene of bacteriophage T4. In a more preferred embodiment, the 3' intron segment has the sequence shown in SEQ ID NO: 3, and the 5' intron segment has the sequence shown in SEQ ID NO: 4. In a preferred embodiment, both the 3' intron segment and the 5' intron segment are derived from a type II intron, preferably from a type II intron of a Clostridium, such as Clostridium tetani, or a type II intron of a Bacillus, such as Bacillus thuringiensis. In a more preferred embodiment, the 3' intron fragment and the 5' intron fragment are derived from a group II intron contained in the nucleotide sequence shown in SEQ ID NO: 5 or 6. In a preferred embodiment, the 3' exon fragment and the 5' exon fragment are derived from the 3' terminal region and the 5' terminal region of a natural exon, respectively, preferably from the cyanobacterium Anabaena pre-tRNA gene or the T4 phage Td gene. In a more preferred embodiment, the 3' exon fragment has the sequence shown in SEQ ID NO: 7, and the 5' exon fragment has the sequence shown in SEQ ID NO: 8. In a preferred embodiment, the 5' internal homology arm has the sequence shown in SEQ ID NO: 9, and the 3' internal homology arm has the sequence shown in SEQ ID NO: 10. In some embodiments, the first spacer is the same as or different from the second spacer.
[0054] In some embodiments, the affinity chromatography medium used in the methods of the present invention is operably linked to an affinity ligand capable of specifically binding to the first tag and / or the second tag. In preferred embodiments, the affinity ligand is selected from the group consisting of poly-X1, poly-X1-X2, poly-X1-X2-X3, and poly-X1-X2-X3-X4, wherein X1, X2, X3, and X4 are independently any one of A, G, C, T, and U. In more preferred embodiments, the affinity ligand is selected from the group consisting of oligo dT, oligo dC, oligo dG, and oligo dU.
[0055] In some embodiments, the affinity chromatography medium used in the method of the present invention is selected from any one of the group consisting of magnetic beads, dextran molecules, polyacrylamide macromolecules, macromolecular cellulose molecules, chitosan materials, modified polylactic acid materials, PET materials, inorganic silicate materials or other high molecular polymers.
[0056] In some embodiments, the methods of the present invention do not include the step of adding RNase R to the mixture comprising circular RNA to remove linear RNA.
[0057] In other embodiments, when the method of the present invention includes the step of adding RNase R to a mixture containing circular RNA to remove linear RNA, the amount of RNase R added is reduced to 30-50%, preferably 40% or 50%, of the amount of enzyme used for purification when the poly (A) tag or a functional variant thereof is not inserted.
[0058] In another aspect, the present invention provides circular RNA prepared using the preparation method of the present invention, wherein the circular RNA is substantially free of non-circular RNA molecules. In preferred embodiments, the circular RNA prepared using the preparation method of the present invention has a purity of greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9% or higher. In other preferred embodiments, the circular RNA prepared using the method of the present invention has a purity of greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9% or higher. In other preferred embodiments, the content of non-ideal nucleic acids (e.g., various intermediate complexes that are not completely circularized and still retain elements such as ribozymes) in the circular RNA prepared using the method of the present invention is no more than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1% or less of the total RNA.
[0059] In another aspect, the present invention provides a linear RNA precursor for use in the methods of the present invention.
[0060] In another aspect, the present invention provides a nucleic acid sequence capable of producing the linear RNA precursor of the present invention by transcription under conditions suitable for transcription. In a preferred embodiment, the nucleic acid sequence of the present invention further comprises, in an operably linked manner, regulatory sequences necessary for the transcription of the linear RNA precursor, including but not limited to promoters, terminators, transcription factor binding sites, untranslated regions (UTRs), enhancers, palindromes, cis-acting elements, trans-acting elements, TATA boxes, CAAT boxes, operators, and transposons.
[0061] In another aspect, the present invention provides a vector comprising the nucleic acid sequence of the present invention. In some embodiments, the vector of the present invention is a linear DNA, a plasmid, a viral nucleic acid fragment or a cell genomic DNA fragment.
[0062] In another aspect, the present invention provides an engineered cell comprising the linear RNA precursor, circular RNA, nucleic acid sequence, or vector of the present invention.
[0063] In another aspect, the present invention provides a composition comprising the linear RNA precursor, circular RNA, nucleic acid sequence, vector, or engineered cell of the present invention.
[0064] In another aspect, the present invention also provides the use of the linear RNA precursor, nucleic acid sequence, vector, or engineered cell of the present invention to prepare circular RNA.
[0065] In another aspect, the present invention also provides a use of a linear RNA precursor, circular RNA, nucleic acid sequence, vector, or engineered cell of the present invention for preparing a drug, cytotoxic agent, or immunomodulatory agent. In some embodiments, the therapeutic drug, cytotoxic agent, or immunomodulatory agent is selected from a virus, pluripotent or multipotent stem cells, iPS cells, engineered immune cells, antibodies or antibody fragments, drug-conjugated antibodies or antibody fragments, chemotherapeutic agents, immunosuppressive or modulatory agents, anti-infective drugs, anticancer agents, hypoglycemic drugs, cardiovascular and cerebrovascular disease drugs, degenerative neurological disease drugs, obesity treatment drugs, hematological disease treatment drugs, respiratory disease treatment drugs, or retroviral disease treatment drugs.
[0066] In another aspect, the present invention also provides a method for administering circular RNA, which comprises administering an effective amount of the circular RNA of the present invention to an organism in need thereof, or administering an effective amount of the circular RNA prepared using the nucleic acid sequence, vector, engineered cell or composition of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 schematically shows a nucleic acid sequence of a linear RNA precursor for preparing circular RNA according to the present application, wherein the following elements are sequentially included in the 5' to 3' direction: a 5' external homology arm, a 3' intron fragment, a 3' exon fragment, a 5' internal homology arm, a first spacer, a miRNA binding region (micro RNA binding sites), a second spacer, a 3' internal homology arm, a 5' exon fragment, a 5' intron fragment, and a 3' external homology arm, wherein the positions of two self-splicing sites are schematically shown by vertical arrows;
[0068] FIG2 shows a schematic diagram of a linear RNA precursor for preparing circular RNA according to the present application;
[0069] FIG3 shows a schematic diagram of another linear RNA precursor for preparing circular RNA according to the present application;
[0070] FIG4 shows a schematic diagram of another linear RNA precursor for preparing circular RNA according to the present application;
[0071] Figures 5A and 5B show gel electrophoresis images of some of the circularization reaction products of the linear RNA precursor according to the present application and the products after RNase R digestion. In the figure, M represents the Riboruler low molecular weight RNA reference; CK represents the comparative example 1 without poly (A) insertion;
[0072] Figure 6 shows the relative IVT yields of the circularization reaction products of the linear RNA precursors according to Preparation Examples 6 to 24 of this application. In the figure, the IVT yield of each preparation example is expressed as a multiple of the control IVT yield and subjected to one-way ANOVA statistical analysis. **** indicates P < 0.0001;
[0073] Figure 7 shows the relative IVT yields of the circularization reaction products of the linear RNA precursors according to Preparation Examples 25 to 29 of this application. In the figure, the IVT yield of each preparation example is expressed as a multiple of the IVT yield of the control and statistically analyzed by one-way ANOVA. *** indicates P < 0.001, **** indicates P < 0.0001;
[0074] 8A to 8D show HPLC graphs of the circularization reaction product of the linear RNA precursor according to the present application before and after RNase R digestion treatment. DETAILED DESCRIPTION
[0075] definition
[0076] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure will have the meanings commonly understood by those of ordinary skill in the art. The meaning and scope of the terms should be clear, however, in the event of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definitions.
[0077] As used herein, the terms "comprise" or "comprising" mean that sequences, compositions, and methods include the recited components or steps, but do not exclude other components or steps. "Consisting essentially of," when used to define sequences, compositions, and methods, should mean excluding any other components or other steps that are clearly important for the technical effect they are intended to achieve. "Consisting of" should mean excluding other components and steps not mentioned.
[0078] As used herein, the terms "linear RNA precursor" or "linear precursor" are used interchangeably and refer to an RNA precursor that is not covalently closed in a ring itself but can produce circular RNA during the cyclization process, which is generally formed by transcription of template DNA, but is not limited thereto. In this article, the linear RNA precursor comprises a complete circular RNA sequence that has not yet formed a ring, and a self-splicing sequence (such as a self-splicing site, an intron ribozyme fragment, and a homology arm, etc.) required for the cyclization of the RNA sequence. These self-splicing sequences are removed from the linear RNA precursor during the cyclization process to produce circular RNA and linear RNA fragments as by-products. In some embodiments, the linear RNA precursor of the present invention can be used to produce circular RNA by incubating in the presence of magnesium ions and guanosine nucleotides or nucleosides at a temperature where RNA cyclization occurs (e.g., between 20°C and 60°C).
[0079] In a preferred embodiment, the linear RNA precursor includes the following elements: a 5' external homology arm, a first tag, a 3' intron fragment, a 3' self-splicing site, a 3' exon fragment, a 5' internal homology arm, an insert fragment, a 3' internal homology arm, a 5' exon fragment, a 5' self-splicing site, a 5' intron fragment, a second tag, and a 3' external homology arm. In some embodiments, the 3' intron fragment, the 3' self-splicing site, and the 3' exon fragment located 5' upstream of the insert fragment, and the 5' exon fragment, the 5' self-splicing site, and the 5' intron fragment located 3' downstream of the insert fragment play a key role in the process of forming circular RNA by splicing of the linear RNA precursor. For example, by leveraging the ribozyme properties of the 3' intron fragment, the 3' self-splicing site connecting the 3' intron fragment and the 3' exon fragment can be cleaved under GTP priming. This cleavage site can further trigger the cleavage of the 5' self-splicing site connecting the 5' intron fragment and the 5' exon fragment, allowing the 5' exon fragment to connect to the 3' exon fragment at the self-splicing site to form a circular RNA. The presence of paired external and internal homology arms can improve the self-splicing and circularization efficiency in the above circularization reaction.
[0080] The nucleotides in the linear RNA precursor of the present invention may be unmodified natural nucleotides or partially modified or fully modified non-natural nucleotides. In some embodiments, the linear RNA precursor of the present invention comprises only naturally occurring nucleotides. In other embodiments, the linear RNA precursor of the present invention comprises one or more modifications that increase stability, such as 2'-O-methyl, fluoro or O-methoxyethyl conjugates, phosphorothioate backbones or 2',4'-cyclic 2'-O-ethyl modifications (Holdt et al., Front Physiol., 9: 1262 (2018); Krutzfeldt et al., Nature, 438 (7068): 685-9 (2005); Crooke et al., Cell Metab 27 (4): 714-739 (2018)), and / or one or more modifications that can reduce the innate immunogenicity of circular RNA molecules in the host, such as at least one N6-methyladenosine (m 6 A)
[0081] As used herein, the term "linear RNA fragment" refers to all RNA products derived from a linear RNA precursor after the linear RNA precursor undergoes a circularization reaction, excluding the resulting circular RNA molecule of interest. This includes, but is not limited to, self-splicing sequences on both sides that are cleaved after the circularization reaction, and various splicing intermediate sequences produced during the circularization reaction. Because not all linear RNA precursor molecules may undergo circularization, "linear RNA fragments" may also include linear RNA precursor sequences remaining in the circularization reaction products. Furthermore, the length of a "linear RNA fragment" is not necessarily shorter than the linear RNA precursor from which it is derived, as the term also encompasses linear and circular high-molecular-weight polymeric RNAs produced by the polymerization of multiple identical or different sequences described above. As used herein, the term "linear RNA fragment" has essentially the same meaning as "byproducts" or "impurities" produced by the circularization reaction.
[0082] As used herein, the terms "circular RNA" or "circRNA" are used interchangeably and refer to polyribonucleotides that form a closed circular structure through covalent bonds. It has been reported that circular RNA is a 3-5' covalently closed RNA ring that does not have a 5' end cap and a 3' end poly (A) tail. Due to the lack of free ends required for exonuclease-mediated degradation, it has the property of resisting RNase degradation and has a longer life or half-life than ordinary linear RNA products (e.g., mature mRNA). Circular RNA can be produced by splicing, and circularization mainly occurs at annotated exon boundaries using conventional splice sites (Starke et al., 2015; Szabo et al., 2015). In a preferred embodiment, such circular RNA is a single-stranded RNA molecule. Unless otherwise specified, the circular RNA molecule herein can have any structure suitable for circular RNA known in the art, but does not contain the same sequence as the first tag and the second tag present in the circular RNA precursor.
[0083] RNA nicking refers to the phenomenon of single-point breaks in an RNA chain. During the preparation of circular RNA, the presence of metal ions (particularly Mg2+) in the reaction system causes random breaks in the circular RNA, resulting in nicked RNA, which has the same molecular weight as the circular RNA and is easily digested by RNases.
[0084] As used herein, the term "intron fragment" refers to a sequence having 75% or greater similarity to a natural type I or type II intron ribozyme (ribozyme or intron ribozyme) or a major active fragment thereof, or a small ribozyme (e.g., satellite RNA, primarily including hammerhead ribozymes, hairpin ribozymes, hepatitis D virus (HDV) RNA, Varkud satellite (VS) ribozyme, and GlmS riboswitch). For example, an exemplary type I intron ribozyme can be the natural intron self-cleaving ribozyme sequence of Anabaena (which is included in the GenBank database with accession number (GenBank:AY768517), with a total sequence length of 313 bases), or a major active fragment of the sequence (e.g., the sequence described in SEQ ID NO. 2 in CN115786374A, with a total sequence length of 246 bases), or equivalents of the above ribozymes with various base substitutions or truncations. Exemplary type II ribozymes can be derived from yeast mitochondrial DNA (as described in the following references: Zimmerly et al., Mob DNA, 2015), the Ll.LtrB intron of Lactococcus lactis, the TeI3c / 4c type II intron of Thermosynechococcus elongatus (as described in the following references: Monat et al., PloS One, 2020; Costa et al., Sscience, 2016), the type II intron of Clostridium such as Clostridium tetani, the type II intron of Bacillus such as Bacillus thuringiensis, or the commercial ribozyme tool Targetron. As used herein, the term "3' intron fragment" refers to a sequence having 75% or greater similarity to the 3' region of the Anabaena ribozyme, and the term "5' intron fragment" refers to a sequence having 75% or greater similarity to the 5' region of the Anabaena ribozyme, as long as the two intron fragments can be spatially located adjacent to each other in the RNA structure to form a complex with complete ribozyme activity (as described in the following reference: Wesselhoeft et al., Nature Communication, 2018). Sequences that can serve as the "3' intron fragment" and "5' intron fragment" in natural intronic ribozymes can be determined by one skilled in the art.
[0085] As used herein, the term "exon fragment" can refer to an exon recognized and spliced by an intronic ribozyme in the intron-exon (PIE) system of a ribozyme, or a fragment of a signal sequence recognized by an intronic ribozyme. For example, a "3' exon fragment" can refer to a sequence having 75% or greater similarity to the 3'-end region of an exon recognized and cleaved by an Anabaena ribozyme, and a "5' exon fragment" can refer to a sequence having 75% or greater similarity to the 5'-end region of an exon recognized and cleaved by an Anabaena ribozyme. As long as the two exon fragments have reversed orientations and can be recognized by the ribozyme and undergo self-splicing (as described in the following reference: Wesselhoeft et al., Nature Communication, 2018). One skilled in the art can determine sequences in natural gene exons that can serve as "3' exon fragments" and "5' exon fragments."
[0086] As used herein, the term "self-splicing site" or "splice site" refers to the position between dinucleotides that, during RNA circularization, undergoes cleavage and generates free ends for circularization.
[0087] As used herein, the term "circularization region" refers to a region that is included in the formed circular RNA but not in the cleaved RNA after linear RNA is prepared into circular RNA in vitro or circular RNA is formed in vivo by splint-mediated method, permuted intron-exon method, RNA ligase-mediated method or other methods, and is irrelevant to the RNA circularization method, the structure and function of the circularization intermediate.
[0088] As used herein, the term "tag" refers to a sequence that can be used to bind to a complementary oligonucleotide (also referred to as an "affinity ligand") on an affinity chromatography matrix, thereby allowing molecules containing the tag to be specifically adsorbed to the affinity chromatography matrix and separated from molecules not containing the tag. The nucleotides comprising the tag can be naturally occurring nucleotides or artificially modified non-natural nucleotides.
[0089] It should be noted that the terms "first," "second," or similar terms used herein and in the text of this article are intended only to distinguish between two elements in a specific context and do not indicate the importance, order, etc. of the elements; the "first" and "second" elements can refer to the same or different objects / concepts. For example, in some cases, "first tag" and "second tag" can refer to the same tag (e.g., polyA) or different tags (e.g., polyA or a functional variant thereof).
[0090] As used herein, the term "poly (A) tag" has a meaning recognized and understood by those of ordinary skill in the art, for example, it refers to a sequence composed of polyadenine nucleotides, which is generally located near the 3' terminal region or the 3' end of a linear messenger RNA molecule (commonly also referred to as a poly (A) tail). It is generally believed that the poly (A) sequence at the 3' end in a linear mRNA molecule generally protects the mRNA from 3' end degradation and plays an important role in cap-dependent protein translation. Patent document 4 also reports that the poly (A) sequence in circular RNA has the effect of reducing its immunogenicity. The poly (A) tag of the present invention is generally composed of about 20 to up to about 100 adenine nucleotides, preferably 30 to 90, more preferably 40 to 80, further preferably 45 to 70, and most preferably 50 to 65 consecutive adenine nucleotides. In some embodiments, the poly(A) tag of the invention is composed of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58 , 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 consecutive adenine nucleotides.
[0091] As used herein, the term "Poly (A) tag" may also refer to a class of poly (A) comprising at least one non-A base, including but not limited to various poly (A) sequences described in PCT application PCT / CN2023 / 079037, Chinese patent application CN112805386A, and U.S. Patent US10717982B2. The above-mentioned application or patent text is incorporated herein in its entirety. Such a poly (A) sequence is also referred to herein as a "functional variant of poly (A)", meaning that the above-mentioned sequence consisting of consecutive adenine nucleotides is interrupted by individual non-A nucleotides, but still has substantially the same functional activity (such as the effect on 3' end stability and / or translation activity) as a sequence consisting of consecutive adenine nucleotides.
[0092] Specifically, the "functional variants of poly(A)" in the present invention include the following different situations:
[0093] (1) a functional variant comprising only one element a, at least one element b, and at least one element c,
[0094] (2) a functional variant comprising only element a, at least one element b, and at least one element d; or
[0095] (3) a functional variant comprising only one element a, at least one element b, and at least one element c and at least one element d,
[0096] Among them, element a is composed of multiple consecutive adenine (A) nucleotides with a length range of ≥20nt; element b is composed of multiple consecutive A nucleotides, and in some embodiments, the length range of element b is 3nt≤b<20nt; element c is composed of a non-A nucleotide, and the nucleotide is selected from thymine (T), cytosine (C), and guanine (G) nucleotides; element d is composed of any two or more consecutive nucleotides, and the nucleotides are selected from A, T, C, and G nucleotides, wherein the nucleotides at the 5' and 3' ends of element d are not A nucleotides, and element d does not contain more than 3 consecutive A nucleotides, and the length range of element d is 2nt≤d≤20nt; element e is composed of one or two consecutive A, and when it exists, it is located and can only be located at the 3' end of the Poly (A) coding sequence, and is adjacent to element d or element c.
[0097] When a Poly(A) contains two or more "element b", "element c", and "element d" at the same time, the sequences of every two elements b may be the same or different, the sequences of every two elements c may be the same or different, and the sequences of every two elements d may be the same or different, as long as they each meet the above definitions of elements a, b, c and d.
[0098] In some embodiments, element a and element b in the poly(A) tag are not adjacent, element c and element d are not adjacent, elements b are not adjacent to each other, elements c are not adjacent to each other, and elements d are not adjacent to each other.
[0099] In some embodiments, the Poly(A) tag further comprises a single element e, which consists of one or two consecutive A's and is located at the 3' end of the Poly(A) tag and adjacent to element d or element c.
[0100] In some embodiments, Poly(A) and its functional variants can be a segment of RNA or a hybrid molecule of DNA and RNA.
[0101] In some embodiments, the sequence structure of the Poly(A) tag is selected from:
[0102] Element a-element c-element b-element c-element b-element c-element b-element c-element b;
[0103] Element b-element c-element b-element c-element a-element d-element b-element c-element b-element c-element b;
[0104] Element b-element c-element b-element c-element b-element d-element a-element c;
[0105] element a-element d-element b-element c-element b-element c-element b; or
[0106] Element b-element c-element b-element c-element b-element d-element a.
[0107] As used herein, the term "binding" refers to the reversible complementary pairing between nucleotide bases through intermolecular interaction forces (e.g., hydrogen bonds), or complementary pairing through chemical bonds, thereby achieving a certain affinity between nucleic acid molecules, between nucleic acids and matrix materials carrying nucleic acids (e.g., organic polymer materials or magnetic beads), or between matrix materials carrying nucleic acids under specific conditions or reactions. For example, a polyA oligosingle-stranded RNA of a certain length (e.g., 50 nt) or a long single-stranded RNA comprising the polyA oligo can "bind" to a dextran matrix with oligo dT through complementary base pairing under certain conditions. Optionally, the hydrogen bond between the polyA and the oligo dT can be broken by changing the environmental conditions (e.g., changing the temperature, strong ionic conditions, or strong acid-base conditions), thereby making the "binding" reversible and further allowing the single-stranded RNA oligo or the long single-stranded RNA comprising the polyA oligo to detach from the dextran matrix. Furthermore, optionally, certain conditions can be used to form an irreversible chemical bond between the polyA and the oligo dT, thereby causing the RNA to irreversibly bind to the dextran matrix.
[0108] As used herein, the term "affinity chromatography medium" refers to materials containing oligonucleotides complementary to the purification tag, or affinity chromatography matrices. Ideal oligonucleotides can be designed based on specific needs. For example, if the tag is poly(A), the ideal oligonucleotide on the medium can be oligo dT. Affinity chromatography matrices can be made from stable materials that are unreactive or inert with nucleotides, with no significant repulsive or attractive intermolecular forces between their molecular surfaces and the nucleotide molecules. For example, the matrix can be dextran molecules, polyacrylamide macromolecules, macromolecular cellulose molecules, chitosan materials, modified polylactic acid materials, PET materials, inorganic silicate materials, or metal materials with a coating. Typically, the oligonucleotides and affinity chromatography matrices are cross-linked via covalent bonds. Furthermore, the oligonucleotides on the medium can be replaced with other molecules that exhibit favorable intermolecular interactions with the nucleotides. For example, if the tag is an oligonucleotide composed of at least one of U, C, and A, hypoxanthine can be cross-linked to the matrix to form a suitable affinity chromatography medium.
[0109] As used herein, the term "isolation" refers to the separation of circular RNA products from linear RNA precursors, linear RNA fragments, various intermediate complexes formed by the splicing process, and the like, thereby simply achieving the effect of purification. Techniques for purifying target polynucleotides and polypeptides are well known in the art and include, for example, ion exchange chromatography, affinity chromatography, and density-based sedimentation. Generally, a substance is purified when it is present in a sample in an amount greater than its naturally occurring amount relative to other components of the sample. In preferred embodiments herein, the circular RNA isolated by affinity chromatography or affinity adsorption has a purity of greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9% or higher, or has a purity of greater than 90%, greater than 91%, greater than 92%, greater than 93%, greater than 94%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, greater than 99%, greater than 99.5%, greater than 99.9% or higher, wherein the content of non-ideal nucleic acids (e.g., various intermediate complexes that are not completely circularized and still retain elements such as ribozymes) therein is no more than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1% or less.
[0110] A linear nucleic acid molecule is said to have a "5'-terminus" (5' end) and a "3'-terminus" (3' end) because nucleic acid phosphodiester bonds are present at the 5' carbon and 3' carbon of the sugar moiety of the substituted mononucleotide. In some examples, the 5'-terminal nucleotide of the nucleic acid molecule is a nucleotide that can form a new phosphodiester bond with the 5' carbon of its ribose sugar, and correspondingly, the 3'-terminal nucleotide is a nucleotide that can form a new phosphodiester bond with the 3' carbon of its ribose sugar.
[0111] As used herein, the term "5' upstream" refers to a position preceding a specific nucleic acid fragment in the 5' to 3' direction, but not within the fragment. The term "3' downstream" refers to a position following a specific nucleic acid fragment in the 5' to 3' direction, but not within the fragment. For example, in a 1 kb nucleic acid sequence, the "5' upstream" of a specific nucleic acid fragment located between bp 100 and 199 can include any position between bp 0 and 99, and its "3' downstream" can include any position between bp 200 and 1000. Based on a similar concept, the term "5' terminal region" refers to a number of consecutive nucleotide residues within a specific nucleic acid fragment in the 5' to 3' direction, starting from the 5' terminal nucleotide. The term "3' terminal region" refers to a number of consecutive nucleotide residues within a specific nucleic acid fragment in the 3' to 5' direction, starting from the 3' terminal nucleotide. Those skilled in the art can easily determine the 5' and 3' terminal nucleotides of a specific fragment by comparing it to the entire nucleic acid molecule. For example, in a 1 kb nucleic acid sequence, the "5' terminal region" of a specific nucleic acid fragment located between bp 100 and bp 199 may refer to bp 100 to bp 119, bp 100 to bp 114, and bp 100 to bp 109, and its "3' terminal region" may refer to bp 170 to bp 199, bp 175 to bp 199, and bp 180 to bp 199. It should be noted that the specific nucleic acid lengths and positions listed in the explanations herein are exemplary, and these terms may be supplemented, expanded, enlarged, replaced, synonymously, or equivalently interpreted in various ways without violating the concepts and spirit of the present disclosure.
[0112] As used herein, the term "structural gene" refers to a gene that can encode various polypeptides or proteins, and a "functional fragment of a structural gene" may refer to a fragment of a broken structural gene that can ultimately be expressed as a polypeptide in whole or in part. Non-limiting examples of functional fragments of structural genes include, for example, protein subunits, protein active centers, protein hybrids of proteins or non-natural catalytic groups, recombinant protein active subunits or active centers, recombinant artificial enzymes or other biological effect units mainly composed of amino acids; "non-coding RNA" may refer to RNA that does not encode polypeptide or protein products, including rRNA, tRNA, snRNA, snoRNA, microRNA, micronRNA sponge (microRNA sponge), miRNA, lncRNA, circRNA, piRNA and other RNAs with known functions, and also includes RNAs with unknown functions.
[0113] As used herein, the term "spacer" refers to a nucleic acid sequence between a functional sequence (e.g., a promoter or ribosome entry sequence) and another functional sequence (e.g., a structural gene or a fragment thereof, or a non-coding RNA or a DNA that transcribes a non-coding RNA), which is used to space the binding factors of the two functional sequences at a certain distance, or to prevent the two functional sequences from affecting each other's spatial structure, or to space the binding factors of a functional sequence at a certain distance from the spatial structure of the other functional sequence itself, which does not encode a polypeptide, protein or non-coding RNA with a biological effect, or the transcription product or translation product at least does not degrade the biological activity of the polypeptide, protein or non-coding RNA. For example, in some embodiments, the spacer can be a scrambled, mononucleotide repeated, or polynucleotide repeated sequence of a specific length (e.g., 9nt).
[0114] As used herein, the term "RNase R" (Ribonuclease R, RNase R) is an exoribonuclease derived from the Escherichia coli RNR superfamily. It can cleave and degrade linear RNA molecules from the 3' to 5' direction, but does not substantially digest circular RNA, lariat structures, or doublet RNA molecules with 3' overhangs lacking 7 nt. Specific examples of RNase R can be found in the following reference: Cheng ZF, Biol Chem, 2002. Those skilled in the art will appreciate that the term "RNase R" may include polypeptides, hybrid enzymes, and polypeptide analogs that have the same enzymatic activity as those produced artificially or recombinantly based on RNase R.
[0115] As used herein, the term "immunogenicity" refers to the potential to induce an immune response to a substance. When an organism's immune system or a certain type of immune cell is exposed to an immunogenic substance, an immune response can be induced.
[0116] As used herein, the term "cyclization efficiency" refers to a measure of the resulting circular polyribonucleotide compared to its linear starting material.
[0117] As used herein, the term "translational efficiency" refers to the rate or amount of protein or peptide produced from ribonucleotide transcripts. In some embodiments, translation efficiency can be expressed as the amount of protein or peptide produced per a given amount of transcript encoding the protein or peptide.
[0118] The term "nucleotide" refers to a ribonucleotide, a deoxyribonucleotide, a modified form thereof, or an analog thereof. Nucleotides include substances including purines (e.g., adenine, hypoxanthine, guanine, and derivatives and analogs thereof) and pyrimidines (e.g., cytosine, uracil, thymine, and derivatives and analogs thereof). Nucleotide analogs include nucleotides having modified residues in the chemical structure of bases, sugars, and / or phosphates, including but not limited to, 5'-position pyrimidine modifications, 8'-position purine modifications, modifications at the cytosine exocyclic amine, and substitutions with 5-bromo-uracil; and 2'-position sugar modifications, including but not limited to sugar-modified ribonucleotides, wherein 2'-OH is substituted with a group such as H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN, wherein R is an alkyl moiety as defined herein. Nucleotide analogs are also intended to include nucleotides having bases such as inosine, quercetin, and xanthine; sugars such as 2'-methylribose; and non-natural phosphodiester linkages such as methylphosphonate, phosphorothioate, and peptide linkages. Nucleotide analogs include 5-methoxyuridine, 1-methylpseudouridine, and 6-methyladenosine.
[0119] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein to describe a polymer of any length (e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, or up to about 10,000 or more bases) composed of nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and that can be produced enzymatically or synthetically (e.g., as described in U.S. Pat. No. 5,948,902 and references cited therein) that can hybridize with naturally occurring nucleic acids in a sequence-specific manner similar to two naturally occurring nucleic acids, for example, can participate in Watson-Crick base pairing interactions. Naturally occurring nucleic acids are composed of nucleotides including guanine, cytosine, adenine, thymine, and uracil (G, C, A, T, and U, respectively).
[0120] As used herein, the terms "ribonucleic acid" and "RNA" refer to a polymer composed of ribonucleotides.
[0121] As used herein, the terms "deoxyribonucleic acid" and "DNA" refer to a polymer composed of deoxyribonucleotides.
[0122] As used herein, two "homology arms" or "homologous regions" are complementary to each other or are complementary when they have a sufficient level of sequence identity to each other's reverse complement sequences to serve as substrates for hybridization reactions. As used herein, a polynucleotide sequence has "homology" when it is identical or shares sequence identity with a reverse complement sequence or "complementary" sequence. The percentage of sequence identity between a homologous region and the reverse complement sequence of the corresponding homologous region can be any percentage of sequence identity that allows hybridization to occur. In some embodiments, an internal duplex-forming region of a polynucleotide of the present invention is capable of forming a duplex with another internal duplex-forming region and does not form a duplex with an external duplex-forming region.
[0123] "Transcription" refers to the formation or synthesis of an RNA molecule by an RNA polymerase using a DNA molecule as a template. The present invention is not limited to the RNA polymerase used for transcription. For example, in some embodiments, a T7-type RNA polymerase can be used. "Translation" refers to the formation of a polypeptide molecule by ribosomes based on an RNA template.
[0124] It should be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting. As used in this specification and the appended claims, unless the content clearly indicates otherwise, the singular forms "a / an" and "the" include plural referents. Thus, for example, reference to "a cell" includes a combination of two or more cells, or the entire culture of cells; reference to "polynucleotides" actually includes many copies of the polynucleotides. Unless expressly provided or obvious from the context, as used herein, the term "or" is understood to be inclusive. Unless otherwise defined herein or and in the remainder of the specification below, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention belongs.
[0125] Unless expressly provided or obvious from the context, otherwise as used herein, the term "about" should be understood to be within the normal tolerance range in the field, for example, within 2 standard deviations of the mean. "About" can be understood to be within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02% or 0.01% of the value. Unless otherwise obvious from the context, all numerical values provided herein are modified by the term "about".
[0126] As used herein, the term "encoding" broadly refers to any process in which information in a polymeric macromolecule is used to direct the production of a second molecule that is different from the first molecule. The second molecule may have a chemical structure that is different in chemical nature from the first molecule.
[0127] "In combination" or "co-administration" refers to administering a therapeutic agent provided herein in conjunction with one or more additional therapeutic agents sufficiently close in time such that the therapeutic agent provided herein enhances the effect of the one or more additional therapeutic agents, and vice versa.
[0128] As used herein, the terms "treat" and "prevent" and words derived therefrom do not necessarily mean 100% or complete treatment or prevention. Rather, there are varying degrees of treatment or prevention that one of ordinary skill in the art recognizes as having potential benefit or therapeutic effect. The treatment or prevention provided by the methods disclosed herein can include treating or preventing one or more conditions or symptoms of a disease. Furthermore, for purposes herein, "prevention" can include delaying the onset of a disease or its symptoms or conditions.
[0129] As used herein, the term "expressed sequence" may refer to a nucleic acid sequence encoding a product such as a peptide or polypeptide, a regulatory nucleic acid, or a non-coding nucleic acid. An exemplary expressed sequence encoding a peptide or polypeptide may comprise a plurality of nucleotide triplets, each of which may encode an amino acid and is referred to as a "codon."
[0130] The term "antibody" (Ab) includes, but is not limited to, glycoprotein immunoglobulins that specifically bind to an antigen. Generally speaking, an antibody may comprise at least two heavy (H) chains and two light (L) chains interconnected by disulfide bonds, or an antigen-binding molecule thereof. Each H chain comprises a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant region. The heavy chain constant region comprises three constant domains, CH1, CH2, and CH3. Each light chain comprises a light chain variable region (abbreviated herein as VL) and a light chain constant region. The light chain constant region comprises one constant domain, CL. The VH and VL regions can be further subdivided into hypervariable regions, called complementarity determining regions (CDRs), interspersed with more conserved regions, called framework regions (FRs). Each VH and VL comprises three CDRs and four FRs, arranged in the following order from amino terminus to carboxyl terminus: FR1, CDR1, FR2, CDR2, FR3, CDR3, and FR4. The variable regions of the heavy and light chains contain binding domains that interact with the antigen. The constant regions of the antibodies may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system (eg, effector cells) and the first component of the classical complement system. Antibodies may include, for example, monoclonal antibodies, recombinantly produced antibodies, monospecific antibodies, multispecific antibodies (including bispecific antibodies), human antibodies, engineered antibodies, humanized antibodies, chimeric antibodies, immunoglobulins, synthetic antibodies, tetrameric antibodies comprising two heavy and two light chain molecules, antibody light chain monomers, antibody heavy chain monomers, antibody light chain dimers, antibody heavy chain dimers, antibody light chain-antibody heavy chain pairs, intrabodies, antibody fusions (sometimes referred to herein as "antibody conjugates"), heteroconjugate antibodies, single domain antibodies, monovalent antibodies, single chain antibodies or single chain Fv (scFv), camelized antibodies, affibodies, Fab fragments, F(ab')2 fragments, disulfide-linked Fv (sdFv), anti-idiotypic (anti-id) antibodies (including, for example, anti-anti-Id antibodies), miniantibodies, domain antibodies, synthetic antibodies (sometimes referred to herein as "antibody mimetics"), and antigen-binding fragments of any of the foregoing. In some embodiments, the antibodies described herein refer to polyclonal antibody populations.
[0131] Immunoglobulin can be derived from any known isotype, including but not limited to IgA, secretory IgA, IgG and IgM. IgG subclasses are also well known to those skilled in the art, including but not limited to human IgG1, IgG2, IgG3 and IgG4. "Isotype" refers to the Ab class or subclass (e.g., IgM or IgG1) encoded by the heavy chain constant region gene. For example, the term "antibody" includes naturally occurring and non-naturally occurring antibodies; monoclonal and polyclonal antibodies; chimeric and humanized antibodies; human or non-human antibodies; fully synthetic antibodies; and single-chain antibodies. Non-human antibodies can be humanized by recombinant methods to reduce their immunogenicity in humans. In the absence of explicit instructions and unless the context indicates otherwise, the term "antibody" also includes the antigen-binding fragment or antigen-binding portion of any of the above-mentioned immunoglobulins, and includes monovalent and divalent fragments or portions, and single-chain antibodies.
[0132] "Antigen binding molecule", "antigen binding portion" or "antibody fragment" refers to any molecule that comprises the antigen binding portion (e.g., CDR) of the antibody from which the molecule is derived. Antigen binding molecules may include antigen complementary determining regions (CDRs). Examples of antibody fragments include, but are not limited to, Fab, Fab', F(ab')2 and Fv fragments, dAbs, linear antibodies, scFv antibodies, and multispecific antibodies formed by antigen binding molecules. Peptibodies (i.e., Fc fusion molecules comprising peptide binding domains) are another example of suitable antigen binding molecules. In some embodiments, the antigen binding molecules bind to antigens on tumor cells. In some embodiments, the antigen binding molecules bind to antigens on cells involved in hyperproliferative diseases or bind to viral or bacterial antigens. In some embodiments, the antigen binding molecules bind to BCMA. In other embodiments, the antigen binding molecules are antibody fragments that specifically bind to an antigen, including one or more complementary determining regions (CDRs) thereof. In other embodiments, the antigen binding molecules are single-chain variable fragments (scFv).
[0133] "Cancer" or "cancer" refers to a broad range of diseases characterized by the uncontrolled growth of abnormal cells in the body. Unregulated cell division and growth lead to the formation of malignant tumors that invade adjacent tissues and may also metastasize to distant parts of the body via the lymphatic system or bloodstream. "Cancer" or "cancerous tissue" may include tumors. Examples of cancers treatable by the methods disclosed herein include, but are not limited to, cancers of the immune system, including lymphomas, leukemias, myelomas, and other white blood cell malignancies. In some embodiments, the methods disclosed herein can be used to reduce the size of tumors originating from, for example, bone cancer, pancreatic cancer, skin cancer, head and neck cancer, cutaneous or intraocular malignant melanoma, uterine cancer, ovarian cancer, rectal cancer, anal cancer, stomach cancer, testicular cancer, uterine cancer, multiple myeloma, Hodgkin's disease, non-Hodgkin's lymphoma (NHL), primary mediastinal large B-cell lymphoma (PMBC), diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL), transformed follicular lymphoma, splenic marginal zone lymphoma (SMZL), esophageal cancer, small intestine cancer, endocrine system cancer, thyroid cancer, parathyroid cancer, kidney cancer, or Cancer of the upper gland, cancer of the urethra, cancer of the penis, chronic or acute leukemias, acute myeloid leukemia, chronic myeloid leukemia, acute lymphoblastic leukemia (ALL) (including non-T cell ALL), chronic lymphocytic leukemia (CLL), solid tumors in children, lymphocytic lymphomas, bladder cancer, kidney cancer or ureteral cancer, central nervous system (CNS) tumors, primary CNS lymphomas, tumor angiogenesis, spinal axis tumors, brain stem gliomas, pituitary adenomas, epidermoid carcinomas, squamous cell carcinomas, T cell lymphomas, environmentally induced cancers including those induced by asbestos, other B cell malignancies, and combinations of the aforementioned cancers. In some embodiments, the methods disclosed herein can be used to reduce the size of tumors originating from, for example, sarcomas and carcinomas, fibrosarcomas, myxosarcoma, liposarcoma, chondrosarcoma, osteogenic sarcoma, Kaposi's sarcoma, soft tissue sarcomas and other sarcomas, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, hepatocellular carcinoma, lung cancer, colorectal cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinomas (e.g., pancreatic, colon, ovarian, lung, breast, stomach, prostate, cervix), or esophageal adenocarcinoma), sweat gland cancer, sebaceous gland cancer, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, bile duct cancer, choriocarcinoma, Wilms' tumor, cervical cancer, testicular tumors, bladder cancer, fallopian tube cancer, endometrial cancer, cervical cancer, vaginal cancer, vulvar cancer, renal pelvis cancer, CNS tumors (such as glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, melanoma, neuroblastoma, and retinoblastoma). Specific cancers may respond to chemotherapy or radiation therapy, or the cancer may be refractory.Refractory cancer refers to cancer that is not amenable to surgical intervention and that does not initially respond to chemotherapy or radiation therapy, or that becomes unresponsive over time.
[0134] As used herein, the term "immune cell" or "lymphocyte" includes natural killer (NK) cells, T cells or B cells. NK cells are a type of cytotoxic (cytotoxic) lymphocytes that represent the main components of the innate immune system. NK cells repel tumors and virus-infected cells. It works through the process of apoptosis or programmed cell death. They are called "natural killers" because they do not need to be activated to kill cells. T cells play a major role in cell-mediated immunity (without antibody involvement).
[0135] The term "genetic engineering" or "engineering" refers to a method of modifying the genome of a cell, including but not limited to deleting a coding or non-coding region or a portion thereof or inserting a coding region or a portion thereof. In some embodiments, the modified cell is a lymphocyte, such as a T cell, which can be obtained from a patient or a donor. The cell can be modified to express an exogenous construct incorporated into the cell genome, such as a chimeric antigen receptor (CAR) or a T cell receptor (TCR).
[0136] As used herein, the expression "sequence identity" or, for example, comprising "a sequence that is 50% identical to" refers to the extent to which sequences are identical on a nucleotide-by-nucleotide or amino acid-by-amino acid basis over the comparison window. Thus, "percentage of sequence identity" can be calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions at which the same nucleic acid base (e.g., A, T, C, G, I) or the same amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity. Included are nucleotides and polypeptides having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any reference sequence described herein, typically wherein the polypeptide variant retains at least one biological activity of the reference polypeptide.
[0137] As used herein, the term "treat" refers to both therapeutic treatment and prophylactic or preventative or preventative measures, wherein the object is to prevent or slow (mitigate) an undesirable pathological change or condition. For purposes of the present invention, beneficial or desired clinical results include, but are not limited to, alleviation of symptoms, reduction in disease severity, delay or slowing of disease progression, improvement or alleviation of the disease state, and remission (whether partial or complete), whether detectable or undetectable.
[0138] As used herein, the term "therapeutically effective amount" refers to an amount of a compound of the present invention that can: (i) treat or prevent a disease or condition described herein, (ii) ameliorate or eliminate one or more diseases or conditions described herein, or (iii) prevent or delay the onset of one or more symptoms of a disease or condition described herein.
[0139] As used herein, the term "non-coding RNA" or "ncRNA" primarily encompasses microRNA (miRNA), small interfering RNA (siRNA), PIWI-interacting RNA (piRNA), tRNA-derived small RNA (tsRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), long non-coding RNA (lncRNA), circular RNA (circRNA), and pseudogenes. miRNAs are short, approximately 22-23 nucleotides long. Their coding genes are transcribed by RNA polymerase II and regulate mRNA expression by binding to the 3' untranslated region (3'UTR) of mRNA. Over 60% of coding genes are potential regulatory targets of miRNAs. lncRNAs are non-coding RNAs longer than 200 nucleotides, and their biogenesis is similar to that of mRNA. LncRNAs play important roles in a variety of biological processes, including cell cycle regulation, chromatin modification, and mRNA translation. CircRNAs, a type of lncRNA, are primarily produced from exon or intron sequences. They are single-stranded circular RNA molecules that can act as competing endogenous RNAs to bind to miRNAs, regulating transcription or influencing the expression of parental genes. piRNAs are small RNAs, approximately 21 to 35 nucleotides in length, that are processed from long single-stranded transcripts. The genomic loci of these transcripts are clustered throughout the genome and transcribed by RNA polymerase II. Approximately 20,000 piRNAs are present in the human genome, primarily expressed in gonadal cells. tsRNAs, derived from the cleavage of tRNA by nucleases, are typically 18 to 40 nucleotides in length and play important roles in regulating translation, maintaining mRNA stability, gene silencing, and reverse transcription.
[0140] As used herein, the term "genome" or "genomic DNA" refers to the heritable information of a host organism. The genomic DNA includes all of the genetic material of a cell or organism, including nuclear DNA (chromosomal DNA), extrachromosomal DNA, and organelle (e.g., mitochondrial) DNA. Preferably, the term "genome" or genomic "DNA" refers to the chromosomal DNA of the nucleus.
[0141] As used herein, where the term "recombinant" is used to describe an organism or cell (eg, a microorganism), it is used to express that the organism or cell comprises at least one "transgene," "transgenic," or "recombinant" polynucleotide, generally described later.
[0142] As used herein, a polynucleotide that is "exogenous" with respect to an individual organism is a polynucleotide that has been introduced into that organism by any means other than sexual hybridization.
[0143] As used herein, the term "promoter" or "RNA enzyme binding site" is a polynucleotide region that initiates transcription of a coding sequence. The promoter is located near the transcription start site of a gene, on the same strand and upstream of the DNA (towards the 5' region of the sense strand). Some promoters are constitutive because they are active in all cases in the cell, while other promoters are regulated to become active in response to a specific stimulus (e.g., inducible promoters). As used herein, the term "promoter activity" and its grammatical equivalents refer to the degree of expression of a nucleotide sequence operably connected to a promoter whose activity is being measured. Promoter activity can be measured directly by determining the amount of RNA transcript produced, for example, by Northern blot analysis, or indirectly by determining the amount of a product encoded by a connected nucleic acid sequence, such as a reporter nucleic acid sequence connected to the promoter.
[0144] As used herein, the term "plasmid" refers to an extrachromosomal element that often carries genes that are not part of the core metabolic machinery of the cell and is typically in the form of a circular double-stranded DNA molecule. These elements can be autonomously replicating sequences, genome-integrating sequences, phage or nucleotide sequences of any origin, linear, circular or supercoiled single-stranded or double-stranded DNA or RNA. Typically, a plasmid contains a functional origin of replication in a host cell (e.g., E. coli) and a selectable marker for detecting host cells containing the plasmid.
[0145] As used herein, the term "transformation" refers to the introduction of one or more exogenous polynucleotides into a host cell by using physical or chemical methods.
[0146] As used herein, the term "expression" is the process of revealing the information encoded within a gene. If a gene encodes a protein, expression includes transcribing DNA into mRNA, processing the mRNA (if necessary) into a mature mRNA product, and translating the mature mRNA into protein.
[0147] The present invention is further described below by way of specific examples, but this is not intended to limit the present invention. Those skilled in the art may make various modifications or adjustments based on the teachings of the present invention without departing from the spirit and scope of the present invention.
[0148] Example
[0149] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.
[0150] The experimental methods in the following examples, unless otherwise specified, are all conventional methods and are carried out in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions. The materials, reagents, instruments, etc. used in the following examples, unless otherwise specified, can all be obtained from commercial channels. The quantitative tests in the following examples, unless otherwise specified, are the average values of three repeated experiments. In the following examples, unless otherwise specified, the nucleotide sequences in the sequence listing are written from left to right in the order from 5' to 3' end, and the amino acid sequences are written from left to right in the order from amino terminus to carboxyl terminus.
[0151] Reagents, instruments, and cell lines used
[0152] Unless otherwise specified, the reagents, instruments, and cell lines used in the present invention are commercially available.
[0153] Important reagents are listed below:
[0154] BspQ I (catalog number ON-123, Shanghai Zhaowei Technology Development Co., Ltd.);
[0155] High Yield T7 RNA Synthesis Kit (Cat. No. ON-040, Shanghai Zhaowei Technology Development Co., Ltd.);
[0156] Monarch RNA Cleanup Kit (Cat. No. T2040L, NEB);
[0157] RNase R (Cat. No. E224, Nearshore Protein);
[0158] Oligo d(T)25 magnetic beads (Cat. No. S1419S, NEB).
[0159] The important instruments are listed below:
[0160] Agilent 5200 Fragment Analyzer System (Agilent, US),
[0161] Agarose gel electrophoresis system (model DYY-6C, Beijing Liuyi Biotechnology Co., Ltd.),
[0162] Nanodrop spectrophotometer (model Nanodrop ONE c, Thermo Fisher Scientific),
[0163] Centrifuge (model Sorvall Legend Micro 17R, Thermo Fisher Scientific)),
[0164] Metal bath (Model HWS-12, Shanghai Yiheng Scientific Instrument Co., Ltd.).
[0165] Example 1: Preparation of linear RNA precursors
[0166] Suzhou Jinweizhi Biotechnology Co., Ltd. was commissioned to synthesize a double-stranded DNA template for the preparation of linear RNA precursors. This double-stranded template was incorporated into the pUC57 vector (company name, product number) through gene synthesis, placing it downstream of the 3' end of the T7 promoter sequence (TAATACGACTCACTATA), obtaining the template plasmid pUC-PIE-miR21 for the synthesis of linear RNA precursors.
[0167] In vitro transcription was performed using a 20 μl system (1X reaction buffer, 7.5 mM each of ATP, GTP, CTP, and UTP, 1 μl of T7 RNA polymerase mix, and 1 μg of the linear template plasmid pUC-PIE-miR21) and incubated at 37°C for 4 hours to produce the linear RNA precursor represented by SEQ ID NO: 11. This precursor contains the elements listed in Table 1 in a 5' to 3' orientation.
[0168] The underlines in SEQ ID NO:11 indicate the 5' outer homology arm near the 5' end and the 3' outer homology arm near the 3' end, respectively. For ease of description, the first cytosine (C) in the 5' outer homology arm is numbered as position 1, and the adjacent guanine (G) is the transcription start point of the T7 promoter (TAATACGACTCACTATA) and is designated as position 0. The last nucleotide in SEQ ID NO:11 from 5' to 3' is designated as position 807, and the adjacent 3' nucleotide is designated as position 808.
[0169] In SEQ ID NO: 11, the symbols "#" and "*" respectively indicate the chemical bond positions targeted by the two-step transesterification reaction of the linear precursor during circularization, and the bold "UG" and "AA" are the dinucleotides that determine the self-splicing sites in the 3' intron fragment and the 5' intron fragment, respectively. For this specific sequence, the nucleotide residues 1 to 151 located 5' upstream of "#" constitute its 5' end arm, and the nucleotide residues 673 to 807 located 3' downstream of "*" constitute its 3' end arm; the circular RNA formed thereby is cyclized by the sequence represented by the nucleotides 152 to 672 located between "#" and "*", that is, the cyclized region specifically includes: 3' exon fragment (positions 152 to 202), 5' internal homology arm (positions 203 to 227), insert fragment (positions 228 to 635), 3' internal homology arm (positions 636 to 657), and 5' exon fragment (positions 658 to 672).
[0170] FIG1 schematically shows the linear RNA precursor, with the name and start and end position numbers of each element indicated above the corresponding position.
[0171] Table 1: PIE system sequence
[0172] Example 2: Linear RNA Precursors with Purification Tags in Single or Double End Arms and Their IVT Yield and Circularization Efficiency Tests
[0173] Using the same method as in Example 1, linear RNA precursors having the nucleotide sequences set forth in SEQ ID NO:26 to SEQ ID NO:49 were prepared as Preparation Examples 1 to 24, respectively. As shown in Table 2 below, compared to the nucleotide sequence set forth in SEQ ID NO:11, the linear RNA precursors of Preparation Examples 1 to 24 each contained a 50 nt poly(A) purification tag at a different position in SEQ ID NO:11. Sequence ID NO:11 prepared in Example 1 was used as a comparative example (CK).
[0174] The insertion position is explained as follows: when the insertion position is 1, poly A is located at the 3' end of the first nucleotide residue of the linear RNA precursor and is directly connected to the first nucleotide residue.
[0175] Table 2:
[0176] Using known methods, Preparation Examples 1 to 5, 7, 11, 24 and Comparative Example 1 were subjected to cyclization reactions under the same conditions. For example, after incubating the in vitro transcription systems of Examples 1 and 2 at 37°C for 4 hours, 2 mm GTP was added and the cells were transferred to 55°C for incubation for 15 minutes. The incubated mixture was then purified using an RNA Cleanup Kit according to the manufacturer's instructions to obtain a cyclization reaction product (-) before RNase R treatment. The RNA content of the cyclization reaction product was determined using the Nanodrop method. 20 U RNase R and 2 μg RNA were added to a 100 μl reaction system, incubated at 37°C for 15 minutes, and a cyclization reaction product (+) after RNase R treatment was obtained. The two reaction products were detected by 1.0% agarose gel under the same conditions. The electrophoresis results of the product with a poly (A) tag on one end arm are shown in Figure 5A, and the electrophoresis results of the product with a poly (A) tag on both ends are shown in Figure 5B.
[0177] RNase R is a 3'-5' exonuclease that can only degrade linear RNA, not circular RNA. Figure 5 shows that after RNase R treatment, distinct bands corresponding to circular RNA were observed in Preparation Examples 1-5, 7, 11, 24 and Comparative Example 1 (CK), indicating that the insertion of poly(A) did not affect the cyclization of these linear RNA precursors. However, in the lanes corresponding to Preparation Examples 1-5, 7, 11, and 24, no distinct secondary bands were observed, either before (-) or after (+) RNase R treatment, except for the primary band. In contrast, in the lane corresponding to Comparative Example 1, used as a control, two distinct, separate bands were observed before (-) RNase R treatment, with the band with the highest mobility disappearing after (+) RNase R treatment, indicating that these bands primarily correspond to nicked RNA impurities generated by random fragmentation of the circular RNA.
[0178] In summary, this example demonstrates that, without being constrained by any theoretical mechanism, inserting purification tags into the two end arms of a linear RNA precursor (particularly within the 5' external homology arm, 5' intron, or 3' external homology arm) generally does not significantly interfere with the cyclization reaction. Instead, it surprisingly improves the efficiency of the cyclization reaction of the linear RNA precursor, significantly increasing the proportion of the desired fully cyclized product in the reaction products while significantly reducing or even eliminating undesirable byproducts such as nicked RNA. This significantly improves the purity of the crude cyclization reaction product, reduces or even eliminates the dependence of the circular RNA purification method on the RNase R enzyme treatment step, and ultimately greatly simplifies the subsequent purification process while not affecting or even significantly improving the quality of the cyclized product.
[0179] Example 3: Linear RNA Precursors with Purification Tags in Double-End Arms and Their IVT Yield and Circularization Efficiency Tests
[0180] The same method as in Example 2 was used to prepare a linear RNA precursor having a nucleotide sequence as shown in SEQ ID NO: 50 as Comparative Example 2. The sequences of the elements in the linear RNA precursor were the same as in Table 1, except for the 5' external homology arm and the 3' external homology arm shown in Table 3.
[0181] Table 3
[0182] Using the same method, linear RNA precursors having the nucleotide sequences shown in SEQ ID NOs: 51 to 55 were prepared as Preparation Examples 25 to 29, respectively. As shown in Table 4 below, compared to Comparative Example 2, the linear RNA precursors of Preparation Examples 25 to 29 each contained a 50 nt poly(A) as a purification tag at a different position in SEQ ID NO: 50.
[0183] Table 4:
[0184] Using the same method as Example 2, the relative yields of Preparation Examples 25 to 29 compared with Comparative Example 2 (CK) under the same conditions were measured, and the results are shown in FIG7 .
[0185] This example still demonstrates that linear RNA precursors with purification tags are capable of producing the desired cyclized products through in vitro cyclization reactions. Changing the specific sequence of the linear RNA precursor, including changing the end arm sequence (e.g., changing the 5' and / or 3' external homology arms), does not affect the in vitro cyclization of the precursor molecule with the purification tag. However, when the purification tag is inserted at certain positions in the linear RNA precursor, the IVT yield will be significantly reduced, specifically at the residue position immediately upstream of the 5' end of the 5' external homology arm (position 0) and at the first residue position at the 5' end of the 5' external homology arm (position 1). However, a reduction in the amount of non-target products produced can be observed in all cases.
[0186] Example 4: Other exemplary linear RNA precursors
[0187] A linear RNA precursor having the element structure shown in Figures 2 to 4 was prepared using the same method as in Example 1. The external homology arms, intronic fragments, coding region fragments, translation initiation elements, exonic fragments, internal homology arms, spacer sequences, and insertion sequences shown in the figures can all be elements known in the art to have corresponding functions, or elements that can be reasonably inferred to have corresponding functions, including but not limited to the various functional elements described in CN112399860A and CN115404240A. The aforementioned applications or patents are hereby incorporated herein in their entirety.
[0188] Example 5: Purification efficiency of affinity chromatography cyclization reaction products
[0189] The affinity chromatography purification efficiency of linear RNA precursors with poly(A) inserted into both end arms was measured relative to a control without poly(A).
[0190] The final reaction solution containing 5 μg of RNA was taken from each of the cyclization reaction products of Preparation Example 24 and Comparative Example 1 for purification. Without adding RNase R, the final solution was incubated with 200 μl of Oligo d(T)25 magnetic beads for 20 minutes to allow the magnetic beads to fully contact and adsorb the RNA with the poly(A) tag. After 20 minutes, the sample was placed on a magnetic stand and allowed to stand for 10 seconds to separate the magnetic beads from the supernatant. The magnetic beads were removed, and the supernatant was recovered using an RNA Cleanup Kit according to the manufacturer's instructions. The purity of the cyclization reaction product (-) before RNase R treatment and the cyclization reaction product (+) after RNase R treatment were respectively detected using an Agilent 5200 fragment analyzer under default settings. The results are shown in Figures 8A to 8D.
[0191] Figure 8 shows that before incubation with Oligo d(T)25 magnetic beads, the circRNA purities in the samples from Comparative Example 1 and Preparation Example 24 were 81.3% and 92.7%, respectively. After incubation with the magnetic beads, the purity of the circRNA from Comparative Example 1 was measured to be 78.1%, while the purity of Preparation Example 24 increased to 100%. These results demonstrate that the insertion of poly(A) into the flanking arms of the linear RNA precursor allows affinity chromatography (e.g., Oligo d(T)25 magnetic beads) to remove impurities present in the circularization reaction system with the target circRNA, yielding highly pure circRNA.
[0192] Notably, even before incubation with Oligo d(T)25 magnetic beads, a significant difference in the purity of the cyclization products of Preparation Example 24 and Comparative Example 1 was observed (92.7% vs 81.3%). This result is consistent with the electrophoresis results observed in Example 3, indicating that poly(A) inserted into the terminal arms can directly improve the purity of the cyclization product, independent of the affinity purification mechanism.
[0193] The sequences used in the above examples of the present application are shown in the sequence listing. It should be understood that these sequences are merely exemplary sequences of the present application's embodiments and are not intended to limit the present application's embodiments. Although DNA sequences are indicated in the electronic sequence listing, the nucleotide sequences in the present application's sequence listing may represent either DNA sequences or RNA sequences. When representing RNA sequences, "T" represents uridine.
Claims
1. A method for preparing circular RNA, the method comprising: Step A: generating a linear RNA precursor, wherein the linear precursor comprises a 5' end arm, a 3' self-splicing site, a circularization region, a 5' self-splicing site and a 3' end arm in a manner operably connected to each other in sequence from 5' to 3' direction, wherein a first tag is inserted into the 5' end arm and a second tag is inserted into the 3' end arm; Step B: placing the linear RNA precursor under conditions suitable for self-splicing of the 3' self-splicing site and the 5' self-splicing site, to obtain a mixture comprising a linear RNA fragment carrying the first tag and / or the second tag and a circular RNA obtained by cyclization of the cyclization region; Step C: contacting the mixture with an affinity chromatography medium capable of simultaneously binding the first tag and the second tag for a period of time sufficient to allow the affinity chromatography medium to bind to the linear RNA fragment containing at least one of the tags; and Step D: separating the affinity chromatography medium and the mixture after contacting the affinity chromatography medium, collecting the supernatant, and obtaining the circular RNA. Wherein, the first tag and / or the second tag are independently selected from a poly(A) tag consisting of 20 to 100, preferably 30 to 90, more preferably 40 to 80, further preferably 45 to 70, and most preferably 50 to 65 consecutive adenine nucleotides or a functional variant thereof, and the first tag and the second tag are not present in the cyclization region and the circular RNA.
2. The method according to claim 1, wherein: The functional variant of the poly(A) tag is to insert one or more non-A bases into the poly(A) tag, preferably insert 1 to 20 non-A bases, more preferably insert 1 to 10 non-A bases.
3. A method according to any preceding claim, wherein: The functional variant of the poly(A) tag comprises (1) a single element a, at least one element b, and at least one element c, (2) only one element a, at least one element b, and at least one element d; or (3) a single element a, at least one element b, at least one element c and at least one element d, wherein the element a is composed of more than 20 consecutive adenine nucleotides, the element b is composed of more than 3 and less than 20 consecutive adenine nucleotides, the element c is composed of a nucleotide selected from uracil nucleotides, cytosine nucleotides, and guanine nucleotides, the element d is composed of more than 2 and less than 20 nucleotides, the nucleotides are arbitrarily selected from adenine nucleotides, uracil nucleotides, cytosine nucleotides, and guanine nucleotides, and the element d does not contain more than 3 consecutive adenine nucleotides, and the 5' and 3' terminal nucleotides are not adenine nucleotides, Wherein, when the Poly(A) tag contains two or more of the element b, the element c or the element d at the same time, the sequences of every two elements b may be the same or different, the sequences of every two elements c may be the same or different, and the sequences of every two elements d may be the same or different. Furthermore, the element a and the element b, the element c and the element d, the elements b, the elements c, and the elements d are not adjacent to each other.
4. The method according to any of the preceding claims, wherein the element a consists of more than 20 and less than 80 consecutive adenine nucleotides, preferably consists of 30 to 70, 35 to 65, 40 to 60, or 45 to 55 consecutive adenine nucleotides, and more preferably consists of 60 consecutive adenine nucleotides.
5. A method according to any preceding claim, wherein the element b consists of 3 to 10, 10 to 19, 12 to 15, 14 to 17, or 16 to 19, preferably 19 consecutive adenine nucleotides.
6. The method according to any one of the preceding claims, wherein the number of the elements b is 2 to 10, preferably 2 to 5, and more preferably 3.
7. A method according to any preceding claim, wherein the element c is a guanine nucleotide.
8. The method according to any of the preceding claims, wherein the number of the elements c is 2 to 10, 3 to 8, 4 to 6, or 2 to 5, preferably 2.
9. The method according to any of the preceding claims, wherein the element d consists of 3 to 18, 5 to 16, 4 to 10, or 6 to 12 nucleotides, preferably consists of 6 nucleotides.
10. The method according to any of the preceding claims, wherein the element d is selected from any of GAUAUC, GUAUAC, GAAUCU, GCAUAUGACU or GAUAUCGUAUAC.
11. The method according to any one of the preceding claims, wherein the number of the elements d is 0 to 5, preferably 1 to 3, more preferably 1.
12. The method according to any of the preceding claims, wherein when element c and element d are present at the same time, the total number of the elements c and d is 2 to 15, preferably 3 to 5, and more preferably 3.
13. A method according to any preceding claim, wherein: The functional variant of the poly(A) tag has any one structure selected from the following structures: component a-component c-component b-component c-component b-component c-component b-component c-component b, Element b-element c-element b-element c-element a-element d-element b-element c-element b-element c-element b, Element b-element c-element b-element c-element b-element d-element a-element c, element a-element d-element b-element c-element b-element c-element b, or Element b-element c-element b-element c-element b-element d-element a.
14. A method according to any preceding claim, wherein: The 5' end arm comprises a 5' external homology arm and a 3' intron fragment in the 5' to 3' direction, the first tag is inserted into the 5' external homology arm, or inserted into the 5' terminal region of the 3' intron fragment close to the 5' external homology arm, the 5' terminal region is preferably 20 nucleotides, more preferably 15 nucleotides, and further preferably 10 nucleotides at the 5' end of the 3' intron fragment, or inserted upstream of the 5' end of the 5' external homology arm, The 3' end arm comprises a 5' intron fragment and a 3' external homology arm in the 5' to 3' direction, the second tag is inserted in the 3' external homology arm, or inserted in the 3' terminal region of the 5' intron fragment close to the 3' external homology arm, the 3' terminal region is preferably 20 nucleotides at the 3' end of the 5' intron fragment, more preferably 15 nucleotides, further preferably 10 nucleotides, or inserted downstream of the 3' end of the 3' external homology arm, Preferably, The first tag is inserted into the 5' outer homology arm, and the second tag is inserted downstream of the 3' end of the 3' outer homology arm; The first tag is inserted in the 5' outer homology arm and the second tag is inserted in the 3' outer homology arm; or The first tag is inserted in the 5' external homology arm, and the second tag is inserted in the 3' terminal region of the 5' intron fragment; More preferably, the first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted downstream of the 3' end of the 3' external homology arm; The first tag is inserted in any position of the 5' outer homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted in the 3' outer homology arm; or The first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted at the 3' terminal region of the 5' intron fragment.
15. A method according to any preceding claim, wherein: The circularization region comprises a 3' coding region fragment, a translation initiation element, and a 5' coding region fragment in sequence in a manner operably linked to each other from 5' to 3' direction.
16. A method according to any preceding claim, wherein: The circularization region comprises a 3' exon fragment, a 5' internal homology arm, an insert fragment, a 3' internal homology arm and a 5' exon fragment in a manner operably connected to each other in sequence from 5' to 3' direction. Optionally, the circularization region comprises a first spacer between the insert fragment and the 5' internal homology arm, and a second spacer between the insert fragment and the 3' internal homology arm.
17. The method according to claim 16, wherein: The insert fragment comprises a translation initiation element, or comprises a translation initiation element and a coding region, wherein the translation initiation element is preferably an IRES sequence.
18. A method according to any preceding claim, wherein: The inserted fragment comprises a structural gene or a functional fragment thereof or a sequence of a non-coding RNA or its complementary sequence, wherein the structural gene encodes a polypeptide, a protein subunit, a protein active center, a protein or a protein hybrid of a non-natural catalytic group, a recombinant protein active subunit or active center, a recombinant artificial enzyme or other biological effect units mainly composed of amino acids, and the non-coding RNA is selected from microRNA (miRNA), small interfering RNA (siRNA), PIWI protein-interacting RNA (piRNA), transfer RNA-derived small RNA (tsRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), long non-coding RNA (lncRNA), pseudogene, ceRNA (competing endogenous RNAs), microRNA sponge or other types of non-mRNA RNA.
19. A method according to any preceding claim, wherein: The length of the 5' and 3' outer homology arms is each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the lengths of the 5' and 3' intron fragments are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, less than 70 nt, less than 80 nt. , less than 90nt, less than 100nt, less than 150nt, less than 200nt, 5-200nt, 10-150nt, 50-200nt, 50-150nt, 5-60nt, 10-55nt, 15-50nt, 20-45nt, 25-40nt, 30-35nt, 10nt, 15nt, 20nt, 25nt, 30nt, 35nt, 40nt, 45nt, or 50nt, Optionally, the length of the 5' and 3' exon fragments are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the lengths of the 5' and 3' internal homology arms are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the length of the insert fragment is greater than 50nt, greater than 100nt, greater than 150nt, greater than 200nt, greater than 250nt, greater than 300nt, greater than 400nt, greater than 500nt, greater than 600nt, greater than 1k nt, greater than 1.5k nt, greater than 2k nt, greater than 3k nt, less than 50nt, less than 100nt, less than 150nt, less than 200nt, less than 250nt, less than 300nt, less than 400nt, less than 500nt, less than 600nt, less than 600nt, less than 1k nt, less than 1.5k nt, less than 2k nt, less than 3k nt, 50-5knt, 50-5k nt, 50-4k nt, 50-3k nt, 50-2k nt, 50-1.5k nt, 50-1k nt, 50~600nt, 100~550nt, 150~500nt, 200~450nt, 250~400nt, 300~350nt.
20. A method according to any preceding claim, wherein: The 5' outer homology arm has a sequence as shown in SEQ ID NO: 1 or 56, and the 3' outer homology arm has a sequence as shown in SEQ ID NO: 2 or 57.
21. A method according to any preceding claim, wherein: The 3' intron fragment and the 5' intron fragment are derived from type I introns, preferably from the cyanobacteria Anabaena pre-tRNA gene or the T4 phage Td gene, more preferably the 3' intron fragment has a sequence as shown in SEQ ID NO: 3, and the 5' intron fragment has a sequence as shown in SEQ ID NO:
4.
22. A method according to any preceding claim, wherein: The 3' intron fragment and the 5' intron fragment are derived from type II introns, preferably from type II introns of Clostridium such as Clostridium tetani, or type II introns of Bacillus such as Bacillus thuringiensis, more preferably the type II intron is a type II intron contained in the nucleotide sequence shown in SEQ ID NO: 5 or 6.
23. A method according to any preceding claim, wherein: The 3' exon fragment and the 5' exon fragment are respectively derived from the 3' terminal region and the 5' terminal region of a natural exon, preferably from the cyanobacteria Anabaena pre-tRNA gene or the T4 phage Td gene, more preferably the 3' exon fragment has a sequence as shown in SEQ ID NO: 7, and the 5' exon fragment has a sequence as shown in SEQ ID NO:
8.
24. A method according to any preceding claim, wherein: Preferably, the 5' internal homology arm has the sequence shown in SEQ ID NO:9, and the 3' internal homology arm has the sequence shown in SEQ ID NO:
10.
25. A method according to any preceding claim, wherein: The first spacer region is the same as or different from the second spacer region.
26. A method according to any preceding claim, wherein: The affinity chromatography medium is operably connected to an affinity ligand capable of specifically binding to the first tag and / or the second tag, preferably the affinity ligand is selected from the group consisting of polymer X1, polymer X1-X2, polymer X1-X2-X3, polymer X1-X2-X3-X4, wherein X1, X2, X3, X4 are independently any one of A, G, C, T, U, more preferably any one of the group consisting of Oligo dT, Oligo dC, Oligo dG, Oligo dU.
27. A method according to any preceding claim, wherein: The affinity chromatography medium is selected from any one of the group consisting of magnetic beads, dextran molecules, polyacrylamide macromolecules, macromolecular cellulose molecules, chitosan materials, modified polylactic acid materials, PET materials, inorganic silicate materials or other high molecular polymers.
28. A method according to any preceding claim, which does not include the step of adding RNase R to the mixture to remove linear RNA.
29. The method according to any of the preceding claims, when it includes the step of adding RNase R to the mixture to remove linear RNA, the amount of RNase R added is reduced to 30-50%, preferably 40% or 50% of the amount of enzyme used for purification when the poly(A) tag or its functional variant is not inserted.
30. The circular RNA prepared according to any one of claims 1 to 29, which substantially does not contain non-circular RNA molecules.
31. A linear RNA precursor for use in the method of any one of claims 1 to 29.
32. A linear RNA precursor, characterized in that The linear RNA precursor comprises a 5' end arm, a 3' self-splicing site, a circularization region, a 5' self-splicing site and a 3' end arm in a manner operably connected to each other in sequence from 5' to 3' direction, wherein a first tag is inserted into the 5' end arm and a second tag is inserted into the 3' end arm; Optionally, the first tag and / or the second tag are independently selected from a poly(A) tag consisting of 20 to 100, preferably 30 to 90, more preferably 40 to 80, further preferably 45 to 70, and most preferably 50 to 65 consecutive adenine nucleotides or a functional variant thereof, and the first tag and the second tag are not present in the cyclization region and the circular RNA.
33. The linear RNA precursor according to claim 32, wherein The functional variant of the poly(A) tag is to insert one or more non-A bases into the poly(A) tag, preferably insert 1 to 20 non-A bases, more preferably insert 1 to 10 non-A bases.
34. The linear RNA precursor according to claim 32 or 33, wherein The functional variant of the poly(A) tag comprises (1) a single element a, at least one element b, and at least one element c, (2) only one element a, at least one element b, and at least one element d; or (3) a single element a, at least one element b, at least one element c and at least one element d, wherein the element a is composed of more than 20 consecutive adenine nucleotides, the element b is composed of more than 3 and less than 20 consecutive adenine nucleotides, the element c is composed of a nucleotide selected from uracil nucleotides, cytosine nucleotides, and guanine nucleotides, the element d is composed of more than 2 and less than 20 nucleotides, the nucleotides are arbitrarily selected from adenine nucleotides, uracil nucleotides, cytosine nucleotides, and guanine nucleotides, and the element d does not contain more than 3 consecutive adenine nucleotides, and the 5' and 3' terminal nucleotides are not adenine nucleotides, Wherein, when the Poly(A) tag contains two or more of the element b, the element c or the element d at the same time, the sequences of every two elements b may be the same or different, the sequences of every two elements c may be the same or different, and the sequences of every two elements d may be the same or different. Furthermore, the element a and the element b, the element c and the element d, the elements b, the elements c, and the elements d are not adjacent to each other.
35. The linear RNA precursor according to any one of claims 32 to 34, wherein the element a consists of more than 20 and less than 80 consecutive adenine nucleotides, preferably consists of 30 to 70, 35 to 65, 40 to 60, or 45 to 55 consecutive adenine nucleotides, more preferably consists of 60 consecutive adenine nucleotides.
36. The linear RNA precursor according to any one of claims 32 to 35, wherein the element b consists of 3 to 10, 10 to 19, 12 to 15, 14 to 17, or 16 to 19, preferably 19 consecutive adenine nucleotides. 37 . The linear RNA precursor according to claim 32 , wherein the number of the element b is 2 to 10, preferably 2 to 5, and more preferably 3.
38. The linear RNA precursor according to any one of claims 32 to 37, wherein the element c is a guanine nucleotide.
39. The linear RNA precursor according to any one of claims 32 to 38, wherein the number of the elements c is 2 to 10, 3 to 8, 4 to 6, or 2 to 5, preferably 2.
40. The linear RNA precursor according to any one of claims 32 to 39, wherein the element d consists of 3 to 18, 5 to 16, 4 to 10, or 6 to 12 nucleotides, preferably consists of 6 nucleotides.
41. The linear RNA precursor according to any one of claims 32 to 40, wherein the element d is selected from any one of GAUAUC, GUAUAC, GAAUCU, GCAUAUGACU or GAUAUCGUAUAC.
42. The linear RNA precursor according to any one of claims 32 to 41, wherein the number of the element d is 0 to 5, preferably 1 to 3, and more preferably 1.
43. The linear RNA precursor according to any one of claims 32 to 42, wherein when element c and element d exist simultaneously, the total number of element c and element d is 2 to 15, preferably 3 to 5, and more preferably 3.
44. The linear RNA precursor according to any one of claims 32 to 43, wherein The functional variant of the poly(A) tag has any one structure selected from the following structures: component a-component c-component b-component c-component b-component c-component b-component c-component b, Element b-element c-element b-element c-element a-element d-element b-element c-element b-element c-element b, Element b-element c-element b-element c-element b-element d-element a-element c, element a-element d-element b-element c-element b-element c-element b, or Element b-element c-element b-element c-element b-element d-element a; Preferably, the functional variant of the poly(A) tag has a nucleotide sequence selected from any one of SEQ ID NOs: 12 to 25.
45. The linear RNA precursor according to any one of claims 32 to 44, wherein The 5' end arm comprises a 5' external homology arm and a 3' intron fragment in the 5' to 3' direction, the first tag is inserted into the 5' external homology arm, or inserted into the 5' terminal region of the 3' intron fragment close to the 5' external homology arm, the 5' terminal region is preferably 20 nucleotides, more preferably 15 nucleotides, and further preferably 10 nucleotides at the 5' end of the 3' intron fragment, or inserted upstream of the 5' end of the 5' external homology arm, The 3' end arm comprises a 5' intron fragment and a 3' external homology arm in the 5' to 3' direction, the second tag is inserted in the 3' external homology arm, or inserted in the 3' terminal region of the 5' intron fragment close to the 3' external homology arm, the 3' terminal region is preferably 20 nucleotides at the 3' end of the 5' intron fragment, more preferably 15 nucleotides, further preferably 10 nucleotides, or inserted downstream of the 3' end of the 3' external homology arm, Preferably, The first tag is inserted into the 5' outer homology arm, and the second tag is inserted downstream of the 3' end of the 3' outer homology arm; The first tag is inserted in the 5' outer homology arm and the second tag is inserted in the 3' outer homology arm; or The first tag is inserted in the 5' external homology arm, and the second tag is inserted in the 3' terminal region of the 5' intron fragment; More preferably, the first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted downstream of the 3' end of the 3' external homology arm; The first tag is inserted in any position of the 5' outer homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted in the 3' outer homology arm; or The first tag is inserted at any position in the 5' external homology arm except the first nucleotide residue at the 5' end, and the second tag is inserted at the 3' terminal region of the 5' intron fragment.
46. The linear RNA precursor according to any one of claims 32 to 45, wherein The circularization region comprises a 3' coding region fragment, a translation initiation element, and a 5' coding region fragment in sequence in a manner operably linked to each other from 5' to 3' direction.
47. The linear RNA precursor according to any one of claims 32 to 46, wherein The circularization region comprises a 3' exon fragment, a 5' internal homology arm, an insert fragment, a 3' internal homology arm and a 5' exon fragment in a manner operably connected to each other in sequence from 5' to 3' direction. Optionally, the circularization region comprises a first spacer between the insert fragment and the 5' internal homology arm, and a second spacer between the insert fragment and the 3' internal homology arm.
48. The linear RNA precursor according to claim 47, wherein The insert fragment comprises a translation initiation element, or comprises a translation initiation element and a coding region, wherein the translation initiation element is preferably an IRES sequence.
49. The linear RNA precursor according to any one of claims 32 to 48, wherein The inserted fragment comprises a coding sequence of a structural gene or a functional fragment thereof, or a sequence of a non-coding RNA or its complementary sequence, wherein the structural gene is selected from a polypeptide, a protein subunit, a protein active center, a protein or a protein hybrid of a non-natural catalytic group, a recombinant protein active subunit or active center / , a recombinant artificial enzyme or other biological effect units mainly composed of amino acids, and the non-coding RNA is selected from microRNA (miRNA), small interfering RNA (siRNA), PIWI protein-interacting RNA (piRNA), transfer RNA-derived small RNA (tsRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), long non-coding RNA (lncRNA), pseudogene, ceRNA (competing endogenous RNAs), microRNA sponge or other types of non-mRNA RNA.
50. The linear RNA precursor according to any one of claims 32 to 49, wherein The length of the 5' and 3' outer homology arms is each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the lengths of the 5' and 3' intron fragments are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 70 nt, less than 80 nt, less than 90 nt, less than 100 nt, less than 150 nt, less than 200 nt, 5-200 nt, 10-150 nt, 50-200 nt, 50-1 50nt, less than 5nt, less than 10nt, less than 15nt, less than 20nt, less than 25nt, less than 30nt, less than 40nt, less than 50nt, less than 60nt, 5-60nt, 10-55nt, 15-50nt, 20-45nt, 25-40nt, 30-35nt, 10nt, 15nt, 20nt, 25nt, 30nt, 35nt, 40nt, 45nt, or 50nt, Optionally, the length of the 5' and 3' exon fragments are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5-60 nt, 10-55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the lengths of the 5' and 3' internal homology arms are each independently greater than 5 nt, greater than 10 nt, greater than 15 nt, greater than 20 nt, greater than 25 nt, greater than 30 nt, greater than 40 nt, greater than 50 nt, greater than 60 nt, less than 5 nt, less than 10 nt, less than 15 nt, less than 20 nt, less than 25 nt, less than 30 nt, less than 40 nt, less than 50 nt, less than 60 nt, 5 to 60 nt, 10 to 55 nt, 15-50 nt, 20-45 nt, 25-40 nt, 30-35 nt, 10 nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, 40 nt, 45 nt, or 50 nt, Optionally, the length of the insert fragment is greater than 50nt, greater than 100nt, greater than 150nt, greater than 200nt, greater than 250nt, greater than 300nt, greater than 400nt, greater than 500nt, greater than 600nt, greater than 1k nt, greater than 1.5k nt, greater than 2k nt, greater than 3k nt, less than 50nt, less than 100nt, less than 150nt, less than 200nt, less than 250nt, less than 300nt, less than 400nt, less than 500nt, less than 600nt, less than 600nt, less than 1k nt, less than 1.5k nt, less than 2k nt, less than 3k nt, 50-5knt, 50-5k nt, 50-4k nt, 50-3k nt, 50-2k nt, 50-1.5k nt, 50-1k nt, 50~600nt, 100~550nt, 150~500nt, 200~450nt, 250~400nt, 300~350nt.
51. The linear RNA precursor according to any one of claims 32 to 50, wherein The 5' outer homology arm has a sequence as shown in SEQ ID NO: 1 or 56, and the 3' outer homology arm has a sequence as shown in SEQ ID NO: 2 or 57.
52. The linear RNA precursor according to any one of claims 32 to 51, wherein The 3' intron fragment and the 5' intron fragment are derived from type I introns, preferably from the cyanobacteria Anabaena pre-tRNA gene or the T4 phage Td gene, more preferably the 3' intron fragment has a sequence as shown in SEQ ID NO: 3, and the 5' intron fragment has a sequence as shown in SEQ ID NO:
4.
53. The linear RNA precursor according to any one of claims 32 to 52, wherein The 3' intron fragment and the 5' intron fragment are derived from type II introns, preferably from type II introns of Clostridium such as Clostridium tetani, or type II introns of Bacillus such as Bacillus thuringiensis, more preferably the type II intron is a type II intron contained in the nucleotide sequence shown in SEQ ID NO: 5 or 6.
54. The linear RNA precursor according to any one of claims 32 to 53, wherein The 3' exon fragment and the 5' exon fragment are respectively derived from the 3' terminal region and the 5' terminal region of a natural exon, preferably from the cyanobacteria Anabaena pre-tRNA gene or the T4 phage Td gene, more preferably the 3' exon fragment has a sequence as shown in SEQ ID NO: 7, and the 5' exon fragment has a sequence as shown in SEQ ID NO:
8.
55. The linear RNA precursor according to any one of claims 32 to 54, wherein Preferably, the 5' internal homology arm has the sequence shown in SEQ ID NO:9, and the 3' internal homology arm has the sequence shown in SEQ ID NO:
10.
56. The linear RNA precursor according to any one of claims 32 to 55, wherein The first spacer region is the same as or different from the second spacer region.
57. The linear RNA precursor according to any one of claims 32 to 56, which has a nucleotide sequence selected from any one of SEQ ID NOs: 26 to 49 and 51 to 55.
58. A nucleic acid sequence capable of being transcribed to produce a linear RNA precursor for use in the method of any one of claims 1 to 29 or a linear RNA precursor of any one of claims 31 to 57, which preferably also comprises, in an operably linked manner, a regulatory sequence necessary for transcribing and producing the linear RNA precursor, such as a promoter, a terminator, a transcription factor binding site, an untranslated region (UTR), an enhancer, a palindromic sequence, a cis-acting element, a trans-acting element, a TATA box, a CAAT box, an operator, or a transposon.
59. A vector comprising the nucleic acid sequence of claim 58.
60. The vector according to claim 59, wherein The vector is a linear DNA, a plasmid, a viral nucleic acid fragment or a cell genome DNA fragment.
61. An engineered cell comprising a linear RNA precursor used in the method of any one of claims 1 to 29, a circular RNA according to claim 30, a linear RNA precursor according to any one of claims 31 to 57, a nucleic acid sequence according to claim 58, or a vector according to claim 59 or 60.
62. A composition comprising a linear RNA precursor used in the method of any one of claims 1 to 29, a circular RNA according to claim 30, a linear RNA precursor according to any one of claims 31 to 57, a nucleic acid sequence according to claim 58, a vector according to claim 59 or 60, or an engineered cell according to claim 61.
63. Use of a linear RNA precursor used in the method of any one of claims 1 to 29, a circular RNA according to claim 30, a linear RNA precursor according to any one of claims 31 to 57, a nucleic acid sequence according to claim 58, a vector according to claim 59 or 60, or an engineered cell according to claim 61 for preparing circular RNA.
64. Use of a linear RNA precursor used in the method of any one of claims 1 to 29, a circular RNA according to claim 30, a linear RNA precursor according to any one of claims 31 to 57, a nucleic acid sequence according to claim 58, a vector according to claim 59 or 60, or an engineered cell according to claim 61 for preparing a drug, a cytotoxic agent or an immunomodulatory preparation.
65. According to the use described in claim 64, the therapeutic drug, cytotoxic agent or immunomodulatory preparation is selected from viruses, pluripotent or multipotent stem cells, iPS, engineered immune cells, antibodies or antibody fragments, antibodies or antibody fragments conjugated to drugs, chemotherapeutic agents, immunosuppressants or modulators, anti-infective drugs, anticancer agents, hypoglycemic drugs, cardiovascular and cerebrovascular disease therapeutic drugs, degenerative neurological disease drugs, obesity therapeutic drugs, hematological disease therapeutic drugs, respiratory disease therapeutic drugs, or retroviral disease therapeutic drugs.
66. A method for administering circular RNA, comprising administering to an organism in need thereof an effective amount of a linear RNA precursor used in the method of any one of claims 1 to 29, a circular RNA according to claim 30, a linear RNA precursor according to any one of claims 31 to 57, a nucleic acid sequence according to claim 58, a vector according to claim 59 or 60, an engineered cell according to claim 61, or a circular RNA prepared using the composition according to claim 62.